I wanted to answer one question.
SPOT OM is an AI operating system that researches, verifies, designs, schedules, and evaluates Instagram content while requiring only two human approvals. Six weeks live, 51 posts, two fabrications caught.
This isn't really about Instagram. Supervising an AI that works alone looks the same in support, finance, and internal copilots. Instagram was just the cheapest proving ground I owned.
The math: Baseline: 14 posts × 40 minutes = 560 minutes a week. Current: 14 approvals × about 30 seconds = about 7 minutes. Timed on my own account.
Press Live at the bottom left. Open Approvals, hit Request changes, and watch the redo come back with the change marked.
The job was nine steps of hand work, and this user has minutes, not hours.
43% of small business owners spend six or more hours a week on social media (mine took nine), and handing it off costs 60 to 72 thousand a year for a manager, or 1 to 5 thousand a month for an agency.
Runs a K-beauty page at night beside a full-time job. Forty minutes a post was eating my life, and that is where this system started.
The owner of INMI, a Seoul skincare brand with no marketing team. Field-tested the research module, the first proof a piece of this works for someone who is not me.
40 min a post, 9 hours a week. The account ran. My life did not.
The other extreme: the AI took every step, rules lived in written instructions.
In week 3, a product name I did not recognize appeared in the schedule. It was not in the watchlist. The model had fabricated the product and marked it as verified, on an account published under my name.
At first I thought I just needed better prompts. I rewrote them stricter, and watched the rules fade all over again. Prompts were not the problem. I was asking memory to do a tool's job. The thing to design was trust.
Every claim now had to pass a verification tool that reads the live sales page, and two human gates went up around the work.
A second invented product appeared. The batch audit caught it, not me. Something slips, it becomes a rule, and the rule gets locked into a tool. I did not plan that pipeline. I noticed it after the second catch, and then I kept it on purpose.
Two gates, 26 rules (9 to start, added one accident at a time), 7 tools. 51 posts later, no accident has repeated.
Honestly: two people have used this system so far, the INMI owner and me. There are no invented interviews on this page.
I tried full autonomy first, and it failed exactly the way I should have expected. These are the things I believe now. I wrote them after the accidents, not before.
The human only makes the expensive decisions.
Trust has to exist before automation gets any.
Anything the system does alone, the user can undo.
When safety and intelligence conflict, safety wins.
Rules go in tools, because prompts forget.
Chat forgets. Rules kept here faded within weeks, and that is how v2 died.
The 26 rules live in skill files. A skill loads fresh on every run, so it cannot fade.
The 3 most dangerous rules are locked inside tools the AI must use, so following them is never the AI's choice.
Claude workflows and skills, Canva (its connector blocks cutouts and new text, which forced the master-template method), Metricool free plan (a 30-day data window), Apify (batches cut after a proxy block), Meta API (comment-keyword DMs). The constraints became the design.
One gate (review only the release): a bad idea burns a whole production cycle before it dies. Three or more gates: review becomes a job, and collapses.
I gave up control of everything in the middle.
A wrong call is expensive at exactly two points: choosing what to make, and putting your name on it. Everything between those two is reversible.
The review held across all 51 posts. Two is a number a person can keep.
Better instructions: chat forgets, and I watched the rules leak within three weeks. Fine-tuning: no data and no budget for one person. Locking rules into tools: chosen.
Flexibility. Tools are slower to change and forgive no exceptions.
Seven was never the goal. There were seven jobs where improvising caused or could cause an accident, and each one got a tool: product verification, word rules, image rules, scoring, the research scan, scheduling, the audit.
No accident happened twice. The safety survives model swaps.
Auto-posting from day one: a great demo, and users fear it too much to use it. Supervision forever: the minutes-a-week promise breaks.
The demo impact of "fully automatic." At the start, every post waits for a person.
Five clean approvals: small enough to reach in two weeks, big enough to call a streak. It is a starting integer, tuned as data grows. It always asks before climbing, and one tap pulls it back.
Nobody uses automation they are afraid of. The version you can take back in one tap is the version that actually gets used.
Walk the prototype end to end as a first-time user, grade every finding blocker, gap, or polish, fix by grade. One thing that audit changed:
The two gates kept the review honest, the rule tools kept false claims at zero, the ladder kept the user in control, and the score kept the system learning.
Every row is labeled: proven, hypothesis, or future. They never mix.
Six weeks in, I keep coming back to one small thing.
Turns out the smallest job I kept, two yeses, was the thing protecting everything else.
v1, running by hand, ate my life. v2, full automation, invented a product. And even after output was solved, reach still fell.
I thought consistency would solve growth. Fifty-one posts proved me wrong. Making and spreading turned out to be different problems.
A confidence score graded by the same model that invents facts is not a safety device. A gate stands outside the model, so it does not fall with the model.
On purpose. Ingredient claims, ad rules, hype language: lies cost the most here. A safety structure that holds in the hardest domain is the one worth carrying anywhere else.
The parts I built to look complete. The hardcoded month calendar, the comfort charts. Honestly those screens were for me, not for a user, and the score and the work log were already doing all the deciding.
The structure is mine, and the model never was. Two gates, rules in tools, earned autonomy, and the work log rebuild on top of any model. That is the skill I am selling.
Attack the reach problem with a short-video lane.
A second account, a second user. Measure the 70% for real.
The product, for solo founders, creators, and small brands.
If I had engineers tomorrow, the first build is the audit replay harness: automatically re-run every past accident against a new model or a new account, turning safety into a regression test.
If you are stretched across things you care about, I built this for you. And if you are building AI products people have to trust, that is the work I want to do.
Ask me about any decision in this system. I can walk you through the why in twenty minutes.
Ashley Kim · ashleykim1229@gmail.com