S SPOT OM
Ashley Kim · Senior Product Designer

I wanted to answer one question.

When an AI runs an entire Instagram account, where should a human still decide?

SPOT OM is an AI operating system that researches, verifies, designs, schedules, and evaluates Instagram content while requiring only two human approvals. Six weeks live, 51 posts, two fabrications caught.

SPOT OM live: the two-step approval screen
99%
less human time
51
posts in six weeks
0
false claims out
2
human decisions per post
What I designedThe whole system: operating model, safety architecture, workflow orchestration, information architecture, approval system, evaluation framework. Alone. No PM, no engineers.
WhenJun 7 to Jul 13, 2026 · 30+ hrs/wk beside a full-time job
StatusLive in production · @spot.kbeauty
Impact

Every metric below includes how it was measured.

This isn't really about Instagram. Supervising an AI that works alone looks the same in support, finance, and internal copilots. Instagram was just the cheapest proving ground I owned.

99%less human time
Before · by hand
9 hrs/wk
Now · SPOT OM
7 min/wk

The math: Baseline: 14 posts × 40 minutes = 560 minutes a week. Current: 14 approvals × about 30 seconds = about 7 minutes. Timed on my own account.

51
posts in six weeks
Measured: published count, Jun 7 to Jul 13, from the posting calendar. Every one scored.
0
false claims published
Measured: every claim is checked against a live sales page before the post is built, and re-checked before every batch goes out. Both catches were replaced before publishing.
2
human decisions per post
Measured: the two gates in the workflow: one yes on the idea, one yes on the finished post.
70%
reusable for a second account
Measured: counted by module, as an estimate. Pipeline, 7 tools, scoring, and safety rules carry. Topic, voice, and watchlist swap. A second account will measure it for real.
01 · The product

Try it yourself, 20 seconds.

Press Live at the bottom left. Open Approvals, hit Request changes, and watch the redo come back with the change marked.

Clickable prototype
The approval card. You sign what you see.
Approval card with cover, caption, DM reply, and what changed in v2
The decision

The whole post on one card: cover, caption, DM reply, time, and what changed since your last look.

Why

Honest review survives when no is as easy as yes. A change request is one click.

Alternatives considered

A modal hides context and blocks comparison. A feed turns review into scrolling, and scrolling into skimming. A split page scatters the signing moment. One card is the only shape that keeps "you sign what you see" true.

Trust ladder screen
The trust ladder. One step per 5 clean approvals, it always asks first, and one tap pulls it back.
Calendar screen with reasons per slot
The calendar. Every card carries why this time, and a click moves, holds, or opens it.
Three things using it taught me
1The trust ladder came from a usability finding. In my walkthrough test the autonomy setting was display-only, and an autonomy label you cannot change makes anxiety, not trust. So the ladder became a consent control instead of a status badge.
2The navigation redesign came from getting lost. The first version had two items both named Notification and one flat setup list. It was my own prototype and I still lost my way in it. So the nav split into RUN and SET UP.
3The approval flow hid a misunderstanding. The idea gate was buttons with no evidence, so there was no way to know why an idea deserved a yes. Now every idea card carries its source account and its view-count proof.
02 · The problem, and the gap

The original pain: one post ate one evening.

The job was nine steps of hand work, and this user has minutes, not hours.

Before · all human

Research
Write
Fact-check
Design the slides
Build the carousel
Schedule
Reply to comments
Analyze
Repeat
40 min per post · 9 hrs a week

Now · SPOT OM

The AI runs all nine steps
Human: yes to the idea
Human: yes to the finished post
Publish, auto-DM, score
7 min a week

43% of small business owners spend six or more hours a week on social media (mine took nine), and handing it off costs 60 to 72 thousand a year for a manager, or 1 to 5 thousand a month for an agency.

Real user 1 · me

Runs a K-beauty page at night beside a full-time job. Forty minutes a post was eating my life, and that is where this system started.

Real user 2 · INMI

The owner of INMI, a Seoul skincare brand with no marketing team. Field-tested the research module, the first proof a piece of this works for someone who is not me.

The landscape: why existing tools are not enough
SchedulersLater, Buffer, Metricool. They hold the publishing time. What to say is still all on you.
Creation toolsCanva and friends. Making gets faster. Research, verification, and judgment stay untouched.
Chat AIChatGPT or Claude alone. It writes, nobody supervises, and it invents with full confidence.
SPOT OMTies all three layers into one operating system and leaves the human only the control. Nothing on this map sells supervision.
ACT 1 · The AI failed

Trust in the automated system broke in week 3. The design started there.

v1 · all by hand

40 min a post, 9 hours a week. The account ran. My life did not.

v2 · full automation

The other extreme: the AI took every step, rules lived in written instructions.

Week 3 · the moment

In week 3, a product name I did not recognize appeared in the schedule. It was not in the watchlist. The model had fabricated the product and marked it as verified, on an account published under my name.

The wrong assumption

At first I thought I just needed better prompts. I rewrote them stricter, and watched the rules fade all over again. Prompts were not the problem. I was asking memory to do a tool's job. The thing to design was trust.

The experiment

Every claim now had to pass a verification tool that reads the live sales page, and two human gates went up around the work.

The structure tested itself · Jun 29

A second invented product appeared. The batch audit caught it, not me. Something slips, it becomes a rule, and the rule gets locked into a tool. I did not plan that pipeline. I noticed it after the second catch, and then I kept it on purpose.

v3 · current

Two gates, 26 rules (9 to start, added one accident at a time), 7 tools. 51 posts later, no accident has repeated.

ACT 2 · Why it failed

Question, method, insight, and what changed in the design.

What actually works in this niche?
Method: a 50-post study of a reference account
Reframe-education posts pull 2x product posts. Personal vlogs die.
→ No-face, education-first identity
Which signals predict growth?
Method: a signal study across engagement metrics
Likes lie. Saves and shares predict demand.
→ SCORE 4/3/2/1 drives next week's plan
Does it work for someone who is not me?
Method: one external user test (the INMI owner)
The research module passed real use.
→ Next validation: a second account, a second user

Honestly: two people have used this system so far, the INMI owner and me. There are no invented interviews on this page.

05 · Design principles

Five principles, written after the accidents.

I tried full autonomy first, and it failed exactly the way I should have expected. These are the things I believe now. I wrote them after the accidents, not before.

1

The human only makes the expensive decisions.

2

Trust has to exist before automation gets any.

3

Anything the system does alone, the user can undo.

4

When safety and intelligence conflict, safety wins.

5

Rules go in tools, because prompts forget.

ACT 3 · The redesign

The structure: AI owns 10, the human owns 2, and 3 rule blocks.

System diagram: how the work moves
HUMAN LANE · 2 DECISIONS Gate 1: the idea Gate 2: release Watchlist scan Apify Research + rank Claude · skills + rules Verify + build verification tools · Canva Publish Metricool · Meta API Score 4/3/2/1 scores feed next week's plan
Responsibilities: who decides what

Human (2)

  • Pick the idea
  • Release the finished post

AI (10)

  • What to research, which topics to shortlist
  • Which verified products to feature
  • Post type, hook angle, voice
  • Image concept, inside locked rules
  • DM keyword and reply, time slot
  • What to make more of next week, from scores

Rule blocks (3)

  • Products it cannot find for sale
  • Words the rules ban
  • Claims that get ad accounts restricted
Prompt architecture: why the rules cannot be forgotten
Chat instructions ✕

Chat forgets. Rules kept here faded within weeks, and that is how v2 died.

Skill files

The 26 rules live in skill files. A skill loads fresh on every run, so it cannot fade.

Hard blocks in tools

The 3 most dangerous rules are locked inside tools the AI must use, so following them is never the AI's choice.

STACKThe stack, and what it forced (open)

Claude workflows and skills, Canva (its connector blocks cutouts and new text, which forced the master-template method), Metricool free plan (a 30-day data window), Apify (batches cut after a proxy block), Meta API (comment-keyword DMs). The constraints became the design.

07 · Major decisions and iterations

Three decisions with a price, and one small one.

DECISION 1Why exactly two review gates.
Alternatives

One gate (review only the release): a bad idea burns a whole production cycle before it dies. Three or more gates: review becomes a job, and collapses.

The price

I gave up control of everything in the middle.

Why this one

A wrong call is expensive at exactly two points: choosing what to make, and putting your name on it. Everything between those two is reversible.

Outcome

The review held across all 51 posts. Two is a number a person can keep.

DECISION 2Rules live in tools, not in prompts.
Alternatives

Better instructions: chat forgets, and I watched the rules leak within three weeks. Fine-tuning: no data and no budget for one person. Locking rules into tools: chosen.

The price

Flexibility. Tools are slower to change and forgive no exceptions.

Why 7 tools

Seven was never the goal. There were seven jobs where improvising caused or could cause an accident, and each one got a tool: product verification, word rules, image rules, scoring, the research scan, scheduling, the audit.

Outcome

No accident happened twice. The safety survives model swaps.

DECISION 3The AI earns its freedom step by step.
Alternatives

Auto-posting from day one: a great demo, and users fear it too much to use it. Supervision forever: the minutes-a-week promise breaks.

The price

The demo impact of "fully automatic." At the start, every post waits for a person.

Why 5 approvals

Five clean approvals: small enough to reach in two weeks, big enough to call a streak. It is a starting integer, tuned as data grows. It always asks before climbing, and one tap pulls it back.

Outcome

Nobody uses automation they are afraid of. The version you can take back in one tap is the version that actually gets used.

4The small one: posts are scored shares 4, saves 3, comments 2, likes 1, the playbook priority order as integers, and that score plans next week by itself.
What I deliberately did not build
A chat interface. Conversation is where rules go to die: chat forgets, and cards and forms don't.
Auto-posting from day one. A great demo, and automation people fear goes unused.
An analytics-first home. Decision quality lives in one score. Comfort charts only cost upkeep.
The UI was tested the same way

Walk the prototype end to end as a first-time user, grade every finding blocker, gap, or polish, fix by grade. One thing that audit changed:

Prototype v2 approval page
Before · caught in the auditThe idea gate was one hardcoded row whose buttons only fired toasts. Two nav items were both named Notification.
Prototype v3 approval page
After · the rebuildReal idea cards with source and evidence, the nav split into RUN and SET UP, and one-click approval for the safest posts.
08 · Validation and results

Each result traces to one decision.

The two gates kept the review honest, the rule tools kept false claims at zero, the ladder kept the user in control, and the score kept the system learning.

Risk matrix: each failure mode and what catches it
Failure modeReal caseWhat catches itRepeats
Invented productsWeek 3: a nonexistent product marked verified. Jun 29: the audit caught a second one.The sales-page verification tool, plus the recurring audit0
Banned words and hypeCure-claim wording in early draftsThe word-rule tool blocks it before a human sees it0
Ad-restriction claimsBefore-after claim phrasing that ad policy bansThe ad-safe rewrite rules0
Scan breaksA proxy block stopped the watchlist scanSmaller batches and a retry routinefixed
Operating metrics, as a founder counts them
Cadence now14 posts a week, all published on schedule
Human timeabout 7 minutes a week, spent on approval cards
Failure rate0 incidents after publishing. Both catches were blocked before release
Tool cost0 dollars a month on free tiers. Own running cost not measured yet
Success metrics: current, target, next experiment
MetricCurrentTargetNext experiment
Weekly reach6back to two digitsthe short-video lane. The paid test set the direction: broad wins, and visits alone bring no follows.
Validated users23+a second account in a second user's hands, which also measures the 70% reuse estimate.
Post-publish incidents0hold at 0replay past audits as a regression test on every model swap.
09 · Scale

From one account to a thousand: what holds

Every row is labeled: proven, hypothesis, or future. They never mix.

Multi-account Hypothesis70% of the modules carry. Each account swaps only its watchlist, voice, and rules. A second account measures this next.
Roles FutureThe two yeses stay with the owner of the name. What gets delegated is drafts and research, never the signature.
Observability HypothesisThe work log is already a per-account audit trail. At 1,000 accounts, that log becomes the monitoring surface.
Incidents ProvenThe accident pipeline has already run twice: something slips, it becomes a rule, and the rule gets locked into a tool. At scale, one account's accident becomes every account's rule.
Model upgrades HypothesisRules live in tools, so safety survives a model swap. Validating a new model is a regression test: replay the past audits.
Escalation ProvenOne tap pulls autonomy back, and hold mode stops everything. At scale the stop button stays per-account.
10 · Business strategy

If this becomes a company, this is the shape.

Why nowModels became capable enough to carry the work, but not honest enough to run unsupervised. That gap is the product window. The social media management market: about 30 billion dollars in 2025, projected 172 billion by 2033 (Grand View Research).
Wedge and expansion HypothesisSolo creators and founders first (I am that segment, and every account the product runs is its own public demo), then small brands, then agencies running N accounts, where the 9-hours-times-N pain is largest.
Pricing evolution FutureFree single account, paid multi-account, agency seats. The anchors already exist: a manager at 60 to 72 thousand a year, agencies at 1 to 5 thousand a month.
The moatNever the model. The 26 rules born from real accidents, the verified product database, the per-niche playbooks, and the audit corpus that regression-tests any new model. Each account's learned voice and watchlist compound into switching cost.
11 · Reflection

Five hard questions.

Six weeks in, I keep coming back to one small thing.

Turns out the smallest job I kept, two yeses, was the thing protecting everything else.

What failed completely?

v1, running by hand, ate my life. v2, full automation, invented a product. And even after output was solved, reach still fell.

Which assumption was wrong?

I thought consistency would solve growth. Fifty-one posts proved me wrong. Making and spreading turned out to be different problems.

Why gates instead of a confidence score?

A confidence score graded by the same model that invents facts is not a safety device. A gate stands outside the model, so it does not fall with the model.

Why skincare, a hard domain?

On purpose. Ingredient claims, ad rules, hype language: lies cost the most here. A safety structure that holds in the hardest domain is the one worth carrying anywhere else.

If I started over, what would I not build?

The parts I built to look complete. The hardcoded month calendar, the comfort charts. Honestly those screens were for me, not for a user, and the score and the work log were already doing all the deciding.

If a big lab ships this tomorrow, what survives?

The structure is mine, and the model never was. Two gates, rules in tools, earned autonomy, and the work log rebuild on top of any model. That is the skill I am selling.

Learned the hard way
1Growth: followers sit at 38 and weekly reach fell from 23 to 6. The next experiment is the short-video lane.
2The operating UI has no user testing yet. n=2, INMI and me, is the honest present.
3The 70% reuse number is a module-count estimate. A second account will measure it.
12 · What is next

One account today. A second operator next. Then a product.

Now

Attack the reach problem with a short-video lane.

Success: weekly reach back to two digits.
Next

A second account, a second user. Measure the 70% for real.

Success: a stranger runs one week without me.
Later

The product, for solo founders, creators, and small brands.

Success: a user's first week costs minutes of setup and two yeses.

If I had engineers tomorrow, the first build is the audit replay harness: automatically re-run every past accident against a new model or a new account, turning safety into a regression test.

What I hope you take away
AI UXSystems thinkingProduct strategyInformation architectureWorkflow designSafety designHuman-AI interactionOperational designRapid iterationFounder mentality

If you are stretched across things you care about, I built this for you. And if you are building AI products people have to trust, that is the work I want to do.

Ask me about any decision in this system. I can walk you through the why in twenty minutes.

Ashley Kim · ashleykim1229@gmail.com