Free until October 1. Lock your foundation and run your first client diagnostic before your Q1 pipeline conversations start.

Join the beta

Autonomy Is Earned, Not Configured

An autonomy ladder for AI agents where 'act on its own' is earned by evidence, never toggled.

Stijn Hendrikse · Sep 10, 2026

ShareLinkedInXEmail

Every AI product ships an autonomy setting. A dropdown, a toggle, a slider from "suggest" to "auto." You pick your comfort level on day one, before the system has done a single piece of work for you, and that number governs what it is allowed to do to your client's positioning at 2am on a Thursday.

That is backwards. A configuration screen asks you to predict trust. Trust is not predictable — it is observable. It accrues from a record of the system being right about your business, in your words, on your accounts.

So we do not let you configure autonomy. Our agents earn it, one rung at a time, and lose it automatically when the evidence stops holding.

Every agent starts on the bottom rung, no matter how good it is

There are three rungs on the ladder:

proposal_only — the agent takes a position and waits. It cannot write to the workspace. It surfaces a finding, an ICP hypothesis, a draft pain statement, and it sits there until you confirm it or overwrite it with your own wording.

draft_for_review — the agent produces the artifact and stages it. The work is done, the diff is visible, and nothing is canonical until a human locks it.

auto — the agent acts inside a bounded surface without waiting for you.

New agents, new surfaces, and new client engagements all start at proposal_only. Not because the model is weak — the frontier models are extraordinary — but because the model has no evidence about this engagement yet. A system that has never been corrected on your client's segmentation has not earned the right to change it.

"Act on its own" requires a critic lane and stable evals — not a checkbox

An agent gets promoted to auto on a given surface only when two conditions hold at once, and both are measured, not asserted.

The surface has a critic lane. Something other than the producing agent reviews the output before it lands — an adversarial reviewer whose job is to find where the work is wrong, thin, or ungrounded. If the surface has no critic, it does not matter how good the producer is. There is no second opinion in the loop, so there is nothing to catch the failure mode nobody anticipated.

The org's evals are stable. Not "we ran evals once." Stable across recent runs on your actual foundations — your ICP, your personas, your locked value props. Evals that wobble mean the system does not yet understand your business well enough to be left alone in a room with it.

And the ladder runs downhill on its own. When evals destabilise — a new market, a repositioning, a client whose vocabulary the system keeps getting wrong — the surface downgrades automatically to draft_for_review. You do not get an alert asking whether you would like to reduce autonomy. The rope shortens, and you find out because work starts arriving for review again.

This is the part that costs us something. It means our own product gets slower for you exactly when your business is changing fastest, which is the moment you most want leverage. We think that trade is correct. Autonomy that survives a repositioning was never grounded in the first place.

Destructive and outbound actions never reach auto. Ever.

There is no rung above draft_for_review for two classes of action, and no eval score unlocks one.

Destructive — deleting, unlocking, or overwriting a locked foundation. Outbound — anything that leaves the building and reaches a client, a prospect, or a public channel under your name.

The reasoning is not squeamishness about AI. It is that these actions are unrecoverable in the dimension that matters. A bad draft costs you fifteen minutes. A confidently wrong email to your client's CEO costs you the relationship you spent a year building, and there is no rollback for the sentence they already read.

Anyone selling you autonomous outbound is selling you their confidence in their model. We would rather sell you a shorter, honest ceiling.

The autonomy ceiling rises with the richness of your foundations

Here is the mechanism underneath all of it: an agent's ceiling is proportional to how much confirmed signal it has to stand on.

An empty workspace earns proposal_only, because an ungrounded agent is just a competent guesser. A workspace with a locked ICP, confirmed personas, value props with the diffs and the "what the AI got right / what it missed" notes attached — that agent can be trusted with more, because more of its judgment is your judgment, played back.

This is why the human lock gate is not a bottleneck we are apologising for. The moment you confirm a finding, or edit it to keep your own wording, you generate the highest-value signal in the system: not just what changed but why. Remove that moment and you have not removed friction — you have removed the only step that produces the evidence autonomy is built from.

The policy, in five lines you can copy

You do not need our product to run this. Steal the policy:

  1. Default every new agent and surface to propose-only. Trust is per-surface and per-account, never global.
  2. Require a critic lane before anything acts. No independent reviewer, no autonomy.
  3. Gate promotion on eval stability over time, not a one-off pass.
  4. Downgrade automatically when evals wobble. Make the reduction the default, not a decision someone has to make.
  5. Hard-ceiling destructive and outbound actions at human review. No exceptions, no override, no enterprise tier that unlocks it.

A slider asks you how much you trust a system you have not watched work. A ladder shows you what it has earned. When someone asks how much rope to give an AI agent, the honest answer is: exactly as much as it has proven it can hold — and not one rung more.

Join the conversation

Put this playbook to work — with the OS built for it.

T2D3 OS turns the method behind this guide into working modules: ICP, personas, positioning, content, and a full GTM plan. Start free.