Free until October 1. Lock your foundation and run your first client diagnostic before your Q1 pipeline conversations start.

Join the beta

Your AI Makes 159 Judgment Calls. Ours Has to Declare Every One in Writing Before It Ships.

Governance of AI agency as a code artifact: the decision registry.

Stijn Hendrikse · Sep 12, 2026

ShareLinkedInXEmail

How many decisions does your AI make on your behalf, and can anyone show you where they are? In T2D3 OS the answer is 159 — every judgment surface declared in writing in a decision registry that lives in the codebase, where an undeclared one fails continuous integration and blocks the build. Most products cannot produce the list at all.

Not "how good is the output." Where are the decisions. The moments the software picked a direction — chose which persona to lead with, decided a finding was strong enough to treat as true, ranked one channel above another, suggested you cut a segment. Every one of those is a judgment call. Every one of them is either yours or the vendor's. Most products cannot produce a list.

We can. There are 159 of them in T2D3 OS, and each one has a written entry in a decision registry that lives in the codebase. If an engineer adds a new judgment surface without declaring it, continuous integration fails. The build does not ship. That is the whole bold claim of this post: AI agency should be a governed code artifact, not a vibe.

The real risk isn't that the AI is wrong — it's that you can't find where it decided

The fractional CMO fear that keeps showing up in our own customer conversations is not "AI will produce bad work." It's subtler and worse: AI will produce plausible work, my client will act on it, and I will not be able to reconstruct why.

That's a professional liability. Your retainer is priced on judgment. The moment your judgment becomes indistinguishable from a model's default, your rate has no floor — and everyone now believes they, or an AI, could have done the marketing anyway.

Undeclared decision surfaces are how that happens quietly. Not through one dramatic hallucination. Through forty small unlabeled choices that were never routed to you.

What each of the 159 entries has to answer

A surface does not get to exist in our product until it declares five things:

Which grounded recommender does the homework. Every recommendation names the prompt that produced it and the evidence it read. No anonymous suggestions.

Whether a devil's advocate scrutinizes it. High-stakes surfaces get an adversarial critic pass before anything reaches you — the system arguing against its own recommendation before you see it.

What triggers the work proactively. A surface either waits to be asked or fires on a condition. It must say which. Silent background agency is the thing we refuse to allow undeclared.

Where feedback lands. When you confirm, edit, or reject, that verdict has to have a destination that changes future behavior. Feedback with nowhere to go is theatre.

How much autonomy is allowed. A tier, explicitly set. Some surfaces draft and wait. None of them act outbound on their own.

Five declarations, 159 surfaces — that is the whole inventory, and it is countable precisely because it is written down rather than described. The point of writing it down is that a practitioner can audit it. Andrea Nicholas, a management consultant who uses the T2D3 framework with her own clients, put the appeal of the underlying method this way: "it's soup to nuts. It's like plug and play, you know, and it's the full breadth of it. It's not just a piece or a superficial layer." A registry that covers every surface, rather than the few a vendor is proud of, is the same instinct applied to AI governance.

And the invariant that makes it real: the registry has a shrink-only baseline. Existing undeclared debt is frozen at its current count and can only go down. New surfaces cannot join it. You cannot ship agency in the dark by accident, because the test suite treats that as a broken build.

Why we made this a test and not a doc

Documentation rots. Governance that lives in a PDF is a claim. Governance that fails CI is a constraint.

This is the same reason our platform makes you confirm findings before the rest of the workspace treats them as true. As one of our own product notes puts it: remove that human confirmation moment and "you have not removed a bottleneck, you have removed the only step that generates the signal everything else runs on." The diff has nothing to diff against without a human conviction on one side of it.

The registry is the structural version of that belief. Your judgment is the product. So the software has to be able to show you, surface by surface, exactly where it is asking for it.

Audit your own product's judgment surfaces: a template

Whether or not you ever buy anything from us, run this on the AI tooling you already use. For each place the product produces a recommendation, ask:

  1. Is this a decision surface? If a human would have had to choose, it is one. Ranking, filtering, summarizing, and "smart defaults" all count.
  2. Can I name the evidence? If I cannot trace the recommendation to specific inputs, it is a guess wearing a confident tone.
  3. Did anything argue the other side? Or did the first plausible answer win by default?
  4. What fired this? User request, or background trigger nobody declared?
  5. Where did my correction go? If editing it changes nothing next time, my judgment is being discarded, not captured.
  6. What could it do without me? Draft, publish, send, spend. Know the tier.

Then count how many surfaces you could actually answer all six for. That number, divided by the total, is your real governance coverage. Most teams discover it is close to zero — and that the vendor cannot tell them either.

Any AI product that can't enumerate its judgment surfaces is asking for your trust on faith

Accepting a vendor's agency on faith is a fine trade for autocomplete. It is an unacceptable trade for the strategy your client's next two quarters run on.

We think the industry standard should be: publish the count, publish the declarations, and fail the build when a surface goes undeclared. We started at 159 because that is what an honest inventory returned.

Ask your other vendors for their number.

Join the conversation

Put this playbook to work — with the OS built for it.

T2D3 OS turns the method behind this guide into working modules: ICP, personas, positioning, content, and a full GTM plan. Start free.