Free until October 1. Lock your foundation and run your first client diagnostic before your Q1 pipeline conversations start.

Join the beta

No More Silent Zeros: Every Judgment Call Gets a Recommendation — and a Devil's Advocate

Most software asks you questions it could help you answer: a target field that defaults to zero, a price that says "$X", three send buttons and no counsel. We just finished a two-day sweep across every module in T2D3 OS so that every judgment surface ships three things — a grounded recommendation with its arithmetic shown, a devil's advocate that argues against it, and a registry entry naming who decides. Here is the pattern, what the skeptics caught in live testing, and why "no recommender" is sometimes the right answer — as long as you write down why.

Stijn Hendrikse · Aug 22, 2026

ShareLinkedInXEmail

Open almost any B2B tool and you will find the same quiet failure: a required number field that defaults to zero. A quarterly target. A price per seat. A marketing budget. The software needs the number to do its job, the human has the judgment to set it — and the interface just sits there, empty, offering nothing. So the number gets typed from memory, or copied from last quarter, or left at zero until the dashboard quietly grades everything against a target nobody ever chose.

We call this the silent default problem, and over the last two days we finished eliminating it from every module in T2D3 OS. Not by making the AI decide — by making it do the homework before the human decides.

The pattern: recommend, scrutinize, confirm

Every judgment surface now ships the same three-part shape:

1. A grounded recommendation that shows its arithmetic. When the Marketing Scorecard needs a target for "Sales Qualified Leads," the recommendation is not "13, trust us." It is: "50 MQLs/mo × 3 months × 5% MQL→SQL ≈ 13 SQLs for the quarter, per your Growth Calculator, matching your pipeline OKR." Every number traces to a source the team already owns — the locked ICP, the active OKRs, the calculator, the uploaded research. And when nothing supports a number, the recommendation says so: target zero, opening with "Assumption: capture a baseline first." An honest zero beats a plausible invention, because a target nobody can trace destroys your credibility in the boardroom.

2. A devil's advocate that argues against it. Before the recommendation reaches you, a second AI pass — with fresh context and a mandate to object — attacks it from the perspectives of the people who will actually test the decision: the CFO who funds the budget, the buyer who reads the pricing page, the recipient who gets the cold email, the audience watching the deck. Crucially, the advocate argues over the same evidence the recommendation was built from. It cannot invent facts; it can only catch the recommendation contradicting them.

3. A human confirmation that teaches the system. Nothing is written until you click. Accept as-is, adjust the number, or dismiss the whole idea — each outcome is captured as a judgment signal that tunes future recommendations. Your edit is not friction; it is the most valuable data the system collects.

What the skeptics caught in live testing

We test every one of these surfaces against the running product with real model calls, and the devil's advocates earned their seats. A sample from this sweep's acceptance runs:

  • Pricing: the recommender derived a $129/user hero tier from the company's implied deal size. The advocate's objection: "That assumes ~11 seats per account — back-derived, never validated — while your locked value props anchor to consolidators running 500 clinics." It attacked the derivation, not the number.
  • Cold outreach: a draft scored 9/10 on tone rules while pivoting from the recipient's actual pain (a hiring spike) to an unsupported compliance pitch. The advocate held the send: "No evidence in the door research that compliance is their problem" — and then quoted the product's own spam rubric back at the draft: "the word 'guarantee' is flagged in your own scoring as a spam trigger."
  • Competitive monitoring: the recommended weekly-digest recipient list included an address whose email domain belonged to a company on the watched competitor list. No text-quality rule catches that; a skeptic cross-referencing two evidence blocks does.
  • Deck building: the recommender pitched a slide as "the whole point of Day 1." The advocate checked the coverage math — the org's data could fill 2 of that slide's 7 content slots — and warned it would render generic in front of the client.

None of these are template objections. Each one cites the specific evidence line it rests on, which is the entire point: when the recommendation shows its work, the skeptic has something real to bite.

The part most teams skip: writing down who decides

The sweep's real deliverable is not sixteen new prompts. It is a decision registry: a machine-checked table where every module declares, for every judgment surface, four things — which grounded recommendation does the homework, whether a devil's advocate scrutinizes it, what triggers the work proactively, and how far the AI may act on its own (always: propose only for money and outbound).

Two properties make this more than documentation:

It is enforced. Our CI fails any change that adds a module surface without a declaration. "We'll add the recommendation later" is no longer a merge-able state.

"No recommender" is a legal answer — with a reason. Some of the most instructive entries declare no AI on purpose. The engagement-health module's self-assessment gate exists precisely so the account manager's own honest judgment enters first; an AI draft there would write the honesty the gate demands. The outreach module's volume governor is deterministic rules with an audit trail, because an LLM recommending send volumes would launder vibes into deliverability. A ranking engine's critic is a coverage-and-spread check computed in code, because a critic with a spec beats a critic with a vibe. The discipline is not "AI everywhere" — it is naming, for every judgment call, who decides, on what evidence, and who argues back.

Why this matters beyond our product

Every AI vendor will tell you their product "recommends." The questions that separate a working system from a demo are the ones this pattern forces:

  1. Can the recommendation show its arithmetic? If not, it is a guess wearing a confident font.
  2. Does anything argue against it before you see it? A single model grading its own homework converges on flattery.
  3. Is "the AI never acts alone" a promise or a property? Ours is a table that CI reads. Yours should be checkable too.
  4. Does your edit teach the system? If accepting and overriding look identical to the vendor, the product is not learning from the only judgment that matters — yours.

The silent zero was never really an interface bug. It was a governance gap: a decision the software forced someone to make with no evidence, no counterargument, and no memory of the outcome. Closing it took us ten modules, sixteen prompts, and a registry that refuses to let the gap reopen. The result is a product where the AI leads with a proposal on every judgment call — and where the human's yes, no, or "close, but here's the real number" is the signal everything else compounds on.

Put this playbook to work — with the OS built for it.

T2D3 OS turns the method behind this guide into working modules: ICP, personas, positioning, content, and a full GTM plan. Start free.