Why our content critic never sees the writer’s reasoning
Blind review for AI-drafted content: the critic that never sees the writer's reasoning.
Stijn Hendrikse · Sep 23, 2026
Our content critic grades every draft blind: it sees the article, the evidence and the locked brand context, but never the writer's explanation of its own choices. It scores five dimensions and enforces a 7.0 pass mark. A 6.9 draft does not ship. The reasoning is hidden so the critic grades the work, not the defense of it.
That is deliberate. An AI writer can produce a persuasive explanation for almost any choice it makes. If the critic sees that reasoning, it may start grading the explanation instead of the work.
Our content pipeline removes that advantage. The critic sees the draft, the evidence behind it and the locked brand context. It scores what reached the page across five dimensions: point of view, evidence, voice, risk and information gain.
The tradeoff is extra review time and compute. The gain is a draft a fractional CMO can put in front of a client without defending how it was produced.
This is part three of our series, "Adversarial Quality Control for Marketing AI."
A persuasive rationale can hide an average draft
Writers know the problem from human editorial reviews.
Someone presents a weak headline, then spends five minutes explaining the strategic thinking behind it. By the end, the room is discussing the strategy rather than the headline customers will read.
AI compounds that problem because its reasoning arrives quickly and sounds complete. The explanation can be more convincing than the resulting article.
We use information asymmetry to counter this. The writer knows why it made each choice. The critic does not. It has to judge the artifact using the same material available to the eventual reader.
That separation costs something. A critic may reject a draft the writer considers salvageable. We accept that cost because the writer will not accompany every article to explain what it meant.
The published work has to carry the argument alone. For a fractional CMO juggling four clients, that matters twice over: nobody is in the room to explain the draft, and roughly 1.5 unbillable hours a day already go to reloading client context before any real work starts.
The 4+1 rubric tests quality beyond fluent prose
Our critic grades four familiar dimensions and one that most editorial reviews miss.
Point of view: Does the draft make a claim a competent reader could disagree with? A summary of accepted advice may read well while adding little value.
Evidence: Can the important claims be traced to supplied interviews, documents, examples or named operating experience? Fluency does not count as proof.
Voice: Does the piece sound like this company? For T2D3 OS, that means an experienced operator putting the cost and mechanism on the table.
Risk: Could the draft damage the author's standing through an invented fact, unsupported promise, exposed client detail or careless comparison?
Then comes the additional test.
Information gain: Does the reader leave with an idea, distinction or mechanism they did not have before?
This fifth score matters because AI can produce accurate content that contributes nothing. It can restate the first page of search results with cleaner transitions. That output is safe, polished and forgettable.
Five dimensions, four of them conventional and one that catches the failure mode nobody else grades for. Andrea Nicholas, a management consultant who uses the T2D3 approach with her own clients, described what makes something worth carrying to a CEO: "I just would go back to the framework and the, the thrust of T2D3. You know, when I explain it to people and I say, here's what it means, that gets their attention." Information gain is the score that asks whether a draft earns that same attention.
For this article, the information gain is the review design itself: conceal the writer's reasoning so the critic cannot be persuaded by it.
A 7.0 threshold only matters when it can stop publication
Many content rubrics are decorative. The reviewer gives a low score, the deadline arrives and the piece ships anyway.
Our critic has an enforced threshold of at least 7.0 on a zero-to-ten scale, averaged across the five dimensions. A 6.9 is a miss, and a near miss remains a miss.
The purpose is not mathematical precision. A score of 7.0 is not objectively good while 6.9 is objectively bad. The threshold creates an operational decision before deadline pressure enters the room.
Above it, the draft can proceed. Below it, another pass begins.
For human review, one approach is to score each dimension from zero to ten, then average the five scores. Keep evidence and risk visible rather than burying them in a general impression.
A practical review sheet can use these questions:
- Point of view: What specific claim would another experienced operator challenge?
- Evidence: Which supplied source supports each consequential claim?
- Voice: Could this paragraph appear under a competitor's logo unchanged?
- Risk: What would legal, a client or the CEO challenge?
- Information gain: What can the reader now explain or do that they could not before?
Set the passing score before anyone reviews the draft. Otherwise, the rubric becomes a vocabulary for rationalizing work that was already approved.
Persistent retries turn criticism into accumulated work
A failed review should not send a near-complete draft back to zero.
Our retry passes persist their work. The next pass retains what already survived and addresses the critic's findings. A draft that scored 6.9 becomes the starting point for revision rather than a discarded generation.
That distinction changes the economics of adversarial review.
Without persistence, stricter criticism creates waste. Each rejection risks another full draft, another voice drift and another review cycle.
With persistence, the critic can be demanding without becoming destructive. The system keeps the sound argument, supported evidence and on-brand sections while revising the weak parts. Only the dimensions that scored below the bar get reopened, which is usually one or two of the five, not all of them.
The same principle works in a human team. Return a scored draft with specific failure points. Do not ask for "a stronger version." Preserve approved sections and reopen only what missed the bar.
Blind review protects the judgment attached to your name
Fractional CMOs sell judgment. The retainer is not the volume of copy produced. It is the decision about which claim is safe, distinctive and worth carrying into the market.
That judgment becomes harder to defend when AI drafts arrive from five tools with five different contexts. A polished paragraph can conceal weak evidence or a generic point of view. With four clients, that is four sets of context to keep straight and five re-entries a day to pay for.
Blind criticism adds an opposing force. The writer generates. The critic challenges. The 7.0 threshold decides whether the work advances. Persistent retries keep the useful work intact.
Start with one draft and five scores. Hide the writer's explanation, set the threshold before review and record why each failed dimension missed it. That review record gives the next pass specific work to improve.