An LLM should never grade its own citations

The two-strangers rule: how we verify every citation in AI-drafted content.

Stijn Hendrikse · Sep 24, 2026

ShareLinkedInXEmail

One bad citation can turn a 60-source GTM analysis into a 60-source manual audit, and for a fractional CMO that erases the speed gain before the CEO meeting. Our rule is simple: an LLM never certifies its own evidence. Two independent checks have to agree first, one deterministic and one isolated, before a citation counts as verified.

Every citation is checked by two independent halves. A deterministic test measures token overlap against the cited excerpt. An isolated judge then sees only the claim and that excerpt. For web claims, both have to agree before we call the citation verified.

Verified means two strangers agreed, not one model praising its own homework.

A self-grading model keeps the mistake it just made

Ask an AI assistant to draft positioning and it can answer fluently from nowhere. Upload 40 documents and the problem changes. It can now answer from everywhere.

The customer interview, last year's pitch deck and a competitor article may all enter the context with equal apparent authority. A polished sentence can hide which source shaped it.

Now ask the same system whether its citation supports that sentence. The grader already knows the intended argument. It may reconstruct what the writer meant instead of checking what the excerpt says.

That is correlated judgment. The writer and grader share context, language patterns and likely failure modes.

A green check from that process proves the model can defend its answer. It does not prove the cited excerpt supports the published claim. Andrea Nicholas, a management consultant who runs the T2D3 approach with two clients, put the underlying standard plainly when she described the framework: "it's soup to nuts. It's like plug and play." A framework a consultant resells under her own name has to hold up under someone else's reading, not its author's.

One citation can contain the right words and the wrong meaning

Consider a real distinction from an interview with Andrea Nicholas.

Andrea said she was "eager to get certified" in the T2D3 framework. A draft could turn that into: "Andrea Nicholas is a certified T2D3 practitioner."

The citation contains "certified", "T2D3" and Andrea's name. A surface check sees strong overlap. The actual excerpt describes future intent, not current status.

That tense change is small on the page and material in a client document.

The reverse problem also appears. A faithful paraphrase may use different words from its source. A semantic judge can recognize the support while a deterministic test finds too little overlap.

Neither check is sufficient alone. Their disagreement is the useful signal.

The two-strangers rule separates matching from judgment

The citation verifier divides the decision into two jobs:

  1. The deterministic half checks textual contact. It compares the claim with the exact cited excerpt through token overlap. It cannot excuse a weak match because the broader document "felt relevant."

  2. The isolated judge checks meaning. It receives only the claim and excerpt. It cannot use the draft's wider argument to fill missing evidence.

For a web claim, both checks need to pass. If one disagrees, the citation stays unverified.

This approach accepts a real tradeoff. Some valid citations will need another look. We would rather inspect an awkward paraphrase than publish a confident claim whose source never made it.

Citation verification does not make a weak source strong

A citation can support a claim perfectly while the source remains stale, speculative or irrelevant to the decision.

We saw this problem when different materials entered one workspace: a customer's buying-moment interview, an agency-polished deck and a saved competitor blog. All three could support sentences. They should not carry equal weight.

That is why citation checking belongs beside evidence grading.

In T2D3 OS, evidence carries a computed grade on a six-tier ladder. The grade reflects its quality score, human confirmation marks and triage status. Weight can decay with age because an interview from 2023 should not outvote one from last month.

Outputs also show their receipts: how many sources grounded the work, the evidence-band mix and how much was unreviewed during generation.

The citation verifier answers, "Does this excerpt support this claim?" Evidence grading answers, "How much authority should this source carry?" A publishable draft needs both answers.

Seven questions expose weak AI citations before a client does

When reviewing AI-written content, we use questions that force each claim back to its evidence:

  • Can the claim stand alone? Split sentences that combine a sourced fact with an unsupported conclusion.
  • Does the citation open to the exact excerpt? A relevant document is weaker than a relevant passage.
  • Does the excerpt support every name, number and date? One unsupported modifier changes the claim.
  • Did possibility become certainty? "Plans to certify" cannot become "is certified."
  • Did the draft add causation? Sequence alone does not prove one event caused another.
  • Was the judgment isolated? A grader that sees the full draft can inherit its assumptions.
  • Can the reviewer see source quality? Citation support and evidence authority are separate decisions.

This checklist takes longer than accepting a model's green check. It takes less time than reopening every source after a CEO finds one bad citation.

Adversarial quality control starts by removing self-approval

Judging panels, devil's advocates and blind critics can improve an AI draft. Citation grading has a narrower responsibility: prove that each published claim stays inside its cited evidence.

That is why this mechanism belongs in our "Adversarial Quality Control for Marketing AI" series. The useful adversary is independent of the original answer.

For your next AI-drafted article, start with the web claims. Give one checker only the claim and excerpt. Add a deterministic overlap test. Publish "verified" only where both agree, then route every disagreement back to a human reviewer.

Join the conversation

Put this playbook to work — with the OS built for it.

T2D3 OS turns the method behind this guide into working modules: ICP, personas, positioning, content, and a full GTM plan. Start free.