Free until October 1. Lock your foundation and run your first client diagnostic before your Q1 pipeline conversations start.

Join the beta

Trust This, Weigh That Less: Ranking Evidence Before the AI Ever Sees It

Why we rank marketing evidence into trust bands before the AI reads a word of it.

Stijn Hendrikse · Aug 23, 2026

ShareLinkedInXEmail

Ask an AI assistant to draft your positioning and it will answer instantly, fluently, and from nowhere. Ask it again after uploading forty documents and you have a different problem: it now answers from everywhere. The interview transcript where a customer described the exact moment they went looking for you, the pitch deck your agency polished last spring, a competitor's blog post someone saved for reference — all of it lands in the model's context with the same implied authority.

That is how most retrieval-augmented systems work today. They are very good at finding relevant text and completely silent about how much to trust it. The model is left to guess, and models guess the way people skim: recency, fluency, and confidence win. A well-written pitch deck routinely outranks a messy transcript of a real buyer talking.

We think the ranking is the product. Before any generation runs in T2D3 OS, the evidence is banded — and the model is told, explicitly, what each band is worth.

Not all evidence deserves an equal vote

Marketers already know this instinctively. Put these four documents side by side:

  • A raw customer interview transcript, where a buyer explains in their own words why they switched.
  • A win/loss note from your sales team, written the day the deal closed.
  • A case study your team wrote about that same customer.
  • A blog post about your category from a vendor with something to sell.

All four are "relevant" to an ICP exercise. They are nowhere near equally trustworthy. The transcript is the buyer speaking. The win/loss note is one hand removed. The case study is the same story after it survived a marketing filter — it tells you what your team wanted to say. And the category blog post is somebody else's positioning wearing an educational costume.

A retrieval system that treats those as interchangeable context isn't neutral. It is quietly voting for whichever document is longest, cleanest, or most recent. When teams complain that AI output feels generic, this is often why: the distinctive, messy, first-hand material got averaged together with the polished, derivative material, and polish won.

The three bands

Every generation in T2D3 OS that draws on your evidence base receives it as a compact digest in three trust bands, with instructions the model reads before your content does. The actual guidance, near verbatim:

Sources are banded by quality. PRIMARY — locked decisions and human-composed documents — and SUPPORTING — reviewed, high-grade signal — carry your core claims; ground firm conclusions in these. EXPLORATORY sources are raw and unreviewed; use them for breadth, fresh angles, and idea generation, never as proof. When PRIMARY and EXPLORATORY conflict, prefer PRIMARY.

Three things are happening in that paragraph that a context dump never does.

First, trust is explicit. The model isn't inferring authority from tone; it is told which sources carry conclusions and which only suggest directions.

Second, raw material still gets in. The fix for untrustworthy context is not deleting it — fresh, unreviewed uploads are where new angles come from. EXPLORATORY material is welcome for breadth and hypothesis; it simply can't serve as proof. You keep the serendipity without the contamination.

Third, conflicts have a rule. When a raw document contradicts a locked decision, the model doesn't split the difference — it prefers the decision, and the contradiction is a finding to surface, not noise to smooth over.

What decides the band

Banding would be arbitrary if a person hand-stamped it. Instead, the bands are computed from machinery that runs underneath every document:

A quality ladder. Each piece of evidence carries a derived grade on a six-tier ladder, computed on read from its quality score, human confirmation marks, and triage status — never hand-assigned. The weighting is tunable per organization, and it decays with age, because an interview from 2023 should not outvote one from last month.

Human verdicts. Every document in the base gets a human vote — signal, noise, or worth keeping as an outlier. A thumbs-down from you re-grades a piece immediately. The algorithm can propose; it can never override your mark.

Conviction re-enters as evidence. When you lock an ICP or a set of value props, that decision itself becomes high-band evidence for everything downstream. Your convictions compound — with guards so an organization's own echo doesn't drown out fresh buyer voice.

Distance from the buyer. A separate axis scores how many hands a quote has passed through. A testimonial in a pitch deck already survived a marketing filter — it tells you about the marketer, not the buyer — so it weighs less than the same claim in a raw transcript.

None of this requires a librarian. Upload the document; the system grades it, you correct it when it's wrong, and the correction sticks.

One seam, thirty generators

Here is the part that changed how we build. The banding isn't a feature of one module — it is the seam through which roughly thirty downstream generators receive evidence. The ICP builder, the persona work, the messaging framework, the content pipeline, the website copy, the competitive analysis: all of them inherit the same weighting for free, because they all read evidence through the same digest.

That has a compounding consequence. When you thumbs-down a stale deck or confirm a great transcript, you aren't tuning one output — you are re-weighting the evidence for every future generation across the whole system. One person's five-second correction quietly improves everyone's next draft.

And because the grounding is recorded, you can check the work. Outputs show their receipts — grounded in N sources, with the band mix visible, including how much of the grounding was unreviewed at generation time. If a draft leaned mostly on EXPLORATORY material, you can see that before you trust it.

How to apply this in any stack

You don't need our product to use the principle. If you are assembling context for any AI tool today:

  1. Sort your documents into three piles before you paste. Decisions and first-hand buyer voice; reviewed, reliable material; everything raw or derivative. Even a crude sort beats none.
  2. Tell the model the piles exist. One paragraph — "treat the first group as authoritative, the last as inspiration only" — measurably changes how it argues.
  3. Prefer directness. When you can choose between a transcript and the case study built from it, choose the transcript. Filtered language produces filtered output.
  4. Date your evidence. Mark anything older than a year and say how to treat it. Models have no native sense of staleness.
  5. Re-litigate conflicts, don't average them. If two sources disagree, that's a question for a human, not a blend for a model.

Do that manually and you'll feel the ceiling quickly: every new document means re-sorting, every teammate keeps their own piles, and nothing you learned last week carries into next week's prompt. That ceiling is the reason we built the banding as infrastructure — graded once, corrected in seconds, inherited everywhere.

The quiet argument underneath

There's a belief hiding in this design: the scarce input to good marketing is no longer writing capacity — it is trustworthy evidence, ranked honestly. Models are abundant. Your customers' actual words, your team's real conviction, the decisions you've committed to: those are scarce, and they deserve better treatment than being tossed into a context window alongside a competitor's blog post.

If you want to see where your own evidence stands, the free GTM diagnostic at t2d3.pro reads your public messaging the way we read evidence — and shows you what it found, banded and sourced.

Put this playbook to work — with the OS built for it.

T2D3 OS turns the method behind this guide into working modules: ICP, personas, positioning, content, and a full GTM plan. Start free.