Free until October 1. Lock your foundation and run your first client diagnostic before your Q1 pipeline conversations start.

Join the beta

From Tailings to Ingot: A Quality Ladder for Marketing Evidence

A six-tier quality ladder for marketing evidence, computed instead of claimed.

Stijn Hendrikse · Aug 20, 2026

ShareLinkedInXEmail

Ask a marketing team where their evidence lives and they'll point at a drive: interview transcripts next to pitch decks, survey exports next to a competitor's ebook, all of it filed, none of it graded. Then they feed that drive to an AI and wonder why the output sounds confident and generic at the same time.

The problem isn't the AI. It's that every document in the pile carries the same weight. A verbatim customer interview and a scraped blog post walk into the prompt as equals, and the model — which has no way to know one is gold and one is filler — averages them.

We think evidence needs what every other refining process has: a ladder. In T2D3 OS, every piece of marketing evidence sits on a six-tier quality ladder — Tailings, Ore, Assay, Concentrate, Ingot, Refined — and the tier decides how much weight it carries in everything the AI drafts. Here's how the ladder works, why no one is allowed to hand-stamp a grade, and how to build a version of it whatever tools you run.

The six tiers

The names come from ore refining, because that's the honest metaphor for what marketing evidence is: raw material of wildly uneven concentration.

  • Tailings — voted out. Duplicates, commentary, stale strategy. Tailings carry zero weight and are gated out of module grounding entirely.

  • Ore — raw, unreviewed material. A scraped page, an auto-captured document nobody has looked at. Usable, lightly weighted.

  • Assay — unreviewed but promising: the classifier's quality score says there's real content here, a cut above raw Ore.

  • Concentrate — human-reviewed evidence. Someone kept it, marked it up, or confirmed it belongs. This is the working tier of a healthy evidence base.

  • Ingot — conviction. A locked decision — an ICP the team committed to, a ratified value prop — re-enters the corpus at this tier, so your settled strategy grounds downstream work as strongly as your best raw evidence.

  • Refined — the apex: human-validated conviction, such as a locked decision the team has additionally endorsed, or evidence deliberately curated to replace its raw sources.

The distribution matters more than any single grade. A corpus that is mostly Ore produces drafts that hedge; a corpus with real Concentrate density and a few Ingots produces drafts that sound like your company because, in a traceable sense, they are.

One clarification before going further: the tier answers how good a piece of evidence is. It deliberately does not answer where it sits in your pipeline or how directly the buyer spoke — those are separate axes in our system, graded separately. Collapsing quality, stage, and directness into one score is how most content-scoring schemes end up meaning nothing.

Derived on read, never hand-stamped

Here's the design decision we'd defend hardest: nobody assigns a tier. There is no dropdown where an eager intern promotes a pitch deck to Refined. A document's tier is computed, at the moment it's read, from facts the system already holds — its classifier quality score, its human votes, its triage verdict, its provenance, its age.

Hand-stamped labels rot. The doc that deserved "high quality" in January is superseded by March, but the label doesn't know that. Derived grades can't rot, because they're recomputed from current facts every time the evidence is used. When a fact changes — you thumb-up a document, a duplicate is found, a newer version arrives — the grade changes with it, immediately, with no batch job to wait for.

The practical rule for any team: grade evidence by observable properties, not by opinion at filing time. "Has a human confirmed this?" and "how directly did the buyer speak?" are properties. "Feels important" is not.

Humans move the ladder

Every rung of the ladder above Assay is reached through human judgment — the ladder is a record of it:

  • A raw scraped document is pinned to Ore until a human reviews it, no matter how well it scores. Volume can't buy trust.

  • A thumbs-up from anyone on the team is human validation: the piece rises to at least Concentrate, instantly.

  • When the team locks a strategic decision, that conviction re-enters the evidence base as an Ingot.

  • A thumbs-up on a locked Ingot lifts it to Refined — conviction, endorsed.

Notice what this means: your team's ordinary behavior — reviewing, keeping, voting, locking — is the grading process. There is no separate librarian job. The ladder converts judgment your team was already exercising into weight the AI can respect.

Evidence stales; conviction doesn't

Every tier's weight also decays with age, down to a floor — because a customer quote from 2023 should ground your 2026 messaging more softly than one from last quarter. But two kinds of pieces are exempt from decay: locked convictions and human-composed syntheses. A decision your team ratified doesn't get quieter just because time passed; it gets retired when you relock it, deliberately.

That split — evidence decays, conviction persists until explicitly revised — mirrors how good strategy teams actually operate, and it's a rule you can adopt in any stack: date-weight your raw inputs, but never let your settled decisions age out silently.

The ladder is yours to tune

One more thing we refuse to hardcode: what "good" means. The weight each tier carries is tunable per organization, live, by an admin — because a product-led startup drowning in usage data and a services firm rich in interview transcripts should not weight the same tiers identically.

Two safeguards keep tuning honest. A progressive gate strengthens the Concentrate tier's privileges as human-review coverage rises — by default, once roughly seventy percent of the corpus has been human-reviewed, unreviewed material is held to a stricter standard, so early-stage teams aren't punished while mature ones aren't diluted. And echo-guards cap how much of any generation's grounding can come from the org's own locked convictions — your beliefs amplify your evidence; they're never allowed to drown it out and turn the system into a mirror.

Building a ladder without our software

You don't need T2D3 OS to benefit from the idea. A minimum viable ladder is three tiers and two habits:

  1. Three folders (or tags): Unreviewed, Reviewed, Decided. Everything lands in Unreviewed. A weekly fifteen-minute pass promotes what someone actually read and kept. Decided is reserved for artifacts the team formally committed to.

  2. Two habits: never ground on Unreviewed alone, and date-stamp everything. When you brief an AI (or an agency), pull from Reviewed and Decided first, and say so in the brief: "weight the interview transcripts above the deck."

That's it. You'll feel the difference in the first serious piece of AI-assisted work, because the model finally knows what you trust.

The deeper point is the one we keep coming back to across T2D3 OS: quality measured, not asserted. Anyone can claim their AI is grounded in "your data." The ladder makes the claim inspectable — this draft stood on four Concentrate pieces and one Ingot, weighted like so, decayed like so. If you've read our piece on the four pillars, Ore, and Concentrate in the GTM diagnostic, this is the machinery underneath those words.

If you're curious where your own evidence would land, start where most teams do: run the free GTM diagnostic at t2d3.pro. It won't grade your drive — but it will show you, from your public message alone, where the thin evidence is already showing through.

Put this playbook to work — with the OS built for it.

T2D3 OS turns the method behind this guide into working modules: ICP, personas, positioning, content, and a full GTM plan. Start free.