Free until October 1. Lock your foundation and run your first client diagnostic before your Q1 pipeline conversations start.

Join the beta

The Document That Got Heavier: Grading Evidence by What It Produced

Evidence that earns its weight: grading marketing sources by the outcomes they produce.

Stijn Hendrikse · Aug 30, 2026

ShareLinkedInXEmail

Every marketing team has a folder of documents it calls "research." A win-loss summary from last spring. Three customer interview transcripts. A positioning deck from an agency engagement two contracts ago. A competitor teardown somebody made for a board meeting. Ask which of these documents is actually good — which one has earned its place in your strategy work — and you will get opinions, shrugs, or a rating somebody assigned the day it was uploaded and never touched again.

Here is the uncomfortable truth: nobody knows. Not because the team is careless, but because the question is unanswerable at upload time. A document's value is not a property of the document. It is a property of what the document does — what work it feeds, and whether that work survives contact with human judgment.

So we stopped asking people to rate documents and started measuring what happens downstream. In T2D3 OS, evidence earns its weight from outcomes. We call the pattern outcome attribution, and the shorthand we use internally is the title of this piece: some documents get heavier over time, because the work they ground keeps getting accepted.

The problem with rating evidence at the door

Most systems that manage knowledge — DAMs, wikis, "research repositories" — grade content at ingestion, if they grade it at all. A human tags it, maybe stars it, and the rating fossilizes. Three things go wrong from that moment.

First, the rating reflects a guess. When a transcript arrives, nobody knows yet whether it contains the phrase that will anchor your messaging or forty minutes of pleasantries. The honest answer at upload time is "we'll see."

Second, the rating goes stale. The market moves, the product moves, the strategy moves. A document that was load-bearing in March can be quietly wrong by October, and its five-star rating never notices.

Third — and this is the one that actually costs money — the rating never learns from use. Your team makes dozens of small judgments every quarter that reveal which evidence is pulling its weight: they accept some AI-drafted work nearly untouched, they tear other drafts apart and rewrite them. Each of those moments says something precise about the sources that fed the draft. In almost every tool on the market, that information evaporates.

How the loop closes

In T2D3 OS, three mechanisms that each exist for their own reasons happen to join into a closed loop.

Lineage. Every AI generation in the system records which uploaded documents grounded it. This exists for transparency — outputs display "grounded in N sources" so you can see what an answer was built on — but it also means we always know, for any piece of drafted strategy, exactly which evidence fed it.

The lock. Strategy work in T2D3 OS isn't done when the AI drafts it; it's done when a human reviews it and locks it — a deliberate conviction step. At the moment of locking, the system captures the edit diff: how much the human changed between the draft and the version they committed to.

Attribution. This is where it gets interesting. A lock accepted nearly as-is is a positive vote on the documents that grounded it — the evidence did its job. A draft that got heavily rewritten before locking is a negative vote — the grounding produced something the human couldn't stand behind. Those verdicts roll up, per document, into a running behavioral record: how many times has this document's downstream work been accepted, how many times reworked?

The result is a weight adjustment on the document itself. The next time the system assembles evidence for a generation, documents with a track record of feeding accepted work count a little more. Documents that keep grounding drafts people rewrite count a little less. Nobody filled in a form. Nobody scheduled a quarterly evidence review. The grading happened as a byproduct of work the team was doing anyway.

The restraint is the design

The obvious failure mode of a system like this is overreach, and most of the design effort went into preventing it. Four guardrails matter.

Only decisive signal moves weight. We don't treat every edit as a verdict. A lock counts as accepted only when the human changed almost nothing — under roughly fifteen percent of the draft. It counts as reworked only when they rewrote half or more. Everything between is deliberately ignored as neutral. Moderate editing is normal, healthy collaboration; treating it as a judgment on the evidence would inject noise into the very system that exists to reduce it.

The adjustment is clamped. Outcome attribution nudges a document's weight; it never takes it over. A document with a great track record gets modestly heavier, not dominant. Evidence selection is still driven primarily by relevance, quality tier, and recency — the multiplier is a thumb on the scale, not a hand.

A human mark always wins. Elsewhere in T2D3 OS, people can explicitly vote a document up or down as signal. That explicit human judgment is never overridden by the behavioral statistic. If you have marked a document as noise, no accumulation of accepted locks will sneak it back into your grounding. The algorithm proposes; the human's mark is law.

It grades real sources only. The attribution pass distinguishes actual uploaded documents from synthetic or derived material in the corpus, so the weight lands on the evidence a team can actually act on — keep, replace, or go get more like it.

What you can apply without our software

The mechanism is ours, but the discipline is portable. Three practices any team can adopt this quarter:

Track lineage manually. When you produce a significant piece of strategy — a messaging framework, an ICP definition, a pricing rationale — write down which two or three sources actually shaped it. A footnote in the doc is enough. You are building the connective tissue that makes evidence auditable at all.

Review outcomes, not libraries. Skip the annual "clean up the research folder" ritual. Instead, when a piece of work gets approved with barely a change, ask: what fed this, and how do we get more evidence like that? When something gets torn apart in review, ask the mirror question. You will learn more from five outcome moments than from fifty document ratings.

Let recency and results retire your evidence. The hardest habit: treat a document's past importance as irrelevant. That foundational deck from 2024 doesn't get grandfather rights. If nothing it feeds survives review anymore, it has told you what it is now.

Why this compounds

The deeper reason we built outcome attribution is not tidiness — it's compounding. T2D3 OS is designed around a loop: human signal grounds AI drafts, human judgment on those drafts sharpens the signal, and sharper signal makes the next draft better. Every mechanism that closes a loop makes the whole system improve a little faster, for everyone in the org, silently.

Evidence that earns its weight is one of the quieter loops, but it changes how a marketing organization relates to its own knowledge. The library stops being a warehouse and becomes something closer to a portfolio — positions that grow or shrink based on performance, managed by results rather than sentiment. And because the grading rides on judgments your team was already making, the cost of maintaining it is exactly zero additional meetings.

The document that got heavier never asked for the promotion. It just kept feeding work that people signed their names to.

If you want to see what your own evidence base looks like when it's graded — where it's deep, where it's thin, and what it's actually producing — the GTM diagnostic at t2d3.pro is a free place to start.

Join the conversation

Put this playbook to work — with the OS built for it.

T2D3 OS turns the method behind this guide into working modules: ICP, personas, positioning, content, and a full GTM plan. Start free.