We Graded Our Entire Content Library Overnight. It Cost About 16 Cents.
Most B2B content libraries are unaudited inventory: hundreds of pages nobody has re-read since they shipped. T2D3 OS's Content Review grades every published piece against your locked ICP, personas, and value props — five axes, verbatim quotes as evidence, a keep / refresh / retire verdict — and rolls it up into one number, the Signal Ratio. We ran it on our own site: 105 pieces graded overnight on a self-hosted model for about 16 cents total. Here is how library grading works, why the empty cells in the coverage map matter more than the full ones, and what our own numbers taught us.
Stijn Hendrikse · Aug 26, 2026
Your content library is probably unaudited inventory. Content Review, in T2D3 OS, grades every published piece against your locked ICP and personas — five axes, verbatim evidence, a keep / refresh / retire verdict — then rolls the library into one number: the Signal Ratio. We graded our own 105 pages overnight for about 16 cents.
That last sentence is the whole argument, so let me earn it. This is the part of the Deep Content system that closes the loop on everything already published: not "should we write this?" (that's the value rating that argues with itself) but "should we still be hosting this?" Most teams never ask, because auditing a few hundred pages by hand costs weeks. When grading a piece costs a fraction of a cent and runs while you sleep, the question stops being expensive — and the answers get uncomfortable fast.
How: a grade is an argument against your own strategy, not a rank
Every content tool on the market will "audit" your library. Almost all of them mean: rank pages by traffic, flag thin word counts, count keywords. That is grading content against the search engine. Content Review grades content against your strategy — the ICP, buyer personas, value propositions, and brand voice you locked in your foundation. A page with great traffic that speaks to nobody you sell to should not survive the review.
Each piece gets five scores, 0–100:
- Targets fit — does it serve one of the queries you actually want to win? Not "does it rank," but "is this a fight we picked."
- Signal fit — does it speak to your ICP and personas: their pains, their vocabulary?
- Uniqueness — does it say anything your own similar pieces don't already say? Overlap with your own corpus is the main retire signal, not overlap with the internet.
- Voice fit — does it sound like your locked brand voice, down to your banned and preferred terms?
- Proof — does the piece earn its claims with quantified specifics, dates, and first-party evidence, or does it assert generically?
The part I insist on: every score must be justified by a verbatim quote from the piece itself. The grader's contract is blunt — "a quote that is not found word-for-word in the piece VOIDS that axis." No quote, no score. That single rule is the difference between a grade you can act on and a vibes-based dashboard. When a piece scores 42 on proof, you see the exact sentence that earned it: the generic assertion, quoted back at you, character for character. You can disagree with the judgment; you cannot claim it wasn't looking at your actual words.
Here's what that looks like on a real page of ours — our customer-acquisition-cost guide, graded overnight. Proof scored 30, and the evidence is the sentence itself: "Optimal CAC gets the fastest return of your money through monthly recurring revenue growth" — a claim with no data, case study, or first-party evidence behind it. Voice fit caught a single word: "Create leverage with your leads" — "leverage" is on our banned-terms list, and the grader quoted the exact sentence that used it. And when the model tried to justify a targets-fit score with the page title instead of a sentence from the body, the system voided the axis: "voided: quote not found in the piece." The rule applies to the grader too.
And when the grader has nothing to judge, it says so with a null, never a fake zero. A library whose brand voice isn't locked yet gets "not scored: brand voice not locked yet" on the voice axis — an honest gap, not a punishing 0 that poisons the average.
The Signal Ratio is a percentage with a denominator, on purpose
The roll-up number is the Signal Ratio: the percentage of pieces that clear the uniqueness bar (60 by default — the threshold is policy, and yours to raise), counted over the pieces whose uniqueness was actually scored. A voided or unscored axis neither passes nor fails — it leaves the denominator entirely, and the dashboard says how many it left out. It answers the one question a CEO will actually repeat in a board meeting: how much of what we've published is still worth someone's time?
We render it with its denominator, always — "85% — 61 of 72 graded pieces at uniqueness ≥ 60," never a naked "85%." A percentage without a denominator is how dashboards lie: 90% of 10 scored pieces and 90% of 400 are different companies. The denominator also keeps the system honest about coverage — if a third of your library couldn't be scored on uniqueness, or was never graded at all, the Signal Ratio says so before it says anything else.
Our own number, first run, no rehearsal: 85% — 61 of the 72 pieces that earned a uniqueness score on www.t2d3.pro, out of 105 graded in total (the other 33 got an honest null on that axis, not a fake zero). Don't let the 85% fool you: the verdict mix behind it is 2 keep, 46 refresh, 57 retire — being singular in our own corpus doesn't spare a piece from arguing generically or speaking to nobody. I'd love to tell you our library — the one this article lives in — came back pristine. Two keeps out of 105 says it did not, and that's the point of a review you don't get to negotiate with.
The empty cells in the coverage map matter more than the full ones
While grading, the reviewer also attributes each piece to the persona tier it serves best (your primary, secondary, or tertiary buyer) and the funnel stage it was written for (awareness, evaluation, or decision). Plot the library on that 3×3 grid and you get the coverage map: where your published words actually landed, versus where your strategy says they should be.
Full cells are nice. Empty cells are the product. An empty cell is a persona your foundation says matters, at a funnel stage where they're making up their mind, with nothing addressed to them — a strategic hole made visible. So every empty cell in T2D3 OS carries a one-click "draft ideas for this cell" that feeds the persona and stage straight into ideation. The map doesn't just describe the gap; it starts the work that closes it.
Ours came back lopsided, which surprised exactly nobody who has read our content strategy guide and then looked at what a decade of blogging actually produces: of our 105 graded pieces, exactly two land at the decision stage — the moment a buyer actually chooses — while 30 crowd our primary persona's evaluation row, and 61 couldn't be confidently attributed to any persona tier at all: pages written for "the market" rather than for anyone in it. Years of writing tilt toward the awareness content that's easiest to write, and away from the decision-stage pieces that close deals. Now we can see the tilt, cell by cell, with a button in each gap.
A grade remembers which strategy it was graded against
Here's the failure mode nobody designs for: you re-position. New ICP, sharper personas. Every content audit you've ever run is now silently wrong — it graded pages against a strategy that no longer exists, and nothing anywhere says so.
Every grade in Content Review carries a foundation fingerprint — a hash of the exact locked ICP, personas, value props, and voice it was judged against. When your foundation changes, the fingerprints stop matching, and the dashboard reports staleness exposure: the share of grades that were issued against a previous version of you. Stale grades aren't deleted or hidden — they're marked, and the next overnight sweep re-grades those pieces against the current foundation. The system self-heals in both directions, too: if the page changes, its content hash changes, and it gets re-graded on the next pass without anyone asking. On our first run, staleness was 0% — every grade issued against the current foundation, which is exactly what a first run should say. The number becomes interesting the day we re-lock our ICP, and unlike every audit spreadsheet we've ever made, this one will notice.
Retirement follows the same evidence discipline. The verdict rules require that a "retire" name its reason — "duplicates our own better content, or serves no target and no persona" — and when it's duplication, name which piece it duplicates. You never get "retire, trust me." You get "retire: this says what that says, and that says it better," with both pieces on the table.
Wow: overnight, on our own hardware, for a fraction of a cent per piece
Now the economics, because they change what a review is.
The grading runs on a model we host ourselves — the same self-hosted fleet that fills our icon library overnight. Measured from our own call logs, not estimated: a full grade — five scored axes, five verbatim quotes, audience attribution, verdict, suggestions — costs $0.0015 per piece (an average of 6,272 tokens of grounding and piece text in, 1,446 tokens of judgment out, on a 120-billion-parameter open-weight model; 148 calls in total including the retries, because we count those too). Our entire 105-piece library graded for $0.16, while we slept.
Compare that to the alternative you already know: a content audit as a consulting deliverable. Two to four weeks, a spreadsheet, a five-figure invoice, obsolete the day your positioning moves. At a tenth of a cent per piece, the audit stops being a project and becomes a pulse — the sweep just runs, nightly, forever, and re-grades whatever changed. This is the principle we build the whole OS on: token prices fall every quarter, so never design a system around rationing judgment. Design it around the assumption that judgment-per-piece is nearly free and human attention is the scarce thing — then spend the machine extravagantly and the human precisely.
Five refresh drafts a week, and none of them publish themselves
A grade that doesn't start work is a report. So the review feeds a deliberately narrow pipe we call the trickle: each week, the system picks the five pieces where a refresh would move the library most — ranked by the grade evidence, not by recency — and drafts the rewrite overnight. The refresh brief is the grade itself: fix only what the evidence named, "preserve verbatim any customer quotes, named facts, numbers," keep the structure, never invent.
Then the drafts stop and wait. Every one lands in the review inbox as a proposal — nothing is ever auto-published. You read the draft next to the original, and you vote: accept, reject, or edit first. And the votes are not just a gate — they're training data. Each accept or reject, with its reason, feeds back into how the next week's drafts get written, the same signal / noise / glint loop that runs everywhere else in the OS. The rewriter gets a little more yours every week.
Our first trickle picked its five refresh candidates — after first declining to run at all, because grading 105 pieces had already spent the grader agent's daily job budget and the budget guard doesn't make exceptions for the boss. The top pick: our customer-acquisition-cost guide, the most singular piece in the corpus on uniqueness (70 — nothing else we've published covers CAC) argued with generic assertions (proof: 30) to nobody in particular (signal fit: 20). Keep the topic, rebuild the evidence, aim it at someone. That's a real editorial decision, surfaced by evidence, waiting on a human vote — which is the division of labor we keep coming back to: the machine drafts and argues, the human decides and teaches.
One more honesty note, because this article is itself inside the system it describes: we ran this piece through our own retrieval lint — the deterministic check that verifies an article is shaped for answer engines (answer-first lead, specific headings, quotes and statistics per section, a real FAQ). It passed — answer-first lead, specific H2s, FAQ, and Last updated all green — and it still printed three warnings: one section leans on its neighbor, and not every section carries as many statistics and sourced quotations as the lint wants. We're publishing with the warnings visible, because that's the point of a checker you don't get to negotiate with either.
FAQ
What does Content Review actually grade content against?
Your locked GTM foundation — ICP, buyer personas, value propositions, brand voice — plus the content targets you chose to win. Not keyword density, not traffic. Five axes (targets fit, signal fit, uniqueness, voice fit, proof), each score justified by a verbatim quote from the piece or voided.
What is a good Signal Ratio?
There's no universal benchmark — the honest answer is "higher than your last one, on a denominator that covers your whole library." The default bar counts pieces with uniqueness ≥ 60 as signal, over the pieces that could actually be scored on that axis; our own first run came in at 85% (61 of 72 scored, from 105 graded), and the verdict mix — 2 keep, 46 refresh, 57 retire — is what the refresh queue exists to move.
Does the system retire or rewrite content automatically?
No. Verdicts and refresh drafts are proposals. The weekly trickle drafts at most a handful of refreshes, every one waits for a human vote, and retirement is a decision you take with the evidence in front of you. Your votes then teach the next round of drafts.
Why grade on a self-hosted model instead of a frontier cloud model?
Volume economics with nothing lost on the task: grading is bounded judgment over a fixed rubric with all evidence supplied, which an open-weight 120B model handles well — at roughly $0.0015 per piece from our measured logs. That's what makes grading the entire library nightly viable instead of sampling it quarterly.
How does the review know when a grade has gone stale?
Every grade stores a fingerprint of the exact foundation version it was judged against and a hash of the page content. If either changes — you re-lock your ICP, or the page is edited — the mismatch is visible immediately and the piece re-grades on the next overnight sweep.
Last updated 2026-08-26.