Your Insights Folder Is Full of Recycled AI Prose. We Built a Detector.
A detector for recycled AI prose in your marketing evidence — and why you need one.
Stijn Hendrikse · Aug 29, 2026
We build an AI-powered marketing operating system. We also built a detector that flags AI-written prose in the evidence our customers upload. If that sounds like a contradiction, sit with it for a second — because the reason we needed it says something uncomfortable about where B2B marketing is right now.
Here's what actually happens when a company hands over its strategy documents. Mixed in with the customer interviews, the sales decks, and the survey exports, there's a growing layer of documents that were themselves written by a language model: the positioning doc a founder drafted with a chatbot last spring, the "brand strategy" a previous agency generated, the market overview somebody pasted from a research assistant. The people uploading them aren't hiding anything — often they've simply forgotten which documents were authored and which were generated.
The problem isn't that these documents exist. The problem is what happens if your marketing system treats them as evidence.
The exhaust-inhaling problem
An AI marketing system is only as differentiated as the material it's grounded in. The whole point of feeding it your customer interviews and your team's hard-won positions is to pull its output away from the generic average and toward what only your company knows.
Recycled model prose does the opposite. A document that a language model wrote is, statistically, a sample of the average — polished, plausible, and very close to what the same model would say for any of your competitors. When that document sneaks into your evidence base labeled "strategy," you are quietly re-importing the mean you were trying to escape. Generate messaging grounded in it and you get an echo of an echo: confident output with no first-party substance underneath.
And the failure is invisible at the moment it happens. The generated positioning doc doesn't look weaker than the real interview sitting next to it — it usually looks better: cleaner structure, tighter phrasing, no awkward human digressions. Every downstream artifact it touches inherits a little of its genericness, and six months later someone asks why all the messaging sounds like everyone else's.
So before any document can carry weight as evidence in T2D3 OS, we estimate one thing about it: does the body of this document read like LLM-generated prose, human-authored writing, or something else entirely (a transcript, a spreadsheet)?
The signature is structural — and once you see it, you can't unsee it
The detector doesn't call a model to judge another model. It's deliberately deterministic: a structural read of the text, the same result every time, explainable down to each signal. We validated it against a real client's document library, where the generated documents shared a signature so consistent it was almost a watermark:
The "X is NOT / X IS" contrast frame. Generated strategy docs love to define a company by negation first — "Acme is NOT a consultancy. Acme IS a platform." — often as parallel bulleted pairs.
Arrow chains. Reasoning compressed into "problem → insight → solution → outcome" sequences, several per page.
Em-dash density. Language models reach for the em-dash far more often than human business writers do. A high count per thousand characters is one of the strongest single tells.
Hyper-parallel bullets. Long runs of bullets with identical grammatical structure — every line "Verb + noun phrase + benefit" — at a density human authors rarely sustain.
Boilerplate hedges. The phrase-level tells you already recognize: "at its core," "more than just," "designed to," "in today's ever-evolving landscape," and their relatives.
Tricolons everywhere. "Faster, smarter, and more aligned." Models over-produce the three-part list the way nervous speakers over-produce filler words.
None of these is damning alone; human writers use every one of them. The detector scores them together, per thousand characters, and maps the total to three bands: human, mixed, or ai.
Just as important is the negative evidence. A transcript — timestamps, speaker turns — contains human speech by definition, so it pins to human no matter how many em-dashes the transcription software inserted. So does a spreadsheet export: currency symbols and digit density are the fingerprint of data, not authored prose. And anything too short to judge is called human, on purpose. A one-line note is not "AI-generated"; it's a note.
The human stays the judge
Two design rules matter more than the scoring math.
First, the score is orthogonal to origin and to intent. A flagged document is not an accusation. Clients genuinely hand over AI-written strategy docs — sometimes knowingly, because that draft was a useful starting point. The flag doesn't say someone did something wrong; it says this text should probably not count as first-party evidence about your buyers.
Second, the verdict is always overridable by a human. Every document shows its band and the specific signals that drove it — "em-dash density (41)," "heavily bulleted (62%)" — and a person who knows the document's real history can overrule the estimate with one click. The heuristic informs the triage decision; it never makes it. That's the same posture we take everywhere in T2D3 OS: the machine recommends with its reasoning visible, the human decides.
What actually happens to a flagged document is triage, not deletion. Maybe it moves out of the evidence lane and into reference material. Maybe the human says "yes, this was generated, but we then edited it heavily and it is our position now" — a perfectly legitimate verdict. The point is that the decision gets made consciously instead of by omission.
How to audit your own folder by hand
You don't need our detector to run the check that matters. Set aside an hour with your "strategy" and "insights" folders and ask, per document:
- Do you know who wrote it? Not which team — which person, with what inputs. If nobody can say, treat it with suspicion.
- Scan for the signature. NOT/IS framing, arrow chains, relentless parallel bullets, the hedge phrases above. Two or three of these together in a business document is a strong hint.
- Look for anything a model couldn't know. Real customer names, actual numbers, a specific lost deal, a verbatim quote. Generated strategy is conspicuously free of falsifiable detail.
- Sort, don't shred. Move suspected generated docs out of your evidence pile into a "drafts and references" pile. They can still be useful — as writing to react to — but they should never again be cited as what your market told you.
Most teams that run this exercise find the same thing our detector finds: the folder is more synthetic than anyone remembered, and the genuinely first-party material — interviews, tickets, survey verbatims — is thinner than assumed. That second discovery is the valuable one, because it tells you what to go collect.
The honest irony
Yes: an AI system, checking documents for AI prose, so that its own AI output stays grounded in human evidence. We're aware. But that loop is exactly the shape of the next few years of marketing. Generation is cheap and getting cheaper; what's scarce is knowing which of your inputs are real. A system that can't tell its evidence from its own exhaust will drift toward the average — fluently, confidently, and in perfect parallel bullets.
If you'd like to see what a grounded read of your company looks like, the free GTM diagnostic at t2d3.pro analyzes your public messaging with every filter visible. And if you want your evidence base managed this way end to end, start with the founding member program.