Quality Report
How T2D3 OS gets better every month
Token prices fall every year. The quality of the work should not depend on that. Every AI instruction in this product is versioned, reviewed and revertible — and when we claim an improvement, it is verified against tests no optimizer is allowed to edit.
Measured August 7, 2026. Updated nightly.
Under governance
Not a prompt folder. Every instruction the AI runs is a catalogued artifact with a version history, an owner and a rollback.
AI instructions governed
Prompts and skills. None running on an unreviewed default.
Reviewed by an AI panel
Three independent models score each one; a consensus pass decides.
Median quality score
Out of 100, on a calibrated rubric. Below 50 means materially deficient.
Tests we are not allowed to change
This is the part that makes the rest of the page mean something. A frozen set of evaluation cases sits outside the improvement loop — no optimizer, automated or human, may edit them. It is how our scores went up can be told apart from we rewrote the exam.
Frozen evaluation cases
Locked. Growth is gap-driven and human-approved, never automatic.
Evaluation cases in total
Run as a regression tripwire before a change reaches you.
We publish the method and the count, never the cases themselves — a published exam is one you can study for.
Promises enforced by the build
These are not intentions. Each is a check that fails our build, so it cannot be skipped on a busy week.
A new model means a new review
When we adopt a frontier model, our build stays red until every prompt routed to that tier has been re-reviewed. Prompts tuned for the old model do not silently inherit the new one.
Specifications must match the code
Each guided skill declares what it reads, writes and calls. A gate compares that against what the code actually does, and the allowed gap only ever shrinks. Today 34 of 34 align.
Routed by fit, never down-tiered to save money
Each task runs on the model that suits it — 10 models across 6 providers in the last 90 days. Falling token prices buy you better work, not a cheaper bill for us.
What your edits teach it
When you change an AI draft before locking it, that edit is the signal. It is distilled into durable guidance and injected into future generations for that module — so the system converges on how you work.
Judgment signals captured (30d)
Aggregated across workspaces. Never attributed to one.
The same measurements are available for your own workspace, so you can see how the system is doing on your data rather than in aggregate.
See your workspace quality report →