T2D3 OS
HomeFor CMOsCost & ROIPricingBlog
Sign inStart free
T2D3 OS

The B2B SaaS go-to-market method, running as software.

hello@t2d3.club
Product
  • All features
  • GTM Foundation
  • T2D3 Playbook
  • Content Studio
  • Website & GEO
  • AI Engine
Platform
  • Command Center
  • Relay
  • Team & Talent
  • Agency Workspace
  • Expert Marketplace
  • The Living GTM
  • Humans are the loop
  • Work on the marketplace
  • Pricing
Library
  • Beyond templates
  • What T2D3 OS is not
  • Learn (blog)
  • Wren's journal
  • Changelog
  • Templates & playbooks
  • Ready-to-run playbooks
  • Masterclass
  • Certification
  • Glossary
  • The playbook
  • Books by Stijn
  • Podcast
  • Newsletter
Free tools
  • Brand Voice lab
  • GTM diagnostic
  • Growth labor calculator
  • Cost calculator
  • ROI calculator
Company
  • About
  • The team
  • Coaching
  • For fractional CMOs
  • Founding cohort
  • Join the journey
  • FAQ
  • Support
© 2026 T2D3. Triple twice, double three times.
PrivacyTerms

Quality Report

How T2D3 OS gets better every month

Token prices fall every year. The quality of the work should not depend on that. Every AI instruction in this product is versioned, reviewed and revertible — and when we claim an improvement, it is verified against tests no optimizer is allowed to edit.

Measured September 22, 2026. Updated nightly.

Under governance

Not a prompt folder. Every instruction the AI runs is a catalogued artifact with a version history, an owner and a rollback.

819

AI instructions governed

Prompts and skills. None running on an unreviewed default.

98%

Reviewed by an AI panel

Three independent models score each one; a consensus pass decides.

74

Median quality score

Out of 100, on a calibrated rubric. Below 50 means materially deficient.

Tests we are not allowed to change

This is the part that makes the rest of the page mean something. A frozen set of evaluation cases sits outside the improvement loop — no optimizer, automated or human, may edit them. It is how our scores went up can be told apart from we rewrote the exam.

30

Frozen evaluation cases

Locked. Growth is gap-driven and human-approved, never automatic.

409

Evaluation cases in total

Run as a regression tripwire before a change reaches you.

We publish the method and the count, never the cases themselves — a published exam is one you can study for.

Promises enforced by the build

These are not intentions. Each is a check that fails our build, so it cannot be skipped on a busy week.

  • A new model means a new review

    When we adopt a frontier model, our build stays red until every prompt routed to that tier has been re-reviewed. Prompts tuned for the old model do not silently inherit the new one.

  • Specifications must match the code

    Each guided skill declares what it reads, writes and calls. A gate compares that against what the code actually does, and the allowed gap only ever shrinks. Today 43 of 43 align.

  • Routed by fit, never down-tiered to save money

    Each task runs on the model that suits it — 12 models across 6 providers in the last 90 days. Falling token prices buy you better work, not a cheaper bill for us.

What your edits teach it

When you change an AI draft before locking it, that edit is the signal. It is distilled into durable guidance and injected into future generations for that module — so the system converges on how you work.

541

Judgment signals captured (30d)

Aggregated across workspaces. Never attributed to one.

The same measurements are available for your own workspace, so you can see how the system is doing on your data rather than in aggregate.

See your workspace quality report →