Every model. One subscription. Zero prompt babysitting.
AI Engine
57 models across 16 providers, ~44 across 12 in active rotation — and growing — behind task-tuned prompts, routed for the best outcome at the best cost, with credits included in your plan.
You should not need subscriptions to five AI tools, a prompt library, and a system for pasting context between them. The OS catalogs 57 models across 16 providers — roughly 44 across 12 in active rotation: Anthropic, OpenAI, Google, and a dozen more, plus image generation — and routes every task to the model that wins for that job, already loaded with your foundation, your evidence, and your locked judgment.
The engine improves itself while you work. 516 task-tuned prompts drive the OS, and 213 of them have already been improved during beta through eval-backed optimization and A/B benchmarking across models. Every generation's cost and outcome is tracked per model and per use case, so each task runs on what performs best at the best price — that's how your plan stays affordable while output quality keeps climbing.
It's a glass box, not a black box: generations show which sources grounded them, drafts are verified before they're saved, and your edits teach the system. You spend AI credits, included with every plan, instead of managing API keys and per-seat AI subscriptions.
- 57
- models routed
- 16
- providers — and growing
- 555
- task-tuned prompts
- 245
- improved during beta
Under the hood
How every task runs
Five stages between your click and a draft you can defend. You see the AI narrate each one while it works.
- 01
Ground
The app assembles the context for you — locked foundation, ICP, personas, evidence, uploads. Context is the single biggest lever for AI performance, and here it's built in, never pasted in.
- 02
Prompt
One of 516 task-tuned prompts takes over, engineered for that specific job.
- 03
Route
The dispatcher picks the winning model from a catalog of 57 — around 44 in active rotation — with automatic failover if a provider stumbles.
- 04
Verify
Output is checked against its evidence before it's allowed to land in your playbook.
- 05
Learn
Your edits and eval results feed the next optimization round. The engine gets better while you work.
20 AI module generators
18 guided GTM skills
75 automated routines
4 AI roles + Max, the assistant
The model fleet
57 models. 16 providers. One subscription.
~44 across 12 in active rotation — new models join as the frontier moves
Anthropic
Claude Fable 5 · Opus 4.8 · Sonnet 5 · Haiku 4.5
OpenAI
GPT-5.5 Pro · GPT-5.5 · GPT-4.1 · GPT Image 2
Gemini 3.1 Pro · 3.5 Flash · Nano Banana Pro · Nano Banana 2
xAI
Grok 4.3 · Grok Imagine
Mistral
Large 2 · Magistral · Codestral
DeepSeek
R1 · V4 Pro · V4 Flash
Alibaba
Qwen 3.7 Plus · Qwen 3.7 Max
Moonshot
Kimi K2.6
Perplexity
Sonar Pro · Sonar
Cohere
Command A+
Sakana
Fugu Ultra · Fugu
Fal.ai
FLUX 2 Pro · Seedream 5
Ideogram
Ideogram 3.0
Recraft
Recraft V4.1 — raster + SVG
OpenRouter
GLM-5.2 · MiniMax M2.5 · DeepSeek V4 on US infra
T2D3 self-hosted
Qwen on our own GPUs — fast, private, cheap
Live from the engine — last 30 days
- 99M+
- tokens routed
- 83k+
- AI calls dispatched
- 72k+
- signal-quality checks
- ~33k
- tokens of context per playbook task
Text, reasoning, vision, and image generation — including inference on our own hardware for the tasks where fast and private beats big. You never manage an API key, and a heavy month never means a new subscription.
The part nobody else does
An engine that optimizes itself
Most AI products pick one model and hope. The OS treats models and prompts as a market: everything is benchmarked, everything is measured, and only winners ship.
Eval-backed prompt optimization
Every prompt carries eval cases. An improvement ships only when it beats the incumbent on evidence — 213 of the 516 prompts have already been improved during beta, and the sweep never stops.
Cross-model benchmarking
The same task runs against multiple models and is judged head-to-head. The winner becomes the default for that use case — so the routing table reflects results, not brand loyalty.
Cost-per-outcome routing
Token cost and output quality are tracked per model, per use case. Tasks run on what wins at the best price — that's how plan prices stay optimal while output quality keeps climbing.
Provider health & failover
Provider incidents are detected automatically and tasks re-route across the fleet — a model outage somewhere on the internet is not your problem.
What's in it
57 models, 16 providers
Claude, GPT, Gemini, Grok, Mistral, DeepSeek, Qwen, Perplexity, and more — ~44 across 12 in active rotation, routed per task, growing as the frontier moves, no API keys to manage.
Prompt optimization & A/B benchmarks
Prompts and models are benchmarked against each other on real tasks with eval evidence — the winner ships. 213 of 516 prompts improved during beta so far.
Cost-optimal routing
Cost and outcome tracked per model per use case, so every task runs on what wins at the best price — keeping plans affordable as usage grows.
Image generation
Visual-brand imagery through leading image models, on-brand from your visual foundation.
Task-tuned prompts
Every module ships with prompts engineered and continuously optimized for its job — you never write or maintain prompts.
Glass-box grounding
Outputs show the sources that grounded them, so you can trust — and defend — what you ship.
Verify before persist
Generated output is checked against its evidence before it lands in your playbook.
Max, the AI assistant
A cross-module assistant that knows your workspace and the T2D3 method.
AI credits in every plan
Metered credits included monthly; heavy months never mean new subscriptions.
Works with
Put AI Engine to work on a real client.
$0 base fee through the open beta. Upgrade when the engagement lands.