Overview
Portfolio-wide AI spend, routing posture and enforcement
Daily AI spend — last 90 days
Actual routing (s = 15%) vs the risk-adjusted optimum (s = 83%). The gap is the prize.
Spend by model
Last 30 days
Spend by department
Last 30 days — totals reconcile to the monthly run-rate
Sticker price ≠ cost: the cheaper-listed model is pricier in 21.8% of head-to-head pairs — by up to 28×. Metering is the only truth.
Driven by hidden reasoning-token volume — arXiv:2603.23971. Hover for the reversal example.
Agent cost explorer 14 agents · click a row for detail
| Agent | Department | Model | Req / day | Avg tok / req | Monthly cost | Cost-of-Pass | Useful-token | Trend | Status |
|---|
Cost-of-Pass = cost per attempt ÷ success rate · Useful-token ratio = valid output tokens ÷ total billed
Substitution-share control
The % of traffic routed from frontier models to the open-weight floor. Govern the substitution.
Blended cost curve — the U-shape executives need to see
Cost falls as routing shifts to the open floor, then the s8 misrouting-risk penalty bends it back up past ~90%.
“Stop forecasting vendor prices — forecast your routing. Substitution share is the dial you control.”
Routing rules
Model-tier-to-task examples
Pricing reference — mid-2026 project-verified anchors
$ per M tokens · blended = 3:1 input:output = (3×in + 1×out)/4
| Model | Tier | Input | Output | Blended |
|---|
“Mid-2026 snapshot, blended 3:1 input:output. Prices and model versions move monthly — re-verify against official pricing pages before external use.”
“GPT-5.5 cached input is $0.50 (standard); the $10.00 in / $45.00 out tier is the long-context (>272K) surcharge, whose cached rate is $1.00. GPT-5.6 Sol (GA July 2026) holds the same $5/$30 standard price. All routing math here uses the $5/$30 standard tier.”
“Claude Sonnet 5 (GA, intro $2.00/$10.00 through Aug 31, 2026, then $3.00/$15.00) supersedes Sonnet 4.6 at the same standard price. Anchors retained at the mid-2026 snapshot.”
Routing math uses tier anchors: frontier blended $10.00, open blended $0.35 — 2026 frontier/open gap ≈ 28.6×.
Live enforcement feed
Control-plane events, today (AAGATE pattern — monitor / warn / block per agent)
Executive summary — Meridian Financial, Q2 2026
Prepared with CostGrid · AI cost POV companion
The EBITDA & exit bridge
At s = 0.80 the token line falls to 30.9% of frontier-default (~69% reduction); net saving after hosting/observability overhead: 60%.
The three-layer stack — where value accrues
The POV’s reframe of the AI value chain
“Governance, security, sovereignty, the accountable boundary. Scarce, sticky — and where margin and exit value accrue.”
“the lever you control… where substitution share lives: model-tier-to-task, evals, caps, fallback.”
“Deflating ~10x/yr toward the cost of electricity… Commoditizing — no durable margin.”
“Value never accrues where prices converge — it accrues where trust is scarce.”
CostGrid operates the top two layers.
Forecast — 2026 → 2030, bifurcation scenario
Frontier vs open blended $/Mtok (log scale, left) · enterprise GenAI spend $B (right)
The gap widens to ~70× — the substitution prize grows every year.
Falling prices don’t shrink the bill — they grow it: spend $69.1B (2026) → $207.3B (2030).
Forecast markers
Top 3 recommendations — AI-cost playbook
“Match model tier to task value — don’t tokenmaxx”
“Install Agentic FinOps: a budget owner per cost row, spend circuit-breakers before agents scale”
“Keep an open-source option — the open-weight floor is your pricing leverage and your migration insurance.”
“Tokens are cheap. Trust is not. Govern the substitution. Price every AI dollar against risk-adjusted EBITDA — and the value it creates at exit.”