Overview

Portfolio-wide AI spend, routing posture and enforcement

Demo — illustrative data, mid-2026 snapshot Meridian Financial
WHY NOW

Daily AI spend — last 90 days

Actual routing (s = 15%) vs the risk-adjusted optimum (s = 83%). The gap is the prize.

Actual (s=15%) At optimum routing (s=83%)

Spend by model

Last 30 days

Spend by department

Last 30 days — totals reconcile to the monthly run-rate

Sticker price ≠ cost: the cheaper-listed model is pricier in 21.8% of head-to-head pairs — by up to 28×. Metering is the only truth.

Driven by hidden reasoning-token volume — arXiv:2603.23971. Hover for the reversal example.

Reversal example: Gemini 3 Flash listed 78% cheaper than GPT-5.2 but came in 22% higher in actual total cost — arXiv:2603.23971. (Citation example only — not a routable model in this demo.)

Agent cost explorer 14 agents · click a row for detail

Agent Department Model Req / day Avg tok / req Monthly cost Cost-of-Pass Useful-token Trend Status

Cost-of-Pass = cost per attempt ÷ success rate · Useful-token ratio = valid output tokens ÷ total billed

Substitution-share control

The % of traffic routed from frontier models to the open-weight floor. Govern the substitution.

Engine: blendedCost(s) — Routing module, scenario model v2
0% 99%
▲ Meridian today: 15% Risk-adjusted optimum: ~83% ▲
Max output-token cap (p95 clamp)
Trims runaway completions.
Route reasoning-heavy tasks only when justified
Reasoning tokens carry a ~31.5× average price premium — arXiv:2603.28576.

Blended cost curve — the U-shape executives need to see

Cost falls as routing shifts to the open floor, then the s8 misrouting-risk penalty bends it back up past ~90%.

“Stop forecasting vendor prices — forecast your routing. Substitution share is the dial you control.”

Routing rules

Model-tier-to-task examples

Ticket classification
Opus 4.8 → DeepSeek V4 Flash (commodity work — open floor)
Contract summarization
Sonnet 4.6 → DeepSeek V4 Pro with mid-tier fallback
M&A diligence analysis
stays Opus 4.8 (accountable core — trust is the scarce layer)

Pricing reference — mid-2026 project-verified anchors

$ per M tokens · blended = 3:1 input:output = (3×in + 1×out)/4

ModelTier InputOutputBlended

“Mid-2026 snapshot, blended 3:1 input:output. Prices and model versions move monthly — re-verify against official pricing pages before external use.”

“GPT-5.5 cached input is $0.50 (standard); the $10.00 in / $45.00 out tier is the long-context (>272K) surcharge, whose cached rate is $1.00. GPT-5.6 Sol (GA July 2026) holds the same $5/$30 standard price. All routing math here uses the $5/$30 standard tier.”

“Claude Sonnet 5 (GA, intro $2.00/$10.00 through Aug 31, 2026, then $3.00/$15.00) supersedes Sonnet 4.6 at the same standard price. Anchors retained at the mid-2026 snapshot.”

Routing math uses tier anchors: frontier blended $10.00, open blended $0.35 — 2026 frontier/open gap ≈ 28.6×.

Only ~1 in 5 firms has mature agent governance · ~35% couldn’t stop a rogue agent (Deloitte) · 97% of AI security incidents lacked access controls (IBM) · EU AI Act fines begin Aug 2026

Live enforcement feed

Control-plane events, today (AAGATE pattern — monitor / warn / block per agent)

Enforcing