Token spend scales with every transaction, and examiners want to know exactly which model touched which decision. Routing and audit aren't optional here.
In financial services, AI spend is not an R&D line. It compounds with transaction volume. A fraud model that costs a fraction of a cent per decision is a rounding error at a thousand transactions a day and a P&L problem at ten million. The firms that win price the model to the decision, not the department.
Meanwhile the regulatory posture hardened: examiners now ask which model made which decision, on what data, under whose approval. "We use a large provider" is not an answer. Model-risk management expects lineage (inputs, model version, output, human touchpoints) for every consequential decision.
The structural answer is routing: small, specialized, private models handle the routine tail (extraction, classification, triage) at fixed, auditable cost, while frontier models are reserved and logged for the hard cases. Margin and auditability turn out to be the same architecture.
SLM-first routing for fraud triage, document extraction, and KYC. Frontier models stay gated, logged, and reserved for the genuinely hard reasoning.
Model-access tiers, data-residency boundaries, and immutable decision logs an examiner can walk through line by line.
A spend & token audit with chargeback by desk, then a governed routing layer that caps cost per decision without slowing the business.
SLM-first by default; escalation to frontier models is explicit, priced, and logged.
Customer data stays inside declared jurisdictions; the routing layer enforces it, not a policy PDF.
Immutable logs mapping every automated decision to model, version, inputs, and human approvals.
Token spend attributed by desk and workflow, so AI cost lands where it is incurred.
Hard per-workflow caps, so no agent loop can run the meter unattended.
With decision lineage: a log tying each consequential output to the model version, inputs, and the human who approved it. We install that as infrastructure. It is not a documentation exercise after the fact.
Yes. The waste is rarely in the hard problems. It is frontier models doing routine extraction and classification. Routing that tail to specialized small models typically cuts the bill dramatically with no user-visible change.
It starts with a spend and token audit: usage by team, workflow, and model, with chargeback design and a 30-day savings plan. The TCO calculator on this site gives you the first approximation in five minutes.