A governed document pipeline is an inference bill in a costume. Optim sits underneath Forge and does three things: it routes every page to the smallest model that can still clear your accuracy floor, it meters cloud and token usage in real time against named owners, and it forecasts the next quarter as a distribution rather than an average.
You do not buy Optim. It ships in the runtime at every tier. It is the reason your cost per page is lower than the vendor quoting you next.
Prices per token keep falling and bills keep climbing, because usage grows faster than prices drop — and one agentic document workload burns what a chatbot never could. A pipeline that reads a whole estate is the largest inference bill most firms will ever run.
You cannot control a real-time cost with a monthly report — and you cannot cut a bill you can’t predict or attribute.
A clean native-PDF statement and a photographed handwritten amendment are not the same problem, and paying frontier prices for both is how document AI budgets get killed in year two. Optim scores each page for difficulty and sends it down the cheapest lane that can still hold the line your risk function drew.
Swipe to follow the routing →
A forty-page pack is rarely uniform. Page 3 is a clean table and page 27 is a fax of a signature block. Routing is decided page by page, so one bad page does not price the whole document at the top rate.
Text layer present or absent, image quality, layout density, table structure, language, and the field set the pack expects from that page. No content leaves your boundary to make the decision.
If the cheaper route cannot meet the confidence threshold set for a field, the page escalates to a larger model — and if that still cannot clear it, the value abstains into the steward queue. Optim is never permitted to move a value past a gate.
A page that fails the cheap lane is re-run, not accepted. The re-run costs a fraction of a cent. A confidently wrong field in a client statement costs considerably more than that.
Estates are repetitive — the same master agreement underlies four hundred amendments. Context that recurs is cached rather than re-sent, which is most of the saving on a large corpus.
Which models occupy the small, mid and large lanes is configuration set at Fitting. In your own VPC that means your endpoints, your contracts and your rates — Optim routes between them, it does not resell them.
Routing is one lever of four. They compound, and they are applied in a deliberate order — cheapest and least costly to quality first. In a representative model, $63.9k → ~$22.4k of monthly inference before any commitment discount.
| Lever | What it does | Effect |
|---|---|---|
| Semantic + context caching | Similar prompts served from cache; repeated context billed at a fraction of fresh input. On a document estate this is the big one — the same master agreement underlies four hundred amendments. | first lever |
| Model routing & right-sizing | Straightforward pages go to a smaller model; escalates only when the confidence gate demands it. | ~80% / answer |
| Batch & off-peak scheduling | Latency-tolerant traffic moves to asynchronous endpoints. A nightly backfill of ten years of archives does not need the interactive path. | ~50% cheaper |
| Provisioned throughput | The forecast sets your 24/7 baseline; the floor is pre-bought and bursts fall back to pay-per-token. | forecast floor |
Caching is first because it is the only lever that gets better with popularity. Every other lever trades something — a smaller model trades headroom, batching trades latency, a commitment trades flexibility. Cache hits trade nothing.
Routing is second because it is measurable per page: the router scores routable traffic against your confidence gates before it moves any of it, and the split is a number your team can read in the command centre.
Not a monthly bill you reconcile afterwards. A live meter inside your boundary that knows which document family, which queue and which tenant caused the spend — while the run is still going.
| Metered | Attributed to | Refresh | Why it matters |
|---|---|---|---|
| Input & output tokens | Document family, pack, run, queue | Live, per run | Tells you which estate is expensive before finance does |
| Model lane mix | Small / mid / large share per family | Live, per run | A drift toward the large lane is the earliest signal your estate quality changed |
| GPU hours & idle time | Node pool, region, tenant | Live | Idle capacity scales to zero between bursts — in your VPC that is your bill, not our margin |
| Cache hit rate | Document family | Live | A falling hit rate explains a rising bill without anyone guessing |
| Cost per page | Family, pack, business unit | Live, rolling | The single number to compare against whatever the last vendor put in a slide |
| Cost per exception | Steward queue, approver | Per run | Shows what a threshold actually costs, so setting one becomes an informed decision |
Document estates are bursty in ways that averages hide: quarter-end, renewal season, a large portfolio transfer, a regulatory deadline. An average forecast is comfortable and wrong exactly when it matters. Optim forecasts a distribution and provisions against the part of it that hurts.
A distribution tells you when the expensive version of next quarter becomes likely — the only moment intervention is still cheap. Quarter-end, renewal season and a large portfolio transfer are all in here.
Every forecast is a range with stated confidence. Capacity is provisioned against the upper band and cost is planned against the middle one, so a busy quarter is a plan rather than an incident.
Short horizon drives autoscaling decisions during a run. Long horizon tells you which volume band you will be in at renewal — before the conversation, not during it.
Learned from your own run history: day-of-month effects, quarter-end spikes, family-level growth and the drift in your lane mix. It is a model of your estate, not an industry benchmark.
Budget thresholds fire when the forecast crosses them, not when the spend does. A projection that breaches next month is something you can still act on.
Model an unonboarded document family against your measured cost per page before committing to it. Useful for deciding which estate to bring on next, and for saying no to one.
Optim predicts from your history. A genuinely new document family or a step change in volume will be outside the distribution until it has run a few times. The bands widen when the model is unsure, and that widening is shown rather than smoothed away.
A cost panel beside the accuracy and throughput panels, broken down by document family and queue. Same screen, same login, same roles.
Every run closes with a measured cost per page and the lane mix that produced it — including the free evaluation, so you can model your full estate before anyone quotes you.
Usage, attribution and forecast series are readable through the same authenticated API as records and lineage, so this data lands in your own FinOps tooling rather than staying in ours.
Forecast breaches and lane-mix drift fire as webhooks into whatever you already use. Optim does not ask to own your alerting.
It was previously sold on its own as a cloud and AI spend platform. It has not been switched off and it has not been abandoned — it has been repositioned. The forecasting, attribution and routing engine now runs underneath Forge, where it does more good and where it is the reason the unit economics work.
If you hold a current Optim commitment: the same support, the same release cadence, the same named contact, for the term you contracted. Renewal will be an honest conversation about where the engine is going, not a migration deck.
Two crowded markets and two buyers cannot be served well by one small team, and we would rather say that plainly than sell you a second product we cannot give proper attention to. If cloud and AI spend governance is your actual problem, the FinOps category has good dedicated vendors and we will point you at them.
Existing Optim customer wondering what this means for your renewal? You will get a straight answer.
Talk to us about an existing agreementForge reads your document estate into validated records with a page reference behind every field, deployed where your regulator requires, with the weights staying yours.