Optim / a product inside Forge / not a separate licence

Optim. The routing and cost engine inside Forge.

A governed document pipeline is an inference bill in a costume. Optim sits underneath Forge and does three things: it routes every page to the smallest model that can still clear your accuracy floor, it meters cloud and token usage in real time against named owners, and it forecasts the next quarter as a distribution rather than an average.

You do not buy Optim. It ships in the runtime at every tier. It is the reason your cost per page is lower than the vendor quoting you next.

In the runtime at every tier Per-page model routing Realtime cloud + token metering Never trades accuracy for cost
01 / your reality today

Two bills are exploding. Neither is governed.

Prices per token keep falling and bills keep climbing, because usage grows faster than prices drop — and one agentic document workload burns what a chatbot never could. A pipeline that reads a whole estate is the largest inference bill most firms will ever run.

20–35%
Cloud spend wastedZombie clusters, idle GPUs, mistagged workloads.
40–60%
Token budget wasteRedundant calls, oversized models, no caps.
~30%
Can attribute to a teamFinance owns budget; engineering owns the console.
24–72h
Before the bill landsDashboards refresh daily. Spend moves by the second.

You cannot control a real-time cost with a monthly report — and you cannot cut a bill you can’t predict or attribute.

The four figures above are published industry ranges for cloud and AI spend, not measurements from a gothink deployment. They describe the problem Optim was built for. The number that matters for your estate is the cost per page from your own run, which is measured rather than cited.
02 / intelligent model routing

Not every page needs the expensive model.

A clean native-PDF statement and a photographed handwritten amendment are not the same problem, and paying frontier prices for both is how document AI budgets get killed in year two. Optim scores each page for difficulty and sends it down the cheapest lane that can still hold the line your risk function drew.

Optim·router·decision per page, against your accuracy floor live
How Optim routes a page to the smallest sufficient model Pages enter a router that classifies each one by difficulty. Clean native text is sent to a small fast model, mixed layouts to a mid-size model, and degraded scans or handwriting to the largest model. The router is bounded by the accuracy floor your confidence gates set: if the cheaper route cannot meet the threshold for a field, the page is escalated to a larger model or abstains into the steward queue. All three lanes converge back into a single typed record. PAGES IN native text mixed layout scan / handwriting ROUTER difficulty scored SMALL MODEL cheapest path · most pages MID MODEL tables, dense layout LARGE MODEL degraded scans, handwriting ACCURACY FLOOR SAME RECORD accuracy: unchanged cost/page: lower route never lowers the bar STEWARD QUEUE a named approver cannot clear the floor → escalate, then abstain

Swipe to follow the routing →

Cheapest route that still clears your gate Escalated because the page needed it Could not clear the floor — abstained to a steward
Decision

Per page, not per document

A forty-page pack is rarely uniform. Page 3 is a clean table and page 27 is a fax of a signature block. Routing is decided page by page, so one bad page does not price the whole document at the top rate.

Signals

What the router looks at

Text layer present or absent, image quality, layout density, table structure, language, and the field set the pack expects from that page. No content leaves your boundary to make the decision.

The hard constraint

The accuracy floor wins, always

If the cheaper route cannot meet the confidence threshold set for a field, the page escalates to a larger model — and if that still cannot clear it, the value abstains into the steward queue. Optim is never permitted to move a value past a gate.

Escalation

Retry costs less than a wrong answer

A page that fails the cheap lane is re-run, not accepted. The re-run costs a fraction of a cent. A confidently wrong field in a client statement costs considerably more than that.

Caching

Repeated context is not re-sent

Estates are repetitive — the same master agreement underlies four hundred amendments. Context that recurs is cached rather than re-sent, which is most of the saving on a large corpus.

Model choice

Your model family, your lanes

Which models occupy the small, mid and large lanes is configuration set at Fitting. In your own VPC that means your endpoints, your contracts and your rates — Optim routes between them, it does not resell them.

03 / inside the build: token economics

Stack the levers, then pre-buy the floor.

Routing is one lever of four. They compound, and they are applied in a deliberate order — cheapest and least costly to quality first. In a representative model, $63.9k → ~$22.4k of monthly inference before any commitment discount.

See what gets metered
LeverWhat it doesEffect
Semantic + context cachingSimilar prompts served from cache; repeated context billed at a fraction of fresh input. On a document estate this is the big one — the same master agreement underlies four hundred amendments.first lever
Model routing & right-sizingStraightforward pages go to a smaller model; escalates only when the confidence gate demands it.~80% / answer
Batch & off-peak schedulingLatency-tolerant traffic moves to asynchronous endpoints. A nightly backfill of ten years of archives does not need the interactive path.~50% cheaper
Provisioned throughputThe forecast sets your 24/7 baseline; the floor is pre-bought and bursts fall back to pay-per-token.forecast floor
Why the most-asked question stops being the most expensive one

Caching is first because it is the only lever that gets better with popularity. Every other lever trades something — a smaller model trades headroom, batching trades latency, a commitment trades flexibility. Cache hits trade nothing.

Routing is second because it is measurable per page: the router scores routable traffic against your confidence gates before it moves any of it, and the split is a number your team can read in the command centre.

The figures above are a representative model, not a quoted saving. They describe how the levers compose on a worked example. What your estate actually costs per page is measured on your own run and printed in your evaluation report — and none of these levers is ever allowed to lower the accuracy floor to reach a number.
04 / realtime cloud and token usage

Every dollar and every token, attributed as it is spent.

Not a monthly bill you reconcile afterwards. A live meter inside your boundary that knows which document family, which queue and which tenant caused the spend — while the run is still going.

MeteredAttributed toRefreshWhy it matters
Input & output tokensDocument family, pack, run, queueLive, per runTells you which estate is expensive before finance does
Model lane mixSmall / mid / large share per familyLive, per runA drift toward the large lane is the earliest signal your estate quality changed
GPU hours & idle timeNode pool, region, tenantLiveIdle capacity scales to zero between bursts — in your VPC that is your bill, not our margin
Cache hit rateDocument familyLiveA falling hit rate explains a rising bill without anyone guessing
Cost per pageFamily, pack, business unitLive, rollingThe single number to compare against whatever the last vendor put in a slide
Cost per exceptionSteward queue, approverPer runShows what a threshold actually costs, so setting one becomes an informed decision
The meter runs inside your boundary and does not phone home. On Enterprise and Sovereign, usage telemetry is written to your own store and stays there — it is how you see your spend, not how we bill you. Licensing is a volume band declared at renewal, not a meter we read. How the licence works →
05 / prediction

Forecast the tail, not the average.

Document estates are bursty in ways that averages hide: quarter-end, renewal season, a large portfolio transfer, a regulatory deadline. An average forecast is comfortable and wrong exactly when it matters. Optim forecasts a distribution and provisions against the part of it that hurts.

Exhibit A· 8-week forecastyour estate · cloud + tokens

The agent forecasts a range, then acts on the tail.

A distribution tells you when the expensive version of next quarter becomes likely — the only moment intervention is still cheap. Quarter-end, renewal season and a large portfolio transfer are all in here.

Fan chart showing a P10 to P90 forecast band widening over eight weeks, with the P90 crossing the budget ceiling in week six. BUDGET CEILING $290K P90 BREACHES THE CEILING · W6 P90 $317k P50 $284k P10 $251k W1W4W8 SIGNAL ORANGE MEANS MONEY AT RISK — NOTHING ELSE
Realized spend feeds back every period, so accuracy compounds on your own usage. Attribution runs against the same numbers, so a moving P90 already has a named owner — a document family, a queue, a business unit. Figures shown are an illustrative run.
Method

P10 / P50 / P90, not a single line

Every forecast is a range with stated confidence. Capacity is provisioned against the upper band and cost is planned against the middle one, so a busy quarter is a plan rather than an incident.

Horizon

Next run, next month, next renewal

Short horizon drives autoscaling decisions during a run. Long horizon tells you which volume band you will be in at renewal — before the conversation, not during it.

Inputs

Seasonality your estate actually has

Learned from your own run history: day-of-month effects, quarter-end spikes, family-level growth and the drift in your lane mix. It is a model of your estate, not an industry benchmark.

Alerts

Told before, not after

Budget thresholds fire when the forecast crosses them, not when the spend does. A projection that breaches next month is something you can still act on.

Planning

What a new estate would cost

Model an unonboarded document family against your measured cost per page before committing to it. Useful for deciding which estate to bring on next, and for saying no to one.

Honest limit

A forecast is not a guarantee

Optim predicts from your history. A genuinely new document family or a step change in volume will be outside the distribution until it has run a few times. The bands widen when the model is unsure, and that widening is shown rather than smoothed away.

06 / where you actually see it

Four places, and none of them is a second console to learn.

01

The command centre

A cost panel beside the accuracy and throughput panels, broken down by document family and queue. Same screen, same login, same roles.

02

Your run report

Every run closes with a measured cost per page and the lane mix that produced it — including the free evaluation, so you can model your full estate before anyone quotes you.

03

The API

Usage, attribution and forecast series are readable through the same authenticated API as records and lineage, so this data lands in your own FinOps tooling rather than staying in ours.

04

Alerts you route

Forecast breaches and lane-mix drift fire as webhooks into whatever you already use. Optim does not ask to own your alerting.

07 / if you came here looking for Optim the product

Optim is a component of the platform, not a product line.

It was previously sold on its own as a cloud and AI spend platform. It has not been switched off and it has not been abandoned — it has been repositioned. The forecasting, attribution and routing engine now runs underneath Forge, where it does more good and where it is the reason the unit economics work.

Unchanged

Existing Optim agreements are honoured in full

If you hold a current Optim commitment: the same support, the same release cadence, the same named contact, for the term you contracted. Renewal will be an honest conversation about where the engine is going, not a migration deck.

Changed

Not sold standalone to new customers

Two crowded markets and two buyers cannot be served well by one small team, and we would rather say that plainly than sell you a second product we cannot give proper attention to. If cloud and AI spend governance is your actual problem, the FinOps category has good dedicated vendors and we will point you at them.

Questions

Ask directly

Existing Optim customer wondering what this means for your renewal? You will get a straight answer.

Talk to us about an existing agreement
The product / what Optim makes cheaper

The pipeline this engine runs underneath.

Forge reads your document estate into validated records with a page reference behind every field, deployed where your regulator requires, with the weights staying yours.