It reads live state, plans a sequence it can defend, writes back through your APIs, then checks what changed. Autonomy is a dial your team turns — one notch at a time.
The gate sits on act, never on sense or plan. The agent may think as hard as it likes.
Primary records on your side of the boundary — billing exports, tag trees, token meters, IDocs, PO tables. No scraped dashboards.
Steps committed up front with the evidence behind each and the option it rejected. Below your bar, it halts and says why.
Executes through your APIs and your IAM, with a reversal path defined before it can fire.
Each class of action starts on the bottom rung and climbs only after your team has evaluated it on your own numbers.
The agent proposes; a person carries it out. Nothing executes. Every disagreement is logged — the cheapest possible place to discover the thing you were wrong about.
Staged, quantified, parked with a named approver. They see exactly what will move and by how much, sign once, and that signature stays stapled to the run permanently.
Unattended inside your thresholds. Cross one and the step parks itself and waits. Either way the run stays reversible and on the record.
Moving an action up a rung is your team’s call, made on your evidence. gothink.ai never widens an agent’s authority on its own initiative, in a release note, or by default.
Detection, attribution, forecasting and drafting are continuous and unattended — a signal is worth little if it arrives a fortnight late.
A commitment purchase, a posted entry, a lifted payment block. Proposed, evidenced, and held for a named human on your audit trail.
Caps on steps, spend, tokens and records touched per run. Allow-listed operations, rate limits on writes, and a kill switch your team owns.
Inputs, retrieved passages, tool calls, tokens, cost, latency and the approver’s name — held in your systems.
The ground-truth set becomes a regression suite that runs on every prompt edit, model upgrade and tool change.
Actions have approval thresholds. Individual data fields have confidence thresholds — and they work the same way: a score on every field, a bar you set per field and per document type, and anything doubtful routed to a steward instead of into the database.
Raise the bar and more goes to review; lower it and more goes straight through. Either way the trace records which happened.
Autonomy gates decide what an agent may do. Loop engineering decides what it may improve. A small maker-and-checker loop runs, scores itself against an objective test, records what worked, and stops when it is provably right — never because it feels right.
Reads the goal and memory, makes a single small change, and is optimistic about its own work.
A separate pass with different instructions catches what the maker talked itself into.
The stop condition is a test or a score that can fail the work — not an opinion.
Every run is recorded and the lessons carry forward — so it starts smarter each session instead of from zero.
Max iterations, patience, max budget and max time bound every unattended run. Cross one and it halts.
Nothing is accepted because the model said so — it is accepted because it passed the test.
Three steps, one quarter, measured against your live numbers.
Twenty-five documents from your own estate, labelled by your own experts. Free, no contract, and you keep the report either way.
Schema mapping, confidence gates and a ground-truth set, at a fixed price with a fixed end date. Exit at the pilot gate having paid the first milestone only.
Go live on the licence with the tuned weights handed over. Releases, packs and support carry on from there.