Home/Autonomy & gates
Deep dive / how much rope, and who holds it

An agent is not a chatbot wearing a finance skin.

It reads live state, plans a sequence it can defend, writes back through your APIs, then checks what changed. Autonomy is a dial your team turns — one notch at a time.

Reversible by design Evidence on every step Raised only by you
01 / the loop

Sense, plan, act, verify — and one place a person stands.

The gate sits on act, never on sense or plan. The agent may think as hard as it likes.

Exhibit A· the agent loopthe gate is on act, only
A four-step loop — sense, plan, act, verify — with a human approval gate placed only between plan and act. 01Sense 02Plan 03Act 04Verify THEN DO IT AGAIN HUMAN GATE
A copilot answers a question and stops. These close the loop — and everything interesting is in the fourth verb.

Sense

Primary records on your side of the boundary — billing exports, tag trees, token meters, IDocs, PO tables. No scraped dashboards.

Plan

Steps committed up front with the evidence behind each and the option it rejected. Below your bar, it halts and says why.

Act — gated

Executes through your APIs and your IAM, with a reversal path defined before it can fire.

02 / the rungs

Nothing arrives switched on at full autonomy.

Each class of action starts on the bottom rung and climbs only after your team has evaluated it on your own numbers.

The agent proposes; a person carries it out. Nothing executes. Every disagreement is logged — the cheapest possible place to discover the thing you were wrong about.

Reversible by designEvidence on every stepRaised only by you

Moving an action up a rung is your team’s call, made on your evidence. gothink.ai never widens an agent’s authority on its own initiative, in a release note, or by default.

03 / what “gated” means in practice

Reading is free. Changing state is not.

Runs freely

Anything that only observes

Detection, attribution, forecasting and drafting are continuous and unattended — a signal is worth little if it arrives a fortnight late.

Target, unmeasuredoverspend surfaced in < 1 min
Approvalnone required
Waits for a person

Anything that changes state

A commitment purchase, a posted entry, a lifted payment block. Proposed, evidenced, and held for a named human on your audit trail.

Who approvesa named person, inside your environment
Before it can firea reversal path exists
Guardrails

Budgets and blast radius

Caps on steps, spend, tokens and records touched per run. Allow-listed operations, rate limits on writes, and a kill switch your team owns.

Trace

Replayable, down to the step

Inputs, retrieved passages, tool calls, tokens, cost, latency and the approver’s name — held in your systems.

Evaluation

Scored before it ships

The ground-truth set becomes a regression suite that runs on every prompt edit, model upgrade and tool change.

The same idea, one layer down: the confidence gate on data

Actions have approval thresholds. Individual data fields have confidence thresholds — and they work the same way: a score on every field, a bar you set per field and per document type, and anything doubtful routed to a steward instead of into the database.

Raise the bar and more goes to review; lower it and more goes straight through. Either way the trace records which happened.

See the confidence gate

04 / the self-improving loop

A loop that gets better on its own — inside your guardrails.

Autonomy gates decide what an agent may do. Loop engineering decides what it may improve. A small maker-and-checker loop runs, scores itself against an objective test, records what worked, and stops when it is provably right — never because it feels right.

The maker

Proposes one change

Reads the goal and memory, makes a single small change, and is optimistic about its own work.

The checker

Reviews in a fresh context

A separate pass with different instructions catches what the maker talked itself into.

The gate

A scored test decides

The stop condition is a test or a score that can fail the work — not an opinion.

Memory

A leaderboard that survives

Every run is recorded and the lessons carry forward — so it starts smarter each session instead of from zero.

Knobs

Four hard stops

Max iterations, patience, max budget and max time bound every unattended run. Cross one and it halts.

The rule

Provably right, not confidently wrong

Nothing is accepted because the model said so — it is accepted because it passed the test.

The ask / a licence, proven on a scoped pilot

Turn the work you keep discovering by hand into a governed capability.

Three steps, one quarter, measured against your live numbers.

1

Score us on your documents

Twenty-five documents from your own estate, labelled by your own experts. Free, no contract, and you keep the report either way.

2

Three-week Fitting

Schema mapping, confidence gates and a ground-truth set, at a fixed price with a fixed end date. Exit at the pilot gate having paid the first milestone only.

3

Run software, not a project

Go live on the licence with the tuned weights handed over. Releases, packs and support carry on from there.