One Agent Run, Start to Stop

From a working map to a checked result

A luminous thread moves from intention to action, strikes an independent evidence surface, and returns revised.A luminous thread moves from intention to action, strikes an independent evidence surface, and returns revised.

Figure 1. Reflection becomes useful when it leaves the echo chamber, meets resistance, and returns changed. Original illustration.

Part II ended with a practical claim: agency is not a personality added to a chatbot. It is a path through state, action, observation, revision, and stopping.

This interlude shows that path once, in full. Build a Factory Loop carries it into a repeatable engineering process.

The problem below is small: a result that a source, calculation, executable test, physical measurement, or other person could reject. What matters is not the answer but what survives around it: what the system believed, what it did, what the world returned, what changed, and why it stopped.

That is the difference between a conversation and a run.

Two records, not one

The state card at the end of Part II was enough to begin but not enough to audit the result. It named the goal, evidence, open question, next action, and check. It did not show the candidate the agent actually produced, the observation a source or test returned, the discrepancy between them, the exact revision, the verdict and stop reason, or the uncertainty left behind.

So a run keeps two related records. The state card is the agent’s small, mutable working map, short enough to read before every action. The scorecard is the human-readable history of the run: the beginning, the evidence, the revision, and the end. Failed checks are appended to it, never overwritten.

Three roles, not necessarily three agents

A run needs three functions. The worker proposes and produces the next bounded change. The checker compares the candidate with a prewritten condition and evidence that can disagree. The owner defines the goal, approves consequential actions, resolves value choices, and can stop the run.

None of this requires an agent framework or three models. The worker can be one chat; the checker a test, calculator, official page, or fresh prompt that receives the candidate and rubric but not the worker’s rationalization; the owner a person. In a real deployment, owner may name a conflict: user, operator, employer, vendor, regulator, and affected people can hold different claims, so the run records who has authority, which interests that authority does not represent, and how an affected person can challenge the result.

A fresh model context is only a separate check, not independent evidence. Factual claims need an authoritative source; arithmetic needs recomputation; code needs executed tests; usability needs the person who will use the result.

The run in six moves

The owner defines the territory first: the user, the desired result, done conditions that can be observed, exclusions, and the approval boundary. “Helpful” and “complete” are not conditions; “the page opens offline” is. Before asking for an answer, the owner writes at least three test cases—an ordinary case, an edge, and a case with missing or conflicting information—with their expected results, so the model cannot invent the grading rule after seeing its own output. The worker then produces one complete candidate through the smallest useful actions, and each returned observation is saved unpolished before the next action is chosen. Reflection touches the world: the candidate runs against the prewritten cases through a check that can return FAIL, UNSUPPORTED, or UNKNOWN. The largest discrepancy gets one revision of the smallest component likely to have caused it, then the same cases rerun; after two revision rounds the run stops regardless. Finally the owner chooses one status—PASS, BLOCKED, HUMAN DECISION, or TIMEBOX—and records the stop reason, the residual uncertainty, the decision that stayed human, and a three-sentence postmortem: what the first map omitted, what evidence changed the work, and what remains outside the system’s authority.

Without that last move, the next person sees a fluent result without knowing whether it passed, merely stopped, or quietly ran out of evidence.

A small agent loop with visible state, outside evidence, an owner, and explicit stop conditions.A small agent loop with visible state, outside evidence, an owner, and explicit stop conditions.

Figure 2. The working loop is small. The scorecard preserves what happened around it.

The run in full: a library-card errand in one visit

This ordinary example came from a sub-agent review of grounded, low-stakes tasks. It uses the official Toronto Public Library card overview, identification requirements, and Toronto Reference Library page. The facts below are an illustrative snapshot checked on August 12, 2026; a real reader must use their own municipality and recheck volatile facts.

TASK
Run / date:
Illustrative planning run, August 12, 2026.

User and problem:
An adult Toronto resident wants a full library card for physical and digital
borrowing without making a second trip.

Goal / desired state:
Produce a short checklist naming eligibility, documents, branch, ordinary
hours, and the live check required before leaving home.

Done means:
1. Eligibility is supported by the official library page.
2. Each document has a stated purpose: name, address, or both.
3. The selected branch and ordinary hours fit the available window.
4. Time-sensitive facts and the library's final authority remain visible.

Out of scope:
Applying online, uploading identification, guaranteeing acceptance, choosing
transportation, or storing document numbers or scans.

CONTROL
Human owner:
The reader making the trip.

Approval required before:
Sharing any personal information or presenting documents to library staff.

Allowed tools and actions:
Read official public pages and write a private checklist. No form submission,
account creation, upload, message, or booking.

STARTING STATE
Known evidence:
The reader reports living in Toronto, holds a valid Canadian passport and a
Toronto Hydro bill dated July 20, 2026, is free Saturday August 15 from
10:00 a.m. to noon, and prefers Toronto Reference Library.

Open questions:
Does the passport suffice by itself? Is the branch ordinarily open then?

Initial candidate or baseline:
"Bring the passport to Toronto Reference Library on Saturday morning."

Next smallest action:
Read the official eligibility and identification pages before checking hours.

ROUND LOG
Action taken:
Opened the three official pages and recorded the relevant statements with the
August 12, 2026 access date.

Observation returned:
Toronto residents qualify for a free card. The passport appears in the name-ID
list; a current bill appears in the address-ID list. Bills must have been issued
within the preceding two months. The branch address is 789 Yonge Street and its
ordinary Saturday hours are 9:00 a.m.–5:00 p.m.; a service notice did not state
that the branch was closed.

Check method:
A separate pass ignored the draft conclusion, reopened each official page, and
labeled every checklist line SUPPORTED, UNSUPPORTED, or TIME-SENSITIVE.

Expected result:
Every stable line supported; same-day branch status explicitly time-sensitive.

Actual result:
The passport-only instruction was UNSUPPORTED. Eligibility, the two-document
combination, address, and ordinary hours were SUPPORTED. Guaranteed opening and
guaranteed registration were UNSUPPORTED.

Discrepancy:
The initial candidate treated a name document as though it also proved address
and turned ordinary hours into a guarantee.

Revision made:
1. Bring the valid passport as name identification.
2. Bring the July 20 Toronto Hydro bill as current address identification,
   assuming it displays the reader's current address.
3. Go to 789 Yonge Street between 10:00 a.m. and noon Saturday.
4. Reopen the official branch page that morning for an unscheduled closure.
5. Let library staff make the real eligibility and document decision.

Rerun:
All five lines were checked again. Four stable claims were supported; the live
status was correctly marked TIME-SENSITIVE rather than silently passed.

CLOSE
Final verdict:
PASS. The planning checklist passed; one live check remains visible.

Stop reason:
All stable completion conditions passed. Another rewrite cannot settle a future
closure; that uncertainty has an owner, source, and time for resolution.

Residual uncertainty:
The utility bill must actually display the current address, branch status can
change, and staff retain the final decision.

Human-owned decision:
The reader decides whether the trip is worthwhile, which documents to carry,
and whether to proceed if the live status changes. Staff decide acceptance.

Handoff / next owner:
The reader takes ownership of the checklist and reopens the branch page on the
morning of August 15 before leaving home.

Three-sentence postmortem:
The first map collapsed name proof and address proof into one plausible object.
The official identification table corrected the checklist, while the branch
page exposed a fact that must be checked later. The agent can prepare the errand;
it cannot guarantee the building or institution will accept the plan.

Two more runs, in brief

Both are illustrative runs written to test the same discipline at a different level, not reports of real deployments.

A study-load page. A Grade 10 student wanted one offline page that totals tonight’s assignment estimates and reports FITS or OVER. Four browser cases were written before any code. Version 0 failed two of them: it used a strict less-than, so an exactly full evening reported OVER, and it turned a blank field into a valid zero. Two revisions, one per cause, fixed both; a ninth assignment was rejected as specified. Verdict: PASS, as a classroom prototype. What stayed human: the estimates themselves, and what to do when the page says OVER. What the first map omitted: equality, and the difference between missing data and real data.

A workflow supervisor. The third run changes level. The artifact does not solve one user’s problem; it routes small jobs to a calculation, research, or drafting workflow, tracks desired and current state, checks results, and returns PASS, REVISE, BLOCKED, or HUMAN DECISION with a trace. It borrows the control-loop idea from Kubernetes controllers and the plan-before-apply idea from Terraform without managing any infrastructure. Five prewritten traces exposed three failures in version 0: a “draft and send” job called a simulated send before approval; a comparison routed to research because “compare” matched first; and an off-by-one let a third round begin. Deterministic rules before routing, a permission gate before dispatch, and a completed-round counter fixed them. Verdict: PASS as a sandboxed prototype. Concurrency, credentials, hostile input, and cost were excluded rather than mistaken for completed work.

The systemic pattern is compact:

read desired state                       five traces must route and exit as expected
observe current state                    which trace failed, and at which step
choose one allowed workflow              calculate, research, draft, or pause
pause before external change             “send” waits for the owner
run one step                             the routed job executes once
check the observation                    trace assertion: route, tool calls, rounds, exit
update state                             completed rounds increment; verdict recorded
stop, escalate, or reconcile once more   two rounds, then TIMEBOX

Kubernetes, Terraform, programming-language runtimes, schedulers, and workflow engines are vastly more sophisticated. What matters here is the change of level: the supervisor does not need to know how every task is solved. It needs to know which system may act, what state must remain visible, what policy constrains the action, how the result is checked, and when control returns to a person.

What the three runs reveal

The app fails in code. The library plan fails against a rule. The supervisor fails in the space between systems.

The checks therefore differ:

LevelCandidateReality that can disagreeHuman retains
One appLocal HTML behaviorExecuted cases and arithmeticMeaning of the result
One errandA document checklistOfficial rules and live branch stateTrip and document decision
System of systemsRoutes, policies, and tracesDeterministic workflow assertionsAuthority and policy

The loop remains recognizable at every scale: desired state, current map, one action, returned evidence, a checked difference, and an explicit exit.

Before moving to Part III

Build a Factory Loop follows one software feature through that process. Take two things from this page. The checker must have access to something the worker cannot invent. And the owner’s decisions must be named before the work starts, not discovered when the run tries to cross them. Part III then turns to the person: acting, noticing what the action brings up, and choosing again.


References and further reading

  1. Shunyu Yao et al., “ReAct: Synergizing Reasoning and Acting in Language Models”, ICLR, 2023.

  2. Jie Huang et al., “Large Language Models Cannot Self-Correct Reasoning Yet”, ICLR, 2024.

  3. Zhibin Gou et al., “CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing”, ICLR, 2024.

  4. Santiago Díaz, Christoph Kern, and Kara Olive, “Google’s Approach for Secure AI Agents”, Google, 2025.

  5. Google Agent Development Kit, “Loop workflow”, on explicit termination and maximum iterations.

  6. Kubernetes Documentation, “Controllers”, on control loops and desired versus current state.

  7. HashiCorp, “terraform plan command”, on comparing current state with configuration and reviewing proposed actions before apply.

  8. Hangfei Lin, “Architecting Efficient Context-Aware Multi-Agent Framework for Production”, Google Developers Blog, December 2025.

Next articleBuild a Factory Loop