Agent or Workflow: How to Decide What Should Be Autonomous

Agent or Workflow: How to Decide What Should Be Autonomous

Most production systems described as AI agents should be workflows with one or two model calls inside them. Autonomy is a cost the task has to justify, not a feature to maximise. Every step a model decides for itself adds tokens, latency, and a new place to fail, and reliability compounds: a ten-step loop where each step succeeds 95 percent of the time finishes correctly only about 60 percent of the time. The first design decision in any agent project is therefore not which framework or which model. It is how much of the path the model should be allowed to choose, and most teams skip it.

The cost of skipping it is visible in the forecasts. Gartner predicted in June 2025 that over 40 percent of agentic AI projects will be canceled by the end of 2027, citing escalating costs, unclear business value, and inadequate risk controls, and added that many use cases positioned as agentic do not require agentic implementations. That last clause is the subject of this guide.

As of the week of September 28, 2026, capability is not the constraint it was: METR's task-length research shows the tasks frontier agents complete at 50 percent reliability doubling in length every several months. But a 50 percent horizon is a research metric, not a production standard. The gap between "can sometimes do it" and "does it every time, auditably, at an acceptable cost" is where the agent-or-workflow decision lives.

Reliability compounds

Key Takeaways

  • There are four shapes, not two: a deterministic workflow with a model step, a routed workflow, an orchestrator with workers, and an open agent loop. Each hands the model more of the path.
  • Choose the least autonomous shape that handles the task's real variation. Autonomy should be bought with evidence, not assumed.
  • Reliability compounds per step. Ten steps at 95 percent each gives roughly 60 percent end to end, and 99 percent each still fails about one run in ten.
  • Open loops cost more in three currencies at once: tokens, latency, and debugging time. Anthropic measured agents at about 4 times the tokens of a chat interaction and multi-agent systems at about 15 times.
  • The highest-value pattern in practice is a workflow that escalates only its ambiguous residue to an agent, inside step and spend caps.
  • Evaluate workflows at the path level and agents at the trajectory level. A single pass rate hides which part is broken.
Scoring rubric

The Vocabulary Teams Lack

"Agent" now covers everything from one prompt with retrieval to a system that plans, calls dozens of tools, and decides when it is finished. Treating those as one category makes the decision impossible.

The cleanest published boundary comes from Anthropic's December 2024 engineering guide, Building Effective Agents. Workflows are systems where models and tools are orchestrated through predefined code paths. Agents are systems where the model dynamically directs its own process and tool usage. The useful question is who decides the path: code, or the model.

Between those poles sit four shapes worth naming precisely.

Shape 1: Deterministic workflow with a model step

Code decides the path; the model fills a slot. The workflow is ordinary software: fetch the record, call the model to classify or extract or draft, validate the output against a schema, write the result. The model never chooses what happens next. Prompt chaining, where the output of one model call feeds the next in a fixed order, belongs here.

Example: an invoice arrives, the model extracts twelve fields into a typed schema, code checks that line items sum to the total, and anything that fails validation goes to a review queue.

Shape 2: Routed workflow

The model picks among fixed paths; code defines the paths. A classification step reads the input and selects one of several predefined branches, each of which is itself a deterministic workflow. The model has one decision, and the set of possible outcomes is enumerable in advance.

Example: a customer email is classified as refund request, address change, complaint, or other. Each category runs its own fixed handler, and "other" goes to a person.

Shape 3: Orchestrator with workers

The model plans; code bounds the plan. An orchestrating model decomposes a task into subtasks and dispatches them to workers, often other model calls with narrow tools, then synthesizes the results. The plan is dynamic, but the tools available, the number of workers, and the budget are fixed by code.

Example: a due-diligence request is split by the orchestrator into company filings, litigation search, and news review. Workers run in parallel with read-only tools, and the orchestrator merges their findings into a memo template.

Shape 4: Open agent loop

The model chooses tools and decides when it is done. The system runs a loop: the model observes, picks an action from its tool set, sees the result, and repeats until it judges the goal met. Neither the number of steps nor their order is known in advance.

Example: a coding agent fixing a failing test reads files, edits code, reruns tests, and stops when it believes the fix is complete.

Workflow escalation

The Shapes Compared

Shape Who decides the path Relative cost per task Reliability profile Auditability Best fit
Deterministic workflow with model step Code Lowest: one or a few model calls High and testable; failures are local to the model step Full: every path is in the code High-volume, known process with an unstructured input
Routed workflow Model picks, code defines options Low: one routing call plus the branch High if routing is accurate; misroutes are the main failure High: branch taken is logged and enumerable Varied inputs that fall into a known set of cases
Orchestrator with workers Model plans within code bounds Medium to high: many parallel calls Moderate; plan quality varies, but bounded Medium: the plan must be logged to be reviewable Decomposable research and analysis, breadth over depth
Open agent loop Model Highest and variable per run Lowest per task; compounds with loop length Lowest: requires trajectory tracing Open-ended tasks where the steps cannot be known in advance

The table reads left to right as a trade. Each row down buys flexibility with cost, variance, and opacity. The right answer is the highest row that still handles the task.

Decision Criteria and How to Score Them

Six criteria decide the shape. They are not independent, but scoring them separately forces the conversation most teams avoid.

Is the path knowable in advance?

This is the dominant criterion. If a competent analyst could write the steps down before seeing the input, the path is knowable and code should own it. OpenAI's A Practical Guide to Building Agents makes the same cut from the other side: it reserves agents for complex judgment-based decisions, rulesets that have grown unmaintainable, and heavy reliance on unstructured data, and says plainly that where none of those apply, a deterministic solution may suffice.

Note the trap in the third condition. Unstructured data justifies a model; it does not by itself justify an agent. Reading a messy document is a Shape 1 problem.

How many steps?

More steps means more compounding. A task that genuinely needs fifteen sequential judgments pushes toward an agent or, better, a redesign that moves most of those judgments into code.

Cost of a wrong step, and is it reversible?

A wrong draft that a person reviews costs seconds. A wrong payment, a deleted record, or an email sent to a customer costs real money and cannot be undone. Irreversible actions pull hard toward workflows, or at minimum toward human approval on the specific irreversible steps. The accountability consequences of getting this wrong are covered in the agent liability gap.

Latency budget

Every loop iteration is a model round trip. Open loops have long, unpredictable latency tails, which rules them out of most synchronous customer-facing paths and leaves them comfortable in batch.

Need for auditability

Regulated decisions such as credit and claims need an explanation of why the system acted. A workflow's explanation is its code plus each model output. An agent's is a trajectory that must be captured, stored, and interpreted, which is possible but expensive, as the guide to agent observability in production sets out.

How often does the task vary?

If ninety percent of inputs follow one pattern, build for it and escalate the rest. If every input is genuinely different, flexibility earns its cost.

The Scoring Rubric

Score each criterion from 0 to 2, where 0 favours a workflow and 2 favours autonomy. Sum the scores.

Criterion 0 (favours workflow) 1 (middle) 2 (favours autonomy)
Path knowable in advance Steps can be written down before seeing input Known set of paths, input picks one Steps depend on what is discovered along the way
Number of model decisions 1 to 2 3 to 6 7 or more, and cannot be reduced
Cost and reversibility of a wrong step Irreversible or expensive Reversible with effort Cheap and easily undone
Latency budget Under 5 seconds, user waiting Under a minute Minutes to hours, asynchronous
Auditability requirement Regulated or externally reviewable Internal review expected Low; outcome matters, not path
Task variation Most inputs follow one pattern A handful of recurring patterns Nearly every input is different

Reading the total:

Total score Recommended shape
0 to 3 Deterministic workflow with a model step
4 to 6 Routed workflow
7 to 9 Orchestrator with workers, bounded by caps
10 to 12 Open agent loop, with step caps, spend caps, and approval on irreversible actions

Two override rules apply regardless of total. Any irreversible action with material cost gets human approval or a workflow, even at a score of 12. And a score of 0 on "path knowable" should cap the answer at Shape 2, because a knowable path handed to a model is orchestration that belongs in code.

Apply the rubric to each sub-task, not just the whole system. Most real processes are a mix: a knowable spine with one or two genuinely open sub-problems.

The Cost Model of Autonomy

Teams underestimate autonomy because they price it at the per-call rate. The real cost arrives in three currencies at once.

Tokens and latency grow with loop length

An open loop re-reads its growing context on every step: the goal, the tool definitions, and every prior observation. Token spend per task therefore grows faster than linearly with step count unless context is actively managed. Anthropic's June 2025 write-up of its multi-agent research system measured agents at about 4 times the tokens of a chat interaction and multi-agent systems at about 15 times, and noted that the multi-agent design was worth it only for tasks valuable enough to pay for that. The same system outperformed a single agent by 90.2 percent on its internal research evaluation, so the cost can buy real capability. The point is that it has to be bought deliberately.

Reliability compounds per step

If each step succeeds independently with probability p, a run of n steps succeeds with probability p to the power n.

Per-step success rate 3 steps 5 steps 10 steps 20 steps
90 percent 73 percent 59 percent 35 percent 12 percent
95 percent 86 percent 77 percent 60 percent 36 percent
99 percent 97 percent 95 percent 90 percent 82 percent

Real steps are not fully independent and agents sometimes recover, so this simplifies. The empirical evidence still agrees. The tau-bench paper from Sierra introduced a pass-to-the-k metric, the probability an agent succeeds on all of k repeated attempts at the same task. GPT-4o in the retail domain scored about 61 percent on a single attempt and fell to roughly 25 percent when it had to succeed on eight attempts in a row. Models have improved since, but the shape of the curve, where consistency lags capability, is the lesson. Production traffic is thousands of repeated attempts.

Debugging an open loop is far harder than debugging a graph

When a workflow fails, the failing node is identifiable and the fix is local. When an agent fails, the question is which of thirty decisions went wrong, and why. Many agent failures trace back to the tool surface rather than the model, as covered in the tool design guide for agents, and a catalogue of the recurring breakdowns is in the analysis of agent framework failure modes in production. Diagnosis hours are a real, usually unbudgeted cost of autonomy.

Patterns That Capture Most of the Value at Lower Risk

The choice is rarely binary. Three patterns deliver most of what teams want from agents with much of the risk removed.

Workflow with agent escalation for the ambiguous residue

Build the deterministic path for the common case and route only what it cannot handle to an agent or a person. The workflow handles the bulk of inputs cheaply and auditably; the agent handles the long tail, where flexibility is needed and volume is small enough to keep cost and review manageable.

Every escalated case is also an example of variation the fixed path missed. When a category becomes common, it gets its own branch.

Bounded loops with step and spend caps

When an open loop is justified, bound it. A maximum step count stops runaway loops. A per-task token or spend ceiling caps the cost of the pathological run. A wall-clock timeout protects latency. When any cap is hit, the task escalates rather than failing silently. Caps turn an open loop's unbounded cost distribution into a bounded one, which is the difference between a system finance can budget and one it cannot.

Human approval keyed to reversibility

Approval gates fail when they fire on everything, because reviewers learn to click through. Classify every action the system can take as read-only, reversible write, or irreversible write, and gate only the third class. A refund above a threshold, an external email, a record deletion, or a payment gets a person. Reading data and drafting do not. For the permission and isolation layer underneath this, see the guide to agent sandboxing.

When to Graduate a Workflow, and When to Demote an Agent

The decision is not permanent. Shapes should move in both directions as evidence arrives.

Signals that a workflow should graduate toward autonomy:

  • The branch count keeps growing, and each new branch handles a small slice of traffic. The ruleset is becoming unmaintainable, which is one of the OpenAI guide's three conditions.
  • The escalation queue is large and diverse, with no dominant category to turn into a new branch.
  • Reviewers of escalated cases are mostly doing multi-step investigation, not quick judgment calls.

Signals that an agent should be demoted to a workflow:

  • Trajectory logs show the agent taking the same sequence of steps on most runs. A path it always takes is a path code should own.
  • Failures cluster on a few specific steps that could be made deterministic.
  • An auditor, regulator, or customer has asked why a decision was made, and the answer took days to reconstruct.

The first demotion signal is the most common. Agents are excellent at discovering a process; once discovered, it usually belongs in code, with the agent kept for cases that still deviate.

How to Evaluate Each Shape

Evaluation strategy follows shape. Using the wrong one produces numbers that look fine and predict nothing.

Workflows get path-level tests. Because the path is fixed, each model step can be tested in isolation against a labelled set: did the classifier route correctly, did the extractor fill the right fields, did the drafter follow the template. End-to-end tests then confirm the steps compose. Failures point to a specific node. The executive framing for building this capability is in the guide to AI evals.

Agents get trajectory evaluation. The final answer is not enough, because an agent can reach a right answer by a dangerous route or a wrong answer by a reasonable one. Trajectory evaluation checks the sequence: did it call appropriate tools, in a sensible order, with valid arguments, without unnecessary steps, and did it stop at the right time. Two metrics matter most: task success measured across repeated runs, in the spirit of tau-bench's pass-to-the-k rather than a single pass, and cost per successful task including the failed runs.

Worked Example: Commercial Card Dispute Handling

Take a hypothetical mid-sized bank processing commercial card disputes. The team was asked to "build an agent" to handle them end to end. Applying the decision properly produces a different system.

Step 1: decompose the process. A dispute arrives by email or portal. It must be classified by reason (fraud, duplicate charge, goods not received, amount mismatch), the transaction must be matched in the ledger, evidence must be gathered, a provisional credit decision must be made, and the merchant-side chargeback must be filed with the network under its rules and deadlines.

Step 2: score each sub-task.

Sub-task Path knowable Decisions Reversibility Latency Audit Variation Total Shape
Intake and classification 1 0 2 1 0 1 5 Routed workflow
Transaction matching 0 0 2 1 0 0 3 Deterministic workflow with model step
Evidence gathering for unclear cases 2 2 2 2 1 2 11 Bounded agent loop, read-only tools
Provisional credit decision 0 0 0 1 0 0 1 Deterministic rules, human approval above threshold
Chargeback filing 0 0 0 2 0 0 2 Deterministic workflow with model drafting

Step 3: assemble the system. The result is a workflow spine with one bounded agent. Classification routes each dispute to a branch. Matching uses a model to extract the transaction reference and amount, with code doing the ledger lookup. Most disputes, such as clear duplicates and exact amount mismatches, never touch an agent. The minority with ambiguous evidence go to an agent that can search the ledger, read correspondence, and query the merchant record, but cannot write anything, is capped at fifteen steps, and hands a structured evidence summary back to the workflow. Provisional credit is rules plus a human above a threshold, because it moves money. Filing is deterministic, because network rules and deadlines are knowable and a missed deadline is irreversible.

The team still has an agent, running on the fraction of disputes that need it, with no write access and a hard cap, inside a system whose every money-moving step is deterministic. That is what most successful production "agents" actually look like. For executives framing the broader programme, the executive guide to agentic AI covers the strategic layer this design sits under.

A Checklist a Team Can Apply This Week

  1. Score each sub-task, not the process as a whole, on the six criteria and assign a shape.
  2. Apply the two overrides: approval or workflow for irreversible material actions, no autonomy for knowable paths.
  3. Classify every action as read-only, reversible write, or irreversible write.
  4. Cap every loop on steps, spend, and wall-clock time, escalating on breach.
  5. Write down the graduation and demotion signals before launch, so the shape moves on evidence rather than opinion.

Frequently Asked Questions

What is the difference between an AI agent and an AI workflow?

In a workflow, code decides the sequence of steps and the model performs specific tasks inside it, such as classifying, extracting, or drafting. In an agent, the model decides which tools to use, in what order, and when the task is finished. The practical distinction is who controls the path: code or the model.

When should you use an AI agent instead of a workflow?

Use an agent when the steps genuinely cannot be known in advance, the task varies widely from input to input, wrong steps are cheap and reversible, and the latency budget allows multiple model round trips. If the path can be written down before seeing the input, a workflow with model calls inside it will usually be cheaper, faster, more reliable, and easier to audit.

Why do AI agents fail more often than workflows?

Reliability compounds across steps. If each step in a ten-step loop succeeds 95 percent of the time, the whole run succeeds only about 60 percent of the time. Open loops also accumulate context, which raises cost and can degrade later decisions, and their failures are harder to diagnose because the error can sit in any of many model-chosen steps.

Is an agentic workflow the same as an agent?

Not quite. "Agentic workflow" usually describes a system where models make some decisions, such as routing or planning, inside a structure code still bounds. That covers the routed-workflow and orchestrator shapes. A fully open agent loop, where the model chooses every action and decides when it is done, is the far end of the same spectrum.

How do you control the cost of an AI agent in production?

Bound the loop with a maximum step count, a per-task spend ceiling, and a timeout, and escalate when any is reached. Route only the ambiguous minority of cases to the agent and keep the common path deterministic. Measure cost per successful task, including failed runs, rather than cost per call.