Which step needs an agent?

One practitioner thread collected for this July article made its position explicit: I charge clients more to NOT build an AI agent. The author described replacing a proposed inventory agent with a workflow using existing reorder thresholds. As a contrast, they described tenant messages that needed interpretation before work could be assigned to vendors.

The inventory example makes the case for keeping known decisions in code. The tenant example raises a further question: needing a model to interpret a message does not establish that it should choose the subsequent process. The thread is a practitioner’s account, with no independent evaluation of the reported results. Its examples are useful for examining where the decision lies.

Consider a hypothetical support team receiving equipment fault reports as emails with attachments. The team wants software to identify the affected product, extract the reported symptoms, and prepare a case for an engineer. The language varies, attachments arrive in different formats, and important details may be missing. A language model could help interpret that material.

The process around the interpretation may still be known. Read the submission, extract the relevant fields, check them against the asset record, and send a draft case for review. Difficult input does not automatically require the model to decide how the whole process runs. The design question is which decisions need to remain open until the system sees the evidence.

Keep the known process explicit

Anthropic’s Building effective agents, published in December 2024, distinguishes predefined execution paths from systems in which a model chooses its process and tool use as it works. Its guidance recommends starting with a simple design and adding complexity when task performance justifies it. This supplies a distinction the thread’s contrast leaves open: the fault-report workflow can contain model calls without giving the model authority over every transition.

In the example, an extraction step could return a product identifier, symptoms, and references to the passages supporting them. Application code would check the proposed identifier against the asset service. If it could not find a match, the workflow would send the case to a reviewer. A known failure condition has a defined destination.

This structure makes some behavior easier to test. A missing product identifier should always take the review path, regardless of how confidently the model describes the fault. The extraction itself remains probabilistic: it may omit a symptom or misread an attachment. Fixing the sequence does not remove the need to evaluate those outputs.

It also makes responsibility easier to locate. An incorrect identifier points to extraction or matching. A case created despite a failed check points to control logic. Those distinctions help a team decide whether to improve the prompt, repair the integration, or change the process. The book’s control patterns describe ways to make those transitions explicit.

Give investigation room to change direction

Now change the task. The engineer wants the system to investigate why a machine repeatedly shuts down. The useful next step depends on what it discovers. A fault code might justify reading a maintenance bulletin; a recent component replacement might justify comparing service records. A fixed extraction pipeline would prepare the evidence but leave most of the investigation to the engineer.

This is a stronger case for an agent. The proposed capability is choosing which evidence to gather next and revising the investigation when a result contradicts its working explanation. That flexibility has a cost: the system may pursue an irrelevant lead, repeat searches, or produce a plausible diagnosis from incomplete evidence. Those are behaviors the evaluation needs to expose.

A fixed workflow can still surround the investigation. The application can identify the machine and establish the permitted data sources before starting the agent. The agent can investigate within those sources and return a draft diagnosis with its evidence. Creating a repair order can remain a separate operation with its own checks and approval.

The practitioner’s inventory example has a stable decision rule. An investigation may not. If almost every fault report requires the same two lookups, code can perform those lookups directly. If the next useful source regularly depends on an earlier finding, forcing every case through an expanding set of branches may become difficult to maintain. The comparison is between the actual implementations and their results on representative cases.

Compare the result against a workflow

For this example, I would build the extraction workflow first and use its unresolved cases to identify what an investigation agent might add. Those cases are useful for development, but the comparison also needs ordinary reports. A system that helps with unusual faults may spend too long investigating routine ones.

The test set should attach evidence to each case. A known historical fault can include the confirmed diagnosis and the records that support it. An incomplete submission can specify that the correct outcome is a request for missing information. Where several investigative paths are reasonable, the evaluation can assess whether the conclusion is supported without requiring one particular sequence of tool calls. This is the role of evaluation ground truth.

The reviewer’s work also needs to enter the comparison. In a study submitted June 3, Shipi Dhanorkar, Samir Passi, and Mihaela Vorvoreanu interviewed 17 experienced developers about their oversight of software agents. The work extended from setting constraints and planning through monitoring execution and reviewing results. Reviewing generated code could itself be difficult. The study is exploratory and concerns software development, so it cannot establish the value of a support agent; its relevance here is the human effort that execution-time measurements would leave out.

Use the same cases and data access for both designs. For each fault report, assess whether the draft identifies the fault correctly or gives the engineer enough evidence to proceed, and record the time spent checking and correcting it. Latency and tool cost belong alongside that review effort. A polished answer may offer little benefit if verifying it takes longer than preparing the case manually.

Inspect the failures before attributing an improvement to autonomy. If the agent succeeds because it has access to a maintenance database the workflow lacks, the experiment has changed both access and control. Give the workflow the relevant retrieval step and compare again. The decision to delegate more choices should rest on the value of those choices.

Limit the decision you delegate

A successful investigation trial supports giving the agent an investigative role. It does not establish that the same agent should schedule repairs or contact customers. Each additional action changes the consequences of an error and needs its own justification.

For the fault-report system, the initial deployment could let the agent read approved records and prepare a diagnosis, with a limit on investigation time and an explicit unresolved outcome. An engineer would review the evidence before authorizing a repair. The June study makes it worth examining that review as work in its own right: the interface must make supporting records accessible, and the trial should measure how much checking remains. If the agent exhausted its allowance, the application would preserve the work and route the case for review.

The practitioner thread argues that clients can benefit from being talked out of an agent. For the fault-report system, that judgment should turn on a narrower claim: does letting the model choose follow-up searches improve case preparation after the engineer’s review effort is counted? If it does, keep that freedom. If a fixed set of lookups performs as well, use it and reserve investigation for the cases that need it.

← Field notes