Skip to main content
Clearlead AI Consulting
All articles
Paul Ferguson

Agents Aren't Chatbots With Extra Steps

A navy speech bubble opens like a doorway to reveal a shopping cart, calendar and magnifying glass against a purple background.

If you are finding it difficult to pin down what an "AI agent" actually is, some of that confusion is understandable. Two systems can look much the same in a chat window, even if one follows steps written in advance and the other chooses what to do next based on what it finds.

There is also the confusion added by marketing departments, who want to put the term "agentic" on any system that calls a separate tool, just because it makes them sound more advanced.

But there are some significant differences between these systems, and I think that understanding them helps explain what we can delegate to them, what they need to work reliably, and when a simpler approach might be more suitable.

From an order update to finding a replacement

Suppose you have ordered something you need by Friday. You open the retailer's chat window and ask, "Where is order #123?" The AI assistant looks up the order in the retailer's system and replies that it is due to arrive the following Monday.

It has answered your question, but you still need to decide what to do about the late delivery. You are now left with some follow-up work:

  • Search for a replacement yourself
  • Compare the options
  • Check whether any would arrive in time

In this interaction, the assistant provides information and leaves that follow-up work to you.

An order-status question leads to an order lookup and the reply that it will arrive on Monday. The customer needs it by Friday and still needs to find a replacement.
Figure 1: The assistant answers the question; finding a replacement is still left to the customer.

You could also ask the assistant to help find a replacement. It would then need to search for suitable products, compare them with what you ordered and check delivery dates. That is a larger task than looking up an order's status, and there is more than one way to build a system that can do it.

A predefined workflow or an agent?

Two common approaches are to use a predefined workflow and an agent. Both could help find that replacement, using the same AI model and access to the retailer's catalogue and delivery information.

There isn't one universally agreed definition of an agent. I find Anthropic's explanation useful because it focuses on how the work is organised: how much of the procedure is specified in advance, and how much is left for the model to decide as it goes.

  • In a predefined workflow, developers specify the procedure, including what happens when a check passes, fails or needs repeating.
  • With an agent, the model chooses its next action, sees the result and decides how to continue towards the goal. Developers still determine which tools it can use and the limits it must work within.

By tools, I mean functions the software makes available, such as searching a catalogue or checking a delivery date. Both designs can use them.

For our late-delivery example, the request would be something like:

"My order will arrive too late. Find an alternative that arrives by Friday, stay within my budget, and ask before spending."

A predefined workflow could search the catalogue, check each candidate's details and delivery date, and offer the first option that meets the requirements. If a candidate fails, it checks the next one; if none works, it hands the case to a person. Those instructions are written in a predefined manner (in software) before the request arrives.

A predefined workflow searches for alternatives and checks whether any candidates remain. An empty search or exhausted list goes to a person. Otherwise it assesses suitability, price and delivery, offers a qualifying product for approval, or loops to the next candidate. AI can assess suitability within the predefined procedure.
Figure 2: AI can assess a product within the workflow; the procedure determines what happens next.

A workflow can still use AI to make judgements within that procedure. For example, a model could assess whether a product is a suitable replacement. The software would then use that assessment to follow the next predefined step: check its delivery date if it is suitable, or move to the next product if it is not. These rules can handle many different orders.

An agent could use the same catalogue, product details and delivery checks, but the model chooses how to investigate. For example, if the first replacement arrives too late, it can search again for products available sooner. And if a product's suitability is unclear, it could read the specification, discover a missing feature and ask whether the customer would accept that compromise.

That is just one possible run. What it learns about each option informs its next decision. In another case, it might have enough information to propose the first product straight away. But we do need to bear in mind that with this flexibility, it is possible for things to go wrong (for example, it may choose a poor replacement product), which is why this needs rigorous testing.

It is important to say that we could also write a workflow covering those same situations. The design choice is whether to specify that investigation procedure ourselves or let the model choose within the limits we set. Repeating a step, often called a loop, does not by itself make a system agentic. Both designs can repeat steps; the difference is how the next step is selected.

The designs can also be combined: an agent could choose which products to investigate, then use a predefined workflow to place the customer's approved order. The workflow would handle the purchase according to the retailer's rules. Either design could sit behind the same chat window, and in reality in a lot of complex systems they will use a mix of purely agentic and predefined workflows.

What keeps the agent running?

When we say the model performs a check (for example, checks the delivery date), it is actually requesting that a piece of software perform that check. Something has to receive the request, run the tool and pass its result back to the model. The software that manages this cycle is commonly called an agent harness.

The user request enters the harness, which supplies the task and results to the model. The model returns a next action or reply. The harness checks permissions and calls tools, receives results, tracks progress and applies run limits. Replies and approval requests return to the user through the harness.
Figure 3: The harness connects the user, model and tools, and manages the cycle of requests and results.

The harness prepares the task information and calls the model. If the model requests a delivery check, the harness runs the permitted tool and supplies the result alongside the information needed to continue, such as the Friday deadline and the options already tried. The model can then choose another search or compose a reply, which the application passes back to the customer. This cycle lets the system continue without a new prompt from the user after every step.

In addition to running tools, the harness needs to keep track of progress and pause for required approval. When a tool fails, it must return enough information for the system to retry, choose another approach or hand the case to a person. Longer tasks may need saved notes or summaries so the model can pick up where it left off.

An agent may keep searching, repeat a failed check or try another approach without getting closer to completing the task. Each attempt can take more time and incur further model or tool costs. The harness therefore needs to enforce limits on the number of actions, retries, elapsed time or cost of a run. When a limit is reached, the system should stop or ask for help, with a record of what it has tried.

A predefined workflow needs software to run its steps and enforce limits too, but with an agent, the model chooses which step to try next, while the harness checks whether that action is allowed, runs it and returns the result to the model.

When is that flexibility useful?

An agent can be useful when a task involves several decisions along the way, and what decisions to make (or the order they occur) can vary from case to case. In these scenarios, the model can use what it finds to decide how to continue. If the same decisions follow a predictable pattern each time, then a predefined workflow may be a better fit.

For example, suppose the total in a monthly sales report differs from the total in the order system. An agent with access to both could first check whether they cover the same dates and count sales in the same way. What it finds would guide the next check: if one total includes cancelled orders, it could check whether those explain the difference; if orders are missing from the report, it could look for what they have in common. It could then explain the discrepancy and show the records supporting its conclusion, giving someone an answer to review without them having to guide each step of the investigation.

For this to work, the system needs access to the relevant records and tools, and we need a way to judge whether its explanation is correct. If someone has to gather all the information for it and then repeat the investigation to check its answer, much of the potential benefit disappears.

If the same known discrepancy happens every month, a fixed reconciliation workflow may be a better solution. Similarly, extracting fields from an invoice, checking them against a purchase order and sending discrepancies for review follows a procedure we can specify in advance. AI might help read the document, while standard software checks the amounts and applies the predefined approval rules.

In general, if I wanted to look for cases that are suitable for agentic work I would look at situations where there are multiple steps and where the order of those steps or the decisions needed varies from case to case. I would also want evidence that the model can reliably choose the right checks, interpret their results and decide how to continue.

What is the system allowed to do?

Giving an agent the freedom to decide what to check next does not mean giving it permission to make purchases or cancel orders. In our delivery example, it can investigate replacement products, but the customer still decides whether to buy one.

Equally, a predefined workflow could issue refunds automatically within agreed rules. Either design can require approval or be authorised to take certain actions on its own.

The harness gives us places to enforce limits, but we still need to decide what the system can do without asking us. That is what I mean by delegated authority. In our replacement example, the customer has asked for help finding an option and has explicitly kept the purchase decision for themselves.

We can divide that work between the model, ordinary software and a person:

  • The model chooses which replacement options to investigate and compares what it finds.
  • The software checks the price against the agreed budget and prevents a purchase without approval.
  • The customer decides whether the proposed replacement is acceptable and authorises the spending.

Either design can gather information and prepare an action while leaving the final decision with a person. In a disputed complaint, for example, AI can organise the evidence while someone decides what a fair resolution looks like.

The aim is to give the agent enough freedom to complete the task while restricting the actions it can take and the information it can access. For our replacement example, it might need access to the customer's order and the product catalogue, but not other customers' records or the ability to change prices. Those restrictions need to be enforced through tool permissions, access controls and approval checks in the software. Instructions to the model explain the rules, but the system also needs to prevent actions outside them.

The agent requests an action and software checks access, limits and required approval. An already authorised action runs within agreed limits. An action needing approval pauses and runs only if approved. A prohibited action is blocked.
Figure 4: Software checks whether an action may run, needs approval or must be blocked.

Retrieved text also needs care: a product description or tool result may contain instructions that try to redirect the model. The system needs to treat that content as information to assess, without letting it change the customer's permissions or bypass the purchase controls.

How do we check that the work was done?

Producing an answer and checking whether it can be trusted are separate jobs. When AI drafts a response, a person can review it before using it. Once a system can place orders or cancel them, we need checks before those actions go through, as well as confirmation that they succeeded.

Suppose the customer approves a replacement and authorises the system to place the new order and cancel the original. If it cancels the original before discovering that the replacement cannot be ordered, a reassuring message at the end is not much help.

Checking every step manually would recreate much of the work we wanted to remove. For the order-support example, I would separate the checks:

  • Before spending: software checks the total price against the budget and requires approval for the selected product, total price and stated delivery date. If any of those change, it asks again.
  • Before cancelling the original: confirm in the order system that the replacement order was accepted with the agreed delivery date. Approval to buy it does not tell us whether the purchase succeeded.
  • When no option works: pass the case to someone who can decide what to offer, with the options already tried and the reason each failed.

If a step fails, stopping the agent is only part of recovery. Suppose the replacement order goes through but the original cannot be cancelled: somebody still needs to resolve the two orders and explain the situation to the customer.

Would an agent help with your process?

Once a task looks suitable, I would test it on a small set of real cases before committing to a wider deployment. Three questions would guide that proof of concept.

1. Does it save useful work?

Choose representative cases, including routine work and difficult exceptions, and describe what an acceptable result looks like. The system can start by proposing what it would do while people continue handling the real cases. Compare its proposals with the actual decisions, allowing for more than one reasonable way to resolve a case. This can reveal errors before it has permission to make changes, although carrying out those actions still needs separate testing.

Compare it with the current process and a simpler predefined workflow. Look at how much work people still have to do: collecting information, correcting mistakes, checking results and taking over unresolved cases. The work also needs to happen often enough to justify the integration and support effort.

2. Can we limit and recover from its actions?

Write down what it may do independently and what requires approval, then test that the software enforces those limits. Include attempts to access records outside its remit or take actions the customer has not authorised. Start with actions whose consequences are limited or reversible, and provide a way to pause or stop a run.

Someone needs to own unresolved cases and recovery, with the evidence, authority and time to act. For our replacement order, that includes handling two live orders if cancellation fails. It also means deciding when repeated problems should lead to reduced permissions or a suspended service.

3. Is it reliable enough at a reasonable cost?

One of the first things I advise people using generative AI to do is compare a few models on examples of their own work. For an agent, keep the harness configuration, tools and permissions comparable too: changing what information a model receives or how many attempts it gets can change the result. I would look at:

  • Quality: does it complete the task correctly, follow the rules and ask for help when needed? Inspect both the outcome and the actions taken. Repeat important cases to see how consistent the results are.
  • Speed: how long does the whole task take, including checks and retries? A quick initial reply is only part of the experience if someone is waiting for the work to finish.
  • Total cost: include model use, paid tools, failed attempts and human review. A model that costs less per call may need more attempts or more correction.

Test cases should include situations where the agent gets stuck, repeats an unsuccessful action or chooses an unnecessary sequence of checks. Look at how often this happens, how much time and cost it adds, and whether the system stops or asks for help as intended. A few difficult cases can cost much more than a straightforward demonstration suggests. The tests should also show whether the limits leave enough room for the agent to complete legitimate work.

When testing the system, check that it actually completed the task. For example, if it reports that it placed a replacement order, confirm that the order actually exists in the system (with the correct product, price and delivery date).

Then, use those results to compare models, and note that we can also have different models perform different tasks (for example, we can use a more cost-effective model to make simple decisions and hand more complex tasks to a more capable model). A cheaper model may do the job well enough, while a more expensive one may need fewer retries or less human correction. What typically matters is the total cost of completing the work to an acceptable standard.

Once the system is in use, record what it did and which model, instructions and permissions it was using, so you can investigate mistakes. When you change the model, harness or tools, run the same checks again to see whether its performance has improved or new problems have appeared.

Putting it all together

What's most appealing about using agents is how much of a task they might be able to carry through without someone having to direct each step. That can be particularly useful when there are several decisions along the way, with the next action depending on what has been found. But ultimately the value depends on whether the system makes those decisions reliably and reduces the work someone still has to do.


If you have a process in mind and want a clear view of whether an agent would help, or there is some other AI question you want to talk through, get in touch.

Let's Talk

Have a question about AI?

Book a free 30-minute call and we will answer it honestly.