The price comparison is an accurate answer to the wrong question.
Every comparison you'll read comes down to "an agent costs several times more per conversation than a chatbot." That's true and it doesn't matter. Model spend is under 3% of what support automation actually costs you; the other 97% is the human who picks up whatever the software didn't finish. On those numbers an agent only has to resolve 0.7 percentage points more than a bot to pay for itself. The real question is not price. It's whether it resolves more, and whether you can prove it.
There is no shortage of articles comparing these two. They're near-identical, they all define the difference correctly, and they all reach the same conclusion, which is that agents are more capable and more expensive.
The second half of that conclusion is doing almost no work. This guide replaces it with the number that decides real budgets: what one resolved customer contact costs, all in.
If you're weighing up an AI chatbot against something that can take action, this is the arithmetic we'd run on the call before quoting either.
What Each One Actually Does
Briefly, because this part is genuinely well covered elsewhere.
The line between them is whether the software touches your systems. A bot that says "your refund policy is 30 days" and an agent that issues the refund are different products, even when they use the same model.
Why the Price Comparison Misleads You
The figure in circulation is that an agent costs somewhere between three and ten times more per conversation than a chatbot. Our own breakdown of what an AI agent costs to run lands in that range: about $0.017 for a six-turn task against roughly $0.002 for a single retrieval call.
Then you put it next to the cost of a person handling the same contact. Six minutes at a loaded rate of twenty-two dollars is $2.20.
Every published price comparison is an accurate answer to a question that decides nothing.
We watch this play out on scoping calls. Someone arrives having read three comparisons, convinced the agent is the expensive option, and asks whether they can start with something cheaper. Then you put the human column on the same page and the conversation changes in about a minute, because the thing they were protecting their budget from turns out to be one and a half cents.
If you take one thing from this article, make it the habit of putting your own handling cost next to any per-conversation price you're quoted. You will find that almost every automation decision in support is really a decision about what proportion of contacts still reach a person, and the software price is decoration on top of it.
What Each One Actually Resolves
This is the number worth arguing about, and the published bands are reasonably consistent.
| Type | Typical resolution | What it can do |
|---|---|---|
| Basic chatbot | 20% to 40% | FAQ retrieval, scripted flows, ticket capture |
| Assistant with business logic | 40% to 60% | Embedded rules, some account context |
| Action-taking agent | 70% to 85% | Reads and writes your systems, completes the task |
Sources: Lorikeet and Fin both publish 2026 benchmark ranges in this shape. Both are vendors in this market, so treat the top of each band as a best case rather than a forecast for your deployment.
The gap between the first and third row is the whole decision. Everything below is about what that gap is worth and whether you'll actually get it.
The Real Comparison: Cost Per Resolved Contact
Same hundred contacts, both routes, everything counted.
| Per 100 contacts | Chatbot at 35% | Agent at 65% |
|---|---|---|
| Model spend | $0.20 | $1.70 |
| Escalations to a person | 65 | 35 |
| Human handling | $143.00 | $77.00 |
| Total | $143.20 | $78.70 |
| Cost per contact | $1.43 | $0.79 |
The agent is 45% cheaper per contact, and the model spend is 2.2% of its total against 0.1% for the bot. At five thousand contacts a month that difference is about $3,200.
The Break-Even Is Smaller Than You Think
Here's the number that settles the price argument for good.
The agent costs $0.015 more per contact in tokens. Each escalation it avoids saves $2.20. Divide one by the other and you get the resolution lift the agent needs before its premium pays for itself.
- 0.68 percentage points. That's it. If the agent resolves 35.7% where the bot resolved 35%, it has already paid for the extra tokens.
- At equal resolution the agent costs $1.50 more per 100 contacts. Not per contact. Per hundred. If both resolve 35%, the entire penalty for choosing the more capable architecture is a dollar fifty.
- Every point above that is roughly $2.20 per hundred contacts saved. Linear, and it does not care about your volume, which only scales both sides equally.
So stop asking whether you can afford an agent. On these numbers almost everyone can. Ask whether this particular agent will resolve more than the bot would have, which is a completely different investigation and a much harder one.
That shift is worth making explicitly, because it changes who you should be talking to and what you should be asking them. A pricing conversation is a procurement exercise: you compare rate cards and pick one. A resolution conversation is an engineering exercise: you look at your actual contact mix, work out which intents a bot genuinely closes, and find out whether the remainder can be reached by software at all.
You will also find the second conversation is the one vendors are least prepared for, which tells you something useful in itself.
When the Agent Costs You More
The honest inverse, and it's a real risk rather than a theoretical one.
An agent that resolves less than your bot costs more on every axis. It burns more tokens, it takes longer, and it still hands the same work to a person. Below roughly 35% resolution on this model it is worse than the chatbot in pure cash terms, and worse again in customer patience.
That is not a hypothetical failure mode. An agent chaining six model decisions at 95% each lands at 73.5% end to end, and a badly scoped one lands much lower. We work through why in when an AI agent is the wrong tool, and the short version is that autonomy applied to a process that was always a flowchart buys you variance you didn't need.
Want this run on your actual support volume?
Send us a month of ticket data and your current handling time. We'll model both routes on your numbers and tell you what the resolution lift would have to be to justify either, including when the honest answer is to keep the bot you have.
Book a free support automation review ↗Containment Is Not Resolution
The measurement trap, and the reason vendor resolution numbers should be read carefully.
Containment counts conversations that didn't reach a human. A customer who asks a question, gets a vague answer, gives up and leaves is contained. They are also unhappy, and they will probably contact you again through a channel you're not measuring.
Resolution counts problems actually solved. Only the second one appears in the cost model above, because only the second one removes work.
The gap between them is large. Gartner surveyed 5,728 customers and found that only 14% of customer service issues are fully resolved in self-service, despite 73% of customers using self-service somewhere in their journey. Even for issues customers described as very simple, only 36% resolved fully.
Sit with the distance between that 14% and the 70-85% in the vendor benchmark table above. Some of it is genuine progress since that survey. Some of it is the two sides measuring different things and using the same word.
- Ask which metric a quoted number is. If a vendor says 70% and cannot immediately say whether that is containment or resolution, it is containment.
- Ask how resolution was verified. Solved means the order was actually refunded or the address actually changed, checked in the system, not inferred from the customer not replying.
- Ask about repeat contacts. A contact that returns within 48 hours was not resolved, whatever the first interaction logged.
- Ask what the customer said. Post-contact satisfaction on automated resolutions is the only measure that catches a confidently wrong answer.
Our guide to evaluating an AI agent before you ship it covers how to establish these numbers yourself rather than accepting them, and the headline rule applies here too: score the end state, not the reply.
What Actually Decides This
Given that price is settled, three things are left.
- Can it reach your systems?The real gateAn agent that can't read the order database is a chatbot with a bigger bill. If the integrations aren't there and can't be built, the resolution lift you're modeling is imaginary.
- Is the work repeatable?Shape, not volumeSupport contact mixes are usually a long tail with a fat head. The head is where automation pays, and you want to know what share of your volume the top ten intents represent before choosing anything.
- Can you verify resolution?The one people skipIf you cannot tell resolved from contained in your own data, you cannot tell whether the thing worked. Fix the measurement before you buy the software, because otherwise you will be told it worked.
The Hybrid That Usually Wins
Almost nobody should build only one of these.
The pattern that works is a fast deterministic layer in front and an agent behind it. Simple, high-volume, read-only questions get answered instantly and cheaply. Anything needing account context, a lookup, or an action goes to the agent. Anything the agent can't finish escalates with the full history attached, so the person doesn't start from nothing.
CollageDepot's support pipeline is built that way. Every incoming email is classified first, sentiment and urgency are scored deterministically, order-related messages trigger a live Shopify lookup, and replies go out in the customer's own language. It handles over 5,000 emails a month across four languages, resolves 65% in under sixty seconds, and routes the rest to a person with the context already gathered.
The number worth noticing there isn't the 65%. It's that the escalated 35% arrive pre-enriched, so even the contacts the automation didn't resolve cost less to handle than they used to. That effect is missing from every model in this article, including ours, and it moves the answer further toward the agent.
We would encourage you to model it, because it is easy to measure and nobody does. Take your current average handling time, then estimate what it becomes when the person opens a ticket that already has the order record, the sentiment score and a drafted reply attached. If six minutes becomes four, you have just improved the economics of every contact the automation failed on, which is the half of the ledger everyone ignores.
There is a softer version of the same point. Escalations that arrive with context are less unpleasant to work, and support teams that spend their day on genuinely interesting cases rather than order-status lookups tend to stay. That does not go in a spreadsheet, and it is one of the more reliable effects we see after a good deployment.
Which Should You Build?
High volume, simple, read-only questions
A chatbot is genuinely the right answer, and cheaper to run and maintain. Order status, opening hours, policy lookups. If your top intents are all "tell me something", you don't need anything that can act.
Anything that requires account context
Agent. The moment the correct reply depends on which customer is asking, retrieval over help articles cannot get there, and every contact becomes an escalation no matter how good the copy is.
A mixed inbox, which is most businesses
Both, in the order above. Put the deterministic layer in front, measure what it actually resolves for a month, then scope the agent around the residue. That sequencing also gives you the baseline you'll need to prove the agent earned its place.
A regulated or high-stakes process
Agent, with the safety controls and approval points designed before the happy path. The cost model still favors automation; what changes is that a wrong action costs far more than $2.20, so the escalation route stops being a fallback and becomes the main design problem.
Frequently Asked Questions
Is an AI agent more expensive than a chatbot?
Per conversation, yes, by roughly $0.015 on a typical support task. Per resolved contact, usually no: on our model the agent came out 45% cheaper because it escalated to a person half as often, and human handling is over 97% of the true cost either way.
How much more does an agent need to resolve to be worth it?
About 0.7 percentage points, on these numbers. That is the point where the extra token spend equals one avoided escalation. Almost any genuine capability improvement clears it, which is why the price comparison is not the useful comparison.
What is a realistic resolution rate?
Published 2026 benchmarks put basic chatbots at 20% to 40%, assistants with business logic at 40% to 60%, and action-taking agents at 70% to 85%. Those come from vendors in the market, so plan against the lower end of the relevant band and measure your own.
What is the difference between containment and resolution?
Containment counts conversations that never reached a human, including the ones where the customer gave up. Resolution counts problems actually solved. Only resolution removes work, and the two get quoted interchangeably. Gartner found only 14% of self-service issues were fully resolved from a customer's point of view.
Can I start with a chatbot and upgrade later?
Yes, and it's usually the right order. The bot gives you a resolution baseline, a map of your actual intent mix, and the knowledge base the agent will need anyway. What doesn't transfer is the integration work, so scope that separately rather than assuming the bot vendor's connectors will carry over.
Does volume change the answer?
Not the direction, only the size. Both sides scale with contacts, so the percentage difference holds. Volume decides whether the build is worth doing at all: at a few hundred contacts a month the absolute saving may not repay the project, however good the ratio looks.



