Georgii EmelianovProduct

The chat box is the bug

Every vendor list of chatbot failures is a list of symptoms with one cause nobody names, because naming it would rule out the thing they are selling you.

Torn white, cool grey-violet, and black paper framing the headline "The chat box is the bug".

A customer opens your app and types: "Why is my bill higher this month?"

It is one of your most common contacts, and it has no one-sentence answer. Your assistant will reply with a paragraph anyway, and then the customer will contact you regardless. That second contact is the one you pay for.

Every published explanation of why chatbots fail customer support misses it. The explanations are not short of detail; they are short of one admission. The assistant did not fail because it misunderstood the question. It failed because the only thing it can hand back is prose. The honest answer here is not prose. It is last month beside this month, the four line items that changed, the date a promotional rate ended, and two things the customer could do about it.

The short answer

Chatbots fail because the format is the constraint. A chat box can only return text, so every answer whose real form is a comparison, a set of numbers, a plan with steps or a map has to be squeezed into a paragraph first. Customers experience that compression as a bot that will not help, and contact you again.

Why chatbots fail customer support, according to the people selling them

The published answers are consistent, detailed, and written almost entirely by companies whose product is a chat box.

eGain's list is the canonical one — twelve reasons, checked 21 August 2026. It is a good list. The bot tries to do too much. It pushes a help page instead of answering. It misreads intent. It hands you to a second bot that has forgotten the conversation. It hides on the site. It takes too long. It sounds nothing like your brand. Competing lists — seven reasons, ten reasons, five — cover the same ground, and the leading item on most of them is weak escalation: the customer gets stuck, and the bot keeps talking.

Then every one of them arrives at the same prescription. Better knowledge. Better reasoning. Better routing. Better analytics. A pilot. In other words: the box is fine, your implementation of it was not, and here is a better box.

It is worth checking whether that prescription is working, and the honest measure is not the vendor's. Gartner surveyed 5,728 customers in December 2023 and found that only 14% of customer service issues are fully resolved in self-service (published 19 August 2024, checked 21 August 2026). That number has survived several generations of better boxes.

When a dozen independent lists of causes all produce the same list of symptoms, the thing they have in common is usually the cause.

One question, two answers

Here is the bill question answered both ways. Same customer, same account, same underlying information.

As a paragraph. "Your bill this month is $94.20, which is $22 more than last month. This is mainly because your promotional discount ended on the 3rd, and there was also a one-off charge for international calls. You can view your full bill in the Billing section, or contact us if you'd like to discuss your plan options."

That is a good answer. It is accurate, it is polite, it names the cause. Now read it as the customer. Two numbers to hold in your head, one date with no context, a charge with no amount, and a pointer to a screen you have to go find. It ends by inviting you to contact support, which is the outcome the whole thing existed to avoid.

As a screen. The same answer, delivered as structure: this month and last month side by side at the top, with the difference called out. Below that, the four line items that changed, each with its own amount, so the $22 is not a claim but an arithmetic the customer can follow. Below that, the promotional rate with its end date in plain language. At the bottom, two options they can tap — see the plan that would have been cheaper, or talk to someone — with the second one carrying the whole context so nobody starts over.

Nothing in the second version is new information. Every fact in it was already in the paragraph or already in your billing system. What changed is that the customer can see the comparison instead of being told about it, and the follow-up question they were about to ask has already been answered on screen.

That is the entire argument, and it is the reason the interface matters more than the model.

What compression actually costs

Go back to the symptom lists and read them again with the format in mind. Most of the items stop looking like separate problems.

"It pushes a link instead of answering." The answer did not fit in the reply, so the bot handed over the address of a place where it does. That is compression, visible.

"It cannot complete actions." The customer needs to choose between three options. A paragraph can describe three options; it cannot offer them. So the bot describes, and the customer goes looking for the buttons.

"Escalation is weak." The bot has run out of room before it has run out of answer. Handing off is the only remaining move, and it happens at the point where the answer got structurally hard rather than at the point where it got genuinely difficult.

"It repeats itself." A customer who cannot see the whole answer asks the next part of it as a new question, and receives another paragraph that overlaps with the last one.

"It doesn't sound like our brand." Prose is the only channel it has, so every ounce of the experience is carried by wording. In your app, most of your brand is carried by layout, spacing and components — none of which the box can use.

Fixing any one of these inside a chat box means writing a better paragraph. There is a ceiling on that, and the industry has been up against it for a while.

The five answer shapes a paragraph cannot hold

The useful test is not "how smart is the assistant" but "what shape is the honest answer." Five shapes come up constantly in account-based service, and none of them survives being flattened into prose:

A comparison. This month against last month, your plan against the one that fits you, the claim you filed against what the policy covers. Comparisons need columns. Read aloud, they become a list of numbers the customer has to re-assemble.

A set of numbers that belong together. Data used, days remaining, amount due, next charge date. Four numbers with labels is a glance. Four numbers in a sentence is a memory test.

A plan with steps. What happens next, in order, with what is needed at each point. Ordering and progress are the information. A paragraph delivers the words and loses the sequence.

Places. Where the nearest one is, what is open, how far. This is a map, and a map described in text is directions to go and find a map.

A set of options. Three things the customer could do, each with a consequence. Options want to be tappable. Described, they are homework.

If your top contact drivers mostly produce these shapes, no amount of model quality will fix your chat box, because the constraint is not comprehension.

Does your version of this fit in a sentence?

This is the qualifying question, and it disqualifies plenty of people.

Go and read your twenty most frequent in-app contacts. For each one, write down the honest, complete answer — the one a good agent would give, not the one a bot gives. Then look at the shape of what you wrote.

If most of them genuinely fit in one or two sentences, a chat bubble is the right tool and you should keep it. "Where is my order" with a date, "am I covered for this" with a yes, "when does my subscription renew" — these are text answers. Adding structure to them is engineering that buys nothing, and any vendor telling you otherwise is selling.

If most of them come out as a comparison, a set of numbers, a sequence or a map, the box is your ceiling, and the model you put behind it is not the variable.

Most companies have a mix, and the split matters more than the average. If four of your top five contacts are shaped answers and the rest are one-liners, that is a strong case. If it is the reverse, it is not, and the honest thing is to say so before anyone builds anything.

Where this is the wrong idea

If your hardest questions end in a purchase, not an explanation. Answering in structure suits "what happened, why, and what now." It suits commerce much less well — browsing already works, a grid of cards is the one shape a chat box can approximate, and nothing here is urgent enough to fund.

If you need the assistant to change things. The approach we take is deliberately read-only: the assistant explains, compares and shows options, and anything that changes an account hands off to the flow you already built. In account-based service that handoff is the correct behaviour and customers barely notice it. In other products it is a dead end, and you should count it as a limitation rather than a design choice.

If you want a number from us first. We do not have one. Nobody has run a controlled deflection study on this approach yet, and the honest position is that the argument above is a structural one rather than a measured one. The Gartner figure earlier is somebody else's research about self-service generally, not a claim about our results. Any vendor quoting you a resolution improvement for a product this new is quoting you a projection.

What it takes to try it

  1. Pick one contact driver, the highest-volume one whose answer is a shape rather than a sentence. Not three; one.
  2. Write the answer out as a screen — on paper, with a pen. If you cannot lay it out on paper, the problem is not the technology.
  3. Check whether your systems already hold everything on that page. They usually do. The information is rarely the missing piece.
  4. Decide who owns the outcome. A contact-rate number for that one driver, agreed before anything is built, measured the way you already measure it.
  5. Give it a named engineer. Not a team — a person who will be the one to say whether it worked. This is the single strongest predictor of whether a pilot like this finishes.

None of that requires a decision about a vendor, and steps one and two are worth doing regardless.

Frequently asked questions

Why do customers dislike chatbots so much?

Because the interaction usually costs them more effort than the alternative. They arrive with a question that has a structured answer, receive a paragraph that partially covers it, and end up asking a follow-up or contacting a human anyway. The dislike is about the second contact, not about the technology.

Are AI agents different, or is it the same problem?

Better models handle more question types correctly, which is real progress. The delivery constraint does not move: if the reply is still a stream of text in a box, a comparison still arrives as sentences. Smarter reasoning behind the same output format changes what the assistant knows, not what the customer can see.

Can't we just make the chatbot's answers better?

Up to a point, and most teams should. Shorter replies, fewer links, better handoffs all help. The ceiling arrives when the answer's honest form has columns or a sequence in it, because at that point writing it better means writing more words, and more words is the thing customers are already complaining about.

What should we do instead of a chatbot?

Keep it for the questions that genuinely fit in a sentence — that is most support queues' largest bucket. For the contacts whose answers have shape, look at whether your app can show them as screens instead. That is a different product from a chat bubble, and it is worth treating as one.

Where to start

Do the twenty-contact exercise. It takes an afternoon, it needs nobody's approval, and it will tell you which of the two arguments above applies to you — including, quite possibly, that your answers do fit in a sentence and you should stop reading vendor blogs.

If they do not fit, the next thing worth reading is why an in-app assistant is a different product from a chat bubble rather than a better one. The rest of the writing here covers how it is built, for the people on your team who will ask.

And if you would rather see the bill question answered as a screen than read a description of one, we will show you the working version. Email hello@uzori.ai with your top contact driver and we will build that one.

← All posts