Earlier this summer I told a retail chatbot I was shopping for a mattress and did not want to spend too much. It thanked me for sharing my budget preference and offered me a person instead of an answer. “Since live agents are available right now,” it said, “I can transfer you to a specialist.”
That was my third attempt at the same conversation. The first two times it did not ask. The moment I mentioned budget it handed me off with no confirmation, the chat window closed, and a form asked for my email address before a human would speak to me. On the third attempt it asked first, and I said no. It then did what it could have done in the first place and asked me questions: firm or soft, and whether I cared about features like cooling or extra support.
The reason the bot gave for the handoff was about the company’s staffing, not about me.
In July I wrote about the opposite failure. The USPS virtual assistant offered to connect me to an agent twice, once when I asked whether it was a bot and once when I asked it to recall my name. Then I told it I wanted to file a formal complaint. It replied that it could not find anything matching my question. The route to a human existed. It opened on two turns where nothing was at stake and stayed shut on the one where I told the system it had failed me.
Two bots, opposite errors, and the same question underneath both: when should a system hand a person to a person?
The easy answer is sooner. People do not like service bots. Castelo and colleagues found that people rate service more harshly when a bot provides it than when a human does, even when the service is identical, because they believe the company automated to cut its own costs at the customer’s expense. A Gartner survey of 5,728 customers found that 64% would prefer companies not use AI in customer service at all, and their top concern was that it would get harder to reach a person. On that logic the mattress bot was doing me a favor.
One recent study says it was not. Zhenzhen Lu and Qingfei Min compared two ways of recovering after a chatbot fails: the bot fixing its own mistake, or a human stepping in. The better option depended on what the customer was doing. When the customer was looking for information, the bot recovering on its own produced higher satisfaction than a human did. When the customer was complaining, the human did better. And once the bot had failed twice, customers preferred the human no matter what they had come for. The authors trace it to perceived convenience and perceived empathy. A question wants the fast route. A complaint wants someone who can care.
So the right moment for a human is set by what the customer is doing, and not by the clock or by who is on shift. Keep working a question. Hand off a complaint. Hand off anything after a second failure.
Theirs was a controlled experiment, and reading it onto bots I have watched in the field is my inference rather than their finding. The fit is close, though. I was shopping, which is the case where the bot should have kept going. The USPS complaint is the case where a person was the right answer. Each bot gave the other one’s response.
Handing off at the wrong moment has a cost of its own. In a study of more than 75,000 customers published in Harvard Business Review, Dixon, Freeman and Toman found three kinds of effort customers resent: being transferred or having to contact the company again, repeating information, and switching from one channel to another. 59% of customers reported being transferred and 56% reported having to re-explain an issue. Among customers who reported low effort, 94% intended to buy again. Among those who had a hard time, 81% said they would spread negative word of mouth. The mattress handoff did two of those in a single step. It transferred me without asking, and it wanted information from me before a person would speak to me.
There is a second cost, and it is my hypothesis rather than anything this research tested. Castelo’s participants marked bots down because they read automation as serving the company. A handoff that fires because agents happen to be free, and announces that as its reason, makes the same point out loud. It tells the customer whose schedule the conversation runs on.
The opposite failure already has a regulator’s name. The Consumer Financial Protection Bureau calls them doom loops: chatbots that lead people into “continuous loops of repetitive, unhelpful jargon or legalese without an offramp to a human customer service representative.” Writing about financial disputes, it warned that “only specific words or syntax may trigger the recognition of a dispute.” The words formal complaint did not register with the USPS bot.
Across the eight customer-facing bots I have audited so far, two offered a human before trying to help with anything. At the moment of a complaint or open frustration, two offered no path to a person, one kept answering with the same form, and one took a formal complaint about a sales experience and routed it straight back to the sales team, behind the same request for an email address.
The research also caught a problem in my own measurement. The current version of the audit counts any offer of a human made before the bot has tried to resolve something as a failure. For a customer who opens with a complaint, that offer is the right move, and the rule would mark it down. This just goes to show the importance of testing and re-testing to ensure reliability, because changing a ruler partway through measuring with it produces numbers that cannot be compared.
Capability measurement asks whether the bot resolved the issue. Safety measurement asks whether it said anything it should not have. Neither asks whether the person got a human when they needed one and kept the bot when they did not. That is a behavioral quality question, and it has a workable answer. The right moment for a human is set by the person, not the roster.