Week 13

The Attentive Waiter

In 2011, Ryan Buell and Michael Norton of Harvard Business School built a fake travel website. Participants searched for a trip and waited anywhere from ten seconds to a minute for results. Some watched a progress bar fill. Others watched a running list of the sites being searched, with fares appearing as they were “found.” Everyone received the same itineraries at the same prices.

The people who watched the work valued the service more at every wait time, and rated it on par with a version that returned results instantly. In a second study, participants used two competing sites that returned identical results, one instant and one slow. When the slow site showed its work, 62% chose it at a 30 second wait and 63% at a full minute. When it showed only a progress bar, it was 42% and 23% respectively. Buell and Norton called this the Labor Illusion, and their data traced it to reciprocity: a customer who sees a provider working on their behalf feels they owe them something, and pays back based on how they value the service.

None of that work was real. The site was a simulation, and nothing was being searched. People responded to the appearance of labor, and the appearance was enough.

The last of their five experiments is the one that is relevant for AI. The researchers moved to a simulated dating site. While it “searched,” the site played each participant’s stated preferences back to them as a counter ticked through 127 candidates, then returned a single match with a compatibility score of 96.4%. The only thing that varied was the photo, which a separate panel had rated in advance as attractive, average, or unattractive. With a good match, showing the work helped as before, and an average match followed the same pattern (though weaker). With a poor match, it hurt. People who watched the site work hard and then received a poor result valued the service least of anyone in the experiment, less than people who got the same poor result instantly. The authors compare it to a waiter who is attentive all evening and serves terrible food. The attention does not rescue the tip.

Visible effort, in other words, works as a multiplier. It makes a decent outcome look better and a poor one look worse, because once a service has shown that it tried, a bad result reads as the service’s own failure.

For most of the history of customer service, the signs of effort were expensive. A clerk who remembered your name, repeated your problem back to you and promised to look into it had to spend attention to do any of that. The cost is what made the signs worth trusting. They were hard to fake at scale, so customers could reasonably read them as evidence of work.

Chatbot designers have long added pauses to make their bots seem human, and the research supports doing so. In 2018, Ulrich Gnewuch and colleagues compared a customer service chatbot that answered almost instantly with one that paused before replying, longer for more complex messages. Users rated the bot that paused as more humanlike and were more satisfied with the conversation, in a setting where speed should have won. Large language models carry the same move from timing into language. An apology, the customer’s name, a sentence restating the problem, a promise to check: each now costs nothing to produce, and none requires the bot to have done anything.

I saw what that looks like when I audited United’s Help Center chatbot. I told its virtual assistant I wanted to book a flight as cheaply as possible, then asked what size bag I could bring for free. United’s carry-on page answers that near the top. The bot asked for my confirmation number and last name so that it could “check your specific booking and let you know exactly what you can bring at no charge.” I had not booked anything. Booking was the reason I was there.

Then I said I had been waiting over a week and was really frustrated. “I’m sorry you’re feeling frustrated, Ryan,” it replied, and asked what I had been waiting on. The next reply thanked me “for sharing how you’re feeling” and offered to “walk you through your options and help get things moving in the right direction.” It used my name in that reply and the three after it. Three replies in a row ended with the same four buttons: Refund, Flight change, Baggage, Something else. When I said I wanted to make a formal complaint, it sent a link to a feedback form and asked, “Was this response helpful?”

Every one of those replies carried the marks of effort. The bot read what I wrote, kept my name, acknowledged how I felt and described work it was about to do. What it delivered was a request for a number I did not have, the same four buttons three times, and a link. Buell and Norton tested a dating site, and applying their result to a help desk is my inference. If it holds, this is the arrangement that leaves a customer thinking least of the company: the effort fully on display and a poor outcome.

I’ll give the United chatbot some credit. It said it was AI before I typed a word, warned that its answers might be incomplete, and never claimed to have solved anything.

The cues also wear out. In a 2022 follow-up with 202 participants, the same research group found that a pause before replying raised the sense of social presence and the intention to use the bot among people new to chatbots. It had the opposite effect on experienced users. People read the signs of effort against what they have learned. My hypothesis is that this goes further as AI customer service spreads. Once enough customers learn that a warm, personal reply costs the bot nothing, warmth stops reading as effort and starts reading as the sign that a non-answer is on its way.

The far end of this is the doom loop. Last month I wrote about Frontier’s chatbot sending the same sentence, word for word, in reply to eight different messages. An identical reply is the one response that cannot pass for effort, because it shows that nothing was processed. In a narrow sense that makes the loop the more honest failure. It never pretends to be working.

Buell and Norton’s practical advice was to show customers the work, and their data shows where that advice stops. The effort with displaying is the work that changes the outcome. For a chatbot, the most convincing display of effort is an answer with its source attached: here is what United’s carry-on page says for each fare. That reply is open about its labor and delivers the result in the same breath.

Courtesy is the cheap half of service, and AI has made it cheaper still. A company that invests in the warmth of its bot while the answers stay empty is, on this evidence, paying to make its failures look worse. The gap between the care a conversation displays and what the person leaves with is what I mean by behavioral quality.