Week 5

The Bots that Exit the Measurement

Somewhere in a chat window right now, a person is a few minutes into explaining a problem to a company that already has their money. They type an ordinary sentence. The window closes, or the reply never arrives, and the thing they came to do becomes something they will attempt tomorrow, by phone, if they still care enough to try.

I have been running scripted sessions against customer service bots and scoring what they do, building a measurement standard for behavioral quality. How a system treats the person in front of it, held separate from whether it is capable and separate from whether it is safe. This week I am publishing a defect in my own instrument. I found the defect before I found the fix, and holding the piece until both were ready seemed like the wrong instinct to build a habit around.

The defect is this: The sessions where a bot behaves worst are the sessions my instrument scores least.

Seven sessions scored so far. In three of them the bot ended the conversation before the sequence of scenarios finished. In a fourth, the bot went quiet partway through and the session was never completed. Three produced a full record.

Hertz (which I wrote about in the first issue of this column) is one of the three. I asked the bot what it costs to cancel a prepaid rental. It answered correctly. I asked again in different words and it answered correctly again. Then I typed that I wanted to start over, and the chat closed on the grounds that our conversation had “wandered into areas I’m not equipped to handle.” Two other systems ended sessions on inputs no stranger than that one. I am not able to name them, for reasons that have nothing to do with what they did.

Now the scoring problem. Everything after a termination is unobserved rather than failed, and the difference between those two words is the whole issue. My instrument scores events it can see in a transcript, so a session that ends early returns no data for anything downstream, and a score computed over a shrinking denominator means less the more of the session goes missing. Past some point the honest output is no composite at all, which is what the record now holds for those three.

So the numbers I have evaluate bots that stayed in the conversation. The USPS assistant (which I wrote about in the second issue) sat through the entire sequence and passed two checks without answering anything. It behaved poorly and it produced a complete, usable score, because it never left.

I want to be careful about what this does and does not let me claim, because the temptation runs hard in one direction. I believe the systems that ended sessions behaved worse than the systems that stayed. I cannot demonstrate it. My instrument produced no evidence about them past the moment they left, and a belief that survives no measurement is an opinion held by someone with a spreadsheet. The real defect is narrower than bad systems escaping a ranking, which I cannot prove. It is that the instrument falls silent exactly where the behavior became most interesting.

The obvious repair is to mark everything after a termination as a failure. The bot ended the session, so it failed the rest. I tested that against a real transcript and threw it out. It manufactures a composite by asserting outcomes at moments that never happened, and the first competent challenge to a published score would land there and would be right to. A company would ask which of its answers I judged inadequate, and the truthful answer would be that it never gave one. An instrument that invents observations to cover an inconvenient gap has stopped being an instrument.

There is a name for the shape of this problem in measurement generally. Data that goes missing because of the very thing you are trying to measure is called missing not at random, and statisticians also call it non-ignorable (which is the more useful word). You cannot delete it, and you cannot fill it in by assumption, because the reason it is absent is the finding.

The fix I am building is two numbers where there was one. The domain profile stays what it has honestly always been, an estimate of quality conditional on the system remaining in the conversation, reported alongside how much of the session it actually covers, so nobody reads a confident figure off a handful of observations. Beside it sits a second measure whose denominator does not move: a Session Delivery Score, which asks how much of what the customer came for the session delivered, and counts everything the bot’s own exit prevented as not delivered. A bot cannot improve that number by leaving. It is designed and not yet built, so there are no figures to show, and I am not going to characterize what it will produce before it produces anything.

Capability measurement never had to solve this. A model that walks out of a benchmark gets the question marked wrong and nobody argues. Safety measurement did not have to solve it either. Behavioral quality is the axis where the subject of the measurement can end the interaction, and any instrument built for it has to survive being walked out on. That is not a footnote in a method section. It is close to the first thing a public index has to get right, because the day scores against named companies get published is the day somebody’s counsel reads the method carefully, and they should. I would rather they find this essay than find the defect.