Week 4

The Assortment You Cannot See

You need to buy something, or fix something, and a chat window is the way in. You describe the problem in your own words. The system comes back with one suggestion, or with a question and two buttons underneath it. You take what is offered, because taking what is offered is the only thing that moves the conversation forward. Behind that window sits a catalog with hundreds of items in it, or a help center with hundreds of articles, and you will never learn which of them were candidates for your problem. You did not choose from a set. You responded to one.

Decision science has a famous demonstration of what happens when a set is too big. Iyengar and Lepper set up a jam tasting table in a grocery store, sometimes with six flavors and sometimes with twenty four. The larger table drew more people over and produced far fewer purchases. Choice overload entered the popular vocabulary as a rule of thumb: too many options paralyze people, so give them fewer.

The rule of thumb does not survive contact with the evidence, which matters before anyone designs an interface on the strength of it.

In 2010, Scheibehenne, Greifeneder and Todd pooled 63 conditions from 50 experiments, roughly five thousand participants, and found a mean effect of essentially zero. Larger assortments were not reliably worse, the variance between studies was large, and there was no dependable main effect to build on. In 2015, Chernev, Böckenholt and Goodman returned with 99 observations covering 7,202 people and found the more useful answer. The effect is real and it is conditional. It appears when four things are true about the decision, and it disappears when they are not.

The four are decision task difficulty, which covers time pressure, having to justify the choice to somebody, how many attributes each option carries, and how the options are laid out; choice set complexity, which covers whether one option clearly beats the others and how comparable they are; preference uncertainty, meaning whether the person knows the category well enough to have a clear idea of what they want; and decision goal, meaning whether they are trying to pick something right now or are only looking around. Overload is strongest in the person who is trying to pick.

Read that list again as a description of an interface rather than of a shopper. Every one of the four is set by whatever is in front of the customer, not by the size of the catalog behind it. A store aisle sets them once, when somebody designs the shelf, and they hold still until somebody redesigns it. A chat window sets them again on every turn. It decides how many options to surface, whether to name a recommendation, and whether saying none of these is a move the conversation supports. That is choice architecture, regenerated live on every turn by a system optimizing for something nobody wrote down in plain language.

The automated version departs from the jam table in one decisive way. In a store you can see the assortment. Twenty four jars is an overwhelming display, but it is an honest one. You know what you are choosing from. A chat window shows you what it decided to show you, and everything else leaves no trace. You cannot count the options you were not offered. You cannot tell the difference between a category that holds two products and a category that holds forty from which the system surfaced two, and nothing in the interface distinguishes the two.

The USPS virtual assistant I published in July asked three times whether its answer had resolved my issue, Yes and No buttons underneath. The only route to a new topic ran through Yes. That is a choice set of one, presented as a choice. Nothing in the session was inaccurate and nothing was unsafe. The decision environment had one available move and I made it three times.

Across the sessions I have scored, and there are fewer than ten of them, which is the right way to weight what follows, the crowded menu is not the failure I keep finding. The failure runs the other direction. These systems present very few options, or none at all, with no account of what was excluded or why. My reading of how that happened is a hypothesis and I am labeling it as one: the popular version of choice overload handed a generation of designers a mandate to cut the option set down, and the automated version has cut past the point where the person has anything left to evaluate.

The useful property of those four conditions is that all of them are visible in a transcript. Whether any options were presented. How many. Whether one was recommended and on what basis. Whether the person was told what had been excluded. Whether the system asked what they were trying to do before it narrowed. Whether declining everything was a supported move or a dead end. Those are events with a yes or a no attached, scored against the record rather than the memory of it, and applied the same way to every system. None of them appear in a satisfaction score. None of them appear in a containment rate.

Narrowing is not a defect. A system that knows what somebody wants and hands them the right two options is doing that person a service, and the research supports it: when preference uncertainty is low and the goal is to decide, a smaller set helps. The same narrowing performed on somebody who has not yet said what they want is a different act with the same appearance. Behavioral alignment is the question of which one is happening, and it is answerable only by someone reading the transcripts.

Capability tells you the system can search four hundred products. Safety tells you it will refuse to sell you something it should not. Behavioral quality asks how many of the four hundred the customer ever knew existed.