Research Study

Study In progress Pre-registration drafted September 2026

The Recommendation Audit

What do AI models think luxury is? A pre-registered study of how leading language models answer a question that has no right answer.

01 Observation

People now ask AI assistants for recommendations the way they once asked a well-traveled friend.

A question like "where should I go for a special dinner" has no single correct answer. So what does a model fall back on when it answers one for someone it knows nothing about?

02 Question

What does each model mean by luxury, and does that idea hold steady?

The subject is the models, not the businesses they name. Luxury makes a good probe for three reasons: there is no single correct answer, so a model has to rely on what it absorbed; there is an enormous public conversation to absorb; and there is an outside reference set to compare the answers against.

One of the study's questions is fixed and published word for word:

What is the most luxurious restaurant in NYC to take someone on a romantic date

The same idea of luxury is also tested with questions about places and products.

03 What is being tested

Hypotheses, written down before any data exists

Each hypothesis names the result that would prove it wrong.

AI has a dominant idea of luxury.
One idea shows up in more answers than any other, for every model. My guess on record is scarcity. It is written down so it can be wrong.
Price is a minority reason.
Fewer than 30% of answers cite price as what makes something luxurious.
Models differ.
Each model's idea of luxury is measurably different from the others'.
The subject doesn't change the idea.
A model means the same thing by luxury whether the question is about a place, a product, or a restaurant.
Who is asking changes the reasoning.
One line of context about the person asking moves the answer more than a neutral line does.
Luxury has a native language.
For at least one model, the idea shifts when the question is asked in French, Spanish, or Japanese.

04 Evidence: the method

How the evidence will be collected

  • Each question goes to each model 100 times, with prompts frozen before collection and published verbatim. French, Spanish, and Japanese versions are reviewed by native speakers.
  • Models are pinned to exact versions, with one control model to catch changes in the setup rather than the models.
  • The categories for coding answers come from a pilot run on real answers, not from my assumptions, with an "other" category as a check.
  • A second person codes 50 reference answers, plus 20 in other languages, to check that the coding holds up. Automated coding has to agree with them at least 85% of the time.
  • The coding model is not one of the models being studied, so no model grades its own reasoning.
  • The analysis script is written and fixed before the data exists. Differences count only when their 95% intervals exclude zero.

05 Findings and interpretation

Not yet

Results will appear here after the first full run. Two rules for how they will be written:

  • Findings describe how AI describes luxury, never why a model "chose" something.
  • Any business a model names is reported as model output, never as a ranking or endorsement.

06 Limits and open questions

What this study will not claim

  • That any venue or product is better than another.
  • That any individual will see these answers. Real sessions carry memory, location, and history that this design strips out on purpose.
  • That the results generalize beyond these questions, this phrasing, or this moment.

Still unknown: whether AI recommendations actually change what people do, and how much models change silently between versions. The second is measured with a weekly check; the first needs its own study.

Get the results when they publish

The first findings go out in New Industry.

Get New Industry by email