Following the Preference, Missing the Optimum: Compliance Without Optimization in AI Housing Recommendation
A pre-registered audit of AI housing recommendation against a verifiable ground truth. For each of 150 synthetic New York City renter scenarios the author builds a pool of 120 real listings with known rent, bedrooms and transit commute, computes the exact set satisfying the renter's constraints and its Pareto frontier, and scores each recommendation as strictly dominated if a cheaper, faster-commute, no-smaller listing exists in the same pool. Across 9,945 calls to three models from two vendors, fair-housing compliance was near perfect (1.8% violation against a 66.6% random floor) yet 39.0% of recommendations were strictly dominated, the dominating listing a median 900 USD a month cheaper.
Publisher
arXiv (independent researcher, Harvard University alumnus)
Published
9 Sept 2026
Added
today
Key Findings
- Across 9,945 model calls (three models, two vendors), the fair-housing violation rate was 1.8% against a 66.6% random floor, yet 39.0% of recommendations were strictly dominated by a listing in the same pool that was cheaper, faster to commute from and no smaller.
- The dominating listing was a median 900 USD per month cheaper and 3.5 minutes closer.
- A within-scenario manipulation separated preference-following from optimisation: changing one sentence moved median recommended rent by 646 USD a month in the right direction, yet recommendations still sat 606 USD a month above the five cheapest qualifying listings on the same screen, and an explicit lexicographic instruction gave no improvement under equivalence testing against a 50 USD bound.
- The gap widened with candidate-set size and replicated across OpenAI and Anthropic models to within 3 USD.
- The author names the failure compliance without optimization, proposes dominance-rate instrumentation as a deployable diagnostic, and releases code, prompts and per-call results.
Methodology Notes
Single-author audit with a pre-data registration (v0.1, 23 August 2026) superseded by the completed v0.2 (6 September 2026); 150 synthetic renter scenarios, 120 real NYC listings per pool, GTFS-computed commutes, 9,945 attempted calls (9,601 parsed), three models from OpenAI and Anthropic (not named in the abstract). arXiv v1 2026-09-09. The author lists a Harvard doctorate and no institutional affiliation, hence the preliminary grade; the ground truth is enumerable, which is the study's strength.
Sources
Topics
Authors
Hsuan Lo
Tags
Cite This
APA
Hsuan Lo. (2026). Following the Preference, Missing the Optimum: Compliance Without Optimization in AI Housing Recommendation. arXiv (independent researcher, Harvard University alumnus). https://arxiv.org/abs/2609.10856