Powered by OpenAIRE graph
Found an issue? Give us feedback
ZENODOarrow_drop_down
ZENODO
Dataset . 2026
License: CC BY
Data sources: Datacite
ZENODO
Dataset . 2026
License: CC BY
Data sources: Datacite
addClaim

AI Recommendation Calibration Study 2026: aggregate results for UK local business recommendations across four large language models

Authors: BILLINGHAM, IAN;

AI Recommendation Calibration Study 2026: aggregate results for UK local business recommendations across four large language models

Abstract

What this dataset is. Aggregate results from a controlled calibration study measuring how four large language model (LLM) products recommend local UK businesses when asked category-level questions such as "best dentist in Bristol" or "emergency plumber Manchester". Between May 2026 and July 2026, 4,800 model responses were collected across a fully crossed factorial design: two business categories × three UK cities × four AI platforms × four location-phrasing modes × five prompt intents per category × ten repeats per combination. Headline findings. 3,583 distinct businesses were recommended across the study. 64.1% of distinct businesses appeared in only a single response. Cross-platform Jaccard similarity of recommendation sets remained below 0.15 across all tested cells, indicating that AI-generated recommendations are highly fragmented rather than converging on a common shortlist. A cross-model reviewer agreement check on 100 random runs (Claude Haiku 4.5 vs. GPT-4o-mini) produced Jaccard 0.61 on entity extraction and 0.63 on signal classification, with per-category Cohen's kappa reported in the methodology PDF. Platforms tested. OpenAI GPT-5.3 Chat, Anthropic Claude Sonnet 4.6, Google Gemini 3 Flash, Perplexity Sonar (search-enabled). All accessed via commercial API endpoints with provider pinning and fallbacks disabled. Files in this deposit. summary_stats.json — study manifest with counts, categories, cities, platforms, headline numbers and aggregation rules. recommendations_by_category.csv — response volume, distinct businesses recommended, average recommendations per response, positive-sentiment share, per category × city × platform. concentration_by_category.csv — top-1 share, top-3 share, HHI, share of businesses recommended only once, per category × city × platform. cross_platform_overlap.csv — pairwise Jaccard and overlap coefficient across platforms, per category × city. signal_frequency.csv — frequency of each reasoning signal across a 16-category taxonomy, per category × platform. methodology.pdf — full methodology, study design, statistical measures, cross-model reviewer agreement, and limitations. Privacy. All published files are aggregates. No business names, competitor names, cited sources, verbatim AI response text, postcodes, neighbourhoods, or evidence quotes are included. Cells backed by fewer than five underlying observations are suppressed to reduce the risk of indirectly identifying individual businesses. Suggested citation. Billingham, I. (2026). AI Recommendation Calibration Study 2026: aggregate results for UK local business recommendations across four large language models [Data set]. Zenodo. Contact. hello@ai-mention.co.uk

Keywords

LLM, AI recommendations, calibration study, recommendation systems, local search, local business, GPT, generative engine optimization, Perplexity, GEO, Claude, United Kingdom, large language models, AI visibility, UK, Gemini

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average