Powered by OpenAIRE graph
Found an issue? Give us feedback
ZENODOarrow_drop_down
ZENODO
Report . 2026
License: CC BY SA
Data sources: Datacite
ZENODO
Report . 2026
License: CC BY SA
Data sources: Datacite
versions View all 2 versions
addClaim

A Structured Comparative Evaluation of Claude Code and Uxia for UX Navigation and Usability Analysis

A comparative evaluation of synthetic user testing and browser-agent-based usability assessment across three live product flows
Authors: Kentish, Peter; Pique Leal, Pablo; Marín Solà, Francesc Xavier; Diaz-Roig, Borja; Perdiguer Eced, Victor;

A Structured Comparative Evaluation of Claude Code and Uxia for UX Navigation and Usability Analysis

Abstract

A side-by-side comparison of three live-product usability tests shows that a purpose-built synthetic testing platform like Uxia produces fundamentally more reliable and richer results than a general-purpose browser agent (Claude browser extension for this experiment) used as a stand-in usability tester. Reliability. Uxia completed all 3 missions, with every tester finishing end-to-end. Claude fully completed only 1 of 3 missions on the first execution due to abnormal stops: one run was blocked at a domain boundary, another was locked out of the product entirely, and both required reruns with manually adjusted permissions. Behavioral fidelity. All Uxia testers behave within the boundaries of human limits. When blocked, the Claude agent went beyond what a human would do, inspecting the DOM, firing React handlers programmatically, and invoking JavaScript tools, then reported “success” on a flow no human could complete. Insight quality. Uxia insights are grounded in what multiple testers actually experienced, each scored for severity, backed by cross-tester frequency and evidence quotes, and paired with a fix suggestion. Claude’s single-run insights are articulate but unscored, single-perspective, and in the worst case describe flows the agent never actually touched. Coverage. Every Uxia test ran with 5 differentiated synthetic testers (10 by default on standard platform plans); Claude produced one run, one persona, no frequency data. The highest-impact findings, like a resume generator leaking a different person’s identity, were only discoverable because Uxia testers carried realistic personas and files through the full flow. Capabilities. Misclick tracking, heatmaps, journey diagrams or UXIA-Q scoring are available in Uxia and have no equivalent in the Claude extension. Claude’s value. The extension narrates think-aloud friction articulately on pages it can freely access, and it handles login credentials correctly (masked and redacted in logs). Bottom line: a browser agent can describe a page; a synthetic testing platform tells you what real users will actually experience, reliably, repeatably, and at scale. Keywords: usability testing; synthetic users; browser agents; AI agents; UX research; behavioral fidelity

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Upload OA version
Are you the author of this publication? Upload your Open Access version to Zenodo!
It’s fast and easy, just two clicks!