
This preprint introduces Alemetria, a research protocol for auditing interpretational behavior in large language models (LLMs). The method shifts evaluation away from direct comparison of outputs and toward validation of interpretational comparability. It introduces a structured validation pipeline (B0-B4), a rule of calculation prohibition, and a distinction between valid and invalid regimes of inference. Alemetria enables the analysis of substantive agreement, semantic divergence, false consensus, interpretational suppression, and instability under prompt variation. The protocol is tested on a pilot corpus of 10 sociological stimuli using GPT, Llama, and YandexGPT. Earlier versions of the method were published as MMST/TVA preprints and focused on divergence metrics (ID, SD) and boundary conditions. The present version integrates these earlier developments into a single operational protocol. Related earlier preprints: 10.5281/zenodo.1881968910.5281/zenodo.1879943110.5281/zenodo.18865186
Настоящий препринт представляет Алеметрию — протокол аудита интерпретационного поведения больших языковых моделей. Метод смещает анализ с прямого сравнения ответов на проверку сопоставимости интерпретаций. Вводится контур B0–B4, включающий тип ответа, валидность, согласие, ограниченное геометрическое описание и устойчивость к изменению формулировки. Ключевой элемент протокола — правило запрета расчёта: при нарушении условий сопоставимости или устойчивости метрики (DLSD, ID, SD) не считаются допустимым результатом. Это позволяет различать согласие, расхождение, ложный консенсус, уклонение, интерпретационное подавление и конфигурационную нестабильность. Данная версия интегрирует ранние работы MMST/TVA в единый операционный протокол. Полные эмпирические таблицы доступны по запросу.
semantic similarity, AI auditing, интерпретационное подавление, LLM Machine, semantic divergence, ложный консенсус, sociological stimuli, алеметрия, interpretational behavior, глобальная семантическая пустота, multi-model triangulation, интерпретационная пустота, Alemetria, prompt sensitivity, TVA, prompt variation, model comparison, интерпретационное поведение, LLM evaluation, false consensus, семантическая дивергенция, interpretational suppression, MMST
semantic similarity, AI auditing, интерпретационное подавление, LLM Machine, semantic divergence, ложный консенсус, sociological stimuli, алеметрия, interpretational behavior, глобальная семантическая пустота, multi-model triangulation, интерпретационная пустота, Alemetria, prompt sensitivity, TVA, prompt variation, model comparison, интерпретационное поведение, LLM evaluation, false consensus, семантическая дивергенция, interpretational suppression, MMST
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
