Powered by OpenAIRE graph
Found an issue? Give us feedback
ZENODOarrow_drop_down
ZENODO
Dataset . 2026
License: CC BY
Data sources: Datacite
ZENODO
Dataset . 2026
License: CC BY
Data sources: Datacite
ZENODO
Dataset . 2026
License: CC BY
Data sources: Datacite
ZENODO
Dataset . 2026
License: CC BY
Data sources: Datacite
addClaim

Auditing Alignment Controllability in LLMs via Political Axes: Reproducibility Package (Code and Data)

Authors: Bućan, Bartol; Sočec, Nikola; Isufi, Sarah; Granić, Morena; Hobor, Luka; Krajna, Agneza; Kovac, Mihael; +1 Authors

Auditing Alignment Controllability in LLMs via Political Axes: Reproducibility Package (Code and Data)

Abstract

Code and data reproducibility package for the AIES 2026 paper "Auditing Alignment Controllability in LLMs via Political Axes." Most political audits of an LLM report a single point on a political compass. We argue that resting point barely matters for a deployed system; what matters is the reachable area, how far and in which directions a model's political answers can be steered by the system prompt. This package contains the raw model responses, the collection and processing pipeline, and the analysis scripts that reproduce every number and figure in the paper. Core study: 7 frontier models (GPT-5, Claude Sonnet 4.5, Grok-4.3, Gemini 2.5 Flash Lite, DeepSeek-Chat v3.1, Kimi K2, Qwen3.6 Max Preview) × 13 system-prompt contexts × 10 replicates × 70 Political Compass items = 63,700 responses, plus saturation, R3 ablation, and two compound-quadrant pilot lanes. Reproducibility: processed data regenerates byte-identical from the raw responses; paper_claim_audit.py confirms 51/51 paper claims hold; continuous integration re-verifies on every push. Licensing: code is MIT, data is CC-BY-4.0. See the repository for full details and the reproduction recipe.

Keywords

political compass, AI auditing, steerability, LLM evaluation, large language models, alignment, controllability, reproducibility

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average