
Large language models increasingly serve as both decision-makers and evaluators in alignment pipelines such as reinforcement learning from human feedback (RLHF) and constitutional AI. While automated judging enables rapid iteration, recent evidence from Anthropic’s stress-testing work demonstrates that frontier-model evaluators disagree roughly 30% of the time when scoring moral trade-offs, raising questions about what these systems actually reward. This preprint investigates the mechanisms driving evaluator disagreement using a controlled Phase 1 dataset of 1,500 moral-dilemma responses generated by five local instruction-tuned models and evaluated by two LLM judges. Our primary finding is a balance penalty: responses that acknowledge competing values receive substantially lower alignment scores than responses that commit to a single value, even when topicality and temperature are held constant. We quantify this effect, analyze its interaction with framing and sampling temperature, and discuss implications for model evaluation and training.
AI alignment, moral resoning, LLM evaluation, RLHF
AI alignment, moral resoning, LLM evaluation, RLHF
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
