Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint
Data sources: ZENODO
addClaim

The Capability Induction Framework: A Systems Approach to LLM Development

Authors: Sea, Beth;

The Capability Induction Framework: A Systems Approach to LLM Development

Abstract

Current large language model alignment pipelines conflate factual accuracy, social calibration, safety, and tone into a single preference-based training signal (RLHF), producing sycophancy as an inevitable structural artifact rather than a tunable parameter. This paper proposes the Capability Induction Framework (CIF), a three-phase developmental training pipeline that separates these capabilities into sequential stages with stage-appropriate evaluation. Phase 1 (Epistemological Grounding) establishes a factual and cultural baseline through curriculum-sequenced pedagogical materials, verified through automated assessment and paired-source discrimination stress testing. Phase 2 (Relational Generalization) trains perspective-taking through expert-supervised conversation and audited summary production. Phase 3 (Dimension-Restricted Calibration) limits RLHF to delivery calibration only, using a narrowed preference signal that cannot corrupt the capabilities established in earlier phases. The framework replaces imposed behavioral identity with emergent disposition, eliminates the conflated training signal that produces sycophancy, and introduces a validated challenge reward mechanism that trains accurate authority-challenging behavior — the structural inverse of sycophantic compliance. The individual mechanisms at each phase are established techniques; the innovation is the separation and sequencing. The economic case is architectural: front-loaded developmental investment eliminates ongoing remediation costs, and a model that is right on the first response reduces the total tokens-to-accurate-outcome ratio even when individual responses cost more to generate. Five falsifiable predictions are specified for proof-of-concept validation at open-weight model scales.

Powered by OpenAIRE graph
Found an issue? Give us feedback