Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
versions View all 4 versions
addClaim

TELOS Academic Paper

Authors: Brunner, Jeffrey;

TELOS Academic Paper

Abstract

TELOS is a runtime AI governance framework achieving 0 observed attack successes cross 2,550 adversarial attacks (95% CI upper bound ~0.15%). While current AI safety systems accept violation rates of 3.7% to 43.9% as unavoidable, TELOS demonstrates that mathematical enforcement of constitutional boundaries can provide substantially stronger defense under black-box threat models. Key Innovation: Primacy Attractors (PAs) - mathematical embedding-space representations of user purpose that enable continuous fidelity measurement with graduated intervention. Architecture: - Layer 1: Baseline similarity pre-filter (hard boundary) - Layer 2: Basin membership detection (purpose drift) - Three-Tier Governance: PA → RAG Context → Human Escalation Validation Results (0/2,550 observed attack successes): - AILuminate Standard Benchmark: 0/1,200 observed (NIST AI RMF aligned, 12 harm categories) - HarmBench: 0/400 observed (Center for AI Safety benchmark) - MedSafetyBench: 0/900 observed (NeurIPS 2024 healthcare attacks) - SB 243-Aligned Evaluation: 0/50 observed (child safety categories) - Combined 95% CI upper bound: ~0.15% Technical Framework: - Lyapunov stability for convergence guarantees - Statistical Process Control for drift detection - Proportional-Integral control for intervention calibration - JSONL governance traces for EU AI Act Article 72 compliance Regulatory Alignment: - EU AI Act (Article 72 post-market monitoring) - NIST AI Risk Management Framework - FDA Quality System Regulation methodology Resources: - Primary validation dataset: Zenodo 18370659 - Governance benchmark: Zenodo 18009153 - SB 243-aligned evaluation: Zenodo 18370504 This paper presents the mathematical foundations, implementation architecture, and empirical validation of TELOS as governance infrastructure for conversational and agentic AI systems.

Keywords

HIPAA, SB 243, proportional control, runtime governance, EU AI Act, AI governance, AI Alignment, AB 3030, child safety, AI safety, statistical process control, SB 53, DMAIC, constitutional AI

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green