Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

A Fault-Tolerance Threshold for Gated Agentic Computation: Reliable Long-Horizon Work from Unreliable Executors

Authors: Matijašević, Ivo;

A Fault-Tolerance Threshold for Gated Agentic Computation: Reliable Long-Horizon Work from Unreliable Executors

Abstract

Autonomous agents (whether language-model based, human, or hybrid) commit errors at an approximately constant rate per unit of work — a constant-hazard baseline we adopt, motivated by reported time-horizon fits rather than claimed as a universal law — so the probability that a long task completes without fault decays exponentially in its length. The task length at which an agent succeeds half the time — its horizon — has become a leading proposed metric of agentic capability, and empirically it has grown exponentially over recent years. We ask a different question: given a fixed agent, how far can verification extend its reliable horizon? We model long-horizon work as a sequence of checkpointed segments, each verified by a stack of imperfect gates combined by majority vote, and we prove a fault-tolerance threshold analogous to von Neumann's reliable-computation-from-unreliable-components result and to the quantum threshold theorem. (i) When gates are conditionally independent and better than chance, a stack of O(log T) gates per checkpoint achieves any target reliability over horizon T, yielding a multiplicative time overhead that is likewise O(log T). (ii) When gates share a blind spot — a stealth mass λ_st of faults invisible to the entire gate family — the reliable horizon is capped at H_gated ≈ H_raw / λ_st regardless of gate count. (iii) The cost-optimal checkpoint interval is s* ≈ √(G/λ), a direct analogue of the Young–Daly checkpoint formula with verification cost in place of crash-recovery cost. The binding constraint on autonomous horizon is therefore not raw model capability but the diversity of the verification stack, a quantity partly estimable from a system's own telemetry and targeted audits. We further decompose the stealth mass into a diversity-reducible part and a specification-bound irreducible floor, and we note that a recent large-scale N-version experiment with coding agents motivates the bounded regime (ii) for realistic systems. We numerically illustrate all three results in simulation and specify a falsifiable experiment on agentic task suites.

Keywords

threshold theorem, long-horizon autonomy, fault tolerance, AI agents, verification, N-version programming

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green