Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Preprint . 2026
License: CC BY
Data sources: ZENODO
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
ZENODO
Preprint . 2026
License: CC BY
Data sources: Datacite
versions View all 2 versions
addClaim

Synthesis: A Federated Capability Ecosystem for Safe AI Self-Extension Through Test-Driven Development and Graduated Trust

Authors: Maio, Anthony D.;

Synthesis: A Federated Capability Ecosystem for Safe AI Self-Extension Through Test-Driven Development and Graduated Trust

Abstract

As AI agents become more capable, there is increasing interest in systems that can extend their own capabilities through code generation and tool use. However, naive code generation approaches produce unreliable outputs that may fail silently, introduce security vulnerabilities, or behave unexpectedly—a phenomenon well-documented in evaluations of large language model code generation. We present Synthesis, a federated capability ecosystem for safe AI self-extension that addresses these challenges through three integrated mechanisms: (1) Test-Driven Synthesis, where comprehensive test suites are generated before implementation code and capabilities must pass all tests before deployment; (2) Graduated Trust, where newly synthesized capabilities start in maximally restricted sandboxes and progressively earn privileges through demonstrated reliability across quantified thresholds; and (3) Composition Over Creation, where the system exhaustively searches a shared Live Exchange and attempts to compose existing verified capabilities before synthesizing new code, creating network effects that benefit all participating agents. The architecture includes a trust bootstrapping protocol that solves the cold-start problem for new deployments through founding validators and pre-verified seed capabilities. Our empirical measurements show realistic success rates (50–70% one-shot, 70–85% after iterative refinement) while maintaining honest metrics about system limitations. Synthesis provides a foundation for AI systems that can safely adapt to new requirements without compromising reliability, security, or auditability.

Keywords

self-extension, artificial intelligence, test-driven-development

  • BIP!
    Impact byBIP!
    selected citations
    These citations are derived from selected sources.
    This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    0
    popularity
    This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
    Average
    influence
    This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
    Average
    impulse
    This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
    Average
Powered by OpenAIRE graph
Found an issue? Give us feedback
selected citations
These citations are derived from selected sources.
This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Citations provided by BIP!
popularity
This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.
BIP!Popularity provided by BIP!
influence
This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).
BIP!Influence provided by BIP!
impulse
This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.
BIP!Impulse provided by BIP!
0
Average
Average
Average
Green