Powered by OpenAIRE graph
Found an issue? Give us feedback
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/ ZENODOarrow_drop_down
image/svg+xml art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos Open Access logo, converted into svg, designed by PLoS. This version with transparent background. http://commons.wikimedia.org/wiki/File:Open_Access_logo_PLoS_white.svg art designer at PLoS, modified by Wikipedia users Nina, Beao, JakobVoss, and AnonMoos http://www.plos.org/
ZENODO
Report
Data sources: ZENODO
addClaim

When Agents Control the Kernel: A JEPA World Model Safety Gate with Empirical False-Negative Decomposition

Authors: Kulshrestha, Vyom;

When Agents Control the Kernel: A JEPA World Model Safety Gate with Empirical False-Negative Decomposition

Abstract

Autonomous agents that modify files, processes, services, and network state require controls where proposed actions become operating-system effects. FerrumOS normalizes provider output into 41 canonical actions and evaluates independent deterministic and action-conditioned JEPA transition branches before capability-gated syscalls. The original study accounts for 13,697 transitions from 3,639 QEMU episodes and reports 81.4% balanced accuracy on a 500-episode authored fixture; a per-action mean is nearly tied, five complete pipelines vary substantially, and all 52 hybrid false negatives are decomposed. Version 1.2 adds a registered v3-v3.4 lineage with four retained failed iterations and a source-held-out final catalog denied to selection scripts. The frozen candidate improves untouched-test rollout error at H=1, H=3, and H=5 by 9.54%, 2.36%, and 0.92%. On a 512-episode source-held-out deterministic simulator, request-bounded v3.4 policy records 256 TP, 0 FN, 256 TN, and 0 FP; the paired source-stratified 95% balanced-accuracy improvement interval over deployed runtime-v2 is +46.35 to +48.82 percentage points. Rules-only and rules+JEPA are identical on the final catalog, so the safety result is attributed to deterministic authority, not incremental learned safety value. Claim boundary: the safety fixtures are authored software simulations, not production incident replays, independent human adjudication, formal verification, or complete safety evidence. The v3.4 candidate is archived but not deployed; runtime authority, timing, and simulator-to-runtime resource effects remain pending.

Powered by OpenAIRE graph
Found an issue? Give us feedback