
Formal verification of neural networks has matured into a real research area. Tools such as alpha-beta-CROWN, ERAN, andMarabou now reliably verify L-infinity robustness properties of ReLU networks at the ResNet scale, with the annual VNNCOMP competition documenting steady year-over-year progress. The dominant framing of what remains undone — "scaleverification to bigger models" — captures only one of three distinct gaps that separate current capability from usefulguarantees on frontier AI systems. This paper makes three claims. First, the scale gap (verifiers handle networks with millionsbut not billions of parameters), the architecture gap (standard techniques handle ReLU well but degrade sharply ontransformers with softmax, attention, and layer normalization), and the specification gap (we lack formal definitions of "safeLLM output" comparable to L-infinity robustness for image classifiers) are independent obstacles that need independentattention. Second, the specification gap is the most underweighted of the three and may be the binding constraint: even withverifiers that scaled to trillion-parameter models, we would not have formal specifications of harmlessness, honesty, or nondeception to verify against. Third, the most productive near-term frontier is not direct verification of LLM weights butverification of the systems that contain LLMs — LLM-as-policy verification, runtime monitors, agent-action guardrails. Wepropose a research agenda that reorients around the specification problem and the system-level frontier rather thancontinuing to pursue parameter-count scaling as the central goal.
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
