
AIFaultBench is a benchmark of 770 real-world AI software faults collected from 105 open-source repositories across 76 organizations, spanning traditional machine learning, deep learning, large language model infrastructure, reinforcement learning, agentic AI systems, and AI tooling. Each fault ships with the original GitHub issue report, a minimal reproduction script, a dependency specification, codebase reconstruction and environment setup scripts, reproduction logs, structured metadata, and a reproduction trajectory. 652 of the 770 faults (85%) are verified reproducible; the remainder document the reasons preventing reproduction.
If you use AIFaultBench, please cite it as below.
agentic AI, reinforcement learning, benchmark, machine learning, automated program repair, bug reproduction, deep learning, large language models, fault localization, debugging, software engineering, mining software repositories
agentic AI, reinforcement learning, benchmark, machine learning, automated program repair, bug reproduction, deep learning, large language models, fault localization, debugging, software engineering, mining software repositories
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
