
AWB is an open benchmark harness that measures AI coding tool + workflow performance on 100 real-repository tasks pinned to commit SHAs. It scores seven sigmoid-normalized dimensions (correctness, cost, speed, code quality, reliability, security, efficiency) and reports a Production Readiness Score plus a Workflow Lift between vanilla and configured tool runs. Trace artifacts use OpenTelemetry GenAI semantic conventions.
If you use AWB in research, please cite this entry.
workflow-evaluation, shipping-discipline, benchmark, codex-cli, ai-coding, codex, llm-evaluation, claude-code
workflow-evaluation, shipping-discipline, benchmark, codex-cli, ai-coding, codex, llm-evaluation, claude-code
| selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | 0 | |
| popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network. | Average | |
| influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically). | Average | |
| impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network. | Average |
