Language Models and Simple, Stupid Bugs

Abstract—With the advent of powerful neural language models, AI-based systems to assist developers in coding tasks are becoming widely available: Copilot is one such system. Copilot uses Codex, a large language model (LLM), to complete code conditioned on a preceding “prompt". Codex, however, is trained on public GitHub repositories, viz., on code that may include bugs and vulnerabilities. Previous studies [1], [2] show Codex reproduces vulnerabilities seen in training. In this study, we examine how prone Codex is to generate an interesting bug category, single statement bugs, commonly referred to as simple, stupid bugs or SStuBs in the MSR community. We find that Codex and similar LLMs do help avoid some SStuBs, but do produce known, verbatim SStuBs as much as 2x as likely thanknown, verbatim correct code. We explore the consequences of the Codex generated SStuBs and propose avoidance strategies that suggest the possibility of reducing the production of known, verbatim SStubs, and increase the possibility of producing known, verbatim fixes.

Related Organizations

University of California, Davis
United States

Keywords

language models, prompting, deep learning, software engineering

7 Research products, page 1 of 1

PySStuBs: Characterizing Single-Statement Bugs in Popular Open-Source Python Projects
2021IsAmongTopNSimilarDocuments
On the Rise and Fall of Simple Stupid Bugs: a Life-Cycle Analysis of SStuBs
2021IsAmongTopNSimilarDocuments
On the Rise and Fall of Simple Stupid Bugs: a Life-Cycle Analysis of SStuBs
2021IsAmongTopNSimilarDocuments
How Effective is Continuous Integration in Indicating Single-Statement Bugs?
2021IsAmongTopNSimilarDocuments
Language Models and Simple, Stupid Bugs
2023HasVersion
Language Models and Simple, Stupid Bugs
2023IsAmongTopNSimilarDocuments
On the Distribution of "Simple Stupid Bugs" in Unit Test Files: An Exploratory Study
2021IsAmongTopNSimilarDocuments

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	0
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Average
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average