Name: CoverUp: Effective High Coverage Test Generation for Python
Keywords: Software Engineering (cs.SE), FOS: Computer and information sciences, Computer Science - Software Engineering, Computer Science - Machine Learning, Computer Science - Programming Languages, Artificial Intelligence (cs.AI), Computer Science - Artificial Intelligence, D.2.5, I.2.0; D.2.5, I.2.0

descriptionPublicationkeyboard_double_arrow_right Article , Preprint 19 Jun 2025Embargo end date: 01 Jan 2024 English Publisher:Association for Computing Machinery (ACM)Journal:Proceedings of the ACM on Software Engineering, volume 2, pages 2,897-2,919 (eissn: 2994-970X,

Authors: Juan Altmayer Pizzorno; Emery D. Berger;

doi: 10.1145/3729398 , 10.48550/arxiv.2403.16218

arXiv: 2403.16218

CoverUp: Effective High Coverage Test Generation for Python

- Summary
- Subjects
- Related research
  (3)
- Metrics

Abstract

Testing is an essential part of software development. Test generation tools attempt to automate the otherwise labor-intensive task of test creation, but generating high-coverage tests remains challenging. This paper proposes CoverUp, a novel approach to driving the generation of high-coverage Python regression tests. CoverUp combines coverage analysis, code context, and feedback in prompts that iteratively guide the LLM to generate tests that improve line and branch coverage. We evaluate our prototype CoverUp implementation across a benchmark of challenging code derived from open-source Python projects and show that CoverUp substantially improves on the state of the art. Compared to CodaMosa, a hybrid search/LLM-based test generator, CoverUp achieves a per-module median line+branch coverage of 80% (vs. 47%). Compared to MuTAP, a mutation- and LLM-based test generator, CoverUp achieves an overall line+branch coverage of 89% (vs. 77%). We also demonstrate that CoverUp’s performance stems not only from the LLM used but from the combined effectiveness of its components.

Related Organizations

University of Massachusetts System
United States
Amazon (United States)
United States
University of Massachusetts Amherst
United States

Keywords

Software Engineering (cs.SE), FOS: Computer and information sciences, Computer Science - Software Engineering, Computer Science - Machine Learning, Computer Science - Programming Languages, Artificial Intelligence (cs.AI), Computer Science - Artificial Intelligence, D.2.5, I.2.0; D.2.5, I.2.0, Machine Learning (cs.LG), Programming Languages (cs.PL)

3 Research products, page 1 of 1

codamosa-dataset software on GitHub
IsRelatedTo
codamosa software on GitHub
IsRelatedTo
coveragepy software on GitHub
IsRelatedTo

Impact byBIP!

	selected citations These citations are derived from selected sources. This is an alternative to the "Influence" indicator, which also reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	4
	popularity This indicator reflects the "current" impact/attention (the "hype") of an article in the research community at large, based on the underlying citation network.	Top 10%
	influence This indicator reflects the overall/total impact of an article in the research community at large, based on the underlying citation network (diachronically).	Average
	impulse This indicator reflects the initial momentum of an article directly after its publication, based on the underlying citation network.	Average

Found an issue? Give us feedback

Top 10%

Average

Green

CoverUp: Effective High Coverage Test Generation for Python

CoverUp: Effective High Coverage Test Generation for Python

3 Research products, page 1 of 1

codamosa-dataset software on GitHub

codamosa software on GitHub

coveragepy software on GitHub