Automated Unit Test Improvement using Large Language Models at Meta
Nadia Alshahwan https://orcid.org/0009-0009-4763-0396 , Jubin Chheda https://orcid.org/0009-0005-0311-7890 , Anastasia Finegenova https://orcid.org/0009-0005-5824-5179 , Beliz Gokkaya https://orcid.org/0009-0003-0197-6806 , Mark Harman https://orcid.org/0000-0002-5864-4488 , Inna Harper https://orcid.org/0009-0008-9359-0949 , Alexandru Marginean https://orcid.org/0009-0001-5311-762X , Shubho Sengupta https://orcid.org/0009-0007-4204-5185 and Eddy Wang https://orcid.org/0009-0009-8825-6986 Meta Platforms Inc.,1 Hacker WayMenlo ParkCaliforniaUSA
Abstract
This paper describes Meta’s TestGen-LLM tool, which uses LLMs to automatically improve existing human-written tests. TestGen-LLM verifies that its generated test classes successfully clear a set of filters that assure measurable improvement over the original test suite, thereby eliminating problems due to LLM hallucination. We describe the deployment of TestGen-LLM at Meta test-a-thons for the Instagram and Facebook platforms. In an evaluation on Reels and Stories products for Instagram, 75% of TestGen-LLM’s test cases built correctly, 57% passed reliably, and 25% increased coverage. During Meta’s Instagram and Facebook test-a-thons, it improved 11.5% of all classes to which it was applied, with 73% of its recommendations being accepted for production deployment by Meta software engineers. We believe this is the first report on industrial scale deployment of LLM-generated code backed by such assurances of code improvement.
原文 arXiv:2402.09171;中英对照 + 大白话阅读 https://aha.fim.ai/paper/2402.09171v1