Measuring Coding Challenge Competence With APPS

Hendrycks, Dan; Basart, Steven; Kadavath, Saurav; Mazeika, Mantas; Arora, Akul; Guo, Ethan; Burns, Collin; Puranik, Samir; He, Horace; Song, Dawn; Steinhardt, Jacob

Computer Science > Software Engineering

arXiv:2105.09938 (cs)

[Submitted on 20 May 2021 (v1), last revised 8 Nov 2021 (this version, v3)]

Title:Measuring Coding Challenge Competence With APPS

Authors:Dan Hendrycks, Steven Basart, Saurav Kadavath, Mantas Mazeika, Akul Arora, Ethan Guo, Collin Burns, Samir Puranik, Horace He, Dawn Song, Jacob Steinhardt

View PDF

Abstract:While programming is one of the most broadly applicable skills in modern society, modern machine learning models still cannot code solutions to basic problems. Despite its importance, there has been surprisingly little work on evaluating code generation, and it can be difficult to accurately assess code generation performance rigorously. To meet this challenge, we introduce APPS, a benchmark for code generation. Unlike prior work in more restricted settings, our benchmark measures the ability of models to take an arbitrary natural language specification and generate satisfactory Python code. Similar to how companies assess candidate software developers, we then evaluate models by checking their generated code on test cases. Our benchmark includes 10,000 problems, which range from having simple one-line solutions to being substantial algorithmic challenges. We fine-tune large language models on both GitHub and our training set, and we find that the prevalence of syntax errors is decreasing exponentially as models improve. Recent models such as GPT-Neo can pass approximately 20% of the test cases of introductory problems, so we find that machine learning models are now beginning to learn how to code. As the social significance of automatic code generation increases over the coming years, our benchmark can provide an important measure for tracking advancements.

Comments:	NeurIPS 2021. Code and the APPS dataset is available at this https URL
Subjects:	Software Engineering (cs.SE); Computation and Language (cs.CL); Machine Learning (cs.LG)
Cite as:	arXiv:2105.09938 [cs.SE]
	(or arXiv:2105.09938v3 [cs.SE] for this version)
	https://doi.org/10.48550/arXiv.2105.09938

Submission history

From: Dan Hendrycks [view email]
[v1] Thu, 20 May 2021 17:58:42 UTC (539 KB)
[v2] Thu, 27 May 2021 19:41:58 UTC (540 KB)
[v3] Mon, 8 Nov 2021 21:16:44 UTC (548 KB)

Computer Science > Software Engineering

Title:Measuring Coding Challenge Competence With APPS

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Software Engineering

Title:Measuring Coding Challenge Competence With APPS

Submission history

Access Paper:

References & Citations

DBLP - CS Bibliography

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators