Datasets: A Community Library for Natural Language Processing

Quentin Lhoest; Albert Villanova del Moral; Yacine Jernite; Abhishek Thakur; Patrick von Platen; Suraj Patil; Julien Chaumond; Mariama Drame; Julien Plu; Lewis Tunstall; Joe Davison; Mario Šaško; Gunjan Chhablani; Bhavitvya Malik; Simon Brandeis; Teven Le Scao; Victor Sanh; Canwen Xu; Nicolas Patry; Angelina Mcmillan-major; Philipp Schmid; Sylvain Gugger; Clément Delangue; Théo Matussière; Lysandre Debut; Stas Bekman; Pierric Cistac; Thibault Goehringer; Victor Mustar; François Lagunas; Alexander M. Rush; Thomas Wolf

doi:10.18653/v1/2021.emnlp-demo.21

Datasets: A Community Library for Natural Language Processing

Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, Angelina McMillan-Major, Philipp Schmid, Sylvain Gugger, Clément Delangue, Théo Matussière, Lysandre Debut, Stas Bekman, Pierric Cistac, Thibault Goehringer, Victor Mustar, François Lagunas, Alexander Rush, Thomas Wolf

Abstract

The scale, variety, and quantity of publicly-available NLP datasets has grown rapidly as researchers propose new tasks, larger models, and novel benchmarks. Datasets is a community library for contemporary NLP designed to support this ecosystem. Datasets aims to standardize end-user interfaces, versioning, and documentation, while providing a lightweight front-end that behaves similarly for small datasets as for internet-scale corpora. The design of the library incorporates a distributed, community-driven approach to adding datasets and documenting usage. After a year of development, the library now includes more than 650 unique datasets, has more than 250 contributors, and has helped support a variety of novel cross-dataset research projects and shared tasks. The library is available at https://github.com/huggingface/datasets.

Anthology ID:: 2021.emnlp-demo.21
Volume:: Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations
Month:: November
Year:: 2021
Address:: Online and Punta Cana, Dominican Republic
Editors:: Heike Adel, Shuming Shi
Venue:: EMNLP
SIG:
Publisher:: Association for Computational Linguistics
Note:
Pages:: 175–184
Language:
URL:: https://aclanthology.org/2021.emnlp-demo.21
DOI:: 10.18653/v1/2021.emnlp-demo.21
Bibkey:
Cite (ACL):: Quentin Lhoest, Albert Villanova del Moral, Yacine Jernite, Abhishek Thakur, Patrick von Platen, Suraj Patil, Julien Chaumond, Mariama Drame, Julien Plu, Lewis Tunstall, Joe Davison, Mario Šaško, Gunjan Chhablani, Bhavitvya Malik, Simon Brandeis, Teven Le Scao, Victor Sanh, Canwen Xu, Nicolas Patry, et al.. 2021. Datasets: A Community Library for Natural Language Processing. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, pages 175–184, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.
Cite (Informal):: Datasets: A Community Library for Natural Language Processing (Lhoest et al., EMNLP 2021)
Copy Citation:
PDF:: https://aclanthology.org/2021.emnlp-demo.21.pdf
Video:: https://aclanthology.org/2021.emnlp-demo.21.mp4
Code: huggingface/datasets
Data: GLUE, SQuAD, Universal Dependencies

PDF Cite Search Code Video