# Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

> Research article (Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024) · cited 40× · AI/ML

**Wikidata**: [openalex:W4402671286](https://www.wikidata.org/wiki/openalex:W4402671286)  
**Source**: https://4ort.xyz/entity/dolma-an-open-corpus-of-three-trillion-tokens-for-language-model-pretraining-research
