简介
We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for some 20 lan
代表成果
- Scientific articles and book chapters
- Oepen, Stephan; Arefyev, Nikolay; Aulamo, Mikko; Bañón, Marta; Buljan, Maja & Burchell, Laurie V. [Show all 29 contributors for this article] (2026). HPLT 3.0: Very Large-Scale Multilingual Resources for LLMs and MT. Mono- and Bi-lingual Data, Multilingual Evaluation, and Pre-Trained Models. In Piperidis,, Stelios; Bel, Núria; Heuvel, Henk van den; Ide, Nancy; Krek, Simon & Toral, Antonio (Ed.), Proceedings of the Fifteenth Language Resources and Evaluation Conference (LREC 2026). European Language Resources Association. ISSN 9782493814494. p. 1409–1434. doi: 10.63317/25xbdofco9od. Full text in Research Archive Show summary We present an ongoing initiative to provide open, very large, high-quality, and richly annotated textual datasets for almost 200 languages. At 30 trillion tokens, this is likely the largest generally available multilingual collection of LLM pre-training data. These datasets are derived from web crawls from different sources and accompanied with a complete, open-source pipeline for document selection from web archives, text extraction from HTML, language identification for noisy texts, exact and near-deduplication, annotation with, among others, register labels, text quality estimates, and personally identifiable information; and final selection and filtering. We report on data quality probes through contrastive and analytical statistics, through manual inspection of samples for some 20 languages, and through end-to-end evaluation of various language model architectures trained on this data. For multilingual LLM evaluation, we provide a comprehensive collection of benchmarks for nine European languages, with special emphasis on natively created tasks, mechanisms to mitigate prompt sensitivity, and refined normalization and aggregation of scores. Additionally, we train and evaluate a family of 57 monolingual encoder–decoder models, as well as about 30 “smallish” monolingual GPT-like reference models. Besides the monolingual data and models, we also present a very large collection of parallel texts automatically mined from this data, together with a novel parallel corpus synthesized via machine translation.
- Fedorova, Mariia; Arefyev, Nikolay; Buljan, Maja; Helcl, Jindřich; Oepen, Stephan & Rønningstad, Egil [Show all 7 contributors for this article] (2026). OpenLID-v3: Improving the Precision of Closely Related Language Identification – An Experience Report. In Scherrer, Yves; Aepli, Noëmi; Blaschke, Verena; Jauhiainen, Tommi; Ljubešić, Nikola; Nakov, Preslav; Tiedemann, Jörg & Zampieri, Marcos (Ed.), Proceedings of the 13th Workshop on NLP for Similar Languages, Varieties and Dialects. Association for Computational Linguistics (ACL). ISSN 9798891763722. p. 275–292. doi: 10.18653/v1/2026.vardial-1.23. Full text in Research Archive Show summary Language identification (LID) is an essential step in building high-quality multilingual datasets from web data. Existing LID tools (such as OpenLID or GlotLID) often struggle to identify closely related languages and to distinguish valid natural language from noise, which contaminates language-specific subsets, especially for low-resource languages. In this work we extend the OpenLID classifier by adding more training data, merging problematic language variant clusters, and introducing a special label for marking noise. We call this extended system OpenLID-v3 and evaluate it against GlotLID on multiple benchmarks. During the development we focus on three groups of closely related languages (Bosnian, Croatian, and Serbian; Romance varieties of Northern Italy and Southern France; and Scandinavian languages) and contribute new evaluation datasets where existing ones are inadequate. We find that ensemble approaches improve precision but also substantially reduce coverage for low-resource languages.
- Buljan, Maja (2023). What quantifying word order freedom can tell us about dependency corpora. In Schneider, Nathan (Eds.), The Seventh International Conference on Dependency Linguistics (Depling, GURT/SyntaxFest 2023) -- Proceedings of the Conference. Association for Computational Linguistics (ACL). ISSN 9781959429326. p. 54–67. Full text in Research Archive Show summary Building upon existing work on word order freedom and syntactic annotation, this paper investigates whether we can differentiate between findings that reveal inherent properties of natural languages and their syntax, and features dependent on annotations used in computing the measures. An existing quantifiable and linguistically interpretable measure of word order freedom in language is applied to take a closer look at the robustness of the basic measure (word order entropy) to variations in dependency corpora used in the analysis. Measures are compared at three levels of generality, applied to corpora annotated according to the Universal Dependencies v1 and v2 annotation guidelines, selecting 31 languages for analysis. Preliminary results show that certain measures, such as subject-object relation order freedom, are sensitive to slight changes in annotation guidelines, while simpler measures are more robust, highlighting aspects of these metrics that should be taken into consideration when using dependency corpora for linguistic analysis and generalisation.
- Buljan, Maja; Nivre, Joakim; Oepen, Stephan & Øvrelid, Lilja (2020). A tale of three parsers: Towards diagnostic evaluation for meaning representation parsing. In Calzolari, Nicoletta (Eds.), Proceedings of the 12th Language Resources and Evaluation Conference. European Language Resources Association. ISSN 9791095546344. p. 1902–1909. Full text in Research Archive Show summary We discuss methodological choices in contrastive and diagnostic evaluation in meaning representation parsing, i.e. mapping from natural language utterances to graph-based encodings of its semantic structure. Drawing inspiration from earlier work in syntactic dependency parsing, we transfer and refine several quantitative diagnosis techniques for use in the context of the 2019 shared task on Meaning Representation Parsing (MRP). As in parsing proper, moving evaluation from simple rooted trees to general graphs brings along its own range of challenges. Specifically, we seek to begin to shed light on relative strenghts and weaknesses in different broad families of parsing techniques. In addition to these theoretical reflections, we conduct a pilot experiment on a selection of top-performing MRP systems and one of the five meaning representation frameworks in the shared task. Empirical results suggest that the proposed methodology can be meaningfully applied to parsing into graph-structured target representations, uncovering hitherto unknown properties of the different systems that can inform future development and cross-fertilization across approaches.
数据校验于 9/6/2026数据来源