CCpdf: Building a High Quality Corpus for Visually Rich Documents from Web Crawl Data

Turski, Michał; Stanisławek, Tomasz; Kaczmarek, Karol; Dyda, Paweł; Graliński, Filip

Computer Science > Computation and Language

arXiv:2304.14953 (cs)

[Submitted on 28 Apr 2023 (v1), last revised 6 Jun 2023 (this version, v2)]

Title:CCpdf: Building a High Quality Corpus for Visually Rich Documents from Web Crawl Data

Authors:Michał Turski, Tomasz Stanisławek, Karol Kaczmarek, Paweł Dyda, Filip Graliński

View PDF

Abstract:In recent years, the field of document understanding has progressed a lot. A significant part of this progress has been possible thanks to the use of language models pretrained on large amounts of documents. However, pretraining corpora used in the domain of document understanding are single domain, monolingual, or nonpublic. Our goal in this paper is to propose an efficient pipeline for creating a big-scale, diverse, multilingual corpus of PDF files from all over the Internet using Common Crawl, as PDF files are the most canonical types of documents as considered in document understanding. We analysed extensively all of the steps of the pipeline and proposed a solution which is a trade-off between data quality and processing time. We also share a CCpdf corpus in a form or an index of PDF files along with a script for downloading them, which produces a collection useful for language model pretraining. The dataset and tools published with this paper offer researchers the opportunity to develop even better multilingual language models.

Comments:	Accepted at ICDAR 2023
Subjects:	Computation and Language (cs.CL)
Cite as:	arXiv:2304.14953 [cs.CL]
	(or arXiv:2304.14953v2 [cs.CL] for this version)
	https://doi.org/10.48550/arXiv.2304.14953

Submission history

From: Michał Turski [view email]
[v1] Fri, 28 Apr 2023 16:12:18 UTC (2,819 KB)
[v2] Tue, 6 Jun 2023 07:35:17 UTC (2,819 KB)

Computer Science > Computation and Language

Title:CCpdf: Building a High Quality Corpus for Visually Rich Documents from Web Crawl Data

Submission history

Access Paper:

References & Citations

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators

Computer Science > Computation and Language

Title:CCpdf: Building a High Quality Corpus for Visually Rich Documents from Web Crawl Data

Submission history

Access Paper:

References & Citations

BibTeX formatted citation

Bookmark

Bibliographic and Citation Tools

Code, Data and Media Associated with this Article

Demos

Recommenders and Search Tools

arXivLabs: experimental projects with community collaborators