No AI summary available for this article.
Why It Matters
Transforming scholarly PDFs into machine-readable fulltext remains a bottleneck for large-scale information systems.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Transforming scholarly PDFs into machine-readable fulltext remains a bottleneck for large-scale information systems. Recent vision-based parsers improve accuracy, but need GPUs and may introduce noise into the extracted text. GROBID, a modular font-stream parser running on CPU, is the de-facto standard for structuring scientific articles and underpins several of the largest open scholarly corpora. We pair it with a lightweight CPU detector localising figure, table, and paratext (header, footer, page number) regions, encoded as typed-area masks whose tokens are routed to GROBID's specialised mo...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.26381v1 · Indexed about 2 hours ago