No AI summary available for this article.
Why It Matters
Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Whether an information extraction pipeline should process page images or parsed text depends on the document, and the answer flips across the layout spectrum. We study this trade-off under a constraint that rules out (closed) cloud services: privacy-sensitive documents processed on-premise by small ($\le 8\mathrm{B}$ parameter) text-only and vision--language models, evaluated on both accuracy and energy over a design space spanning input representation, model family, and inference configuration. Benchmarking on the near-plain-text Kleister-NDA contracts and the layout-rich VRDU forms, we find...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.31341v1 · Indexed about 22 hours ago