Tools & document processing
Documents in, structure out: parsing tooling measured under real workloads, with benchmarks kept distinct from parsers.
Shared lane
This lane is carried by both hubs of this survey.
Entries are repositories — libraries, benchmarks, specifications, whatever a ledger admitted as evidence. Some of it is polished; most of it is parts.
Tasks in this category
- Parse PDFs & documents 12 of 86 entries
- Multi-format document intelligence Planned 0 of 86 entries
Parse PDFs & documents
Campaigns placed on this task
Multiple runs of this task exist; each is its own sealed record — no supersession is implied.
-
2026-07-28__pdf-parsers
run type (the campaign's own designation) high-recall-map 12 of 86 entries ingested
- Ledger rows
- 16
- Screening rows
- 70
- Repos in ledger
- 13
-
2026-08-02__pdf-parsers-practical-evaluation
run type (the campaign's own designation) high-recall-map Not yet ingested
- Ledger rows
- 40
- Screening rows
- 116
- Repos in ledger
- 28
Entries — 12 of 86 on this hub
-
Tools & document processing Live
allenai/olmocr
Parse PDFs & documents
Editorial draft OCR pipeline for scans and multicolumn documents with natural reading-order output; the ledger tempers it with local-setup caveats.
Forks 1594 Language Python Stars 19284
Snapshot · retrieved UTC
-
Tools & document processing Live
datalab-to/chandra
Parse PDFs & documents
Editorial draft An OCR model aimed at regulated documents, with open local weights; the ledger's Apple-feasibility caveats travel with it.
Forks 1226 Language Python Stars 12012
Snapshot · retrieved UTC
-
Tools & document processing Live
datalab-to/marker
Parse PDFs & documents
Editorial draft Document conversion with hybrid native and OCR routing, JSON polygons, and Apple CPU/MPS defaults.
Forks 2749 Language Python Stars 38588
Snapshot · retrieved UTC
-
Tools & document processing Live
datalab-to/surya
Parse PDFs & documents
Editorial draft OCR and layout components producing line boxes, polygons and reading order; cited four times in the parsing ledger.
Forks 1524 Language Python Stars 21234
Snapshot · retrieved UTC
-
Tools & document processing Live
docling-project/docling
Parse PDFs & documents
Editorial draft A permissively licensed document-parsing SDK with provenance boxes and offline operation.
Forks 4583 Language Python Stars 64459
Snapshot · retrieved UTC
-
Tools & document processing Live
opendatalab/MinerU
Parse PDFs & documents
Editorial draft Document parser with coordinate-bearing JSON output and an official Apple path; recorded as the incumbent in this campaign's ledger.
Forks 6498 Language Python Stars 77208
Snapshot · retrieved UTC
-
Tools & document processing Live
opendatalab/OmniDocBench
Parse PDFs & documents
Editorial draft The document-parsing benchmark this lane measures against — not a runnable parser, and the entry says so. Seven ledger rows cite it.
Forks 191 Language Python Stars 1959
Snapshot · retrieved UTC
-
Tools & document processing Live
opendataloader-project/opendataloader-pdf
Parse PDFs & documents
Editorial draft A Java PDF loader with accessibility-oriented JSON and reading order; the ledger records empty-output cases on its own project pages.
Forks 2700 Language Java Stars 28299
Snapshot · retrieved UTC
-
Tools & document processing Live
PaddlePaddle/Paddle
Parse PDFs & documents
Editorial draft The deep-learning framework underlying PaddleOCR; kept as lineage context for the OCR row.
Forks 6014 Language C++ Stars 24045
Snapshot · retrieved UTC
-
Tools & document processing Live
PaddlePaddle/PaddleOCR
Parse PDFs & documents
Editorial draft OCR pipeline with structured coordinates and reading order, under Apache-2.0; cited three times in the ledger.
Forks 11162 Language Python Stars 87297
Snapshot · retrieved UTC
-
Tools & document processing Live
pymupdf/pymupdf4llm
Parse PDFs & documents
Editorial draft Lightweight extraction of native PDF geometry and reading order for model pipelines; version-mapping caveats are recorded.
Forks 240 Language Python Stars 2078
Snapshot · retrieved UTC
-
Tools & document processing Live
run-llama/liteparse
Parse PDFs & documents
Editorial draft A small Apache-2.0 geometry and OCR baseline in Rust; its own documentation bounds it away from dense multicolumn work.
Forks 823 Language Rust Stars 11993
Snapshot · retrieved UTC
Tasks with no entry today
1 of 2 tasks in this category carry no entry today. Each is listed with its state: ingested with nothing admitted, on the record and not yet ingested, or planned and not yet run.
- Multi-format document intelligence Planned No campaign has run yet. Grounded in crosswalk row backlog 20 (internal crosswalk id).