Compare commits

...
Author SHA1 Message Date
Abimael Martell 0b017ee7e6 docs: sync benchmark references 2026-07-31 21:10:59 -06:00
Abimael Martell 1ba770769a docs(readme): refresh remaining parser results 2026-07-31 21:06:50 -06:00
Abimael Martell 377937c7ce docs(readme): include stored parser results 2026-07-31 20:54:13 -06:00
Abimael Martell 7dbf340b07 docs(readme): refresh parser benchmark results 2026-07-31 20:50:50 -06:00
Abimael Martell 7b7960ee73 chore(pypi): bump pdf-inspector to 0.2.6 (#188) 2026-07-31 17:12:34 -06:00
Abimael Martell 7188667045 chore(wasm): bump @firecrawl/pdf-inspector-wasm to 0.1.3 (#190) 2026-07-31 16:43:52 -06:00
Abimael Martell 6ff104409a chore(npm): bump @firecrawl/pdf-inspector to 1.11.2 (#189) 2026-07-31 16:37:50 -06:00
Abimael Martell 5e8f1570f6 chore(crate): bump pdf-inspector to 0.1.7 (#187) 2026-07-31 16:25:57 -06:00
tomsideguide b31e4b1727 fix(extractor): don't flag gid Differences names covered by ToUnicode (#186)
* fix(extractor): don't flag gid Differences names covered by ToUnicode

Pages were marked as having unresolvable gid-encoded fonts whenever any
font's /Differences array used gidNNNN glyph names, and when every page
carried such a font the whole document's markdown was suppressed.
LibreOffice exports do exactly this: subset fonts get /gidNNNN names in
Differences alongside a complete ToUnicode CMap that decodes them, so
ordinary text documents lost their entire markdown output even though
extraction decoded every glyph.

Track the character codes behind the gid names and only flag the font
when its ToUnicode CMap addresses none of them. Partially mapped codes
stay unflagged: an emoji ZWJ sequence maps whole on its first code, and
the remaining component-glyph codes are subset leftovers, not damage.
Fonts without ToUnicode, or whose CMap ignores the gid codes, are
flagged as before, and the downstream garbage/encoding checks still
catch partial breakage.

* fix(extractor): require a usable ToUnicode mapping to clear the gid flag

A mapping to U+FFFD (or an empty string) is rejected by extraction as
an invalid CMap result, so it must not count as decodable when deciding
whether gid-named Differences codes are resolvable.
2026-07-31 15:54:19 -06:00
14 changed files with 200 additions and 55 deletions
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector"
version = "0.1.6"
version = "0.1.7"
edition = "2021"
autobins = false
authors = ["Firecrawl Team"]
+8 -8
View File
@@ -28,17 +28,17 @@ Evaluated on the [opendataloader-bench](https://github.com/opendataloader-projec
| Engine | Overall | Reading Order (NID) | Tables (TEDS) | Headings (MHS) | Speed (200 docs) |
|---|---|---|---|---|---|
| pdf-inspector | **0.875** | **0.915** | **0.814** | 0.788 | **2.8s** |
| liteparse | 0.870 | 0.908 | 0.693 | **0.811** | 13.9s |
| opendataloader | 0.843 | 0.912 | 0.489 | 0.760 | 9.8s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 15.5s |
| markitdown | 0.583 | 0.879 | 0.000 | 0.000 | 6.7s |
| pdf-inspector | **0.875** | **0.915** | **0.814** | 0.788 | **0.470s** |
| liteparse | 0.873 | 0.913 | 0.693 | **0.811** | 0.750s |
| opendataloader | 0.831 | 0.902 | 0.489 | 0.739 | 2.569s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 17.117s |
| markitdown | 0.589 | 0.844 | 0.273 | 0.000 | 16.165s |
Results were refreshed on July 16, 2026, on an Apple M4 Pro. Engine versions were pdf-inspector 0.1.6, LiteParse 2.6.0, OpenDataLoader 2.1.1, PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.4. Speed is the median of three complete corpus runs.
Results were refreshed on July 31, 2026, on an Apple M4 Pro. Engine versions were pdf-inspector 0.2.6, LiteParse 2.10.1, OpenDataLoader 2.2.1, PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.5. Speed is the median of five alternating or rotating complete corpus runs after an excluded warm-up run, with each parser processing documents sequentially in a single process.
For context, engines that use OCR or model-based document parsing (docling, marker, mineru) score 0.83-0.88 overall but take 2-180 minutes on the same corpus — pdf-inspector reaches the top of that range without either, in 2.8 seconds.
The complete parser configuration, per-document predictions, evaluator output, and generated charts are available in the [reproducible results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
**Best fit:** Native-text PDFs where speed, reading order, and table structure matter. pdf-inspector delivered the highest overall, reading-order, and table scores, along with the fastest complete run in this benchmark. That makes it a strong local default for reports, research papers, financial documents, invoices, and legal PDFs that need clean, structured Markdown without adding OCR latency or infrastructure.
**Best fit:** Native-text PDFs where speed, reading order, and table structure matter. In this comparison, pdf-inspector delivered the higher overall, reading-order, and table scores, along with the fastest complete run. That makes it a strong local default for reports, research papers, financial documents, invoices, and legal PDFs that need clean, structured Markdown without adding OCR latency or infrastructure.
Use the [paired benchmark harness](docs/benchmarking.md) to compare two local builds against the exact same corpus and evaluator revision.
+8 -5
View File
@@ -30,11 +30,14 @@ overwrite one another.
## Published comparison protocol
The public benchmark table was refreshed on July 16, 2026, on an Apple M4 Pro
using pdf-inspector 0.1.6, LiteParse 2.6.0, OpenDataLoader 2.1.1,
PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.4. Every engine processed the same 200
PDFs with OCR disabled. Reported speed is the median of three complete corpus
runs; quality scores come from the benchmark evaluator over all 200 outputs.
The public benchmark table was refreshed on July 31, 2026, on an Apple M4 Pro
using pdf-inspector 0.2.6, LiteParse 2.10.1, OpenDataLoader 2.2.1,
PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.5. Every engine processed the same 200
PDFs sequentially in a single process with OCR disabled. Reported speed is the
median of five alternating or rotating complete corpus runs after an excluded
warm-up run; quality scores come from the benchmark evaluator over all 200
outputs. Raw timings, predictions, evaluations, and charts are available in the
[results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
## Optional backend evidence probe
+6 -6
View File
@@ -18,13 +18,13 @@ Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|---|---|---|---|---|---|
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **2.8s** |
| liteparse | 0.870 | 0.908 | 0.693 | **0.811** | 13.9s |
| opendataloader | 0.843 | 0.912 | 0.489 | 0.760 | 9.8s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 15.5s |
| markitdown | 0.583 | 0.879 | 0.000 | 0.000 | 6.7s |
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **0.470s** |
| liteparse | 0.873 | 0.913 | 0.693 | **0.811** | 0.750s |
| opendataloader | 0.831 | 0.902 | 0.489 | 0.739 | 2.569s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 17.117s |
| markitdown | 0.589 | 0.844 | 0.273 | 0.000 | 16.165s |
Refreshed July 16, 2026, on Apple M4 Pro; speed is the median of three complete corpus runs. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark).
Refreshed July 31, 2026, on Apple M4 Pro; speed is the median of five complete corpus runs after an excluded warm-up. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark), with raw timings and artifacts in the [results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
## Install
+6 -6
View File
@@ -18,13 +18,13 @@ Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|---|---|---|---|---|---|
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **2.8s** |
| liteparse | 0.870 | 0.908 | 0.693 | **0.811** | 13.9s |
| opendataloader | 0.843 | 0.912 | 0.489 | 0.760 | 9.8s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 15.5s |
| markitdown | 0.583 | 0.879 | 0.000 | 0.000 | 6.7s |
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **0.470s** |
| liteparse | 0.873 | 0.913 | 0.693 | **0.811** | 0.750s |
| opendataloader | 0.831 | 0.902 | 0.489 | 0.739 | 2.569s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 17.117s |
| markitdown | 0.589 | 0.844 | 0.273 | 0.000 | 16.165s |
Refreshed July 16, 2026, on Apple M4 Pro; speed is the median of three complete corpus runs. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark).
Refreshed July 31, 2026, on Apple M4 Pro; speed is the median of five complete corpus runs after an excluded warm-up. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark), with raw timings and artifacts in the [results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
## Install
+1 -1
View File
@@ -830,7 +830,7 @@ checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe"
[[package]]
name = "pdf-inspector"
version = "0.1.6"
version = "0.1.7"
dependencies = [
"env_logger",
"log",
+6 -6
View File
@@ -18,13 +18,13 @@ Built by [Firecrawl](https://firecrawl.dev) for hybrid OCR pipelines — extract
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|---|---|---|---|---|---|
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **2.8s** |
| liteparse | 0.870 | 0.908 | 0.693 | **0.811** | 13.9s |
| opendataloader | 0.843 | 0.912 | 0.489 | 0.760 | 9.8s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 15.5s |
| markitdown | 0.583 | 0.879 | 0.000 | 0.000 | 6.7s |
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **0.470s** |
| liteparse | 0.873 | 0.913 | 0.693 | **0.811** | 0.750s |
| opendataloader | 0.831 | 0.902 | 0.489 | 0.739 | 2.569s |
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 17.117s |
| markitdown | 0.589 | 0.844 | 0.273 | 0.000 | 16.165s |
Refreshed July 16, 2026, on Apple M4 Pro; speed is the median of three complete corpus runs. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark).
Refreshed July 31, 2026, on Apple M4 Pro; speed is the median of five complete corpus runs after an excluded warm-up. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark), with raw timings and artifacts in the [results branch](https://github.com/firecrawl/opendataloader-bench/tree/abi/pdf-parser-benchmark-results).
## Install
+4 -4
View File
@@ -1,6 +1,6 @@
{
"name": "@firecrawl/pdf-inspector",
"version": "1.11.1",
"version": "1.11.2",
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
"main": "index.js",
"types": "index.d.ts",
@@ -49,8 +49,8 @@
"@napi-rs/cli": "^3.4.1"
},
"optionalDependencies": {
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.11.1",
"@firecrawl/pdf-inspector-darwin-arm64": "1.11.1",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.11.1"
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.11.2",
"@firecrawl/pdf-inspector-darwin-arm64": "1.11.2",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.11.2"
}
}
+1 -1
View File
@@ -6,7 +6,7 @@ build-backend = "maturin"
name = "pdf-inspector"
# Bump this to publish to PyPI — CI publishes automatically when the version
# changes on main (same flow as napi/package.json for npm).
version = "0.2.5"
version = "0.2.6"
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
readme = "docs/python.md"
license = { text = "MIT" }
+1 -1
View File
@@ -162,7 +162,7 @@ pub(crate) fn extract_page_text_items(
let fonts = doc.get_page_fonts(page_id).unwrap_or_default();
// Build font encoding maps from Differences arrays
let (font_encodings, has_gid_fonts) = build_font_encodings(doc, &fonts);
let (font_encodings, has_gid_fonts) = build_font_encodings(doc, &fonts, font_cmaps);
// Build font width info for accurate text positioning
let font_widths = build_font_widths(doc, &fonts);
+154 -12
View File
@@ -497,9 +497,13 @@ pub(crate) fn get_operand_bytes(obj: &Object) -> Option<&[u8]> {
/// Build encoding maps for all fonts on a page.
/// Returns `(encodings, has_gid_fonts)` where `has_gid_fonts` is true when
/// any font uses raw glyph ID names (gidNNNNN) that can't be decoded.
/// Gid names whose codes the font's own ToUnicode CMap maps are decodable
/// and do not set the flag (LibreOffice subsets write /gidNNNN Differences
/// names alongside a complete ToUnicode CMap).
pub(crate) fn build_font_encodings(
doc: &Document,
fonts: &std::collections::BTreeMap<Vec<u8>, &lopdf::Dictionary>,
cmaps: &FontCMaps,
) -> (PageFontEncodings, bool) {
let mut encodings = PageFontEncodings::new();
let mut has_gid_fonts = false;
@@ -508,7 +512,9 @@ pub(crate) fn build_font_encodings(
let resource_name = String::from_utf8_lossy(font_name).to_string();
if let Some(result) = parse_font_encoding(doc, font_dict) {
if result.gid_glyph_count > 0 {
if !result.gid_codes.is_empty()
&& !tounicode_maps_codes(font_dict, cmaps, &result.gid_codes)
{
has_gid_fonts = true;
}
if !result.map.is_empty() {
@@ -520,6 +526,34 @@ pub(crate) fn build_font_encodings(
(encodings, has_gid_fonts)
}
/// True when the font's ToUnicode CMap maps the gid-named character codes,
/// so the Differences entries still decode through the CMap.
fn tounicode_maps_codes(font_dict: &lopdf::Dictionary, cmaps: &FontCMaps, codes: &[u8]) -> bool {
let Some(obj_ref) = font_dict
.get(b"ToUnicode")
.ok()
.and_then(|o| o.as_reference().ok())
else {
return false;
};
let Some(entry) = cmaps.get_by_obj(obj_ref.0) else {
return false;
};
// At least one gid code usably mapped means the CMap addresses these
// codes; remaining unmapped codes are subset leftovers (e.g. the
// component glyphs of an emoji ZWJ sequence mapped whole on its first
// code). A mapping is usable only when extraction would accept it —
// empty or U+FFFD results are rejected there as invalid. Fonts whose
// CMap ignores the gid codes entirely stay flagged, and the downstream
// garbage/encoding checks still catch partial damage.
codes.iter().any(|&code| {
entry
.primary
.lookup(code as u16)
.is_some_and(|s| !s.is_empty() && !s.contains('\u{FFFD}'))
})
}
/// Parse font encoding from a font dictionary
pub(crate) fn parse_font_encoding(
doc: &Document,
@@ -558,11 +592,10 @@ pub(crate) fn parse_font_encoding(
/// Result of parsing an encoding dictionary's Differences array.
pub(crate) struct EncodingResult {
pub map: FontEncodingMap,
/// Number of glyph names matching the `gidNNNNN` pattern (raw glyph IDs).
/// These indicate a font with unresolvable encoding — the glyph IDs
/// reference the original font's glyph table, but without the original
/// font's cmap there is no way to map them to Unicode.
pub gid_glyph_count: u32,
/// Character codes whose glyph names match the `gidNNNNN` pattern (raw
/// glyph IDs). These reference the original font's glyph table and are
/// only decodable when the font's ToUnicode CMap maps the code.
pub gid_codes: Vec<u8>,
}
/// Parse an encoding dictionary with Differences array
@@ -588,7 +621,7 @@ pub(crate) fn parse_encoding_dictionary(
let mut encoding_map = FontEncodingMap::new();
let mut current_code: u8 = 0;
let mut ligature_count = 0u32;
let mut gid_glyph_count = 0u32;
let mut gid_codes: Vec<u8> = Vec::new();
for item in diff_array {
match item {
@@ -614,7 +647,7 @@ pub(crate) fn parse_encoding_dictionary(
&& glyph_name.len() >= 4
&& glyph_name[3..].chars().all(|c| c.is_ascii_digit())
{
gid_glyph_count += 1;
gid_codes.push(current_code);
}
if let Some(ch) = mapped_char {
encoding_map.insert(current_code, ch);
@@ -638,16 +671,16 @@ pub(crate) fn parse_encoding_dictionary(
);
}
if gid_glyph_count > 0 {
if !gid_codes.is_empty() {
debug!(
" Differences: {} gid-encoded glyphs (unresolvable without original font)",
gid_glyph_count
" Differences: {} gid-encoded glyphs (decodable only via ToUnicode)",
gid_codes.len()
);
}
Some(EncodingResult {
map: encoding_map,
gid_glyph_count,
gid_codes,
})
}
@@ -1938,4 +1971,113 @@ mod tests {
false
));
}
fn gid_font_doc(bfchar: Option<&str>) -> (Document, lopdf::ObjectId) {
use lopdf::Stream;
let mut doc = Document::with_version("1.4");
let cmap = format!(
"/CIDInit /ProcSet findresource begin
12 dict begin
begincmap
1 begincodespacerange
<00> <FF>
endcodespacerange
1 beginbfchar
{}
endbfchar
endcmap
CMapName currentdict /CMap defineresource pop
end
end",
bfchar.unwrap_or_default()
);
let tounicode_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {},
cmap.into_bytes(),
)));
let enc_id = doc.add_object(dictionary! {
"Type" => "Encoding",
"Differences" => vec![
1.into(),
Object::Name(b"gid1283".to_vec()),
Object::Name(b"gid1464".to_vec()),
],
});
let mut font = dictionary! {
"Type" => "Font",
"Subtype" => "TrueType",
"BaseFont" => "ABCDEF+OpenSymbol",
"Encoding" => Object::Reference(enc_id),
};
if bfchar.is_some() {
font.set("ToUnicode", Object::Reference(tounicode_id));
}
let font_id = doc.add_object(font);
let page_id = doc.add_object(dictionary! {
"Type" => "Page",
"Resources" => dictionary! {
"Font" => dictionary! { "F1" => Object::Reference(font_id) },
},
"MediaBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
});
let pages_id = doc.add_object(dictionary! {
"Type" => "Pages",
"Count" => Object::Integer(1),
"Kids" => vec![Object::Reference(page_id)],
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"Pages" => Object::Reference(pages_id),
});
doc.trailer.set("Root", Object::Reference(catalog_id));
(doc, page_id)
}
fn gid_flagged(bfchar: Option<&str>) -> bool {
let (doc, page_id) = gid_font_doc(bfchar);
let cmaps = FontCMaps::from_doc(&doc);
let fonts = doc.get_page_fonts(page_id).unwrap();
let (_, has_gid_fonts) = build_font_encodings(&doc, &fonts, &cmaps);
has_gid_fonts
}
#[test]
fn gid_differences_with_covering_tounicode_are_not_flagged() {
// LibreOffice subsets write /gidNNNN Differences names alongside a
// ToUnicode CMap that decodes those codes; the page must not be
// flagged as unresolvable (which would suppress the whole document's
// markdown when every page carries such a font).
assert!(!gid_flagged(Some("<01> <2022>\n<02> <25E6>")));
}
#[test]
fn gid_differences_with_partial_tounicode_are_not_flagged() {
// An emoji ZWJ sequence maps whole on its first code; the remaining
// component-glyph codes are subset leftovers, not damage.
assert!(!gid_flagged(Some(
"<01> <D83DDC68200DD83DDC69200DD83DDC67>"
)));
}
#[test]
fn gid_differences_without_tounicode_are_flagged() {
assert!(
gid_flagged(None),
"gid glyphs without ToUnicode are unresolvable"
);
}
#[test]
fn gid_differences_with_disjoint_tounicode_are_flagged() {
// A ToUnicode that never addresses the gid codes leaves them
// unresolvable.
assert!(gid_flagged(Some("<10> <0041>")));
}
#[test]
fn gid_differences_with_replacement_char_tounicode_are_flagged() {
// A mapping to U+FFFD is not usable — extraction rejects it as an
// invalid CMap result — so it must not clear the gid flag.
assert!(gid_flagged(Some("<01> <FFFD>\n<02> <FFFD>")));
}
}
+1 -1
View File
@@ -162,7 +162,7 @@ fn extract_form_xobject_text_inner(
// Get fonts from the Form's Resources
let form_fonts = get_form_fonts(doc, &stream.dict);
let (font_encodings, _has_gid_fonts) = build_font_encodings(doc, &form_fonts);
let (font_encodings, _has_gid_fonts) = build_font_encodings(doc, &form_fonts, font_cmaps);
// Build font width info for the form
let font_widths = build_font_widths(doc, &form_fonts);
+2 -2
View File
@@ -724,7 +724,7 @@ checksum = "d6790f58c7ff633d8771f42965289203411a5e5c68388703c06e14f24770b41e"
[[package]]
name = "pdf-inspector"
version = "0.1.6"
version = "0.1.7"
dependencies = [
"env_logger",
"include_dir",
@@ -740,7 +740,7 @@ dependencies = [
[[package]]
name = "pdf-inspector-wasm"
version = "0.1.2"
version = "0.1.3"
dependencies = [
"console_error_panic_hook",
"js-sys",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector-wasm"
version = "0.1.2"
version = "0.1.3"
edition = "2021"
authors = ["Firecrawl Team"]
description = "Browser WebAssembly bindings for pdf-inspector"