Compare commits
merge into: mu-ref/pdf-inspector:abi/key-value-table-fallback
mu-ref/pdf-inspector:main
mu-ref/pdf-inspector:feat/mc-weave-class
mu-ref/pdf-inspector:fix/parallel-prose-table-veto
mu-ref/pdf-inspector:feat/edge-furniture-stripping
mu-ref/pdf-inspector:feat/recursive-region-segmentation
mu-ref/pdf-inspector:fix/detector-stream-inflate
mu-ref/pdf-inspector:fix/multi-column-reading-order
mu-ref/pdf-inspector:perf/ocr-adaptive-staged-engine
mu-ref/pdf-inspector:feat/base-font-names
mu-ref/pdf-inspector:feat/native-formula-latex
mu-ref/pdf-inspector:fix/masthead-scan-ocr-routing
mu-ref/pdf-inspector:abi/fix-npm-arm64-release
mu-ref/pdf-inspector:abi/release-1.15.0
mu-ref/pdf-inspector:abi/ocr-launch-polish
mu-ref/pdf-inspector:abi/ocr-bindings
mu-ref/pdf-inspector:abi/ocr-ci
mu-ref/pdf-inspector:abi/ocr-api-finalize
mu-ref/pdf-inspector:abi/ocr-release-hardening
mu-ref/pdf-inspector:abi/ocr-adaptive-fallback
mu-ref/pdf-inspector:abi/ocr-native-recovery
mu-ref/pdf-inspector:abi/ocr-runtime
mu-ref/pdf-inspector:abi/table-reference-rows
mu-ref/pdf-inspector:abi/ocr-assembly
mu-ref/pdf-inspector:abi/local-ocr-api
mu-ref/pdf-inspector:abi/local-ocr-fusion
mu-ref/pdf-inspector:abi/local-ocr
mu-ref/pdf-inspector:abi/local-ocr-contracts
mu-ref/pdf-inspector:abi/local-ocr-oar
mu-ref/pdf-inspector:abi/local-ocr-routing
mu-ref/pdf-inspector:markdown/dehyphenate-line-breaks
mu-ref/pdf-inspector:chore/release-1.14.2
mu-ref/pdf-inspector:fix/cluster-rects-disjoint-quadratic
mu-ref/pdf-inspector:fix/detector-tj-quadratic-scan
mu-ref/pdf-inspector:fix/tounicode-bfrange-expansion-budget
mu-ref/pdf-inspector:fix/encoding-cidrange-expansion-budget
mu-ref/pdf-inspector:tables/running-header-footer-veto
mu-ref/pdf-inspector:fix/content-decode-operation-cap
mu-ref/pdf-inspector:extractor/merge-small-caps-runs
mu-ref/pdf-inspector:xobjects/text-line-matrix-and-missing-operators
mu-ref/pdf-inspector:fix/cid-w-range-expansion-budget
mu-ref/pdf-inspector:fix/form-xobject-expansion-budget
mu-ref/pdf-inspector:waku/session-a51f9c8e
mu-ref/pdf-inspector:chore/release-1.14.1
mu-ref/pdf-inspector:fix/regions-invisible-text-layer
mu-ref/pdf-inspector:docs-structure-elements
mu-ref/pdf-inspector:abi/unify-package-versions
mu-ref/pdf-inspector:expose-structure-headings
mu-ref/pdf-inspector:abi/bump-package-versions
mu-ref/pdf-inspector:cursor/fix-structtree-alias-dos-84e6
mu-ref/pdf-inspector:cursor/clamp-column-histogram-bins-3fe5
mu-ref/pdf-inspector:cursor/node-async-bindings-47a5
mu-ref/pdf-inspector:cursor/security-md-reporting-channels-3fe5
mu-ref/pdf-inspector:cursor/harden-glyph-name-parsing-f3eb
mu-ref/pdf-inspector:cursor/fix-tounicode-char-boundary-panic-e329
mu-ref/pdf-inspector:cursor/fix-acroform-kids-cycle-dos-84e6
mu-ref/pdf-inspector:fix-embedded-drop-cap
mu-ref/pdf-inspector:table-script-filter
mu-ref/pdf-inspector:abi/refresh-site-benchmark
mu-ref/pdf-inspector:split-font-metrics
mu-ref/pdf-inspector:abi/fix-strikeout-false-positives
mu-ref/pdf-inspector:split-layout-heuristics
mu-ref/pdf-inspector:abi/fix-digit-only-text-runs
mu-ref/pdf-inspector:legacy-tex-pdf-extraction
mu-ref/pdf-inspector:abi/update-benchmark-results
mu-ref/pdf-inspector:abi/release-wasm-0.1.3
mu-ref/pdf-inspector:abi/release-npm-1.11.2
mu-ref/pdf-inspector:abi/release-pypi-0.2.6
mu-ref/pdf-inspector:abi/release-crate-0.1.7
mu-ref/pdf-inspector:fix/gid-differences-tounicode
mu-ref/pdf-inspector:abi/fix-wasm-detection-strategy
mu-ref/pdf-inspector:abi/organize-bindings
mu-ref/pdf-inspector:abi/wasm-demo
mu-ref/pdf-inspector:abi/wasm-bindings
mu-ref/pdf-inspector:abi/refresh-oss-site
mu-ref/pdf-inspector:abi/refresh-benchmark-table
mu-ref/pdf-inspector:abi/backend-evidence-probe
mu-ref/pdf-inspector:abi/paired-benchmark-harness
mu-ref/pdf-inspector:abi/region-reading-order
mu-ref/pdf-inspector:abi/table-candidate-scoring
mu-ref/pdf-inspector:fix/link-clip-center-y
mu-ref/pdf-inspector:bump-1104
mu-ref/pdf-inspector:eng5058-exclusive-items
mu-ref/pdf-inspector:abi/bump-npm-1.10.3
mu-ref/pdf-inspector:eng5034-decoration-geometry
mu-ref/pdf-inspector:bump-1102
mu-ref/pdf-inspector:feat/granular-ocr-reasons
mu-ref/pdf-inspector:fix/coverage-fallback-heading-gate
mu-ref/pdf-inspector:feat/password-flag
mu-ref/pdf-inspector:docs/refresh-benchmark
mu-ref/pdf-inspector:feat/landing-page-v2
mu-ref/pdf-inspector:abi/bump-napi-1.10.1
mu-ref/pdf-inspector:feat/tracked-letterspacing
mu-ref/pdf-inspector:fix/heading-sparse-page-density
mu-ref/pdf-inspector:fix/heading-classification-2
mu-ref/pdf-inspector:fix/row-stripe-prose-guard
mu-ref/pdf-inspector:fix/heading-classification
mu-ref/pdf-inspector:fix/urw-medi-bold
mu-ref/pdf-inspector:chore/bump-1.10.0
mu-ref/pdf-inspector:feat/descriptor-styles-and-strikeout
mu-ref/pdf-inspector:add-mit-license
mu-ref/pdf-inspector:refactor/consolidate-text-quality-detectors
mu-ref/pdf-inspector:fix/detect-shifted-cipher-tounicode
mu-ref/pdf-inspector:fix/issue-118-garbled-text
mu-ref/pdf-inspector:abi/formatting-round-2b
mu-ref/pdf-inspector:abi/underline-flag
mu-ref/pdf-inspector:abi/fix-navigation-running-headers
mu-ref/pdf-inspector:abi/fix-docugram-column-order
mu-ref/pdf-inspector:abi/ocr-reason-signal
mu-ref/pdf-inspector:abi/legacy-npm-publish
mu-ref/pdf-inspector:abi/pdf-text-quality-ocr
mu-ref/pdf-inspector:abi/setup-crates-publishing
mu-ref/pdf-inspector:abi/add-crates-readme
mu-ref/pdf-inspector:abi/use-lopdf-crate
mu-ref/pdf-inspector:abi/fix-bold-abstract-headings
mu-ref/pdf-inspector:abi/key-value-continuations
mu-ref/pdf-inspector:abi/key-value-table-fallback
mu-ref/pdf-inspector:abi/vector-region-borderless-tables
mu-ref/pdf-inspector:abi/vector-table-confidence
mu-ref/pdf-inspector:feat/emit-image-xobject-bboxes
mu-ref/pdf-inspector:extract-tables/region-text-density-floor
mu-ref/pdf-inspector:abimaelmartell/ci-job-investigation
mu-ref/pdf-inspector:abimaelmartell/security-policy
mu-ref/pdf-inspector:extract-tables/reject-partial-extraction
mu-ref/pdf-inspector:detect-lines/columns-from-horizontal-segments
mu-ref/pdf-inspector:extract-tables/use-vector-detectors
mu-ref/pdf-inspector:detect-lines/accept-full-page-grids
mu-ref/pdf-inspector:detect-rects/long-cells-in-multi-row-grids
mu-ref/pdf-inspector:napi-probe-doc
mu-ref/pdf-inspector:codex/multiline-indent-cell-grid
mu-ref/pdf-inspector:fix-wireless-vector-grid-prose
mu-ref/pdf-inspector:abimaelmartell/wired-grids
mu-ref/pdf-inspector:fix-cid-mojibake-and-wide-tsr-cells
mu-ref/pdf-inspector:codex/doc51-multirow-fallback-guard
mu-ref/pdf-inspector:codex/fix-doc128-vector-grid
mu-ref/pdf-inspector:tables-expand-multi-row-cells
mu-ref/pdf-inspector:tables/vector-grid-region-napi
mu-ref/pdf-inspector:fix/tsr-auto-fallback-bugs
mu-ref/pdf-inspector:feat/tsr-auto-fallback
mu-ref/pdf-inspector:fix/tsr-exclusive-item-cell-assignment
mu-ref/pdf-inspector:codex/pdf-container-repair
mu-ref/pdf-inspector:fix/tsr-stage2-same-line-only
mu-ref/pdf-inspector:fix/tsr-stage2-orphan-assignment
mu-ref/pdf-inspector:fix-tsr-overlapping-cell-bboxes
mu-ref/pdf-inspector:tsr-aware-table-extraction
mu-ref/pdf-inspector:codex/fix-tagged-table-header-recovery
mu-ref/pdf-inspector:abimaelmartell/promote-implicit-header-rows
mu-ref/pdf-inspector:abimaelmartell/prose-in-framed-region
mu-ref/pdf-inspector:abimaelmartell/merged-cell-tangent-fix
mu-ref/pdf-inspector:abimaelmartell/wrapped-bold-list-lead
mu-ref/pdf-inspector:abimaelmartell/fix-mythos-lists
mu-ref/pdf-inspector:abimaelmartell/issue-49-review
mu-ref/pdf-inspector:abimaelmartell/fix-list-bullets
mu-ref/pdf-inspector:abimaelmartell/table-kind-enum
mu-ref/pdf-inspector:abimaelmartell/exclude-toc-from-tables-meta
mu-ref/pdf-inspector:abimaelmartell/toc-hierarchical-indent
mu-ref/pdf-inspector:abimaelmartell/split-interleaved-tiny-text
mu-ref/pdf-inspector:abimaelmartell/review-pdf-16
mu-ref/pdf-inspector:abimaelmartell/npm-windows-binaries
mu-ref/pdf-inspector:abimaelmartell/fix-toc-false-positive
mu-ref/pdf-inspector:abimaelmartell/npm-cli-bin
mu-ref/pdf-inspector:fix/actualtext-ligature-position
mu-ref/pdf-inspector:feat/formula-latex-recovery
mu-ref/pdf-inspector:fix/classifier-heuristics-decodable-fonts
mu-ref/pdf-inspector:abimaelmartell/formula-extraction
mu-ref/pdf-inspector:abimaelmartell/arxiv-ocr-false-pos
mu-ref/pdf-inspector:fix/numeric-column-table-detection
mu-ref/pdf-inspector:abimaelmartell/auto-npm-publish
mu-ref/pdf-inspector:abimaelmartell/pages-classification
mu-ref/pdf-inspector:abimaelmartell/heuristic-layout-detect
mu-ref/pdf-inspector:v1.15.0
mu-ref/pdf-inspector:v1.14.2
mu-ref/pdf-inspector:packages-2026-08-10
mu-ref/pdf-inspector:v0.7.0
mu-ref/pdf-inspector:v0.6.0
mu-ref/pdf-inspector:v0.5.0
mu-ref/pdf-inspector:v0.4.3
mu-ref/pdf-inspector:v0.4.2
mu-ref/pdf-inspector:v0.4.1
mu-ref/pdf-inspector:v0.4.0
mu-ref/pdf-inspector:v0.3.6
mu-ref/pdf-inspector:v0.3.5
mu-ref/pdf-inspector:v0.3.4
mu-ref/pdf-inspector:v0.3.3
mu-ref/pdf-inspector:v0.3.2
mu-ref/pdf-inspector:v0.3.1
mu-ref/pdf-inspector:v0.3.0
mu-ref/pdf-inspector:v0.2.3
mu-ref/pdf-inspector:v0.2.2
mu-ref/pdf-inspector:v0.2.1
mu-ref/pdf-inspector:v0.2.0
...
pull from: mu-ref/pdf-inspector:fix/classifier-heuristics-decodable-fonts
mu-ref/pdf-inspector:main
mu-ref/pdf-inspector:feat/mc-weave-class
mu-ref/pdf-inspector:fix/parallel-prose-table-veto
mu-ref/pdf-inspector:feat/edge-furniture-stripping
mu-ref/pdf-inspector:feat/recursive-region-segmentation
mu-ref/pdf-inspector:fix/detector-stream-inflate
mu-ref/pdf-inspector:fix/multi-column-reading-order
mu-ref/pdf-inspector:perf/ocr-adaptive-staged-engine
mu-ref/pdf-inspector:feat/base-font-names
mu-ref/pdf-inspector:feat/native-formula-latex
mu-ref/pdf-inspector:fix/masthead-scan-ocr-routing
mu-ref/pdf-inspector:abi/fix-npm-arm64-release
mu-ref/pdf-inspector:abi/release-1.15.0
mu-ref/pdf-inspector:abi/ocr-launch-polish
mu-ref/pdf-inspector:abi/ocr-bindings
mu-ref/pdf-inspector:abi/ocr-ci
mu-ref/pdf-inspector:abi/ocr-api-finalize
mu-ref/pdf-inspector:abi/ocr-release-hardening
mu-ref/pdf-inspector:abi/ocr-adaptive-fallback
mu-ref/pdf-inspector:abi/ocr-native-recovery
mu-ref/pdf-inspector:abi/ocr-runtime
mu-ref/pdf-inspector:abi/table-reference-rows
mu-ref/pdf-inspector:abi/ocr-assembly
mu-ref/pdf-inspector:abi/local-ocr-api
mu-ref/pdf-inspector:abi/local-ocr-fusion
mu-ref/pdf-inspector:abi/local-ocr
mu-ref/pdf-inspector:abi/local-ocr-contracts
mu-ref/pdf-inspector:abi/local-ocr-oar
mu-ref/pdf-inspector:abi/local-ocr-routing
mu-ref/pdf-inspector:markdown/dehyphenate-line-breaks
mu-ref/pdf-inspector:chore/release-1.14.2
mu-ref/pdf-inspector:fix/cluster-rects-disjoint-quadratic
mu-ref/pdf-inspector:fix/detector-tj-quadratic-scan
mu-ref/pdf-inspector:fix/tounicode-bfrange-expansion-budget
mu-ref/pdf-inspector:fix/encoding-cidrange-expansion-budget
mu-ref/pdf-inspector:tables/running-header-footer-veto
mu-ref/pdf-inspector:fix/content-decode-operation-cap
mu-ref/pdf-inspector:extractor/merge-small-caps-runs
mu-ref/pdf-inspector:xobjects/text-line-matrix-and-missing-operators
mu-ref/pdf-inspector:fix/cid-w-range-expansion-budget
mu-ref/pdf-inspector:fix/form-xobject-expansion-budget
mu-ref/pdf-inspector:waku/session-a51f9c8e
mu-ref/pdf-inspector:chore/release-1.14.1
mu-ref/pdf-inspector:fix/regions-invisible-text-layer
mu-ref/pdf-inspector:docs-structure-elements
mu-ref/pdf-inspector:abi/unify-package-versions
mu-ref/pdf-inspector:expose-structure-headings
mu-ref/pdf-inspector:abi/bump-package-versions
mu-ref/pdf-inspector:cursor/fix-structtree-alias-dos-84e6
mu-ref/pdf-inspector:cursor/clamp-column-histogram-bins-3fe5
mu-ref/pdf-inspector:cursor/node-async-bindings-47a5
mu-ref/pdf-inspector:cursor/security-md-reporting-channels-3fe5
mu-ref/pdf-inspector:cursor/harden-glyph-name-parsing-f3eb
mu-ref/pdf-inspector:cursor/fix-tounicode-char-boundary-panic-e329
mu-ref/pdf-inspector:cursor/fix-acroform-kids-cycle-dos-84e6
mu-ref/pdf-inspector:fix-embedded-drop-cap
mu-ref/pdf-inspector:table-script-filter
mu-ref/pdf-inspector:abi/refresh-site-benchmark
mu-ref/pdf-inspector:split-font-metrics
mu-ref/pdf-inspector:abi/fix-strikeout-false-positives
mu-ref/pdf-inspector:split-layout-heuristics
mu-ref/pdf-inspector:abi/fix-digit-only-text-runs
mu-ref/pdf-inspector:legacy-tex-pdf-extraction
mu-ref/pdf-inspector:abi/update-benchmark-results
mu-ref/pdf-inspector:abi/release-wasm-0.1.3
mu-ref/pdf-inspector:abi/release-npm-1.11.2
mu-ref/pdf-inspector:abi/release-pypi-0.2.6
mu-ref/pdf-inspector:abi/release-crate-0.1.7
mu-ref/pdf-inspector:fix/gid-differences-tounicode
mu-ref/pdf-inspector:abi/fix-wasm-detection-strategy
mu-ref/pdf-inspector:abi/organize-bindings
mu-ref/pdf-inspector:abi/wasm-demo
mu-ref/pdf-inspector:abi/wasm-bindings
mu-ref/pdf-inspector:abi/refresh-oss-site
mu-ref/pdf-inspector:abi/refresh-benchmark-table
mu-ref/pdf-inspector:abi/backend-evidence-probe
mu-ref/pdf-inspector:abi/paired-benchmark-harness
mu-ref/pdf-inspector:abi/region-reading-order
mu-ref/pdf-inspector:abi/table-candidate-scoring
mu-ref/pdf-inspector:fix/link-clip-center-y
mu-ref/pdf-inspector:bump-1104
mu-ref/pdf-inspector:eng5058-exclusive-items
mu-ref/pdf-inspector:abi/bump-npm-1.10.3
mu-ref/pdf-inspector:eng5034-decoration-geometry
mu-ref/pdf-inspector:bump-1102
mu-ref/pdf-inspector:feat/granular-ocr-reasons
mu-ref/pdf-inspector:fix/coverage-fallback-heading-gate
mu-ref/pdf-inspector:feat/password-flag
mu-ref/pdf-inspector:docs/refresh-benchmark
mu-ref/pdf-inspector:feat/landing-page-v2
mu-ref/pdf-inspector:abi/bump-napi-1.10.1
mu-ref/pdf-inspector:feat/tracked-letterspacing
mu-ref/pdf-inspector:fix/heading-sparse-page-density
mu-ref/pdf-inspector:fix/heading-classification-2
mu-ref/pdf-inspector:fix/row-stripe-prose-guard
mu-ref/pdf-inspector:fix/heading-classification
mu-ref/pdf-inspector:fix/urw-medi-bold
mu-ref/pdf-inspector:chore/bump-1.10.0
mu-ref/pdf-inspector:feat/descriptor-styles-and-strikeout
mu-ref/pdf-inspector:add-mit-license
mu-ref/pdf-inspector:refactor/consolidate-text-quality-detectors
mu-ref/pdf-inspector:fix/detect-shifted-cipher-tounicode
mu-ref/pdf-inspector:fix/issue-118-garbled-text
mu-ref/pdf-inspector:abi/formatting-round-2b
mu-ref/pdf-inspector:abi/underline-flag
mu-ref/pdf-inspector:abi/fix-navigation-running-headers
mu-ref/pdf-inspector:abi/fix-docugram-column-order
mu-ref/pdf-inspector:abi/ocr-reason-signal
mu-ref/pdf-inspector:abi/legacy-npm-publish
mu-ref/pdf-inspector:abi/pdf-text-quality-ocr
mu-ref/pdf-inspector:abi/setup-crates-publishing
mu-ref/pdf-inspector:abi/add-crates-readme
mu-ref/pdf-inspector:abi/use-lopdf-crate
mu-ref/pdf-inspector:abi/fix-bold-abstract-headings
mu-ref/pdf-inspector:abi/key-value-continuations
mu-ref/pdf-inspector:abi/key-value-table-fallback
mu-ref/pdf-inspector:abi/vector-region-borderless-tables
mu-ref/pdf-inspector:abi/vector-table-confidence
mu-ref/pdf-inspector:feat/emit-image-xobject-bboxes
mu-ref/pdf-inspector:extract-tables/region-text-density-floor
mu-ref/pdf-inspector:abimaelmartell/ci-job-investigation
mu-ref/pdf-inspector:abimaelmartell/security-policy
mu-ref/pdf-inspector:extract-tables/reject-partial-extraction
mu-ref/pdf-inspector:detect-lines/columns-from-horizontal-segments
mu-ref/pdf-inspector:extract-tables/use-vector-detectors
mu-ref/pdf-inspector:detect-lines/accept-full-page-grids
mu-ref/pdf-inspector:detect-rects/long-cells-in-multi-row-grids
mu-ref/pdf-inspector:napi-probe-doc
mu-ref/pdf-inspector:codex/multiline-indent-cell-grid
mu-ref/pdf-inspector:fix-wireless-vector-grid-prose
mu-ref/pdf-inspector:abimaelmartell/wired-grids
mu-ref/pdf-inspector:fix-cid-mojibake-and-wide-tsr-cells
mu-ref/pdf-inspector:codex/doc51-multirow-fallback-guard
mu-ref/pdf-inspector:codex/fix-doc128-vector-grid
mu-ref/pdf-inspector:tables-expand-multi-row-cells
mu-ref/pdf-inspector:tables/vector-grid-region-napi
mu-ref/pdf-inspector:fix/tsr-auto-fallback-bugs
mu-ref/pdf-inspector:feat/tsr-auto-fallback
mu-ref/pdf-inspector:fix/tsr-exclusive-item-cell-assignment
mu-ref/pdf-inspector:codex/pdf-container-repair
mu-ref/pdf-inspector:fix/tsr-stage2-same-line-only
mu-ref/pdf-inspector:fix/tsr-stage2-orphan-assignment
mu-ref/pdf-inspector:fix-tsr-overlapping-cell-bboxes
mu-ref/pdf-inspector:tsr-aware-table-extraction
mu-ref/pdf-inspector:codex/fix-tagged-table-header-recovery
mu-ref/pdf-inspector:abimaelmartell/promote-implicit-header-rows
mu-ref/pdf-inspector:abimaelmartell/prose-in-framed-region
mu-ref/pdf-inspector:abimaelmartell/merged-cell-tangent-fix
mu-ref/pdf-inspector:abimaelmartell/wrapped-bold-list-lead
mu-ref/pdf-inspector:abimaelmartell/fix-mythos-lists
mu-ref/pdf-inspector:abimaelmartell/issue-49-review
mu-ref/pdf-inspector:abimaelmartell/fix-list-bullets
mu-ref/pdf-inspector:abimaelmartell/table-kind-enum
mu-ref/pdf-inspector:abimaelmartell/exclude-toc-from-tables-meta
mu-ref/pdf-inspector:abimaelmartell/toc-hierarchical-indent
mu-ref/pdf-inspector:abimaelmartell/split-interleaved-tiny-text
mu-ref/pdf-inspector:abimaelmartell/review-pdf-16
mu-ref/pdf-inspector:abimaelmartell/npm-windows-binaries
mu-ref/pdf-inspector:abimaelmartell/fix-toc-false-positive
mu-ref/pdf-inspector:abimaelmartell/npm-cli-bin
mu-ref/pdf-inspector:fix/actualtext-ligature-position
mu-ref/pdf-inspector:feat/formula-latex-recovery
mu-ref/pdf-inspector:fix/classifier-heuristics-decodable-fonts
mu-ref/pdf-inspector:abimaelmartell/formula-extraction
mu-ref/pdf-inspector:abimaelmartell/arxiv-ocr-false-pos
mu-ref/pdf-inspector:fix/numeric-column-table-detection
mu-ref/pdf-inspector:abimaelmartell/auto-npm-publish
mu-ref/pdf-inspector:abimaelmartell/pages-classification
mu-ref/pdf-inspector:abimaelmartell/heuristic-layout-detect
mu-ref/pdf-inspector:v1.15.0
mu-ref/pdf-inspector:v1.14.2
mu-ref/pdf-inspector:packages-2026-08-10
mu-ref/pdf-inspector:v0.7.0
mu-ref/pdf-inspector:v0.6.0
mu-ref/pdf-inspector:v0.5.0
mu-ref/pdf-inspector:v0.4.3
mu-ref/pdf-inspector:v0.4.2
mu-ref/pdf-inspector:v0.4.1
mu-ref/pdf-inspector:v0.4.0
mu-ref/pdf-inspector:v0.3.6
mu-ref/pdf-inspector:v0.3.5
mu-ref/pdf-inspector:v0.3.4
mu-ref/pdf-inspector:v0.3.3
mu-ref/pdf-inspector:v0.3.2
mu-ref/pdf-inspector:v0.3.1
mu-ref/pdf-inspector:v0.3.0
mu-ref/pdf-inspector:v0.2.3
mu-ref/pdf-inspector:v0.2.2
mu-ref/pdf-inspector:v0.2.1
mu-ref/pdf-inspector:v0.2.0
| Author | SHA1 | Message | Date | |
|---|---|---|---|---|
|
|
90e36e8dfd |
fix(detector): respect resource shadowing when resolving page-content fonts
The previous ObjectId-based fix correctly scoped Form XObject fonts
but still violated PDF resource inheritance for page content. When a
page overrides /F1 from a parent /Pages node (different font dict for
the same name), get_page_resources returns the page's own /Resources
plus all ancestor /Resources dicts. The old code called
resolve_font_names_to_ids on each one and added every match to
used_font_ids — both font ObjectIds ended up in the used set even
though only the page's /F1 is actually visible to that page's content.
Per ISO 32000-1 §7.7.3.4, resource names are inherited with
shadowing semantics: the most-specific (deepest, closest to the page)
definition wins.
Fix:
- New lookup_font_id helper resolves a single name in a single dict.
- New resolve_with_shadowing iterates names, checking the page's own
/Resources first, then walking ancestors in most-specific-first
order (which is the order lopdf's get_page_resources returns).
First hit wins via a labeled `continue 'name` — subsequent
ancestors are skipped for that name.
- analyze_page_content's flat resolution loop replaced with one call
to resolve_with_shadowing.
Audit:
- XObject path is correct: each Form XObject already resolves names
against its OWN /Resources (XObjects don't inherit from page tree).
- font_map population is correct: keyed by ObjectId, so collecting
from all dicts builds the full available-fonts catalog. The bug
was only in the used-set resolution.
- Confirmed lopdf returns ancestors in most-specific-first order
(page → parent → grandparent → root), matching the shadowing
direction used here.
Tests added (3):
- page /F1 undecodable shadows parent's decodable /F1 → MUST flag
- page /F1 decodable shadows parent's undecodable /F1 → MUST NOT flag
- no override: page inherits parent's decodable /F1 → MUST NOT flag
Validation:
- cargo test --release: 462 tests pass (356 lib + 104 integration + 2 doc)
- cargo clippy --lib --bin detect-pdf -- -D warnings: clean
- external eval: 9/9, 0 regressions, 6/6 false positives resolved,
61/61 scanned pages still correctly flagged
|
||
|
|
dd2621c488 |
fix(detector): scope font lookups by ObjectId + handle indirect Form Resources
Addresses two more reviewer findings on the previous decodable-font commit.
P1 — Resource-name scoping bug
The previous fix keyed used_font_names and font_map by raw resource
names like b"F1". PDF resource names are scoped to each resource
dictionary: a Form XObject can legally define its own /F1 that points
to a completely different font from the page's /F1. Because
collect_fonts_from_resource_dict skipped duplicates with
`if font_map.contains_key(name)`, the first definition won and later
Tf /F1 usages in different scopes resolved against the wrong font.
This could reintroduce both the undecodable-Identity-H false flag
and the decodable-CID false unflag depending on which side of the
collision happened to be inserted first.
Fix: switch the lookup mechanism from font names to font ObjectIds.
- font_map: HashMap<ObjectId, FontInfo> (was Vec<u8> keys)
- used_font_ids: HashSet<ObjectId> (was Vec<u8> names)
- new resolve_font_names_to_ids() runs immediately after each
content scan, against the resource dict in scope, to translate
the per-scope name set into ObjectIds.
Each Form XObject's content stream now resolves /F1 against THAT
XObject's own Resources, so name collisions are impossible by design.
Inline (no-ID) font dicts are skipped — extremely rare in practice
and have no stable key.
P2 — Indirect Form /Resources skipped
scan_xobjects_in_resources used `.as_dict()` on the Form's /Resources
entry, which returns None for indirect references. PDFs frequently
store /Resources as `X 0 R`, in which case font collection and
recursion were both skipped — even though the Tf usages inside the
XObject content had already been recorded.
Fix: handle Object::Reference(r) in addition to Object::Dictionary(d)
by resolving via doc.get_dictionary. Audited the rest of the file —
the other /Resources access points (analyze_page_images,
collect_images_from_resources) already handled both cases.
Tests added (4):
- P1 same-name-different-font (page undecodable, XObject decodable):
must NOT flag — XObject's text is decodable in its own scope.
- P1 inverse (page decodable, XObject undecodable, content uses
XObject /F1): MUST flag — undecodable text exists in real scope.
- P2 indirect Form /Resources: font discovery must still work when
/Resources is a `X 0 R` reference rather than inline.
- Combined regression: indirect Resources + name collision.
Validation:
- cargo test --release: 459 tests pass (353 lib + 104 integration + 2 doc)
- cargo clippy --lib --bin detect-pdf -- -D warnings: clean
- external eval (9 PDFs): 9/9 pass, 6/6 false positives resolved,
0 regressions, 61/61 truly-scanned pages still flagged
The behavior on the eval set is identical — confirms the correctness
fix isn't masking any change in classifier outcomes.
|
||
|
|
fe3b026a46 |
fix(detector): make decodable-font checks usage-based and XObject-aware
Addresses two reviewer concerns on the previous heuristic fix:
P1 — resource-based check could create an inverse bug
page_has_identity_h_no_tounicode and page_has_decodable_text_fonts
iterated all fonts in the page Resources dict, including unused fonts.
A page whose actual text was rendered exclusively in an undecodable
Identity-H font but whose Resources also listed an unused decodable
Type1 would be wrongly unflagged.
Fix: parse Tf operator operands during content stream scanning to
collect the set of font names actually referenced. The font checks
now filter to only USED fonts via a new used_fonts_have_*
family of functions operating on (used_font_names, font_map).
P2 — checks didn't follow text into Form XObjects
analyze_page_content correctly recurses through Form XObjects via
scan_xobjects_in_resources, but the font checks only looked at the
page's top-level Resources/Font. Pages that render text through Form
XObjects (corporate templates, header/footer overlays) had their
XObject font resources missed entirely.
Fix: scan_xobjects_in_resources now propagates the used_font_names
set AND collects fonts from each Form XObject's own Resources into
the shared font_map. The usage-based check sees the full picture:
page-level fonts + every nested XObject's fonts, intersected with
fonts actually referenced by Tf operators anywhere in the content.
Implementation:
- New extract_font_name_before_tf helper (parses /Name immediately
preceding Tf).
- New FontInfo struct caches font properties per-name.
- New collect_fonts_from_resource_dict + new used_fonts_have_*
functions are pure filters over (used_names, font_map).
- analyze_page_content threads used_font_names + font_map through
page content scan and XObject recursion, then runs the new checks.
- Old resource-based functions kept as #[cfg(test)] for the existing
unit-test interface.
- Phase 3 uncached-page loop now goes through analyze_page_content
so it also gets the usage-based + XObject-aware behavior.
Tests added (8):
- extract_font_name_before_tf basic + long-name parsing
- scan_content_for_text_operators collects used font names
- P1 — unused decodable font in Resources doesn't save a page
whose used font is undecodable
- P1 — both fonts used → decodable font correctly prevents flag
- P2 — decodable font inside Form XObject correctly unflags
- P2 — undecodable font only in XObject still flags even with
unused decodable font at page level
- P2 — has_decodable_text_fonts populated from XObject fonts
Validation:
- 349 lib + 104 integration + 2 doc tests pass (was 341)
- cargo clippy --lib --bin detect-pdf -- -D warnings: clean
- External eval: 9/9 PDFs pass, 6/6 false positives resolved,
0 regressions, 61/61 scanned pages still correctly flagged
- No eval delta — confirms previous fix wasn't relying on the
resource-based bug for any of the eval PDFs
|
||
|
|
ee3c0967b2 |
test(detector): add unit tests for the three classifier fixes
Adds 10 unit tests covering the heuristic changes: - has_vector_text alphanum guard - real text + decorative paths → not flagged - true outlined glyphs (low alphanum) → still flagged - page_has_identity_h_no_tounicode supplementary-font handling - undecodable Identity-H + decodable Type1 → not flagged (new) - undecodable Identity-H alone → still flagged (regression) - page_has_decodable_text_fonts (new helper) - Type1 → true - Type0 with ToUnicode → true - undecodable Identity-H only → false - looks_like_scan with has_decodable_text_fonts override - CID-encoded decodable text → not flagged as scan - same metrics with no decodable fonts → still flagged - decodable fonts but text_ops < 10 (page-number overlay) → still flagged |
||
|
|
f4bdb35935 |
fix(detector): correct false flags for CID-encoded text and supplementary fonts (0.7.4)
The page classifier was over-aggressively flagging Mixed-PDF pages as needing OCR in three distinct cases. Each is fixed at the root in analyze_page_content / page_has_identity_h_no_tounicode / the looks_like_scan check. 1. has_vector_text false positives on dense layouts path_ops > text_ops*200 fired on pages with decorative paths (column borders, dividers) alongside real selectable text. Added a unique_alphanum_chars < 30 guard: real outlined-text pages have very few unique alphanum chars (each glyph is a path), while pages with real text + decorations have many. 2. Identity-H without ToUnicode flagged whole pages on supplementary fonts page_has_identity_h_no_tounicode would flag a page if any single Type0 font lacked ToUnicode and had no fallback CMap, even when the page's actual text came from other decodable fonts (Type1 with ToUnicode, etc.). Rewrote to track both undecodable Identity-H fonts AND other decodable fonts, only flagging when no decodable text font is present. 3. CID-encoded text with ToUnicode misclassified as scan looks_like_scan checked unique_alphanum_chars < 10 on raw string operand bytes. CID-encoded fonts (Type0 with ToUnicode) emit 2-byte CID values that aren't ASCII alphanum, so the metric is blind to them even when the text is fully decodable. Added a has_decodable_text_fonts signal: when a page has decodable fonts AND >= 10 text ops, the low alphanum count is treated as a CID encoding artifact rather than evidence of a scan. Validated against a broad PDF corpus: - 6 known false-positive pages now correctly classified as text - 22 previously-missed scan pages (cover/blank/photo) now correctly flagged for OCR - 0 regressions on truly-scanned PDFs (61/61 pages stay flagged) - All 437 existing tests pass; clippy clean Bumps NAPI package to 0.7.4. |
2 changed files with 1856 additions and 88 deletions
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "firecrawl-pdf-inspector",
|
||||
"version": "0.7.3",
|
||||
"version": "0.7.4",
|
||||
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
||||
"main": "index.js",
|
||||
"types": "index.d.ts",
|
||||
|
||||
+1855
-87
File diff suppressed because it is too large
Load Diff
Reference in New Issue
Block a user
Blocking a user prevents them from interacting with repositories, such as opening or commenting on pull requests or issues. Learn more about blocking a user.