Compare commits

...
Author SHA1 Message Date
Abimael MartellandCursor 7fb9c64130 fix(extractor): scan inline-image EI with the full PDF whitespace set
A missed EI terminator used to consume the rest of the stream and drop
later operators from the decode cap. If EI is absent, keep scanning.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 15:00:43 -07:00
Abimael MartellandCursor 22b1b37123 fix(extractor): treat NUL and form-feed as PDF whitespace in op counting
Names must stop on the full PDF whitespace set so a following operator is
not absorbed into /Name, which would undercount and skip the decode cap.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 14:55:48 -07:00
Abimael MartellandCursor 7ea39eeca2 fix(extractor): cap content-stream decode before allocating operators
The 1M operation limit ran after lopdf materialized the full vector, so a
compact page of q/Q pairs could still abort under memory pressure.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 14:33:20 -07:00
Abimael Martell 89dd20d02c fix(xobjects): track text line matrix and handle T*/TL/'/"/Tc/Tw in Form XObjects (#369)
* fix(xobjects): track text line matrix and handle T*/TL/'/"/Tc/Tw in Form XObjects

The Form XObject text extractor in xobjects.rs is a separate hand-rolled
implementation of the operator state machine in content_stream.rs, and it had
drifted well out of parity:

- No text line matrix (TLM). `Td`/`TD` were applied to the text matrix already
  advanced by `Tj`/`TJ`, so every line began where the previous line *ended*
  instead of at the line start. Lines marched off the right edge and were
  dropped as off-page.
- `T*` was not handled at all, so it never advanced to the next line.
- `TL`, `'` and `"` were missing, and `TD` never set the leading as a side
  effect.
- `Tc`/`Tw` were hardcoded to 0.0 when computing advance widths, drifting
  positions and inserting spurious spaces.
- Text state (Tc/Tw/TL/Tf) is part of the graphics state but was not saved or
  restored by `q`/`Q`.

This matters well beyond an edge case: producers that emit a page stream of
just `q /X Do Q` and put all content in a Form XObject are common in
print-to-PDF and typesetting workflows, so this parser is on the hot path for
whole classes of real documents.

Measured on nycourts.gov 199AD3d.pdf (1370 pages, PDFlib producer, every page
wrapped in a Form XObject, 1331 pages using T*), word recall against a
pdftotext reference goes from 19.2% to 97.7% — 116k extracted words to 582k
against a 578k-word reference. On a 10-page subset, sequence similarity goes
from 27.3% to 97.7%, against 99.8% for Mistral OCR.

pdf-evals: 195 passed / 7 failed, byte-identical to the origin/main baseline
with the same failure list — no regressions.

Adds 7 unit tests covering the line-matrix-relative `Td`, `T*`, `TD` setting
leading, `'`, `"`, `Tc` advance widths, and `q`/`Q` text-state restore, all
driven through a page whose content is only `q /X1 Do Q`.

* fix(xobjects): restore fill colour across q/Q in Form XObjects

A white fill set inside a q/Q pair leaked past the Q, so all subsequent
text was treated as invisible and dropped. Save and restore fill_is_white
with the rest of the graphics state.

This was the cause of several long-standing extraction failures where
whole passages went missing or degraded into per-character garbage:
cambridge_excerpt (+8.5KB of recovered text), MTUAeroEngines (+7.4KB),
2025_findings-acl_668 (+1.4KB), HTM_02-01_Part_A (+1.3KB), and
HuttoISDWorkPerks / ebgt7isj04ophcq, which both went from exploded
per-character tables to clean prose.

Adds a regression test that fails without the restore (the text after Q
is dropped entirely).

Reported by cubic on #369.
2026-08-12 14:32:50 -07:00
Abimael MartellandCursor 3d33ff3dbd fix(extractor): bound CID /W range expansion (#372)
Type0 /W parsing and the Unicode-CID heuristic expanded every CID in every
range. Repeating a full-width [0 65535 w] entry therefore re-materialized
the same 65,536-key domain on every copy, growing a temporary vector and
HashMap work without bound.

Cap expansion at the 16-bit CID domain, collect unique CIDs for the
median heuristic, and stop width assignment once that many entries have
been written. Legitimate compact /W arrays are unchanged.

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 14:27:29 -07:00
Abimael MartellandCursor 75e9b09593 fix(extractor): bound Form XObject expansion per page (#370)
* fix(extractor): bound Form XObject expansion with invocation and operation budgets

Nested Form XObjects were only limited by recursion depth (5). An acyclic
graph where each form invokes the next N times still expands to N^depth
work, so a small PDF can force millions of nested /Do evaluations.

Share a per-page FormWalkBudget that caps 10,000 Form invocations and
1,000,000 operations walked across those expansions. Extraction stops
when either cap is hit. Legitimate shallow nesting is unchanged.

Co-authored-by: Cursor <cursoragent@cursor.com>

* fix(extractor): share Form XObject budget across invisible-layer retry

extract_page_text_items created a fresh FormWalkBudget on every call, so
the invisible-text retry could consume a second full expansion budget for
the same page. Own the budget at the call site and pass it into both
passes.

Also charge form operations independently of the invocation cap so a form
that was already admitted can finish its stream.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Cursor <cursoragent@cursor.com>
2026-08-12 14:03:19 -07:00
Abimael MartellandClaude Fable 5 f4aab3b36f chore(release): bump package versions to 1.14.1 (#354)
Releases fix(regions) #351 — invisible (Tr 3) OCR text layers served from
the region extractor instead of falling back to GPU OCR.

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
2026-08-11 14:18:41 -07:00
16 changed files with 1186 additions and 87 deletions
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector"
version = "1.14.0"
version = "1.14.1"
edition = "2021"
autobins = false
authors = ["Firecrawl Team"]
+2 -2
View File
@@ -851,7 +851,7 @@ checksum = "384b8ab6d37215f3c5301a95a4accb5d64aa607f1fcb26a11b5303878451b4fe"
[[package]]
name = "pdf-inspector"
version = "1.14.0"
version = "1.14.1"
dependencies = [
"env_logger",
"include_dir",
@@ -867,7 +867,7 @@ dependencies = [
[[package]]
name = "pdf-inspector-napi"
version = "1.14.0"
version = "1.14.1"
dependencies = [
"napi",
"napi-build",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector-napi"
version = "1.14.0"
version = "1.14.1"
edition = "2021"
[lib]
+6 -6
View File
@@ -8,12 +8,12 @@
"@napi-rs/cli": "^3.4.1",
},
"optionalDependencies": {
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.0",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.0",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.0",
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.0",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.0",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.0",
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.1",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.1",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.1",
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.1",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.1",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.1",
},
},
},
+7 -7
View File
@@ -1,6 +1,6 @@
{
"name": "@firecrawl/pdf-inspector",
"version": "1.14.0",
"version": "1.14.1",
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
"main": "index.js",
"types": "index.d.ts",
@@ -52,11 +52,11 @@
"@napi-rs/cli": "^3.4.1"
},
"optionalDependencies": {
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.0",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.0",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.0",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.0",
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.0",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.0"
"@firecrawl/pdf-inspector-linux-x64-gnu": "1.14.1",
"@firecrawl/pdf-inspector-linux-x64-musl": "1.14.1",
"@firecrawl/pdf-inspector-linux-arm64-gnu": "1.14.1",
"@firecrawl/pdf-inspector-linux-arm64-musl": "1.14.1",
"@firecrawl/pdf-inspector-darwin-arm64": "1.14.1",
"@firecrawl/pdf-inspector-win32-x64-msvc": "1.14.1"
}
}
+1 -1
View File
@@ -6,7 +6,7 @@ build-backend = "maturin"
name = "pdf-inspector"
# Keep package versions in sync with `python3 scripts/version.py <version>`.
# CI publishes automatically when the synchronized change lands on main.
version = "1.14.0"
version = "1.14.1"
description = "Fast PDF inspection, classification, and text extraction with smart scanned vs text-based detection"
readme = "docs/python.md"
license = { text = "MIT" }
+1 -1
View File
@@ -975,7 +975,7 @@ result = pdf_inspector.<span class="fn">process_pdf</span>(<span class="str">"do
<script>
(() => {
const MAX_FILE_SIZE = 25 * 1024 * 1024;
const WASM_MODULE_URL = "https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@1.14.0/pdf_inspector_wasm.js";
const WASM_MODULE_URL = "https://cdn.jsdelivr.net/npm/@firecrawl/pdf-inspector-wasm@1.14.1/pdf_inspector_wasm.js";
const input = document.querySelector("#pdf-input");
const dropZone = document.querySelector("#drop-zone");
const filePanel = document.querySelector("#demo-file");
+328
View File
@@ -0,0 +1,328 @@
//! Bounded content-stream decoding.
//!
//! `lopdf::content::Content::decode` materializes every operator before any
//! caller can apply a limit. A compact page of `q Q` pairs can therefore
//! allocate hundreds of megabytes and abort. Count operators first (without
//! allocating `Operation` objects) and skip decode when the cap is exceeded.
use crate::PdfError;
use lopdf::content::Content;
/// Maximum content-stream operators decoded for a page or a single Form
/// XObject. Matches the previous post-decode skip threshold.
pub(crate) const MAX_PAGE_OPERATIONS: usize = 1_000_000;
/// Decode `data` unless it contains more than `max_operations` operators.
///
/// Returns `Ok(None)` when the stream exceeds the cap, so callers can skip
/// extraction without first allocating the operation vector.
pub(crate) fn decode_content_bounded(
data: &[u8],
max_operations: usize,
) -> Result<Option<Content>, PdfError> {
if content_exceeds_operation_limit(data, max_operations) {
return Ok(None);
}
Content::decode(data)
.map(Some)
.map_err(|e| PdfError::Parse(e.to_string()))
}
fn content_exceeds_operation_limit(data: &[u8], max_operations: usize) -> bool {
count_content_operators(data, max_operations.saturating_add(1)) > max_operations
}
/// Count operators using the same token rules as lopdf's content parser,
/// stopping at `limit`. Does not allocate `Operation` / `Object` values.
fn count_content_operators(data: &[u8], limit: usize) -> usize {
let mut i = 0;
let mut count = 0;
while i < data.len() && count < limit {
skip_content_space(data, &mut i);
if i >= data.len() {
break;
}
if data[i] == b'%' {
skip_comment(data, &mut i);
continue;
}
match data[i] {
b'(' => i = skip_literal_string(data, i),
b'<' => {
if data.get(i + 1) == Some(&b'<') {
i += 2;
} else {
i = skip_hex_string(data, i);
}
}
b'>' => {
i += 1;
if data.get(i) == Some(&b'>') {
i += 1;
}
}
b'[' | b']' => i += 1,
b'/' => skip_name(data, &mut i),
b'+' | b'-' | b'.' => skip_number(data, &mut i),
b if b.is_ascii_digit() => skip_number(data, &mut i),
b if is_operator_byte(b) => {
let start = i;
i += 1;
while i < data.len() && is_operator_byte(data[i]) {
i += 1;
}
let token = &data[start..i];
if token == b"true" || token == b"false" || token == b"null" {
continue;
}
count += 1;
if token == b"BI" && (i >= data.len() || is_content_space(data[i])) {
i = skip_inline_image_after_bi(data, i);
}
}
_ => i += 1,
}
}
count
}
fn is_content_space(b: u8) -> bool {
// PDF whitespace (ISO 32000): NUL, tab, LF, FF, CR, space. Names must
// stop on these so a following operator is not absorbed into `/Name`.
matches!(b, b'\0' | b'\t' | b'\n' | b'\x0c' | b'\r' | b' ')
}
fn is_operator_byte(b: u8) -> bool {
b.is_ascii_alphabetic() || matches!(b, b'*' | b'\'' | b'"')
}
fn is_delimiter(b: u8) -> bool {
matches!(
b,
b'(' | b')' | b'<' | b'>' | b'[' | b']' | b'{' | b'}' | b'/' | b'%'
)
}
fn skip_content_space(data: &[u8], i: &mut usize) {
while *i < data.len() && is_content_space(data[*i]) {
*i += 1;
}
}
fn skip_comment(data: &[u8], i: &mut usize) {
while *i < data.len() && data[*i] != b'\n' && data[*i] != b'\r' {
*i += 1;
}
}
fn skip_literal_string(data: &[u8], mut i: usize) -> usize {
let mut depth = 1i32;
i += 1;
while i < data.len() && depth > 0 {
match data[i] {
b'\\' => {
i += 1;
if i < data.len() {
i += 1;
}
}
b'(' => {
depth += 1;
i += 1;
}
b')' => {
depth -= 1;
i += 1;
}
_ => i += 1,
}
}
i
}
fn skip_hex_string(data: &[u8], mut i: usize) -> usize {
i += 1;
while i < data.len() && data[i] != b'>' {
i += 1;
}
if i < data.len() {
i += 1;
}
i
}
fn skip_name(data: &[u8], i: &mut usize) {
*i += 1;
while *i < data.len() && !is_content_space(data[*i]) && !is_delimiter(data[*i]) {
*i += 1;
}
}
fn skip_number(data: &[u8], i: &mut usize) {
if *i < data.len() && matches!(data[*i], b'+' | b'-') {
*i += 1;
}
while *i < data.len() && data[*i].is_ascii_digit() {
*i += 1;
}
if *i < data.len() && data[*i] == b'.' {
*i += 1;
while *i < data.len() && data[*i].is_ascii_digit() {
*i += 1;
}
}
}
/// After a `BI` operator, skip inline-image data through `EI`.
/// Uses the same PDF whitespace set as `is_content_space`. If `EI` is not
/// found, leave the cursor in place so later operators are still counted
/// (undercounting would let decode allocate the full vector).
fn skip_inline_image_after_bi(data: &[u8], mut i: usize) -> usize {
skip_content_space(data, &mut i);
let rest = &data[i..];
if let Some(pos) = rest.windows(4).position(|w| {
is_content_space(w[0]) && w[1] == b'E' && w[2] == b'I' && is_content_space(w[3])
}) {
return i + pos + 3;
}
i
}
#[cfg(test)]
mod tests {
use super::*;
fn lopdf_op_count(data: &[u8]) -> usize {
Content::decode(data)
.map(|c| c.operations.len())
.unwrap_or(0)
}
/// DoS safety: never report fewer operators than lopdf would allocate.
/// Overcount is acceptable (skip a page); undercount would re-open decode.
fn assert_count_does_not_undercount(data: &[u8]) {
let ours = count_content_operators(data, usize::MAX);
match Content::decode(data) {
Ok(content) => assert!(
ours >= content.operations.len(),
"undercount: ours={ours} lopdf={} for {:?}",
content.operations.len(),
String::from_utf8_lossy(data)
),
Err(_) => {}
}
}
#[test]
fn operator_count_matches_lopdf_for_typical_streams() {
let samples: &[&[u8]] = &[
b"q 1 0 0 1 0 0 cm BT /F1 12 Tf 72 720 Td (Hello) Tj ET Q",
b"q Q q Q",
b"BT /F1 12 Tf 12 TL 1 0 0 1 100 512 Tm (first) Tj (struck) ' ET",
b"1 0 0 rg 0 0 10 10 re f",
b"true false null q",
b"% comment\nq Q\n",
b"[ (a) 1 (b) ] TJ",
b"1 0 0 1 0 0 cm /Im0 Do",
];
for data in samples {
assert_eq!(
count_content_operators(data, usize::MAX),
lopdf_op_count(data),
"count mismatch for {}",
String::from_utf8_lossy(data)
);
}
}
#[test]
fn strings_and_comments_are_not_operators() {
let data = b"(q Q Tj) Tj % q Q\nET";
assert_eq!(
count_content_operators(data, usize::MAX),
lopdf_op_count(data)
);
assert_eq!(count_content_operators(data, usize::MAX), 2); // Tj, ET
}
#[test]
fn inline_image_counts_as_one_operator() {
let data = b"BI /W 2 /H 2 /CS /RGB /BPC 8 ID \x00\x01\x02\x03 EI q";
assert_eq!(
count_content_operators(data, usize::MAX),
lopdf_op_count(data)
);
assert_eq!(count_content_operators(data, usize::MAX), 2); // BI, q
}
#[test]
fn inline_image_ei_accepts_pdf_whitespace() {
let tab = b"BI /W 1 /H 1 ID \xff\tEI\t q Q";
let nul = b"BI /W 1 /H 1 ID \xff\x00EI\x00 q Q";
let ff = b"BI /W 1 /H 1 ID \xff\x0cEI\x0c q Q";
for data in [tab.as_slice(), nul.as_slice(), ff.as_slice()] {
assert_count_does_not_undercount(data);
assert!(
count_content_operators(data, usize::MAX) >= 3,
"BI plus following q Q must remain visible after EI, got {} for {:?}",
count_content_operators(data, usize::MAX),
String::from_utf8_lossy(data)
);
}
}
#[test]
fn decode_is_skipped_when_operator_cap_is_exceeded() {
let mut data = Vec::new();
for _ in 0..20 {
data.extend_from_slice(b"q Q\n");
}
assert!(decode_content_bounded(&data, 10).unwrap().is_none());
let decoded = decode_content_bounded(&data, 50).unwrap().unwrap();
assert_eq!(decoded.operations.len(), 40);
}
#[test]
fn name_whitespace_does_not_swallow_following_operator() {
// NUL / form-feed end a name (PDF whitespace). Absorbing `q` into
// `/x` would undercount and let decode allocate the operator vector.
let mut nul_sep = Vec::new();
let mut ff_sep = Vec::new();
for _ in 0..8_000 {
nul_sep.extend_from_slice(b"/x\x00q");
ff_sep.extend_from_slice(b"/x\x0cq");
}
assert_count_does_not_undercount(&nul_sep);
assert_count_does_not_undercount(&ff_sep);
assert!(count_content_operators(&ff_sep, usize::MAX) >= 8_000);
}
#[test]
fn edge_streams_do_not_undercount_vs_lopdf() {
let samples: &[&[u8]] = &[
b".5 0 0 .5 0 0 cm",
b"+1 -2 3.0 rg",
b"<0041> Tj",
b"(unbalanced",
b"BI /W 1 /H 1 ID \xff\xff no EI here q Q q Q",
b"q\x00Q\x00q\x00Q",
b"/F1\x0c12 Tf (Hi) Tj",
b"{ 1 2 add } cvx",
];
for data in samples {
assert_count_does_not_undercount(data);
}
}
#[test]
fn million_q_pairs_are_rejected_without_decode() {
let mut data = Vec::with_capacity((MAX_PAGE_OPERATIONS + 1) * 2);
for _ in 0..=MAX_PAGE_OPERATIONS {
data.extend_from_slice(b"q\n");
}
assert!(content_exceeds_operation_limit(&data, MAX_PAGE_OPERATIONS));
assert!(decode_content_bounded(&data, MAX_PAGE_OPERATIONS)
.unwrap()
.is_none());
}
}
+41 -15
View File
@@ -19,7 +19,7 @@ use super::fonts::{
CMapDecisionCache, FontStyleCache,
};
use super::underline::UnderlineLine;
use super::xobjects::{extract_form_xobject_text, get_page_xobjects, XObjectType};
use super::xobjects::{extract_form_xobject_text, get_page_xobjects, FormWalkBudget, XObjectType};
use super::{get_number, image_bbox_from_ctm, multiply_matrices};
/// Strip PDF comments (% to end of line) from content stream bytes.
@@ -149,9 +149,8 @@ pub(crate) fn extract_page_text_items(
font_cmaps: &FontCMaps,
include_invisible: bool,
style_cache: &mut FontStyleCache,
form_budget: &mut FormWalkBudget,
) -> Result<(PageExtraction, bool, bool, bool), PdfError> {
use lopdf::content::Content;
let mut items = Vec::new();
let mut rects: Vec<PdfRect> = Vec::new();
let mut clip_rects: Vec<PdfRect> = Vec::new();
@@ -255,18 +254,20 @@ pub(crate) fn extract_page_text_items(
// Content::decode parser, causing it to skip operators like ET and Q.
let content_data = strip_pdf_comments(&content_data);
let content = Content::decode(&content_data).map_err(|e| PdfError::Parse(e.to_string()))?;
const MAX_OPERATIONS: usize = 1_000_000;
if content.operations.len() > MAX_OPERATIONS {
log::warn!(
"page {}: skipping extraction — {} operations exceeds limit ({})",
page_num,
content.operations.len(),
MAX_OPERATIONS
);
return Ok(((Vec::new(), Vec::new(), Vec::new()), false, false, false));
}
let content = match super::content_decode::decode_content_bounded(
&content_data,
super::content_decode::MAX_PAGE_OPERATIONS,
)? {
Some(content) => content,
None => {
log::warn!(
"page {}: skipping extraction — content stream exceeds {} operations",
page_num,
super::content_decode::MAX_PAGE_OPERATIONS
);
return Ok(((Vec::new(), Vec::new(), Vec::new()), false, false, false));
}
};
// Graphics state tracking
let mut ctm = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0]; // Current Transformation Matrix
@@ -916,6 +917,7 @@ pub(crate) fn extract_page_text_items(
&ctm,
&mut cmap_decisions,
style_cache,
form_budget,
);
items.extend(form_items);
}
@@ -1285,6 +1287,12 @@ pub(crate) fn extract_page_text_items(
}
}
if form_budget.was_truncated() {
log::warn!(
"page {page_num}: Form XObject expansion truncated (invocation or operation budget reached); nested form text may be incomplete"
);
}
// Underline detection reads only painted ink: `re` rects confirmed by
// a paint operator plus filled-subpath rects — never clip-only rects,
// which draw nothing.
@@ -1544,6 +1552,7 @@ mod tests {
&font_cmaps,
false,
&mut FontStyleCache::new(),
&mut FormWalkBudget::new(),
)
.unwrap();
items
@@ -1773,6 +1782,7 @@ BT /F1 12 Tf 0 1 -1 0 240 100 Tm (WORLD) Tj ET
&font_cmaps,
false,
&mut FontStyleCache::new(),
&mut FormWalkBudget::new(),
)
.unwrap();
let ((items, rects, lines), _has_gid, _coords_rotated, _skipped_invisible) = result;
@@ -1863,6 +1873,7 @@ BT 30 700 Tm <41> Tj ET";
&font_cmaps,
false,
&mut FontStyleCache::new(),
&mut FormWalkBudget::new(),
)
.unwrap();
let text = items
@@ -1926,4 +1937,19 @@ BT 30 700 Tm <41> Tj ET";
let output = strip_pdf_comments(input);
assert_eq!(output, b"(x\\\\) Tj \nET\n");
}
#[test]
fn oversized_content_stream_skips_extraction() {
let mut content =
Vec::with_capacity((super::super::content_decode::MAX_PAGE_OPERATIONS + 1) * 2);
for _ in 0..=super::super::content_decode::MAX_PAGE_OPERATIONS {
content.extend_from_slice(b"q\n");
}
content.extend_from_slice(b"BT /F1 12 Tf 72 720 Td (Hello) Tj ET\n");
let items = extract_simple_items(&content);
assert!(
items.is_empty(),
"pages over the operator cap must not be decoded"
);
}
}
+107 -16
View File
@@ -482,7 +482,11 @@ pub(crate) fn parse_cid_w_array(
widths: &mut HashMap<u16, u16>,
) {
let mut i = 0;
let mut assigned = 0usize;
while i < w_array.len() {
if assigned >= crate::tounicode::MAX_CID_W_EXPANSION {
return;
}
let start_cid = match &w_array[i] {
Object::Integer(n) => *n as u16,
Object::Real(n) => *n as u16,
@@ -501,12 +505,14 @@ pub(crate) fn parse_cid_w_array(
Object::Array(arr) => {
// [c [w1 w2 ...]] — consecutive widths starting at c
for (j, w_obj) in arr.iter().enumerate() {
let w = match w_obj {
Object::Integer(n) => *n as u16,
Object::Real(n) => *n as u16,
_ => continue,
};
widths.insert(start_cid + j as u16, w);
if !assign_cid_width(
widths,
start_cid.wrapping_add(j as u16),
w_obj,
&mut assigned,
) {
return;
}
}
i += 1;
}
@@ -514,12 +520,14 @@ pub(crate) fn parse_cid_w_array(
// Could be a reference to an array
if let Ok(Object::Array(arr)) = doc.get_object(*r) {
for (j, w_obj) in arr.iter().enumerate() {
let w = match w_obj {
Object::Integer(n) => *n as u16,
Object::Real(n) => *n as u16,
_ => continue,
};
widths.insert(start_cid + j as u16, w);
if !assign_cid_width(
widths,
start_cid.wrapping_add(j as u16),
w_obj,
&mut assigned,
) {
return;
}
}
i += 1;
} else {
@@ -542,8 +550,8 @@ pub(crate) fn parse_cid_w_array(
continue;
}
};
for cid in start_cid..=end {
widths.insert(cid, w);
if !assign_cid_width_range(widths, start_cid, end, w, &mut assigned) {
return;
}
i += 1;
}
@@ -561,8 +569,8 @@ pub(crate) fn parse_cid_w_array(
continue;
}
};
for cid in start_cid..=end {
widths.insert(cid, w);
if !assign_cid_width_range(widths, start_cid, end, w, &mut assigned) {
return;
}
i += 1;
}
@@ -573,6 +581,45 @@ pub(crate) fn parse_cid_w_array(
}
}
fn assign_cid_width(
widths: &mut HashMap<u16, u16>,
cid: u16,
w_obj: &Object,
assigned: &mut usize,
) -> bool {
let w = match w_obj {
Object::Integer(n) => *n as u16,
Object::Real(n) => *n as u16,
_ => return true,
};
if *assigned >= crate::tounicode::MAX_CID_W_EXPANSION {
return false;
}
widths.insert(cid, w);
*assigned += 1;
true
}
fn assign_cid_width_range(
widths: &mut HashMap<u16, u16>,
start: u16,
end: u16,
w: u16,
assigned: &mut usize,
) -> bool {
if start > end {
return true;
}
for cid in start..=end {
if *assigned >= crate::tounicode::MAX_CID_W_EXPANSION {
return false;
}
widths.insert(cid, w);
*assigned += 1;
}
true
}
/// Compute the width of a string in text space units,
/// given raw bytes and font width info.
/// Returns width in text space units (font_units * units_scale * font_size).
@@ -2313,4 +2360,48 @@ end",
// invalid CMap result — so it must not clear the gid flag.
assert!(gid_flagged(Some("<01> <FFFD>\n<02> <FFFD>")));
}
#[test]
fn parse_cid_w_array_range_and_consecutive() {
use super::parse_cid_w_array;
use lopdf::{Document, Object};
use std::collections::HashMap;
let doc = Document::new();
let mut widths = HashMap::new();
let w = vec![
Object::Integer(10),
Object::Integer(12),
Object::Integer(500),
Object::Integer(20),
Object::Array(vec![Object::Integer(100), Object::Integer(200)]),
];
parse_cid_w_array(&doc, &w, &mut widths);
assert_eq!(widths.get(&10), Some(&500));
assert_eq!(widths.get(&11), Some(&500));
assert_eq!(widths.get(&12), Some(&500));
assert_eq!(widths.get(&20), Some(&100));
assert_eq!(widths.get(&21), Some(&200));
}
#[test]
fn parse_cid_w_array_repeated_full_ranges_stay_bounded() {
use super::parse_cid_w_array;
use crate::tounicode::MAX_CID_W_EXPANSION;
use lopdf::{Document, Object};
use std::collections::HashMap;
let doc = Document::new();
let mut widths = HashMap::new();
let mut w = Vec::new();
for _ in 0..5_000 {
w.push(Object::Integer(0));
w.push(Object::Integer(65535));
w.push(Object::Integer(500));
}
parse_cid_w_array(&doc, &w, &mut widths);
assert!(widths.len() <= MAX_CID_W_EXPANSION);
assert_eq!(widths.get(&0), Some(&500));
assert_eq!(widths.get(&65535), Some(&500));
}
}
+3
View File
@@ -3,6 +3,7 @@
//! This module extracts text with position information for structure detection.
mod base14;
mod content_decode;
pub(crate) mod content_stream;
mod fonts;
mod layout;
@@ -37,6 +38,7 @@ pub(crate) use layout::group_prefiltered_items_into_lines_with_thresholds_and_re
pub(crate) use layout::is_newspaper_layout;
pub(crate) use layout::ColumnRegion;
pub use layout::{group_into_lines, group_into_lines_preserving_all_text};
pub(crate) use xobjects::FormWalkBudget;
// ---------------------------------------------------------------------------
// Public API
@@ -282,6 +284,7 @@ fn extract_positioned_text_impl(
font_cmaps,
include_invisible,
&mut style_cache,
&mut FormWalkBudget::new(),
);
let ((mut items, mut rects, mut lines), has_gid_fonts, coords_rotated, _skipped_invisible) =
match page_result {
+600 -21
View File
@@ -16,6 +16,76 @@ use super::{get_number, image_bbox_from_ctm, multiply_matrices};
const MAX_FORM_XOBJECT_DEPTH: u8 = 5;
/// Upper bound on Form XObject invocations during a single page extraction.
/// Depth alone is not enough: an acyclic DAG where each form invokes the next
/// N times expands to N^depth work before the depth cap is reached.
const MAX_FORM_XOBJECT_INVOCATIONS: usize = 10_000;
/// Upper bound on content-stream operations walked across all Form XObject
/// expansions for a page. Nested forms are decoded independently of the
/// page-level operation cap, so this keeps total form work in the same
/// ballpark as that page cap.
const MAX_FORM_XOBJECT_OPERATIONS: usize = 1_000_000;
/// Shared budget for Form XObject expansion on a page. Bounds both nested DAG
/// expansion and repeated sibling `/Do` invocations of the same form.
pub(crate) struct FormWalkBudget {
invocations: usize,
operations: usize,
max_invocations: usize,
max_operations: usize,
truncated: bool,
}
impl FormWalkBudget {
pub(crate) fn new() -> Self {
Self::with_limits(MAX_FORM_XOBJECT_INVOCATIONS, MAX_FORM_XOBJECT_OPERATIONS)
}
fn with_limits(max_invocations: usize, max_operations: usize) -> Self {
Self {
invocations: 0,
operations: 0,
max_invocations,
max_operations,
truncated: false,
}
}
fn exhausted(&mut self) -> bool {
if self.invocations >= self.max_invocations || self.operations >= self.max_operations {
self.truncated = true;
true
} else {
false
}
}
fn charge_invocation(&mut self) -> bool {
if self.exhausted() {
return false;
}
self.invocations += 1;
true
}
/// Charge one walked content-stream operator. Independent of the
/// invocation cap so a form that was already admitted can finish its
/// stream (up to the operation cap).
fn charge_operation(&mut self) -> bool {
if self.operations >= self.max_operations {
self.truncated = true;
return false;
}
self.operations += 1;
true
}
pub(crate) fn was_truncated(&self) -> bool {
self.truncated
}
}
pub(crate) enum XObjectType {
Image,
Form(ObjectId),
@@ -109,6 +179,7 @@ fn collect_xobjects_from_dict(
}
/// Extract text items from a Form XObject
#[allow(clippy::too_many_arguments)]
pub(crate) fn extract_form_xobject_text(
doc: &Document,
form_id: ObjectId,
@@ -117,6 +188,7 @@ pub(crate) fn extract_form_xobject_text(
parent_ctm: &[f32; 6],
cmap_decisions: &mut CMapDecisionCache,
style_cache: &mut FontStyleCache,
budget: &mut FormWalkBudget,
) -> Vec<TextItem> {
extract_form_xobject_text_inner(
doc,
@@ -127,6 +199,7 @@ pub(crate) fn extract_form_xobject_text(
cmap_decisions,
style_cache,
0,
budget,
)
}
@@ -140,11 +213,14 @@ fn extract_form_xobject_text_inner(
cmap_decisions: &mut CMapDecisionCache,
style_cache: &mut FontStyleCache,
depth: u8,
budget: &mut FormWalkBudget,
) -> Vec<TextItem> {
use lopdf::content::Content;
let mut items = Vec::new();
if !budget.charge_invocation() {
return items;
}
// Get the Form XObject stream
let Ok(Object::Stream(stream)) = doc.get_object(form_id) else {
return items;
@@ -156,8 +232,12 @@ fn extract_form_xobject_text_inner(
Err(_) => stream.content.clone(),
};
// Decode the content stream
let Ok(content) = Content::decode(&content_data) else {
// Decode the content stream. Cap before lopdf materializes the operator
// vector — the walk budget cannot help if decode itself allocates first.
let Ok(Some(content)) = super::content_decode::decode_content_bounded(
&content_data,
super::content_decode::MAX_PAGE_OPERATIONS,
) else {
return items;
};
@@ -246,19 +326,55 @@ fn extract_form_xobject_text_inner(
let mut current_font = String::new();
let mut current_font_size: f32 = 12.0;
let mut text_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
// Text line matrix (TLM) — Td/TD/T* move relative to the start of the
// current line, not to the position left by the last show operator.
let mut line_matrix = [1.0f32, 0.0, 0.0, 1.0, 0.0, 0.0];
let mut text_leading: f32 = 0.0; // TL parameter (text-space units)
let mut char_spacing: f32 = 0.0; // Tc parameter
let mut word_spacing: f32 = 0.0; // Tw parameter
let mut in_text_block = false;
let mut fill_is_white = false;
let mut ctm = base_ctm;
let mut ctm_stack: Vec<[f32; 6]> = Vec::new();
// Text state (Tc/Tw/TL/Tf) and the fill colour are part of the graphics
// state and must be saved/restored by q/Q alongside the CTM.
#[derive(Clone)]
struct GraphicsState {
ctm: [f32; 6],
char_spacing: f32,
word_spacing: f32,
text_leading: f32,
current_font: String,
current_font_size: f32,
fill_is_white: bool,
}
let mut ctm_stack: Vec<GraphicsState> = Vec::new();
for op in &content.operations {
if !budget.charge_operation() {
break;
}
match op.operator.as_str() {
"q" => {
ctm_stack.push(ctm);
ctm_stack.push(GraphicsState {
ctm,
char_spacing,
word_spacing,
text_leading,
current_font: current_font.clone(),
current_font_size,
fill_is_white,
});
}
"Q" => {
if let Some(saved) = ctm_stack.pop() {
ctm = saved;
ctm = saved.ctm;
char_spacing = saved.char_spacing;
word_spacing = saved.word_spacing;
text_leading = saved.text_leading;
current_font = saved.current_font;
current_font_size = saved.current_font_size;
fill_is_white = saved.fill_is_white;
}
}
"cm" => {
@@ -276,7 +392,7 @@ fn extract_form_xobject_text_inner(
let xobj_name = String::from_utf8_lossy(name).to_string();
match form_xobjects.get(&xobj_name) {
Some(XObjectType::Form(nested_id)) => {
if depth < MAX_FORM_XOBJECT_DEPTH {
if depth < MAX_FORM_XOBJECT_DEPTH && !budget.exhausted() {
let nested_items = extract_form_xobject_text_inner(
doc,
*nested_id,
@@ -286,6 +402,7 @@ fn extract_form_xobject_text_inner(
cmap_decisions,
style_cache,
depth + 1,
budget,
);
items.extend(nested_items);
}
@@ -321,6 +438,7 @@ fn extract_form_xobject_text_inner(
"BT" => {
in_text_block = true;
text_matrix = [1.0, 0.0, 0.0, 1.0, 0.0, 0.0];
line_matrix = text_matrix;
}
"ET" => {
in_text_block = false;
@@ -333,12 +451,33 @@ fn extract_form_xobject_text_inner(
current_font_size = get_number(&op.operands[1]).unwrap_or(12.0);
}
}
"TL" => {
// Set text leading (used by T*, ', and ")
if let Some(tl) = op.operands.first().and_then(get_number) {
text_leading = tl;
}
}
"Tc" => {
if let Some(tc) = op.operands.first().and_then(get_number) {
char_spacing = tc;
}
}
"Tw" => {
if let Some(tw) = op.operands.first().and_then(get_number) {
word_spacing = tw;
}
}
"Td" | "TD" => {
// Move text position: TLM = T(tx,ty) x TLM; Tm = TLM
if op.operands.len() >= 2 {
let tx = get_number(&op.operands[0]).unwrap_or(0.0);
let ty = get_number(&op.operands[1]).unwrap_or(0.0);
text_matrix[4] += tx * text_matrix[0] + ty * text_matrix[2];
text_matrix[5] += tx * text_matrix[1] + ty * text_matrix[3];
line_matrix[4] += tx * line_matrix[0] + ty * line_matrix[2];
line_matrix[5] += tx * line_matrix[1] + ty * line_matrix[3];
text_matrix = line_matrix;
if op.operator == "TD" {
text_leading = -ty;
}
}
}
"Tm" => {
@@ -347,8 +486,20 @@ fn extract_form_xobject_text_inner(
text_matrix[i] =
get_number(operand).unwrap_or(if i == 0 || i == 3 { 1.0 } else { 0.0 });
}
line_matrix = text_matrix;
}
}
"T*" => {
// Move to start of next line: equivalent to `0 -TL Td`
let tl = if text_leading != 0.0 {
text_leading
} else {
current_font_size * 1.2
};
line_matrix[4] += (-tl) * line_matrix[2];
line_matrix[5] += (-tl) * line_matrix[3];
text_matrix = line_matrix;
}
"g" => {
if let Some(gray) = op.operands.first().and_then(get_number) {
fill_is_white = gray > 0.95;
@@ -384,17 +535,33 @@ fn extract_form_xobject_text_inner(
_ => fill_is_white = false,
}
}
"Tj" => {
if in_text_block && !op.operands.is_empty() {
"Tj" | "'" | "\"" => {
// `'` = move to next line then show; `"` = set word/char spacing,
// move to next line, then show (string is the last operand).
if op.operator != "Tj" {
if op.operator == "\"" && op.operands.len() >= 3 {
word_spacing = get_number(&op.operands[0]).unwrap_or(word_spacing);
char_spacing = get_number(&op.operands[1]).unwrap_or(char_spacing);
}
let tl = if text_leading != 0.0 {
text_leading
} else {
current_font_size * 1.2
};
line_matrix[4] += (-tl) * line_matrix[2];
line_matrix[5] += (-tl) * line_matrix[3];
text_matrix = line_matrix;
}
if let (true, Some(show_operand)) = (in_text_block, op.operands.last()) {
if fill_is_white {
if let Some(font_info) = font_widths.get(&current_font) {
if let Some(raw_bytes) = get_operand_bytes(&op.operands[0]) {
if let Some(raw_bytes) = get_operand_bytes(show_operand) {
let w_ts = compute_string_width_ts(
raw_bytes,
font_info,
current_font_size,
0.0,
0.0,
char_spacing,
word_spacing,
);
text_matrix[4] += w_ts * text_matrix[0];
text_matrix[5] += w_ts * text_matrix[1];
@@ -403,7 +570,7 @@ fn extract_form_xobject_text_inner(
continue;
}
if let Some(text) = extract_text_from_operand(
&op.operands[0],
show_operand,
&current_font,
font_base_names.get(&current_font).map(|s| s.as_str()),
font_cmaps,
@@ -419,13 +586,13 @@ fn extract_form_xobject_text_inner(
* type3_scales.get(&current_font).copied().unwrap_or(1.0);
let (x, y) = (combined[4], combined[5]);
let width = if let Some(font_info) = font_widths.get(&current_font) {
if let Some(raw_bytes) = get_operand_bytes(&op.operands[0]) {
if let Some(raw_bytes) = get_operand_bytes(show_operand) {
let w_ts = compute_string_width_ts(
raw_bytes,
font_info,
current_font_size,
0.0,
0.0,
char_spacing,
word_spacing,
);
text_matrix[4] += w_ts * text_matrix[0];
text_matrix[5] += w_ts * text_matrix[1];
@@ -548,8 +715,8 @@ fn extract_form_xobject_text_inner(
raw_bytes,
fi,
current_font_size,
0.0,
0.0,
char_spacing,
word_spacing,
);
}
}
@@ -683,3 +850,415 @@ pub(crate) fn get_form_fonts<'a>(
fonts
}
#[cfg(test)]
mod tests {
use super::*;
use crate::extractor::content_stream::extract_page_text_items;
use lopdf::{dictionary, Dictionary, Stream};
/// Build an acyclic Form XObject DAG: `levels` form objects, each non-leaf
/// invoking the next form `branches` times. The leaf draws a single `(X)`.
/// Returns `(doc, root_form_id)`.
fn form_dag(branches: usize, levels: usize) -> (Document, ObjectId) {
assert!(levels >= 2);
let mut doc = Document::new();
let font_id = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Helvetica",
});
let ids: Vec<ObjectId> = (0..levels).map(|_| doc.new_object_id()).collect();
for level in 0..levels {
let stream = if level + 1 == levels {
Stream::new(
dictionary! {
"Type" => "XObject",
"Subtype" => "Form",
"BBox" => vec![0.into(), 0.into(), 100.into(), 100.into()],
"Resources" => dictionary! {
"Font" => dictionary! {
"F1" => Object::Reference(font_id),
},
},
},
b"BT /F1 10 Tf 10 10 Td (X) Tj ET\n".to_vec(),
)
} else {
let next_name = format!("Fm{}", level + 1);
let content = format!("/{next_name} Do\n").repeat(branches);
let mut xobjects = Dictionary::new();
xobjects.set(next_name, Object::Reference(ids[level + 1]));
let mut resources = Dictionary::new();
resources.set("XObject", Object::Dictionary(xobjects));
let mut dict = dictionary! {
"Type" => "XObject",
"Subtype" => "Form",
"BBox" => vec![0.into(), 0.into(), 100.into(), 100.into()],
};
dict.set("Resources", Object::Dictionary(resources));
Stream::new(dict, content.into_bytes())
};
doc.set_object(ids[level], Object::Stream(stream));
}
(doc, ids[0])
}
fn page_invoking_form(mut doc: Document, form_id: ObjectId) -> (Document, ObjectId) {
let content_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {},
b"/Fm0 Do\n".to_vec(),
)));
let page_id = doc.add_object(dictionary! {
"Type" => "Page",
"Contents" => Object::Reference(content_id),
"Resources" => dictionary! {
"XObject" => dictionary! {
"Fm0" => Object::Reference(form_id),
},
},
"MediaBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
});
let pages_id = doc.add_object(dictionary! {
"Type" => "Pages",
"Count" => Object::Integer(1),
"Kids" => vec![Object::Reference(page_id)],
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"Pages" => Object::Reference(pages_id),
});
doc.trailer.set("Root", Object::Reference(catalog_id));
(doc, page_id)
}
fn extract_form(
doc: &Document,
form_id: ObjectId,
budget: &mut FormWalkBudget,
) -> Vec<TextItem> {
extract_form_xobject_text(
doc,
form_id,
1,
&FontCMaps::from_doc(doc),
&[1.0, 0.0, 0.0, 1.0, 0.0, 0.0],
&mut CMapDecisionCache::new(),
&mut FontStyleCache::new(),
budget,
)
}
#[test]
fn nested_form_still_extracts_leaf_text() {
let (doc, root) = form_dag(1, 3);
let items = extract_form(&doc, root, &mut FormWalkBudget::new());
assert_eq!(items.len(), 1);
assert_eq!(items[0].text, "X");
}
#[test]
fn acyclic_form_dag_within_budget_keeps_all_leaves() {
// 4 sibling invocations across 4 nested levels → 4^4 leaf drawings.
// Default budgets are far above 256, so legitimate nesting is intact.
let (doc, root) = form_dag(4, 5);
let items = extract_form(&doc, root, &mut FormWalkBudget::new());
assert_eq!(items.len(), 4usize.pow(4));
assert!(items.iter().all(|item| item.text == "X"));
}
#[test]
fn acyclic_form_dag_stops_at_invocation_budget() {
// Same DAG as above would draw 256 leaves; a tiny invocation cap must
// stop expansion rather than walking the full tree.
let (doc, root) = form_dag(4, 5);
let mut budget = FormWalkBudget::with_limits(20, MAX_FORM_XOBJECT_OPERATIONS);
let items = extract_form(&doc, root, &mut budget);
assert!(
items.len() < 4usize.pow(4),
"invocation budget must truncate DAG expansion; got {} items",
items.len()
);
assert!(budget.was_truncated());
}
#[test]
fn form_operations_stop_at_budget() {
let mut doc = Document::new();
let font_id = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Helvetica",
});
let mut content = b"q Q\n".repeat(50);
content.extend_from_slice(b"BT /F1 10 Tf 10 10 Td (X) Tj ET\n");
let form_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {
"Type" => "XObject",
"Subtype" => "Form",
"BBox" => vec![0.into(), 0.into(), 100.into(), 100.into()],
"Resources" => dictionary! {
"Font" => dictionary! {
"F1" => Object::Reference(font_id),
},
},
},
content,
)));
let mut budget = FormWalkBudget::with_limits(MAX_FORM_XOBJECT_INVOCATIONS, 10);
let items = extract_form(&doc, form_id, &mut budget);
assert!(
items.is_empty(),
"operation budget must stop before the trailing text show"
);
assert!(budget.was_truncated());
}
#[test]
fn page_level_form_dag_stays_within_production_budget() {
// A page-level `/Do` of an 8-wide, 6-level Form DAG would expand to
// 8^5 = 32_768 leaf drawings without a budget. The production
// invocation cap must keep extraction bounded.
let (doc, root) = form_dag(8, 6);
let (doc, page_id) = page_invoking_form(doc, root);
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
&mut FormWalkBudget::new(),
)
.unwrap();
assert!(
items.len() <= MAX_FORM_XOBJECT_INVOCATIONS,
"page-level Form expansion must stay within the invocation cap; got {}",
items.len()
);
assert!(
!items.is_empty(),
"budget must still allow some nested form text through"
);
}
#[test]
fn shared_form_budget_spans_two_extraction_passes() {
// The invisible-layer retry calls extract_page_text_items twice for
// the same page; both passes must share one budget.
let (doc, root) = form_dag(1, 2);
let (doc, page_id) = page_invoking_form(doc, root);
let font_cmaps = FontCMaps::from_doc(&doc);
// Root + leaf = 2 invocations on the first pass.
let mut budget = FormWalkBudget::with_limits(2, MAX_FORM_XOBJECT_OPERATIONS);
let ((first, _, _), _, _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
&mut budget,
)
.unwrap();
assert_eq!(first.iter().filter(|item| item.text == "X").count(), 1);
assert!(!budget.was_truncated());
let ((second, _, _), _, _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
true,
&mut FontStyleCache::new(),
&mut budget,
)
.unwrap();
assert!(
second.iter().all(|item| item.text != "X"),
"second pass must not get a fresh invocation budget"
);
assert!(budget.was_truncated());
}
/// Build a document whose page draws *all* of its content through a single
/// Form XObject — the shape emitted by print-to-PDF producers like PDFlib,
/// where the page stream itself is only `q /X1 Do Q`.
fn doc_with_form_content(form_content: &[u8]) -> (Document, ObjectId) {
let mut doc = Document::new();
let widths: Vec<Object> = (0..=255).map(|_| 600.into()).collect();
let font_id = doc.add_object(dictionary! {
"Type" => "Font",
"Subtype" => "Type1",
"BaseFont" => "Helvetica",
"FirstChar" => 0,
"LastChar" => 255,
"Widths" => Object::Array(widths),
});
let form_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {
"Type" => "XObject",
"Subtype" => "Form",
"BBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
"Resources" => dictionary! {
"Font" => dictionary! { "F1" => Object::Reference(font_id) },
},
},
form_content.to_vec(),
)));
let content_id = doc.add_object(Object::Stream(Stream::new(
dictionary! {},
b"q /X1 Do Q".to_vec(),
)));
let page_id = doc.add_object(dictionary! {
"Type" => "Page",
"Contents" => Object::Reference(content_id),
"Resources" => dictionary! {
"XObject" => dictionary! { "X1" => Object::Reference(form_id) },
},
"MediaBox" => vec![0.into(), 0.into(), 612.into(), 792.into()],
});
let pages_id = doc.add_object(dictionary! {
"Type" => "Pages",
"Count" => Object::Integer(1),
"Kids" => vec![Object::Reference(page_id)],
});
let catalog_id = doc.add_object(dictionary! {
"Type" => "Catalog",
"Pages" => Object::Reference(pages_id),
});
doc.trailer.set("Root", Object::Reference(catalog_id));
(doc, page_id)
}
fn form_items(form_content: &[u8]) -> Vec<TextItem> {
let (doc, page_id) = doc_with_form_content(form_content);
let font_cmaps = FontCMaps::from_doc(&doc);
let ((items, _, _), _, _, _) = extract_page_text_items(
&doc,
page_id,
1,
&font_cmaps,
false,
&mut FontStyleCache::new(),
&mut FormWalkBudget::new(),
)
.unwrap();
items
}
fn find<'a>(items: &'a [TextItem], text: &str) -> &'a TextItem {
items
.iter()
.find(|item| item.text == text)
.unwrap_or_else(|| {
let found: Vec<&String> = items.iter().map(|i| &i.text).collect();
panic!("no item {text:?} in {found:?}")
})
}
#[test]
fn t_star_inside_form_moves_to_next_line() {
// T* was previously unhandled inside Form XObjects, so every line after
// the first piled onto the preceding baseline and drifted right.
let items =
form_items(b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm (first) Tj T* (second) Tj ET");
let first = find(&items, "first");
let second = find(&items, "second");
assert!((first.y - 700.0).abs() < 0.1, "first y = {}", first.y);
assert!((second.y - 688.0).abs() < 0.1, "second y = {}", second.y);
assert!((second.x - 100.0).abs() < 0.1, "second x = {}", second.x);
}
#[test]
fn td_inside_form_is_relative_to_line_start_not_shown_text() {
// Td moves relative to the text *line* matrix. Applying it to the
// matrix already advanced by Tj marched each line off the right edge.
let items = form_items(b"BT /F1 12 Tf 1 0 0 1 100 700 Tm (AAAAA) Tj 0 -12 Td (B) Tj ET");
let b = find(&items, "B");
assert!((b.x - 100.0).abs() < 0.1, "B x = {} (expected 100)", b.x);
assert!((b.y - 688.0).abs() < 0.1, "B y = {}", b.y);
}
#[test]
fn td_inside_form_sets_leading_for_later_t_star() {
// `TD` sets the leading to -ty as a side effect; a following T* must
// reuse it.
let items = form_items(
b"BT /F1 12 Tf 1 0 0 1 100 700 Tm (one) Tj 0 -15 TD (two) Tj T* (three) Tj ET",
);
assert!((find(&items, "two").y - 685.0).abs() < 0.1);
let three = find(&items, "three");
assert!((three.y - 670.0).abs() < 0.1, "three y = {}", three.y);
assert!((three.x - 100.0).abs() < 0.1, "three x = {}", three.x);
}
#[test]
fn quote_operator_inside_form_moves_to_next_line() {
let items = form_items(b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm (first) Tj (second) ' ET");
let second = find(&items, "second");
assert!((second.y - 688.0).abs() < 0.1, "second y = {}", second.y);
assert!((second.x - 100.0).abs() < 0.1, "second x = {}", second.x);
}
#[test]
fn double_quote_operator_inside_form_sets_spacing_and_moves() {
// `aw ac (string) "` — set word spacing and char spacing, then T* and show.
let items =
form_items(b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm (first) Tj 0 0 (second) \" ET");
let second = find(&items, "second");
assert!((second.y - 688.0).abs() < 0.1, "second y = {}", second.y);
assert!((second.x - 100.0).abs() < 0.1, "second x = {}", second.x);
}
#[test]
fn char_spacing_inside_form_widens_advance() {
// Tc was hardcoded to 0 in the form parser, so advance widths drifted.
// 2 glyphs x 600/1000 x 12pt = 14.4, plus 2 x Tc(2.0) = 18.4.
let items = form_items(b"BT /F1 12 Tf 1 0 0 1 100 700 Tm 2 Tc (AB) Tj ET");
let ab = find(&items, "AB");
assert!((ab.width - 18.4).abs() < 0.1, "AB width = {}", ab.width);
}
#[test]
fn q_restores_fill_colour_inside_form() {
// A white fill set inside q/Q must not leak past the Q — otherwise the
// following black text is treated as invisible and dropped entirely.
let items = form_items(
b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm q 1 g (hidden) Tj Q T* (visible) Tj ET",
);
assert!(
items.iter().any(|item| item.text == "visible"),
"text after Q was dropped: {:?}",
items.iter().map(|i| &i.text).collect::<Vec<_>>()
);
assert!(
!items.iter().any(|item| item.text == "hidden"),
"white-filled text should still be suppressed"
);
}
#[test]
fn q_restores_text_state_inside_form() {
// Tc/TL live in the graphics state; `Q` must roll them back.
let items =
form_items(b"BT /F1 12 Tf 12 TL 1 0 0 1 100 700 Tm q 30 TL (a) Tj Q T* (b) Tj ET");
let b = find(&items, "b");
assert!(
(b.y - 688.0).abs() < 0.1,
"b y = {} (leading should restore to 12)",
b.y
);
}
}
+12 -1
View File
@@ -826,7 +826,10 @@ pub fn extract_text_in_regions_mem(
let height = get_page_height(&doc, page_id).unwrap_or(792.0);
page_heights.insert(*page_num, height);
// Extract text items for this page
// Extract text items for this page. The Form XObject budget is shared
// with the invisible-layer retry below so one page cannot consume two
// full expansion budgets.
let mut form_budget = extractor::FormWalkBudget::new();
let ((mut items, _rects, _lines), mut has_gid, mut coords_rotated, skipped_invisible) =
extractor::content_stream::extract_page_text_items(
&doc,
@@ -835,6 +838,7 @@ pub fn extract_text_in_regions_mem(
&font_cmaps,
false,
&mut style_cache,
&mut form_budget,
)?;
// OCR-layer fallback: scanned pages often carry their text as an
// invisible (Tr 3) layer behind the page raster. The visible-only
@@ -863,6 +867,7 @@ pub fn extract_text_in_regions_mem(
&font_cmaps,
true,
&mut style_cache,
&mut form_budget,
)
{
let inv_alnum = non_placeholder_alnum(&inv_items);
@@ -1045,6 +1050,7 @@ pub fn extract_tables_in_regions_mem(
&font_cmaps,
false,
&mut style_cache,
&mut extractor::FormWalkBudget::new(),
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
@@ -1356,6 +1362,7 @@ pub fn detect_vector_grid_in_region_mem(
&font_cmaps,
false,
&mut extractor::FontStyleCache::new(),
&mut extractor::FormWalkBudget::new(),
)?;
text_utils::fix_letterspaced_items(&mut items);
@@ -1550,6 +1557,7 @@ mod vector_grid_tests {
&cmaps,
false,
&mut crate::extractor::FontStyleCache::new(),
&mut crate::extractor::FormWalkBudget::new(),
)
.unwrap();
@@ -1593,6 +1601,7 @@ mod vector_grid_tests {
&cmaps,
false,
&mut crate::extractor::FontStyleCache::new(),
&mut crate::extractor::FormWalkBudget::new(),
)
.unwrap();
@@ -2330,6 +2339,7 @@ pub fn extract_tables_with_structure_cells_mem(
&font_cmaps,
false,
&mut style_cache,
&mut extractor::FormWalkBudget::new(),
)?;
let threshold = text_utils::fix_letterspaced_items(&mut items);
if threshold > 0.10 {
@@ -3132,6 +3142,7 @@ fn detect_tsr_quality_issue(
&font_cmaps,
false,
&mut extractor::FontStyleCache::new(),
&mut extractor::FormWalkBudget::new(),
)?;
let adaptive_threshold = text_utils::fix_letterspaced_items(&mut items);
let coords = if coords_rotated {
+73 -12
View File
@@ -1832,6 +1832,11 @@ fn merge_cmaps(mut base: ToUnicodeCMap, overlay: ToUnicodeCMap) -> ToUnicodeCMap
base
}
/// Upper bound on CID `/W` range expansion. The CID domain is 16-bit, so more
/// than 65,536 unique keys cannot exist; repeating full-width ranges must not
/// re-expand the same domain.
pub(crate) const MAX_CID_W_EXPANSION: usize = 65_536;
/// Check if a CIDFont's /W (widths) array contains CID values that look like
/// Unicode codepoints rather than low-value GIDs.
///
@@ -1843,20 +1848,23 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
_ => return false,
};
// The /W array format: [cid [w1 w2 ...]] or [cid_start cid_end w]
// We extract all CID values (the first element of each group).
let mut cids: Vec<u16> = Vec::new();
// The /W array format: [cid [w1 w2 ...]] or [cid_start cid_end w].
// Collect unique CIDs only: repeating a full-width range must not grow a
// temporary vector (or the sort) with the range length on every copy.
let mut seen = HashSet::new();
let mut i = 0;
while i < w_arr.len() {
while i < w_arr.len() && seen.len() < MAX_CID_W_EXPANSION {
if let Ok(cid) = w_arr[i].as_i64() {
cids.push(cid as u16);
// Skip the width data
let start = cid as u16;
if i + 1 < w_arr.len() {
match &w_arr[i + 1] {
Object::Array(widths) => {
// [cid [w1 w2 ...]] — CIDs are cid, cid+1, ..., cid+len-1
for j in 1..widths.len() {
cids.push((cid as u16).wrapping_add(j as u16));
for j in 0..widths.len() {
if seen.len() >= MAX_CID_W_EXPANSION {
break;
}
seen.insert(start.wrapping_add(j as u16));
}
i += 2;
}
@@ -1864,9 +1872,7 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
// [cid_start cid_end w] — range of CIDs
if i + 2 < w_arr.len() {
if let Ok(cid_end) = w_arr[i + 1].as_i64() {
for c in (cid as u16)..=(cid_end as u16) {
cids.push(c);
}
record_unique_cid_range(start, cid_end as u16, &mut seen);
}
i += 3;
} else {
@@ -1875,6 +1881,7 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
}
}
} else {
seen.insert(start);
i += 1;
}
} else {
@@ -1882,10 +1889,11 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
}
}
if cids.is_empty() {
if seen.is_empty() {
return false;
}
let mut cids: Vec<u16> = seen.into_iter().collect();
cids.sort_unstable();
let median = cids[cids.len() / 2];
// Unicode text CIDs are typically >= 0x20 (space) with letters at 0x41+.
@@ -1894,6 +1902,18 @@ pub(crate) fn cid_values_look_like_unicode(cid_font_dict: &lopdf::Dictionary) ->
median >= 0x41
}
fn record_unique_cid_range(start: u16, end: u16, seen: &mut HashSet<u16>) {
if start > end {
return;
}
for cid in start..=end {
if seen.len() >= MAX_CID_W_EXPANSION {
return;
}
seen.insert(cid);
}
}
/// Build a ToUnicodeCMap from predefined CID→Unicode mapping based on CIDSystemInfo.
///
/// Supports Adobe-Korea1 (Korean) character collection. Can be extended for
@@ -3298,4 +3318,45 @@ endbfrange
"An indirect /Subtype naming CIDFontType2 must still reach the remap"
);
}
#[test]
fn cid_values_look_like_unicode_letter_range() {
let mut dict = lopdf::Dictionary::new();
dict.set(
"W",
Object::Array(vec![
Object::Integer(0x41),
Object::Integer(0x5A),
Object::Integer(500),
]),
);
assert!(cid_values_look_like_unicode(&dict));
}
#[test]
fn cid_values_look_like_unicode_low_gids() {
let mut dict = lopdf::Dictionary::new();
dict.set(
"W",
Object::Array(vec![
Object::Integer(0),
Object::Array(vec![Object::Integer(500); 10]),
]),
);
assert!(!cid_values_look_like_unicode(&dict));
}
#[test]
fn cid_values_look_like_unicode_repeated_full_ranges_stay_bounded() {
// Repeating `[0 65535 w]` must not materialize 65,536 CIDs per copy.
let mut w = Vec::new();
for _ in 0..5_000 {
w.push(Object::Integer(0));
w.push(Object::Integer(65535));
w.push(Object::Integer(500));
}
let mut dict = lopdf::Dictionary::new();
dict.set("W", Object::Array(w));
assert!(cid_values_look_like_unicode(&dict));
}
}
+2 -2
View File
@@ -724,7 +724,7 @@ checksum = "d6790f58c7ff633d8771f42965289203411a5e5c68388703c06e14f24770b41e"
[[package]]
name = "pdf-inspector"
version = "1.14.0"
version = "1.14.1"
dependencies = [
"env_logger",
"include_dir",
@@ -740,7 +740,7 @@ dependencies = [
[[package]]
name = "pdf-inspector-wasm"
version = "1.14.0"
version = "1.14.1"
dependencies = [
"console_error_panic_hook",
"js-sys",
+1 -1
View File
@@ -1,6 +1,6 @@
[package]
name = "pdf-inspector-wasm"
version = "1.14.0"
version = "1.14.1"
edition = "2021"
authors = ["Firecrawl Team"]
description = "Browser WebAssembly bindings for pdf-inspector"