fix(extractor): cap content-stream decode before allocating operators (#373)
* fix(extractor): cap content-stream decode before allocating operators The 1M operation limit ran after lopdf materialized the full vector, so a compact page of q/Q pairs could still abort under memory pressure. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(extractor): treat NUL and form-feed as PDF whitespace in op counting Names must stop on the full PDF whitespace set so a following operator is not absorbed into /Name, which would undercount and skip the decode cap. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(extractor): scan inline-image EI with the full PDF whitespace set A missed EI terminator used to consume the rest of the stream and drop later operators from the decode cap. If EI is absent, keep scanning. Co-authored-by: Cursor <cursoragent@cursor.com> --------- Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
co-authored by
Cursor
parent
ec6e54afb8
commit
076183e2e4
@@ -3,6 +3,7 @@
|
||||
//! This module extracts text with position information for structure detection.
|
||||
|
||||
mod base14;
|
||||
mod content_decode;
|
||||
pub(crate) mod content_stream;
|
||||
mod fonts;
|
||||
mod layout;
|
||||
|
||||
Reference in New Issue
Block a user