feat(extractor): geometric underline detection on TextItem (#116)
* feat(extractor): geometric underline detection on TextItem (ENG-5015) PDFs carry no underline font flag — underlines are stroked horizontal lines or thin filled rects drawn under the baseline. Correlate those graphics (already parsed from the content stream) with text items in a post-pass: a rule within ~0.35em below the baseline covering >=60% of an item's width marks is_underline. Exposed through the napi and python bindings. Verified on real docs: 4/4 underlined sentences flagged on a Japanese report, links/headings flagged on 8 of 10 underline-bearing eval docs, zero flags on docs without underlines. Known FP source (table cell borders) documented — downstream applies inline styling only to plain-text regions. napi 1.9.8 -> 1.9.9. Co-authored-by: Cursor <cursoragent@cursor.com> * fix(extractor): underline rules only from painted rects, normalized extents (review) Two review fixes: (1) normalize rect extents before the thickness/width checks — `re` operands pass through the CTM so width/height can be negative, which missed negative-width rules and let negative-height bands pass as thin; (2) only feed painted rects to underline detection — `re` rects now wait in a pending list until a paint operator (S/s, f/F/ f*, B/B*/b/b*) confirms them, and `re W n` clip-only paths are discarded at `n`, so invisible clip boundaries no longer underline nearby text. Marking moved into content_stream where paint state lives (pre-rotation, consistent device space). Co-authored-by: Cursor <cursoragent@cursor.com> * fix(extractor): harden underline detection * feat(cli): export positioned text item json --------- Co-authored-by: Cursor <cursoragent@cursor.com>
This commit is contained in:
co-authored by
Cursor
parent
30eddade77
commit
422a2ff118
+1
-1
@@ -1,6 +1,6 @@
|
||||
{
|
||||
"name": "@firecrawl/pdf-inspector",
|
||||
"version": "1.9.8",
|
||||
"version": "1.9.9",
|
||||
"description": "Fast PDF classification and text extraction. Detect text-based vs scanned PDFs, extract text by region with quality checks. Native Rust performance via napi-rs.",
|
||||
"main": "index.js",
|
||||
"types": "index.d.ts",
|
||||
|
||||
@@ -80,6 +80,9 @@ pub struct TextItem {
|
||||
pub page: u32,
|
||||
pub is_bold: bool,
|
||||
pub is_italic: bool,
|
||||
/// Underline detected geometrically (drawn rule/thin rect under the
|
||||
/// baseline) — PDFs carry no underline font flag.
|
||||
pub is_underline: bool,
|
||||
pub item_type: ItemType,
|
||||
/// URL for link items, `None` for other types.
|
||||
pub link_url: Option<String>,
|
||||
@@ -290,6 +293,7 @@ pub fn extract_text_with_positions(
|
||||
page: item.page,
|
||||
is_bold: item.is_bold,
|
||||
is_italic: item.is_italic,
|
||||
is_underline: item.is_underline,
|
||||
item_type,
|
||||
link_url,
|
||||
}
|
||||
|
||||
Reference in New Issue
Block a user