* feat(extractor): descriptor/embedded-font style flags + geometric strikeout detection
Two style-recall gaps, both invisible to the existing name-based
heuristics:
1. Subset fonts with opaque BaseFont names ("Tc1", "AAAAAB+Amplitude")
defeat is_italic_font/is_bold_font. New descriptor_style_flags reads
the FontDescriptor (ItalicAngle beyond 4 degrees, Flags bit 7 Italic,
bit 19 ForceBold) and, when the descriptor claims upright, falls back
to the embedded font file: ttf-parser's OS/2 fsSelection + post
italicAngle for sfnt fonts, and the CFF Name INDEX PostScript name
for bare-CFF FontFile3 (descriptor rewritten to ItalicAngle 0 while
embedding "Amplitude-LightItalic" was observed in the wild).
ORed into is_bold/is_italic at item creation (content streams and
form XObjects).
2. No strikeout signal existed. New is_strikeout on TextItem, detected
in the same pass as underline: same rules pipeline (stroked lines /
thin filled rects, table-ruling suppression), different vertical
window — a rule crossing the glyphs at 12-55% of the em above the
baseline instead of sitting at it. Exposed through napi and python
bindings and pdf2md --items-json.
Verified on public ParseBench corpus docs: previously-missed italic
council titles and bold CJK itinerary headings now flagged (render-
checked); 24/508 docs gain flags, none lose any; 35 strikeout items
detected corpus-wide, disjoint from underline.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB
* fix(review): quote-op advance width, Ts text rise, doc-level font style cache (PR #125 review)
Address three valid findings from review:
- The ' (move-to-next-line-and-show-text) operator emitted zero-width
items and never advanced the text matrix, so geometric underline/
strikeout detection (which requires width > 0) could never mark its
text, and following show ops overlapped it. Reuse Tj's advance-width
computation and matrix advance.
- Ts (text rise) was dropped entirely: raised/lowered runs kept the
unshifted baseline, so rules drawn at the risen glyph position missed
the strike/underline windows. Track rise in the text state (saved and
restored with q/Q) and shift the rendering position through the text
matrix's y column; advances stay on the unshifted matrix per spec.
- descriptor_style_flags re-decompressed and re-parsed the same embedded
font program on every page whenever the descriptor left a style flag
unset (the common case). Add a document-scoped FontStyleCache keyed by
the FontFile2/FontFile3 object id, threaded through page and form
extraction alongside the existing CMapDecisionCache.
The fourth finding (Form XObject rules never reach geometric detection)
is real but pre-existing for underline and needs the form walker to grow
path/paint tracking plus a new return type; deferred as a follow-up.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB
* fix(review): ActualText items render at their glyphs' text rise (PR #125 review)
The EMC-built ActualText item used the captured text matrix without the
rise adjustment the ordinary Tj/TJ/' emission sites apply, so a tagged
run shown with Ts landed on the unshifted baseline — off the strikeout/
underline windows and inconsistent with untagged runs. The rise is
captured together with the first-glyph matrix (and at BDC for the
entry-position fallback): the item must render at the rise of its
GLYPHS, not whatever rise is set by EMC time.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB
* fix(review): capture ActualText glyph position after the quote op's line move (PR #125 review)
The `'` handler skipped the entire suppressed-extraction block, so a
tagged span whose show op is `'` never captured its glyph matrix/rise —
the EMC item fell back to the BDC-entry matrix, which sits on the
PREVIOUS line (the `'` line move happens after BDC) with no rise. The
capture now happens right after the line move, matching the Tj/TJ
paths.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB
* fix(review): style-boundary gate on subscript merge + strikeout suppression coverage (PR #125 review)
merge_subscript_items absorbed a script digit into its parent
regardless of underline/strikeout flags — dropping the digit's own mark
or widening the parent's over it. The merged item carries one flag, so
differing marks now break the merge, mirroring merge_text_items'
style-boundary rule (pre-existing for underline as well).
Also extends the table-suppression test to assert is_strikeout is
cleared alongside is_underline.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01L3U6BKYCS73DVA83odAfYB
---------
Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
170 lines
4.9 KiB
Python
170 lines
4.9 KiB
Python
"""Type stubs for pdf_inspector."""
|
|
|
|
from typing import Optional
|
|
|
|
class PdfResult:
|
|
"""Result of processing a PDF file."""
|
|
pdf_type: str
|
|
"""'text_based', 'scanned', 'image_based', or 'mixed'."""
|
|
markdown: Optional[str]
|
|
page_count: int
|
|
processing_time_ms: int
|
|
pages_needing_ocr: list[int]
|
|
title: Optional[str]
|
|
confidence: float
|
|
is_complex_layout: bool
|
|
pages_with_tables: list[int]
|
|
pages_with_columns: list[int]
|
|
has_encoding_issues: bool
|
|
|
|
class PdfClassification:
|
|
"""Lightweight PDF classification result."""
|
|
pdf_type: str
|
|
"""'text_based', 'scanned', 'image_based', or 'mixed'."""
|
|
page_count: int
|
|
pages_needing_ocr: list[int]
|
|
"""0-indexed page numbers that need OCR."""
|
|
confidence: float
|
|
|
|
class TextItem:
|
|
"""A positioned text item extracted from a PDF."""
|
|
text: str
|
|
x: float
|
|
y: float
|
|
width: float
|
|
height: float
|
|
font: str
|
|
font_size: float
|
|
page: int
|
|
is_bold: bool
|
|
is_italic: bool
|
|
is_underline: bool
|
|
is_strikeout: bool
|
|
item_type: str
|
|
|
|
class RegionText:
|
|
"""Extracted text for a single region."""
|
|
text: str
|
|
needs_ocr: bool
|
|
"""True when the text should not be trusted."""
|
|
|
|
class PageRegionTexts:
|
|
"""Extracted text for one page's regions."""
|
|
page: int
|
|
"""0-indexed page number."""
|
|
regions: list[RegionText]
|
|
|
|
class PageMarkdown:
|
|
"""Per-page markdown extraction result."""
|
|
page: int
|
|
"""0-indexed page number."""
|
|
markdown: str
|
|
"""Formatted markdown for this page (empty string when needs_ocr is True)."""
|
|
needs_ocr: bool
|
|
"""True when text on this page is unreliable and OCR should be used instead."""
|
|
|
|
class PagesExtractionResult:
|
|
"""Per-page markdown output with document-wide layout classification."""
|
|
pages: list[PageMarkdown]
|
|
"""Per-page markdown results, in the order requested."""
|
|
pages_with_tables: list[int]
|
|
"""1-indexed pages where tables were detected."""
|
|
pages_with_columns: list[int]
|
|
"""1-indexed pages where multi-column layout was detected."""
|
|
pages_needing_ocr: list[int]
|
|
"""1-indexed pages that need OCR."""
|
|
is_complex: bool
|
|
"""True if any page has tables or multi-column layout."""
|
|
|
|
def process_pdf(path: str, pages: Optional[list[int]] = None) -> PdfResult:
|
|
"""Process a PDF: detect type, extract text, convert to Markdown."""
|
|
...
|
|
|
|
def process_pdf_bytes(data: bytes, pages: Optional[list[int]] = None) -> PdfResult:
|
|
"""Process a PDF from bytes in memory."""
|
|
...
|
|
|
|
def detect_pdf(path: str) -> PdfResult:
|
|
"""Fast detection only — no text extraction."""
|
|
...
|
|
|
|
def detect_pdf_bytes(data: bytes) -> PdfResult:
|
|
"""Fast detection from bytes."""
|
|
...
|
|
|
|
def classify_pdf(path: str) -> PdfClassification:
|
|
"""Lightweight classification — type, page count, and OCR pages (0-indexed)."""
|
|
...
|
|
|
|
def classify_pdf_bytes(data: bytes) -> PdfClassification:
|
|
"""Lightweight classification from bytes."""
|
|
...
|
|
|
|
def extract_text(path: str) -> str:
|
|
"""Extract plain text from a PDF."""
|
|
...
|
|
|
|
def extract_text_bytes(data: bytes) -> str:
|
|
"""Extract plain text from PDF bytes."""
|
|
...
|
|
|
|
def extract_text_with_positions(path: str, pages: Optional[list[int]] = None) -> list[TextItem]:
|
|
"""Extract text with position information."""
|
|
...
|
|
|
|
def extract_text_with_positions_bytes(data: bytes, pages: Optional[list[int]] = None) -> list[TextItem]:
|
|
"""Extract text with position information from bytes."""
|
|
...
|
|
|
|
def extract_text_in_regions(
|
|
path: str,
|
|
page_regions: list[tuple[int, list[list[float]]]],
|
|
) -> list[PageRegionTexts]:
|
|
"""Extract text within bounding-box regions from a PDF file.
|
|
|
|
Args:
|
|
path: Path to the PDF file.
|
|
page_regions: List of (page_0indexed, [[x1, y1, x2, y2], ...]) tuples.
|
|
"""
|
|
...
|
|
|
|
def extract_text_in_regions_bytes(
|
|
data: bytes,
|
|
page_regions: list[tuple[int, list[list[float]]]],
|
|
) -> list[PageRegionTexts]:
|
|
"""Extract text within bounding-box regions from PDF bytes.
|
|
|
|
Args:
|
|
data: PDF file contents as bytes.
|
|
page_regions: List of (page_0indexed, [[x1, y1, x2, y2], ...]) tuples.
|
|
"""
|
|
...
|
|
|
|
def extract_pages_markdown(
|
|
path: str,
|
|
pages: Optional[list[int]] = None,
|
|
) -> PagesExtractionResult:
|
|
"""Extract formatted markdown for pages of a PDF, with layout classification.
|
|
|
|
Args:
|
|
path: Path to the PDF file.
|
|
pages: Optional list of 0-indexed pages. When ``None`` (default), every
|
|
page is returned in document order. Otherwise, output matches the
|
|
caller-supplied order.
|
|
|
|
Returns:
|
|
PagesExtractionResult with per-page markdown and document-wide layout
|
|
classification (tables, columns, OCR needs).
|
|
"""
|
|
...
|
|
|
|
def extract_pages_markdown_bytes(
|
|
data: bytes,
|
|
pages: Optional[list[int]] = None,
|
|
) -> PagesExtractionResult:
|
|
"""Extract formatted markdown for pages of a PDF from bytes.
|
|
|
|
See :func:`extract_pages_markdown` for details.
|
|
"""
|
|
...
|