Files
pdf-inspector/pdf_inspector.pyi
T
Sunil a410d5aa08 fix: sync Python type stubs (.pyi) with actual bindings and fix doc bugs (#250)
## Bug 1: Python type stubs missing 5 fields/classes

The pdf_inspector.pyi file was out of sync with the actual Python
bindings exposed via #[pyo3(get)] in src/python.rs. This breaks
IDE autocompletion and type checking (PyCharm, VS Code/Pylance, mypy)
for all Python users.

Added:
- PdfResult.ocr_reasons_by_page (python.rs:35)
- PageOcrReasons class with page and 
easons fields (python.rs:67-86)
- RegionText.ocr_reason (python.rs:136)
- PageMarkdown.ocr_reason (python.rs:193)
- PagesExtractionResult.ocr_reasons_by_page (python.rs:226)

## Bug 2: PdfResult.pages_needing_ocr indexing undocumented

PdfResult.pages_needing_ocr is 1-indexed (per python.rs:30) but
neither the .pyi stubs nor docs/python.md annotated this, while the
same field on PdfClassification was annotated as 0-indexed. Users
mixing both APIs would get wrong page numbers.

## Bug 3: README.md duplicate bullet character

The Markdown features table listed * twice in bullet prefixes.
The first should be ullet (U+2022), matching the actual source code in
src/markdown/mod.rs:34 which uses: ullet, dash, sterisk, circle, illed-circle, open-circle.

## Bug 4: docs/python.md missing fields in type reference

The Types section was missing PageOcrReasons, RegionText class
definition, ocr_reason fields, and ocr_reasons_by_page fields.

## Evidence

Cross-referenced every #[pyo3(get)] attribute in src/python.rs
against the .pyi declarations and docs/python.md type reference.
2026-08-04 12:36:30 -07:00

186 lines
5.5 KiB
Python

"""Type stubs for pdf_inspector."""
from typing import Optional
class PdfResult:
"""Result of processing a PDF file."""
pdf_type: str
"""'text_based', 'scanned', 'image_based', or 'mixed'."""
markdown: Optional[str]
page_count: int
processing_time_ms: int
pages_needing_ocr: list[int]
"""1-indexed page numbers that need OCR."""
ocr_reasons_by_page: list["PageOcrReasons"]
"""Machine-readable OCR reasons by 1-indexed page."""
title: Optional[str]
confidence: float
is_complex_layout: bool
pages_with_tables: list[int]
pages_with_columns: list[int]
has_encoding_issues: bool
class PageOcrReasons:
"""OCR reasons for a single 1-indexed page."""
page: int
"""1-indexed page number."""
reasons: list[str]
"""Machine-readable OCR reason identifiers."""
class PdfClassification:
"""Lightweight PDF classification result."""
pdf_type: str
"""'text_based', 'scanned', 'image_based', or 'mixed'."""
page_count: int
pages_needing_ocr: list[int]
"""0-indexed page numbers that need OCR."""
confidence: float
class TextItem:
"""A positioned text item extracted from a PDF."""
text: str
x: float
y: float
width: float
height: float
font: str
font_size: float
page: int
is_bold: bool
is_italic: bool
is_underline: bool
is_strikeout: bool
item_type: str
class RegionText:
"""Extracted text for a single region."""
text: str
needs_ocr: bool
"""True when the text should not be trusted."""
ocr_reason: Optional[str]
"""Machine-readable OCR reason when the cause is known."""
class PageRegionTexts:
"""Extracted text for one page's regions."""
page: int
"""0-indexed page number."""
regions: list[RegionText]
class PageMarkdown:
"""Per-page markdown extraction result."""
page: int
"""0-indexed page number."""
markdown: str
"""Formatted markdown for this page (empty string when needs_ocr is True)."""
needs_ocr: bool
"""True when text on this page is unreliable and OCR should be used instead."""
ocr_reason: Optional[str]
"""Machine-readable OCR reason when the cause is known."""
class PagesExtractionResult:
"""Per-page markdown output with document-wide layout classification."""
pages: list[PageMarkdown]
"""Per-page markdown results, in the order requested."""
pages_with_tables: list[int]
"""1-indexed pages where tables were detected."""
pages_with_columns: list[int]
"""1-indexed pages where multi-column layout was detected."""
pages_needing_ocr: list[int]
"""1-indexed pages that need OCR."""
ocr_reasons_by_page: list[PageOcrReasons]
"""Machine-readable OCR reasons by 1-indexed page."""
is_complex: bool
"""True if any page has tables or multi-column layout."""
def process_pdf(path: str, pages: Optional[list[int]] = None) -> PdfResult:
"""Process a PDF: detect type, extract text, convert to Markdown."""
...
def process_pdf_bytes(data: bytes, pages: Optional[list[int]] = None) -> PdfResult:
"""Process a PDF from bytes in memory."""
...
def detect_pdf(path: str) -> PdfResult:
"""Fast detection only — no text extraction."""
...
def detect_pdf_bytes(data: bytes) -> PdfResult:
"""Fast detection from bytes."""
...
def classify_pdf(path: str) -> PdfClassification:
"""Lightweight classification — type, page count, and OCR pages (0-indexed)."""
...
def classify_pdf_bytes(data: bytes) -> PdfClassification:
"""Lightweight classification from bytes."""
...
def extract_text(path: str) -> str:
"""Extract plain text from a PDF."""
...
def extract_text_bytes(data: bytes) -> str:
"""Extract plain text from PDF bytes."""
...
def extract_text_with_positions(path: str, pages: Optional[list[int]] = None) -> list[TextItem]:
"""Extract text with position information."""
...
def extract_text_with_positions_bytes(data: bytes, pages: Optional[list[int]] = None) -> list[TextItem]:
"""Extract text with position information from bytes."""
...
def extract_text_in_regions(
path: str,
page_regions: list[tuple[int, list[list[float]]]],
) -> list[PageRegionTexts]:
"""Extract text within bounding-box regions from a PDF file.
Args:
path: Path to the PDF file.
page_regions: List of (page_0indexed, [[x1, y1, x2, y2], ...]) tuples.
"""
...
def extract_text_in_regions_bytes(
data: bytes,
page_regions: list[tuple[int, list[list[float]]]],
) -> list[PageRegionTexts]:
"""Extract text within bounding-box regions from PDF bytes.
Args:
data: PDF file contents as bytes.
page_regions: List of (page_0indexed, [[x1, y1, x2, y2], ...]) tuples.
"""
...
def extract_pages_markdown(
path: str,
pages: Optional[list[int]] = None,
) -> PagesExtractionResult:
"""Extract formatted markdown for pages of a PDF, with layout classification.
Args:
path: Path to the PDF file.
pages: Optional list of 0-indexed pages. When ``None`` (default), every
page is returned in document order. Otherwise, output matches the
caller-supplied order.
Returns:
PagesExtractionResult with per-page markdown and document-wide layout
classification (tables, columns, OCR needs).
"""
...
def extract_pages_markdown_bytes(
data: bytes,
pages: Optional[list[int]] = None,
) -> PagesExtractionResult:
"""Extract formatted markdown for pages of a PDF from bytes.
See :func:`extract_pages_markdown` for details.
"""
...