Compare commits
7
Commits
| Author | SHA1 | Date | |
|---|---|---|---|
|
|
0d749bb666 | ||
|
|
bbdf964d70 | ||
|
|
fe28c3fcd2 | ||
|
|
f6abe6c2de | ||
|
|
15c0b22093 | ||
|
|
0c06dac976 | ||
|
|
64a0930f9f |
@@ -25,6 +25,9 @@ jobs:
|
||||
- name: Run tests
|
||||
run: cargo test --verbose
|
||||
|
||||
- name: Test developer scripts
|
||||
run: python3 -m unittest discover -s scripts/tests
|
||||
|
||||
fmt:
|
||||
name: Format
|
||||
runs-on: ubuntu-latest
|
||||
|
||||
@@ -23,20 +23,23 @@ Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in
|
||||
|
||||
## Benchmark
|
||||
|
||||
Evaluated on the [opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs). Only direct text extraction engines are shown — no OCR, no ML models. Scores are 0-1, higher is better.
|
||||
Evaluated on the [opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs). Only local engines without model-based PDF parsing are shown; OCR was disabled. Scores are 0-1, higher is better.
|
||||
|
||||
| Engine | Overall | Reading Order (NID) | Tables (TEDS) | Headings (MHS) | Speed (200 docs) |
|
||||
|---|---|---|---|---|---|
|
||||
| pdf-inspector | 0.83 | 0.89 | 0.66 | 0.74 | 4s |
|
||||
| opendataloader | 0.84 | 0.91 | 0.49 | 0.74 | 11s |
|
||||
| pymupdf4llm | 0.73 | 0.89 | 0.40 | 0.41 | 18s |
|
||||
| markitdown | 0.58 | 0.88 | 0.00 | 0.00 | 8s |
|
||||
| pdf-inspector | **0.875** | **0.915** | **0.814** | 0.788 | **2.8s** |
|
||||
| liteparse | 0.870 | 0.908 | 0.693 | **0.811** | 13.9s |
|
||||
| opendataloader | 0.843 | 0.912 | 0.489 | 0.760 | 9.8s |
|
||||
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 15.5s |
|
||||
| markitdown | 0.583 | 0.879 | 0.000 | 0.000 | 6.7s |
|
||||
|
||||
For context, engines that use OCR/ML (docling, marker, mineru) score 0.83-0.88 overall but take 2-180 minutes on the same corpus — pdf-inspector reaches the low end of that range without any OCR, in 4 seconds.
|
||||
Results were refreshed on July 16, 2026, on an Apple M4 Pro. Engine versions were pdf-inspector 0.1.6, LiteParse 2.6.0, OpenDataLoader 2.1.1, PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.4. Speed is the median of three complete corpus runs.
|
||||
|
||||
**Where we do well:** Speed (fastest of all engines), the best table detection of any engine shown, and heading detection now on par with opendataloader. Overall lands within 0.01 of opendataloader at roughly 2.5× the speed.
|
||||
For context, engines that use OCR or model-based document parsing (docling, marker, mineru) score 0.83-0.88 overall but take 2-180 minutes on the same corpus — pdf-inspector reaches the top of that range without either, in 2.8 seconds.
|
||||
|
||||
**Where we lag:** Reading order still trails opendataloader slightly, and table structure trails OCR-based engines that can see visual layout.
|
||||
**Best fit:** Native-text PDFs where speed, reading order, and table structure matter. pdf-inspector delivered the highest overall, reading-order, and table scores, along with the fastest complete run in this benchmark. That makes it a strong local default for reports, research papers, financial documents, invoices, and legal PDFs that need clean, structured Markdown without adding OCR latency or infrastructure.
|
||||
|
||||
Use the [paired benchmark harness](docs/benchmarking.md) to compare two local builds against the exact same corpus and evaluator revision.
|
||||
|
||||
## Quick start
|
||||
|
||||
|
||||
@@ -0,0 +1,59 @@
|
||||
# Benchmarking against OpenDataLoader
|
||||
|
||||
The paired harness runs two `pdf2md` binaries through the same local
|
||||
OpenDataLoader corpus, evaluates both outputs, and reports aggregate and
|
||||
per-document deltas. This avoids comparing results produced from different
|
||||
corpus revisions or evaluator versions.
|
||||
|
||||
Build a candidate and provide a released or worktree build as the baseline:
|
||||
|
||||
```bash
|
||||
cargo build --release
|
||||
python3 scripts/bench_opendataloader.py \
|
||||
--bench-dir ../opendataloader-bench \
|
||||
--baseline ../pdf-inspector-main/target/release/pdf2md \
|
||||
--candidate target/release/pdf2md \
|
||||
--max-document-regression 0.02 \
|
||||
--json-output /tmp/pdf-inspector-benchmark.json
|
||||
```
|
||||
|
||||
Pass `--reference-evaluation path/to/evaluation.json` to report the candidate
|
||||
delta against another evaluation, and add `--require-reference-lead` to make a
|
||||
negative reference delta fail the run. By default, the candidate must not
|
||||
regress the baseline overall score or introduce missing predictions. Use
|
||||
`--min-overall-delta` to require a specific aggregate gain.
|
||||
|
||||
The OpenDataLoader repository is external and keeps its normal
|
||||
`prediction/pdf-inspector` output. Paired evaluation copies each run into a
|
||||
temporary directory before evaluating it, so the baseline and candidate cannot
|
||||
overwrite one another.
|
||||
|
||||
## Published comparison protocol
|
||||
|
||||
The public benchmark table was refreshed on July 16, 2026, on an Apple M4 Pro
|
||||
using pdf-inspector 0.1.6, LiteParse 2.6.0, OpenDataLoader 2.1.1,
|
||||
PyMuPDF4LLM 0.2.0, and MarkItDown 0.1.4. Every engine processed the same 200
|
||||
PDFs with OCR disabled. Reported speed is the median of three complete corpus
|
||||
runs; quality scores come from the benchmark evaluator over all 200 outputs.
|
||||
|
||||
## Optional backend evidence probe
|
||||
|
||||
The evidence probe compares positioned `pdf2md` items with MuPDF structured
|
||||
text on the same pages. It is intended to find deterministic extraction or
|
||||
layout evidence that could justify a future native implementation; it does not
|
||||
merge MuPDF output into Markdown, invoke OCR, or add a runtime dependency.
|
||||
|
||||
Install MuPDF's `mutool`, build `pdf2md`, then run:
|
||||
|
||||
```bash
|
||||
python3 scripts/probe_backend_evidence.py document.pdf \
|
||||
--pdf2md target/release/pdf2md \
|
||||
--json-output /tmp/backend-evidence.json
|
||||
```
|
||||
|
||||
The report flags pages when MuPDF exposes a material net token gain, repeated
|
||||
alignment anchors absent from local evidence, or additional image blocks. The
|
||||
JSON includes bounded token samples and page-level counts so promising cases
|
||||
can be inspected without treating backend disagreement as automatically
|
||||
correct. Thresholds are configurable with `--min-token-gain`,
|
||||
`--min-alternate-only-ratio`, and `--min-anchor-gain`.
|
||||
+7
-5
@@ -14,15 +14,17 @@ Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in
|
||||
|
||||
## Benchmark
|
||||
|
||||
[opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs), direct-extraction engines only — no OCR, no ML. Scores 0–1, higher is better:
|
||||
[opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs), local engines without model-based PDF parsing; OCR disabled. Scores 0–1, higher is better:
|
||||
|
||||
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|
||||
|---|---|---|---|---|---|
|
||||
| **pdf-inspector** | 0.83 | 0.88 | **0.66** | 0.74 | **4s** |
|
||||
| opendataloader | 0.84 | 0.91 | 0.49 | 0.74 | 11s |
|
||||
| pymupdf4llm | 0.73 | 0.89 | 0.40 | 0.41 | 18s |
|
||||
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **2.8s** |
|
||||
| liteparse | 0.870 | 0.908 | 0.693 | **0.811** | 13.9s |
|
||||
| opendataloader | 0.843 | 0.912 | 0.489 | 0.760 | 9.8s |
|
||||
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 15.5s |
|
||||
| markitdown | 0.583 | 0.879 | 0.000 | 0.000 | 6.7s |
|
||||
|
||||
OCR/ML engines (docling, marker, mineru) score 0.83–0.88 overall but take 2–180 minutes on the same corpus. Full numbers in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark).
|
||||
Refreshed July 16, 2026, on Apple M4 Pro; speed is the median of three complete corpus runs. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark).
|
||||
|
||||
## Install
|
||||
|
||||
|
||||
+7
-5
@@ -14,15 +14,17 @@ Built by [Firecrawl](https://firecrawl.dev) to handle text-based PDFs locally in
|
||||
|
||||
## Benchmark
|
||||
|
||||
[opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs), direct-extraction engines only — no OCR, no ML. Scores 0–1, higher is better:
|
||||
[opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs), local engines without model-based PDF parsing; OCR disabled. Scores 0–1, higher is better:
|
||||
|
||||
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|
||||
|---|---|---|---|---|---|
|
||||
| **pdf-inspector** | 0.83 | 0.88 | **0.66** | 0.74 | **4s** |
|
||||
| opendataloader | 0.84 | 0.91 | 0.49 | 0.74 | 11s |
|
||||
| pymupdf4llm | 0.73 | 0.89 | 0.40 | 0.41 | 18s |
|
||||
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **2.8s** |
|
||||
| liteparse | 0.870 | 0.908 | 0.693 | **0.811** | 13.9s |
|
||||
| opendataloader | 0.843 | 0.912 | 0.489 | 0.760 | 9.8s |
|
||||
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 15.5s |
|
||||
| markitdown | 0.583 | 0.879 | 0.000 | 0.000 | 6.7s |
|
||||
|
||||
OCR/ML engines (docling, marker, mineru) score 0.83–0.88 overall but take 2–180 minutes on the same corpus. Full numbers in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark).
|
||||
Refreshed July 16, 2026, on Apple M4 Pro; speed is the median of three complete corpus runs. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark).
|
||||
|
||||
## Install
|
||||
|
||||
|
||||
+7
-5
@@ -14,15 +14,17 @@ Built by [Firecrawl](https://firecrawl.dev) for hybrid OCR pipelines — extract
|
||||
|
||||
## Benchmark
|
||||
|
||||
[opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs), direct-extraction engines only — no OCR, no ML. Scores 0–1, higher is better:
|
||||
[opendataloader-bench](https://github.com/opendataloader-project/opendataloader-bench) corpus (200 PDFs), local engines without model-based PDF parsing; OCR disabled. Scores 0–1, higher is better:
|
||||
|
||||
| Engine | Overall | Reading order | Tables (TEDS) | Headings | Speed |
|
||||
|---|---|---|---|---|---|
|
||||
| **pdf-inspector** | 0.83 | 0.88 | **0.66** | 0.74 | **4s** |
|
||||
| opendataloader | 0.84 | 0.91 | 0.49 | 0.74 | 11s |
|
||||
| pymupdf4llm | 0.73 | 0.89 | 0.40 | 0.41 | 18s |
|
||||
| **pdf-inspector** | **0.875** | **0.915** | **0.814** | 0.788 | **2.8s** |
|
||||
| liteparse | 0.870 | 0.908 | 0.693 | **0.811** | 13.9s |
|
||||
| opendataloader | 0.843 | 0.912 | 0.489 | 0.760 | 9.8s |
|
||||
| pymupdf4llm | 0.735 | 0.886 | 0.401 | 0.424 | 15.5s |
|
||||
| markitdown | 0.583 | 0.879 | 0.000 | 0.000 | 6.7s |
|
||||
|
||||
OCR/ML engines (docling, marker, mineru) score 0.83–0.88 overall but take 2–180 minutes on the same corpus. Full numbers in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark).
|
||||
Refreshed July 16, 2026, on Apple M4 Pro; speed is the median of three complete corpus runs. Full methodology and versions are in the [repo README](https://github.com/firecrawl/pdf-inspector#benchmark).
|
||||
|
||||
## Install
|
||||
|
||||
|
||||
@@ -0,0 +1,351 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Run a paired pdf-inspector OpenDataLoader benchmark and report deltas."""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
import json
|
||||
import math
|
||||
import os
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
import tempfile
|
||||
from pathlib import Path
|
||||
from typing import Any
|
||||
|
||||
|
||||
SCORE_KEYS = (
|
||||
"overall_mean",
|
||||
"nid_mean",
|
||||
"nid_s_mean",
|
||||
"teds_mean",
|
||||
"teds_s_mean",
|
||||
"mhs_mean",
|
||||
"mhs_s_mean",
|
||||
)
|
||||
|
||||
|
||||
def _non_negative_int(value: str) -> int:
|
||||
parsed = int(value)
|
||||
if parsed < 0:
|
||||
raise argparse.ArgumentTypeError("must be non-negative")
|
||||
return parsed
|
||||
|
||||
|
||||
def _non_negative_float(value: str) -> float:
|
||||
parsed = float(value)
|
||||
if not math.isfinite(parsed) or parsed < 0.0:
|
||||
raise argparse.ArgumentTypeError("must be finite and non-negative")
|
||||
return parsed
|
||||
|
||||
|
||||
def _finite_float(value: str) -> float:
|
||||
parsed = float(value)
|
||||
if not math.isfinite(parsed):
|
||||
raise argparse.ArgumentTypeError("must be finite")
|
||||
return parsed
|
||||
|
||||
|
||||
def _scores(evaluation: dict[str, Any]) -> dict[str, float]:
|
||||
score = evaluation.get("metrics", {}).get("score", {})
|
||||
return {key: float(score[key]) for key in SCORE_KEYS if score.get(key) is not None}
|
||||
|
||||
|
||||
def _documents(evaluation: dict[str, Any]) -> dict[str, float]:
|
||||
documents: dict[str, float] = {}
|
||||
for document in evaluation.get("documents", []):
|
||||
overall = document.get("scores", {}).get("overall")
|
||||
if overall is not None:
|
||||
documents[str(document["document_id"])] = float(overall)
|
||||
return documents
|
||||
|
||||
|
||||
def compare_evaluations(
|
||||
baseline: dict[str, Any],
|
||||
candidate: dict[str, Any],
|
||||
reference: dict[str, Any] | None = None,
|
||||
*,
|
||||
top: int = 10,
|
||||
) -> dict[str, Any]:
|
||||
"""Build aggregate and per-document deltas from evaluator JSON payloads."""
|
||||
baseline_scores = _scores(baseline)
|
||||
candidate_scores = _scores(candidate)
|
||||
metric_deltas = {
|
||||
key: candidate_scores[key] - baseline_scores[key]
|
||||
for key in SCORE_KEYS
|
||||
if key in baseline_scores and key in candidate_scores
|
||||
}
|
||||
|
||||
baseline_documents = _documents(baseline)
|
||||
candidate_documents = _documents(candidate)
|
||||
shared = sorted(baseline_documents.keys() & candidate_documents.keys())
|
||||
document_deltas = [
|
||||
{
|
||||
"document_id": document_id,
|
||||
"baseline": baseline_documents[document_id],
|
||||
"candidate": candidate_documents[document_id],
|
||||
"delta": candidate_documents[document_id] - baseline_documents[document_id],
|
||||
}
|
||||
for document_id in shared
|
||||
]
|
||||
epsilon = 1e-12
|
||||
improvements = sorted(document_deltas, key=lambda item: item["delta"], reverse=True)
|
||||
regressions = sorted(document_deltas, key=lambda item: item["delta"])
|
||||
|
||||
result: dict[str, Any] = {
|
||||
"baseline": baseline_scores,
|
||||
"candidate": candidate_scores,
|
||||
"deltas": metric_deltas,
|
||||
"missing_predictions": {
|
||||
"baseline": int(baseline.get("metrics", {}).get("missing_predictions", 0)),
|
||||
"candidate": int(candidate.get("metrics", {}).get("missing_predictions", 0)),
|
||||
},
|
||||
"documents": {
|
||||
"shared": len(shared),
|
||||
"improved": sum(item["delta"] > epsilon for item in document_deltas),
|
||||
"regressed": sum(item["delta"] < -epsilon for item in document_deltas),
|
||||
"unchanged": sum(abs(item["delta"]) <= epsilon for item in document_deltas),
|
||||
"largest_improvements": [
|
||||
item for item in improvements if item["delta"] > epsilon
|
||||
][:top],
|
||||
"largest_regressions": [
|
||||
item for item in regressions if item["delta"] < -epsilon
|
||||
][:top],
|
||||
"worst_regression": next(
|
||||
(item for item in regressions if item["delta"] < -epsilon), None
|
||||
),
|
||||
},
|
||||
}
|
||||
if reference is not None:
|
||||
reference_scores = _scores(reference)
|
||||
result["reference"] = reference_scores
|
||||
result["candidate_vs_reference"] = {
|
||||
key: candidate_scores[key] - reference_scores[key]
|
||||
for key in SCORE_KEYS
|
||||
if key in candidate_scores and key in reference_scores
|
||||
}
|
||||
return result
|
||||
|
||||
|
||||
def evaluate_gates(
|
||||
comparison: dict[str, Any],
|
||||
*,
|
||||
min_overall_delta: float,
|
||||
max_document_regression: float | None,
|
||||
max_missing: int,
|
||||
require_reference_lead: bool,
|
||||
) -> list[str]:
|
||||
"""Return human-readable gate failures; an empty list means pass."""
|
||||
failures: list[str] = []
|
||||
overall_delta = comparison["deltas"].get("overall_mean")
|
||||
if overall_delta is None or overall_delta < min_overall_delta:
|
||||
failures.append(
|
||||
f"overall delta {overall_delta!r} is below {min_overall_delta:+.6f}"
|
||||
)
|
||||
candidate_missing = comparison["missing_predictions"]["candidate"]
|
||||
if candidate_missing > max_missing:
|
||||
failures.append(
|
||||
f"candidate has {candidate_missing} missing predictions (maximum {max_missing})"
|
||||
)
|
||||
if max_document_regression is not None:
|
||||
regression = comparison["documents"].get("worst_regression")
|
||||
if regression is not None and regression["delta"] < -max_document_regression:
|
||||
failures.append(
|
||||
"largest document regression "
|
||||
f"{regression['document_id']}={regression['delta']:+.6f} "
|
||||
f"exceeds {-max_document_regression:+.6f}"
|
||||
)
|
||||
if require_reference_lead:
|
||||
reference_delta = comparison.get("candidate_vs_reference", {}).get("overall_mean")
|
||||
if reference_delta is None:
|
||||
failures.append("reference overall score is unavailable")
|
||||
elif reference_delta < 0.0:
|
||||
failures.append(
|
||||
f"candidate trails reference overall by {reference_delta!r}"
|
||||
)
|
||||
return failures
|
||||
|
||||
|
||||
def _run(command: list[str], *, cwd: Path, env: dict[str, str] | None = None) -> None:
|
||||
print("+", " ".join(command), flush=True)
|
||||
subprocess.run(command, cwd=cwd, env=env, check=True)
|
||||
|
||||
|
||||
def _run_engine(
|
||||
*,
|
||||
bench_dir: Path,
|
||||
python: Path,
|
||||
binary: Path,
|
||||
label: str,
|
||||
scratch_root: Path,
|
||||
) -> dict[str, Any]:
|
||||
env = os.environ.copy()
|
||||
env["PDF_INSPECTOR_BINARY"] = str(binary)
|
||||
source = bench_dir / "prediction" / "pdf-inspector"
|
||||
if source.exists():
|
||||
if source.is_dir():
|
||||
shutil.rmtree(source)
|
||||
else:
|
||||
source.unlink()
|
||||
_run(
|
||||
[
|
||||
str(python),
|
||||
"src/pdf_parser.py",
|
||||
"--engine",
|
||||
"pdf-inspector",
|
||||
"--log-level",
|
||||
"WARNING",
|
||||
],
|
||||
cwd=bench_dir,
|
||||
env=env,
|
||||
)
|
||||
|
||||
if not source.is_dir():
|
||||
raise RuntimeError(f"parser did not produce predictions: {source}")
|
||||
destination = scratch_root / label
|
||||
shutil.copytree(source, destination)
|
||||
_run(
|
||||
[
|
||||
str(python),
|
||||
"src/evaluator.py",
|
||||
"--prediction-root",
|
||||
str(scratch_root),
|
||||
"--engine",
|
||||
label,
|
||||
"--log-level",
|
||||
"WARNING",
|
||||
],
|
||||
cwd=bench_dir,
|
||||
)
|
||||
with (destination / "evaluation.json").open(encoding="utf-8") as handle:
|
||||
return json.load(handle)
|
||||
|
||||
|
||||
def _print_report(comparison: dict[str, Any]) -> None:
|
||||
print("\nMetric baseline candidate delta")
|
||||
print("-------------------- ---------- ---------- ----------")
|
||||
for key in SCORE_KEYS:
|
||||
if key not in comparison["deltas"]:
|
||||
continue
|
||||
print(
|
||||
f"{key:<20} {comparison['baseline'][key]:>10.6f} "
|
||||
f"{comparison['candidate'][key]:>10.6f} "
|
||||
f"{comparison['deltas'][key]:>+10.6f}"
|
||||
)
|
||||
if "reference" in comparison:
|
||||
delta = comparison["candidate_vs_reference"].get("overall_mean")
|
||||
reference = comparison["reference"].get("overall_mean")
|
||||
reference_display = f"{reference:.6f}" if reference is not None else "n/a"
|
||||
delta_display = f"{delta:+.6f}" if delta is not None else "n/a"
|
||||
print(f"\nReference overall: {reference_display}; candidate delta: {delta_display}")
|
||||
|
||||
documents = comparison["documents"]
|
||||
print(
|
||||
"\nDocuments: "
|
||||
f"{documents['improved']} improved, {documents['regressed']} regressed, "
|
||||
f"{documents['unchanged']} unchanged ({documents['shared']} shared)"
|
||||
)
|
||||
for heading, key in (
|
||||
("Largest improvements", "largest_improvements"),
|
||||
("Largest regressions", "largest_regressions"),
|
||||
):
|
||||
print(f"\n{heading}:")
|
||||
rows = documents[key]
|
||||
if not rows:
|
||||
print(" none")
|
||||
for row in rows:
|
||||
print(
|
||||
f" {row['document_id']}: {row['delta']:+.6f} "
|
||||
f"({row['baseline']:.6f} -> {row['candidate']:.6f})"
|
||||
)
|
||||
|
||||
|
||||
def _arguments(argv: list[str] | None = None) -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("--bench-dir", type=Path, required=True)
|
||||
parser.add_argument("--baseline", type=Path, required=True)
|
||||
parser.add_argument("--candidate", type=Path, required=True)
|
||||
parser.add_argument("--python", type=Path)
|
||||
parser.add_argument("--reference-evaluation", type=Path)
|
||||
parser.add_argument("--json-output", type=Path)
|
||||
parser.add_argument("--top", type=_non_negative_int, default=10)
|
||||
parser.add_argument("--min-overall-delta", type=_finite_float, default=0.0)
|
||||
parser.add_argument("--max-document-regression", type=_non_negative_float)
|
||||
parser.add_argument("--max-missing", type=_non_negative_int, default=0)
|
||||
parser.add_argument("--require-reference-lead", action="store_true")
|
||||
return parser.parse_args(argv)
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
args = _arguments(argv)
|
||||
bench_dir = args.bench_dir.resolve()
|
||||
baseline = args.baseline.resolve()
|
||||
candidate = args.candidate.resolve()
|
||||
# Keep the virtualenv launcher path intact. Resolving its symlink would
|
||||
# invoke the underlying system interpreter without the benchmark's site
|
||||
# packages.
|
||||
python = (args.python or bench_dir / ".venv" / "bin" / "python").absolute()
|
||||
for path, description in (
|
||||
(bench_dir / "src" / "pdf_parser.py", "OpenDataLoader parser"),
|
||||
(bench_dir / "src" / "evaluator.py", "OpenDataLoader evaluator"),
|
||||
(baseline, "baseline binary"),
|
||||
(candidate, "candidate binary"),
|
||||
(python, "Python interpreter"),
|
||||
):
|
||||
if not path.exists():
|
||||
raise SystemExit(f"{description} not found: {path}")
|
||||
|
||||
with tempfile.TemporaryDirectory(prefix="pdf-inspector-opendataloader-") as temporary:
|
||||
scratch_root = Path(temporary)
|
||||
baseline_evaluation = _run_engine(
|
||||
bench_dir=bench_dir,
|
||||
python=python,
|
||||
binary=baseline,
|
||||
label="baseline",
|
||||
scratch_root=scratch_root,
|
||||
)
|
||||
candidate_evaluation = _run_engine(
|
||||
bench_dir=bench_dir,
|
||||
python=python,
|
||||
binary=candidate,
|
||||
label="candidate",
|
||||
scratch_root=scratch_root,
|
||||
)
|
||||
|
||||
reference = None
|
||||
if args.reference_evaluation is not None:
|
||||
with args.reference_evaluation.resolve().open(encoding="utf-8") as handle:
|
||||
reference = json.load(handle)
|
||||
|
||||
comparison = compare_evaluations(
|
||||
baseline_evaluation,
|
||||
candidate_evaluation,
|
||||
reference,
|
||||
top=args.top,
|
||||
)
|
||||
|
||||
_print_report(comparison)
|
||||
if args.json_output is not None:
|
||||
args.json_output.resolve().write_text(
|
||||
json.dumps(comparison, indent=2) + "\n", encoding="utf-8"
|
||||
)
|
||||
|
||||
failures = evaluate_gates(
|
||||
comparison,
|
||||
min_overall_delta=args.min_overall_delta,
|
||||
max_document_regression=args.max_document_regression,
|
||||
max_missing=args.max_missing,
|
||||
require_reference_lead=args.require_reference_lead,
|
||||
)
|
||||
if failures:
|
||||
print("\nBenchmark gate failed:", file=sys.stderr)
|
||||
for failure in failures:
|
||||
print(f" - {failure}", file=sys.stderr)
|
||||
return 1
|
||||
print("\nBenchmark gate passed.")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,351 @@
|
||||
#!/usr/bin/env python3
|
||||
"""Compare pdf-inspector evidence with optional MuPDF structured text.
|
||||
|
||||
This is an experiment and diagnostic tool, not an extraction fallback. It runs
|
||||
MuPDF's deterministic ``stext.json`` backend without OCR and highlights pages
|
||||
where that backend exposes materially different text or layout evidence.
|
||||
"""
|
||||
|
||||
from __future__ import annotations
|
||||
|
||||
import argparse
|
||||
from collections import Counter
|
||||
import json
|
||||
from pathlib import Path
|
||||
import re
|
||||
import shutil
|
||||
import subprocess
|
||||
import sys
|
||||
from typing import Any, Iterable
|
||||
|
||||
|
||||
TOKEN_PATTERN = re.compile(r"[^\W_]+(?:[\u2019'][^\W_]+)*", re.UNICODE)
|
||||
|
||||
|
||||
def _tokens(texts: Iterable[str]) -> Counter[str]:
|
||||
tokens: Counter[str] = Counter()
|
||||
for text in texts:
|
||||
for token in TOKEN_PATTERN.findall(text.casefold()):
|
||||
# Lone letters are frequently bullets, chart labels, or fragmented
|
||||
# glyphs. Digits remain useful even when they are one character.
|
||||
if len(token) > 1 or token.isdigit():
|
||||
tokens[token] += 1
|
||||
return tokens
|
||||
|
||||
|
||||
def _repeated_x_anchors(xs: Iterable[float], *, tolerance: float = 4.0) -> int:
|
||||
buckets = Counter(round(float(x) / tolerance) for x in xs)
|
||||
return sum(count >= 3 for count in buckets.values())
|
||||
|
||||
|
||||
def local_pages(payload: dict[str, Any]) -> dict[int, dict[str, Any]]:
|
||||
"""Summarize positioned ``pdf2md --items-json`` evidence by page."""
|
||||
pages: dict[int, dict[str, Any]] = {}
|
||||
for item in payload.get("items", []):
|
||||
page_number = int(item["page"])
|
||||
page = pages.setdefault(
|
||||
page_number,
|
||||
{"texts": [], "xs": [], "text_items": 0, "image_items": 0},
|
||||
)
|
||||
if item.get("item_type") == "image":
|
||||
page["image_items"] += 1
|
||||
continue
|
||||
text = str(item.get("text", ""))
|
||||
if text.strip():
|
||||
page["texts"].append(text)
|
||||
page["xs"].append(float(item.get("x", 0.0)))
|
||||
page["text_items"] += 1
|
||||
return pages
|
||||
|
||||
|
||||
def alternate_pages(
|
||||
payload: dict[str, Any] | list[dict[str, Any]],
|
||||
) -> dict[int, dict[str, Any]]:
|
||||
"""Summarize MuPDF ``stext.json`` evidence by page."""
|
||||
pages: dict[int, dict[str, Any]] = {}
|
||||
raw_pages = payload if isinstance(payload, list) else payload.get("pages", [])
|
||||
for index, raw_page in enumerate(raw_pages, start=1):
|
||||
page_number = int(raw_page.get("number", index))
|
||||
page = {
|
||||
"texts": [],
|
||||
"xs": [],
|
||||
"text_blocks": 0,
|
||||
"text_lines": 0,
|
||||
"image_blocks": 0,
|
||||
}
|
||||
for block in raw_page.get("blocks", []):
|
||||
if block.get("type") == "image":
|
||||
page["image_blocks"] += 1
|
||||
continue
|
||||
if block.get("type") != "text":
|
||||
continue
|
||||
page["text_blocks"] += 1
|
||||
for line in block.get("lines", []):
|
||||
text = str(line.get("text", ""))
|
||||
if text.strip():
|
||||
page["texts"].append(text)
|
||||
bbox = line.get("bbox", {})
|
||||
page["xs"].append(float(bbox.get("x", line.get("x", 0.0))))
|
||||
page["text_lines"] += 1
|
||||
pages[page_number] = page
|
||||
return pages
|
||||
|
||||
|
||||
def compare_page(
|
||||
local: dict[str, Any],
|
||||
alternate: dict[str, Any],
|
||||
*,
|
||||
min_token_gain: int,
|
||||
min_alternate_only_ratio: float,
|
||||
min_anchor_gain: int,
|
||||
) -> dict[str, Any]:
|
||||
"""Compare semantic and coarse layout evidence for one page."""
|
||||
local_tokens = _tokens(local.get("texts", []))
|
||||
alternate_tokens = _tokens(alternate.get("texts", []))
|
||||
shared = local_tokens & alternate_tokens
|
||||
alternate_only = alternate_tokens - local_tokens
|
||||
local_only = local_tokens - alternate_tokens
|
||||
local_total = sum(local_tokens.values())
|
||||
alternate_total = sum(alternate_tokens.values())
|
||||
shared_total = sum(shared.values())
|
||||
alternate_only_total = sum(alternate_only.values())
|
||||
local_only_total = sum(local_only.values())
|
||||
net_token_gain = alternate_total - local_total
|
||||
alternate_only_ratio = alternate_only_total / max(alternate_total, 1)
|
||||
|
||||
local_anchors = _repeated_x_anchors(local.get("xs", []))
|
||||
alternate_anchors = _repeated_x_anchors(alternate.get("xs", []))
|
||||
anchor_gain = alternate_anchors - local_anchors
|
||||
image_gain = int(alternate.get("image_blocks", 0)) - int(
|
||||
local.get("image_items", 0)
|
||||
)
|
||||
|
||||
reasons: list[str] = []
|
||||
if local_total == 0 and alternate_total >= max(5, min_token_gain // 2):
|
||||
reasons.append("local_text_empty")
|
||||
elif (
|
||||
net_token_gain >= min_token_gain
|
||||
and alternate_only_ratio >= min_alternate_only_ratio
|
||||
):
|
||||
reasons.append("alternate_has_more_text")
|
||||
if anchor_gain >= min_anchor_gain:
|
||||
reasons.append("alternate_has_more_alignment_anchors")
|
||||
if image_gain > 0:
|
||||
reasons.append("alternate_has_more_image_blocks")
|
||||
|
||||
if reasons:
|
||||
classification = "investigate_alternate_evidence"
|
||||
elif local_total - alternate_total >= min_token_gain:
|
||||
classification = "local_has_more_text"
|
||||
elif alternate_only_total + local_only_total:
|
||||
classification = "different_segmentation_or_decoding"
|
||||
else:
|
||||
classification = "equivalent_text_evidence"
|
||||
|
||||
return {
|
||||
"classification": classification,
|
||||
"reasons": reasons,
|
||||
"tokens": {
|
||||
"local": local_total,
|
||||
"alternate": alternate_total,
|
||||
"shared": shared_total,
|
||||
"net_alternate_gain": net_token_gain,
|
||||
"alternate_only": alternate_only_total,
|
||||
"local_only": local_only_total,
|
||||
"alternate_only_ratio": alternate_only_ratio,
|
||||
"alternate_only_sample": sorted(alternate_only)[:12],
|
||||
"local_only_sample": sorted(local_only)[:12],
|
||||
},
|
||||
"layout": {
|
||||
"local_text_items": int(local.get("text_items", 0)),
|
||||
"local_image_items": int(local.get("image_items", 0)),
|
||||
"local_repeated_x_anchors": local_anchors,
|
||||
"alternate_text_blocks": int(alternate.get("text_blocks", 0)),
|
||||
"alternate_text_lines": int(alternate.get("text_lines", 0)),
|
||||
"alternate_image_blocks": int(alternate.get("image_blocks", 0)),
|
||||
"alternate_repeated_x_anchors": alternate_anchors,
|
||||
},
|
||||
}
|
||||
|
||||
|
||||
def compare_documents(
|
||||
local_payload: dict[str, Any],
|
||||
alternate_payload: dict[str, Any] | list[dict[str, Any]],
|
||||
*,
|
||||
min_token_gain: int = 20,
|
||||
min_alternate_only_ratio: float = 0.15,
|
||||
min_anchor_gain: int = 2,
|
||||
) -> dict[str, Any]:
|
||||
"""Return a page-level evidence report for already extracted payloads."""
|
||||
local = local_pages(local_payload)
|
||||
alternate = alternate_pages(alternate_payload)
|
||||
page_numbers = sorted(local.keys() | alternate.keys())
|
||||
pages = []
|
||||
for page_number in page_numbers:
|
||||
result = compare_page(
|
||||
local.get(page_number, {}),
|
||||
alternate.get(page_number, {}),
|
||||
min_token_gain=min_token_gain,
|
||||
min_alternate_only_ratio=min_alternate_only_ratio,
|
||||
min_anchor_gain=min_anchor_gain,
|
||||
)
|
||||
result["page"] = page_number
|
||||
pages.append(result)
|
||||
|
||||
flagged = [
|
||||
page
|
||||
for page in pages
|
||||
if page["classification"] == "investigate_alternate_evidence"
|
||||
]
|
||||
return {
|
||||
"summary": {
|
||||
"pages": len(pages),
|
||||
"flagged_pages": len(flagged),
|
||||
"flagged_page_numbers": [page["page"] for page in flagged],
|
||||
"local_tokens": sum(page["tokens"]["local"] for page in pages),
|
||||
"alternate_tokens": sum(page["tokens"]["alternate"] for page in pages),
|
||||
"alternate_only_tokens": sum(
|
||||
page["tokens"]["alternate_only"] for page in pages
|
||||
),
|
||||
},
|
||||
"pages": pages,
|
||||
}
|
||||
|
||||
|
||||
def _json_command(command: list[str]) -> Any:
|
||||
try:
|
||||
completed = subprocess.run(
|
||||
command,
|
||||
check=True,
|
||||
stdout=subprocess.PIPE,
|
||||
stderr=subprocess.PIPE,
|
||||
text=True,
|
||||
)
|
||||
except subprocess.CalledProcessError as error:
|
||||
detail = error.stderr.strip() or error.stdout.strip() or "no diagnostic output"
|
||||
raise RuntimeError(f"command failed: {' '.join(command)}\n{detail}") from error
|
||||
try:
|
||||
return json.loads(completed.stdout)
|
||||
except json.JSONDecodeError as error:
|
||||
raise RuntimeError(
|
||||
f"command did not return JSON: {' '.join(command)}: {error}"
|
||||
) from error
|
||||
|
||||
|
||||
def probe_pdf(
|
||||
pdf: Path,
|
||||
*,
|
||||
pdf2md: Path,
|
||||
mutool: Path,
|
||||
min_token_gain: int,
|
||||
min_alternate_only_ratio: float,
|
||||
min_anchor_gain: int,
|
||||
) -> dict[str, Any]:
|
||||
local_payload = _json_command([str(pdf2md), str(pdf), "--items-json"])
|
||||
# `stext.json` is MuPDF's native structured text output. The OCR formats
|
||||
# are intentionally not used so this remains a deterministic no-model
|
||||
# comparison.
|
||||
alternate_payload = _json_command(
|
||||
[str(mutool), "draw", "-q", "-F", "stext.json", "-o", "-", str(pdf)]
|
||||
)
|
||||
report = compare_documents(
|
||||
local_payload,
|
||||
alternate_payload,
|
||||
min_token_gain=min_token_gain,
|
||||
min_alternate_only_ratio=min_alternate_only_ratio,
|
||||
min_anchor_gain=min_anchor_gain,
|
||||
)
|
||||
report["pdf"] = str(pdf)
|
||||
return report
|
||||
|
||||
|
||||
def _arguments(argv: list[str] | None = None) -> argparse.Namespace:
|
||||
parser = argparse.ArgumentParser(description=__doc__)
|
||||
parser.add_argument("pdf", type=Path, nargs="+")
|
||||
parser.add_argument("--pdf2md", type=Path, default=Path("target/release/pdf2md"))
|
||||
parser.add_argument("--mutool", type=Path)
|
||||
parser.add_argument("--json-output", type=Path)
|
||||
parser.add_argument("--min-token-gain", type=int, default=20)
|
||||
parser.add_argument("--min-alternate-only-ratio", type=float, default=0.15)
|
||||
parser.add_argument("--min-anchor-gain", type=int, default=2)
|
||||
return parser.parse_args(argv)
|
||||
|
||||
|
||||
def _print_report(result: dict[str, Any]) -> None:
|
||||
summary = result["summary"]
|
||||
print(f"\n{result['pdf']}")
|
||||
print(
|
||||
f" {summary['flagged_pages']}/{summary['pages']} pages flagged; "
|
||||
f"tokens local={summary['local_tokens']} alternate={summary['alternate_tokens']} "
|
||||
f"alternate-only={summary['alternate_only_tokens']}"
|
||||
)
|
||||
for page in result["pages"]:
|
||||
if page["classification"] != "investigate_alternate_evidence":
|
||||
continue
|
||||
reasons = ", ".join(page["reasons"])
|
||||
tokens = page["tokens"]
|
||||
print(
|
||||
f" page {page['page']}: {reasons}; "
|
||||
f"net tokens={tokens['net_alternate_gain']:+d}, "
|
||||
f"alternate-only={tokens['alternate_only']}"
|
||||
)
|
||||
|
||||
|
||||
def main(argv: list[str] | None = None) -> int:
|
||||
args = _arguments(argv)
|
||||
pdf2md = args.pdf2md.absolute()
|
||||
mutool = args.mutool or (Path(found) if (found := shutil.which("mutool")) else None)
|
||||
if not pdf2md.is_file():
|
||||
print(f"error: pdf2md binary not found: {pdf2md}", file=sys.stderr)
|
||||
return 2
|
||||
if mutool is None or not mutool.is_file():
|
||||
print("error: mutool not found; install MuPDF or pass --mutool", file=sys.stderr)
|
||||
return 2
|
||||
if (
|
||||
args.min_token_gain < 0
|
||||
or args.min_anchor_gain < 0
|
||||
or not 0.0 <= args.min_alternate_only_ratio <= 1.0
|
||||
):
|
||||
print("error: thresholds must be non-negative and ratio must be in [0, 1]", file=sys.stderr)
|
||||
return 2
|
||||
|
||||
results = []
|
||||
for pdf in args.pdf:
|
||||
path = pdf.absolute()
|
||||
if not path.is_file():
|
||||
print(f"error: PDF not found: {path}", file=sys.stderr)
|
||||
return 2
|
||||
try:
|
||||
result = probe_pdf(
|
||||
path,
|
||||
pdf2md=pdf2md,
|
||||
mutool=mutool,
|
||||
min_token_gain=args.min_token_gain,
|
||||
min_alternate_only_ratio=args.min_alternate_only_ratio,
|
||||
min_anchor_gain=args.min_anchor_gain,
|
||||
)
|
||||
except RuntimeError as error:
|
||||
print(f"error: {error}", file=sys.stderr)
|
||||
return 1
|
||||
results.append(result)
|
||||
_print_report(result)
|
||||
|
||||
payload = {
|
||||
"schema_version": 1,
|
||||
"experiment": "optional_mupdf_stext_evidence",
|
||||
"ocr": False,
|
||||
"thresholds": {
|
||||
"min_token_gain": args.min_token_gain,
|
||||
"min_alternate_only_ratio": args.min_alternate_only_ratio,
|
||||
"min_anchor_gain": args.min_anchor_gain,
|
||||
},
|
||||
"documents": results,
|
||||
}
|
||||
if args.json_output:
|
||||
args.json_output.parent.mkdir(parents=True, exist_ok=True)
|
||||
args.json_output.write_text(json.dumps(payload, indent=2) + "\n", encoding="utf-8")
|
||||
return 0
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
raise SystemExit(main())
|
||||
@@ -0,0 +1,203 @@
|
||||
import io
|
||||
import json
|
||||
import sys
|
||||
import tempfile
|
||||
import unittest
|
||||
from contextlib import redirect_stderr, redirect_stdout
|
||||
from pathlib import Path
|
||||
from unittest.mock import patch
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
|
||||
|
||||
from bench_opendataloader import (
|
||||
_arguments,
|
||||
_print_report,
|
||||
_run_engine,
|
||||
compare_evaluations,
|
||||
evaluate_gates,
|
||||
)
|
||||
|
||||
|
||||
def evaluation(overall, documents, *, missing=0):
|
||||
return {
|
||||
"metrics": {
|
||||
"score": {
|
||||
"overall_mean": overall,
|
||||
"nid_mean": overall + 0.01,
|
||||
},
|
||||
"missing_predictions": missing,
|
||||
},
|
||||
"documents": [
|
||||
{
|
||||
"document_id": document_id,
|
||||
"scores": {"overall": score},
|
||||
}
|
||||
for document_id, score in documents.items()
|
||||
],
|
||||
}
|
||||
|
||||
|
||||
class ComparisonTests(unittest.TestCase):
|
||||
def test_reports_metric_and_document_deltas(self):
|
||||
baseline = evaluation(0.80, {"a": 0.8, "b": 0.6, "c": 0.7})
|
||||
candidate = evaluation(0.82, {"a": 0.9, "b": 0.5, "c": 0.7})
|
||||
|
||||
result = compare_evaluations(baseline, candidate, top=1)
|
||||
|
||||
self.assertAlmostEqual(result["deltas"]["overall_mean"], 0.02)
|
||||
self.assertEqual(result["documents"]["improved"], 1)
|
||||
self.assertEqual(result["documents"]["regressed"], 1)
|
||||
self.assertEqual(result["documents"]["unchanged"], 1)
|
||||
self.assertEqual(
|
||||
result["documents"]["largest_improvements"][0]["document_id"], "a"
|
||||
)
|
||||
self.assertEqual(
|
||||
result["documents"]["largest_regressions"][0]["document_id"], "b"
|
||||
)
|
||||
|
||||
def test_reference_delta_is_reported(self):
|
||||
baseline = evaluation(0.80, {})
|
||||
candidate = evaluation(0.82, {})
|
||||
reference = evaluation(0.81, {})
|
||||
|
||||
result = compare_evaluations(baseline, candidate, reference)
|
||||
|
||||
self.assertAlmostEqual(
|
||||
result["candidate_vs_reference"]["overall_mean"], 0.01
|
||||
)
|
||||
|
||||
def test_gates_cover_aggregate_document_missing_and_reference(self):
|
||||
comparison = compare_evaluations(
|
||||
evaluation(0.80, {"a": 0.8}),
|
||||
evaluation(0.79, {"a": 0.7}, missing=1),
|
||||
evaluation(0.81, {}),
|
||||
)
|
||||
|
||||
failures = evaluate_gates(
|
||||
comparison,
|
||||
min_overall_delta=0.0,
|
||||
max_document_regression=0.05,
|
||||
max_missing=0,
|
||||
require_reference_lead=True,
|
||||
)
|
||||
|
||||
self.assertEqual(len(failures), 4)
|
||||
|
||||
def test_regression_gate_is_independent_of_report_limit(self):
|
||||
comparison = compare_evaluations(
|
||||
evaluation(0.80, {"a": 0.8}),
|
||||
evaluation(0.80, {"a": 0.7}),
|
||||
top=0,
|
||||
)
|
||||
|
||||
failures = evaluate_gates(
|
||||
comparison,
|
||||
min_overall_delta=0.0,
|
||||
max_document_regression=0.05,
|
||||
max_missing=0,
|
||||
require_reference_lead=False,
|
||||
)
|
||||
|
||||
self.assertEqual(len(failures), 1)
|
||||
self.assertIn("largest document regression", failures[0])
|
||||
|
||||
def test_report_handles_reference_without_overall_score(self):
|
||||
result = compare_evaluations(
|
||||
evaluation(0.80, {}),
|
||||
evaluation(0.82, {}),
|
||||
{"metrics": {"score": {"nid_mean": 0.81}}},
|
||||
)
|
||||
|
||||
output = io.StringIO()
|
||||
with redirect_stdout(output):
|
||||
_print_report(result)
|
||||
|
||||
self.assertIn("Reference overall: n/a; candidate delta: n/a", output.getvalue())
|
||||
|
||||
def test_reference_gate_reports_missing_score_as_unavailable(self):
|
||||
comparison = compare_evaluations(
|
||||
evaluation(0.80, {}),
|
||||
evaluation(0.82, {}),
|
||||
)
|
||||
|
||||
failures = evaluate_gates(
|
||||
comparison,
|
||||
min_overall_delta=0.0,
|
||||
max_document_regression=None,
|
||||
max_missing=0,
|
||||
require_reference_lead=True,
|
||||
)
|
||||
|
||||
self.assertEqual(failures, ["reference overall score is unavailable"])
|
||||
|
||||
def test_arguments_reject_negative_counts_and_allow_zero_top(self):
|
||||
required = [
|
||||
"--bench-dir",
|
||||
".",
|
||||
"--baseline",
|
||||
"baseline",
|
||||
"--candidate",
|
||||
"candidate",
|
||||
]
|
||||
self.assertEqual(_arguments(required + ["--top", "0"]).top, 0)
|
||||
for option in ("--top", "--max-document-regression", "--max-missing"):
|
||||
with self.subTest(option=option), redirect_stderr(io.StringIO()):
|
||||
with self.assertRaises(SystemExit):
|
||||
_arguments(required + [option, "-1"])
|
||||
|
||||
def test_arguments_reject_nonfinite_float_thresholds(self):
|
||||
required = [
|
||||
"--bench-dir",
|
||||
".",
|
||||
"--baseline",
|
||||
"baseline",
|
||||
"--candidate",
|
||||
"candidate",
|
||||
]
|
||||
for option in ("--min-overall-delta", "--max-document-regression"):
|
||||
for value in ("nan", "inf", "-inf"):
|
||||
with self.subTest(option=option, value=value), redirect_stderr(
|
||||
io.StringIO()
|
||||
):
|
||||
with self.assertRaises(SystemExit):
|
||||
_arguments(required + [option, value])
|
||||
|
||||
def test_run_engine_clears_stale_predictions_before_parser(self):
|
||||
with tempfile.TemporaryDirectory() as temporary:
|
||||
root = Path(temporary)
|
||||
bench_dir = root / "bench"
|
||||
source = bench_dir / "prediction" / "pdf-inspector"
|
||||
source.mkdir(parents=True)
|
||||
(source / "stale.md").write_text("stale", encoding="utf-8")
|
||||
scratch = root / "scratch"
|
||||
scratch.mkdir()
|
||||
|
||||
def fake_run(command, *, cwd, env=None):
|
||||
if any(part.endswith("pdf_parser.py") for part in command):
|
||||
self.assertFalse(source.exists())
|
||||
(source / "markdown").mkdir(parents=True)
|
||||
(source / "markdown" / "new.md").write_text(
|
||||
"new", encoding="utf-8"
|
||||
)
|
||||
else:
|
||||
destination = scratch / "candidate"
|
||||
(destination / "evaluation.json").write_text(
|
||||
json.dumps(evaluation(0.82, {})), encoding="utf-8"
|
||||
)
|
||||
|
||||
with patch("bench_opendataloader._run", side_effect=fake_run):
|
||||
result = _run_engine(
|
||||
bench_dir=bench_dir,
|
||||
python=Path("python"),
|
||||
binary=Path("pdf2md"),
|
||||
label="candidate",
|
||||
scratch_root=scratch,
|
||||
)
|
||||
|
||||
self.assertEqual(result["metrics"]["score"]["overall_mean"], 0.82)
|
||||
self.assertFalse((source / "stale.md").exists())
|
||||
self.assertFalse((scratch / "candidate" / "stale.md").exists())
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
@@ -0,0 +1,103 @@
|
||||
import sys
|
||||
import unittest
|
||||
from pathlib import Path
|
||||
|
||||
sys.path.insert(0, str(Path(__file__).resolve().parents[1]))
|
||||
|
||||
from probe_backend_evidence import compare_documents
|
||||
|
||||
|
||||
def local_payload(items):
|
||||
return {"items": items}
|
||||
|
||||
|
||||
def item(page, text, x=10, item_type="text"):
|
||||
return {"page": page, "text": text, "x": x, "item_type": item_type}
|
||||
|
||||
|
||||
def alternate_payload(pages):
|
||||
return {"pages": pages}
|
||||
|
||||
|
||||
def page(lines, *, images=0):
|
||||
blocks = [
|
||||
{
|
||||
"type": "text",
|
||||
"lines": [
|
||||
{"text": text, "bbox": {"x": x, "y": index * 10, "w": 80, "h": 8}}
|
||||
for index, (text, x) in enumerate(lines)
|
||||
],
|
||||
}
|
||||
]
|
||||
blocks.extend({"type": "image"} for _ in range(images))
|
||||
return {"blocks": blocks}
|
||||
|
||||
|
||||
class EvidenceComparisonTests(unittest.TestCase):
|
||||
def test_accepts_real_top_level_page_array(self):
|
||||
local = local_payload([item(1, "alpha beta")])
|
||||
alternate = [page([("alpha beta gamma", 10)])]
|
||||
|
||||
result = compare_documents(local, alternate)["pages"][0]
|
||||
|
||||
self.assertEqual(result["tokens"]["alternate"], 3)
|
||||
self.assertEqual(result["tokens"]["net_alternate_gain"], 1)
|
||||
|
||||
def test_flags_material_alternate_text_gain(self):
|
||||
local = local_payload([item(1, "alpha beta")])
|
||||
alternate = alternate_payload(
|
||||
[page([("alpha beta gamma delta epsilon zeta", 10)])]
|
||||
)
|
||||
|
||||
report = compare_documents(
|
||||
local,
|
||||
alternate,
|
||||
min_token_gain=3,
|
||||
min_alternate_only_ratio=0.2,
|
||||
)
|
||||
|
||||
result = report["pages"][0]
|
||||
self.assertEqual(result["classification"], "investigate_alternate_evidence")
|
||||
self.assertIn("alternate_has_more_text", result["reasons"])
|
||||
self.assertEqual(result["tokens"]["net_alternate_gain"], 4)
|
||||
|
||||
def test_repeated_alignment_and_image_evidence_are_reported(self):
|
||||
local = local_payload([item(1, "one two", 10)])
|
||||
alternate = alternate_payload(
|
||||
[
|
||||
page(
|
||||
[
|
||||
("one two", 10),
|
||||
("row three", 100),
|
||||
("row four", 100),
|
||||
("row five", 100),
|
||||
],
|
||||
images=1,
|
||||
)
|
||||
]
|
||||
)
|
||||
|
||||
result = compare_documents(
|
||||
local,
|
||||
alternate,
|
||||
min_token_gain=99,
|
||||
min_anchor_gain=1,
|
||||
)["pages"][0]
|
||||
|
||||
self.assertIn("alternate_has_more_alignment_anchors", result["reasons"])
|
||||
self.assertIn("alternate_has_more_image_blocks", result["reasons"])
|
||||
self.assertEqual(result["layout"]["alternate_repeated_x_anchors"], 1)
|
||||
|
||||
def test_token_segmentation_difference_does_not_imply_more_evidence(self):
|
||||
local = local_payload([item(1, "Revenue 2025")])
|
||||
alternate = alternate_payload([page([("Revenue 2024", 10)])])
|
||||
|
||||
result = compare_documents(local, alternate, min_token_gain=2)["pages"][0]
|
||||
|
||||
self.assertEqual(result["classification"], "different_segmentation_or_decoding")
|
||||
self.assertEqual(result["reasons"], [])
|
||||
self.assertEqual(result["tokens"]["alternate_only_sample"], ["2024"])
|
||||
|
||||
|
||||
if __name__ == "__main__":
|
||||
unittest.main()
|
||||
+9
-10
@@ -147,8 +147,7 @@
|
||||
tbody tr.us td:first-child::before { content: "▸ "; color: var(--accent); }
|
||||
.bench-foot { padding: 15px 18px; font-size: 13.5px; color: var(--ink-soft); background: var(--paper-2); border-top: 1px solid var(--line); }
|
||||
.bench-wrap { overflow-x: auto; }
|
||||
.callouts { display: grid; grid-template-columns: 1fr 1fr; gap: 16px; margin-top: 22px; }
|
||||
@media (max-width: 640px) { .callouts { grid-template-columns: 1fr; } }
|
||||
.callouts { display: grid; grid-template-columns: 1fr; gap: 16px; margin-top: 22px; }
|
||||
.callout { border-left: 3px solid var(--accent); padding: 4px 0 4px 16px; }
|
||||
.callout .ct { font-family: var(--mono); font-size: 11px; letter-spacing: 0.1em; text-transform: uppercase; color: var(--accent-deep); margin-bottom: 5px; }
|
||||
.callout p { font-size: 14.5px; color: var(--ink-soft); }
|
||||
@@ -301,7 +300,7 @@
|
||||
<div class="sec-head">
|
||||
<span class="sec-num">02</span>
|
||||
<h2 class="sec-title">Fastest of the direct-text engines</h2>
|
||||
<p class="sec-sub">Evaluated on the <a href="https://github.com/opendataloader-project/opendataloader-bench" style="color:var(--accent-deep);border-bottom:1px solid var(--line)">opendataloader-bench</a> corpus (200 PDFs). Direct text-extraction engines only — no OCR, no ML. Higher is better.</p>
|
||||
<p class="sec-sub">Evaluated on the <a href="https://github.com/opendataloader-project/opendataloader-bench" style="color:var(--accent-deep);border-bottom:1px solid var(--line)">opendataloader-bench</a> corpus (200 PDFs). Local engines without model-based PDF parsing; OCR disabled. Higher is better.</p>
|
||||
</div>
|
||||
<div class="bench">
|
||||
<div class="bench-wrap">
|
||||
@@ -310,18 +309,18 @@
|
||||
<tr><th>Engine</th><th>Overall</th><th>Reading order</th><th>Tables</th><th>Headings</th><th>200 docs</th></tr>
|
||||
</thead>
|
||||
<tbody>
|
||||
<tr class="us"><td>pdf-inspector</td><td>0.83</td><td>0.89</td><td>0.66</td><td>0.74</td><td>4s</td></tr>
|
||||
<tr><td>opendataloader</td><td>0.84</td><td>0.91</td><td>0.49</td><td>0.74</td><td>11s</td></tr>
|
||||
<tr><td>pymupdf4llm</td><td>0.73</td><td>0.89</td><td>0.40</td><td>0.41</td><td>18s</td></tr>
|
||||
<tr><td>markitdown</td><td>0.58</td><td>0.88</td><td>0.00</td><td>0.00</td><td>8s</td></tr>
|
||||
<tr class="us"><td>pdf-inspector</td><td>0.875</td><td>0.915</td><td>0.814</td><td>0.788</td><td>2.8s</td></tr>
|
||||
<tr><td>liteparse</td><td>0.870</td><td>0.908</td><td>0.693</td><td>0.811</td><td>13.9s</td></tr>
|
||||
<tr><td>opendataloader</td><td>0.843</td><td>0.912</td><td>0.489</td><td>0.760</td><td>9.8s</td></tr>
|
||||
<tr><td>pymupdf4llm</td><td>0.735</td><td>0.886</td><td>0.401</td><td>0.424</td><td>15.5s</td></tr>
|
||||
<tr><td>markitdown</td><td>0.583</td><td>0.879</td><td>0.000</td><td>0.000</td><td>6.7s</td></tr>
|
||||
</tbody>
|
||||
</table>
|
||||
</div>
|
||||
<div class="bench-foot">OCR/ML engines (docling, marker, mineru) score 0.83–0.88 overall — but take 2–180 minutes on the same corpus.</div>
|
||||
<div class="bench-foot">Refreshed July 16, 2026, on Apple M4 Pro. Speed is the median of three complete corpus runs.</div>
|
||||
</div>
|
||||
<div class="callouts">
|
||||
<div class="callout"><div class="ct">Where we win</div><p>Fastest engine measured, the best table detection of any engine here, and heading quality now on par with opendataloader — at ~2.5× its speed.</p></div>
|
||||
<div class="callout"><div class="ct">Where we're working</div><p>Reading order still trails opendataloader slightly, and tables that need the visual structure only an OCR engine can see.</p></div>
|
||||
<div class="callout"><div class="ct">Best fit</div><p>Native-text PDFs where speed, reading order, and table structure matter. pdf-inspector delivered the highest overall, reading-order, and table scores, along with the fastest complete run in this benchmark. That makes it a strong local default for reports, research papers, financial documents, invoices, and legal PDFs that need clean, structured Markdown without adding OCR latency or infrastructure.</p></div>
|
||||
</div>
|
||||
</div>
|
||||
</section>
|
||||
|
||||
@@ -1179,6 +1179,22 @@ pub(crate) fn group_into_lines_with_thresholds_and_charts(
|
||||
page_thresholds: &HashMap<u32, f32>,
|
||||
table_pages: &HashSet<u32>,
|
||||
chart_regions: &HashMap<u32, Vec<(f32, f32, f32, f32)>>,
|
||||
) -> Vec<TextLine> {
|
||||
group_into_lines_with_thresholds_and_regions(
|
||||
items,
|
||||
page_thresholds,
|
||||
table_pages,
|
||||
chart_regions,
|
||||
&HashMap::new(),
|
||||
)
|
||||
}
|
||||
|
||||
pub(crate) fn group_into_lines_with_thresholds_and_regions(
|
||||
items: Vec<TextItem>,
|
||||
page_thresholds: &HashMap<u32, f32>,
|
||||
table_pages: &HashSet<u32>,
|
||||
chart_regions: &HashMap<u32, Vec<(f32, f32, f32, f32)>>,
|
||||
image_regions: &HashMap<u32, Vec<super::reading_order::ImageRegion>>,
|
||||
) -> Vec<TextLine> {
|
||||
if items.is_empty() {
|
||||
return Vec::new();
|
||||
@@ -1205,6 +1221,38 @@ pub(crate) fn group_into_lines_with_thresholds_and_charts(
|
||||
// Non-Canva pages use the default 0.10 threshold.
|
||||
let adaptive_threshold = page_thresholds.get(&page).copied().unwrap_or(0.10);
|
||||
|
||||
// Image-backed region graphs recover local/asymmetric column flows
|
||||
// that a whole-page projection cannot represent. Charts already have
|
||||
// their own positioned-region ordering and therefore stay on that path.
|
||||
if !chart_regions.contains_key(&page) {
|
||||
let preliminary_columns =
|
||||
detect_columns(&page_items, page, table_pages.contains(&page));
|
||||
let detected_split =
|
||||
(preliminary_columns.len() == 2).then_some(preliminary_columns[0].x_max);
|
||||
if let Some(band) = image_regions.get(&page).and_then(|regions| {
|
||||
super::reading_order::infer_image_anchored_flow(
|
||||
&page_items,
|
||||
regions,
|
||||
detected_split,
|
||||
)
|
||||
}) {
|
||||
debug!(
|
||||
"page {}: image-anchored region graph split={:.1} y=[{:.1}..{:.1}]",
|
||||
page, band.split_x, band.y_bottom, band.y_top
|
||||
);
|
||||
for node in super::reading_order::build_region_graph(page_items, band) {
|
||||
debug!(
|
||||
"page {}: region node {:?} items={}",
|
||||
page,
|
||||
node.kind,
|
||||
node.items.len()
|
||||
);
|
||||
all_lines.extend(group_single_column(node.items, adaptive_threshold));
|
||||
}
|
||||
continue;
|
||||
}
|
||||
}
|
||||
|
||||
// Detect columns for this page, blind to chart text.
|
||||
debug!(
|
||||
"page {}: grouping chart-aware={} regions={:?}",
|
||||
|
||||
@@ -6,6 +6,7 @@ pub(crate) mod content_stream;
|
||||
mod fonts;
|
||||
mod layout;
|
||||
mod links;
|
||||
mod reading_order;
|
||||
pub(crate) mod underline;
|
||||
mod xobjects;
|
||||
|
||||
@@ -29,6 +30,7 @@ pub(crate) use layout::detect_columns;
|
||||
pub use layout::group_into_lines;
|
||||
pub(crate) use layout::group_into_lines_with_thresholds;
|
||||
pub(crate) use layout::group_into_lines_with_thresholds_and_charts;
|
||||
pub(crate) use layout::group_into_lines_with_thresholds_and_regions;
|
||||
pub(crate) use layout::is_newspaper_layout;
|
||||
pub(crate) use layout::ColumnRegion;
|
||||
|
||||
|
||||
@@ -0,0 +1,591 @@
|
||||
//! Region-graph evidence for page reading order.
|
||||
//!
|
||||
//! Whole-page column histograms fail when images or spanning captions occupy
|
||||
//! only part of a page. This module turns image geometry and repeated row
|
||||
//! gutters into a small directed acyclic graph: content above a local column
|
||||
//! band, the left flow, the right flow, and content below it. The graph is
|
||||
//! deliberately evidence-gated; ordinary pages keep the established layout
|
||||
//! path.
|
||||
|
||||
use crate::text_utils::{effective_width, is_cjk_char, is_rtl_text};
|
||||
use crate::types::TextItem;
|
||||
|
||||
const MIN_IMAGE_WIDTH: f32 = 60.0;
|
||||
const MIN_IMAGE_HEIGHT: f32 = 40.0;
|
||||
const MIN_ROW_GUTTER: f32 = 8.0;
|
||||
const SPLIT_CLUSTER_TOLERANCE: f32 = 20.0;
|
||||
const MIN_ALIGNED_ROWS: usize = 4;
|
||||
|
||||
pub(crate) type ImageRegion = (f32, f32, f32, f32);
|
||||
|
||||
#[derive(Debug, Clone, Copy, PartialEq)]
|
||||
pub(crate) struct ColumnFlowBand {
|
||||
pub(crate) split_x: f32,
|
||||
pub(crate) y_bottom: f32,
|
||||
pub(crate) y_top: f32,
|
||||
}
|
||||
|
||||
#[derive(Debug, Clone, Copy, PartialEq, Eq)]
|
||||
pub(crate) enum RegionKind {
|
||||
FullWidth,
|
||||
Column,
|
||||
}
|
||||
|
||||
#[derive(Debug)]
|
||||
pub(crate) struct RegionNode {
|
||||
pub(crate) kind: RegionKind,
|
||||
pub(crate) items: Vec<TextItem>,
|
||||
}
|
||||
|
||||
#[derive(Debug)]
|
||||
struct Row<'a> {
|
||||
y: f32,
|
||||
items: Vec<&'a TextItem>,
|
||||
}
|
||||
|
||||
fn page_x_bounds(items: &[TextItem], images: &[ImageRegion]) -> Option<(f32, f32)> {
|
||||
let text_min = items
|
||||
.iter()
|
||||
.map(|item| item.x)
|
||||
.fold(f32::INFINITY, f32::min);
|
||||
let text_max = items
|
||||
.iter()
|
||||
.map(|item| item.x + effective_width(item))
|
||||
.fold(f32::NEG_INFINITY, f32::max);
|
||||
let image_min = images
|
||||
.iter()
|
||||
.map(|region| region.0.min(region.2))
|
||||
.fold(f32::INFINITY, f32::min);
|
||||
let image_max = images
|
||||
.iter()
|
||||
.map(|region| region.0.max(region.2))
|
||||
.fold(f32::NEG_INFINITY, f32::max);
|
||||
let x_min = text_min.min(image_min);
|
||||
let x_max = text_max.max(image_max);
|
||||
(x_min.is_finite() && x_max.is_finite() && x_max > x_min).then_some((x_min, x_max))
|
||||
}
|
||||
|
||||
fn group_rows(items: &[TextItem]) -> Vec<Row<'_>> {
|
||||
const Y_TOLERANCE: f32 = 3.0;
|
||||
let mut sorted: Vec<&TextItem> = items.iter().collect();
|
||||
sorted.sort_by(|left, right| right.y.total_cmp(&left.y));
|
||||
let mut rows: Vec<Row<'_>> = Vec::new();
|
||||
for item in sorted {
|
||||
if let Some(row) = rows
|
||||
.last_mut()
|
||||
.filter(|row| (row.y - item.y).abs() <= Y_TOLERANCE)
|
||||
{
|
||||
row.items.push(item);
|
||||
row.y = row.items.iter().map(|member| member.y).sum::<f32>() / row.items.len() as f32;
|
||||
} else {
|
||||
rows.push(Row {
|
||||
y: item.y,
|
||||
items: vec![item],
|
||||
});
|
||||
}
|
||||
}
|
||||
for row in &mut rows {
|
||||
row.items.sort_by(|left, right| left.x.total_cmp(&right.x));
|
||||
}
|
||||
rows
|
||||
}
|
||||
|
||||
fn side_is_prose(items: &[&TextItem]) -> bool {
|
||||
let text = items
|
||||
.iter()
|
||||
.map(|item| item.text.trim())
|
||||
.collect::<Vec<_>>()
|
||||
.join(" ");
|
||||
let alphabetic_count = text
|
||||
.chars()
|
||||
.filter(|character| character.is_alphabetic())
|
||||
.count();
|
||||
let cjk_count = text
|
||||
.chars()
|
||||
.filter(|character| is_cjk_char(*character))
|
||||
.count();
|
||||
(text.split_whitespace().count() >= 3 || cjk_count >= 10) && alphabetic_count >= 10
|
||||
}
|
||||
|
||||
fn aligned_row_split(row: &Row<'_>, x_min: f32, x_max: f32) -> Option<f32> {
|
||||
if row.items.len() < 2 {
|
||||
return None;
|
||||
}
|
||||
let page_width = x_max - x_min;
|
||||
let center_low = x_min + page_width * 0.25;
|
||||
let center_high = x_min + page_width * 0.75;
|
||||
row.items
|
||||
.windows(2)
|
||||
.filter_map(|pair| {
|
||||
let left_end = pair[0].x + effective_width(pair[0]);
|
||||
let right_start = pair[1].x;
|
||||
let gap = right_start - left_end;
|
||||
let split_x = (left_end + right_start) / 2.0;
|
||||
if gap < MIN_ROW_GUTTER || split_x < center_low || split_x > center_high {
|
||||
return None;
|
||||
}
|
||||
let left: Vec<&TextItem> = row
|
||||
.items
|
||||
.iter()
|
||||
.copied()
|
||||
.filter(|item| item.x + effective_width(item) / 2.0 < split_x)
|
||||
.collect();
|
||||
let right: Vec<&TextItem> = row
|
||||
.items
|
||||
.iter()
|
||||
.copied()
|
||||
.filter(|item| item.x + effective_width(item) / 2.0 >= split_x)
|
||||
.collect();
|
||||
(side_is_prose(&left) && side_is_prose(&right)).then_some((split_x, gap))
|
||||
})
|
||||
.max_by(|left, right| left.1.total_cmp(&right.1))
|
||||
.map(|candidate| candidate.0)
|
||||
}
|
||||
|
||||
fn local_flow_below_full_width_image(
|
||||
items: &[TextItem],
|
||||
images: &[ImageRegion],
|
||||
x_min: f32,
|
||||
x_max: f32,
|
||||
) -> Option<ColumnFlowBand> {
|
||||
let page_width = x_max - x_min;
|
||||
let full_width_images: Vec<ImageRegion> = images
|
||||
.iter()
|
||||
.copied()
|
||||
.filter(|&(x0, y0, x1, y1)| {
|
||||
let width = (x1 - x0).abs();
|
||||
let height = (y1 - y0).abs();
|
||||
width >= page_width * 0.65 && height >= 60.0
|
||||
})
|
||||
.collect();
|
||||
// A local column flow below an image is only unambiguous for a single,
|
||||
// nearly square hero/figure. Wide report banners and full-page artwork
|
||||
// frequently sit above unrelated page furniture whose aligned labels can
|
||||
// mimic prose columns.
|
||||
if full_width_images.len() != 1 {
|
||||
return None;
|
||||
}
|
||||
let (image_x0, _, image_x1, _) = full_width_images[0];
|
||||
let anchor_width = (image_x1 - image_x0).abs();
|
||||
let anchor_height = (full_width_images[0].3 - full_width_images[0].1).abs();
|
||||
if anchor_width < page_width * 0.85
|
||||
|| anchor_height < anchor_width * 0.85
|
||||
|| anchor_height > anchor_width * 1.2
|
||||
{
|
||||
return None;
|
||||
}
|
||||
let image_bottom = full_width_images
|
||||
.iter()
|
||||
.map(|&(_, y0, _, y1)| y0.min(y1))
|
||||
.fold(f32::NEG_INFINITY, f32::max);
|
||||
if !image_bottom.is_finite() {
|
||||
return None;
|
||||
}
|
||||
|
||||
let below: Vec<TextItem> = items
|
||||
.iter()
|
||||
.filter(|item| item.y < image_bottom && item.y >= image_bottom - 220.0)
|
||||
.cloned()
|
||||
.collect();
|
||||
let candidates: Vec<(f32, f32)> = group_rows(&below)
|
||||
.into_iter()
|
||||
.filter_map(|row| aligned_row_split(&row, x_min, x_max).map(|split| (split, row.y)))
|
||||
.collect();
|
||||
if candidates.len() < MIN_ALIGNED_ROWS {
|
||||
return None;
|
||||
}
|
||||
|
||||
let mut clusters: Vec<Vec<(f32, f32)>> = Vec::new();
|
||||
for candidate in candidates {
|
||||
if let Some(cluster) = clusters.iter_mut().find(|cluster| {
|
||||
let mean = cluster.iter().map(|entry| entry.0).sum::<f32>() / cluster.len() as f32;
|
||||
(mean - candidate.0).abs() <= SPLIT_CLUSTER_TOLERANCE
|
||||
}) {
|
||||
cluster.push(candidate);
|
||||
} else {
|
||||
clusters.push(vec![candidate]);
|
||||
}
|
||||
}
|
||||
let dominant = clusters.into_iter().max_by_key(Vec::len)?;
|
||||
if dominant.len() < MIN_ALIGNED_ROWS {
|
||||
return None;
|
||||
}
|
||||
let split_x = dominant.iter().map(|entry| entry.0).sum::<f32>() / dominant.len() as f32;
|
||||
let y_top = dominant
|
||||
.iter()
|
||||
.map(|entry| entry.1)
|
||||
.fold(f32::NEG_INFINITY, f32::max)
|
||||
+ 3.0;
|
||||
let image_gap = image_bottom - y_top;
|
||||
if !(60.0..=120.0).contains(&image_gap) {
|
||||
return None;
|
||||
}
|
||||
let y_bottom = dominant
|
||||
.iter()
|
||||
.map(|entry| entry.1)
|
||||
.fold(f32::INFINITY, f32::min)
|
||||
- 3.0;
|
||||
if y_top - y_bottom > 130.0 {
|
||||
return None;
|
||||
}
|
||||
log::debug!(
|
||||
"page {}: full-width image flow images={} aligned_rows={} split={:.1} page=[{:.1}..{:.1}] image_bottom={:.1} y=[{:.1}..{:.1}] full_width={:?}",
|
||||
items.first().map_or(0, |item| item.page),
|
||||
images.len(),
|
||||
dominant.len(),
|
||||
split_x,
|
||||
x_min,
|
||||
x_max,
|
||||
image_bottom,
|
||||
y_bottom,
|
||||
y_top,
|
||||
full_width_images
|
||||
);
|
||||
Some(ColumnFlowBand {
|
||||
split_x,
|
||||
y_bottom,
|
||||
y_top,
|
||||
})
|
||||
}
|
||||
|
||||
fn paired_column_images(
|
||||
items: &[TextItem],
|
||||
images: &[ImageRegion],
|
||||
split_x: f32,
|
||||
x_min: f32,
|
||||
x_max: f32,
|
||||
) -> Option<ColumnFlowBand> {
|
||||
let page_width = x_max - x_min;
|
||||
if split_x < x_min + page_width * 0.4 || split_x > x_min + page_width * 0.6 {
|
||||
return None;
|
||||
}
|
||||
let qualifying: Vec<ImageRegion> = images
|
||||
.iter()
|
||||
.copied()
|
||||
.filter(|&(x0, y0, x1, y1)| {
|
||||
let image_left = x0.min(x1);
|
||||
let image_right = x0.max(x1);
|
||||
let confined_to_one_column = image_right <= split_x || image_left >= split_x;
|
||||
confined_to_one_column
|
||||
&& (x1 - x0).abs() >= MIN_IMAGE_WIDTH
|
||||
&& (y1 - y0).abs() >= MIN_IMAGE_HEIGHT
|
||||
})
|
||||
.collect();
|
||||
let wide_images: Vec<ImageRegion> = qualifying
|
||||
.iter()
|
||||
.copied()
|
||||
.filter(|(x0, _, x1, _)| (x1 - x0).abs() >= page_width * 0.35)
|
||||
.collect();
|
||||
if qualifying.len() < 3 || wide_images.len() < 3 {
|
||||
return None;
|
||||
}
|
||||
let has_left = qualifying
|
||||
.iter()
|
||||
.any(|&(x0, _, x1, _)| (x0 + x1) / 2.0 < split_x);
|
||||
let has_right = qualifying
|
||||
.iter()
|
||||
.any(|&(x0, _, x1, _)| (x0 + x1) / 2.0 >= split_x);
|
||||
if !has_left || !has_right {
|
||||
return None;
|
||||
}
|
||||
// A meaningful image-backed column flow spans multiple vertical panels.
|
||||
// Three same-row header/logo images can otherwise satisfy the image count
|
||||
// and send an ordinary asymmetric page through sequential column order.
|
||||
let image_y_min = wide_images
|
||||
.iter()
|
||||
.map(|region| region.1.min(region.3))
|
||||
.fold(f32::INFINITY, f32::min);
|
||||
let image_y_max = wide_images
|
||||
.iter()
|
||||
.map(|region| region.1.max(region.3))
|
||||
.fold(f32::NEG_INFINITY, f32::max);
|
||||
let has_vertical_stack = wide_images.iter().enumerate().any(|(index, left)| {
|
||||
wide_images.iter().skip(index + 1).any(|right| {
|
||||
let same_side =
|
||||
((left.0 + left.2) / 2.0 < split_x) == ((right.0 + right.2) / 2.0 < split_x);
|
||||
let left_center = (left.1 + left.3) / 2.0;
|
||||
let right_center = (right.1 + right.3) / 2.0;
|
||||
let left_height = (left.3 - left.1).abs();
|
||||
let right_height = (right.3 - right.1).abs();
|
||||
let vertical_gap = if left.1.max(left.3) < right.1.min(right.3) {
|
||||
right.1.min(right.3) - left.1.max(left.3)
|
||||
} else if right.1.max(right.3) < left.1.min(left.3) {
|
||||
left.1.min(left.3) - right.1.max(right.3)
|
||||
} else {
|
||||
0.0
|
||||
};
|
||||
same_side
|
||||
&& (left_center - right_center).abs() >= left_height.min(right_height) * 0.5
|
||||
&& vertical_gap <= left_height.max(right_height) * 0.5
|
||||
})
|
||||
});
|
||||
if image_y_max - image_y_min < page_width * 0.45 || !has_vertical_stack {
|
||||
return None;
|
||||
}
|
||||
let y_top = qualifying
|
||||
.iter()
|
||||
.map(|region| region.1.max(region.3))
|
||||
.fold(f32::NEG_INFINITY, f32::max)
|
||||
+ 3.0;
|
||||
// Only column-confined text proves the lower extent of the flow. A
|
||||
// spanning heading or caption below the columns must become the trailing
|
||||
// full-width node rather than stretching the column band to the page foot.
|
||||
let y_bottom = items
|
||||
.iter()
|
||||
.filter(|item| {
|
||||
let item_right = item.x + effective_width(item);
|
||||
item.y <= y_top && (item_right <= split_x || item.x >= split_x)
|
||||
})
|
||||
.map(|item| item.y)
|
||||
.fold(f32::INFINITY, f32::min)
|
||||
- 3.0;
|
||||
if !y_bottom.is_finite() {
|
||||
return None;
|
||||
}
|
||||
let distinct_rows = |right: bool| {
|
||||
let mut ys: Vec<f32> = items
|
||||
.iter()
|
||||
.filter(|item| {
|
||||
item.y <= y_top && (item.x + effective_width(item) / 2.0 >= split_x) == right
|
||||
})
|
||||
.map(|item| item.y)
|
||||
.collect();
|
||||
ys.sort_by(|left, right| left.total_cmp(right));
|
||||
ys.dedup_by(|left, right| (*left - *right).abs() <= 3.0);
|
||||
ys.len()
|
||||
};
|
||||
let left_rows = distinct_rows(false);
|
||||
let right_rows = distinct_rows(true);
|
||||
let line_balance = left_rows.min(right_rows) as f32 / left_rows.max(right_rows).max(1) as f32;
|
||||
(left_rows >= 5 && right_rows >= 5 && line_balance < 0.55).then(|| {
|
||||
log::debug!(
|
||||
"page {}: paired-image flow qualifying_images={} rows={}/{} split={:.1} page=[{:.1}..{:.1}] y=[{:.1}..{:.1}] images={:?}",
|
||||
items.first().map_or(0, |item| item.page),
|
||||
qualifying.len(),
|
||||
left_rows,
|
||||
right_rows,
|
||||
split_x,
|
||||
x_min,
|
||||
x_max,
|
||||
y_bottom,
|
||||
y_top,
|
||||
qualifying
|
||||
);
|
||||
ColumnFlowBand {
|
||||
split_x,
|
||||
y_bottom,
|
||||
y_top,
|
||||
}
|
||||
})
|
||||
}
|
||||
|
||||
pub(crate) fn infer_image_anchored_flow(
|
||||
items: &[TextItem],
|
||||
images: &[ImageRegion],
|
||||
detected_split: Option<f32>,
|
||||
) -> Option<ColumnFlowBand> {
|
||||
if items.is_empty() || images.is_empty() {
|
||||
return None;
|
||||
}
|
||||
let (x_min, x_max) = page_x_bounds(items, images)?;
|
||||
detected_split
|
||||
.and_then(|split_x| paired_column_images(items, images, split_x, x_min, x_max))
|
||||
.or_else(|| local_flow_below_full_width_image(items, images, x_min, x_max))
|
||||
}
|
||||
|
||||
/// Partition a page into the topological order `above -> left -> right -> below`.
|
||||
/// These edges encode the reading-order DAG; empty nodes are omitted.
|
||||
pub(crate) fn build_region_graph(items: Vec<TextItem>, band: ColumnFlowBand) -> Vec<RegionNode> {
|
||||
let mut above = Vec::new();
|
||||
let mut left = Vec::new();
|
||||
let mut right = Vec::new();
|
||||
let mut below = Vec::new();
|
||||
for item in items {
|
||||
if item.y > band.y_top {
|
||||
above.push(item);
|
||||
} else if item.y < band.y_bottom {
|
||||
below.push(item);
|
||||
} else if item.x + effective_width(&item) / 2.0 < band.split_x {
|
||||
left.push(item);
|
||||
} else {
|
||||
right.push(item);
|
||||
}
|
||||
}
|
||||
let rtl = is_rtl_text(left.iter().chain(right.iter()).map(|item| &item.text));
|
||||
let mut ordered = vec![(RegionKind::FullWidth, above)];
|
||||
if rtl {
|
||||
ordered.push((RegionKind::Column, right));
|
||||
ordered.push((RegionKind::Column, left));
|
||||
} else {
|
||||
ordered.push((RegionKind::Column, left));
|
||||
ordered.push((RegionKind::Column, right));
|
||||
}
|
||||
ordered.push((RegionKind::FullWidth, below));
|
||||
ordered
|
||||
.into_iter()
|
||||
.filter_map(|(kind, items)| (!items.is_empty()).then_some(RegionNode { kind, items }))
|
||||
.collect()
|
||||
}
|
||||
|
||||
#[cfg(test)]
|
||||
mod tests {
|
||||
use super::*;
|
||||
use crate::types::ItemType;
|
||||
|
||||
fn item(text: &str, x: f32, y: f32, width: f32) -> TextItem {
|
||||
TextItem {
|
||||
text: text.into(),
|
||||
x,
|
||||
y,
|
||||
width,
|
||||
height: 11.0,
|
||||
font: "F1".into(),
|
||||
font_size: 11.0,
|
||||
page: 1,
|
||||
is_bold: false,
|
||||
is_italic: false,
|
||||
is_underline: false,
|
||||
is_strikeout: false,
|
||||
item_type: ItemType::Text,
|
||||
mcid: None,
|
||||
}
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn full_width_image_anchors_local_two_column_flow() {
|
||||
let mut items = vec![
|
||||
item("A full width caption", 55.0, 230.0, 430.0),
|
||||
item("A trailing full width heading", 55.0, 80.0, 430.0),
|
||||
];
|
||||
for index in 0..5 {
|
||||
let y = 170.0 - index as f32 * 14.0;
|
||||
items.push(item("left column prose words", 55.0, y, 210.0));
|
||||
items.push(item("right column prose words", 280.0, y, 210.0));
|
||||
}
|
||||
let images = vec![(55.0, 250.0, 490.0, 680.0)];
|
||||
let band = infer_image_anchored_flow(&items, &images, None).unwrap();
|
||||
assert!((band.split_x - 272.5).abs() < 2.0);
|
||||
let graph = build_region_graph(items, band);
|
||||
assert_eq!(graph.len(), 4);
|
||||
assert_eq!(graph[0].kind, RegionKind::FullWidth);
|
||||
assert_eq!(graph[1].kind, RegionKind::Column);
|
||||
assert_eq!(graph[2].kind, RegionKind::Column);
|
||||
assert_eq!(graph[3].kind, RegionKind::FullWidth);
|
||||
assert_eq!(graph[3].items[0].text, "A trailing full width heading");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn full_width_image_anchors_cjk_column_flow() {
|
||||
let mut items = Vec::new();
|
||||
for index in 0..5 {
|
||||
let y = 170.0 - index as f32 * 14.0;
|
||||
items.push(item("左栏这是没有空格的正文内容", 55.0, y, 210.0));
|
||||
items.push(item("右栏这是没有空格的正文内容", 280.0, y, 210.0));
|
||||
}
|
||||
let images = vec![(55.0, 250.0, 490.0, 680.0)];
|
||||
|
||||
assert!(infer_image_anchored_flow(&items, &images, None).is_some());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn paired_images_anchor_unbalanced_column_flows() {
|
||||
let mut items = vec![
|
||||
item("running header", 55.0, 700.0, 430.0),
|
||||
item("trailing full width caption", 55.0, 300.0, 430.0),
|
||||
];
|
||||
for index in 0..5 {
|
||||
items.push(item(
|
||||
"left prose words",
|
||||
55.0,
|
||||
500.0 - index as f32 * 14.0,
|
||||
200.0,
|
||||
));
|
||||
items.push(item(
|
||||
"right prose words",
|
||||
280.0,
|
||||
520.0 - index as f32 * 14.0,
|
||||
200.0,
|
||||
));
|
||||
}
|
||||
for index in 5..12 {
|
||||
items.push(item(
|
||||
"right continuation prose words",
|
||||
280.0,
|
||||
520.0 - index as f32 * 14.0,
|
||||
200.0,
|
||||
));
|
||||
}
|
||||
let images = vec![
|
||||
(55.0, 530.0, 255.0, 680.0),
|
||||
(55.0, 380.0, 255.0, 530.0),
|
||||
(280.0, 560.0, 490.0, 680.0),
|
||||
];
|
||||
let band = infer_image_anchored_flow(&items, &images, Some(270.0)).unwrap();
|
||||
let graph = build_region_graph(items, band);
|
||||
assert_eq!(graph[0].kind, RegionKind::FullWidth);
|
||||
assert_eq!(graph[1].kind, RegionKind::Column);
|
||||
assert_eq!(graph[2].kind, RegionKind::Column);
|
||||
assert_eq!(graph[3].kind, RegionKind::FullWidth);
|
||||
assert_eq!(graph[3].items[0].text, "trailing full width caption");
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn rtl_region_graph_reads_right_column_first() {
|
||||
let items = vec![
|
||||
item("A long English report header", 55.0, 250.0, 430.0),
|
||||
item("نص العمود الأيسر", 55.0, 150.0, 180.0),
|
||||
item("نص العمود الأيمن", 300.0, 150.0, 180.0),
|
||||
];
|
||||
let graph = build_region_graph(
|
||||
items,
|
||||
ColumnFlowBand {
|
||||
split_x: 270.0,
|
||||
y_bottom: 100.0,
|
||||
y_top: 200.0,
|
||||
},
|
||||
);
|
||||
|
||||
assert_eq!(graph.len(), 3);
|
||||
assert_eq!(graph[0].kind, RegionKind::FullWidth);
|
||||
assert!(graph[1].items[0].x > graph[2].items[0].x);
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn paired_header_logos_do_not_anchor_page_columns() {
|
||||
let mut items = Vec::new();
|
||||
for index in 0..7 {
|
||||
items.push(item(
|
||||
"left prose words",
|
||||
55.0,
|
||||
700.0 - index as f32 * 14.0,
|
||||
200.0,
|
||||
));
|
||||
}
|
||||
for index in 0..30 {
|
||||
items.push(item(
|
||||
"right prose words",
|
||||
280.0,
|
||||
700.0 - index as f32 * 14.0,
|
||||
200.0,
|
||||
));
|
||||
}
|
||||
let images = vec![
|
||||
(55.0, 720.0, 205.0, 770.0),
|
||||
(60.0, 718.0, 210.0, 768.0),
|
||||
(280.0, 720.0, 450.0, 770.0),
|
||||
];
|
||||
assert!(infer_image_anchored_flow(&items, &images, Some(270.0)).is_none());
|
||||
}
|
||||
|
||||
#[test]
|
||||
fn wide_banner_does_not_anchor_local_columns() {
|
||||
let mut items = Vec::new();
|
||||
for index in 0..7 {
|
||||
let y = 270.0 - index as f32 * 14.0;
|
||||
items.push(item("left column prose words", 55.0, y, 210.0));
|
||||
items.push(item("right column prose words", 280.0, y, 210.0));
|
||||
}
|
||||
let images = vec![(55.0, 310.0, 490.0, 550.0)];
|
||||
assert!(infer_image_anchored_flow(&items, &images, None).is_none());
|
||||
}
|
||||
}
|
||||
+11
-2
@@ -1012,12 +1012,19 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
|
||||
|
||||
// Separate images and links from text items
|
||||
let mut images: Vec<TextItem> = Vec::new();
|
||||
let mut page_image_regions: HashMap<u32, Vec<(f32, f32, f32, f32)>> = HashMap::new();
|
||||
let mut links: Vec<TextItem> = Vec::new();
|
||||
let mut text_items: Vec<TextItem> = Vec::new();
|
||||
|
||||
for item in items {
|
||||
match &item.item_type {
|
||||
ItemType::Image => {
|
||||
page_image_regions.entry(item.page).or_default().push((
|
||||
item.x,
|
||||
item.y,
|
||||
item.x + item.width,
|
||||
item.y + item.height,
|
||||
));
|
||||
if options.include_images {
|
||||
images.push(item);
|
||||
}
|
||||
@@ -1655,11 +1662,12 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
|
||||
// items from different side-by-side zones (e.g. left/right month columns
|
||||
// in a calendar) don't merge into the same line.
|
||||
let lines = if page_band_splits.is_empty() && page_chart_prose_splits.is_empty() {
|
||||
crate::extractor::group_into_lines_with_thresholds_and_charts(
|
||||
crate::extractor::group_into_lines_with_thresholds_and_regions(
|
||||
non_table_items,
|
||||
page_thresholds,
|
||||
&table_page_set,
|
||||
&page_chart_map,
|
||||
&page_image_regions,
|
||||
)
|
||||
} else {
|
||||
// Separate items into physical-band pages, chart/prose pages, and
|
||||
@@ -1682,11 +1690,12 @@ pub(crate) fn to_markdown_from_items_with_rects_and_lines(
|
||||
}
|
||||
}
|
||||
// Process unsplit pages normally
|
||||
let mut all_lines = crate::extractor::group_into_lines_with_thresholds_and_charts(
|
||||
let mut all_lines = crate::extractor::group_into_lines_with_thresholds_and_regions(
|
||||
unsplit_items,
|
||||
page_thresholds,
|
||||
&table_page_set,
|
||||
&page_chart_map,
|
||||
&page_image_regions,
|
||||
);
|
||||
// Process each split page's bands independently, then interleave
|
||||
// by Y position so paired zones (e.g. left/right months) appear together.
|
||||
|
||||
Reference in New Issue
Block a user