Tagged PDFs carry a structure tree with real heading roles (H1..H6), and
the core already parses it (structure_tree::StructTree) and threads MCIDs
onto TextItem — but neither surfaced through the bindings.
- Expose TextItem.mcid (Option<i64>) through the napi and pyo3 bindings,
matching the core field added with the marked-content extractor.
- Add StructRole::name(), the inverse of from_name, so roles have a
stable string form.
- Add extract_structure_elements / extract_structure_elements_mem to the
core: one (page, mcid, role) entry per marked-content reference, sorted
by (page, mcid), empty for untagged PDFs. Pages are 1-indexed to match
TextItem.page, so results join directly against
extract_text_with_positions output.
- Bind it as extractStructureElements (napi) and
extract_structure_elements / extract_structure_elements_bytes (pyo3),
with type-stub updates in pdf_inspector.pyi.
- Cover the join in Rust integration tests, napi test.mjs, and pytest,
using the existing firecrawl_docs_tagged.pdf fixture (tagged) and
thermo-freon12.pdf (untagged).
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>