feat(F-09): rich parsing — formula + figure extraction

- codex/parsing/mathpix.py: pix2tex (local, CPU) primary + MathPix API
  optional; bbox heuristic h>15px, math-char-count>5; singleton model cache
- codex/parsing/figures.py: pymupdf embedded-image extraction → PNG;
  caption detection via proximity + "Figure/Fig./Abbildung" prefix
- codex/models.py: FormulaChunk + FigureChunk dataclasses (R-10/R-11)
- codex/ingest.py: --rich flag wires formula+figure extraction post-ingest
- codex/cli.py: search_app sub-typer (paper + formula subcommands),
  --rich flag on ingest; wiki_app from F-12 preserved intact
- codex/config.py: mathpix_app_id/key, pix2tex_fallback, figures_dir
- infra/schema.sql: formulas + figures tables with HNSW pgvector indexes
- pyproject.toml: pymupdf>=1.24, pix2tex>=0.1.4
- tests/parsing/test_mathpix.py + test_figures.py: 31 tests (mock pix2tex
  + MathPix HTTP, real pymupdf on synthetic PDF)

Gate: 158 passed, ruff clean, mypy clean (20 files)
Requirements: R-10 R-11 R-12 R-13 R-14 → done

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
This commit is contained in:
Tarik Moussa
2026-06-14 01:16:24 +02:00
parent d2f9141a5c
commit 1df9be6563
13 changed files with 1930 additions and 26 deletions

View File

@@ -78,3 +78,39 @@ class CodeLink:
note: str | None = None
id: int | None = None
added_at: datetime | None = None
@dataclass
class FormulaChunk:
"""Maps to the ``formulas`` table (F-09 Rich Parsing).
Stores a single extracted mathematical formula (LaTeX) from a PDF page.
``id`` is set by the database (BIGSERIAL).
``embedding`` is reserved for future pgvector similarity search.
"""
paper_id: str
page: int
raw_latex: str
context: str
id: int | None = None
eq_label: str | None = None
embedding: list[float] | None = None
@dataclass
class FigureChunk:
"""Maps to the ``figures`` table (F-09 Rich Parsing).
Stores metadata for a single extracted figure from a PDF page.
``image_path`` points to the saved PNG on disk.
``id`` is set by the database (BIGSERIAL).
``embedding`` is reserved for future pgvector similarity search.
"""
paper_id: str
page: int
image_path: str
caption: str
id: int | None = None
embedding: list[float] | None = None