Ersetzt die private LAN-IP des Jetson durch jetson.local in Doku,
Shell-Skripten, .env.example und Docstrings. Funktional betroffen ist
nur ein os.environ.setdefault-Fallback im Spike; gesetzte Env-Vars
(OLLAMA_BASE_URL, GROBID_URL) haben weiterhin Vorrang. uv.lock bleibt
unveraendert (enthaelt Versionsnummern, keine Adressen).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Resolved 4 conflicts: cli.py (R-D citing-coverage warning + main's small-corpus JSON/hub-title changes both kept); ingest.py (C-7 paper.id=caller-id pin first, then DQ-2 abstract recovery); test_ingest.py + test_openalex.py (kept both branches' tests).
C-7/DQ-5 overlap (both canonicalise papers.id): kept both layers — openalex._canonical_id normalizes fetch_paper's id (DQ-5); ingest pins papers.id to the caller's bare id (C-7); they converge on the bare canonical id. Fixed the auto-merged test_fetch_paper_success that had contradictory bare/URL assertions. 402 tests pass; ruff + mypy clean.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
GROBID is only needed during reference extraction (ra_grobid_backfill.py),
not for normal corpus operations (search, embeddings, wiki). Idle it still
holds ~3.5GB on the 8GB Jetson, so keep it stopped by default.
- infra/docker-compose.yml: grobid restart unless-stopped -> "no" (no auto-start
on boot; db keeps unless-stopped). Start/stop papers-grobid around an R-A run.
- docs/infra/grobid-jetson.md: add "Run on-demand" section with the start/stop
workflow, and record the refextract head-to-head evaluation on the Jetson that
led to keeping GROBID: far worse DOI recall on this math corpus (lutz thesis
101 usable IDs vs refextract 2) and slower per PDF despite a smaller footprint.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Enable roadmap R-A (GROBID reference extraction) end-to-end and run it
against local Jetson infra instead of a Mac/Rosetta emulation.
Backfill:
- Add scripts/ra_grobid_backfill.py: references-only, idempotent citation
backfill (dry-run default). Deliberately not ingest_paper(source_path=pdf),
which re-OCRs and replaces the DQ-3-clean .txt chunks; this touches only the
citations table (ON CONFLICT DO NOTHING). 7 tests.
- Make extract_references timeout configurable (large theses exceed the 60s
default, especially against a slower GROBID).
Jetson infra:
- Bump grobid/grobid 0.8.2 -> 0.9.0-crf. The -crf tag is the only GROBID
variant published as an arm64 multi-arch manifest; every 0.8.x tag and the
full deep-learning image are amd64-only, which is why GROBID never ran on the
aarch64 Jetson. Enable the 4g memory cap + init/ulimits per GROBID docs.
- Add docs/infra/grobid-jetson.md runbook: arm64 image rationale, the two host
prerequisites (docker-group membership + the Compose v2 CLI plugin, both
missing on the Jetson), service-scoped deploy, and the GET /api/isalive check.
Verified live 2026-06-17: native arm64 pull, isalive=true, end-to-end
extract_references on lutz-2024-thesis.pdf = 116 refs (101 with DOI/arXiv).
docs(audit): mark R-A core done (citing coverage 27/29 -> 29/29; edges 920 -> 1022).
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Now that 'codex migrate' applies the idempotent schema via a privileged role,
recommend it (with MIGRATION_DATABASE_URL) as the primary schema-sync fix in the
preflight abort message, keeping the manual owner-ALTER as a fallback.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
M-1: schema.sql could not be re-applied (leading CREATE TABLE/INDEX lacked
IF NOT EXISTS), apply_schema was never called, and the live migration just
failed on a missing chunks.section column. Fixes:
- schema.sql: all CREATE TABLE/INDEX now use IF NOT EXISTS — the whole file is
re-applyable as a no-op.
- codex migrate: new CLI command that applies schema.sql. Connects via
MIGRATION_DATABASE_URL (falls back to DATABASE_URL) and catches
InsufficientPrivilege with guidance — because the app role is DML-only and
core tables are owned by 'postgres' (the privilege dimension found during the
live migration incident).
- config: migration_database_url (optional privileged connection).
- db: apply_schema docstring corrected (idempotency now true) + privilege note.
- T-2: static test that schema.sql is fully idempotent; migrate privilege-error
test.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The helper TRUNCATEd papers CASCADE and only then ran ingest_all.sh, which
failed on the missing chunks.section column (audit M-1) — leaving the corpus
wiped. Add a preflight that verifies chunks.section exists BEFORE the
destructive step and aborts early (no TRUNCATE) with the owner-level ALTER to
run, since the app user cannot alter postgres-owned tables.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
One-time, guarded migration for the C-7 id-format change: a plain re-ingest
would duplicate every paper (new bare-id row beside the old URL-id PK), so this
wipes (TRUNCATE papers CASCADE) and rebuilds via ingest_all.sh.
Safety: requires the SSH tunnel, prints a BEFORE snapshot, gates the TRUNCATE
behind an explicit 'MIGRATE' confirmation, then prints AFTER verification —
C-7 (url_form_ids should be 0) and C-1 via the real resolver-based
discovery_leads() (ingested papers leaked should be 0). Uses .venv psycopg
(psql is not installed); DATABASE_URL is sourced, never echoed.
Joins PR #13 (Wave 2). Read-only verification SQL validated against the live DB.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Fixes preprint/journal duplicate problem: arXiv papers were ingested
under their preprint OpenAlex ID, but corpus citations reference the
published-version OpenAlex ID → those citations were invisible in the
graph (appeared as dangling).
Changes:
- infra/schema.sql: CREATE TABLE paper_identifiers (paper_id, openalex_id)
with UNIQUE index; seeded from papers.openalex_id on migration
- codex/graph.py: build_citation_graph LEFT JOINs paper_identifiers as
a second resolution path: COALESCE(p.id, pi.paper_id, c.cited_id)
- codex/ingest.py: every ingest inserts openalex_id into paper_identifiers
(ON CONFLICT DO NOTHING) — aliases can be added manually alongside
Live DB: 35 rows seeded + 4 published-version aliases added:
math/0503219 → W2163787581 (DCG 2007)
math/0306167 → W2136126748 (Commun Analysis Geom 2004)
math/0203250 → W1567166970 (Trans AMS 2003)
1005.2698 → W1511400044 (Geom & Topol 2015)
Effect: dangling 574→570; Bobenko+Springborn 2007 now IN-KB rank 9.
Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>