docs(audit): live-DB verification — confirm C-7 & C-1, refine M-1

Read-only queries vs the live Jetson corpus (papers=36, citations=811):
- C-7 CONFIRMED: 35/36 papers.id are DOI-URL form (https://doi.org/...),
  incl. arXiv papers as arXiv-DOI URLs. Bare-ID lookups fail for 35/36.
- C-1 CONFIRMED: discover.py mis-flags 13 ingested papers as dangling;
  graph.py correctly resolves 57 cross-cite edges → the two paths diverge.
- M-1 REFINED HIGH→MED: paper_identifiers exists (39 rows), so no live crash;
  structural migration-tooling gap remains.
- C-2 partially refuted: 39 alias rows > 35 → ~4 diverging aliases, not inert.

Adjusted tally: 7 HIGH, 12 MED, 26 LOW.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This commit is contained in:
Tarik Moussa
2026-06-15 10:40:46 +02:00
parent 2d1b0b5e23
commit 33fb91d478

View File

@@ -440,8 +440,9 @@ CLI/config are otherwise solid and well-documented.
# Executive Summary — whole-repo audit complete (loop-1 + Parts AG) # Executive Summary — whole-repo audit complete (loop-1 + Parts AG)
**Tally: 45 findings — 8 HIGH, 11 MED, 26 LOW/INFO.** Test suite 330 green, **Tally: 45 findings — 8 HIGH, 11 MED, 26 LOW/INFO** (after live-DB verification:
`ruff`/`mypy` clean. The codebase is well-structured and idiomatic; the defects **7 HIGH, 12 MED** — M-1 reclassified, see the verification section at the bottom).
Test suite 330 green, `ruff`/`mypy` clean. The codebase is well-structured and idiomatic; the defects
cluster into five root themes, not random scatter. cluster into five root themes, not random scatter.
### The 8 HIGH findings ### The 8 HIGH findings
@@ -486,3 +487,51 @@ Every `codex/` module was read. Not deep-reviewed: `nougat.py` (54-line `.pdf`
wrapper, skimmed), and the exact runtime DB state (URL-form ids in C-7, and the wrapper, skimmed), and the exact runtime DB state (URL-form ids in C-7, and the
`paper_identifiers` existence in M-1) is **inferred from code** and flagged `paper_identifiers` existence in M-1) is **inferred from code** and flagged
"confirm on the Jetson DB" — I could not query the live database from here. "confirm on the Jetson DB" — I could not query the live database from here.
---
# Live DB Verification — 2026-06-15 (read-only, via existing SSH tunnel)
Ran read-only queries against the live Jetson corpus (`papers=36`,
`citations=811`). This **confirms the two identity HIGHs with hard numbers** and
**corrects an over-statement in M-1**.
### C-7 — CONFIRMED ✅ (severity stands: HIGH)
`papers.id` form distribution: **`doi_url=35`, `bare_arxiv=1`, everything else 0**
(of 36). Samples: `https://doi.org/10.48550/arxiv.math/0603097`,
`https://doi.org/10.48550/arxiv.2310.17529`. Even arXiv papers are stored as their
arXiv-DOI **URL**, never the bare `math/0603097` / `2301.x` the schema, `models.Paper`,
and `ingest_all.sh` assume. The single bare id is the one paper OpenAlex 404'd
(fell to `ingest.py:94`). → bare-ID lookups fail for 35/36 papers. Confirmed.
### C-1 — CONFIRMED ✅ with impact (severity stands: HIGH)
- `citations.cited_id` is uniformly **`oa_url=811`** (all OpenAlex URLs) — the live
corpus was ingested via the OpenAlex path, so the heterogeneous bare-DOI/S2 mix
theorised in Part B is **not present here** (it would appear only via S2/GROBID,
i.e. `.pdf`/fallback ingest).
- **`discover.py` wrongly flags 13 already-ingested papers as "dangling / not
ingested"** (their `cited_id` is an OpenAlex URL whose paper *is* in the corpus
via `openalex_id`). Meanwhile **`graph.py` correctly resolves 57 cross-citation
edges** via the `openalex_id` JOIN. The two "dangling" implementations
measurably disagree — exactly the C-1 divergence, now quantified.
### M-1 — REFINED ⤵ (downgrade acute severity; structural finding stands)
`paper_identifiers` **exists on the live DB with 39 rows** — the catastrophic
"graph JOIN crashes with relation-does-not-exist" scenario **does not occur**;
the table was applied (manually or via DB re-create). The *structural* gap stands:
`apply_schema` is still non-idempotent and never invoked, so the **next** schema
change has no migration path — but the acute "live graph is broken" framing was
**over-stated**. Re-rank M-1 from HIGH → **MED** (latent migration debt, not an
active outage).
### C-2 — PARTIALLY REFUTED ⤵
`paper_identifiers` has **39 rows vs ~35 papers-with-openalex_id**, so ~4 papers
carry a diverging alias → the `pi` JOIN is **not** fully inert (it resolves a
handful of cases the `p` JOIN can't). Keep as LOW/INFO: the table earns its keep
for a few papers, though the cost/benefit at this corpus size is still marginal.
### Net adjustment to the tally
HIGH **8 → 7** (M-1 → MED). C-7 and C-1 are now **confirmed**, not inferred — they
should lead remediation. The dominant root theme (**un-normalized identity**) is
empirically validated: 35/36 ids in the "wrong" form, 13 papers mis-classified by
the older code path, 57 edges rescued by the newer one.