# AUDIT — Full Repository (iterative) **Auditor:** Audit-Loop (Opus) **Started:** 2026-06-15 **Scope:** entire `main` tree, audited iteratively module-by-module. **Companion doc:** [`AUDIT-2026-06-15-loop-1.md`](AUDIT-2026-06-15-loop-1.md) — the F-15 / dev-loop-1 findings (S-1…U-4). This document continues the **same finding-ID register** (so `C-4` here follows `C-3` there) and covers the rest of the repo. > Severity: **HIGH** fix before relying · **MED** fix soon · **LOW** polish · > **INFO** accepted/no-action. Findings only — remediation planned separately. ## Coverage tracker | Part | Area | Status | |------|------|--------| | (loop-1) | F-15 graph, db/schema, ingest, discover, mcp, provenance-scan, cli(search+graph), config(graph) | ✅ documented | | **A** | Grounding guard (`wiki.py` `_parse_claims`/`_run_grounding_guard`, quarantine) — shared by `synthesis.py` | ✅ documented | | **B** | `sources/` (arxiv, openalex, semanticscholar) | ✅ documented | | **C** | `parsing/` (grobid, nougat, mathpix, figures, tex) | ✅ documented | | D | `synthesis.py` orchestration (find_connections/gaps/improvements, conjectures, write_leads) | ⏳ next | | E | `wiki.py` rest (retrieve, compile_concept, cross-refs, index/log, embed) | ☐ | | F | `embed.py`, `quality.py`, `models.py` | ☐ | | G | `cli.py` rest (ingest/wiki/synthesis/provenance commands), `config.py` rest | ☐ | --- ## Part A — Grounding guard (the "zero-fact-leak" core) The guard lives in `wiki.py` (`_parse_claims:216`, `_run_grounding_guard:299`) and is imported by `synthesis.py:57-59`, so **one implementation gates both** the wiki pages and the synthesis leads. A claim is "grounded" iff the **last sentence** before its `[BibKey #loc]` citation contains a run of 5 consecutive non-stopword tokens that appears as a substring of any chunk of the same bibkey. ### C-4 — Zero-claim pages bypass quarantine and publish as "compiled" · **HIGH** - **Evidence:** `wiki.py:730` — `if total_claims > 0 and grounding_rate < threshold:` quarantines; the `else` (`:737-742`) **publishes to `wiki/` and records the hash** as successfully compiled. - **Bug:** when the LLM output has **no parseable citations** (`total_claims == 0`, so `grounding_rate == 0.0`), the `total_claims > 0` guard is false → it takes the *publish* branch. The quarantine gate is skipped exactly when grounding evidence is **weakest** (zero citations). - **Trigger paths:** (a) the LLM emits uncited prose; (b) **LLM outage** — the module header guarantees "LLM unavailable → empty ConceptPage, no crash" (`wiki.py:13-14`), so a transient Ollama outage writes near-empty stub pages to `wiki/`, marks them compiled, and stores their hash. On the next `changed_only` run the hash matches → the stub is **skipped forever** until a forced recompile. A research wiki silently fills with empty/uncited pages. ### C-5 — "Every claim is grounded" holds only for the last cited sentence · **MED** - **Evidence:** `_CLAIM_RE` (`wiki.py:211-213`) captures `text = [^\[]+?` (the span back to the previous `]`/start); `_last_sentence` (`:279-291`) then keeps only the final sentence; `_run_grounding_guard` checks that sentence only. - **Two holes in the guarantee:** 1. **Paragraph passengers** — any sentence that is *not* the last one before a citation is never grounding-checked and never `⚠`-marked. "Hallucination one. Hallucination two. Grounded tail [Bib]." passes with only the tail verified. 2. **Uncited text is invisible** — a sentence with no `[BibKey]` is not a Claim at all (`_parse_claims` only yields cited spans), so it is never grounded nor marked. The LLM can emit fully uncited assertions into a published page. - **Impact:** the effective invariant is "*the last sentence immediately before each citation has lexical overlap with the cited paper*" — materially weaker than the N-02 "zero-fact-leak enforcement" framing. The module **header is honest** ("Substring-Match MVP", `wiki.py:4,13`); the commit/ADR language is not. Recommend aligning the claim or strengthening the guard (check every sentence; flag uncited sentences). ### C-6 — Substring match does not enforce the documented "5 consecutive tokens" · **LOW** - **Evidence:** `wiki.py:347-349` — `gram = " ".join(content[i:i+5])`; `if any(gram in src for src in sources)` where `src` is the space-joined content-word string of the chunk. - **Issue:** a *substring* test on the joined string lets the first/last gram token match as a **prefix/suffix of a longer chunk word** (e.g. gram `"cat dog …"` substring-matches inside `"scat dogma …"`). The docstring claims matching "as 5 consecutive non-stopword **tokens**" — true sequence matching would compare token lists with a sliding window (or add boundary sentinels). Low real-world impact at 5-gram length, but the claimed invariant is not what the code enforces. ### R-1 — Tokenizer brittleness causes false-negative over-quarantine · **MED** - **Evidence:** `_content_words` (`wiki.py:294-296`) lowercases and `.split()`s on whitespace only — **no punctuation stripping**. `_last_sentence` splits on `.!?` (`:276`). - **Impact:** punctuation stays glued to tokens, so a faithful paraphrase grounds only on *exact* surface form — `"volume,"` ≠ `"volume"`, `"L(gamma1)"` must match character-for-character, and a trailing decimal like `"V = 1.5"` makes `_last_sentence` return the fragment `"5"` (split at the `.`), which then has `< 5` content words → **forced ungrounded** (`:339-344`). Math-heavy and terse claims are systematically marked ungrounded even when verbatim, inflating the ungrounded ratio and pushing real pages into quarantine (and into C-4's blind spot when it drives `total_claims` interactions). The passing test `test_grounding_guard_…` only succeeds because its fixture reproduces the chunk text *exactly*. ### R-2 — Reference-list filter can drop theorem-dense chunks · **LOW** - **Evidence:** `_retrieve_chunks` (`wiki.py:188-199`) skips a chunk when `>60 %` of its lines match `^\s*(\[\d+\]|[A-Z][a-z]+,?\s+[A-Z]\.)`. - **Issue:** the second alternative (`[A-Z][a-z]+,?\s+[A-Z]\.`) also matches prose like `"Theorem A."`, `"Lemma B."`, `"Section C."`. A chunk formatted as a list of theorems/lemmas can exceed 60 % and be discarded as a "reference list" — silently dropping exactly the mathematical content the wiki wants. ### R-3 — Conflict detection has near-zero precision · **LOW** - **Evidence:** `_detect_conflicts` (`wiki.py:367-…`) flags a pair `(bk1, bk2)` if **each** has *any* chunk containing a hedging keyword (`but|however|in contrast|…`, `_CONFLICT_KEYWORDS:361-364`). - **Issue:** it never checks the keywords relate to a shared topic or to each other. In a corpus where most papers contain "however"/"but" somewhere, this flags ~O(n²) pairs as "conflicting" — noise surfaced in `CompileReport`. Labeled "heuristic signal only", but as-is it is not actionable. **Part A net:** 1 HIGH (C-4), 2 MED (C-5, R-1), 3 LOW (C-6, R-2, R-3). The guard is a reasonable lexical MVP, but (a) it has a concrete publish-bypass bug for zero-citation output, and (b) the "zero-fact-leak" framing overstates what a last-sentence substring check can guarantee. --- ## Part B — `sources/` (external API clients) Solid foundations: all three clients set httpx timeouts; OpenAlex and S2 retry 429/5xx with exponential back-off (tenacity); S2 enforces a ≥1 req/s floor with a monotonic clock under a lock. The findings are about **ID format consistency** and a couple of dead/inert paths. ### C-7 — `papers.id` is stored in OpenAlex *URL* form, breaking the bare-ID contract · **HIGH** - **Evidence:** `openalex.py:94` — `_map_paper` sets `id = data.get("doi") or data.get("id") or ""`. Real OpenAlex returns `doi` as a **full URL** (`https://doi.org/10.x`) and `id` as `https://openalex.org/W…` (the repo's own test confirms the URL form for `openalex_id`, `test_openalex.py:92`). `ingest.py:82` takes `paper.id` from this unchanged and upserts it as `papers.id` (`ingest.py:138`). - **Why it's currently masked:** the test fixture `_SAMPLE_WORK["doi"]` is a **bare** DOI `"10.1145/3592430"` (`test_openalex.py:18`) — not the URL form real OpenAlex emits — and **no test asserts `paper.id`** at all. So the suite is green while production stores `papers.id = "https://doi.org/10.1145/3592430"`. - **Impact:** 1. **Bare-ID lookups fail.** `ingest_all.sh` ingests `"10.1145/1964921.1964997"`, but `papers.id` becomes the URL; `codex graph related 10.1145/…`, `discover citing 10.1145/…`, search-by-id, etc. find nothing. 2. **Second ID-mismatch axis (compounds C-1).** `cited_id` from the S2/GROBID paths is a **bare** DOI/arXiv (`semanticscholar.py:98-99`, `ingest.py:187`), so a bare-DOI reference can never equal a URL-form `papers.id` — bare-DOI cross-citations stay "dangling" regardless of the F-15 W-ID resolution (which only matches `papers.openalex_id`, the URL form). - **Action:** confirm with `SELECT id FROM papers LIMIT 5` on the Jetson DB; if URL-form, normalize `papers.id` to a canonical bare form at ingest (and the `cited_id` writers to match), then add a representative fixture + `paper.id` assertion (see T-3). ### T-3 — OpenAlex test fixture is unrepresentative and skips `paper.id` · **MED** - **Evidence:** `_SAMPLE_WORK["doi"] = "10.1145/3592430"` (bare) vs the real `https://doi.org/…` URL form; `test_fetch_paper_*` assert title/year/authors/ abstract/`openalex_id` but **never** `paper.id`. - **Impact:** the id-mapping — the exact surface of C-7 — is untested, and the fixture actively hides the URL-form behaviour. Any normalization fix needs this closed first or it will regress invisibly. ### R-5 — S2 `fetch_references` breaks on legacy arXiv IDs containing `/` · **LOW** - **Evidence:** `semanticscholar.py:80` interpolates `paper_id` into the path `…/paper/{paper_id}/references`. The ingest fallback passes `arXiv:math/0603097` (`ingest.py:91`) — the embedded `/` yields a malformed path → 404 → `fetch_references` silently returns `[]`. Legacy-arXiv papers that miss the OpenAlex path get **zero references** with no error. ### R-4 — `arxiv.fetch_source` no-op try/except + no retry · **LOW** - **Evidence:** `arxiv.py:41-44` — `try: httpx.get(...) except httpx.RequestError: raise` catches and immediately re-raises, doing nothing. Unlike OpenAlex/S2, the arXiv client has **no** retry/back-off, so a transient network blip aborts ingest. Minor resilience asymmetry. ### C-9 — `semanticscholar.fetch_recommendations` is unreachable app code · **LOW** - **Evidence:** referenced only by `tests/sources/test_semanticscholar.py`; no `codex/` module or CLI calls it (`grep fetch_recommendations`). Tested but never wired — dead feature surface (decide: wire a `discover recommend` command, or drop). **Part B net:** 1 HIGH (C-7, ID-form contract — the big one; confirm on live DB), 1 MED (T-3, fixture masks it), 3 LOW (R-4, R-5, C-9). The retry/timeout/rate-limit hygiene is genuinely good; the weakness is the **un-normalized, heterogeneous ID space** (URL vs bare, W-ID vs DOI vs arXiv vs S2-hash) threaded through `papers.id` and `citations.cited_id` — the same root cause behind C-1 and the F-15 ID dance. --- ## Part C — `parsing/` (tex, grobid, nougat, mathpix, figures) > **Scope reality:** the live corpus is ingested entirely as **`.txt`** > (`ingest_all.sh` maps every paper to a `*.txt` file), so the `.tex`-clean, > GROBID, Nougat, Mathpix, and Figures paths are **latent** — only > `tex.chunk_text` + `quality.*` run on the live data. The findings below are > real but mostly fire only once `.tex`/`.pdf --rich` ingest is used. > > **Not a finding:** the broad `except Exception` in `mathpix.py:109,151` and > `figures.py:125,152` are *justified* per-item graceful degradation > (`fitz.open` fails → return empty; pix2tex/image fails → skip that item), not > swallowed bugs. ### C-10 — Formulas/figures key on the PDF *filename*, not `paper.id` · **MED** (latent, `.pdf --rich`) - **Evidence:** `figures.py:119` and the mathpix extractor set `paper_id = Path(pdf_path).stem`. `ingest.py:289,303` insert `f.paper_id` / `fig.paper_id` **unchanged** into `formulas`/`figures`, whose schema declares `paper_id TEXT REFERENCES papers(id)` (`schema.sql:97,115`). - **Impact:** the filename stem (e.g. `springborn-2008-weighted-delaunay…`) will not equal `papers.id` (an OpenAlex URL/bare id — see C-7), so the FK insert **fails and aborts rich ingest**, or (without FK enforcement) writes orphaned rows. Identity is derived from the *file*, not the *paper*. Fix: pass `paper.id` into the extractors, or have ingest overwrite `f.paper_id = paper.id` before insert. ### R-7 — `_clean_latex` misses `\(…\)` / `\[…\]` math + strips all math from chunks · **LOW** (`.tex` path) - **Evidence:** `tex.py:21-56` strips `%comment`, `$…$`, `$$…$$`, `equation`/`align` envs, `\cite/\label/\ref` — but **not** the `\(…\)` / `\[…\]` math delimiters. Papers using those leave raw `\( … \)` markup in the chunk text. - **Impact:** (a) leftover math markup pollutes embeddings; (b) *all* math is removed from chunk text, so chunk-level semantic search can never match formula content — it relies entirely on the separate formula FTS (F-09), which is itself latent here. Acceptable as a separation-of-concerns choice, but worth stating. ### R-6 — PDF-path external calls are unwrapped / un-retried · **LOW** (`.pdf` path) - **Evidence:** `ingest.py:185` calls `extract_references()` outside any `try`; GROBID (`grobid.py:54,133`) and the mathpix HTTP calls `raise_for_status` with **no** retry/back-off (unlike OpenAlex/S2). A transient GROBID/Mathpix 5xx aborts the whole PDF ingest. - **INFO:** `grobid.py:56` uses stdlib `ET.fromstring` on the server's XML. Fine for a trusted self-hosted GROBID; switch to `defusedxml` if it is ever pointed at an untrusted endpoint. **Part C net:** 1 MED (C-10), 2 LOW (R-6, R-7), all **latent** for the current `.txt` corpus. `nougat.py` was skimmed only (54 lines, `.pdf`-only wrapper) — no deep review yet; flagged in the ledger. Parsing hygiene is otherwise good. > **Request (a) — second audit pass — is now complete:** §8 of the loop-1 doc > (grounding guard ✓ Part A, `sources/` ✓ Part B, `parsing/` ✓ Part C) is closed. > Remaining whole-repo parts D–G are the broader sweep.