33 KiB
DATA-QUALITY AUDIT — codex knowledge base (handoff for a fresh session)
Status: baseline scan done 2026-06-15. DQ-1 RESOLVED 2026-06-16 (S2 supplement
applied to live DB + wired into ingest). DQ-4 MEASURED 2026-06-16 — retrieval is
strong after fixing a P0 codex search paper crash. DQ-2 RESOLVED 2026-06-16
(1911.00966 fully recovered from arXiv; s00454 abstract from S2; book chapter is a
documented limit; recovery wired into ingest). DQ-3 RESOLVED 2026-06-16
(content fidelity clean — full coverage, only OCR junk dropped). All four
assessment items (DQ-1..DQ-4) DONE. DQ-5 RESOLVED 2026-06-17 (DOI id normalized
to bare canonical form in openalex._map_paper; live re-ingest of
10.1007/s00454-019-00132-8 confirmed idempotent → unblocks R-A/R-C).
Roadmap R-A..R-E added 2026-06-16 (free-source acquisition levers — see Roadmap section).
Author: Audit-Loop (Opus). Intended reader: a cold session with no memory
of the audit conversation — everything needed to continue is in this file.
Why this exists — the second audit axis
The code/pipeline audit (AUDIT-2026-06-15-loop-1.md, -full-repo.md) verified the
loader is correct: IDs are canonical, the citation graph is consistent, the
ingest pipeline does what it claims. It did not ask whether the information
in the database is actually good. A correct loader can faithfully load thin,
incomplete, or unrepresentative data. This document is that second axis — a
data-quality assessment of the live corpus, independent of code correctness.
How to reach the data (self-contained)
- Live DB: Jetson PostgreSQL via an SSH tunnel (open it first):
ssh -f -N -L 5433:localhost:5432 alfred@192.168.178.103 - Credentials:
.env.jetson-ingest(gitignored) holdsDATABASE_URL(postgresql://researcher:…@localhost:5433/papers). Theresearcherrole is DML-only (read freely; no DDL). - Run probes:
set -a; source .env.jetson-ingest; set +athenPYTHONPATH=. .venv/bin/pythonwithpsycopg+psycopg.rows.dict_row. Reusable domain helpers:codex.quality(is_quality_chunk,_bib_score,classify_section),codex.graph,codex.discover. - Corpus snapshot (2026-06-15, post canonical-id migration): 29 papers, 1507 chunks, 590 citations.
Baseline — what is already GOOD (measured, no action needed)
- Chunk content is clean. Re-running the F-16 quality gate over all stored
chunks: 0 / 1507 fail
is_quality_chunk(none too-short, none low-alpha / OCR-ish, none bibliography-like). The gate did its job on the.txtingest. - Chunk sizing healthy. chars: min 478 / median 2913 / max 4154; 0 exact duplicate chunk bodies.
sectioncolumn is low-signal (98 %body; 13bibliography, 9theorem, 2intro, 1proof). This confirms code-audit R-12 — word-window chunks rarely start at a section header. Not a data defect; the column is just weak.
Findings (open investigation items)
DQ-1 — Citation coverage: 12 / 29 papers (41 %) had ZERO citations · RESOLVED 2026-06-16
Resolution (what was done):
- Verified upstream, not a bug. Queried the OpenAlex API directly for all 11
zero-citation papers that have an
openalex_id: every one returnsreferenced_works_count = 0server-side, W-ids all match what we stored. They aretype: preprint(arXiv) /article/dissertation— works OpenAlex indexes without a parsed bibliography. Confirmed source-coverage limit. - Semantic Scholar supplement applied to the live DB (idempotent backfill,
ON CONFLICT (citing_id,cited_id) DO NOTHING): +330 citations (590 → 920). Zero-out-edge papers 12 → 2; in-corpus citing coverage 17/29 → 27/29; graph 489→704 nodes, 590→920 edges; discovery-lead targets 472→677. 9 of the new edges are in-corpus (e.g.2305.10988 → 10.1007/s00454-019-00132-8). - Wired into ingest so future re-ingests self-heal: when OpenAlex returns an
empty
referenced_works,ingest.pynow falls back to_s2_reference_supplement()(citing_id rewritten to canonicalpapers.id, DOI cited-ids lowercased). Same helper drove the one-off backfill. - Bug fixed (would have crashed the batch):
semanticscholar.fetch_referencesuseddata.get("data", []), but S2 returns HTTP 200{"data": null}for a known paper with no parsed refs (e.g.depositonce-5415) →TypeError. Changed todata.get("data") or []. Removed a now-redundant discarded S2 probe in the OpenAlex-404 arXiv branch. New regression test added. - Irreducible remainder (documented limit): the 2
depositoncetheses (10.14279/depositonce-20357,-5415) stay at zero — neither OpenAlex nor S2 indexes references for TU-Berlin institutional DOIs. Accept as inherent. - Files touched (uncommitted in working tree):
codex/sources/semanticscholar.py,codex/ingest.py,tests/ingest/test_ingest.py. Full suite: 331 passed.
Original finding (for context):
- Measured: 12 papers contribute no out-edges; the 590-edge citation graph comes from only ~17 papers. The whole F-15 layer (PageRank / coupling / co-citation / discovery leads) therefore runs on ~59 % of the corpus.
- Sharpening — it is mostly a source-coverage limit, not an ingest bug:
11 of the 12 zero-citation papers do have an
openalex_id, i.e. ingest calledopenalex.fetch_citations(openalex_id)and OpenAlex returned an emptyreferenced_workslist. Only1911.00966has noopenalex_id(OpenAlex 404 → arXiv/S2 fallback, which also yielded nothing). - Zero-citation ids:
1005.2698,1505.01341,1911.00966(no oa),2206.13461,2305.10988,2310.17529,2601.22903,math/0001176,math/0503219,math/0603097,10.14279/depositonce-20357,10.14279/depositonce-5415. (Pattern: arXiv preprints + the twodepositoncetheses — works OpenAlex indexes without parsed references.) - To investigate next:
- Confirm it is OpenAlex coverage, not a
fetch_citationsbug: for 2–3 of the ids, hitGET https://api.openalex.org/works/<openalex_id>and check whetherreferenced_worksis genuinely empty server-side. - If genuinely empty: enrich via the Semantic Scholar references path
(
semanticscholar.fetch_references("arXiv:<id>")) as a supplement for OpenAlex-empty papers — S2 often has references where OpenAlex doesn't. Note S2 cited_ids are bare DOI/arXiv, so they flow through the existingRESOLVED_CITATIONS_SQLresolver (audit C-1) fine. - Decide whether the F-15
graph_min_corpus_sizewarning should also flag low citing-paper coverage, not just paper count.
- Confirm it is OpenAlex coverage, not a
- Acceptance: either (a) citation coverage materially improves after an S2 supplement, or (b) documented as an inherent OpenAlex-coverage limit with the graph caveated accordingly.
DQ-2 — Metadata gaps: 3 no-abstract, 1 fully metadata-less · RESOLVED 2026-06-16
Resolution (what was done):
1911.00966fully recovered from the arXiv Atom API (OpenAlex 404'd on it): title "A discrete version of Liouville's theorem on conformal maps", Pinkall & Springborn, 2019, 384-char abstract → bibkeyPinkallSpringborn2019auto-generated, abstract embedded (non-zero). It was the worst node in the corpus — now@cite-able, in wiki grounding, and its 10 chunks are visible to chunk search (previously filtered out bybibkey IS NOT NULL). Citations were already added in DQ-1 (27 S2 refs).10.1007/s00454-019-00132-8abstract recovered from S2 (1186 chars) and re-embedded (targetedUPDATE, not full re-ingest — see DQ-5).10.1007/978-3-642-17413-1_7(book chapter): documented inherent limit. No abstract in OpenAlex, S2, or Crossref (verified). Left as the single remaining zero-vector paper; user may add an abstract by hand.- Corpus now: 1 paper without abstract (the book chapter), down from 3+1. Zero-vector search pollution (DQ-4 secondary finding) reduced from 2 tail fillers to 1.
- Wired into ingest so this self-heals: OpenAlex-404 on an arXiv id now
recovers metadata via
arxiv.fetch_metadata(new); an empty abstract is supplemented via_recover_abstract→ S2 then Crossref (newcodex/sources/crossref.py). Newsemanticscholar.fetch_abstract. - Files touched:
codex/sources/arxiv.py,codex/sources/crossref.py(new),codex/sources/semanticscholar.py,codex/ingest.py, + tests. Full suite: 342.
Original finding (for context):
- Measured:
- No abstract (3):
10.1007/978-3-642-17413-1_7,10.1007/s00454-019-00132-8,1911.00966. These get a zero-vectorabstract_emb, so paper-level semantic search can't place them. - No bibkey / year (1):
1911.00966only — and its authors array is empty too. This paper is fully degraded (OpenAlex 404 →Paper(id, title="")fallback): no abstract, no bibkey, no year, empty authors, no citations. With no bibkey it cannot be@cited and is invisible to wiki grounding (which keys on bibkey).
- No abstract (3):
- To investigate next:
1911.00966— confirm the arXiv id is correct and whether OpenAlex/S2 has it under a different id (DOI?); if recoverable, re-ingest to populate metadata. If not, decide: keep as a degraded node or drop.- The 2 no-abstract DOIs — check if OpenAlex has an
abstract_inverted_indexthat the mapper missed, or if the abstract is genuinely absent upstream.
- Acceptance: every paper has at least a bibkey + non-zero abstract embedding, or the exceptions are documented with rationale.
DQ-3 — Content fidelity (chunks vs source .txt): VERIFIED CLEAN · RESOLVED 2026-06-16
Resolution (read-only; re-ran chunk_text + is_quality_chunk on all 29
source files and compared to stored chunks):
- No paper is silently gutted. Coverage (stored chunk chars / source chars)
is 109–114% for every paper — the >100% is exactly the 64/512-word overlap
inflation, i.e. full retention.
stored == keptfor all 29, so the live DB matches what the current pipeline produces. - Only 18 chunks dropped corpus-wide, all
low_alpha(0 short, 0 bib), concentrated in10.1007/0-387-29555-0_13(6) anddepositonce-5415(10). Inspected: they are garbled OCR tables from scanned-book sources (alpha 0.30–0.36, e.g."(u: 1) 7 < u D:= ud3424 — else H®: out.v. Family 1…") — exactly the artefacts the alpha-ratio gate targets. No real math content lost (the handoff's feared failure mode did not occur; math prose stays > 0.40 alpha). - Head/tail alignment exact on sampled papers: chunk 0 begins at the source head (title/authors/abstract captured), last chunk ends at the source tail → no front/back truncation.
- Note: 0 bibliography drops across the corpus → the upstream
.txtextraction had already stripped reference lists, which is why DQ-1's citation graph relies on the API/GROBID sources rather than text-mined refs. - Acceptance met for all 29 papers (not just a sample). No action needed.
- ⚠ SCOPE LIMIT — DQ-3 covers
.txt → chunksONLY. It treats each source.txtas ground truth and verifies the ingest pipeline preserves it. It does not measure the layer above — PDF/source →.txt— which was done by an upstream tool inConformalLabpp(outside codex-py) and is unexamined. There is already evidence that layer is lossy: the 18 dropped chunks were garbled OCR tables already present in the.txt(i.e. the PDF→txt step garbled the table/formula regions), and the bibliographies were stripped upstream. So "VERIFIED CLEAN" means the loader faithfully preserves whatever the.txtcontains, not that no information was lost from the source paper. See DQ-6.
DQ-6 — PDF/source → .txt extraction fidelity: NOT MEASURED · MED, open
- Open question: how much did the upstream PDF→
.txtextraction (done inConformalLabpp, outside codex-py) drop or garble vs the original papers? Math/equations, tables, multi-column layouts, and figure captions are the usual casualties of PDF text extraction — and we have direct evidence of garble (DQ-3's low_alpha drops were already-garbled tables in the.txt). - Checkable: the 36 original PDFs are present in
…/ConformalLabpp/papers/(map each paper'ssource_path.txtto its sibling.pdfby filename stem). - To investigate: for a sample (incl. the known-bad scanned book
0-387-29555-0_13and a born-digital arXiv PDF), extract PDF text independently and compare length / page-coverage to the.txt; spot-check whether specific theorem statements & equations survive. Flag papers where the.txtis missing large spans. - Acceptance: a per-sample source→txt coverage estimate, or a list of papers
whose
.txtis materially degraded (candidates for re-extraction — note roadmap R-A/R-C would re-ingest from PDF/.texdirectly and largely sidestep this).
Original open question (for context):
- Open question: does the stored chunk set faithfully reconstruct each source
file, or did the chunker /
filter_chunkssilently drop material (e.g. an abstract, a section, math-heavy passages)? The whole corpus was ingested from.txt(PAPERS_DIR=/Users/tarikmoussa/Desktop/ConformalLabpp/papers/txt), so the.tex/.pdfpaths never ran. - To investigate next: for a sample of ~5 papers, compare
len("".join(stored chunks))againstlen(source.txt)(coverage %), and eyeball the first/last chunk vs the file head/tail. Flag papers where coverage is low (content lost) — F-16 dropping >X % of a paper is a signal. - Acceptance: sampled papers retain ≳ the expected fraction of source text; no paper is silently gutted by the quality gate.
DQ-4 — Retrieval quality · MEASURED 2026-06-16 — good, after fixing a P0 crash
P0 bug found + fixed: codex search paper was crash-broken on the live DB.
The CLI passed the query embedding uncast, so pgvector raised
operator does not exist: vector <-> double precision[] on every paper search.
The working call sites (mcp_server.py, wiki.py) cast %(emb)s::vector; the
CLI did not. Fixed both <-> occurrences in codex/cli.py search_paper; added a
regression guard in tests/cli/test_cli.py asserting the cast (mocked-DB unit
tests can't catch the type error — this is why it survived the code audit).
Relevance verdict (post-fix): strong. Six domain queries, both surfaces
(paper-level abstract search + chunk-level). rank-1 was the exactly-correct
paper in 6/6 queries for chunk search and 5/6 for paper-level (Laplace–Beltrami
put the exact paper at rank-2 behind the closely-related "Polygon Laplacian Made
Simple" — acceptable). Examples: "combinatorial Yamabe flow" → math/0306167
(exact, all 5 top chunks that paper); "variational principle for Delaunay
triangulations" → math/0603097 (exact); "circle packing rigidity" →
2601.22903 (exact). Three of the DQ-1-rescued papers (2305.10988,
1005.2698, depositonce-5415) now surface as top hits. The KB is trustworthy
for lookups.
Secondary finding → reinforces DQ-2: zero-vector abstract pollution. Papers
with no abstract get a zero abstract_emb, which sits at constant L2 distance
≈1.000 from every query and appears as filler in the paper-level tail whenever
< 5 strongly-relevant papers exist. 1911.00966 (fully degraded, DQ-2) showed up
at distance 1.000 in 4/6 queries. Harmless at rank>1 with a visible 1.000 score,
but noise — and a concrete reason to fix DQ-2. Chunk-level search is immune (it
filters bibkey IS NOT NULL, and 1911.00966 has no bibkey/chunks).
Note: the MCP search docstring claims "hybrid dense + FTS" but the SQL is
dense-only; score = 1.0 - dist can go negative (bge-m3 L2 ranges 0–2). Ranking
is fine (monotonic); only the absolute score label is misleading.
Original open question (for context):
- Open question: end-to-end, do real queries return relevant results? Clean chunks + a working index don't guarantee useful retrieval.
- To investigate next: run a handful of domain queries through the actual
search path (
codex search paper "<q>"and the chunk-level MCPsearch), e.g. "discrete conformal map", "circle packing rigidity", "combinatorial Yamabe flow", "discrete Laplace-Beltrami operator", and judge whether the top-5 hits are on-topic and point at the right papers. Cross-check a couple against--cite-boostto see if the boost helps or hurts on this corpus. - Acceptance: a short relevance table (query → top-5 → on/off-topic) good enough to trust the KB for lookups, or a list of failure modes to fix.
DQ-5 — ingest_paper is NOT idempotent for DOI papers · RESOLVED 2026-06-17 (found 2026-06-16)
Resolution (what was done):
- Root-cause fix, minimal + contained: added
openalex._normalize_doi()and applied it in_map_paperso the DOI branch yields the canonical bare, lower-cased form (stripshttps://doi.org//http://doi.org//doi:; lower-cases — DOIs are case-insensitive, matchingingest._norm_cited_id). The no-DOI fallback (OpenAlex W-id URL) is left untouched._map_paper→fetch_paper→ingest_paperis the only consumer chain, so blast radius is one function. - Checked nothing depended on the URL form. It was the opposite: the full-URL
id silently defeated
_s2_id_for(paper.id)and_recover_abstract'spaper.id.startswith("10.")(ahttps://…string is neither10.- nordoi:-prefixed), so S2/Crossref recovery never fired for DOI papers. The fix repairs those paths too. - Tests (offline, mocked DB): closed audit T-3 — the OpenAlex fixture now
emits the real URL-form DOI and asserts
paper.idis the bare form; added_normalize_doiunit tests (https/http/doi:/upper-case/already-bare) and a W-id-fallback test; added ingest regressiontest_ingest_doi_paper_upserts_bare _canonical_id(runs the real_map_paperthrough ingest, asserts the papers upsert carries the bare id into%(id)sand usesON CONFLICT (id)). Verified it has teeth (fails on the pre-fix code). Full suite 348 passed; ruff + mypy --strict clean. - Live idempotency confirmed (tunnel up): re-ingested
10.1007/s00454-019-00132-8against the Jetson DB → noUniqueViolation;IngestResult(paper_id='10.1007/s00454-019-00132-8', citations_upserted=45). After-state: UPDATE in place —idandopenalex_idunchanged,added_atunchanged (proves no fresh INSERT), exactly 1 row for the DOI (no URL-form duplicate), citations 45→45, total papers 29→29. - Unblocks R-A (GROBID re-ingest) and R-C (.tex re-ingest) — both re-ingest
existing papers. Files touched:
codex/sources/openalex.py,tests/sources/test_openalex.py,tests/ingest/test_ingest.py. - Note (out of scope, pre-existing): arXiv papers stored as bare ids
(
2305.10988) are still not round-trip idempotent if OpenAlex returns an arXiv-DOI/W-id instead of the bare arXiv id — a separate canonicalization concern, not this UniqueViolation. The DOI path (this finding) is fixed.
Original finding (for context):
- Symptom: re-ingesting an existing DOI paper crashes with
UniqueViolation: papers_openalex_id_key. Surfaced trying to re-ingest10.1007/s00454-019-00132-8during the DQ-2 backfill. - Root cause:
openalex._map_papersetsPaper.id = data["doi"], which is the full URLhttps://doi.org/10.1007/…. The stored canonical id is the bare10.1007/…(post the M-1 canonical-id migration). SoINSERT … ON CONFLICT (id)finds no id match, attempts a fresh INSERT, and trips the separate unique constraint onopenalex_id(already held by the canonical-id row). - Confirmed:
openalex.fetch_paper("10.1007/s00454-019-00132-8").id == "https://doi.org/10.1007/s00454-019-00132-8"(≠ stored bare id). - Impact: any re-ingest of a DOI paper aborts. This blocks roadmap R-A (GROBID on PDFs) and R-C (.tex ingest) — both re-ingest existing papers. Also implies new DOI ingests store full-URL ids, inconsistent with the bare-id corpus (latent split-identity risk).
- Proposed fix (handle with care — M-1 incident territory): normalize the DOI
in
_map_paperto the canonical bare, lower-cased form (striphttps://doi.org/), matching what the M-1 migration produced. Add a re-ingest idempotency test over a DOI paper. Did not fix inline this session — the canonical-id derivation is exactly where the M-1 migration incident occurred, so it needs its own focused change + a live re-ingest check. - Workaround used for DQ-2: targeted
UPDATE … SET abstract, abstract_embinstead of full re-ingest.
Roadmap — data-acquisition levers (post-audit, all free sources)
These are forward-looking acquisition improvements, distinct from the DQ-1..DQ-4 audit findings (which assessed existing data). They came out of a strategy discussion on 2026-06-16: for this Math/CS corpus the valuable data (references, abstracts, full text) is almost entirely open (arXiv, OpenAlex, Crossref, S2), so the quality lever is combining more free sources, not buying paywall access. Paywall/uni-login was explicitly considered and deferred (see R-E). Each item is self-contained so a cold session can pick it up.
R-A — Run GROBID reference extraction on PDFs · CORE DONE 2026-06-17 (full sweep optional)
Resolution (what was done) — citing coverage 27/29 → 29/29 (100%); edges 920 → 1022:
- References-only backfill, not full re-ingest. Added
scripts/ra_grobid_backfill.py: per paper,
run
grobid.extract_references(pdf), normalize cited-ids to the canonical bare form (DOIs de-URL'd + lower-cased;arXiv:stripped), INSERT intocitations(ON CONFLICT DO NOTHING). Deliberately avoidsingest_paper(source_path=pdf), whose PDF branch re-runs Nougat +DELETE FROM chunksand would overwrite the DQ-3-verified-clean.txtchunks (and needs Nougat). Idempotent; dry-run by default; 7 tests; ruff/mypy clean. - GROBID provisioning: the pinned
grobid/grobid:0.8.2is amd64-only, but the Jetson and the dev Mac are both arm64 (this is why GROBID never ran). Ran the lighter CRF image (grobid/grobid:0.8.2-crf) on the Mac under Podman+Rosetta (PDFs are local; citations written over the SSH DB tunnel). Madeextract_referencestimeout configurable (theses need >60s emulated). A separate task is provisioning a native arm64 GROBID on the Jetson (docs/infra/grobid-jetson.md). - Both
depositoncetheses closed:…depositonce-20357(lutz-2024) +96 edges (5 in-corpus);…depositonce-5415(sechelmann-2016) +6 edges — its 147 MB/173-page scan blew pdfalto's 120s limit, so the 9 bibliography pages were extracted with PyMuPDF first (77 refs, only 6 with a DOI/arXiv). No paper has zero out-edges now. - Optional remaining (secondary acceptance): sweep the other 27 already-covered
papers for extra GROBID refs via
python scripts/ra_grobid_backfill.py --write(the 2 book PDFs need the same bib-page extraction as sechelmann).
Original plan (for context):
- ✓ Unblocked (DQ-5 RESOLVED 2026-06-17): DOI re-ingest is now idempotent, so
the GROBID re-ingest no longer aborts on
papers_openalex_id_key. - What: the whole corpus was ingested from
.txt(PAPERS_DIR=…/papers/txt), so the GROBID PDF path never ran. Re-ingest with a PDFsource_pathper paper socodex.parsing.grobid.extract_referencesparses each bibliography. - Why: extracts references the APIs lack — a complement to DQ-1's S2
supplement. Crucially, the 2
depositoncetheses that stayed at zero in DQ-1 (no API references anywhere) almost certainly have a references section in their open-access PDFs → GROBID could finally give them out-edges. Closes the last 2/29. - How: no new code — ingest already merges + dedups
api_citations + pdf_citations(codex/ingest.py:288). GROBID server is already configured (GROBID_URL=http://192.168.178.103:8070in.env.jetson-ingest). Download free arXiv PDFs, re-ingest withsource_path. - Acceptance: the 2 theses gain references; published papers gain GROBID-parsed refs beyond OpenAlex/S2 coverage; citing coverage 27/29 → 29/29.
R-B — Add Crossref as a third citation + abstract source · MED lever, MED effort
- What: new
codex.sources.crossrefclient hitting Crossref REST/works/{doi}for thereferencearray andabstract. - Why: Crossref carries publisher-deposited references for many DOIs (free,
polite pool via
mailto) — a third leg after OpenAlex/S2. It also recovers abstracts OpenAlex is contractually barred from redistributing (likely the 2 Springer DOIs missing abstracts in DQ-2). cited_ids are DOIs → flow through the existing graph resolver unchanged. - How: mirror
codex/sources/openalex.py(tenacity retry, polite pool). Extend the ingest fallback chain to OpenAlex-empty → S2 → Crossref (same_s2_reference_supplementshape). Add unit tests with mocked HTTP. - Acceptance: Crossref recovers refs and/or an abstract for ≥1 paper where OpenAlex+S2 are empty.
R-C — Prefer arXiv .tex source over .txt for chunk fidelity · DONE 2026-06-17
Resolution (what was done) — all 13 arXiv papers re-ingested from .tex:
- Multi-file flatten:
arxiv.fetch_sourcenow inlines\input/\include(tex.flatten_inputs) so multi-file projects yield the full body, not the primary file's include skeleton (e.g. math/0603097 = 15 files / 12 includes; 2305.10988 = 6 / 5). Also handles legacy single-file bare-gzipped.tex(no tar wrapper) for old math/* papers (commit121b582). - Section-aware chunking: new
tex.chunk_sectionsreturns(title, chunk)pairs (chunks never span a\section); the ingest.texpath labels the storedsectionfrom the real heading instead of runningclassify_sectionon a header-less word-window. Quality-gated as pairs so labels stay aligned. - Unblocked by the arXiv half of DQ-5 (commit
f504ca6):_canonical_idstrips the10.48550/arxiv.DOI prefix so an arXiv paper's id is the bare arXiv id — without it, re-ingest trippedpapers_openalex_id_key(confirmed live, then fixed; idempotent re-ingest verified). - Outcome:
sectioncolumn gained signal — ~34/329 arXiv chunks now intro/theorem/proof (was ~allbody); math/0603097{body:31}→{body:29, proof:9}, 2310.17529 gained proof+intro, etc. Math papers with descriptively titled sections staybody(classify_section's fixed vocab) — modest but real (R-12). No paper gutted (chunk counts stable/▲); F-16 gate unchanged. - Files:
codex/parsing/tex.py,codex/sources/arxiv.py,codex/ingest.py,codex/sources/openalex.py,scripts/rc_tex_reingest.py(driver, dry-run default), + tests. Full suite 370 passed; ruff/mypy clean. Branchfeat/tex-ingest(commits1da81c7,121b582,f504ca6).
Original plan (for context):
- What: re-ingest from arXiv LaTeX source through the existing
codex.parsing.tex.latex_to_textpath instead of flattened.txt. - Why: highest-fidelity input — preserves math, structure, and section
boundaries. Directly serves DQ-3 (content fidelity) and audit R-12 (the
sectioncolumn is low-signal because word-window.txtchunks rarely start at a real header;.texheaders fix this). - How: download free arXiv source tarballs, feed
.texto ingest withsource_path. Handle multi-file projects (main.tex+ includes). - Acceptance:
sectioncolumn gains signal (fewerbody-only); DQ-3 coverage vs source improves; no regression in the F-16 quality gate.
R-D — F-15: warn on low citing-paper coverage, not just paper count · SMALL
- What: the open DQ-1 sub-item:
graph_min_corpus_sizeonly flags total paper count. Add a warning when the share of papers with ≥1 out-edge is low (e.g. < 80%), since that is what actually starves PageRank/coupling. - Why: at 17/29 the graph silently ran on 59% of the corpus with no signal. Post-DQ-1 it's 27/29 (93%) — but the guard should catch future regressions.
- Acceptance:
codex graph reportsurfaces citing-coverage % and warns below threshold.
R-E — Paywall / institutional access · DEFERRED — only for a paywall-only expansion
- Decision (2026-06-16): a uni login does not help the current corpus — every DQ-1/DQ-2 gap is in openly-available works (arXiv preprints, OA theses), and all metadata/citation sources are free APIs. Revisit only if the corpus expands to papers that exist solely behind a paywall with no preprint; then institutional access → full-text PDF → GROBID refs (R-A) + clean chunks (R-C).
- ToS caveat: manual download of individual papers you have legitimate access to is fine; automated bulk download through a uni proxy (EZproxy/Shibboleth) violates most publisher terms and can get the whole institution's access revoked. Do not wire a publisher login into the ingest pipeline.
R-F — Richer section labels: store real \section titles · DONE 2026-06-17
Resolution (what was done):
- Why: R-C labelled
.texchunks by mapping the real\sectionheading throughclassify_section's fixed vocab, so descriptively-titled Math sections collapsed tobody(only 34/329 = 10 % non-body). A research pass confirmed thesectioncolumn is write-only (nothing filters on the vocab —mcp_server.pyandwiki.pydon't even SELECT it), so storing free-text titles is safe. - Change: new
quality.section_label(title, content)(+_clean_title) — keep the controlled bucket (intro/theorem/proof/abstract/bibliography) when the heading maps to one, else store the cleaned real title ("the flip algorithm", "main results", …). The.texingest path uses it;.pdf/.txt/run_quality_passkeepclassify_sectionunchanged._clean_titlestrips LaTeX commands/accents/ties + leading numbering, lower-cases, truncates at a word boundary. ~1 helper + 1 ingest line; no schema change (column already freeTEXT). - Result (live): re-applied to the 13 arXiv papers → non-
body34/329 (10 %) → 314/329 (95 %). Onlymath/0306167(old AMS-TeX, no\section{}) staysbody. - Tests:
section_label/_clean_titleunits +.texingest stores a descriptive heading verbatim. Full suite 377 passed; ruff/mypy clean. Branchfeat/section-labels(offfeat/tex-ingest). - Pending (infra, not code): an accent-cleanup polish (
m\"obius→mobius) landed in_clean_title, and a quick in-placeUPDATEre-clean (no re-embed) was started, but the Jetson went offline (ping 100 % loss) before it finished — a few live labels still show LaTeX accents/ties. Re-run the in-place re-clean when the Jetson is back.
Priority for the next session
DQ-1— DONE 2026-06-16 (live DB at 920 citations, 27/29 citing coverage; remedy wired into ingest).DQ-4— DONE 2026-06-16 (P0search papercrash fixed; relevance verified strong; surfaced the zero-vector pollution that motivates DQ-2).DQ-2— DONE 2026-06-16 (1911.00966 fully recovered; s00454 abstract from S2; book chapter documented; recovery wired into ingest).DQ-3— DONE 2026-06-16 (content fidelity clean; full coverage, only 18 OCR-junk chunks dropped, no math lost). Assessment axis DQ-1..DQ-4 complete.DQ-5— DONE 2026-06-17 (DOI id normalized to bare form inopenalex._map_paper; live re-ingest of10.1007/s00454-019-00132-8idempotent — UPDATE in place,added_atunchanged; suite 348 green). Unblocks R-A/R-C.- Roadmap levers (see section above — now all unblocked): R-A (GROBID on
arXiv PDFs — closes the last 2 zero-citation theses), R-B (Crossref refs —
abstract half already shipped in DQ-2), R-C (
.texingest — serves DQ-3), R-D (F-15 coverage warning, quick). R-E (paywall) deferred.
DQ-1..DQ-4 are data/assessment work; DQ-5 + roadmap R-A..R-D add new code
(canonical-id fix, re-ingest scripts, Crossref references, a .tex path, a graph
warning) on top of the already-remediated pipeline (see the AUDIT-* docs and PRs
#12–#14). The DQ-1/DQ-2 ingest wiring + Crossref/arXiv abstract recovery already
landed this session.
Appendix — reproduce the baseline scan
import os, psycopg, statistics
from collections import Counter, defaultdict
from psycopg.rows import dict_row
from codex.config import Settings
from codex.quality import is_quality_chunk, _bib_score
s = Settings()
with psycopg.connect(os.environ["DATABASE_URL"], row_factory=dict_row) as c:
papers = c.execute("SELECT id, bibkey, title, abstract, year, authors, openalex_id FROM papers").fetchall()
chunks = c.execute("SELECT paper_id, content, section FROM chunks").fetchall()
cit = c.execute("SELECT citing_id, count(*) n FROM citations GROUP BY citing_id").fetchall()
# zero-citation papers (DQ-1)
cited = {r["citing_id"] for r in cit}
zero = [p["id"] for p in papers if p["id"] not in cited]
# metadata gaps (DQ-2): p["abstract"] empty / p["bibkey"] None / not p["authors"]
# F-16 re-gate (baseline): [ch for ch in chunks if not is_quality_chunk(ch["content"], settings=s)]
# section dist (R-12): Counter(ch["section"] or "(null)" for ch in chunks)