2026-06-04 23:31:15 +02:00
2026-06-04 11:19:10 +02:00
2026-06-04 11:19:10 +02:00
2026-06-04 11:19:10 +02:00
2026-06-04 11:19:10 +02:00
2026-06-04 11:19:10 +02:00
2026-06-04 11:19:10 +02:00

codex — Personal Paper Knowledge Base

A self-hostable tool for managing scientific papers: ingest PDFs and arXiv sources, build a citation graph, and link C++ implementations back to the papers that describe them.


Quick start

# 1. Start Postgres (pgvector) + GROBID
docker compose -f infra/docker-compose.yml up -d

# 2. Copy and edit environment variables
cp .env.example .env
$EDITOR .env

# 3. Install Python dependencies (requires uv)
uv sync

# 4. Apply the database schema (first run only)
uv run python -c "
from codex.db import get_conn, apply_schema
with get_conn() as conn:
    apply_schema(conn)
print('Schema applied.')
"

Environment variables

Variable Default Description
DATABASE_URL postgresql://researcher:change_me@localhost:5432/papers libpq connection string
GROBID_URL http://localhost:8070 GROBID HTTP API base URL
OLLAMA_BASE_URL http://localhost:11434 Local Ollama endpoint (optional)
EMBEDDING_MODEL BAAI/bge-m3 sentence-transformers model name
EMBEDDING_DIM 1024 Embedding vector dimension (must match model)
OPENALEX_MAILTO (empty) E-mail for OpenAlex Polite Pool (required for automated use)

See .env.example for a documented template.


Three-layer data model

All data lives in a single Postgres instance with the pgvector extension.

Layer 1 — Semantics (papers, chunks)

Papers are stored with metadata and an abstract-level dense embedding (BGE-M3, 1024 dimensions). Full-text is split into chunks, each with its own dense embedding and a Postgres full-text (GIN) index.
Hybrid search combines nearest-neighbour vector search with keyword (FTS) retrieval for robust handling of exact mathematical terminology.

Layer 2 — Citations (citations)

Directed edges in the citation graph. The cited_id column has no foreign-key constraint on purpose: edges pointing to papers that have not yet been ingested are kept as-is. Those "dangling" targets are your discovery leads — papers frequently cited by your collection that you have not yet read.

-- Top discovery leads
SELECT cited_id, count(*) AS pull
FROM   citations
WHERE  cited_id NOT IN (SELECT id FROM papers)
GROUP  BY cited_id
ORDER  BY pull DESC
LIMIT  20;

Maps C++ symbols (qualified names or file.cpp:line references) to the papers they implement. The workflow:

  1. Tag C++ source with @cite <bibkey> in Doxygen comments.
  2. Run codex provenance sync --lib-path <path> to scan and resolve tags.
  3. Run codex provenance export-bib <out.bib> to generate a .bib file containing only the cited subset of your collection.

The exported .bib is a derived view of the master catalogue — regenerate it at any time; it is not the source of truth.


CLI reference (coming in F-07)

codex ingest <id>               # ingest one paper by arXiv ID or DOI
codex ingest-file <ids.txt>     # bulk ingest from a file of IDs
codex search "<query>" [--hybrid]
codex discover leads
codex provenance sync --lib-path <path>
codex provenance export-bib <out.bib>
codex ask "<question>"          # optional LLM Q&A via Ollama

Development

uv run ruff check . && uv run ruff format --check .   # lint
uv run mypy codex/                                     # type-check
uv run pytest                                          # tests
Description
conformalap knowledge base — Python implementation (F-01–F-08)
Readme 2 MiB
Languages
Python 97.5%
Shell 2.5%