diff --git a/docs/infra/jetson-ssh-stability.md b/docs/infra/jetson-ssh-stability.md new file mode 100644 index 0000000..2358f21 --- /dev/null +++ b/docs/infra/jetson-ssh-stability.md @@ -0,0 +1,66 @@ +# Infra TODO — stable long-running ops against the Jetson (tmux + autossh) + +**Status:** OPEN (infra). **Filed:** 2026-06-17. **Priority:** MED — it slows live ops +but has a known workaround (short/idempotent ops). + +## Problem (observed) + +Live data ops run on the **Mac** and reach the Jetson Postgres through an SSH +port-forward tunnel (`ssh -f -N -L 5433:localhost:5432 alfred@192.168.178.103`, +`DATABASE_URL` → `localhost:5433`). During the R-A / R-C / R-F work this tunnel +**dropped repeatedly mid-run**: any operation longer than a few minutes (corpus +re-ingest, the bge-m3 embed loop, the GROBID sweep) had a high chance of failing +partway when the tunnel died, leaving the work half-applied. At one point the Jetson +went **fully offline** (ping 100 % loss, `ssh:22` timeout) mid-operation. + +`ingest_paper()` and the backfill scripts open one DB connection for the whole run, so +a tunnel drop raises `OperationalError: connection refused` and aborts — partial state, +manual retry. Adding `ServerAliveInterval` to the tunnel helped only marginally. + +## Recommended fix + +**1. Run long ops *on the Jetson* inside `tmux` (primary fix).** +The Jetson is where Postgres (and GROBID) live, so running ingest there means the DB is +`localhost` — **no tunnel at all** — and `tmux` keeps the process alive across SSH +disconnects (detach / reattach to monitor). This removes the tunnel from the critical +path for every long-running write. + +- [ ] Confirm the `codex` env is deployable on the Jetson (Python 3.12+, deps incl. the + bge-m3 embedder; the Orin Nano GPU should handle it — GROBID already runs there). + If not, stand up a venv / container for it. +- [ ] Standard pattern: `ssh alfred@jetson`, then + `tmux new -s codex` → `set -a; source .env; set +a` → + `PYTHONPATH=. python scripts/.py --write` → detach (`Ctrl-b d`). + Reattach later with `tmux attach -t codex`. The job survives the SSH drop. +- [ ] Point the Jetson-side `DATABASE_URL` at `localhost:5432` directly (no `:5433`). +- [ ] Document this as the default for any multi-minute write (re-ingest, sweeps, + embed loops) in `docs/audit/DATA-QUALITY-2026-06-15.md` "How to reach the data". + +**2. `autossh` for when a tunnel *is* needed (secondary).** +For interactive/dev use from the Mac (psql, quick probes, read-only queries), replace +the plain `ssh -f -N -L` with `autossh -M 0 -f -N -o ServerAliveInterval=15 -o +ServerAliveCountMax=3 -L 5433:localhost:5432 alfred@jetson` so a dropped tunnel +auto-reconnects. NB: autossh only restarts the *tunnel* — an in-flight transaction +still fails, so this is for short/interactive use, not long writes (use tmux-on-Jetson +for those). + +**3. (Optional, code) make backfill scripts resilient.** +Lower-value once (1)/(2) are in place, but: open a fresh `get_conn()` per paper (not one +for the whole run) and/or wrap the per-item write in a small retry, so a transient drop +costs one item, not the whole batch. The R-A backfill already does per-paper connections; +`rc_tex_reingest.py` opens a connection per paper too — but `ingest_paper` itself holds +one connection for its whole body, which is the unit that fails. + +## Why this matters + +The chronic flakiness directly cost time this session and left one cosmetic task +unfinished (the R-F accent re-clean — see DATA-QUALITY-2026-06-15.md R-F "Pending"). The +**robust pattern that already works** is short, idempotent operations (per-paper, or a +single quick `UPDATE`); tmux-on-Jetson makes even the long ops safe. + +## Acceptance + +- A documented, repeatable way to run a full corpus re-ingest that survives an SSH drop + (tmux-on-Jetson), verified by detaching/reattaching across a disconnect. +- The DATA-QUALITY doc's "How to reach the data" section recommends tmux-on-Jetson for + long writes and `autossh` for interactive tunnels.