Data-quality roadmap: R-A–R-F + DQ-5 fixes (citations 590→1075, section signal 10%→95%) #16
@@ -104,7 +104,36 @@ docker compose logs -f grobid # wait for "Started ... in N seconds" (mode
|
|||||||
|
|
||||||
The `deploy.resources.limits.memory: 4g` cap in the compose file is honored by
|
The `deploy.resources.limits.memory: 4g` cap in the compose file is honored by
|
||||||
Compose v2 here (no swarm needed). 3 GB suffices for references; 4 GB leaves
|
Compose v2 here (no swarm needed). 3 GB suffices for references; 4 GB leaves
|
||||||
headroom alongside Postgres.
|
headroom alongside Postgres. The `restart: "no"` policy means GROBID does **not**
|
||||||
|
auto-start on boot — see "Run on-demand" below.
|
||||||
|
|
||||||
|
---
|
||||||
|
|
||||||
|
## Run on-demand (stop when idle)
|
||||||
|
|
||||||
|
GROBID is only needed **during reference extraction** (`scripts/ra_grobid_backfill.py`);
|
||||||
|
nothing else in the stack uses it. Idle it still holds ~3.5 GB, so on the 8 GB Jetson
|
||||||
|
keep it stopped and start it only for an R-A run:
|
||||||
|
|
||||||
|
```bash
|
||||||
|
ssh alfred@192.168.178.103 'docker start papers-grobid' # ~30s model load
|
||||||
|
curl -s --retry 30 --retry-delay 2 --retry-connrefused \
|
||||||
|
http://192.168.178.103:8070/api/isalive # wait for: true
|
||||||
|
|
||||||
|
# ... run R-A from the Mac ...
|
||||||
|
PYTHONPATH=. python scripts/ra_grobid_backfill.py --write
|
||||||
|
|
||||||
|
ssh alfred@192.168.178.103 'docker stop papers-grobid' # reclaim ~3.5 GB
|
||||||
|
```
|
||||||
|
|
||||||
|
First-time deploy uses `docker compose up -d grobid` (above); afterwards the
|
||||||
|
container persists, so `docker start/stop papers-grobid` is all you need.
|
||||||
|
|
||||||
|
> Why on-demand rather than a lighter tool: refextract was evaluated head-to-head
|
||||||
|
> on the Jetson and rejected — far worse DOI recall on this math corpus (lutz
|
||||||
|
> thesis: GROBID **101** usable IDs vs refextract **2**), and it was *slower* per
|
||||||
|
> PDF despite a smaller footprint. GROBID stays; running it on-demand removes its
|
||||||
|
> only real downside (idle RAM).
|
||||||
|
|
||||||
---
|
---
|
||||||
|
|
||||||
|
|||||||
@@ -44,7 +44,11 @@ services:
|
|||||||
init: true # reap zombie children (GROBID docs recommend --init)
|
init: true # reap zombie children (GROBID docs recommend --init)
|
||||||
ulimits:
|
ulimits:
|
||||||
core: 0 # disable core dumps (GROBID docs: --ulimit core=0)
|
core: 0 # disable core dumps (GROBID docs: --ulimit core=0)
|
||||||
restart: unless-stopped
|
# On-demand only: GROBID is needed solely during reference extraction
|
||||||
|
# (scripts/ra_grobid_backfill.py), not for normal corpus operations — start it
|
||||||
|
# for an R-A run, stop it after to reclaim ~3.5GB. "no" = no auto-start on boot.
|
||||||
|
# See docs/infra/grobid-jetson.md ("Run on-demand").
|
||||||
|
restart: "no"
|
||||||
# GROBID is RAM-hungry; cap usage on the constrained Jetson (Orin Nano, 8GB
|
# GROBID is RAM-hungry; cap usage on the constrained Jetson (Orin Nano, 8GB
|
||||||
# RAM shared with Postgres). 3GB suffices for references; 4g leaves headroom.
|
# RAM shared with Postgres). 3GB suffices for references; 4g leaves headroom.
|
||||||
# Honored by `docker compose up` under Compose v2 (no swarm needed).
|
# Honored by `docker compose up` under Compose v2 (no swarm needed).
|
||||||
|
|||||||
Reference in New Issue
Block a user