# Finding Orchestration — Sessions × Models × Review Gates **Updated:** 2026-05-31 **Purpose:** a structured plan to work through every audit finding across **multiple short sessions**, each run by the **model best suited** to the work, with an **external Opus reviewer session after every implementation session**. This lets you pick up any "Session" below cold, hand it to the named model, and know exactly what it covers and how its output gets validated. > Companion: the per-finding detail lives in the 11 audit documents in this > folder (see [`README.md`](README.md)). This file is the *meta-plan* — the > order, the model assignment, and the review cadence. --- ## How to use this 1. Pick the next **⬜ pending** session from the [sequence](#session-sequence). 2. Start a session with the **assigned model**, point it at the listed findings in their audit docs (each finding has `file:line`, fix, acceptance criteria). 3. When it finishes, start the paired **🔍 Opus review session** — it validates the diff (correctness, value-identity, tests, no parity regression) and only then is the batch "done". 4. Mark the session ✅ here and move on. Rationale for the cadence: implementation models (Sonnet/Haiku) are cheaper and fast for well-specified work; Opus is reserved for (a) findings that need real numerical/architectural judgement and (b) the **review gate**, because an independent high-capability pass catches subtle parity/numerics regressions that a self-review from the implementing model tends to miss. --- ## Model-assignment heuristic | Model | Take findings that are… | Examples | |---|---|---| | **Haiku** | mechanical, local, low-risk; docs, citations, renames, named-constant extraction | M-series, N4/N6, A1–A3, doc files | | **Sonnet** | normal coding with a clear acceptance criterion; tests, error-handling, CI, small API changes | V-series, C-series, I-series, H-series | | **Opus** | numerics, math correctness, architecture, public-API/irreversible decisions — **and every review gate** | B1-port, N1/N3/N5/N7, H2 refactor, A4/A5, all 🔍 reviews | > Decision rule isn't "how severe" but "how much must be *understood* to get it > right." A 🔴 one-line switch can be Sonnet; a 🟡 numerical reformulation is Opus. --- ## Master finding table Status: ✅ done · ⬜ open (actionable) · ⏸ deferred (intentional) · ⛔ blocked (G0) · 👤 needs human expert | ID | Audit | Sev | Status | Model | Session | |----|-------|-----|--------|-------|---------| | B1 (HyperIdeal) | api-perf | 🔴 | ✅ | Sonnet | S1 | | B1 (Inv-Dist port) | api-perf | 🔴 | ✅ | Opus | S1 | | V3 | input-val | 🟡 | ✅ | Sonnet | S1 | | C1 | test-cov | 🔴 | ✅ | Sonnet | S1 | | N4, N6 | numerics | 🟡/🔵 | ✅ | Haiku | S1 | | M1, M2, M4 | math-cite | 🟡/🔵 | ✅ | Haiku | S1 | | C2, C3 | test-cov | 🔴 | ✅ | Sonnet | S1 | | V1, V2, V4 | input-val | 🟡 | ✅ | Sonnet | S1 | | I2, I3, I4 | test-cov | 🟡 | ✅ | Sonnet | S1 | | A1, A2, A3 | api-perf | 🔴/🟡 | ✅ | Haiku | S1 | | N5 | numerics | 🟡 | ✅ | Opus | S1 | | N3 | numerics | 🟡 | ✅ | Opus | S1 | | H2, B2, B3, B4, B5 | test-cov/api-perf | 🔵/🟡 | ✅ | Opus | S1 | | I1 | test-cov | 🟡 | ✅ | Opus | S2 | | H1 | test-cov | 🔵 | ✅ | Opus | S2 | | N7 | numerics | 🔵 | ✅ | Opus | S2 | | **H3** | test-cov | 🔵 | ⬜ | Sonnet→🔍Opus | **S3** | | **H4** | test-cov | 🔵 | ⬜ | Sonnet→🔍Opus | **S3** | | **H5** | test-cov | 🔵 | ⬜ | Sonnet→🔍Opus | **S3** | | **V5, V6** | input-val | 🔵 | ⬜ | Sonnet→🔍Opus | **S3** | | **N2** | numerics | 🟡 | ⬜ | Haiku→🔍Opus | **S4** | | **thread-safety doc** | thread-safety | 🟡 | ⬜ | Haiku→🔍Opus | **S4** | | **M3** | math-cite | 🟡 | ⬜ | Haiku→🔍Opus | **S4** | | **I5** | test-cov | 🟡 | ⬜ | Sonnet→🔍Opus | **S5** | | A4, A5 | api-perf | 🟡 | ⏸ | Opus | S6 (after G0/G1) | | N1 | numerics | 🔴 | ✅* | — | (subsumed by B1; doc note folds into N2/S4) | | G1–G12 | cgal | ⛔/🔴 | ⛔ | Opus | (after G0) | | D1, D2 | dep-license | 🔴 | ⛔ | Opus | (after G0) | | M5 | math-cite | 🔵 | 👤 | — | human domain expert | | G0 | cgal | ⛔ | 🔄 in progress | owner | **author contacted by email — awaiting reply** | \* N1 (FD-step vs Newton tol) is largely subsumed by the B1 block-FD work; the remaining piece is the tolerance-coupling note, folded into N2 (Session S4). --- ## Session sequence ### ✅ S1 — Audit quick-wins + numerics + refactor (DONE, 2026-05-31) Branch `fix/b1-v3-c1-quick-wins`, 9 commits, 298/298 CGAL tests green. - **Sonnet:** B1(HyperIdeal)/V3/C1; C2/C3, V1/V2/V4, I2/I3/I4. - **Haiku:** N4/N6; M1/M2/M4; A1–A3. - **Opus:** B1 inv-dist port; N5; N3; H2+B2/B3/B4/B5. - **🔍 Opus review:** validated all Sonnet/Haiku commits (value-identity, error handling, coverage gate, alias forwarding) — all correct. ### ✅ S2 — Solver result diagnostics (DONE, 2026-05-31, Opus impl + self-review) Done directly in the warm post-refactor Opus session (all three findings live in the just-written `newton_core`/`NewtonResult`/`NewtonLinearSolver`). Commit `90e966a`; 301/301 CGAL tests green. **I1** status enum + **H1** iteration fix + **N7** `sparse_qr_fallback_used` / `min_ldlt_pivot` diagnostics. Follow-up idea (optional): propagate `NewtonStatus` up into the CGAL `Conformal_map_result`.
original S2 plan #### S2 — Solver result diagnostics (Sonnet → 🔍 Opus) - **I1** — give `NewtonResult` a status enum (`Converged`/`MaxIterations`/`LinearSolverFailed`/`LineSearchStalled`). The `newton_core` refactor centralised every exit point, so this lands in one place now. Update the 5 wrappers + any callers reading `converged`. - **H1** — `res.iterations` is stale when the solver breaks on `!ok`/stall; set it correctly at each `break` in `newton_core`. - **N7** *(Opus within this session)* — optional conditioning diagnostic (smallest LDLT pivot / cheap cond estimate) surfaced in the result; pairs naturally with the I1 status enum (near-singular ⇒ a distinct status/flag). - **🔍 Opus review:** confirm the enum doesn't break the CGAL-layer result mapping and that `converged` semantics are unchanged for existing callers.
### ⬜ S3 — Robustness & test-gap closure (Sonnet → 🔍 Opus) - **H3** — `enforce_gauss_bonnet` reports the magnitude of the applied correction. - **H4** — test `reduce_to_fundamental_domain` boundary `Im(τ)==0.0`. - **H5** — integration test for the degenerate-triangle path in the Newton solve. - **V5** — make the hand-rolled XML reader *reject* reformatted-but-valid XML instead of silently mis-reading it into zeros. - **V6** — DOF-vector-vs-mesh size check on load. - **🔍 Opus review:** verify the new rejections don't break valid round-trips. ### ⬜ S4 — Documentation & citations (Haiku → 🔍 Opus) - **N2** — `doc/math/tolerances.md`: every numerical threshold, its role, and the constraints between them (incl. the N1 FD-step ≪ Newton-tol coupling). - **thread-safety** — document the same-mesh concurrency contract (the code has no mutable static state; this is purely the written contract). - **M3** — re-verify the future-dated / arXiv-only citations against their published versions; update `references.md`. - **🔍 Opus review:** sanity-check the tolerance relationships are stated correctly. ### ⬜ S5 — Coverage truthfulness (Sonnet → 🔍 Opus) - **I5** — wire `coverage.sh` to the full CGAL suite (not just the 26 fast tests, ~9.6%), then remove `SKIP_COVERAGE_GATE=1` so the agreed 80%/70%/90% gate goes live (C3 ramp-up → enforced). - **🔍 Opus review:** confirm the measured numbers are honest and the gate fails correctly when coverage drops. ### ⏸ S6 — Public CGAL surface + packaging (Opus, **only after G0/G1**) Gated by the author's reply on porting/relicensing rights. - **A4, A5** — public entry-point renames + unified `Circle_packing_result`. - **G1–G12, D1, D2** — license headers, CGAL package layout, concept doc, user manual, CGAL test harness, dependency/data provenance. - **G0 status:** *author contacted by email (2026-05-31), awaiting reply.* No public release / submission / relicensing until a written answer exists. ### 👤 Out of model scope - **M5** — the large hand-derivations need a domain-expert prose review (the numerical results are already validated; the prose is not). --- ## Review-gate checklist (what every 🔍 Opus session checks) - [ ] Builds clean; full CGAL suite green (no count regression). - [ ] No Java golden-vector / parity test perturbed (HardJava defaults intact). - [ ] Numeric changes are value-identical where claimed, or justified + tested. - [ ] New public surface (result types, enums, API) is intentional and documented. - [ ] Commit message attributes the implementing model (`Co-Authored-By: Claude `). - [ ] Finding marked ✅ in the master table above with the commit ref.