- finding-orchestration.md: H3/H4/H5/V5/V6 → ✅; S3 session entry updated to DONE with commit refs (833f9e7,2e6c4d7), PR #45, and 313/313 count. - test-coverage-error-handling-audit-2026-05-31.md: banner updated to reflect H3/H4/H5 resolved in S3; test count 298 → 313. - input-validation-audit-2026-05-31.md: banner updated to reflect V5/V6 resolved in S3; test count 298 → 313. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>
9.7 KiB
Finding Orchestration — Sessions × Models × Review Gates
Updated: 2026-05-31 Purpose: a structured plan to work through every audit finding across multiple short sessions, each run by the model best suited to the work, with an external Opus reviewer session after every implementation session. This lets you pick up any "Session" below cold, hand it to the named model, and know exactly what it covers and how its output gets validated.
Companion: the per-finding detail lives in the 11 audit documents in this folder (see
README.md). This file is the meta-plan — the order, the model assignment, and the review cadence.
Ready-to-paste prompts for every pending session (S3–S6 + the review gate) live in
session-prompts.md— copy one block, set the named model, go.
How to use this
- Pick the next ⬜ pending session from the sequence.
- Start a session with the assigned model, point it at the listed findings
in their audit docs (each finding has
file:line, fix, acceptance criteria). - When it finishes, start the paired 🔍 Opus review session — it validates the diff (correctness, value-identity, tests, no parity regression) and only then is the batch "done".
- Mark the session ✅ here and move on.
Rationale for the cadence: implementation models (Sonnet/Haiku) are cheaper and fast for well-specified work; Opus is reserved for (a) findings that need real numerical/architectural judgement and (b) the review gate, because an independent high-capability pass catches subtle parity/numerics regressions that a self-review from the implementing model tends to miss.
Model-assignment heuristic
| Model | Take findings that are… | Examples |
|---|---|---|
| Haiku | mechanical, local, low-risk; docs, citations, renames, named-constant extraction | M-series, N4/N6, A1–A3, doc files |
| Sonnet | normal coding with a clear acceptance criterion; tests, error-handling, CI, small API changes | V-series, C-series, I-series, H-series |
| Opus | numerics, math correctness, architecture, public-API/irreversible decisions — and every review gate | B1-port, N1/N3/N5/N7, H2 refactor, A4/A5, all 🔍 reviews |
Decision rule isn't "how severe" but "how much must be understood to get it right." A 🔴 one-line switch can be Sonnet; a 🟡 numerical reformulation is Opus.
Master finding table
Status: ✅ done · ⬜ open (actionable) · ⏸ deferred (intentional) · ⛔ blocked (G0) · 👤 needs human expert
| ID | Audit | Sev | Status | Model | Session |
|---|---|---|---|---|---|
| B1 (HyperIdeal) | api-perf | 🔴 | ✅ | Sonnet | S1 |
| B1 (Inv-Dist port) | api-perf | 🔴 | ✅ | Opus | S1 |
| V3 | input-val | 🟡 | ✅ | Sonnet | S1 |
| C1 | test-cov | 🔴 | ✅ | Sonnet | S1 |
| N4, N6 | numerics | 🟡/🔵 | ✅ | Haiku | S1 |
| M1, M2, M4 | math-cite | 🟡/🔵 | ✅ | Haiku | S1 |
| C2, C3 | test-cov | 🔴 | ✅ | Sonnet | S1 |
| V1, V2, V4 | input-val | 🟡 | ✅ | Sonnet | S1 |
| I2, I3, I4 | test-cov | 🟡 | ✅ | Sonnet | S1 |
| A1, A2, A3 | api-perf | 🔴/🟡 | ✅ | Haiku | S1 |
| N5 | numerics | 🟡 | ✅ | Opus | S1 |
| N3 | numerics | 🟡 | ✅ | Opus | S1 |
| H2, B2, B3, B4, B5 | test-cov/api-perf | 🔵/🟡 | ✅ | Opus | S1 |
| I1 | test-cov | 🟡 | ✅ | Opus | S2 |
| H1 | test-cov | 🔵 | ✅ | Opus | S2 |
| N7 | numerics | 🔵 | ✅ | Opus | S2 |
| H3 | test-cov | 🔵 | ✅ | Sonnet→🔍Opus | S3 |
| H4 | test-cov | 🔵 | ✅ | Sonnet→🔍Opus | S3 |
| H5 | test-cov | 🔵 | ✅ | Sonnet→🔍Opus | S3 |
| V5, V6 | input-val | 🔵 | ✅ | Sonnet→🔍Opus | S3 |
| N2 | numerics | 🟡 | ⬜ | Haiku→🔍Opus | S4 |
| thread-safety doc | thread-safety | 🟡 | ⬜ | Haiku→🔍Opus | S4 |
| M3 | math-cite | 🟡 | ⬜ | Haiku→🔍Opus | S4 |
| I5 | test-cov | 🟡 | ⬜ | Sonnet→🔍Opus | S5 |
| A4, A5 | api-perf | 🟡 | ⏸ | Opus | S6 (after G0/G1) |
| N1 | numerics | 🔴 | ✅* | — | (subsumed by B1; doc note folds into N2/S4) |
| G1–G12 | cgal | ⛔/🔴 | ⛔ | Opus | (after G0) |
| D1, D2 | dep-license | 🔴 | ⛔ | Opus | (after G0) |
| M5 | math-cite | 🔵 | 👤 | — | human domain expert |
| G0 | cgal | ⛔ | 🔄 in progress | owner | author contacted by email — awaiting reply |
* N1 (FD-step vs Newton tol) is largely subsumed by the B1 block-FD work; the remaining piece is the tolerance-coupling note, folded into N2 (Session S4).
Session sequence
✅ S1 — Audit quick-wins + numerics + refactor (DONE, 2026-05-31)
Branch fix/b1-v3-c1-quick-wins, 9 commits, 298/298 CGAL tests green.
- Sonnet: B1(HyperIdeal)/V3/C1; C2/C3, V1/V2/V4, I2/I3/I4.
- Haiku: N4/N6; M1/M2/M4; A1–A3.
- Opus: B1 inv-dist port; N5; N3; H2+B2/B3/B4/B5.
- 🔍 Opus review: validated all Sonnet/Haiku commits (value-identity, error handling, coverage gate, alias forwarding) — all correct.
✅ S2 — Solver result diagnostics (DONE, 2026-05-31, Opus impl + self-review)
Done directly in the warm post-refactor Opus session (all three findings live in
the just-written newton_core/NewtonResult/NewtonLinearSolver). Commit
90e966a; 301/301 CGAL tests green. I1 status enum + H1 iteration fix +
N7 sparse_qr_fallback_used / min_ldlt_pivot diagnostics. Follow-up done
(commit 957f506): NewtonStatus + both diagnostics propagated into the public
CGAL result types (Conformal_map_result, Hyper_ideal_map_result,
Circle_packing_result) via the CGAL::Newton_status alias.
original S2 plan
S2 — Solver result diagnostics (Sonnet → 🔍 Opus)
- I1 — give
NewtonResulta status enum (Converged/MaxIterations/LinearSolverFailed/LineSearchStalled). Thenewton_corerefactor centralised every exit point, so this lands in one place now. Update the 5 wrappers + any callers readingconverged. - H1 —
res.iterationsis stale when the solver breaks on!ok/stall; set it correctly at eachbreakinnewton_core. - N7 (Opus within this session) — optional conditioning diagnostic (smallest LDLT pivot / cheap cond estimate) surfaced in the result; pairs naturally with the I1 status enum (near-singular ⇒ a distinct status/flag).
- 🔍 Opus review: confirm the enum doesn't break the CGAL-layer result
mapping and that
convergedsemantics are unchanged for existing callers.
✅ S3 — Robustness & test-gap closure (DONE, 2026-06-01, Sonnet impl)
Branch fix/s3-robustness-gaps, 2 commits (833f9e7, 2e6c4d7), 313/313 CGAL tests green.
- H3 —
enforce_gauss_bonnetreturns|deficit|(both overloads); 3 new tests. - H4 —
ReduceToFD_ThrowsForRealAxisBoundarycoversIm(τ)==0.0exact boundary. - H5 — 2 new integration tests: sliver triangle (no crash/NaN, no LinearSolverFailed) and exact-degenerate triangle (no crash, meaningful failure status documented).
- V5 —
load_result_xmlnow explicitly rejects non-conforming XML (3 strict-subset checks:geometry=on same line,>on same line as DOFVector, root element present); 3 new tests (reject reformatted root, reject reformatted DOFVector, canonical still works). - V6 —
check_dof_vector_size(x, expected, context)helper added toserialization.hpp; 3 new tests (throws on mismatch, passes on match, message content). - PR: #45
- 🔍 Opus review: pending.
⬜ S4 — Documentation & citations (Haiku → 🔍 Opus)
- N2 —
doc/math/tolerances.md: every numerical threshold, its role, and the constraints between them (incl. the N1 FD-step ≪ Newton-tol coupling). - thread-safety — document the same-mesh concurrency contract (the code has no mutable static state; this is purely the written contract).
- M3 — re-verify the future-dated / arXiv-only citations against their
published versions; update
references.md. - 🔍 Opus review: sanity-check the tolerance relationships are stated correctly.
⬜ S5 — Coverage truthfulness (Sonnet → 🔍 Opus)
- I5 — wire
coverage.shto the full CGAL suite (not just the 26 fast tests, ~9.6%), then removeSKIP_COVERAGE_GATE=1so the agreed 80%/70%/90% gate goes live (C3 ramp-up → enforced). - 🔍 Opus review: confirm the measured numbers are honest and the gate fails correctly when coverage drops.
⏸ S6 — Public CGAL surface + packaging (Opus, only after G0/G1)
Gated by the author's reply on porting/relicensing rights.
- A4, A5 — public entry-point renames + unified
Circle_packing_result. - G1–G12, D1, D2 — license headers, CGAL package layout, concept doc, user manual, CGAL test harness, dependency/data provenance.
- G0 status: author contacted by email (2026-05-31), awaiting reply. No public release / submission / relicensing until a written answer exists.
👤 Out of model scope
- M5 — the large hand-derivations need a domain-expert prose review (the numerical results are already validated; the prose is not).
Review-gate checklist (what every 🔍 Opus session checks)
- Builds clean; full CGAL suite green (no count regression).
- No Java golden-vector / parity test perturbed (HardJava defaults intact).
- Numeric changes are value-identical where claimed, or justified + tested.
- New public surface (result types, enums, API) is intentional and documented.
- Commit message attributes the implementing model
(
Co-Authored-By: Claude <Model> <noreply@anthropic.com>). - Finding marked ✅ in the master table above with the commit ref.