Files
ConformalLabpp/doc/reviewer/agentic-system-design.md
Tarik Moussa 3c2aad1596
All checks were successful
C++ Tests / test-fast (pull_request) Successful in 2m16s
C++ Tests / quality-gates (pull_request) Has been skipped
C++ Tests / test-cgal (pull_request) Has been skipped
docs(roadmap): add feature-development agentic system + cross-link
Forward counterpart to the audit/remediation system: a DAG-scheduled pipeline
that works the planned phases (phases.md) and research questions
(research-track.md), with a port-vs-research router, a Theorist/Spec-author
(Opus) that derives discrete formulas + designs validation, a research
spike→go/no-go gate, an item-type-specific validation strategy (golden oracle /
invariant / convergence / cross-library), DAG prerequisite gating (chains A/B),
and reuse of the audit system's Integrator/CI/review machinery.  The two systems
compose into a loop (build forward, audit back).  Reciprocal cross-link added to
agentic-system-design.md.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
2026-06-01 00:19:07 +02:00

219 lines
13 KiB
Markdown
Raw Blame History

This file contains ambiguous Unicode characters

This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.

# Agentic system design — multi-model audit-to-code pipeline
**Purpose:** turn the static plan in [`finding-orchestration.md`](finding-orchestration.md)
into a running **multi-agent system** that works the planned sessions, interleaves
**reviews and fresh audits**, and lands findings in the code **cleanly** (atomic,
attributed, tested, parity-safe).
This design is grounded in what Session 1+2 actually did: Haiku renamed/constants,
Sonnet did validation/error-handling, **Opus** did the numerics + the **review gate**
that validated the cheaper models' work, then diagnosed a CI OOM and sequenced the
#36→#39 rebase. The system below generalises exactly that flow.
> **Companion:** the *forward* pipeline that builds the library (phases, port, research) is [`../roadmap/feature-dev-agentic-system.md`](../roadmap/feature-dev-agentic-system.md). Feature-dev lands code; this audit system hardens it — they compose into a loop.
---
## 1. Design principles (the lessons that shape the architecture)
1. **Assign by cognitive load, not severity.** A 🔴 one-line switch is Haiku/Sonnet
work; a 🟡 numerical reformulation is Opus. The dispatcher routes on *"how much
must be understood to get it right"*.
2. **Reviewer ≠ implementer, and the reviewer is the strong model.** Every change
passes an **independent Opus review gate**. Self-review misses parity/numeric
regressions; a fresh high-capability pass catches them (it caught nothing wrong
this session — which is the point: it *certified* the cheap work).
3. **The plan doc is the blackboard.** Agents are stateless sessions that
read → act → write `finding-orchestration.md`. No message bus; the shared state
is human-inspectable markdown.
4. **One finding → one commit → one test → one attribution.** Traceability and clean
revertability. Renames ship with `[[deprecated]]` aliases; refactors ship with
value-identity assertions.
5. **Gate everything irreversible to a human.** Legal/provenance (G0), public-API
stabilization (A4/A5), and the final merge are human decisions the system
*surfaces*, never makes.
6. **Separate code failures from infra failures.** A CI red is not automatically a
code bug — a dedicated diagnostic step distinguishes `cc1plus Killed` (OOM) from a
real defect before anyone touches the code.
7. **Exploit warm context.** Batch findings that touch the same code while a session
is warm (S2's I1/H1/N7 were done by Opus *immediately* after the `newton_core`
refactor instead of as a cold pickup).
---
## 2. Roles
| Role | Model | Owns | Triggered by | Output |
|---|---|---|---|---|
| **Orchestrator / Planner** | Sonnet | the backlog + dispatch | new findings, finished session | next session spec + launch prompt; updated status board |
| **Mechanic** | **Haiku** | renames, named constants, doc/citation/format fixes | dispatched mechanical findings | atomic commits + `[[deprecated]]` aliases |
| **Engineer** | **Sonnet** | tests, error handling, CI, small/safe API changes, robustness | dispatched standard findings | commits + per-finding tests |
| **Specialist** | **Opus** | numerics, math correctness, architecture refactors, public-API/irreversible calls | dispatched judgment findings | commits + tests + design notes |
| **Reviewer (Gatekeeper)** | **Opus** | the review gate after every implementation session | session marked "implemented" | APPROVE / CHANGES-REQUESTED + fixes |
| **Auditor (External reviewer)** | **Opus** | generating *new* audits with `file:line` + fix + acceptance | new module, schedule, or "re-audit" | new audit doc → new backlog findings |
| **Integrator / Release** | Sonnet | git mechanics: branch, PR, rebase/conflict, CI watch, merge sequencing | approved session | PR; rebases; merge (after human gate) |
| **Infra Diagnostician** | Sonnet→Opus | CI-failure triage (code vs runner/OOM), infra fixes | any CI red | "code vs infra" verdict + fix (e.g. `-j1`) |
| **Human (Owner)** | — | legal/provenance, API stabilization, final merge | surfaced decision | decision recorded in the plan |
> Tiering rule of thumb: **Haiku ≈ "transcribe the fix"**, **Sonnet ≈ "engineer the
> fix from a clear spec"**, **Opus ≈ "decide what the fix should be" + "judge someone
> else's fix"**.
---
## 3. Workflow (state machine, per session)
```
┌──────────────────────────────────────────────────────────┐
│ BACKLOG (audit findings, status ⬜/⏸/⛔ in the plan doc) │
└───────────────┬──────────────────────────────────────────┘
│ Orchestrator picks next ⬜, routes by load
┌────────── PLAN ──────────┐ emits a session-prompts.md block
│ model + findings + branch │────────────────────────────────────────┐
└───────────┬──────────────┘ │
▼ │
IMPLEMENT (Mechanic | Engineer | Specialist) │
· atomic commit per finding · test per finding │
· deprecated aliases / value-identity asserts │
▼ │
SELF-VERIFY (build + full test suite green locally) │
▼ │
┌──── REVIEW GATE (Opus, independent) ────┐ │
│ build · parity · value-identity · │── CHANGES-REQUESTED ─────┘
│ public-surface · attribution · plan ✅ │
└───────────┬─────────────────────────────┘ APPROVE
INTEGRATE (Integrator) → open PR
CI ──red──► INFRA DIAGNOSTICIAN ──code?──► back to IMPLEMENT
│ └─infra?─► fix infra (e.g. -j1), re-run
green
HUMAN MERGE GATE ──► MERGE ──► Orchestrator marks ✅, updates plan
(every N sessions / per module) AUDITOR re-audit ──► new BACKLOG findings
```
Key edges proven this session: the **CHANGES-REQUESTED loop** (none needed — Sonnet/Haiku
work was clean), the **INFRA branch** (the `cc1plus` OOM → `-j1`, *not* a code change),
and the **multi-PR sequencing** at INTEGRATE (#36 merged first, #39 rebased to drop the
duplicate A1A3).
---
## 4. Shared state & artifacts (the blackboard)
- **`finding-orchestration.md`** — backlog + status board + per-finding model/session.
The single source of truth every agent reads/writes.
- **Audit docs** (`*-audit-*.md`) — the finding *specs*: `file:line`, fix, acceptance
criteria. The Auditor writes them; implementers consume them.
- **`session-prompts.md`** — launch templates the Orchestrator fills in per session.
- **PRs + commits** — the audit trail; commit trailer `Co-Authored-By: Claude <Model>`
records *which* model did *which* finding.
- **Decision log** (a section in the plan) — human gates: G0 status ("authors emailed,
awaiting reply"), API-stabilization decisions, merge approvals.
Why blackboard over a chat-bus: sessions are short-lived and stateless; the durable
state is the markdown; humans can read and override it at any point.
---
## 5. Gates & guardrails
**Automated gates (must pass to advance):**
- Build clean; **full test suite green, no count regression**.
- **No parity/golden-vector test perturbed** (HardJava-style defaults intact).
- New public surface (result types, enums, file formats) is intentional + documented.
- Each finding has a commit + a test; renames carry `[[deprecated]]` aliases.
**Human gates (system surfaces, human decides):**
- **Legal/provenance (G0)** — porting/relicensing rights. Blocks G1G12, D1/D2, A4/A5.
- **Public-API stabilization** — once in a released/CGAL API, renames need a deprecation
cycle; lock names only when the owner says so.
- **Final merge** — the agent prepares a green, mergeable PR; the human clicks merge
(or explicitly delegates it).
**Blocked-dependency tracking:** findings gated by an open human decision are marked ⛔
in the plan and never dispatched (S6 stays blocked until G0 resolves).
---
## 6. How phases, reviews and audits interleave (cadence)
- **Per session:** implement → **review gate** (always). The review is *inline* with
the cadence, not a separate phase you forget.
- **Per PR:** CI + infra-triage + human merge gate.
- **Per module / every N sessions / on request:** an **Auditor** pass produces a fresh
audit, refilling the backlog. This is the loop that keeps the system honest as the
code grows — new code gets audited, new findings get planned, the cycle repeats.
- **Continuous:** the Orchestrator keeps the plan's status board current so any human
glance shows "done / in-flight / blocked".
So the rhythm is: **audit → plan → (implement → review)\* → integrate → merge → re-audit.**
---
## 7. How to actually run it with Claude
Three implementation levels, cheapest first:
1. **Manual sessions (today).** The Orchestrator (you, or a Sonnet session) hands a
`session-prompts.md` block to a fresh session set to the named model. The review
gate is a second session set to Opus. This already works — it is exactly Session 1+2.
2. **Claude Code subagents.** Define each role as a subagent with a pinned model and a
tight tool allow-list (Mechanic: edit+bash; Reviewer: read+bash, no write-to-main;
Integrator: bash/git/PR API). A top-level "Orchestrator" session dispatches via the
Task tool and reconciles the plan doc.
3. **Claude Agent SDK (autonomous).** A supervisor process loops the state machine:
reads the plan, spawns role-agents (model per role), runs the gates as code
(build/test/CI checks), and stops at human gates. The blackboard (the markdown +
git) is the durable state across runs.
Pin models per role explicitly; do **not** let a cheap role silently escalate — that is
what the review gate and the dispatcher's load-routing are for.
---
## 8. Cost & efficiency strategy
- **Route down aggressively, gate up always.** Most findings go to Haiku/Sonnet; only
the Opus review gate and genuine-judgment findings pay for Opus.
- **Batch by warm context.** Group findings touching the same files into one session
(S2's three findings shared `newton_core`); avoid cold re-derivation.
- **Make the reviewer cheap to satisfy.** Atomic commits + per-finding tests + a fixed
checklist make the Opus review fast (it reads a small, well-scoped diff).
- **Fail fast on infra.** The diagnostician prevents wasted Opus cycles "fixing" a
phantom code bug that was really an OOM.
---
## 9. Worked example (this session, mapped to the roles)
| What happened | Role(s) |
|---|---|
| A1A3 renames, N4/N6 constants, M-citations | **Mechanic (Haiku)** |
| V1V4 error handling, C2/C3 coverage gate, I-tests | **Engineer (Sonnet)** |
| B1 inv-dist block-FD port, N3/N5, `newton_core` (H2+B2/B4/B5), S2 (I1/H1/N7) | **Specialist (Opus)** |
| Validated the Haiku/Sonnet commits (value-identity, alias forwarding, coverage logic) | **Reviewer (Opus)** |
| Built the plan, model assignment, session prompts | **Orchestrator** |
| `#36` merge → rebase `#39` (drop dup A1A3) → PRs | **Integrator** |
| `test-fast` red → proved `cc1plus` OOM via a pure-main probe → `-j1` fix | **Infra Diagnostician** |
| G0 porting rights: "authors emailed, awaiting reply"; merge approval | **Human gate** |
The audits themselves (the 11 docs) were a prior **Auditor (Opus)** pass — the input the
whole pipeline consumes.
---
## 10. Failure modes the design absorbs
- **Cheap model gets it subtly wrong** → caught by the Opus review gate.
- **Refactor silently changes numerics** → caught by parity tests + value-identity asserts.
- **CI red panic** → diagnostician separates infra (OOM) from code before any "fix".
- **Two PRs overlap** (#36/#39) → Integrator sequences merge + rebase deterministically.
- **A finding needs a human/legal call** → marked ⛔, never auto-dispatched.
- **Backlog goes stale as code grows** → periodic Auditor re-audit refills it.
```