Skip to content

Latest commit

 

History

History
501 lines (344 loc) · 21.8 KB

File metadata and controls

501 lines (344 loc) · 21.8 KB

Loop 3: Agent-Centric (Implementation)

Page baselines live in the current checkout's .codoop-flow/ui-snapshots/; configuration defaults to .codoop-flow/codoop_flow.toml with legacy root fallback. Init inventories pages; ticket checks/reuses relevant baselines for proposals. Approval never replaces implemented UI. After verification and review, execute refreshes baselines in the isolated worktree and finish commits them with code on the ticket branch, without copying them back to the main checkout. Retain old HTML and mark unavailable captures stale with reasons. HTML does not replace the ui_capture screenshot gate. See shared rules.

Overview

Loop 3 is the third and final loop in codoop-flow's Triple-Loop Model. It is a fully automated, end-to-end ticket execution pipeline that reliably implements, verifies, reviews, and ships software tickets.

What it solves: Getting an AI coding agent to implement a ticket without hallucinating infrastructure decisions or shipping untested code. Loop 3 splits labor deliberately: intelligence (agent) writes code and reasons about design; determinism (guardrail CLI) manages git, runs tests, and archives artifacts. The only thing the agent can do wrong without breaking the pipeline is write bad code — and that's caught by review personas.

Position in pipeline: Reads Loop 2 outputs from docs/tickets/pending/<ticket_id>/ → executes in an isolated git worktree on branch dev/<ticket_id> → produces code commit + updated living docs + archives to docs/tickets/done/ (or failed/).


Quick Start

In any AI coding tool, say:

Use the codoop-execute skill to run a ticket against /path/to/repo/.codoop-flow/codoop_flow.toml

Or schedule it to run continuously:

/loop 5m run the codoop-execute skill against /path/to/repo/.codoop-flow/codoop_flow.toml

The skill picks the oldest pending ticket, builds it in an isolated worktree, runs tests, gathers review feedback, self-heals on failure, and ships the result when ready.


Workflow

Step 1 — Pick (CLI)

python3 <SKILL>/../../runtime/codoop-flow/codoop_tools.py --config <toml> pick

The CLI:

  • Checks in_progress/ first. If a ticket is already there, rebuilds its worktree (if deleted) and reports it for resume (don't start a new one)
  • If pending/ is empty, reports "no pending tickets"
  • Otherwise, picks the oldest pending ticket, creates an isolated git worktree on branch dev/<ticket_id>, and returns full ticket metadata

Output JSON:

{
  "picked": true,
  "ticket_id": "ticket_001",
  "title": "Add user search feature",
  "ticket_dir": "/abs/path/to/in_progress/ticket_001",
  "worktree": "/abs/path/to/worktrees/ticket_001",
  "branch": "dev/ticket_001",
  "modules": ["backend", "web"],
  "ui_capture": false,
  "screenshot_dir": null
}

The agent's job: Parse the JSON. If already in_progress, resume. If no pending tickets, stop. If picked is true, proceed to Build.

Step 2 — Build (Agent)

The agent reads the ticket package from ticket_dir:

  • module_prd.md — 100% business description
  • spec.md — API contract, data schema, UI interactions
  • preview.html — when present, the human-reviewed local visual flow and key interactions; use it as design context, not as production code
  • plan.md — step-by-step execution plan
  • todo.md — atomic checkbox tasks

The agent also reads project architectural boundaries from docs/tech/project-structure.md and docs/tech/tech-standards.md.

Implementation discipline: Use the /skill incremental-implementation workflow: implement one thin vertical slice, test, verify, then move to the next. Work through todo.md items in order, checking them off as you go (- [x]).

Edit-scope rule: A Scope heading in spec.md or bug_report.md is guidance: prefer it, but make the smallest necessary root-cause change outside it when a required gate demands it, and record the reason. Do not infer a hard file allowlist from Scope text.

Step 3 — Verify (independent steps + CLI)

Before editing, the agent records a baseline by running the declared validation as independent ordered steps: lint, build, focused test, and UI capture when required. A chained shell command such as lint && build && test is not a valid verification plan. A failed lint step must not skip independent build, focused-test, capture, or review evidence.

Each step records its command, stdout/stderr, exit code, and any screenshots. For every diagnostic, the agent stores a normalized fingerprint: module, step/command, file path, line (when available), rule/error code, and normalized error text. After the change it runs the same steps and compares those fingerprints with git diff --name-only:

  • New/changed diagnostics, or diagnostics in changed files, are ticket failures.
  • An exactly matching diagnostic outside changed files is a baseline_blocker. It remains fully reported but cannot spend a healing attempt, require a fix, or by itself fail the ticket.

There is no broad lint ignore or permanent allowlist. A different file, line, rule/code, normalized message, or a new diagnostic fails again.

python3 <SKILL>/../../runtime/codoop-flow/codoop_tools.py --config <toml> verify <ticket_id>

UI screenshot gate (only if ui_capture: true) checks ticket_dir/public/qa-screenshots/ for at least one file with a recognized image extension (.png, .jpg, .jpeg, .webp, .gif). It fails if none exist.

Output JSON:

{
  "ticket_id": "ticket_001",
  "ok": true,
  "reasons": []
}

Or on failure:

{
  "ticket_id": "ticket_001",
  "ok": false,
  "reasons": ["ui_capture ticket produced no screenshots"]
}

Exit code: 0 if ok: true, 1 if ok: false.

Step 4 — Self-Heal (Agent, on ticket failure)

On verify failure, the agent applies /skill debugging-and-error-recovery:

  1. Read the reported ticket failure, never a baseline blocker
  2. Fix only the root cause with a minimal change; follow Scope guidance and report a necessary exception
  3. Re-run the independent validation steps and verify

Budget: Up to max_healing_attempts retries (default 3, per ticket in metadata.json). Only ticket failures and review rejections count; baseline blockers do not. If still failing after exhausting retries → Fail Path.

When all functional gates, UI evidence, and review pass with only baseline blockers remaining, completion is allowed and the finish handoff includes baseline_warnings with the exact fingerprints and outputs. If a target configuration explicitly requires repository-wide green, record blocked_by_baseline instead of failed; keep the worktree and evidence for recovery.

Step 5 — Review (Agent, after verify passes)

The agent runs review personas from <SKILL>/../../runtime/codoop-flow/agents/ against the git diff in the worktree.

Approval is unanimous — any Critical or Important finding blocks release.

Static personas (always run — 3 people):

  1. code-reviewer — evaluates correctness, readability, architecture, security, performance. Output: APPROVE or REQUEST CHANGES (findings categorized as Critical / Important / Suggestion). Critical or Important = REJECT.

  2. security-auditor — identifies exploitable vulnerabilities. Maps findings to OWASP Top 10 and OWASP LLM Top 10. Severity: Critical / High / Medium / Low / Info. Critical or High = REJECT.

  3. test-engineer — analyzes test strategy, coverage, pyramid level, test quality. Output: Gaps identified, recommended tests by priority.

Dynamic UI/UX personas (only if ui_capture: true — 2 additional people):

  1. evidence-collector — screenshot-obsessed QA. Gets screenshot_dir to inspect rendered screens. Uses Playwright output, before/after interaction screenshots. Requires visual proof for every claim. Default assumption: 3–5 issues minimum on first implementation. Automatic fail triggers: zero-issues claims, perfect scores.

  2. reality-checker — integration and deployment readiness. Gets screenshot_dir to cross-validate QA findings. Defaults to "NEEDS WORK" unless overwhelming evidence supports READY. Checks end-to-end user journeys, spec compliance, performance (>3s load = fail).

If any reviewer rejects: Findings are fed back to the agent for a fix → re-verify → re-review (still within healing budget).

Step 6 — Experience Walkthrough (Agent, optional and non-blocking)

After technical approval, a runnable ticket with user-visible behavior may load codoop-ux-walkthrough. It passes the ticket's PRD role, goal, scope, and acceptance criteria as task context to an independently chosen persona, then saves experience_report.md in the ticket directory. The report is a qualitative product insight for human review: it never blocks release, triggers self-heal, changes code, or creates a new ticket. Pure infrastructure, refactoring, and internal-only tickets skip this step.

Step 7 — Ship Living Docs (Agent, after unanimous approval)

Before finishing, the agent syncs living documentation inside the worktree:

  • Update docs/prd/ with changed business logic
  • Update docs/tech/project-structure.md for new/moved files
  • Append concise entry to docs/tech/changelog.md

Style: second person, present tense, active voice, one concept per section, no broken code examples.

Step 8 — Finish (CLI)

python3 <SKILL>/../../runtime/codoop-flow/codoop_tools.py --config <toml> finish <ticket_id> --message "<conventional commit>"

The CLI:

  • Stages all changes excluding generated noise (__pycache__, *.pyc, node_modules, etc.)
  • Commits with the provided message (or template fallback)
  • Moves in_progress/<ticket_id> to done/<ticket_id>
  • Removes the worktree with git worktree remove --force

Output JSON:

{
  "ticket_id": "ticket_001",
  "state": "done",
  "committed": true,
  "commit": "abc1234deadbeef..."
}

Pushing is your call. The agent tells you "branch dev/<ticket_id> is ready to push; decide whether to push to origin."

Fail Path — Budget Exhausted

python3 <SKILL>/../../runtime/codoop-flow/codoop_tools.py --config <toml> fail <ticket_id> --report "<summary>"

The CLI:

  • Moves in_progress/<ticket_id> to failed/<ticket_id>
  • Writes the report to failed/<ticket_id>/healing_report.md
  • Retains the worktree, including uncommitted changes, for recovery

The ticket and report travel to failed/ for manual human intervention.


Guardrail CLI — All Subcommands

All commands require --config <toml>. All output is JSON to stdout.

status

Output: Lists all ticket names in each pipeline stage.

{
  "pending": ["ticket_a", "ticket_b"],
  "in_progress": ["ticket_c"],
  "done": ["ticket_d", "ticket_e"],
  "failed": []
}

Exit 0 always.

pick

Behavior: Picks oldest pending ticket or resumes in_progress. Creates worktree on dev/<ticket_id>.

Output: Full ticket metadata (see Step 1).

Exit 0 always.

verify <ticket_id>

Input: Ticket ID (positional arg).

Behavior: Runs the deterministic UI screenshot hard gate (if applicable). The agent runs lint/build/focused-test commands as the independent validation steps described above; their results are classified against the baseline.

Output: Success or failure with specific reasons.

Exit 0 if OK, exit 1 if any gate fails.

finish <ticket_id> --message "<msg>"

Input: Ticket ID, optional --message.

Behavior: Stages (excluding noise), commits, archives to done/, removes worktree. When --message is omitted, the fallback message is <prefix>(<module>): <title> [<id>], where <prefix> is fix for ticket_type: "fix" and feat otherwise.

Output: Commit SHA and final state.

Exit 0 on success.

fail <ticket_id> --report "<text>"

Input: Ticket ID, optional --report.

Behavior: Archives to failed/, writes healing_report.md, releases the lease, and retains the worktree (including uncommitted changes) for human recovery. The report records the worktree path and branch.

Output: Path to report.

Exit 0 on success.

resume <ticket_id>

Input: Failed ticket ID, after human approval.

Behavior: Moves failed/<ticket_id> to in_progress/, mints a new lease, and reuses the retained worktree without reset. The former healing_report.md is preserved as healing_report.previous*.md. If that worktree is gone, the CLI recreates it from the ticket branch and reports worktree_recreated:true; uncommitted recovery changes cannot be restored.

Use the returned lease token and continue the execute loop at Build with a new self-healing budget. Do not route a failed ticket through pending/.


Worktree Isolation

Each ticket executes in its own isolated git worktree — a full independent checkout on a separate branch, at a separate path.

Branch naming: dev/<ticket_id> (e.g., dev/ticket_001).

Lifecycle:

  1. On first pick: git worktree add -b dev/<ticket_id> <path> (creates branch and attaches worktree)
  2. On resume: git reset --hard HEAD to scrub any dirty state
  3. On finish: git worktree remove --force <path> (best-effort cleanup; git worktree prune reconciles refs on next run). On fail, retain the worktree for recovery.

Worktree root location: <worktree_root>/<ticket_id> (default: ~/codoop_tickets/worktrees/ticket_001). Set via codoop_flow.toml.


Self-Heal Mechanism

Trigger: a new/changed diagnostic, a diagnostic in a changed file, a failed functional/UI gate, or a review persona rejection. Exact unchanged diagnostics outside changed files are baseline blockers, not triggers.

Budget: max_healing_attempts retries (default 3, set per ticket in metadata.json).

Process per attempt:

  1. Read the ticket failure to identify the root cause
  2. Fix the root cause (not symptoms), minimal change; follow Scope guidance and report a necessary exception
  3. Re-run all independent validation steps and verify
  4. If still failing, retry (if budget remains)

Budget exhausted: Call fail with denoised summary. Ticket moves to failed/.


Review Personas

Static Group (always runs)

1. code-reviewer (runtime/codoop-flow/agents/code-reviewer.md)

  • Checks: Correctness (logic, specs), Readability (naming, structure), Architecture (patterns, abstraction), Security (no obvious vulns), Performance (efficient algorithms)
  • Verdict: APPROVE or REQUEST CHANGES
  • Blocking: Critical (must fix) and Important (should fix)
  • Non-blocking: Suggestion

2. security-auditor (runtime/codoop-flow/agents/security-auditor.md)

  • Checks: Input Handling, Auth/Authz, Data Protection, Infrastructure, Third-Party Integrations, AI/LLM Features
  • Maps to: OWASP Top 10 and OWASP LLM Top 10
  • Severity: Critical / High / Medium / Low / Info
  • Blocking: Critical and High

3. test-engineer (runtime/codoop-flow/agents/test-engineer.md)

  • Checks: Test pyramid (unit/integration/E2E), behavior vs. implementation, test quality, coverage gaps
  • Output: Coverage analysis and recommended tests by priority

Dynamic UI/UX Group (when ui_capture: true)

4. evidence-collector (runtime/codoop-flow/agents/testing-evidence-collector.md)

  • Gets screenshot_dir to inspect rendered screens
  • Runs Playwright capture, reviews test-results.json
  • Requires visual proof for every claim; no fantasy approvals
  • Default assumption: 3–5+ issues on first implementation

5. reality-checker (runtime/codoop-flow/agents/testing-reality-checker.md)

  • Gets screenshot_dir to cross-validate QA findings
  • End-to-end journey analysis, cross-device consistency, performance checks (>3s load = fail)
  • Default status: NEEDS WORK; only READY with overwhelming evidence

Unanimous approval required: ALL active personas must APPROVE / PASS. Any Critical/Important from any persona = REJECT.


UI Capture Mode

Triggered by: "ui_capture": true in metadata.json.

What it requires:

  • Delivery must place screenshots in ticket_dir/public/qa-screenshots/
  • At least one file with extension .png, .jpg, .jpeg, .webp, or .gif must exist after tests run
  • Absence of screenshots = verify fails

Extra personas: evidence-collector and reality-checker (the dynamic UI/UX group).

Screenshot archival: Screenshots travel with the ticket to done/ or failed/ for human inspection.


Configuration

codoop_flow.toml

Field Type Required Default Meaning
target_repo string (path) Yes — Absolute path to target git repo. Expanded via expanduser().resolve().
worktree_root string (path) No ~/codoop_tickets/worktrees Directory where per-ticket worktrees are created. Expanded via expanduser().
output_language string No "auto" BCP 47 language tag for user-facing replies, delivery reports, and delegated agent prose. "auto" follows the user's current language.
user_role string No "general" developer, product_manager, designer, operations, founder, or general. Affects only live conversation, not reports, code, or delegated agent prose.

The ticket pipeline directories (pending/, in_progress/, done/, failed/) are derived from <target_repo>/docs/tickets/. They are created by codoop setup but never configured in the TOML.

metadata.json (Loop 3 consumption)

Field Type Required Default Meaning
ticket_id string Yes — Unique identifier; matches directory name and branch name dev/<ticket_id>
title string Yes — Human-readable title; used in fallback commit message
ticket_type string No "feature" "feature" or "fix"; selects the fallback commit prefix (feat/fix)
modules list[string] Yes — Modules this ticket touches: backend, web, mobile, desktop
coding_engine string or null No null Informational; which AI tool handles this ticket
max_healing_attempts int No 3 Max self-heal retries; agent counts (CLI does not enforce)
ui_capture bool No false If true: activates screenshot gate and adds UI personas

Integration

Inputs (reading from Loop 2)

Loop 3 picks a ticket from docs/tickets/pending/<ticket_id>/ and consumes:

  • metadata.json — drives all scheduler decisions
  • module_prd.md + spec.md — progressively disclosed to the agent at startup
  • plan.md + todo.md — agent reads and checks off items as completed
  • public/qa-screenshots/ — (for UI tickets) created at runtime; read by review personas

Outputs (after finish)

On success (finish):

  • Committed on branch dev/<ticket_id>: all changes + living doc updates + conventional commit message
  • Archived to done/<ticket_id>/: full ticket directory + public/qa-screenshots/ (if UI ticket)
  • Commit SHA returned; branch ready for push (human decides whether to push)
  • If baseline blockers remain, the handoff includes baseline_warnings and the archived verification report retains their exact fingerprints and output

On strict-baseline block:

  • State is blocked_by_baseline, never failed; the worktree and verification evidence are retained so the ticket can be resumed after the repository debt is cleared

On failure (fail):

  • Nothing committed; worktree discarded
  • Archived to failed/<ticket_id>/: full ticket directory + newly written healing_report.md with denoised failure summary

What Gets Committed and Archived

On finish

Committed to dev/<ticket_id> branch:

  • All staged changes in the worktree (excluding generated noise)
  • Living doc updates (docs/prd/, docs/tech/)
  • Commit message: Conventional Commit format (type: scope: subject, optional body)

Archived to done/<ticket_id>/:

  • The entire ticket directory: metadata.json, module_prd.md, spec.md, plan.md, todo.md (all checkboxes ticked), and for UI tickets, public/qa-screenshots/ with all visual evidence
  • experience_report.md when an optional user experience walkthrough ran; it is advisory product input for human review, not a release verdict

On fail

Not committed: No code commit runs.

Archived to failed/<ticket_id>/:

  • The ticket directory as-is
  • Newly written healing_report.md with denoised failure summary from the agent

Skeleton Tests

The project includes tests/test_skeleton.py which exercises only the deterministic guardrails (no AI, no mocking). Each test gets a fresh temporary git repo and worktrees directory.

Key tests:

  • test_pick_moves_and_creates_worktree — pick moves ticket from pending to in_progress and creates worktree
  • test_verify_has_no_test_command_gate — legacy test_command entries are ignored by verify
  • test_ui_capture_gate — UI ticket without screenshots fails; with screenshots passes
  • test_finish_commits_and_archives — finish commits on dev/<id>, archives to done/, removes worktree
  • test_ticket_lifecycle — Human-Centric path: init → fill → validate → promote

Run with: python tests/test_skeleton.py (no pytest needed).


Key Design Principles

  • Determinism Over Cleverness — The CLI is small, fully deterministic, never guesses. The agent owns all intelligence.
  • Independent Validation + Screenshot Gate — the agent records lint, build, focused-test, and UI evidence independently; the CLI enforces the required UI screenshot gate.
  • Unanimous Approval — Review is only done when all personas agree. Any Critical/Important blocks.
  • Self-Heal Within Budget — Failures trigger retry (if budget permits), not immediate failure. Budget exhaustion moves to failed/ for human intervention.
  • Isolated Worktrees — Each ticket gets its own branch and checkout path; main repo is never touched.
  • Living Docs Kept in Sync — After code ships, living docs (docs/prd/, docs/tech/) are updated so they remain authoritative.