Skip to content

Add diff-behavior skill: run the same scenarios on trunk and head, report every observable difference #15

Description

@thisguymartin

Problem

A PR's diff shows what changed in the code. What a reviewer actually needs is what changed in behavior, and whether all of it was intended. multi-phase-plan already requires this for its PRs (the "regression lane against trunk": same scenario on trunk and head, dual-sided perf gate). But it lives inside one playbook, so a plain feature or bug-fix PR, a migrate run, or a dependency bump never gets it.

blast-radius is close but different: it hunts for one thing outside the diff that could break and proves it safe. Diff-behavior is the systematic version: drive the whole load-bearing surface on both builds and list what differs.

Proposed skill

Skill Use it when
diff-behavior You want to know what a change did from the outside: run the same scenarios against trunk and head, list every observable difference (output, UI, logs, timing, errors), and flag the ones the PR didn't claim.

Flow: two builds -> one scenario set -> run both -> normalize -> diff -> classify intended / unintended / noise.

  1. Two builds. Trunk (or a named base) and head, each in its own worktree, built the same way. Same env, same seed data, same config. Never the primary checkout.
  2. Scenario set. From the project's verification skill's feature map when it exists; otherwise from how on the touched surface plus the PR's own claims. Include the load-bearing scenarios even if the diff looks unrelated (that's the point). Each scenario names what it captures: stdout, HTTP response, screenshot, log lines, a metric, exit code.
  3. Run both. swarm partition: one worker per scenario, each runs trunk then head, captures both. Interleave for timing scenarios so drift affects both sides equally (same rule as perf-issue). Same seed, same clock where the app allows.
  4. Normalize and diff. Strip timestamps, ids, ordering that the app doesn't guarantee. Diff what's left. Screenshots via image diff (visual-parity mechanics). Metrics as trunk baseline -> head value -> ratio.
  5. Classify. For each difference: intended (matches a claim in the PR description or brief), unintended (no claim covers it), or noise (normalizer gap; fix the normalizer and rerun, don't hand-wave). The unintended list is the deliverable.
  6. Report. Table: scenario -> trunk -> head -> difference -> class -> evidence path. Post to the PR when asked. Zero unintended differences is a real result and says so.

Output: the report, the scenario set (reusable), captured artifacts for both sides, receipts.

Real-world use cases

1. "Small" dependency bump. A Go HTTP router minor version. Diff-behavior runs the 15 API scenarios on both. 14 identical. One: a trailing-slash route now returns 301 instead of 200. Nobody's tests covered it; the PR description said "no behavior change". Classified unintended, the PR gets a fix before merge.

2. Bug fix that fixes too much. PR fixes CSV export timezone. Diff-behavior shows the CSV scenario changed as intended, and the JSON export scenario also changed (same shared formatter). The author didn't know the formatter was shared. Intended or not is now the author's explicit call, with evidence.

3. Migration proof. After a migrate run swapping moment for date-fns, diff-behavior over every date-rendering scenario shows three differences: two intended (locale format), one unintended (a null date now renders "Invalid Date" instead of blank). Found before a user did.

4. Perf claim. PR says "30% faster search". Diff-behavior's timing scenario, interleaved, 20 runs each: trunk p95 210ms, head p95 160ms, ratio 0.76. Claim confirmed with numbers a reviewer can rerun. If head lacks a baseline (new feature), it says so instead of inventing a ratio, per multi-phase-plan's rule.

5. Overnight gate. nightwatch / autopilot-stack can call diff-behavior as part of the STACK-READY verdict so the morning report shows behavior differences per PR, not just "tests green".

What already exists and how this differs

  • multi-phase-plan regression lane and dual-sided perf gate: same idea, locked inside one playbook. Diff-behavior extracts it so any PR can use it; multi-phase-plan should then reference it.
  • blast-radius: finds the one fact outside the diff. Diff-behavior enumerates the surface.
  • visual-parity: pixel diffs for a styling migration. Diff-behavior reuses its image-diff mechanics as one capture type.
  • swarm: the fan-out primitive.
  • create-verification-skill: the best source of scenarios; diff-behavior should say when a project lacks one.

Guardrails

  • Both sides always run. A head-only run is not a diff.
  • Noise is fixed in the normalizer, never waved off. A rerun after a normalizer fix is the rule.
  • No claim of a ratio between unlike scenarios.
  • Workers in worktrees; no implicit timeout; dropouts named by provider/model/receipt, never substituted.

Acceptance

  • plugins/pstack/skills/diff-behavior/SKILL.md, shared tree, with capture types, normalizer rules, and the report table format.
  • multi-phase-plan.md regression-lane text references the skill instead of restating it (small edit; re-ground against upstream per UPSTREAM.md).
  • README row and docs/reference.md entry.
  • Bun tests, strict typecheck, static invariants, plugin validation pass.
  • Live test on the installed candidate in each affected harness: scratch repo with a PR that changes one scenario intentionally and one by accident; diff-behavior must classify both correctly with artifacts for each side. Record installed version, surface, action, observed result in the PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions