You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
A PR's diff shows what changed in the code. What a reviewer actually needs is what changed in behavior, and whether all of it was intended. multi-phase-plan already requires this for its PRs (the "regression lane against trunk": same scenario on trunk and head, dual-sided perf gate). But it lives inside one playbook, so a plain feature or bug-fix PR, a migrate run, or a dependency bump never gets it.
blast-radius is close but different: it hunts for one thing outside the diff that could break and proves it safe. Diff-behavior is the systematic version: drive the whole load-bearing surface on both builds and list what differs.
Proposed skill
Skill
Use it when
diff-behavior
You want to know what a change did from the outside: run the same scenarios against trunk and head, list every observable difference (output, UI, logs, timing, errors), and flag the ones the PR didn't claim.
Flow: two builds -> one scenario set -> run both -> normalize -> diff -> classify intended / unintended / noise.
Two builds. Trunk (or a named base) and head, each in its own worktree, built the same way. Same env, same seed data, same config. Never the primary checkout.
Scenario set. From the project's verification skill's feature map when it exists; otherwise from how on the touched surface plus the PR's own claims. Include the load-bearing scenarios even if the diff looks unrelated (that's the point). Each scenario names what it captures: stdout, HTTP response, screenshot, log lines, a metric, exit code.
Run both.swarm partition: one worker per scenario, each runs trunk then head, captures both. Interleave for timing scenarios so drift affects both sides equally (same rule as perf-issue). Same seed, same clock where the app allows.
Normalize and diff. Strip timestamps, ids, ordering that the app doesn't guarantee. Diff what's left. Screenshots via image diff (visual-parity mechanics). Metrics as trunk baseline -> head value -> ratio.
Classify. For each difference: intended (matches a claim in the PR description or brief), unintended (no claim covers it), or noise (normalizer gap; fix the normalizer and rerun, don't hand-wave). The unintended list is the deliverable.
Report. Table: scenario -> trunk -> head -> difference -> class -> evidence path. Post to the PR when asked. Zero unintended differences is a real result and says so.
Output: the report, the scenario set (reusable), captured artifacts for both sides, receipts.
Real-world use cases
1. "Small" dependency bump. A Go HTTP router minor version. Diff-behavior runs the 15 API scenarios on both. 14 identical. One: a trailing-slash route now returns 301 instead of 200. Nobody's tests covered it; the PR description said "no behavior change". Classified unintended, the PR gets a fix before merge.
2. Bug fix that fixes too much. PR fixes CSV export timezone. Diff-behavior shows the CSV scenario changed as intended, and the JSON export scenario also changed (same shared formatter). The author didn't know the formatter was shared. Intended or not is now the author's explicit call, with evidence.
3. Migration proof. After a migrate run swapping moment for date-fns, diff-behavior over every date-rendering scenario shows three differences: two intended (locale format), one unintended (a null date now renders "Invalid Date" instead of blank). Found before a user did.
4. Perf claim. PR says "30% faster search". Diff-behavior's timing scenario, interleaved, 20 runs each: trunk p95 210ms, head p95 160ms, ratio 0.76. Claim confirmed with numbers a reviewer can rerun. If head lacks a baseline (new feature), it says so instead of inventing a ratio, per multi-phase-plan's rule.
5. Overnight gate.nightwatch / autopilot-stack can call diff-behavior as part of the STACK-READY verdict so the morning report shows behavior differences per PR, not just "tests green".
What already exists and how this differs
multi-phase-plan regression lane and dual-sided perf gate: same idea, locked inside one playbook. Diff-behavior extracts it so any PR can use it; multi-phase-plan should then reference it.
blast-radius: finds the one fact outside the diff. Diff-behavior enumerates the surface.
visual-parity: pixel diffs for a styling migration. Diff-behavior reuses its image-diff mechanics as one capture type.
swarm: the fan-out primitive.
create-verification-skill: the best source of scenarios; diff-behavior should say when a project lacks one.
Guardrails
Both sides always run. A head-only run is not a diff.
Noise is fixed in the normalizer, never waved off. A rerun after a normalizer fix is the rule.
No claim of a ratio between unlike scenarios.
Workers in worktrees; no implicit timeout; dropouts named by provider/model/receipt, never substituted.
Acceptance
plugins/pstack/skills/diff-behavior/SKILL.md, shared tree, with capture types, normalizer rules, and the report table format.
multi-phase-plan.md regression-lane text references the skill instead of restating it (small edit; re-ground against upstream per UPSTREAM.md).
README row and docs/reference.md entry.
Bun tests, strict typecheck, static invariants, plugin validation pass.
Live test on the installed candidate in each affected harness: scratch repo with a PR that changes one scenario intentionally and one by accident; diff-behavior must classify both correctly with artifacts for each side. Record installed version, surface, action, observed result in the PR.
Problem
A PR's diff shows what changed in the code. What a reviewer actually needs is what changed in behavior, and whether all of it was intended.
multi-phase-planalready requires this for its PRs (the "regression lane against trunk": same scenario on trunk and head, dual-sided perf gate). But it lives inside one playbook, so a plain feature or bug-fix PR, amigraterun, or a dependency bump never gets it.blast-radiusis close but different: it hunts for one thing outside the diff that could break and proves it safe. Diff-behavior is the systematic version: drive the whole load-bearing surface on both builds and list what differs.Proposed skill
diff-behaviorFlow: two builds -> one scenario set -> run both -> normalize -> diff -> classify intended / unintended / noise.
howon the touched surface plus the PR's own claims. Include the load-bearing scenarios even if the diff looks unrelated (that's the point). Each scenario names what it captures: stdout, HTTP response, screenshot, log lines, a metric, exit code.swarmpartition: one worker per scenario, each runs trunk then head, captures both. Interleave for timing scenarios so drift affects both sides equally (same rule asperf-issue). Same seed, same clock where the app allows.visual-paritymechanics). Metrics as trunk baseline -> head value -> ratio.Output: the report, the scenario set (reusable), captured artifacts for both sides, receipts.
Real-world use cases
1. "Small" dependency bump. A Go HTTP router minor version. Diff-behavior runs the 15 API scenarios on both. 14 identical. One: a trailing-slash route now returns 301 instead of 200. Nobody's tests covered it; the PR description said "no behavior change". Classified unintended, the PR gets a fix before merge.
2. Bug fix that fixes too much. PR fixes CSV export timezone. Diff-behavior shows the CSV scenario changed as intended, and the JSON export scenario also changed (same shared formatter). The author didn't know the formatter was shared. Intended or not is now the author's explicit call, with evidence.
3. Migration proof. After a
migraterun swappingmomentfordate-fns, diff-behavior over every date-rendering scenario shows three differences: two intended (locale format), one unintended (a null date now renders "Invalid Date" instead of blank). Found before a user did.4. Perf claim. PR says "30% faster search". Diff-behavior's timing scenario, interleaved, 20 runs each: trunk p95 210ms, head p95 160ms, ratio 0.76. Claim confirmed with numbers a reviewer can rerun. If head lacks a baseline (new feature), it says so instead of inventing a ratio, per
multi-phase-plan's rule.5. Overnight gate.
nightwatch/autopilot-stackcan call diff-behavior as part of the STACK-READY verdict so the morning report shows behavior differences per PR, not just "tests green".What already exists and how this differs
multi-phase-planregression lane and dual-sided perf gate: same idea, locked inside one playbook. Diff-behavior extracts it so any PR can use it; multi-phase-plan should then reference it.blast-radius: finds the one fact outside the diff. Diff-behavior enumerates the surface.visual-parity: pixel diffs for a styling migration. Diff-behavior reuses its image-diff mechanics as one capture type.swarm: the fan-out primitive.create-verification-skill: the best source of scenarios; diff-behavior should say when a project lacks one.Guardrails
Acceptance
plugins/pstack/skills/diff-behavior/SKILL.md, shared tree, with capture types, normalizer rules, and the report table format.multi-phase-plan.mdregression-lane text references the skill instead of restating it (small edit; re-ground against upstream perUPSTREAM.md).docs/reference.mdentry.