Skip to content

Add bisect skill: find the breaking commit by running real behavior, not by reading diffs #12

Description

@thisguymartin

Problem

"It worked last week" is one of the most common tasks handed to an agent, and the usual agent move is to read recent diffs and guess. pstack's forensics playbooks (runtime-forensics, trace-forensics) diagnose from a live process or an artifact at one point in time. Nothing walks history with a real check at each step.

git bisect is the right tool, but it needs a reliable pass/fail predicate. Projects that ran create-verification-skill already have one. The skill that connects the two doesn't exist.

Proposed skill

Skill Use it when
bisect Something observable worked at a known-good point and doesn't now, and you want the exact commit that broke it, proven by running the behavior at each step rather than reading diffs.

Flow: define the predicate -> prove it flips between good and bad -> git bisect run in a dedicated worktree -> land on one commit -> blast-radius the fix.

  1. Predicate first. Turn the report into one executable check that exits 0 on pass, 1 on fail, 125 on "can't test this commit" (build broken, migration missing). Prefer, in order: an existing test that reproduces it; a scenario from the project's verification skill; a new minimal script. If the predicate can't be made deterministic, stop and say so. A flaky predicate makes bisect lie.
  2. Anchor. Find good and bad refs. If the user only knows "last week", test tagged releases or dated commits until one passes. Prove the predicate fails at bad and passes at good before starting; record both runs.
  3. Run. git bisect start bad good && git bisect run <predicate> in its own worktree so the primary checkout never moves. Handle 125 for commits that don't build. Log every step (commit, result, elapsed) to the decision trail.
  4. Confirm. Check out the found commit and its parent, run the predicate on both, show the flip. One commit, cited.
  5. Explain, then hand off. Read the culprit diff now that you know where to look. Say what it changed and why that breaks the predicate. Route the fix through bug-fix (with the predicate as the failing test for tdd) and run blast-radius on the culprit, since a commit that broke this may have broken siblings.

Output: culprit SHA, parent SHA, predicate script path, the two confirming runs, step log, and the suggested fix route.

Real-world use cases

1. Perf regression nobody noticed. A Go service's p95 on /search went from 40ms to 300ms sometime in the last 200 commits. Predicate: run the existing benchmark, fail if p95 > 100ms. Bisect lands on a commit that swapped a map lookup for a linear scan in a "small refactor". Reading diffs would have skipped it; the commit message said "cleanup".

2. CLI flag stopped working. mytool export --format=csv now writes JSON. Predicate: run the command, grep the first byte. Eight bisect steps, three of which hit 125 because a dependency bump broke the build mid-history. Culprit: a flag-parsing library upgrade that changed precedence. Fix goes through bug-fix with the predicate as the test.

3. Visual regression. A dashboard's chart lost its legend. Predicate: the verification skill's screenshot scenario plus an image diff against a saved baseline (visual-parity mechanics). Bisect finds a CSS module rename that stopped matching a selector.

4. Can't make it deterministic. A race that shows 1 in 30 runs. Bisect refuses to run with that predicate and says why, then points at swarm gauntlet mode to tighten the repro first. Better than a confident wrong answer.

What already exists and how this differs

  • runtime-forensics / trace-forensics: diagnose now. Bisect diagnoses when.
  • tdd: bisect produces the failing test tdd wants.
  • blast-radius: bisect calls it after finding the culprit.
  • create-verification-skill: the best source of predicates. Bisect should say when a project would benefit from one.

Guardrails

  • Never bisect in the primary checkout. Dedicated worktree only.
  • No timeout invented per step. Long builds are long. Liveness via the retained handle.
  • The predicate must be shown flipping before the run starts. No exceptions.
  • Bisect does not fix. It hands off to bug-fix.

Acceptance

  • plugins/pstack/skills/bisect/SKILL.md, shared tree; Codex mapping for worktree and background exec via codex-tools.md.
  • Predicate template (exit code contract) in the skill.
  • README row and docs/reference.md entry.
  • Bun tests, strict typecheck, static invariants, plugin validation pass.
  • Live test on the installed candidate in each affected harness: seed a scratch repo with a known breaking commit among 20+, run bisect, confirm it lands on the right SHA with both confirming runs. Record installed version, surface, action, observed result in the PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions