Skip to content

Add nightwatch skill: work a batch of issues unattended and wake up to verified PRs #11

Description

@thisguymartin

Problem

pstack has the pieces for unattended work but not the whole loop. autonomous-run drives one task to its exit condition. autopilot-full / autopilot-stack run a program of PRs once someone has written the briefs. pause-safely leaves a checkpoint. What's missing is the operator-facing entry: "here are N issues, I'm going to bed, have PRs and a trail ready by morning".

Today that means hand-writing N briefs, starting an autopilot, and hoping the trail is readable. Each piece exists; the glue doesn't.

Depends on intake (#10) for the issue -> brief step.

Proposed skill

Skill Use it when
nightwatch You have a list of issues and a block of unattended time, and want each one worked in its own worktree, verified, and handed back as a reviewable PR stack plus one trail you can read over coffee.

Flow: issues -> intake (brief + classify + worktree) -> runnable / needs-human split -> autopilot-stack over the runnable set -> swarm verify each head -> morning report.

  1. Intake pass. Run intake on every issue. Split into runnable unattended (clear scope, observable exit condition, verification plan possible) and needs a human (open product questions). Post the questions on the parked issues so they're answered by morning.
  2. Capacity and cost check. Count lanes per configured role, estimate token spend from the brief sizes, and state it before starting. Nightwatch never changes the model sheet or picks cheaper models on its own (AGENTS.md: no weaker-model fallback). If the estimate is over what the operator said they'd spend, park the lowest-priority issues and say so.
  3. Run. Hand the runnable briefs to autopilot-stack (default) or autopilot-full when the operator granted landing authority up front. One owner per PR, one worktree per owner, root owns topology.
  4. Verify. Each STACK-READY head gets the swarm verdict per autopilot-stack step 4: unit, live, perf boxes. Nothing is reported as done without a live check.
  5. Wake chain. Audit ticks via /loop dynamic mode as autopilot already does. No sleep, no invented deadline.
  6. Morning report. One file: a table of issues -> PR -> verdict -> what to look at first, the parked issues with their questions, dropouts by provider/model/receipt, and total elapsed/tokens/cost from receipts. Trail per show-me-your-work.

Real-world use cases

1. The Friday night batch. Six small issues on a TypeScript service: two bug reports with repros, a dependency bump, a missing index, a log-format change, a flaky test. /pstack:nightwatch 51 52 53 54 55 56. Intake parks ericlitman#54 (the index) because the issue doesn't say which query it's for, posts the question, and runs the other five. Morning: five PRs in a stack, each with a verifier verdict, the flaky-test PR says BLOCKED: reproduces 1/40 runs, cause not proven rather than a guess, and ericlitman#54 has a comment waiting for an answer.

2. Post-incident cleanup. After an outage, the team files eight follow-ups (add timeout here, add metric there, delete dead flag). They're all mechanical but nobody wants to spend a day on them. Nightwatch runs them as a stack; the on-call reviews the stack Monday and lands it bottom-to-top.

3. Cost-capped run on gateway lanes. A solo dev with only DeepSeek + MiniMax keys says "spend at most $3 tonight". Nightwatch estimates from brief sizes, runs the four issues that fit, parks the fifth with the estimate attached, and the morning report shows actual spend from receipts multiplied by the LANES.md price table.

4. Something breaks at 3am. The Grok CLI loses auth halfway through. The affected lane drops out with an unauthenticated receipt. Nightwatch proceeds N-1, never substitutes a model, and the report names the lane and the receipt path.

What already exists and how this differs

  • autopilot-stack / autopilot-full: the engine. Nightwatch is the operator wrapper that feeds them briefs and produces the report.
  • autonomous-run: one task. Nightwatch is a batch.
  • orchestrate: a whole project handed to one lead; assumes briefs. Nightwatch starts from issues.
  • pause-safely: nightwatch uses it for the parked / partially-done state so a cold-start agent can resume.

Guardrails

  • No implicit timeout. Audit ticks observe; they never cancel a healthy lane.
  • No weaker-model fallback. A budget cap parks work; it never downgrades a lane.
  • Never merges unless autopilot-full was chosen and landing authority was explicitly granted.
  • Product questions go back to the issue as comments, not guessed.
  • Parent resolves every route once; workers never detect the harness.

Acceptance

  • plugins/pstack/skills/nightwatch/SKILL.md, shared tree, with the morning-report format.
  • README row and docs/reference.md entry; USAGE.md gets a short "overnight" example.
  • Bun tests, strict typecheck, static invariants, plugin validation pass.
  • Live test: a real batch of at least two issues in a scratch repo on the installed candidate in each affected harness, morning report produced, at least one parked issue with a posted question. Record installed version, surface, action, observed result in the PR.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions