Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
Original file line number Diff line number Diff line change
@@ -0,0 +1,16 @@
---
schema_version: 1
id: "iss-2610080618506115"
slug: "route-work-to-model-tiers-by-role-per-task-as-well-as-per"
severity: "minor"
category: "future-work-seed"
source: "user-observation"
found_during: "the guard-refusal session of 2026-10-07"
origin: researcher-authored
production_mode: hand-written
found_at: "model routing (internal/core/oracle)"
refines: [itd-2609170822093401]
remedy: "Measure first: in a lab, or by keeping Opus live and running Haiku on the same tasks in the shadow for a later evaluation, establish which tasks the quick and default tiers handle as well as the thinking tier; then add per-task routing over role-named tiers (thinking, default, quick) that each model family translates into its own models."
---

Route work to model tiers by role, per task as well as per agent: thinking (Opus today), default (Sonnet) and quick (Haiku), so that, for example, a summary written for the product thinker goes to the quick tier rather than the thinking tier. The tier names are what abcd stores; each model family maps them to its own models, so a move to another vendor changes the mapping and nothing else. This refines the planned per-agent tier table (itd-2609170822093401), which has economy and frontier but no middle tier and no per-task routing. Before a cheaper tier takes over any task it should earn it: keep the thinking tier doing the live work and run the quick tier on the same task in the shadow, recording both outputs, so a later evaluation can show where the quick tier is good enough.
Original file line number Diff line number Diff line change
@@ -0,0 +1,25 @@
---
schema_version: 1
id: "iss-2610080620372731"
slug: "when-an-agent-or-sub-agent-loses-its-network-connection"
severity: "major"
category: "bug"
source: "user-observation"
found_during: "the guard-refusal session of 2026-10-07"
origin: researcher-authored
production_mode: hand-written
found_at: "agents and sub-agents in autonomous runs (internal/core/implement)"
remedy: "On any lost connection (the host's model calls, a sub-agent, or a tool's network call such as git push or gh), a lane keeps doing offline work and, at a network step, waits on one shared probe recorded in the run state and run log: retry at 1, 5 and 10 minutes, then hourly for up to 8 hours, then stop the run and notify the product thinker; a lane whose agent died restarts as a fresh agent from its last commit, with uncommitted edits saved aside for review."
---

When an agent or sub-agent loses its network connection mid-task, the work stalls or fails instead of waiting the outage out, and this keeps happening on the product thinker's machine. An agent that notices a lost connection should pause and retry on a widening schedule: wait one minute and try again, then five minutes, then ten, then an hour. Sibling agents should also learn of the outage without depending on each other over the network, since they may have lost the connection too: a channel on the machine itself, such as the run state and run log abcd already shares between sessions, would let one agent record the outage and the others hold off instead of each failing on its own.

## The product thinker's decisions (interview of 2026-10-09)

- **What it covers:** every kind of lost connection, because which one strikes first is not yet known: the host's own model calls stalling, a sub-agent coming back failed mid-task, and a tool's network call (git push, gh, the merge-queue call, a download) failing. The first lab run records which kinds actually occur.
- **After the hour:** retries continue hourly for up to 8 hours after the hourly stage begins, so an outage at 23:00 is still retried in the morning; then the run stops cleanly and reports what was done and what is left.
- **Who waits:** lanes carry on with offline work (editing, tests); a lane that reaches a network step waits on one shared probe recorded in the run state and run log, instead of each lane retrying on its own.
- **A lane whose agent died:** when the network returns, a fresh agent starts from the lane's last commit; uncommitted edits are saved aside for review, never built on.
- **Telling the product thinker:** a notification only when the 8-hour limit is reached and the run stops; shorter outages appear in the end-of-run report with how long each lasted and what was retried.

Proof belongs in a lab that cuts the network during a live run and shows each of the three kinds waited out, the shared probe holding the other lanes, a dead lane restarted from its last commit, and the give-up path ending in one notification.
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
---
schema_version: 1
id: "iss-2610080621450495"
slug: "consider-a-collection-of-agents-for-abcd-s-autonomous-runs"
severity: "minor"
category: "future-work-seed"
source: "user-observation"
found_during: "the guard-refusal session of 2026-10-07"
origin: researcher-authored
production_mode: hand-written
found_at: "autonomous runs: the implement loop and the drain (internal/core/implement)"
remedy: "Run a lab comparing abcd's current in-session lanes with a collection of agents, each worker an independent terminal session in a store worktree, supervised by a zero-token shell watcher that wakes the lead only on a worker event; measure the lead's context and token use, wall-clock time, and how each survives a lost connection and a restart, then decide whether the implement loop and drain adopt it."
---

Consider running abcd's autonomous runs as a collection of agents, a shape seen in an outside open-source agent tool: the person talks to one lead agent; the lead hands each task to a worker running as its own terminal session in its own git worktree; and a plain shell watcher, which costs no model tokens, sleeps on the workers and wakes the lead only when one finishes, fails or needs a decision. Today abcd's lanes run as sub-agents inside the host session, and the lead keeps checking on them itself (re-arming watchers on pull requests, polling queue state), which spends its context on waiting. Independent worker sessions would also contain a lost connection or a crash to one worker instead of the whole run, and state kept on disk lets every worker resume after a restart. The same tool can run workers on another machine over SSH, which is worth weighing later.

Two personal-agent products that large AI vendors launched in September 2026 point the same way on three counts and differ on one. Both keep working after the person leaves and come back only when something changes or needs approval, which is the watcher's job here. Both run each agent in its own isolated computer, a cloud virtual machine where this proposal uses a worktree and a terminal session. Both gate consequential actions outside the agent itself: one runs a separate gatekeeper agent on the same machine, kept apart from the worker at the system level, through which every outward action must pass; the other checks such actions against the person's rules before deciding whether they need approval. Both also show the person an activity view or audit trail of everything done and planned. They differ in shape: each is one persistent agent per person that juggles several projects, with people and agents meeting on shared pages, not a lead dispatching a collection of parallel workers, and they run in the vendor's cloud, not on the person's machine.

Three of their choices are worth carrying into the lab. A gatekeeper that sits outside the worker, as abcd's guard already does per tool call, could extend to every outward act a worker makes: a push, a pull request, a comment. An activity view of what each worker has done and plans next is the natural first page for the product thinker's dashboard. And a worker that runs away from the person's machine keeps going when that machine loses its connection, which bears on iss-2610080620372731. The broader coordination survey that iss-2608230943533581 asks for is where this comparison belongs once it runs.
Original file line number Diff line number Diff line change
@@ -0,0 +1,15 @@
---
schema_version: 1
id: "iss-2610080623391354"
slug: "consider-a-browser-review-loop-for-the-product-thinker-seen"
severity: "minor"
category: "future-work-seed"
source: "user-observation"
found_during: "the guard-refusal session of 2026-10-07"
origin: researcher-authored
production_mode: hand-written
found_at: "the product thinker's dashboard and abcd's interviews (commands/dashboard.md)"
remedy: "Prototype on the dashboard a review page for one HTML artifact with element-anchored annotations, a feedback queue the agent collects by long-polling, and revision markers; trial it on one real interview or criteria walk against the terminal questions, measuring the person's effort and the rounds needed, before deciding whether interviews hand long material to it."
---

Consider a browser review loop for the product thinker, seen in an outside open-source agent tool: the agent writes an HTML artifact (a mockup, a report, a criteria walk, an intent's press release); the person opens it in the browser, annotates exact elements or text, attaches images, and queues feedback; the agent collects the whole batch with one long-polling call and writes the next revision, with the changed parts marked as revisions. The same tool audits each render for layout faults (clipped text, unreachable controls) and lists them as issues the person can send back. Today abcd's only channel from the person back to the agent is the terminal question tool, held to twenty-four rows at eighty columns, so long material is split across many questions or pushed into an HTML report read separately, with the answers still typed in the terminal. abcd's dashboard already runs a server reachable only from the person's own devices over Tailscale, but serves nothing of the project yet, so it is the natural host for such a loop. Feedback arriving from the browser is a new path into the agent's context and must be treated as data, never instruction.
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
---
schema_version: 1
id: "iss-2610090816296964"
slug: "a-decision-s-full-options-do-not-fit-the-question-box"
severity: "minor"
category: "ux"
source: "user-observation"
found_during: "an abcd question asked in another session on 2026-10-09"
origin: researcher-authored
production_mode: hand-written
found_at: "asking the person: the question tool and the guard's question check (internal/core/question)"
remedy: "Adopt a no-repetition rule now: the text before a question says only what its options cannot (what is being decided and where it stands), never restating an option, written into the GRILL asking rules and the abcd:question-drafter agent; then run a lab with the drafter on the quick tier (Haiku 4.5) rewriting until the row count fits, compared with the default tier on the same material for fit rate, lost meaning and cost (iss-2610080618506115)."
---

When abcd asks the person to choose, the question keeps running past the host's question box: a question over 24 rows at 80 columns was refused in a live session on the 0.13.2 plugin, and the agent cut its options down to a few words each to get it through. The product thinker, reading the refused draft, saw that its long text added nothing: a paragraph before the options described each option again, while the options' own labels and descriptions were already perfectly clear. The problem was repetition, not too little room, so the fix is a rule against it rather than a second channel for the overflow; a separate text file for the full options was considered and dropped on the product thinker's ruling of 2026-10-09.

The rule: the text before a question says only what its options cannot, which is what is being decided and where it stands now, and never restates an option's meaning, gain or cost. It belongs in the GRILL asking rules and in the abcd:question-drafter agent, which v0.13.3 ships to draft questions within the limit; the guard can count rows but cannot judge repetition, so the drafter is where it is enforced. v0.13.3 also stops refusing a question whose only fault is its length (iss-2610070637562567), which removes the refusal but not the repetition.

The lab that follows runs the drafter on the quick tier (Haiku 4.5), rewriting until the row count, computed exactly by abcd, fits the window, and compares its drafts with the default tier's on the same material: how often each fits, whether any meaning is lost, and the cost. Its result feeds the model-tier idea in iss-2610080618506115.
Loading