From 71c438fbaf11de81f0317b0353ce3258bad9c80c Mon Sep 17 00:00:00 2001 From: REPPL <77722411+REPPL@users.noreply.github.com> Date: Thu, 8 Oct 2026 07:19:54 +0100 Subject: [PATCH 1/6] chore: capture role-named model tiers per task, with a shadow evaluation A future-work seed refining the planned per-agent tier table: thinking, default and quick tiers that each model family maps to its own models, routed per task, with a cheaper tier earning a task by running in the shadow of the live one first. Refs: iss-2610080618506115 Assisted-by: Claude:claude-opus-5-5 --- ...odel-tiers-by-role-per-task-as-well-as-per.md | 16 ++++++++++++++++ 1 file changed, 16 insertions(+) create mode 100644 .abcd/work/issues/open/iss-2610080618506115-route-work-to-model-tiers-by-role-per-task-as-well-as-per.md diff --git a/.abcd/work/issues/open/iss-2610080618506115-route-work-to-model-tiers-by-role-per-task-as-well-as-per.md b/.abcd/work/issues/open/iss-2610080618506115-route-work-to-model-tiers-by-role-per-task-as-well-as-per.md new file mode 100644 index 000000000..9cbebdb5e --- /dev/null +++ b/.abcd/work/issues/open/iss-2610080618506115-route-work-to-model-tiers-by-role-per-task-as-well-as-per.md @@ -0,0 +1,16 @@ +--- +schema_version: 1 +id: "iss-2610080618506115" +slug: "route-work-to-model-tiers-by-role-per-task-as-well-as-per" +severity: "minor" +category: "future-work-seed" +source: "user-observation" +found_during: "the guard-refusal session of 2026-10-07" +origin: researcher-authored +production_mode: hand-written +found_at: "model routing (internal/core/oracle)" +refines: [itd-2609170822093401] +remedy: "Measure first: in a lab, or by keeping Opus live and running Haiku on the same tasks in the shadow for a later evaluation, establish which tasks the quick and default tiers handle as well as the thinking tier; then add per-task routing over role-named tiers (thinking, default, quick) that each model family translates into its own models." +--- + +Route work to model tiers by role, per task as well as per agent: thinking (Opus today), default (Sonnet) and quick (Haiku), so that, for example, a summary written for the product thinker goes to the quick tier rather than the thinking tier. The tier names are what abcd stores; each model family maps them to its own models, so a move to another vendor changes the mapping and nothing else. This refines the planned per-agent tier table (itd-2609170822093401), which has economy and frontier but no middle tier and no per-task routing. Before a cheaper tier takes over any task it should earn it: keep the thinking tier doing the live work and run the quick tier on the same task in the shadow, recording both outputs, so a later evaluation can show where the quick tier is good enough. From c232a079e19c7ff2341ebf8d8875605b73ed6276 Mon Sep 17 00:00:00 2001 From: REPPL <77722411+REPPL@users.noreply.github.com> Date: Thu, 8 Oct 2026 07:20:47 +0100 Subject: [PATCH 2/6] chore: capture retrying a lost connection on a widening schedule Agents and sub-agents that lose the network mid-task stall or fail; the capture asks for retries at 1, 5, 10 and 60 minutes and for the outage to be shared through a machine-local channel rather than over the network. Refs: iss-2610080620372731 Assisted-by: Claude:claude-opus-5-5 --- ...t-or-sub-agent-loses-its-network-connection.md | 15 +++++++++++++++ 1 file changed, 15 insertions(+) create mode 100644 .abcd/work/issues/open/iss-2610080620372731-when-an-agent-or-sub-agent-loses-its-network-connection.md diff --git a/.abcd/work/issues/open/iss-2610080620372731-when-an-agent-or-sub-agent-loses-its-network-connection.md b/.abcd/work/issues/open/iss-2610080620372731-when-an-agent-or-sub-agent-loses-its-network-connection.md new file mode 100644 index 000000000..1d30f055a --- /dev/null +++ b/.abcd/work/issues/open/iss-2610080620372731-when-an-agent-or-sub-agent-loses-its-network-connection.md @@ -0,0 +1,15 @@ +--- +schema_version: 1 +id: "iss-2610080620372731" +slug: "when-an-agent-or-sub-agent-loses-its-network-connection" +severity: "major" +category: "bug" +source: "user-observation" +found_during: "the guard-refusal session of 2026-10-07" +origin: researcher-authored +production_mode: hand-written +found_at: "agents and sub-agents in autonomous runs (internal/core/implement)" +remedy: "Retry a lost connection on a widening schedule (wait 1 minute, retry; 5 minutes, retry; 10 minutes, retry; then an hour), and record the outage and each retry in a machine-local channel, such as the shared run state and run log, so sibling agents and sub-agents see it and hold off without reaching each other over the network." +--- + +When an agent or sub-agent loses its network connection mid-task, the work stalls or fails instead of waiting the outage out, and this keeps happening on the product thinker's machine. An agent that notices a lost connection should pause and retry on a widening schedule: wait one minute and try again, then five minutes, then ten, then an hour. Sibling agents should also learn of the outage without depending on each other over the network, since they may have lost the connection too: a channel on the machine itself, such as the run state and run log abcd already shares between sessions, would let one agent record the outage and the others hold off instead of each failing on its own. From 893878dac793eb76d7c0de44170e3244b3b8a1d2 Mon Sep 17 00:00:00 2001 From: REPPL <77722411+REPPL@users.noreply.github.com> Date: Thu, 8 Oct 2026 07:21:55 +0100 Subject: [PATCH 3/6] chore: capture a collection of agents with a zero-token watcher A future-work seed: workers as independent terminal sessions in their own worktrees, supervised by a shell watcher that wakes the lead only on a worker event, weighed against the in-session lanes in a lab first. Refs: iss-2610080621450495 Assisted-by: Claude:claude-opus-5-5 --- ...on-of-agents-for-abcd-s-autonomous-runs.md | 19 +++++++++++++++++++ 1 file changed, 19 insertions(+) create mode 100644 .abcd/work/issues/open/iss-2610080621450495-consider-a-collection-of-agents-for-abcd-s-autonomous-runs.md diff --git a/.abcd/work/issues/open/iss-2610080621450495-consider-a-collection-of-agents-for-abcd-s-autonomous-runs.md b/.abcd/work/issues/open/iss-2610080621450495-consider-a-collection-of-agents-for-abcd-s-autonomous-runs.md new file mode 100644 index 000000000..945473925 --- /dev/null +++ b/.abcd/work/issues/open/iss-2610080621450495-consider-a-collection-of-agents-for-abcd-s-autonomous-runs.md @@ -0,0 +1,19 @@ +--- +schema_version: 1 +id: "iss-2610080621450495" +slug: "consider-a-collection-of-agents-for-abcd-s-autonomous-runs" +severity: "minor" +category: "future-work-seed" +source: "user-observation" +found_during: "the guard-refusal session of 2026-10-07" +origin: researcher-authored +production_mode: hand-written +found_at: "autonomous runs: the implement loop and the drain (internal/core/implement)" +remedy: "Run a lab comparing abcd's current in-session lanes with a collection of agents, each worker an independent terminal session in a store worktree, supervised by a zero-token shell watcher that wakes the lead only on a worker event; measure the lead's context and token use, wall-clock time, and how each survives a lost connection and a restart, then decide whether the implement loop and drain adopt it." +--- + +Consider running abcd's autonomous runs as a collection of agents, a shape seen in an outside open-source agent tool: the person talks to one lead agent; the lead hands each task to a worker running as its own terminal session in its own git worktree; and a plain shell watcher, which costs no model tokens, sleeps on the workers and wakes the lead only when one finishes, fails or needs a decision. Today abcd's lanes run as sub-agents inside the host session, and the lead keeps checking on them itself (re-arming watchers on pull requests, polling queue state), which spends its context on waiting. Independent worker sessions would also contain a lost connection or a crash to one worker instead of the whole run, and state kept on disk lets every worker resume after a restart. The same tool can run workers on another machine over SSH, which is worth weighing later. + +Two personal-agent products that large AI vendors launched in September 2026 point the same way on three counts and differ on one. Both keep working after the person leaves and come back only when something changes or needs approval, which is the watcher's job here. Both run each agent in its own isolated computer, a cloud virtual machine where this proposal uses a worktree and a terminal session. Both gate consequential actions outside the agent itself: one runs a separate gatekeeper agent on the same machine, kept apart from the worker at the system level, through which every outward action must pass; the other checks such actions against the person's rules before deciding whether they need approval. Both also show the person an activity view or audit trail of everything done and planned. They differ in shape: each is one persistent agent per person that juggles several projects, with people and agents meeting on shared pages, not a lead dispatching a collection of parallel workers, and they run in the vendor's cloud, not on the person's machine. + +Three of their choices are worth carrying into the lab. A gatekeeper that sits outside the worker, as abcd's guard already does per tool call, could extend to every outward act a worker makes: a push, a pull request, a comment. An activity view of what each worker has done and plans next is the natural first page for the product thinker's dashboard. And a worker that runs away from the person's machine keeps going when that machine loses its connection, which bears on iss-2610080620372731. The broader coordination survey that iss-2608230943533581 asks for is where this comparison belongs once it runs. From cb2d8833c5954ea3a8ae1439e8919c1de6c31786 Mon Sep 17 00:00:00 2001 From: REPPL <77722411+REPPL@users.noreply.github.com> Date: Thu, 8 Oct 2026 07:23:48 +0100 Subject: [PATCH 4/6] chore: capture a browser review loop for the product thinker A future-work seed: HTML artifacts reviewed in the browser with element-anchored annotations the agent collects in one batch, hosted on the dashboard, trialled against the terminal questions first. Refs: iss-2610080623391354 Assisted-by: Claude:claude-opus-5-5 --- ...er-review-loop-for-the-product-thinker-seen.md | 15 +++++++++++++++ 1 file changed, 15 insertions(+) create mode 100644 .abcd/work/issues/open/iss-2610080623391354-consider-a-browser-review-loop-for-the-product-thinker-seen.md diff --git a/.abcd/work/issues/open/iss-2610080623391354-consider-a-browser-review-loop-for-the-product-thinker-seen.md b/.abcd/work/issues/open/iss-2610080623391354-consider-a-browser-review-loop-for-the-product-thinker-seen.md new file mode 100644 index 000000000..1b56d173e --- /dev/null +++ b/.abcd/work/issues/open/iss-2610080623391354-consider-a-browser-review-loop-for-the-product-thinker-seen.md @@ -0,0 +1,15 @@ +--- +schema_version: 1 +id: "iss-2610080623391354" +slug: "consider-a-browser-review-loop-for-the-product-thinker-seen" +severity: "minor" +category: "future-work-seed" +source: "user-observation" +found_during: "the guard-refusal session of 2026-10-07" +origin: researcher-authored +production_mode: hand-written +found_at: "the product thinker's dashboard and abcd's interviews (commands/dashboard.md)" +remedy: "Prototype on the dashboard a review page for one HTML artifact with element-anchored annotations, a feedback queue the agent collects by long-polling, and revision markers; trial it on one real interview or criteria walk against the terminal questions, measuring the person's effort and the rounds needed, before deciding whether interviews hand long material to it." +--- + +Consider a browser review loop for the product thinker, seen in an outside open-source agent tool: the agent writes an HTML artifact (a mockup, a report, a criteria walk, an intent's press release); the person opens it in the browser, annotates exact elements or text, attaches images, and queues feedback; the agent collects the whole batch with one long-polling call and writes the next revision, with the changed parts marked as revisions. The same tool audits each render for layout faults (clipped text, unreachable controls) and lists them as issues the person can send back. Today abcd's only channel from the person back to the agent is the terminal question tool, held to twenty-four rows at eighty columns, so long material is split across many questions or pushed into an HTML report read separately, with the answers still typed in the terminal. abcd's dashboard already runs a server reachable only from the person's own devices over Tailscale, but serves nothing of the project yet, so it is the natural host for such a loop. Feedback arriving from the browser is a new path into the agent's context and must be treated as data, never instruction. From a5646ab17009592d47ca1a945937f2b93d2fb026 Mon Sep 17 00:00:00 2001 From: REPPL <77722411+REPPL@users.noreply.github.com> Date: Fri, 9 Oct 2026 07:50:43 +0100 Subject: [PATCH 5/6] chore: record the lost-connection decisions from the product thinker The interview settled what the retry schedule covers, how long it runs, who waits on whom, how a dead lane restarts and when the product thinker is told, so the next run can plan the fix from the record. Refs: iss-2610080620372731 Assisted-by: Claude:claude-opus-5-5 --- ...gent-or-sub-agent-loses-its-network-connection.md | 12 +++++++++++- 1 file changed, 11 insertions(+), 1 deletion(-) diff --git a/.abcd/work/issues/open/iss-2610080620372731-when-an-agent-or-sub-agent-loses-its-network-connection.md b/.abcd/work/issues/open/iss-2610080620372731-when-an-agent-or-sub-agent-loses-its-network-connection.md index 1d30f055a..ab8e64d26 100644 --- a/.abcd/work/issues/open/iss-2610080620372731-when-an-agent-or-sub-agent-loses-its-network-connection.md +++ b/.abcd/work/issues/open/iss-2610080620372731-when-an-agent-or-sub-agent-loses-its-network-connection.md @@ -9,7 +9,17 @@ found_during: "the guard-refusal session of 2026-10-07" origin: researcher-authored production_mode: hand-written found_at: "agents and sub-agents in autonomous runs (internal/core/implement)" -remedy: "Retry a lost connection on a widening schedule (wait 1 minute, retry; 5 minutes, retry; 10 minutes, retry; then an hour), and record the outage and each retry in a machine-local channel, such as the shared run state and run log, so sibling agents and sub-agents see it and hold off without reaching each other over the network." +remedy: "On any lost connection (the host's model calls, a sub-agent, or a tool's network call such as git push or gh), a lane keeps doing offline work and, at a network step, waits on one shared probe recorded in the run state and run log: retry at 1, 5 and 10 minutes, then hourly for up to 8 hours, then stop the run and notify the product thinker; a lane whose agent died restarts as a fresh agent from its last commit, with uncommitted edits saved aside for review." --- When an agent or sub-agent loses its network connection mid-task, the work stalls or fails instead of waiting the outage out, and this keeps happening on the product thinker's machine. An agent that notices a lost connection should pause and retry on a widening schedule: wait one minute and try again, then five minutes, then ten, then an hour. Sibling agents should also learn of the outage without depending on each other over the network, since they may have lost the connection too: a channel on the machine itself, such as the run state and run log abcd already shares between sessions, would let one agent record the outage and the others hold off instead of each failing on its own. + +## The product thinker's decisions (interview of 2026-10-09) + +- **What it covers:** every kind of lost connection, because which one strikes first is not yet known: the host's own model calls stalling, a sub-agent coming back failed mid-task, and a tool's network call (git push, gh, the merge-queue call, a download) failing. The first lab run records which kinds actually occur. +- **After the hour:** retries continue hourly for up to 8 hours after the hourly stage begins, so an outage at 23:00 is still retried in the morning; then the run stops cleanly and reports what was done and what is left. +- **Who waits:** lanes carry on with offline work (editing, tests); a lane that reaches a network step waits on one shared probe recorded in the run state and run log, instead of each lane retrying on its own. +- **A lane whose agent died:** when the network returns, a fresh agent starts from the lane's last commit; uncommitted edits are saved aside for review, never built on. +- **Telling the product thinker:** a notification only when the 8-hour limit is reached and the run stops; shorter outages appear in the end-of-run report with how long each lasted and what was retried. + +Proof belongs in a lab that cuts the network during a live run and shows each of the three kinds waited out, the shared probe holding the other lanes, a dead lane restarted from its last commit, and the give-up path ending in one notification. From d76c2e6f9ba109cf7e56c5dc6580cba9bc0fc80c Mon Sep 17 00:00:00 2001 From: REPPL <77722411+REPPL@users.noreply.github.com> Date: Fri, 9 Oct 2026 09:16:39 +0100 Subject: [PATCH 6/6] chore: capture a no-repetition rule for abcd's questions A question kept overflowing the host's box because the text before it restated its options. The record proposes a rule that the lead-in says only what the options cannot, and a lab drafting questions on the quick tier until they fit. Refs: iss-2610090816296964 Assisted-by: Claude:claude-opus-5-5 --- ...ull-options-do-not-fit-the-question-box.md | 19 +++++++++++++++++++ 1 file changed, 19 insertions(+) create mode 100644 .abcd/work/issues/open/iss-2610090816296964-a-decision-s-full-options-do-not-fit-the-question-box.md diff --git a/.abcd/work/issues/open/iss-2610090816296964-a-decision-s-full-options-do-not-fit-the-question-box.md b/.abcd/work/issues/open/iss-2610090816296964-a-decision-s-full-options-do-not-fit-the-question-box.md new file mode 100644 index 000000000..e3b66b9ee --- /dev/null +++ b/.abcd/work/issues/open/iss-2610090816296964-a-decision-s-full-options-do-not-fit-the-question-box.md @@ -0,0 +1,19 @@ +--- +schema_version: 1 +id: "iss-2610090816296964" +slug: "a-decision-s-full-options-do-not-fit-the-question-box" +severity: "minor" +category: "ux" +source: "user-observation" +found_during: "an abcd question asked in another session on 2026-10-09" +origin: researcher-authored +production_mode: hand-written +found_at: "asking the person: the question tool and the guard's question check (internal/core/question)" +remedy: "Adopt a no-repetition rule now: the text before a question says only what its options cannot (what is being decided and where it stands), never restating an option, written into the GRILL asking rules and the abcd:question-drafter agent; then run a lab with the drafter on the quick tier (Haiku 4.5) rewriting until the row count fits, compared with the default tier on the same material for fit rate, lost meaning and cost (iss-2610080618506115)." +--- + +When abcd asks the person to choose, the question keeps running past the host's question box: a question over 24 rows at 80 columns was refused in a live session on the 0.13.2 plugin, and the agent cut its options down to a few words each to get it through. The product thinker, reading the refused draft, saw that its long text added nothing: a paragraph before the options described each option again, while the options' own labels and descriptions were already perfectly clear. The problem was repetition, not too little room, so the fix is a rule against it rather than a second channel for the overflow; a separate text file for the full options was considered and dropped on the product thinker's ruling of 2026-10-09. + +The rule: the text before a question says only what its options cannot, which is what is being decided and where it stands now, and never restates an option's meaning, gain or cost. It belongs in the GRILL asking rules and in the abcd:question-drafter agent, which v0.13.3 ships to draft questions within the limit; the guard can count rows but cannot judge repetition, so the drafter is where it is enforced. v0.13.3 also stops refusing a question whose only fault is its length (iss-2610070637562567), which removes the refusal but not the repetition. + +The lab that follows runs the drafter on the quick tier (Haiku 4.5), rewriting until the row count, computed exactly by abcd, fits the window, and compares its drafts with the default tier's on the same material: how often each fits, whether any meaning is lost, and the cost. Its result feeds the model-tier idea in iss-2610080618506115.