Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
933 commits
Select commit Hold shift + click to select a range
2daa32e
run: declared checks are the review contract, judged by observable pr…
santoshkumarradha Sep 19, 2026
b208ccc
test(connect): the second-door tests use dynamic ports, not the fixed…
santoshkumarradha Sep 19, 2026
98cc254
session: a run works in a copy of its own, and its work comes home wh…
santoshkumarradha Sep 19, 2026
692d4a1
tui3: a wait whose stream has stopped arriving paints at the spinner …
santoshkumarradha Sep 19, 2026
608c459
tui3: a task's page says only what is true of it (#1240)
santoshkumarradha Sep 19, 2026
2f6559a
chat: name unread config.json keys once before defaults apply (#1239)
santoshkumarradha Sep 19, 2026
1cbcb4b
test(exec): identify a leaf-end survivor by process identity, not a b…
santoshkumarradha Sep 19, 2026
8f990d5
config: concurrent profile writes in one process compose (#1242)
santoshkumarradha Sep 19, 2026
5d9c47b
run: regression test for seating the check after a self-finished work…
santoshkumarradha Sep 19, 2026
72bbd34
session: every run a conversation has made stays readable (#1234)
santoshkumarradha Sep 19, 2026
12840ee
scripts: a repeatable hosted drive of a run on the real binary (#1241)
santoshkumarradha Sep 19, 2026
7eab247
config: unread notice skips keys the product itself retired (#1245)
santoshkumarradha Sep 19, 2026
5a7748c
tui3: what is typed after a press on a run's row goes to its page, ne…
santoshkumarradha Sep 19, 2026
1e332c3
session: make the claim-unspent revision test deterministic by forcin…
santoshkumarradha Sep 19, 2026
45137b0
fix(connect): a reconnect reuses its registered identity instead of i…
santoshkumarradha Sep 19, 2026
443c661
test(connect): stop the second-door Slack test racing on a freed port…
santoshkumarradha Sep 19, 2026
d38345f
config: two codeaf processes writing the profile at once keep each ot…
santoshkumarradha Sep 19, 2026
8ed67b7
exec: the job shell tmux tests read one marker line, not a full env d…
santoshkumarradha Sep 19, 2026
c3ab6e0
tui3: a press never waits behind a run's summary, and an open page fo…
santoshkumarradha Sep 19, 2026
7b8be4b
session: a person's stop reaches a run (#1250)
santoshkumarradha Sep 19, 2026
992779e
repo: remove a cell's findings note that rode in with #1239
santoshkumarradha Sep 19, 2026
4bf3673
exec: the skill path shell test reads one marker line (#1252)
santoshkumarradha Sep 19, 2026
eae7055
build: make check cross-compiles every platform the release ships (#1…
santoshkumarradha Sep 19, 2026
7ab8820
chatv3: the standing ticker stops and is joined when the process clos…
santoshkumarradha Sep 19, 2026
7cefdf1
tui3: no frame reads a run's rows, whatever made the read due (#1255)
santoshkumarradha Sep 19, 2026
24e4385
tui3: a run's step rows show the work, cut from the commands as they …
santoshkumarradha Sep 19, 2026
37a9b3b
tui3: a task you stopped reads stopped wherever its state word is dra…
santoshkumarradha Sep 19, 2026
1c21396
session: usage writers have an owner and are closed when the process …
santoshkumarradha Sep 19, 2026
8866dda
test(session): the claim-unspent revision is typed through Submit, th…
santoshkumarradha Sep 19, 2026
09b50f8
tui3: a call that did not run has no row, and a refused action is one…
santoshkumarradha Sep 19, 2026
137bb8c
scripts: the heavy-suite lock is a kernel file lock, so one pid names…
santoshkumarradha Sep 19, 2026
761cdab
chatv3: the pool judge sweep is stopped and joined at close (#1267)
santoshkumarradha Sep 19, 2026
42daffa
test(chatv3): the goroutine law says truthfully what the two model wa…
santoshkumarradha Sep 19, 2026
b71454e
run: a time limit a person sets ends a run by the same road as the co…
santoshkumarradha Sep 19, 2026
ced4491
tui3: the dim line under a step row is only ever the row's own (#1265)
santoshkumarradha Sep 19, 2026
d3762f4
run: a task's step numbers continue across a wake (#1261)
santoshkumarradha Sep 19, 2026
da3368d
run: a dollar limit holds while a worker is still working (#1268)
santoshkumarradha Sep 19, 2026
0d3f3ed
plandb: a parked task is finished only by a worker woken for it (#1271)
santoshkumarradha Sep 19, 2026
8eb26dc
enginehost: a launch that lands while a host is coming up reaches it …
santoshkumarradha Sep 19, 2026
ec9953e
scripts: the heavy-suite lock follows the suite, not the helper that …
santoshkumarradha Sep 19, 2026
2b02957
run: a worker that has stopped making progress is told once, then end…
santoshkumarradha Sep 19, 2026
70b3a3c
catalog: the model catalog's own warm is cancelled and joined at clos…
santoshkumarradha Sep 19, 2026
c44b488
run: no run answers done while a finished piece of work has had no re…
santoshkumarradha Sep 19, 2026
057d01f
approval: a backslash pair inside double quotes is text, read by one …
santoshkumarradha Sep 19, 2026
8832349
session: a woken turn's fixture holds its script like every other tur…
santoshkumarradha Sep 19, 2026
cd2bf32
chat: closing a process stops its place sweep before it can change a …
santoshkumarradha Sep 19, 2026
5240a37
session, run: a run's dollars reach the conversation's own total (#1280)
santoshkumarradha Sep 19, 2026
2a04d63
run, session: an ending a person caused is not a fault, and it names …
santoshkumarradha Sep 19, 2026
b293a72
suite lock: the lock follows the suite, not what the suite leaves beh…
santoshkumarradha Sep 19, 2026
bbdd7cd
exec: a cancelled command is ended by the call that ran it (#1282)
santoshkumarradha Sep 19, 2026
3a154ce
session: a run is handed what is left of every dollar limit (#1281)
santoshkumarradha Sep 19, 2026
c37f467
session: a long declared check says why it was refused (#1283)
santoshkumarradha Sep 20, 2026
cb3f0f3
session: a usage flush says whether it drained (#1285)
santoshkumarradha Sep 20, 2026
1357788
standing: a row that needs you names what it needed and opens the ite…
santoshkumarradha Sep 20, 2026
f796212
tui3: a run task's page draws its brief through the reader the transc…
santoshkumarradha Sep 20, 2026
3d95e70
session: a run is work, and its life is the conversation's (#1291)
santoshkumarradha Sep 20, 2026
8b68293
session: a standing firing runs behind the gate a task node runs behi…
santoshkumarradha Sep 20, 2026
aef0deb
task: Trace settings to permission gate wiring
agentfield-bot Sep 20, 2026
70f7baf
task: Inventory every settings row for the UX review
agentfield-bot Sep 20, 2026
33a32b7
Merge branch 'task/inventory-every-settings-row-for-60437e' into feat…
agentfield-bot Sep 20, 2026
c975972
session: a run stopped on its own question says so (#1293)
santoshkumarradha Sep 20, 2026
db7ca98
task: Implement fzf v2 fuzzy matcher universally
agentfield-bot Sep 20, 2026
3d3b754
session: work nothing is driving reads as interrupted, not failed (#1…
santoshkumarradha Sep 20, 2026
1a692da
session: a run's working copy is written down, so it can be found aga…
santoshkumarradha Sep 20, 2026
952e003
manual: the interrupted page no longer offers two answers nobody can …
santoshkumarradha Sep 20, 2026
718f6b1
fuzzy: land the remaining surface cutover (autonomy, connections, mem…
santoshkumarradha Sep 20, 2026
de66900
fuzzy: state the gap price the code and tests charge
santoshkumarradha Sep 20, 2026
70116df
Merge branch 'task/implement-fzf-v2-fuzzy-matcher-u-71bbf5' into feat…
agentfield-bot Sep 20, 2026
83de531
docs: a footprint benchmark, its method, and one day's numbers (#1126)
santoshkumarradha Sep 20, 2026
1e0537e
docs: resume cost joins the footprint benchmark (#1317)
santoshkumarradha Sep 20, 2026
b123705
feat(session): a run nothing is driving can be carried on, in the cop…
santoshkumarradha Sep 20, 2026
9420891
chore(scripts): three proof steps that could not report failure, in o…
santoshkumarradha Sep 20, 2026
af1cdd4
tui3: a standing row stopped for permission says the line, not asks: …
santoshkumarradha Sep 20, 2026
91d38b1
session: a sign-in and a harness ask say what they are waiting for (#…
santoshkumarradha Sep 20, 2026
e098ca1
tui3: the front tab and the tab beside it agree about one conversatio…
santoshkumarradha Sep 20, 2026
1094d73
task: Fix the terminal cursor park for the settings search box
agentfield-bot Sep 20, 2026
32b4127
One fuzzy matcher for every picker, subtle match highlighting, and th…
santoshkumarradha Sep 20, 2026
9b38c2d
session: a landing already accepted stops saying your call (#1322)
santoshkumarradha Sep 20, 2026
3e2214a
The commit co-author links to the CodeAF account, and assisted-by nam…
santoshkumarradha Sep 20, 2026
d57f26d
scripts: a heavy suite takes both locks so a stale tree can see it (#…
santoshkumarradha Sep 20, 2026
bc116d3
docs: the trailer entry names its pull request and takes the required…
santoshkumarradha Sep 20, 2026
9f8f761
task: Skills: per-node attachment in briefs
agentfield-bot Sep 21, 2026
e46edfa
task: Verify standing-tasks findings
agentfield-bot Sep 21, 2026
1142f90
task: Finish skills store record: continue the cut-off branch
agentfield-bot Sep 21, 2026
12b4b6e
docs: changelog entry for skills tranche 1
santoshkumarradha Sep 21, 2026
7aded50
task: Finish skills shelf and use_skill: run the right suite, land th…
agentfield-bot Sep 21, 2026
d2c2723
task: Finish skills shelf and use_skill (windowed catalog + depth-gat…
santoshkumarradha Sep 21, 2026
21660ea
session: fix ActivateSkill test callers after the skills merge
santoshkumarradha Sep 21, 2026
7a13426
exec, manual: the drafted-with footer links to the codeaf repository …
santoshkumarradha Sep 21, 2026
e87b147
docs: the change entry for the drafted-with link (#1329)
santoshkumarradha Sep 21, 2026
777bb12
skills: compose per-node skill attachments at brief time and render them
agentfield-bot Sep 21, 2026
b3f2dda
cmd: wire the active shelf into the chat engine's plan builds
agentfield-bot Sep 21, 2026
7b2fc92
task: Skills tranche 1: integrate, wire live flow, prepare draft PR
agentfield-bot Sep 21, 2026
f316c06
task: Skills tranche 1: integrate, wire live flow, prepare draft PR
santoshkumarradha Sep 21, 2026
c4f5e3c
docs: drop pr-draft.md now that the PR body carries it
santoshkumarradha Sep 21, 2026
d39a4e3
session: a held landing offers its answer and stops demanding one (#1…
santoshkumarradha Sep 21, 2026
8219b14
task: Fix round 1: address skills review findings
agentfield-bot Sep 21, 2026
d1d874a
audit-notes: move the standing-visibility investigation out of the PR…
santoshkumarradha Sep 21, 2026
a51cc6e
session: align the catalog header test with the reviewed rewording
santoshkumarradha Sep 21, 2026
af61b0c
task: Fix round 3: final fixes for skills tranche 1
agentfield-bot Sep 21, 2026
48bd211
task: Make imported SKILL.md skills readable through codeaf's skill d…
agentfield-bot Sep 21, 2026
7ed8d2f
skills: import foreign SKILL.md folders in place
santoshkumarradha Sep 21, 2026
3e6a3ba
Merge branch 'task/import-existing-skills-claude-co-b9c535' into feat…
agentfield-bot Sep 21, 2026
9ca7b44
import foreign skills before the first message on every v3 launch
santoshkumarradha Sep 21, 2026
26bbabc
skills: preserve user scope without a project root
santoshkumarradha Sep 21, 2026
d6de383
the world this task started from: Resolve skills tranche merge into s…
agentfield-bot Sep 21, 2026
72e711c
merge skills tranche and existing-skill import
santoshkumarradha Sep 21, 2026
619860c
session, cmd, bench: the worker harness is the belt a task runs on (#…
santoshkumarradha Sep 21, 2026
2984917
merge skills tranche and existing-skill import
santoshkumarradha Sep 21, 2026
6d97ef7
merge skills tranche and existing-skill import
santoshkumarradha Sep 21, 2026
ae14cf7
docs: the bash-belt entry names only surfaces the checker knows
santoshkumarradha Sep 21, 2026
88bbd5b
manual: the use_skill page says what happened instead of naming it
santoshkumarradha Sep 21, 2026
2d5eeaf
docs: the bash-belt entry names only surfaces the checker knows
santoshkumarradha Sep 21, 2026
bdf8e2f
session: a person can put a skill in front of this conversation by hand
santoshkumarradha Sep 21, 2026
fb67f07
The node belt is the default again, on a measured comparison (#1336) …
santoshkumarradha Sep 21, 2026
dd188e8
session: one message carries the skills its own words choose, attachm…
santoshkumarradha Sep 21, 2026
cef7dd4
docs: the change entry for the message-chosen skills block
santoshkumarradha Sep 21, 2026
e66650f
tui3: /skill opens the shelf as a picker and carries the attached ski…
santoshkumarradha Sep 21, 2026
43f334e
session: a missed skill name says what the shelf holds instead of not…
santoshkumarradha Sep 21, 2026
e1b36b8
session: the prompts say what a skill is and when to open one before …
santoshkumarradha Sep 21, 2026
a2a803d
manual: the use_skill page says what a missed name answers
santoshkumarradha Sep 21, 2026
c7a6de3
docs: change entry for the shelf-on-the-page wave
santoshkumarradha Sep 21, 2026
65581fd
session: the skills a turn carried ride as names, not as a sentence
santoshkumarradha Sep 21, 2026
b656cd0
tui3: test the /skill picker, write its manual page and change entry,…
santoshkumarradha Sep 21, 2026
ea967cd
session: an absent skills list means unknown, not none
santoshkumarradha Sep 21, 2026
1eeb74c
session, chat: a pass that compacted nothing stops taking the convers…
santoshkumarradha Sep 21, 2026
72aa535
tui3: draw carried skills from event field
santoshkumarradha Sep 21, 2026
cddf198
docs: explain skills carried by a turn
santoshkumarradha Sep 21, 2026
3183985
docs: add skills row change entry
santoshkumarradha Sep 21, 2026
904bdb3
docs: explain skills carried by a turn
santoshkumarradha Sep 21, 2026
7bbc70a
provider, taxonomy: a quiet reply from the only machine there is gets…
santoshkumarradha Sep 21, 2026
8a83f41
chat: a replay draws the record once, whatever was already on the scr…
santoshkumarradha Sep 21, 2026
14043ab
Merge remote-tracking branch 'origin/cell/skillpick' into skills-preview
santoshkumarradha Sep 21, 2026
a7af389
Merge remote-tracking branch 'origin/cell/skillsay' into skills-preview
santoshkumarradha Sep 21, 2026
06dcf11
Merge remote-tracking branch 'origin/santos/dev2' into skills-preview
santoshkumarradha Sep 21, 2026
5d30bad
Merge remote-tracking branch 'origin/cell/c375' into skills-preview
santoshkumarradha Sep 22, 2026
95f4fa2
tui3: give /skill its fate at home and fit its row
santoshkumarradha Sep 22, 2026
d79c12a
the bash worker's page stops naming the shelf verb
santoshkumarradha Sep 22, 2026
ee93e0a
Merge remote-tracking branch 'origin/task/bashtask-skill-rule' into s…
santoshkumarradha Sep 22, 2026
99f301b
session, taxonomy: a conversation against one machine waits for it in…
santoshkumarradha Sep 22, 2026
09815c9
Merge remote-tracking branch 'origin/santos/dev2' into skills-preview
santoshkumarradha Sep 22, 2026
90233e1
build: the pull-request gate runs on santos/dev2 as well as dev (#1362)
santoshkumarradha Sep 22, 2026
abf8032
e2e: the tmux suite reads the surface again (#1360)
santoshkumarradha Sep 22, 2026
b38a29f
docs: the belt DOE, and the chat as manager over the plan store (#1354)
santoshkumarradha Sep 22, 2026
b18b619
session, run, cmd: the worker harness is the default belt again, with…
santoshkumarradha Sep 22, 2026
cb612d3
tui3: the skills a turn used outlive the run that used them
santoshkumarradha Sep 22, 2026
1f4972c
Revert "tui3: the skills a turn used outlive the run that used them"
santoshkumarradha Sep 22, 2026
863d5ee
docs: the wait-forever entry names the pull request that merged it
santoshkumarradha Sep 22, 2026
49971de
repo: drop the local state a wave swept in with it
santoshkumarradha Sep 22, 2026
d91bb63
Merge remote-tracking branch 'origin/santos/dev2' into HEAD
santoshkumarradha Sep 22, 2026
55dd97f
session, run, chat: notes are a channel into a running worker, not a …
santoshkumarradha Sep 22, 2026
9321a85
config: persistedCount says what it destroys (#1369)
santoshkumarradha Sep 22, 2026
89dacca
session, chat: the chat sees the plan when the person speaks (#1357)
santoshkumarradha Sep 22, 2026
518f289
session: the picture hand refuses a prompt that is not there (#1372)
santoshkumarradha Sep 22, 2026
77566dc
connect: the Slack thread limit is written down once (#1371)
santoshkumarradha Sep 22, 2026
32ef528
tui3: the stopped frame is compared against a clock in the test's han…
santoshkumarradha Sep 22, 2026
2f94ca1
tui3: the stopped frame's helper says it exists in order to be shared…
santoshkumarradha Sep 22, 2026
3f46952
chat, session: skills that are switched off say so instead of looking…
santoshkumarradha Sep 22, 2026
7fb2644
Merge origin/dev into santos/dev2: dev's thirteen commits since the fork
santoshkumarradha Sep 23, 2026
e731437
run, tests: a parent wakes on its children's landings, and the pool t…
santoshkumarradha Sep 23, 2026
dca2f3f
remote: the standing firing test asks for its lane once
santoshkumarradha Sep 23, 2026
942eef6
Merge origin/dev: sign in to Codex in the browser (#1336)
santoshkumarradha Sep 23, 2026
525b77f
suite-lock: the wrapper-killed test waits for the lock to be named
santoshkumarradha Sep 23, 2026
3a9c5e4
tui3, chat, pool: task-page keys, notes, tab mark and unowned warms (…
santoshkumarradha Sep 23, 2026
b998c87
session, config: the chat manages a live run in the side list's words…
santoshkumarradha Sep 23, 2026
af913d6
run, session: a run's ending is written where the next request reads …
santoshkumarradha Sep 23, 2026
c61615a
session, provider, suite-lock: the one machine is waited on, a pool i…
santoshkumarradha Sep 23, 2026
80bdef3
Merge origin/dev into santos/dev2: /cost keeps its columns under the …
santoshkumarradha Sep 23, 2026
7c3a499
Merge origin/dev into santos/dev2: a Codex model carries its 272k window
santoshkumarradha Sep 23, 2026
e2f7587
task: admit a deliverable path that is not on this machine (#1418)
santoshkumarradha Sep 23, 2026
10beaa5
attribution: always sign, and name only the bare model (#1412)
santoshkumarradha Sep 23, 2026
d71b388
task: a path is a place only where its own directory is (#1421)
santoshkumarradha Sep 23, 2026
19e86f7
fix(review): scope matching, worker panic resilience, check env leak,…
santoshkumarradha Sep 23, 2026
1ce4cac
do, run, session: codeaf do edits in place, commits nothing, and stop…
santoshkumarradha Sep 23, 2026
147382a
tui3: a run's rows and its page look like the old task rows (#1420)
santoshkumarradha Sep 23, 2026
b8765e1
One run per batch: simultaneous hand-offs join one run, a failed run …
santoshkumarradha Sep 23, 2026
03e301c
session: a read sweep the chat will not hand off is handed off by the…
santoshkumarradha Sep 23, 2026
af99c4e
Merge remote-tracking branch 'origin/dev' into task/sync-dev-1419
santoshkumarradha Sep 23, 2026
b67c162
Engine status and stop, the older engine gives up the slot, a held co…
santoshkumarradha Sep 23, 2026
bdb08cf
Merge remote-tracking branch 'origin/santos/dev2' into task/sync-dev-…
santoshkumarradha Sep 23, 2026
c0fe652
crew: route the worker, planner and checker per task
santoshkumarradha Sep 24, 2026
b4aa1b1
crew: wire the chat, headless doors and tests to the per-task router
santoshkumarradha Sep 24, 2026
6983cf3
docs: describe the per-task crew everywhere; retire preset vocabulary
santoshkumarradha Sep 24, 2026
e934dcc
crew: carry effort and redo over the wire; fix suite fallout
santoshkumarradha Sep 24, 2026
a872c5a
docs(model-pool): rewrite the Pareto crewing paper
santoshkumarradha Sep 24, 2026
6f1f39f
Skills from Claude Code plugins and Codex, with memory off, chosen by…
santoshkumarradha Sep 24, 2026
008363c
fix: synchronize run worker step boundaries before further actions
santoshkumarradha Sep 24, 2026
ff7ced3
crew: allowed-rule edits, offers per model and a restorable crew state
santoshkumarradha Sep 24, 2026
a3940bd
tui3: /crew is an interactive overlay panel
santoshkumarradha Sep 24, 2026
6194dcc
docs: the /crew panel and the settings seats row
santoshkumarradha Sep 24, 2026
36ed835
Merge origin/santos/dev2 into feat/pareto-crew
santoshkumarradha Sep 24, 2026
c0f7206
Revert "Merge origin/santos/dev2 into feat/pareto-crew"
santoshkumarradha Sep 24, 2026
1d64c74
config: models.crew.free_routes profile row
santoshkumarradha Sep 24, 2026
8831410
config: record models.crew.free_routes in the profile-key ledger
santoshkumarradha Sep 24, 2026
1f16183
crew: a providers row on the /crew panel
santoshkumarradha Sep 24, 2026
0fe86d3
tui3: the free-routes switch as the /crew providers list's last line
santoshkumarradha Sep 24, 2026
90da8d8
tui3: /crew providers fixes: narrow list, /connect return, routable g…
santoshkumarradha Sep 24, 2026
a6810bf
crewroute: canonical model identity, free routes, seat capability fil…
santoshkumarradha Sep 24, 2026
ceab2b8
provider: a route-failure taxonomy read from refusals
santoshkumarradha Sep 24, 2026
7f2146b
config: route health and the crew degradation ladder
santoshkumarradha Sep 24, 2026
050c61f
session: seats fall down their ladder inside the task
santoshkumarradha Sep 24, 2026
437d4f7
tui3: one crew line updated in place after the started line
santoshkumarradha Sep 24, 2026
f83eed0
manual: the crew line, the seat ladder and one-rung redo
santoshkumarradha Sep 24, 2026
2e33f08
config: a crew seat reads and writes text
santoshkumarradha Sep 24, 2026
1fc970a
tui3: the /crew seat hint names only a reachable model
santoshkumarradha Sep 24, 2026
537c9a7
session: a task kept on its branch counts as accepted
santoshkumarradha Sep 24, 2026
f30a82d
crew: the free rung, account-aware rescue and a stopped action
santoshkumarradha Sep 24, 2026
ad1825c
tui3: a failed crew line leads with its state; lines reset per conver…
santoshkumarradha Sep 24, 2026
87ff3f4
manual: the stopped crew line, free rescue and free pins
santoshkumarradha Sep 24, 2026
faaea7a
crew: free rescue on thin models, credit probe, health-aware helpers,…
santoshkumarradha Sep 25, 2026
8cf1c5a
session: seat ladder stops on what the task saw; helpers avoid dead r…
santoshkumarradha Sep 25, 2026
eb09dbd
tui3: a pinned free pool is reachable whatever the free-routes switch…
santoshkumarradha Sep 25, 2026
9320b67
crewroute: calibrated estimates, exact pins, decision explanations
santoshkumarradha Sep 25, 2026
e617902
crew: learned cost factor, exact-id catalog reads, compact decision rows
santoshkumarradha Sep 25, 2026
ed02bd5
session: a run call on a tier the crew does not seat passes the healt…
santoshkumarradha Sep 25, 2026
d886abb
manual: how the crew estimate is made and what a pin sends
santoshkumarradha Sep 25, 2026
9adbff6
crew: rescue prefers general models; one-clause free notice; credit a…
santoshkumarradha Sep 25, 2026
6e900dc
provider: a call may hand a rate limit back instead of waiting it out
santoshkumarradha Sep 25, 2026
13792ad
crewroute: tiny and domain-tuned names, rivals at a similar cost, net…
santoshkumarradha Sep 25, 2026
4ddde28
config: rescue floor, indexes merged per model, migration keeps every…
santoshkumarradha Sep 25, 2026
1a7d45b
session: seats never wait out a limit, walk at most three free pools,…
santoshkumarradha Sep 25, 2026
b737313
manual: bounded free walk, net-move line, pins kept through migration
santoshkumarradha Sep 25, 2026
e3367cf
crewroute: label spellings, plain defect titles, complex-bugfix worke…
santoshkumarradha Sep 25, 2026
8cf1b09
config: crew spend lines, unpriced models, route answers, pin order
santoshkumarradha Sep 25, 2026
28462e4
session: price seat calls before sending, rest failed sends, route gate
santoshkumarradha Sep 25, 2026
2245424
tui3: plain-tier marks on crew surfaces; manual for cap, ceiling, reach
santoshkumarradha Sep 25, 2026
dee766c
crewroute: carry the whole reading through RouteCrew; evidence floor
santoshkumarradha Sep 25, 2026
5abfc84
session: crew guards on the default profile; helpers capped; crew sto…
santoshkumarradha Sep 25, 2026
cce59b0
docs(model-pool): ship the paper and the router prior only
santoshkumarradha Sep 25, 2026
49b65ed
crewroute: drop the test pinned to measured per-crew costs
santoshkumarradha Sep 25, 2026
a712d53
crewroute, session, config: state routing rules without measurements
santoshkumarradha Sep 25, 2026
40c6b21
docs(model-pool): keep the built paper only
santoshkumarradha Sep 25, 2026
699df7c
feat(crew): score every catalog model with fitted weights; per-task $…
santoshkumarradha Sep 25, 2026
3e0ec84
fix(crew): index-first ability, risk-weighed seats, checker floor, re…
santoshkumarradha Sep 25, 2026
1edb45b
feat(crew): support upgrades go to the checker; worker ability floor …
santoshkumarradha Sep 25, 2026
72e33b9
Merge dev into feat/pareto-crew
santoshkumarradha Sep 25, 2026
bd6dd30
config: test that a fresh profile starts on auto with the default limits
santoshkumarradha Sep 25, 2026
867f62c
docs: describe the routed crew fully in the manual, README and LIMITS
santoshkumarradha Sep 25, 2026
e5ffb16
crewroute: time a decision by its fastest batch
santoshkumarradha Sep 25, 2026
e0ed1f5
config: a pinned thinking level rides the pin's send
AbirAbbas Sep 25, 2026
ff4cde1
config: auto in the reflex or small-work row reads the default and mi…
AbirAbbas Sep 25, 2026
1030df4
do: -yes-spend passes the crew's daily cap as the door promises
AbirAbbas Sep 25, 2026
b2995d9
crew: the cap refusal names only the ways on that work
AbirAbbas Sep 25, 2026
f96b332
crew: a zero or unknown amount of crew money is not drawn
AbirAbbas Sep 25, 2026
c3d3f63
manual: the Providers tab draws the crew as one seats row
AbirAbbas Sep 25, 2026
34598f1
Merge origin/dev (#1429, #1485) into the Pareto crew
AbirAbbas Sep 25, 2026
195cf13
Merge origin/dev (#1494) into the Pareto crew
AbirAbbas Sep 25, 2026
805aba0
README: link the Pareto Crewing paper from the crew section
agentfield-bot Sep 25, 2026
316d92b
Merge remote-tracking branch 'origin/feat/pareto-crew' into review/sp…
AbirAbbas Sep 25, 2026
3e44d5f
Merge origin/dev (#1511) into the Pareto crew
AbirAbbas Sep 25, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 2 additions & 2 deletions PERF.md
Original file line number Diff line number Diff line change
Expand Up @@ -1136,8 +1136,8 @@ and an explicit zero temperature or output limit is preserved.
Explicit choices remain explicit. `CODEAF_REASONING`,
`CODEAF_EXEC_REASONING`, a model or crew value with `:low`, `:medium` or
`:high`, a saved task or standing-work rung, and an embedder's `ai.Option` still
travel. The three shipped crew presets contain bare model ids and add no effort
level. `cmd/harness-design` is a development command with explicit CLI-sized
travel. A crew seat nobody pinned is routed to a bare model id and adds no
effort level. `cmd/harness-design` is a development command with explicit CLI-sized
requests and retains its caps.

The local process remains bounded independently of provider generation:
Expand Down
38 changes: 22 additions & 16 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -160,22 +160,28 @@ Every number, the method and the limits: [docs/benchmarks/deepswe](docs/benchmar

## The right model for each call

One session, many models. The model you talk to is one seat. Five more, the
crew, take the calls you did not type:
One session, many models. The model you talk to is one seat. Every task you hand
off runs on a crew of three more, picked for that task from what kind of work it
is — a bug fix, a complex fix, open-ended work, or something else:

| seat | what it answers |
| seat | what it does |
| --- | --- |
| reflex | memory, titles, the safety gate. Near free, reads every turn. |
| small work | digests, task names, yes-or-no checks |
| worker | every task you hand off. Most of the bill. |
| careful work | checks on finished work, the brief a task is shaped into, vision |
| mastermind | plans runs and designs subharnesses |

`/crew frugal`, `balanced` or `max` sets all five in one word. Any seat can be
pinned.

Every finished task is graded by the check it already had to pass. Work that
keeps failing on the worker seat moves up to careful work on its own, and each
| worker | does the work. Most of the bill. |
| planner | plans runs and designs subharnesses |
| checker | checks finished work before it lands |

The crew is auto by default. Each seat is scored from its catalog metadata by
learned weights, moved by how this install's own tasks ended, over every model your
connected providers serve. The method is in the paper,
[Pareto Crewing](docs/design/model-pool/pareto-crewing.pdf). `/crew` shows it: the seats, the allowed models, the
providers, a per-task limit ($5 unless set) and an optional daily cap, and today's
spend. `/crew pin checker <model>` pins one seat, `/crew models open` limits every
seat to open-weight models, and `/task --best` or `/task --cheap` moves one task.
Each task says its crew and what it cost against the estimate, and `/redo
stronger` runs it again a step up. Two cheap rows, reflex and small work, take
the small calls you did not type: memory, titles, digests, the safety gate.

Every finished task is graded by the check it already had to pass, and each
request goes to the provider that has been fastest for that kind of call.
`codeaf models` prints the ratings.

Expand All @@ -186,15 +192,15 @@ ChatGPT plan, Ollama and any OpenAI-compatible endpoint.

## Model Pool

The picker can choose models from what other installs have found. It is on by
Installs can pool what they have measured about models. It is on by
default: what an install sends is computed, text-free numbers about the models
it ran (role, model, a number, which model judged, door, size bucket, day) under
a per-install nonce, never code,
prompts, paths or an identity, and `codeaf pool status` shows exactly what is
waiting to go. Turn it off with `model_pool = off` on the settings sheet or
`CODEAF_MODEL_POOL=off`; `read` uses the pool and sends nothing, and
`CODEAF_TELEMETRY=off` caps it at `read` along with the usage counts. The relay
publishes a signed index the crew picker reads under `picked from = learn`. The index is mirrored on the `model-pool` branch at
publishes a signed index of what the installs measured. The index is mirrored on the `model-pool` branch at
`pool/index.json`. The design is [Pareto Crewing](docs/design/model-pool/pareto-crewing.pdf);
the relay's code is under `relay/`, with a [runbook](docs/design/model-pool/RUNBOOK.md)
that includes running your own.
Expand Down
30 changes: 10 additions & 20 deletions audit-notes/settings-inventory.md
Original file line number Diff line number Diff line change
Expand Up @@ -39,7 +39,7 @@ sentence of the registry's own Hint (`settings.go:663`).
rows; a search → every tab's matches under faint headings, and the tab bar
follows the first match.
- Providers is led by `modelsSection` (`settings.go:700`): your model,
provider, speed guard, routing, prompt profile, crew, then the five tier
provider, speed guard, routing, prompt profile, then the five tier
rows in `roles.Tiers` order (reflex, low, worker, high, mastermind), then
pinned roles. Every other row follows in registry order
(`tabRows`, `settings.go:1430`). The registry builds the 10 model rows
Expand Down Expand Up @@ -229,12 +229,11 @@ rows land between the roles section and "looking"):
| lane.guard | speed guard | toggle | bool | on (settings.go:935) | "an answer that is slow to start is asked of the next-best provider as well, and you read whichever replies first. One extra call, under a tenth of spend." | registry `settings.go:2055`, skin `settings.go:703` |
| routing | routing | cycle | choice: simple/latency/price/off | simple | "one model is served by many providers. simple is the one it ships with and sends no preference of ours — no pinned provider means the router's own default answers, and a pinned provider is the whole request; latency asks for the fastest and demotes one that keeps being slow; price asks for the cheapest; off asks for nothing, measures nothing, and leaves the two rows above it with no provider to name. a change here takes effect on your next message." | registry `settings.go:1999`, skin `settings.go:779` |
| prompt.profile | prompt profile | cycle | choice: auto/lean/full | auto | "how much codeaf tells the model before you type. auto reads the model's context window and goes lean under 32,000 tokens; lean and full say so yourself, for a provider that reports a window its model does not really have." | registry `settings.go:2022`, skin `settings.go:710` |
| models.crew | crew | cycle | choice: frugal/balanced/max (+ reads custom) | balanced — derived; the five shipped tiers are exactly the balanced preset (`settings.go:1017-1031`) | dynamic, built by `crewAbout` (`settings.go:716`): "the five below, chosen as one word: frugal — …; balanced — …; max — …. Answer one yourself and this reads custom." (the preset sentences come from `config.CrewLineFor`) | registry `settings.go:2269`, skin `settings.go:322` |
| models.tiers.reflex | reflex | select | model | mistralai/mistral-nemo (shipped, settings.go:1017) | "near-free · reads every turn — memory, titles, safety" | registry `settings.go:2297`, skin `settings.go:334` |
| models.tiers.low | small work | select | model | deepseek/deepseek-v4-flash-0731 (shipped) | "cheap · the small calls — names, digests, the safety gate" | registry `settings.go:2306`, skin `settings.go:364` |
| models.tiers.worker | worker | select | model | z-ai/glm-5.3-flash (shipped) | "does the work · every task, its parts, every run node — most of the bill" | registry `settings.go:2317`, skin `settings.go:368` |
| models.tiers.high | careful work | select | model | moonshotai/kimi-k3 (shipped) | "careful · checks what must not be wrong — audits, briefs, vision" | registry `settings.go:2327`, skin `settings.go:372` |
| models.tiers.mastermind | mastermind | text | model, may carry :low/:medium/:high | moonshotai/kimi-k3 (shipped) | "thinks · plans runs and designs harnesses — add :low, :medium or :high" | registry `settings.go:2338`, skin `settings.go:380` |
| models.tiers.worker | worker | select | model | empty — auto, routed per task | "does the work · every task, its parts, every run node — most of the bill" | registry `settings.go:2317`, skin `settings.go:368` |
| models.tiers.high | checker | select | model | empty — auto, routed per task | "checks what must not be wrong — audits, briefs, vision" | registry `settings.go:2327`, skin `settings.go:372` |
| models.tiers.mastermind | planner | text | model, may carry :low/:medium/:high | empty — auto, routed per task | "plans runs and designs harnesses — add :low, :medium or :high" | registry `settings.go:2338` |
| models.roles | pinned roles | text | text (role:model pairs) | none | "exceptions to the five rows above, one per role: title:openai/gpt-5-mini." | registry `settings.go:2347`, skin `settings.go:386` |
| model.plan | planning | select | model | blank, "follows execution" | firstSentence of hint: "the model that plans and reviews the work." (full hint `settings.go:2965`: "the model that plans and reviews the work. Empty follows the work model.") | registry `settings.go:2686` (modelRow), skin `settings.go:683` (init) |
| model.work | execution | select | model | the work model | "the model that does the work." (full hint: "the model that does the work. It changes on the next job.") | registry modelRow; skin init |
Expand All @@ -249,7 +248,6 @@ rows land between the roles section and "looking"):
| effort | thinking | cycle | choice: auto + the effort rungs (auto, plus five explicit levels; `EffortChoices`, `config/effort.go:29`) | auto (effort.Ship = None, `effort/effort.go:63`) | "how hard the model thinks, unless something nearer the work says otherwise. ctrl+v moves the rung of whatever you stand on — the rung beside the model above the message box for one conversation, a task, or a standing item — and ctrl+t in /model dials one model. This row answers for everything nobody dialled." | registry `settings.go:1789`, skin `settings.go:757` |
| document_engine | reading | cycle | choice: auto/local/free/ocr | auto (config.go:62) | "which rung reads your documents. auto walks local, then free, then paid OCR." | registry `settings.go:1798`, skin `settings.go:680` |
| api_key | openrouter key | text | text, secret | not set | "the key codeaf talks to models with. A missing default key opens connect openrouter in your browser; paste a replacement here if needed. A change lands on this conversation at once." | registry `settings.go:1829`, skin `settings.go:476` |
| models.crew.source | model family | cycle | choice: open/all | open (crew.go:86) | "which models the crew word draws from: open weights, or the whole catalog with closed and frontier models in it. Open is the default." | registry `settings.go:2284`, skin `settings.go:341` |
| reply.guard | reply guard | cycle | choice: on/off | on | "on cuts a reply that has come apart — one line or one letter repeated, alphabets mixed inside words — throws it away and asks once more. Code blocks are never judged." | registry `settings.go:2417`, skin `settings.go:427` |

### The roles section (generated, not registry rows)
Expand All @@ -266,8 +264,8 @@ registered from internal/session init functions. 23 rows in 5 groups:
| roles · reflex | reflex |
| roles · small work | title, caption, router, consolidate, task-name, job-name, intake, guardian, sentinel, spellout |
| roles · worker | worker |
| roles · careful work | careful, auditor, repair, shaper, vision |
| roles · mastermind | planner, designer, mark-reader, handoff, division, router-confirm |
| roles · checker | careful, auditor, repair, shaper, vision |
| roles · planner | planner, designer, mark-reader, handoff, division, router-confirm |

Each row shows the model the role resolves to right now; a pinned row also says
"pinned". The one line under a row (verbatim): unpinned → "<description> ·
Expand Down Expand Up @@ -378,12 +376,11 @@ branch can, because its registry tops out at 77 (78 with split_pct).
|---|---|---|
| model_pool | CategoryModels, choice, label "model pool" (dev settings.go:1911) | fde49a587, #1194 |
| models.pool.public_key | CategoryModels, text, label "pool key" (dev settings.go:1926) | fde49a587, #1194 |
| models.crew.pick | CategoryModels, choice, label "picked from" (dev settings.go:2400) | fde49a587, #1194 |
| telemetry | CategoryInterface, bool, label "telemetry", default on (dev settings.go:2627) | cc8bea7e9, #1095 |

The `Key:` diff between this branch's build and dev's is exactly those four
The `Key:` diff between this branch's build and dev's is exactly those three
rows and nothing else: every key this branch has, dev has too, and no key dev
has is missing here but those four. The four-row difference is fixed rows that
has is missing here but those three. The three-row difference is fixed rows that
landed on dev after the fork, not conditional rows and not rows this branch
retired. The deletion comment at settings.go:3046 names two rows gone outright
(`practice_demand_pct`, `propose_new_skills`), and an earlier draft of this
Expand Down Expand Up @@ -465,16 +462,9 @@ for the word "budget" will not find it — though the search does match the key
the panel moved the row to Spending as "per conversation" — a person reading
the Session tab's promise ("this conversation and only this conversation")
will not find the row that bounds this conversation.
- THE CREW ROW'S NEIGHBOUR MOVED. The skin comment on "model family"
(models.crew.source) claims "it sits directly under the crew word"
(settings.go:337), but `modelsSection` does not lead it, so on the drawn
Providers tab it reads after the openrouter-key row (and after the roles and
services sections when they stand), ~25 rows below the crew word it changes.
The comment describes registry order, not the tab's reading order — a
redesign should either lead it or fix the comment.
- "thinking" (effort) vs mastermind's ":high". Two rows both about how hard a
- "thinking" (effort) vs the planner row's ":high". Two rows both about how hard a
model thinks; effort's about says it "answers for everything nobody dialled"
and names ctrl+v and /model's ctrl+t, and the mastermind row's about says
and names ctrl+v and /model's ctrl+t, and the planner row's about says
"add :low, :medium or :high". The distinction (install-wide rung vs one
tier's level) is stated only in the about lines.
- The Connections tab name does double duty: it is the accounts tab, and the
Expand Down
10 changes: 5 additions & 5 deletions cmd/codeaf/chat.go
Original file line number Diff line number Diff line change
Expand Up @@ -166,14 +166,14 @@ func buildBrain(w *chatWindow, session string, opts brainOptions) (*chatBrain, e
// then let the machine form print underneath, which said one fact twice.
return nil, err
}
// A tier row that says auto is answered from this catalog (config.AutoModels),
// so the word is wired BEFORE the seats handed in are applied — a door whose
// ladder answered the word before this line resolved it from nothing. The
// read is the same non-blocking one, never a fetch.
// The crew router picks its seats from this catalog (config.CrewCatalog),
// so it is wired BEFORE the seats handed in are applied — a door whose
// ladder routed before this line would have routed from nothing. The read
// is the same non-blocking one, never a fetch.
modelCatalog := catalog.LoadLazy(context.Background(), catalog.Options{
BaseURL: settings.BaseURL, APIKey: settings.APIKey, Dir: settings.ProfileDir,
})
config.AutoModels = modelCatalog.ModelsNow
seatCrewRows(modelCatalog.ModelsNow)
wirePoolIndex(settings.ProfileDir)
if opts.seats != nil {
applySeats(&settings, *opts.seats)
Expand Down
23 changes: 20 additions & 3 deletions cmd/codeaf/chatv3.go
Original file line number Diff line number Diff line change
Expand Up @@ -4,6 +4,7 @@ import (
"context"
"errors"
"fmt"
"github.com/Agent-Field/codeaf/internal/crewroute"
"io"
"os"
"path/filepath"
Expand Down Expand Up @@ -937,15 +938,15 @@ func openV3Launch(proc *v3Process, opts v3Options) (*v3Launch, error) {
// engine's own nothing. The ask is built once and a client is made from it
// per call, each billed to the judge's own seat.
taskLanded := poolJudgeHook(settings, settings.ProfileDir, workspace,
config.AutoModels, poolJudgeAsk(settings, settings.ProfileDir), time.Now, "task")
config.CrewCatalog, poolJudgeAsk(settings, settings.ProfileDir), time.Now, "task")
// The runs a live process would have judged but a process death left unjudged,
// and the headless doors that never had this hook: at start, on a goroutine
// nobody waits on, judge the resumed session's own final-state nodes and the
// pending file's rows, each exactly once, bounded so it never holds the prompt. The
// process tracker cancels and joins it at close.
poolErrandGoCtx(settings.ProfileDir, "pool/judge-sweep", func(ctx context.Context) {
poolJudgeSweepRun(ctx, settings, settings.ProfileDir, found.Place.Tasks(),
config.AutoModels, poolJudgeAsk(settings, settings.ProfileDir), time.Now)
config.CrewCatalog, poolJudgeAsk(settings, settings.ProfileDir), time.Now)
})

cfg := session.Config{
Expand Down Expand Up @@ -1627,6 +1628,16 @@ func applyV3Governance(cfg session.Config, profileDir string, yolo, oneModel boo
cfg.ApprovalPolicy = policy
cfg.RolesSource = source
cfg.OneModel = oneModel
// EVERY TASK THIS CONVERSATION STARTS IS ROUTED ITS OWN CREW — worker,
// planner, checker picked for that task from the profile's allowed models
// and pins (internal/config's RouteCrew, internal/session's taskcrew.go).
// Under `--one-model` there is no crew: every call rides the conversation's
// model, which is what the flag says, so no router is handed over.
if !oneModel {
cfg.RouteCrew = func(ask config.CrewAsk) (crewroute.Decision, error) {
return config.RouteCrew(profileDir, ask)
}
}
cfg.SpendRailUSD = rail
// The fallback chain reads PROFILE-ONLY, like the search keys below and
// unlike the three rows above it. A repository that could answer this could
Expand Down Expand Up @@ -2153,7 +2164,7 @@ func v3SearchSeam(brain *store.Store) tui3.SearchStore {
// IT IS LIVE. It used to be resolved once, at boot, on the argument that two
// calls in one conversation must not answer to different settings — and the
// crew is what makes that argument the wrong way round. A person who types
// `/crew max` because the planner is not thinking hard enough has said something
// `/crew pin planner …` because the planner is not thinking hard enough has said something
// about the run they are about to start, not about the next launch, and a source
// that made them restart to be heard would be a knob that does nothing on the
// surface that offers it.
Expand Down Expand Up @@ -2284,6 +2295,12 @@ func (c *v3Crew) snapshot() (map[string]string, error) {
// the cheapest question in it. The mastermind tier: it plans adaptive runs
// and designs saved harnesses, and a repository that could point it at a
// model would be spending a visitor's credit on the run it asked for.
// A CHECKER NO PROJECT NAMED IS THE CREW'S: its pin, or the router's
// standing pick ([config.TierModelAt]) — never an empty row that would
// fall to the conversation's model.
if strings.TrimSpace(high) == "" {
high = config.TierModelAt(c.profileDir, config.ModelTierHigh)
}
values := map[string]string{
roles.TierKey(roles.TierReflex): config.TierModelAt(c.profileDir, config.ModelTierReflex),
roles.TierKey(roles.TierMastermind): config.TierModelAt(c.profileDir, config.ModelTierMastermind),
Expand Down
Loading
Loading