feat(agent): coordinate shared local inference#97
Conversation
|
Important Review skippedAuto reviews are disabled on this repository. Please check the settings in the CodeRabbit UI or the ⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Run ID: You can disable this status message by setting the Use the checkbox below for a quick retry:
✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_b4884de9-270b-40c2-94b2-542be1ad900b) |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 6a3fe3e40b
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_45beb517-0ddc-4f4e-a427-fb42f135d258) |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_952ae4df-3c2e-4938-aacf-86f7401eb5c6) |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ea42165ac2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_373706a5-e317-49f8-9f3d-ec619c4dd1f6) |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 1d2db7b781
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_7f7ef354-45d0-4dd1-83c5-9c6f55e38472) |
|
Caution Review failedAn error occurred during the review process. Please try again later. ✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_2c86a194-1ca3-4925-8f2b-47d7ef30d0eb) |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 7deb4979b1
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_944fae22-f320-4ab1-9570-3885e618d825) |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_083c3507-9ee1-4ff7-976b-dfc93c5a0c99) |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_6df17067-9728-4486-a968-54c749665594) |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_476877b3-848c-4234-bf8f-0d48ae16c675) |
|
Pushed What changed:
Verification:
I intentionally did not launch the 27B model; the globally linked |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_4dcb66cd-8095-48d0-b87e-e85788615d05) |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_3f29558c-aad2-4fb5-b2b0-166e198d72c6) |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 6d8e56e3df
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_15453506-7051-4cb6-babc-2896e7a06a6a) |
|
Follow-up multimodal reliability fix is pushed in The supplied Qwen3.5-MoE trace narrowed the failure to the first 2,048-token cold-prefill chunk, at the layer-8 residual materialization barrier. The changed four-image request contained 10,016 merged image tokens / 40,064 raw vision patches. Our shared Qwen3.5 vision encoder was still creating one global 40,064 x 40,064 block-diagonal mask and retaining that lazy 27-layer vision graph into language prefill. This update:
Verification:
No additional model process was launched while the user's agent session was resident. |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_ae45d667-cfee-4f86-9813-d734a3354df7) |
|
Updated in 0885568 with the cumulative multi-image vision memory fix. What changed:
Validation:
No real model was launched during validation. |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 0885568348
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_b5da6746-6d51-45de-b4a2-50faa2050d70) |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: efd3d03f01
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_dab47d34-3855-4000-a724-c64e58811676) |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_3cb59587-a67f-4a98-9cf3-f83b20232e70) |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: a8992b938d
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_51e61bc3-6c73-4be4-bb74-72db9c71d008) |
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_1273fe0a-7fcd-4fd3-b148-d82a8af26c87) |
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: 13a86112f2
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
Bugbot couldn't run - usage limit reachedBugbot is counted against Cursor usage for this user or team, and this run hit a usage or spend limit. A user or team admin can review and increase usage limits in the Cursor dashboard. (requestId: serverGenReqId_7e9d43c1-b3cf-4198-9b69-24811500b234) |
Summary
MlxModelHost, model registry, and physical paged KV pool instead of loading one model per child processWhy the allocator exhausted
The previous fixed/default pool could advertise a 262,144-token trained window while physically holding about 104,848 tokens. Pi therefore continued sending turns after the hot KV pool was full, producing
BlockAllocator exhausted. The loaded model now exposes both trained and physical limits, and both TypeScript and native paged entry points reject or clamp before allocator mutation.Shared-host behavior
Parallel subagents may run up to four independent Pi tool loops, but all inference is serialized through one resident model host. Each subagent keeps independent in-memory conversation, settings, cwd, tools, skills, and project context. Model weights and the paged pool are loaded only once for the shared model.
SSD tier boundary
This PR provides and verifies the durable cold-cache storage/restore layer, but does not enable Qwen eviction yet. Qwen3.5/3.6 hybrid continuation also needs matching GDN and MTP side state; restoring full-attention K/V alone would be incorrect. Wiring eviction is intentionally deferred until that sidecar can be captured and restored atomically.
Top-level Pi still needs an upstream
SettingsManagerinjection/update hook to scale its fixed compaction reserve for unusually small physical windows. In-process subagents are scaled here; exact preflight remains the correctness backstop everywhere.Verification
cargo check -p mlx-coremlx-coreandmlx-paged-attnwith warnings deniedNote
High Risk
Changes span serialized multi-agent inference, physical KV/context limits, native cache replay semantics, and paged-attention override cloning—errors can cause wrong model routing, context corruption, or allocator failures under load.
Overview
Agent CLI and Pi handoff now strips
--trace/--trace-dir, configuresMLX_NODE_LOG/MLX_NODE_LOG_FILEbefore the native addon loads, and prepends--models mlx/*on most runs so Pi cannot pick a cloud default while still restoring fork/session models. Context capacity is checked inChatSession(including image-expanded prompts) and in HTTP streaming handlers so oversize requests return JSONcontext_length_exceededbefore SSE or native cache mutation; outputmaxNewTokensis clamped to the physical window.Session recovery resets native caches and cold-replays preserved history after failed or abandoned delta/tool/stream paths instead of issuing another delta on possibly desynced KV state. Paged config overrides move behind
PagedConfigOverrideManagerwith separate launch-claude vs agent policies (Gemma draft hiding, MTP sidecar symlinks, collision-safe roots, cleanup races).Native (
mlx-core) adds optionalcache_owner_id/cache_root_owner_idonChatConfig, structured MLX eval errors anddeep_copy, decode-profiler token/TTFT accounting fixes, VLM paged-prefix downgrade when GDN sidecars are missing, image-aware cache-key helpers, and DSpark adaptive AR-fallback hooks. Gemma4 e2e expectations shift: pure images warm-continue; audio cold-replays withcachedTokens === 0.imagecrate enables GIF decoding.Reviewed by Cursor Bugbot for commit eba950d. Bugbot is set up for automated code reviews on this repo. Configure here.
Qwen3.5 / Qwen3.6 dense and MoE performance update
D256/BS16long-context decode: dense24Q/4KVand MoE16Q/2KV, including two-row MTP verificationmetal-rs 0.33/objc 0.2/block 0.1.6with a narrow ownership-aware adapter overobjc2 0.6.4andobjc2-metal 0.3.2Measurements
Controlled 16K / 128-token runs on the local Qwen3.6-27B MXFP4+MXFP8 dense model, with one loaded model and cold cache before each trial:
The bracketing AR weighted decode rate was 19.8 tok/s; MTP reached 91.4% of it with a mean commit of 1.826 tokens/cycle. This confirms the profiler fix and rules out the earlier 12 tok/s report as the healthy short-context baseline.
For Qwen3.6-35B-A3B without MTP, the pre-fix mlx-agent trace reached about 584 tok/s prefill but only 24.2 tok/s decode near 91K context. The corrected oMLX comparison supplied for the same M5 Max class is 547.3 tok/s prefill and 44.9 tok/s decode at 128K, so prefill is already in range and long-context decode is the remaining gap.
Isolated
16Q/2KVraw-attention A/B at 112K cached tokens improved q1 from 3.770 ms to 1.441 ms and q2 from 4.294 ms to 2.148 ms. The release addon is rebuilt; full-model post-fix 35B-A3B validation is pending the next mlx-agent trace.Additional verification
mlx-core,mlx-paged-attn, andmlx-metalblock 0.1.6,objc 0.2.7, andmetal 0.33are absentGemma4 QAT conversion and agent inference update
--config-dirplus Gemma4 mmproj remapping so authoritative config/tokenizer assets and unified vision/audio tensors can be carried into the converted checkpointdraft/;mlx agentcontinues to prefer paged AR and requiresMLX_AGENT_ENABLE_GEMMA_DRAFT=1for the flat speculative pathThe grouped Gemma4 D512 route remains default-off and diagnostic until sequential full-model A/B establishes a safe context/device crossover. The latest 47K-context agent trace confirmed prompt-cache and physical paging correctness at 607.7 tok/s bulk prefill and 15.3 tok/s decode; this update does not claim that the default long-context decode gap is closed yet.
The final unresolved review thread is also addressed: paged override clones now preserve only loader-known Qwen3.5 nested MTP sidecars and reject escaping paths while keeping unrelated directories and Gemma4
draft/hidden.Additional verification
git diff --checkmlx agentsession is active locally