Skip to content

Add memory-aware concurrent indexing for large repositories - #1925

Open
zhiyuzhang001-a11y wants to merge 7 commits into
DeusData:mainfrom
zhiyuzhang001-a11y:codex/m32-shared-provider-upstream
Open

Add memory-aware concurrent indexing for large repositories#1925
zhiyuzhang001-a11y wants to merge 7 commits into
DeusData:mainfrom
zhiyuzhang001-a11y:codex/m32-shared-provider-upstream

Conversation

@zhiyuzhang001-a11y

Copy link
Copy Markdown

Summary

  • scope daemon index workers to the canonical request repository instead of inheriting the daemon starter workspace boundary
  • add bounded two-pass indexing so large repositories do not retain the full extraction set in memory
  • add a global memory-aware scheduler for concurrent indexing across projects, while preserving exact output parity

Motivation

This enables multiple project-scoped MCP clients to share the Provider safely. Small daily repositories can run concurrently; large jobs are admitted according to measured source bytes and memory budget instead of starting without a global bound.

Validation

  • rebased onto current upstream main
  • git diff --check passes
  • ASan/UBSan focused suites: 788 passed, 1 platform skip
  • suites: subprocess, daemon_application, discover, pipeline, extraction
  • the same three commits previously passed the full local Provider suite: 7477 passed, 4 platform skips
  • downstream Codebase Atlas product suite: 237/237
  • independent root acceptance suite: 47/47

No release or binary distribution is included in this PR.

@github-actions

Copy link
Copy Markdown

Thanks for opening this — it has been seen, and it is queued.

This note is automated, but it is not a brush-off: it exists so you know where your PR stands instead of having to guess from silence.

Current review status: working through a backlog. 0.9.1-rc.1 is out, so the release freeze that held reviews is over — but it left a large queue of open pull requests behind it, and we are reading through them oldest-first. The background is in discussion #1144.

What that means for this PR, concretely:

  • It will not be closed for inactivity. No stale bot touches pull requests here.
  • It may still sit a while before a human reads it. That is on us, not on you.
  • Older PRs are read first, so a recent one is not being skipped — it is behind a queue.

Things that will genuinely speed it up whenever review does happen:

  • Keep it rebased on main — the tree is moving quickly right now, and a conflicting branch cannot be reviewed as the diff you intended.
  • Get CI green, or say which failures you believe are pre-existing.
  • Keep the change to one claim. Bundled features and refactors get split before they get merged, which costs you a round trip.
  • Every commit needs a sign-off (git commit -s) — CI enforces DCO.

If this fixes a bug, a reproduction we can run is worth more than a description of the symptom.

Thanks for contributing, and sorry in advance for the wait.

Signed-off-by: Zhiyu <zhiyuzhang001@gmail.com>
Signed-off-by: Zhiyu <zhiyuzhang001@gmail.com>
Signed-off-by: Zhiyu <zhiyuzhang001@gmail.com>
@zhiyuzhang001-a11y
zhiyuzhang001-a11y force-pushed the codex/m32-shared-provider-upstream branch from f7032c5 to 1e15c63 Compare August 30, 2026 03:06
Signed-off-by: Zhiyu <zhiyuzhang001@gmail.com>
Signed-off-by: Zhiyu <zhiyuzhang001@gmail.com>
Signed-off-by: Zhiyu <zhiyuzhang001@gmail.com>
Signed-off-by: Zhiyu <zhiyuzhang001@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant