diff --git a/README.md b/README.md index 6a12b19..797faa4 100644 --- a/README.md +++ b/README.md @@ -2,7 +2,7 @@ English | [简体中文](README.zh-CN.md) -> Current protocol version: `0.1.20` (in development); workflow version: `0.1.3` +> Current protocol version: `0.1.21` (in development); workflow version: `0.1.3` Polaris is a repo-native engineering workflow for coding agent hosts. It stores requirements, plans, implementation results, independent reviews, validation evidence, and task state in Git, then uses deterministic gates to prevent requirement drift, stale evidence, and agents declaring their own work complete. @@ -89,9 +89,13 @@ codegraph init polaris code-intelligence add codegraph --repo . ``` -Run these commands from the target repository as appropriate. `codegraph init` creates the `.codegraph/` marker; without it Polaris uses source and Git directly and creates no stage record. Polaris writes a Code Intelligence record only when it actually performs a Provider status, sync, or explore operation. Polaris can only read CodeGraph status, explore indexed relationships, and perform one bounded `codegraph sync` at a declared stage boundary. It never installs, initializes, starts, configures, reconfigures, waits for, or manages CodeGraph or its watcher/daemon/MCP configuration. +Run these commands from the target repository as appropriate. `codegraph init` creates the `.codegraph/` marker; without it Polaris uses source and Git directly and creates no stage record. Vendoring registers the project-scoped `polaris-codegraph` proxy in `.codex/config.toml` and `.mcp.json` without replacing unrelated settings. The host may require project trust or first-use approval; that approval remains the user's decision. -CodeGraph's watcher and connection reconciliation are the primary freshness mechanisms. Polaris records a limited conclusion at the time it checks: `CURRENT_AT_CHECK`, `PARTIAL_STALE`, `INDEX_STALE`, `NOT_VERIFIED`, or `UNAVAILABLE`; it never claims commit-exact graph freshness. A `PARTIAL_STALE` response names specific files: read each current file directly (`READ_SOURCE`), or inspect the registered Git diff if it was deleted (`INSPECT_GIT_DIFF`). For `INDEX_STALE` or `NOT_VERIFIED`, treat the graph only as a lead and search the repository plus Git (`SEARCH_SOURCE`). Validation remains graph-free and relies on source, Git, builds, tests, static checks, and Human Checks. +Polaris stages call only `polaris_codegraph_explore`. The proxy checks status, may internally perform one bounded `codegraph sync` when requested and pending, runs one explore, rechecks status, and returns a freshness envelope before graph content. There is no separate stage status/sync MCP call. `CURRENT` means `NON_AUTHORITATIVE_CONTEXT`; `STALE` and `UNKNOWN` mean `NAVIGATION_ONLY` and require the named source/Git fallback; `UNAVAILABLE` means no graph. A current named file uses `READ_SOURCE`, a deleted file uses `INSPECT_GIT_DIFF`, and an index-wide or unsafe result uses `SEARCH_SOURCE`. Validation remains graph-free and relies on source, Git, builds, tests, static checks, and Human Checks. + +The repository owner, not Polaris, owns CodeGraph installation, initialization, configuration, raw MCP registration, watcher, and daemon. Polaris never starts, configures, reconfigures, waits for, or manages them. Raw `codegraph_explore` or `codegraph explore` remains available out-of-band but cannot back `CURRENT` Polaris evidence. New records are v3 projections of the retained proxy bundle and completed fallbacks; v1/v2 are historical only. CodeGraph remains optional and never becomes a workflow gate. + +Protocol `0.1.21` adds the project-scoped Polaris CodeGraph proxy, host adapter v3 registration, and auditable Code Intelligence record v3 while leaving Workflow at `0.1.3`. Record v1 and v2 are immutable historical evidence only; new evidence is projected from a retained proxy bundle into v3. ## v0.1 scope diff --git a/README.zh-CN.md b/README.zh-CN.md index 871a53d..c7d1397 100644 --- a/README.zh-CN.md +++ b/README.zh-CN.md @@ -2,7 +2,7 @@ [English](README.md) | 简体中文 -> 当前协议版本:`0.1.20`(开发中);Workflow 版本:`0.1.3` +> 当前协议版本:`0.1.21`(开发中);Workflow 版本:`0.1.3` Polaris 是运行在 Coding Agent 宿主上的仓库原生工程工作流。它把需求、计划、实现、独立审查、验证和任务状态保存在 Git 仓库中,并通过确定性门禁防止需求漂移、证据过期和 Agent 自行宣布完成。 @@ -89,9 +89,13 @@ codegraph init polaris code-intelligence add codegraph --repo . ``` -`codegraph init` 创建 `.codegraph/` marker。只有目标仓库已经有这个 marker 且项目策略允许时,Polaris 才会使用 CodeGraph;没有 marker 时直接使用源码和 Git,不生成阶段 record。只有实际执行 Provider `status`、`sync` 或 `explore` 操作时才写 Code Intelligence record。Polaris 只会读取 `status`、查询 `explore`,以及只在声明的阶段边界至多执行一次有界 `codegraph sync`;它绝不安装、初始化、启动、配置、重新配置、等待或管理 CodeGraph、watcher、daemon 或 MCP 配置。 +`codegraph init` 创建 `.codegraph/` marker;没有 marker 时 Polaris 直接使用源码和 Git,不生成阶段 record。Vendoring 会在 `.codex/config.toml` 与 `.mcp.json` 中非破坏地注册项目级 `polaris-codegraph` 代理,并保留其他设置。宿主可能要求信任项目或首次使用确认;是否批准仍由用户决定。 -CodeGraph 的 watcher 和连接时 reconciliation 是正常情况下的实时更新机制。Polaris 只记录检查时的有限结论:`CURRENT_AT_CHECK`、`PARTIAL_STALE`、`INDEX_STALE`、`NOT_VERIFIED` 或 `UNAVAILABLE`,不会宣称与 Git commit 精确一致。`PARTIAL_STALE` 会精确列出待同步文件:当前普通文件必须直接读取并记录 `READ_SOURCE`;已删除文件必须检查注册 subject 的 Git diff 并记录 `INSPECT_GIT_DIFF`。`INDEX_STALE` 或 `NOT_VERIFIED` 时,图只能作为导航线索,Agent 必须通过仓库搜索和 Git 证据记录 `SEARCH_SOURCE`。Provider 不可用、status 不可读或 sync 失败都不阻断阶段;Validation 不调用 CodeGraph,仍以源码、Git、构建、测试、静态检查和 Human Check 为准。 +Polaris 阶段只调用 `polaris_codegraph_explore`。代理先检查 status,按请求且确有 pending 时至多执行一次有界 `codegraph sync`,再执行一次 explore、复查 status,并保证 freshness envelope 位于图内容之前;阶段没有独立的 status/sync MCP 调用。`CURRENT` 表示 `NON_AUTHORITATIVE_CONTEXT`;`STALE` 与 `UNKNOWN` 表示 `NAVIGATION_ONLY`,必须完成 envelope 指定的源码/Git 回退;`UNAVAILABLE` 表示没有图内容。当前具名文件使用 `READ_SOURCE`,已删除文件使用 `INSPECT_GIT_DIFF`,索引级或不安全结果使用 `SEARCH_SOURCE`。Validation 不调用 CodeGraph,仍以源码、Git、构建、测试、静态检查和 Human Check 为准。 + +CodeGraph 的安装、初始化、配置、raw MCP 注册、watcher 与 daemon 归仓库所有者,而不是 Polaris。Polaris 绝不启动、配置、重新配置、等待或管理这些能力。raw `codegraph_explore` 或 `codegraph explore` 仍可作为带外工具使用,但不能支持 Polaris 的 `CURRENT` 证据。新 record 必须由保留的代理 bundle 与已完成回退投影为 v3;v1/v2 仅供历史读取。CodeGraph 始终可选,永远不是 Workflow 门禁。 + +协议 `0.1.21` 新增项目级 Polaris CodeGraph 代理、Host Adapter v3 注册和可审计的 Code Intelligence record v3,Workflow 仍为 `0.1.3`。record v1/v2 仅作为不可变历史证据读取;新证据必须由保留的代理 bundle 投影为 v3。 ## v0.1 边界 diff --git a/VERSION b/VERSION index baa9837..7906299 100644 --- a/VERSION +++ b/VERSION @@ -1 +1 @@ -0.1.20 +0.1.21 diff --git a/docs/USAGE.md b/docs/USAGE.md index 9cc12c0..ed1f6da 100644 --- a/docs/USAGE.md +++ b/docs/USAGE.md @@ -2,7 +2,7 @@ 本文面向希望在受支持 Coding Agent 宿主中使用 Polaris 管理软件工程任务的项目成员。当前内置 Codex 与 Claude Code 适配器;本文从首次接入讲到日常提出需求、独立 Implementation、进度查询、Review、验证、恢复与升级。 -> 当前协议版本:v0.1.20;Workflow 版本:v0.1.3。Polaris v0.1 是仓库原生的 Skills、宿主 worker 定义与 Python 脚本集合,并提供一个只分发到这些脚本的 `polaris` CLI;不提供后台服务或图形界面。 +> 当前协议版本:v0.1.21;Workflow 版本:v0.1.3。Polaris v0.1 是仓库原生的 Skills、宿主 worker 定义与 Python 脚本集合,并提供一个只分发到这些脚本的 `polaris` CLI;不提供后台服务或图形界面。 ## 1. 先理解 Polaris 保存什么 @@ -175,17 +175,17 @@ Claude Code 应加载 `.claude/skills/engineering-task/SKILL.md`。R1/R2 Impleme ### 3.6 添加新的宿主适配器(维护者) -1. 新建 `hosts//adapter.json`,使用 `adapter_version: 2`,并按 `schemas/host-adapter.schema.json` 声明 Skill 目标、调用前缀、真实入口、能力、入口 frontmatter、overlay、appendix 与专用文件。 +1. 新建 `hosts//adapter.json`,使用 `adapter_version: 3`,并按 `schemas/host-adapter.schema.json` 声明 Skill 目标、调用前缀、真实入口、能力、入口 frontmatter、overlay、appendix、专用文件与唯一 `project_mcp` 注册。 2. `capabilities` 必须显式声明 `structured_user_input / worker_create / worker_status / worker_resume / stable_worker_identity`。`worker_status` 和稳定身份依赖 worker 创建;续接同时依赖创建与稳定身份。声明可创建 worker 的宿主必须提供入口 Skill appendix,写清创建、身份、查询和续接机制。 3. 只把宿主能力差异放入该目录:metadata 放在 overlay,worker 创建/身份/等待/续接规则放在 `skill-appendices/engineering-task.md`,原生 agent 或仓库规则放在 `files` 清单中。 4. 不要在共享 `skills/`、Workflow、Authority schema 或三个生命周期脚本中新增宿主名分支。共享 Skill 引用另一 Skill 时使用 `{{skill:}}`。 5. 运行完整测试,并在真实宿主中 smoke test 入口发现、显式触发边界、隔离 worker、handoff 拒绝和同一 Implementer 续接 Documentation Sync。 -`entry_skill` 必须对应 canonical `skills//SKILL.md`。Overlay 只能在已知 Skill 下增加 canonical 源中不存在的普通文件,不能提供 `SKILL.md`、覆盖任何同路径内容或包含未知 Skill;appendix 也只能使用 `.md`。适配器源树、manifest、overlay、appendix、专用源文件和目标写入路径都禁止 symlink,所有目标必须留在仓库内。不同宿主不能声明重叠目标。若新宿主无法用 v2 的“文件复制 + Skill 渲染 + 能力声明 + 执行附录”表达,应先升级适配器契约,而不是在核心脚本里写例外。 +`entry_skill` 必须对应 canonical `skills//SKILL.md`。Overlay 只能在已知 Skill 下增加 canonical 源中不存在的普通文件,不能提供 `SKILL.md`、覆盖任何同路径内容或包含未知 Skill;appendix 也只能使用 `.md`。适配器源树、manifest、overlay、appendix、专用源文件、MCP 配置目标和启动器路径都禁止 symlink,所有目标必须留在仓库内。`project_mcp` 固定 server ID、`python3` 启动器、vendored 脚本及 `--repo .`;不同宿主不能声明重叠目标。若新宿主无法用 v3 契约表达,应先升级适配器契约,而不是在核心脚本里写例外。 ### 3.7 可选 Code Intelligence -Polaris v0.1 只支持 [colbymchenry/codegraph](https://github.com/colbymchenry/codegraph) 作为正式 Code Intelligence Provider。用户拥有安装、初始化和宿主 MCP 配置;在目标仓库中自行按顺序运行: +Polaris v0.1 只支持 [colbymchenry/codegraph](https://github.com/colbymchenry/codegraph) 作为正式 Code Intelligence Provider。用户拥有 CodeGraph 的安装、初始化、配置、raw MCP 注册、watcher 与 daemon;在目标仓库中自行按顺序运行: ```text codegraph install @@ -193,9 +193,11 @@ codegraph init polaris code-intelligence add codegraph --repo . ``` -前两个命令绝不会由 Polaris 执行;`codegraph init` 创建 `.codegraph/`,它是 Polaris 允许查询的前提。最后一个命令只创建或更新 `.polaris/code-intelligence.json`,将模式设为 `auto_optional`、将 CodeGraph 放到 Provider 优先级首位,并保留已有 `include` / `exclude` 规则。命令可幂等重跑,未知 Provider 或非法旧配置会在写入前拒绝。 +前两个命令绝不会由 Polaris 执行;`codegraph init` 创建 `.codegraph/`,它是 Polaris 允许查询的前提。最后一个命令只创建或更新 `.polaris/code-intelligence.json`,将模式设为 `auto_optional`、将 CodeGraph 放到 Provider 优先级首位,并保留已有 `include` / `exclude` 规则。命令可幂等重跑,未知 Provider 或非法旧配置会在写入前拒绝。Polaris 不会启动、配置、重新配置或等待 CodeGraph。 -Python CLI 无法直接查看 Codex 或 Claude Code 当前会话中的 MCP 工具,因此成功只表示 Provider 已加入 Polaris;返回的 `runtime_status` 为 `checked_by_next_workflow`。下一次 Workflow 只有在仓库已经存在 `.codegraph/` 时才会检查实际能力;缺少 marker 或策略禁用时直接继续源码搜索、读取、构建、测试和 Review,不生成阶段 record。只有实际执行 `status`、`sync` 或 `explore` 后才写 record;操作失败时如实记录 `UNAVAILABLE` 或 `NOT_VERIFIED`,但不阻断阶段。 +Vendoring 会非破坏地把项目级 `polaris-codegraph` 代理注册到 Codex 的 `.codex/config.toml` 和 Claude Code 的 `.mcp.json`,只管理同名条目并把配置列入安装清单的 `preserved_files`。已有无关设置与服务器会保留;损坏配置、同名冲突、越界路径或 symlink 会在覆盖前拒绝。宿主可能在首次启动时要求信任项目或批准 MCP;这是用户决定,Polaris 不绕过。 + +Python CLI 无法直接查看 Codex 或 Claude Code 当前会话中的 MCP 工具,因此配置命令成功只表示 Provider 已加入 Polaris;返回的 `runtime_status` 为 `checked_by_next_workflow`。下一次 Workflow 只有在仓库已经存在 `.codegraph/` 时才会调用项目代理;缺少 marker 或策略禁用时直接继续源码搜索、读取、构建、测试和 Review,不生成阶段 record。 不执行该命令时仍保留默认自动发现。`.polaris/code-intelligence.json` 也可用于禁用 Provider、调整优先级或限制索引范围。例如: @@ -217,11 +219,13 @@ Python CLI 无法直接查看 Codex 或 Claude Code 当前会话中的 MCP 工 } ``` -CodeGraph watcher 与连接时 reconciliation 是常规实时更新机制。Polaris 只在 Planning、Implementation、Review 的阶段入口和最终 Documentation Sync 的有界点读取 status;仅 status 指出 pending changes 时,才至多运行一次 `codegraph sync` 并至多复查一次 status。Polaris 只会 `status`、`explore` 和这一次有界 `sync`,不会等待 watcher、循环查询、启动 daemon 或改写 MCP 配置。 +CodeGraph watcher 与连接时 reconciliation 是常规实时更新机制。Polaris 阶段只调用 `polaris_codegraph_explore`:代理在同一有界窗口内检查 status,按调用参数且确有 pending 时至多运行一次 `codegraph sync`,执行一次 explore,再复查 status。阶段没有独立的 status/sync MCP 工具,也不会等待 watcher、轮询、重试、启动 daemon 或改写用户的 raw MCP 配置。Documentation Sync 仅在 supported source 变化时执行一次查询,使用 `sync_if_needed: true`,并把 query 限制到 changed source paths 与 documented symbols。 + +代理结果的第一个内容块总是 freshness envelope。`CURRENT / NON_AUTHORITATIVE_CONTEXT` 表示图可作为非权威上下文;`STALE / NAVIGATION_ONLY` 表示已知失效;`UNKNOWN / NAVIGATION_ONLY` 表示无法证明新鲜度;`UNAVAILABLE / NO_GRAPH` 表示没有图输出。任何状态都不宣称与 Git commit 严格一致,`UNKNOWN` 绝不能当作 current。raw `codegraph_explore` 或 `codegraph explore` 仍可由用户带外调用,但不能支持 Polaris 的 `CURRENT` 证据。 -精简 record 保存在任务的 `code-intelligence/rNNN/*.json`,包含 Provider、阶段、目标 commit/diff、查询目的、响应哈希、新鲜度、stale point 与实际源码回退证据;原始 MCP 响应只允许进入 ignored 的 `runtime/code-intelligence/`。新鲜度只表示检查时的有限结论:`CURRENT_AT_CHECK`、`PARTIAL_STALE`、`INDEX_STALE`、`NOT_VERIFIED` 或 `UNAVAILABLE`,不宣称与 Git commit 严格一致。 +`STALE` 或 `UNKNOWN` 必须先完成 envelope 指定的源码/Git fallback。当前具名普通文件直接读取并记录 `READ_SOURCE` 与当前 SHA-256;安全但已删除的路径检查注册 subject 的 Git diff,记录 `INSPECT_GIT_DIFF`、null observed SHA-256 与 base/head/diff hashes;不安全路径或索引级失效执行 `SEARCH_SOURCE`,记录有限、受限的当前文件路径与 SHA-256。图不能扩大冻结 scope、替代源码或决定 Review verdict,Validation 完全不调用 CodeGraph。 -`PARTIAL_STALE` 会精确列出 pending 文件。若列出的受限路径仍是当前普通文件,Agent 必须直接读取它并记录 `READ_SOURCE`;若已删除,必须检查注册 subject 的 Git diff 并记录 `INSPECT_GIT_DIFF`。`INDEX_STALE` 或 `NOT_VERIFIED` 表示整个图只能作为导航线索,Agent 必须以仓库搜索和 Git 证据回退并记录 `SEARCH_SOURCE`。没有 `.codegraph/`、Provider 故障或 sync 失败都不阻塞阶段;图不能扩大冻结 scope、替代源码或决定 Review verdict,Validation 完全不调用 CodeGraph。 +每次代理调用都会把精确响应和 bundle 留在 ignored 的 `runtime/code-intelligence/`。完成 fallback 后,Agent 写只含 summary、已确认 symbols 和 source_fallbacks 的 annotations JSON,再运行 `record_code_intelligence.py --repo . --bundle --annotations ` 投影不可变 v3 record。不得手写 record;没有代理调用就省略 record。v1/v2 record 仅作为不可变历史证据读取。 ## 4. Polaris 仓库自举 @@ -606,8 +610,9 @@ polaris migrate --repo . 1. `workflow/migrations.json` 是支持路径的唯一、append-only 注册表;历史步骤必须保留,以便校验已提交的迁移记录。一次命令只允许从当前项目版本迁移到 vendored 版本的一个显式相邻步骤,不推断、不跨级。 2. 注册步骤同时绑定源/目标 `polaris_version` 与 `workflow_version`。Migration protocol v2 支持仅更新版本,也支持显式替换冻结 workflow 并映射任务状态。 3. `0.1.19 → 0.1.20` 使用 `replace_version_and_workflow` 与 `append_mapped_workflow_event`:冻结 workflow 更新到 `0.1.3`,旧 `IMPLEMENTED` / `DOCS_SYNCED` 映射到 `IMPLEMENTING`,旧 `REVIEWED` 映射到 `VALIDATING`;旧 R0/R1 `VERIFIED` 也映射回 `VALIDATING`,以便通过 `PASS_AND_CLOSE` 重新提交关闭产物,R2 `VERIFIED` 保持不变。迁移事件记录源/目标状态及旧版本;旧 `events.jsonl` 行不可修改。 -4. `.polaris/migrations/MIG--to-.json` 先写为 `IN_PROGRESS`,全部投影更新后改为 `COMPLETED`。迁移锁会记录迁移/任务身份、主机名和 PID;若进程在中间终止,同一主机重新执行命令会接管已死亡的同迁移锁、验证并复用已经追加的事件,不会重复迁移。活跃进程、其他迁移或来源不明的锁不会被自动删除。 -5. 迁移完成后脚本自动运行项目校验;`validate_project.py` 会拒绝未完成记录、缺失/伪造的任务迁移事件或版本不一致。 +4. `0.1.20 → 0.1.21` 只替换协议版本,Workflow 保持 `0.1.3`。迁移会校验并清点 canonical v1/v2 Code Intelligence 历史记录的路径与 SHA-256,保持原字节不变;中断恢复前会重算清单,任何变化都会拒绝继续。v1/v2 此后仅可作为历史证据读取。 +5. `.polaris/migrations/MIG--to-.json` 先写为 `IN_PROGRESS`,全部投影更新后改为 `COMPLETED`。迁移锁会记录迁移/任务身份、主机名和 PID;若进程在中间终止,同一主机重新执行命令会接管已死亡的同迁移锁、验证并复用已经追加的事件,不会重复迁移。活跃进程、其他迁移或来源不明的锁不会被自动删除。 +6. 迁移完成后脚本自动运行项目校验;`validate_project.py` 会拒绝未完成记录、缺失/伪造的任务迁移事件或版本不一致。 没有注册路径时不要手改版本号。应先取得包含所需相邻步骤的 Polaris 版本,逐级完成并分别提交;任何失败都先保留 `.polaris/migrations/` 和事件现场,修复原因后重跑同一迁移命令。 @@ -631,6 +636,8 @@ v0.1.19 将正式 Provider 固定为 [colbymchenry/codegraph](https://github.com v0.1.20 / Workflow v0.1.3 删除没有独立治理边界的中间状态和事件;`START_IMPLEMENTATION` 与 `START_REVIEW` 各自原子注册所需产物,Review 接受后直接进入 `VALIDATING`,R0/R1 使用 `PASS_AND_CLOSE`。本机进度改为可选遥测;未执行 Provider 操作时不再生成 Code Intelligence record。 +v0.1.21 新增项目级 Polaris CodeGraph MCP 代理、Host Adapter v3 注册与 Code Intelligence record v3;Workflow 仍为 v0.1.3,CodeGraph 仍为可选且不参与门禁。v1/v2 record 仅作为不可变历史证据读取。 + ## 13. 失败探索与卡点 如果一个技术方向被证据否定,不要让结论只留在聊天中。记录任务内探索: diff --git a/docs/superpowers/plans/2026-08-18-codegraph-freshness.md b/docs/superpowers/plans/2026-08-18-codegraph-freshness.md index d055214..27b2a98 100644 --- a/docs/superpowers/plans/2026-08-18-codegraph-freshness.md +++ b/docs/superpowers/plans/2026-08-18-codegraph-freshness.md @@ -1,5 +1,11 @@ # CodeGraph Freshness Integration Implementation Plan +> **Historical plan:** This v2 plan has been superseded by +> `docs/superpowers/plans/2026-08-19-codegraph-polaris-mcp-proxy.md`. +> Current stages must use only `polaris_codegraph_explore`; the direct +> status/sync/raw-explore instructions below are retained solely as migration +> history and cannot support new Polaris `CURRENT` evidence. + > **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. **Goal:** Make `colbymchenry/codegraph` the only formal `codegraph` Provider, keep its graph current with watcher-aware one-shot sync, and record precise stale points that force bounded source fallback. diff --git a/docs/superpowers/plans/2026-08-19-codegraph-polaris-mcp-proxy.md b/docs/superpowers/plans/2026-08-19-codegraph-polaris-mcp-proxy.md new file mode 100644 index 0000000..1e67b79 --- /dev/null +++ b/docs/superpowers/plans/2026-08-19-codegraph-polaris-mcp-proxy.md @@ -0,0 +1,937 @@ +# CodeGraph Polaris MCP Proxy Implementation Plan + +> **For agentic workers:** REQUIRED SUB-SKILL: Use superpowers:subagent-driven-development (recommended) or superpowers:executing-plans to implement this plan task-by-task. Steps use checkbox (`- [ ]`) syntax for tracking. + +**Goal:** Add a project-scoped Polaris MCP proxy that performs one bounded CodeGraph freshness window, delivers an adjacent freshness envelope, and records auditable v3 evidence without making CodeGraph a workflow gate. + +**Architecture:** Extend the existing CodeGraph CLI adapter with reusable status/sync/explore primitives, then place a host-neutral proxy orchestration module above it. A thin standard-library stdio MCP entry point exposes only `polaris_codegraph_explore`; host adapter v3 renders project-local registrations for Codex and Claude Code. New v3 records copy and validate the proxy bundle while frozen v1/v2 schemas remain readable historical formats. + +**Tech Stack:** Python 3.10+ standard library (`argparse`, `hashlib`, `json`, `subprocess`, `unittest`; `tomllib` when available), JSON Schema through Polaris's existing validator, JSON-RPC 2.0/MCP stdio protocol revision `2025-11-25`, Git, GitHub Actions. + +**Spec:** `docs/superpowers/specs/2026-08-18-codegraph-polaris-mcp-proxy-design.md` + +## Global Constraints + +- Polaris protocol and package version becomes exactly `0.1.21`. +- Workflow graph version remains exactly `0.1.3`; do not change states or transitions. +- Host adapter manifest version becomes exactly v3. +- New Code Intelligence writes use record v3; v1 and v2 remain immutable historical read formats. +- `colbymchenry/codegraph` remains the only official CodeGraph Provider. +- Runtime code remains Python-standard-library-only; add no package dependency. +- The proxy performs no sleep, polling, retry, CodeGraph installation, initialization, configuration, daemon management, or watcher management. +- One invocation performs at most one pre-status, one sync, one post-sync status, one explore, and one post-query status. +- Raw `codegraph_explore` MCP and unrestricted shell access remain available but cannot back `CURRENT` Polaris evidence. +- Validation remains graph-free and CI must pass without CodeGraph installed. +- `CURRENT` is non-authoritative context; `STALE` and `UNKNOWN` are navigation-only and require exact source/Git fallback evidence. +- Project mismatch, unsafe paths, malformed warnings, missing proof, and unavailable capabilities fail closed. + +--- + +## File Structure + +- `scripts/internal/codegraph_adapter.py`: low-level bounded CodeGraph CLI calls and fail-safe response classification. +- `scripts/internal/code_intelligence_proxy.py`: stage resolution, query-window orchestration, delivery-state merge, runtime bundle persistence, and envelope rendering. +- `scripts/code_intelligence_mcp.py`: newline-delimited JSON-RPC/MCP stdio dispatcher only. +- `scripts/internal/code_intelligence_protocol.py`: v1/v2/v3 record selection and semantic validation. +- `schemas/code-intelligence-record-v2.schema.json`: frozen copy of the current v2 schema. +- `schemas/code-intelligence-record.schema.json`: current v3 schema. +- `schemas/code-intelligence-record-annotations.schema.json`: Agent-supplied summaries, symbols, and completed fallback evidence used when projecting a bundle. +- `scripts/internal/project_mcp_registration.py`: non-destructive Codex TOML and Claude JSON registration rendering/validation. +- `scripts/internal/host_adapters.py`: adapter v3 manifest validation and safe project MCP target resolution. +- `scripts/vendor_project.py`, `scripts/init_project.py`, `scripts/validate_project.py`: transactionally install and validate the project-local registration. +- `hosts/codex/adapter.json`, `hosts/claude-code/adapter.json`: declarative host registration metadata. +- `skills/*`, `templates/AGENTS.md`, `README*.md`, `docs/USAGE.md`, `plan.md`: one shared human/Agent behavior contract. +- `workflow/migrations.json`, `scripts/internal/migration_protocol.py`, version/template files: adjacent `0.1.20` to `0.1.21` migration and frozen v2 inventory. +- `tests/test_codegraph.py`: adapter, proxy, MCP, record, migration, Skill, and optional real-CLI coverage. +- `tests/test_core.py`: host registration, vendoring, version, validation, and rollback coverage. + +--- + +### Task 1: Bounded CodeGraph CLI primitives and fail-safe response classification + +**Files:** +- Modify: `scripts/internal/codegraph_adapter.py` +- Test: `tests/test_codegraph.py` + +**Interfaces:** +- Consumes: existing `inspect_status(repo, descriptor, runner, timeout_seconds) -> dict` and CodeGraph descriptor `cli.*_args`. +- Produces: `run_explore(repo, descriptor, query, *, runner, timeout_seconds) -> dict`, `synchronize_observed_status(repo, descriptor, initial, *, runner, status_timeout_seconds, sync_timeout_seconds) -> dict`, and stricter `classify_response(repo, response, checked_at=None) -> dict`. + +- [ ] **Step 1: Write failing tests for one-shot explore and observed-status sync** + +Add tests that assert the exact call sequence and shared repository cwd: + +```python +def test_explore_and_observed_sync_are_bounded_to_one_repo(self) -> None: + calls = [] + + def runner(command, **kwargs): + calls.append((command, kwargs["cwd"])) + if command[1:3] == ["status", "--json"]: + return completed(healthy_status(self.repo)) + if command[1:] == ["sync", "--quiet"]: + return completed("synced\n") + return completed("graph response\n") + + pending = json.loads(healthy_status(self.repo)) + pending["pendingChanges"]["modified"] = 1 + initial = codegraph_adapter._status_result( + self.repo, pending, "2026-08-19T00:00:00Z", "a" * 64 + ) + synchronized = codegraph_adapter.synchronize_observed_status( + self.repo, load_providers(ROOT)["codegraph"], initial, runner=runner + ) + explored = codegraph_adapter.run_explore( + self.repo, load_providers(ROOT)["codegraph"], "find symbol A", runner=runner + ) + self.assertEqual(synchronized["sync"]["status"], "SUCCESS") + self.assertEqual(explored["status"], "SUCCESS") + self.assertEqual(explored["response_sha256"], hashlib.sha256(b"graph response\n").hexdigest()) + self.assertTrue(all(cwd == self.repo for _command, cwd in calls)) + self.assertEqual(sum(command[1] == "sync" for command, _cwd in calls), 1) + self.assertEqual(sum(command[1] == "explore" for command, _cwd in calls), 1) +``` + +- [ ] **Step 2: Run the new focused test and confirm RED** + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests.test_explore_and_observed_sync_are_bounded_to_one_repo -v` + +Expected: `AttributeError` for `synchronize_observed_status` or `run_explore`. + +- [ ] **Step 3: Implement the reusable primitives and refactor the legacy wrapper** + +Use these exact public shapes: + +```python +def run_explore(repo, descriptor, query, *, runner=subprocess.run, timeout_seconds=60): + checked_at = _checked_at() + if not isinstance(query, str) or not query.strip(): + return {"status": "FAILED", "checked_at": checked_at, "response": None, + "response_sha256": None, "error": "CodeGraph query must not be blank"} + try: + completed = _run_cli( + repo, descriptor, "explore_args", timeout_seconds, runner, + extra_args=[query], + ) + raw, digest = _stdout_and_hash(completed) + except (KeyError, OSError, TypeError, UnicodeError, ValueError, + subprocess.TimeoutExpired) as error: + return {"status": "FAILED", "checked_at": checked_at, "response": None, + "response_sha256": None, "error": _error_summary(error)} + if completed.returncode != 0: + return {"status": "FAILED", "checked_at": checked_at, "response": None, + "response_sha256": digest, + "error": f"CodeGraph explore exited with {completed.returncode}"} + return {"status": "SUCCESS", "checked_at": checked_at, "response": raw, + "response_sha256": digest, "error": None} +``` + +Change `_run_cli(..., extra_args: list[str] | None = None)` to append only the supplied list. Move the current sync body into `synchronize_observed_status(...)`; keep `sync_if_needed(...)` as `inspect_status(...)` followed by that function so the existing CLI remains compatible. + +- [ ] **Step 4: Write failing tests for suspicious warning forms** + +Replace the old permissive expectations with: + +```python +def test_suspicious_or_wrapped_freshness_warnings_are_not_verified(self) -> None: + samples = ( + "warning: graph may be stale\n", + "quoted: ⚠️ CodeGraph auto-sync is DISABLED — the index is frozen.\n", + "\ufeff⚠️ CodeGraph auto-sync is DISABLED — the index is frozen.\n", + " pending-sync required\n", + ) + for response in samples: + with self.subTest(response=response): + result = codegraph_adapter.classify_response(self.repo, response) + self.assertEqual(result["classification"], "NOT_VERIFIED") + self.assertEqual(result["stale_points"][0]["reason"], "STATUS_UNREADABLE") + self.assertEqual( + result["response_sha256"], + hashlib.sha256(response.encode("utf-8")).hexdigest(), + ) +``` + +- [ ] **Step 5: Run the warning test and confirm RED** + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests.test_suspicious_or_wrapped_freshness_warnings_are_not_verified -v` + +Expected: at least one sample classifies as `NONE`. + +- [ ] **Step 6: Implement conservative suspicious-signal classification** + +After exact supported-banner parsing and before returning `NONE`, reject case-insensitive `warning`, `stale`, `pending-sync`, `pending sync`, `out-of-date`, or any `⚠` marker as `NOT_VERIFIED`. Preserve the exact response digest in every branch. + +- [ ] **Step 7: Run adapter tests** + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests.test_explore_and_observed_sync_are_bounded_to_one_repo tests.test_codegraph.CodeGraphTests.test_suspicious_or_wrapped_freshness_warnings_are_not_verified tests.test_codegraph.CodeGraphTests.test_pending_changes_sync_once_and_recheck_once tests.test_codegraph.CodeGraphTests.test_response_banner_marks_only_named_files_stale -v` + +Expected: all PASS; no test observes more than one sync or explore. + +- [ ] **Step 8: Commit Task 1** + +```bash +git add scripts/internal/codegraph_adapter.py tests/test_codegraph.py +git commit -m "feat: add bounded CodeGraph query primitives" +``` + +--- + +### Task 2: Proxy query-window engine and immutable runtime bundles + +**Files:** +- Create: `scripts/internal/code_intelligence_proxy.py` +- Modify: `scripts/internal/task_layout.py` +- Modify: `scripts/internal/code_intelligence_protocol.py` +- Test: `tests/test_codegraph.py` + +**Interfaces:** +- Consumes: Task 1's `inspect_status`, `synchronize_observed_status`, `run_explore`, and `classify_response`. +- Produces: `resolve_stage_context(repo, task_id, stage) -> dict`, `execute_proxy_query(repo, task_id, stage, query_id, purpose, query, sync_if_needed, *, runner=subprocess.run) -> dict`, and `render_freshness_envelope(bundle) -> str`. + +- [ ] **Step 1: Write failing stage-context and path-confinement tests** + +Assert these canonical runtime locations: + +```python +def test_proxy_stage_context_uses_record_name_and_sequential_query_ids(self) -> None: + context = code_intelligence_proxy.resolve_stage_context( + self.repo, "TASK-0001", "PLANNING" + ) + self.assertEqual(context["work_item_revision"], 1) + self.assertEqual(context["artifact_attempt"], None) + self.assertEqual(context["reviewer_slot"], None) + self.assertEqual(context["record_name"], "planning") + expected = ( + self.repo + / ".polaris/tasks/TASK-0001/runtime/code-intelligence/planning/CIQ-001.json" + ) + self.assertEqual( + code_intelligence_proxy.proxy_bundle_path( + self.repo, "TASK-0001", context, "CIQ-001" + ), + expected, + ) +``` + +Add table cases for Implementation/Documentation Sync using the current implementation handoff attempt, and Review using the current review handoff plus the next reviewer slot. Reject a stage inconsistent with task state, `CIQ-000`, a skipped ID, an existing bundle, a symlinked runtime component, and more than `CIQ-999`. + +- [ ] **Step 2: Run stage-context tests and confirm RED** + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests.test_proxy_stage_context_uses_record_name_and_sequential_query_ids -v` + +Expected: import failure for `internal.code_intelligence_proxy`. + +- [ ] **Step 3: Implement stage resolution and bundle layout** + +Add `code_intelligence_proxy_bundle` to `TASK_PATH_PATTERNS`: + +```python +"code_intelligence_proxy_bundle": ( + "runtime/code-intelligence/{record_name}/{query_id}.json" +), +"code_intelligence_proxy_response": ( + "runtime/code-intelligence/{record_name}/{query_id}.response.txt" +), +``` + +Extend `task_relative_path(..., query_id: str = "CIQ-001")` and pass +`query_id=query_id` into its single `pattern.format(...)` call so both helpers +remain governed by the task-layout registry. + +`resolve_stage_context` must validate the task, current revision, and stage-specific artifact, then return exactly: + +```python +{ + "task_id": task_id, + "work_item_revision": revision, + "stage": stage, + "artifact_attempt": attempt_or_none, + "reviewer_slot": slot_or_none, + "record_name": _record_name(record_identity), + "target": {"base_commit": base, "head_commit": head, "diff_hash": diff_hash}, +} +``` + +Planning derives `base_commit` from the frozen work item and uses null head/diff. Implementation and Documentation Sync use the current implementation handoff/subject. Review uses the current review handoff subject and chooses slot 1 before slot 2. Reuse existing artifact validators rather than trusting raw JSON fields. + +- [ ] **Step 4: Write failing query-window classification tests** + +Use a scripted runner and assert all four outcomes: + +```python +class ScriptedCodeGraphRunner: + def __init__(self, responses): + self.responses = list(responses) + self.calls = [] + + def __call__(self, command, **kwargs): + self.calls.append((command, kwargs)) + if not self.responses: + raise AssertionError(f"unexpected extra CodeGraph call: {command}") + return self.responses.pop(0) + +def test_proxy_window_requires_clean_pre_and_post_status_for_current(self) -> None: + runner = ScriptedCodeGraphRunner([ + completed(healthy_status(self.repo)), + completed("graph bytes\n"), + completed(healthy_status(self.repo)), + ]) + result = code_intelligence_proxy.execute_proxy_query( + self.repo, "TASK-0001", "PLANNING", "CIQ-001", + "locate affected symbols", "symbol A", False, runner=runner, + ) + bundle = result["bundle"] + self.assertEqual(bundle["delivery"]["state"], "CURRENT") + self.assertEqual(bundle["delivery"]["usage"], "NON_AUTHORITATIVE_CONTEXT") + self.assertEqual(bundle["delivery"]["record_status"], "CURRENT_AT_CHECK") + self.assertEqual([call[0][1] for call in runner.calls], ["status", "explore", "status"]) +``` + +Add cases for pending pre-status without sync (`STALE/INDEX_STALE` but explore once), pending pre-status with a successful single sync, pending post-status downgrade, malformed pre-status (`UNKNOWN` and no explore), explore failure (`UNKNOWN` without graph), stale banner, suspicious banner, project mismatch, disabled policy/no marker/missing executable (`UNAVAILABLE` and no Provider call), and unsafe response path (discard graph). + +- [ ] **Step 5: Run query-window tests and confirm RED** + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests.test_proxy_window_requires_clean_pre_and_post_status_for_current -v` + +Expected: missing `execute_proxy_query`. + +- [ ] **Step 6: Implement the exact bundle and conservative merge** + +Persist this versioned shape with `write_json_atomic` and reject overwrite: + +```python +{ + "bundle_version": 1, + "proxy": {"server_id": "polaris-codegraph", "tool": "polaris_codegraph_explore"}, + "provider": {"id": "codegraph", "descriptor_version": 2}, + "repository": {"project_id": project_id, "root_sha256": sha256(str(repo.resolve()))}, + "task_context": stage_context, + "query": { + "id": query_id, "purpose": purpose, "text": query, + "status": "SUCCESS|FAILED|UNAVAILABLE", "response_sha256": digest_or_none, + "error": error_or_none, + }, + "pre_status": status_observation, + "sync": sync_observation_or_none, + "post_sync_status": status_observation_or_none, + "response_classification": classification_or_none, + "post_query_status": status_observation_or_none, + "delivery": { + "state": "CURRENT|STALE|UNKNOWN|UNAVAILABLE", + "record_status": "CURRENT_AT_CHECK|PARTIAL_STALE|INDEX_STALE|NOT_VERIFIED|UNAVAILABLE", + "reason": finite_reason, + "checked_at": timestamp, + "usage": "NON_AUTHORITATIVE_CONTEXT|NAVIGATION_ONLY|NO_GRAPH", + "required_fallback": "NONE|READ_SOURCE|INSPECT_GIT_DIFF|SEARCH_SOURCE", + "stale_points": stale_points, + "error": finite_error_or_none, + }, + "response_path": task_relative_response_path_or_none, +} +``` + +Only `CURRENT` may use `NON_AUTHORITATIVE_CONTEXT`; only `UNAVAILABLE` may use `NO_GRAPH`. Any known stale signal wins over clean observations. Any missing/unreadable proof becomes `UNKNOWN`. Save the exact UTF-8 graph response before the bundle and verify its digest. If project identity or response integrity is unsafe, remove `response_path` and do not return graph text. + +- [ ] **Step 7: Implement and test the bounded envelope** + +`render_freshness_envelope` must emit only finite scalar fields between exact start/end markers, with pending counts from the most conservative successful status and a task-relative bundle path. Add an assertion that the first returned character sequence is `[POLARIS_CODEGRAPH_FRESHNESS]` and that diagnostics are truncated to 240 characters. + +- [ ] **Step 8: Run proxy tests** + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests -k proxy -v` + +Expected: all proxy tests PASS and all fake runner call counts match their exact maxima. + +- [ ] **Step 9: Commit Task 2** + +```bash +git add scripts/internal/code_intelligence_proxy.py scripts/internal/task_layout.py scripts/internal/code_intelligence_protocol.py tests/test_codegraph.py +git commit -m "feat: bind CodeGraph queries to freshness windows" +``` + +--- + +### Task 3: Standard-library stdio MCP server + +**Files:** +- Create: `scripts/code_intelligence_mcp.py` +- Test: `tests/test_codegraph.py` + +**Interfaces:** +- Consumes: Task 2's `execute_proxy_query` and `render_freshness_envelope`. +- Produces: executable MCP server supporting `initialize`, `notifications/initialized`, `ping`, `tools/list`, and `tools/call` for exactly one tool. + +- [ ] **Step 1: Write a subprocess MCP transcript test** + +Send one compact JSON object per input line: + +```python +messages = [ + {"jsonrpc": "2.0", "id": 1, "method": "initialize", "params": { + "protocolVersion": "2025-11-25", "capabilities": {}, + "clientInfo": {"name": "test", "version": "1"}, + }}, + {"jsonrpc": "2.0", "method": "notifications/initialized"}, + {"jsonrpc": "2.0", "id": 2, "method": "tools/list", "params": {}}, +] +completed = subprocess.run( + [sys.executable, SCRIPTS / "code_intelligence_mcp.py", "--repo", self.repo], + input="".join(json.dumps(item) + "\n" for item in messages), + text=True, capture_output=True, check=False, +) +responses = [json.loads(line) for line in completed.stdout.splitlines()] +self.assertEqual(responses[0]["result"]["protocolVersion"], "2025-11-25") +self.assertEqual(responses[0]["result"]["capabilities"], {"tools": {"listChanged": False}}) +self.assertEqual([item["name"] for item in responses[1]["result"]["tools"]], + ["polaris_codegraph_explore"]) +self.assertNotIn("codegraph_explore", {item["name"] for item in responses[1]["result"]["tools"]}) +``` + +Also assert the initialized notification produces no response and stdout contains no non-JSON lines. + +- [ ] **Step 2: Run the transcript test and confirm RED** + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests.test_mcp_server_initializes_and_lists_one_proxy_tool -v` + +Expected: script missing. + +- [ ] **Step 3: Implement the MCP lifecycle and one tool schema** + +Use newline-delimited UTF-8 JSON-RPC per the MCP `2025-11-25` stdio transport. The tool schema must set `additionalProperties: false`, require all six approved arguments, constrain task/stage/query IDs, and never expose a repository argument: + +```python +TOOL = { + "name": "polaris_codegraph_explore", + "description": "Run one bounded Polaris CodeGraph freshness window.", + "inputSchema": { + "type": "object", + "required": ["task_id", "stage", "query_id", "purpose", "query", "sync_if_needed"], + "additionalProperties": False, + "properties": { + "task_id": {"type": "string", "pattern": r"^TASK-[0-9]{4}$"}, + "stage": {"type": "string", "enum": ["PLANNING", "IMPLEMENTATION", "DOCUMENTATION_SYNC", "REVIEW"]}, + "query_id": {"type": "string", "pattern": r"^CIQ-[0-9]{3}$"}, + "purpose": {"type": "string", "minLength": 1, "maxLength": 240}, + "query": {"type": "string", "minLength": 1, "maxLength": 8000}, + "sync_if_needed": {"type": "boolean"}, + }, + }, +} +``` + +Return parse/invalid-request/method errors as JSON-RPC `-32700`, `-32600`, `-32601`, and `-32602`. Return user-correctable tool execution/input failures as `tools/call` results with `isError: true`. Never write logs to stdout. + +- [ ] **Step 4: Write failing envelope-order and error tests** + +Mock `execute_proxy_query` in-process and assert a successful tool call returns two text blocks—the envelope first and raw graph second. A stale/unknown graph remains `isError: false`; Provider failure has only the envelope; invalid stage/query is `isError: true`; unknown tool name is `-32602`; requests before initialization are rejected. + +- [ ] **Step 5: Implement tool-call result formatting** + +Return: + +```python +{ + "content": [ + {"type": "text", "text": render_freshness_envelope(bundle)}, + *([{"type": "text", "text": graph_text}] if graph_text is not None else []), + ], + "structuredContent": {"bundle": bundle}, + "isError": False, +} +``` + +The structured bundle must match the saved runtime bundle exactly. Ensure `json.dumps(..., ensure_ascii=False, separators=(",", ":"))` produces one physical stdout line. + +- [ ] **Step 6: Run MCP tests** + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests -k mcp_server -v` + +Expected: all PASS; no output before the envelope and no repository property in the tool schema. + +- [ ] **Step 7: Commit Task 3** + +```bash +git add scripts/code_intelligence_mcp.py tests/test_codegraph.py +git commit -m "feat: expose the Polaris CodeGraph MCP proxy" +``` + +--- + +### Task 4: Code Intelligence v3 record projection and validation + +**Files:** +- Create: `schemas/code-intelligence-record-v2.schema.json` +- Create: `schemas/code-intelligence-record-annotations.schema.json` +- Modify: `schemas/code-intelligence-record.schema.json` +- Modify: `scripts/internal/code_intelligence_protocol.py` +- Modify: `scripts/record_code_intelligence.py` +- Modify: `templates/task-sources/code-intelligence-record.json` +- Test: `tests/test_codegraph.py` + +**Interfaces:** +- Consumes: Task 2 bundle v1 and existing v1/v2 validators/fallback rules. +- Produces: `validate_historical_v2_record_value(...)`, `_validate_v3_record_value(...)`, and `record_proxy_bundle(repo, task_id, bundle_path, annotations, root=None) -> dict`. + +- [ ] **Step 1: Freeze v2 and write failing v3 projection tests** + +Copy the current schema byte-for-byte to `code-intelligence-record-v2.schema.json`, then add a test that records a Task 2 bundle through: + +```python +result = record_proxy_bundle( + self.repo, + "TASK-0001", + bundle_path, + { + "summary": "Located the affected symbol.", + "symbols": [{"path": "src/a.py", "line": 1, "name": "A"}], + "source_fallbacks": [], + }, + ROOT, +) +recorded = json.loads(Path(result["path"]).read_text(encoding="utf-8")) +self.assertEqual(recorded["record_version"], 3) +self.assertEqual(recorded["proxy"]["server_id"], "polaris-codegraph") +self.assertEqual(recorded["query_window"]["pre_status"]["pending_changes"], + {"added": 0, "modified": 0, "removed": 0}) +self.assertEqual(recorded["delivery"]["state"], "CURRENT") +``` + +- [ ] **Step 2: Run the projection test and confirm RED** + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests.test_v3_record_projects_exact_proxy_bundle -v` + +Expected: `record_proxy_bundle` missing. + +- [ ] **Step 3: Define the v3 and annotations schemas** + +The current schema must require these top-level fields and reject extras: + +```json +[ + "record_version", "task_id", "work_item_revision", "stage", + "artifact_attempt", "reviewer_slot", "provider", "repository", "target", + "status", "proxy", "query", "query_window", "delivery", + "source_fallbacks", "recorded_at" +] +``` + +Use `record_version: {"const": 3}`. `proxy` requires `server_id: "polaris-codegraph"`, `tool: "polaris_codegraph_explore"`, and a 64-hex `evidence_bundle_sha256`. `repository` requires nonblank `project_id` and 64-hex `root_sha256`. `query_window` requires `pre_status`, nullable `sync`, nullable `post_sync_status`, nullable `response_classification`, and nullable `post_query_status`. Every successful status observation requires `checked_at`, `response_sha256`, exact nonnegative `pending_changes`, `stale_reasons`, and null error; failed/unavailable observations prohibit pending counts and hashes. Reuse the current stale-point and source-fallback definitions without weakening them. + +The annotations schema permits only `summary`, `symbols`, and `source_fallbacks`; the recorder derives identity, target, status, observations, hashes, and timestamps from the bundle/current task. + +- [ ] **Step 4: Implement bundle projection and v3 semantic invariants** + +`record_proxy_bundle` must confine the bundle below the current task's runtime directory, reject symlinks, validate bundle v1, hash its exact bytes, confirm its repository/task/stage context, and build the record. Map bundle delivery to record status exactly: + +```python +RECORD_STATUS = { + "CURRENT": "USED", + "STALE": "USED", + "UNKNOWN": "FAILED", + "UNAVAILABLE": "UNAVAILABLE", +} +``` + +The v3 validator must reject: a non-proxy provider, project/root mismatch, target mismatch, non-sequential query ID for that stage record, response hash mismatch, `CURRENT` without two successful zero-pending effective pre/post observations, stale without an explicit reason/fallback, unknown without a verification error, unavailable with any attempted operation, successful sync without post-sync status, more than one sync, and missing exact fallback evidence. + +- [ ] **Step 5: Preserve historical validators and make v3 the only new write** + +Dispatch exactly: + +```python +if version == 1: + return validate_legacy_record_value(...) +if version == 2: + return _validate_v2_record_value(..., schema_name="code-intelligence-record-v2.schema.json") +if version == 3: + return _validate_v3_record_value(...) +raise RuleFailure("unsupported Code Intelligence record version") +``` + +`record(...)` must reject versions 1 and 2 with `new Code Intelligence records must use record_version 3`. Add `validate_historical_v2_record_value` so migration can validate canonical prior-revision v2 records without requiring the task's current revision. + +- [ ] **Step 6: Update the recording CLI and template** + +Support only: + +```text +record_code_intelligence.py TASK-0001 --bundle --annotations +record_code_intelligence.py --select-provider ... +``` + +Reject the old unrestricted `--input` new-write path. Update the template to a valid v3 unavailable example carrying proxy/repository/query-window/delivery fields and no attempted graph operation. + +- [ ] **Step 7: Add mutation tests for every v3 invariant** + +Deep-copy a valid v3 record, mutate one field per subtest, and assert `RuleFailure` for nonzero pending `CURRENT`, missing post-query status, response hash mismatch, bundle hash shape, repository mismatch, contradictory delivery/usage, stale without fallback, unsafe fallback result path, sync without post-sync observation, and v2 new-write attempt. Verify frozen v1/v2 samples remain byte-identical and readable. + +- [ ] **Step 8: Run record tests** + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests -k v3 -v` + +Expected: all v3 tests PASS; existing v1/v2 historical tests PASS after updating only expected new-write messages. + +- [ ] **Step 9: Commit Task 4** + +```bash +git add schemas/code-intelligence-record-v2.schema.json schemas/code-intelligence-record-annotations.schema.json schemas/code-intelligence-record.schema.json scripts/internal/code_intelligence_protocol.py scripts/record_code_intelligence.py templates/task-sources/code-intelligence-record.json tests/test_codegraph.py +git commit -m "feat: record auditable CodeGraph proxy evidence" +``` + +--- + +### Task 5: Host adapter v3 and non-destructive project MCP registration + +**Files:** +- Create: `scripts/internal/project_mcp_registration.py` +- Modify: `schemas/host-adapter.schema.json` +- Modify: `hosts/codex/adapter.json` +- Modify: `hosts/claude-code/adapter.json` +- Modify: `scripts/internal/host_adapters.py` +- Modify: `scripts/vendor_project.py` +- Modify: `scripts/init_project.py` +- Modify: `scripts/validate_project.py` +- Test: `tests/test_core.py` + +**Interfaces:** +- Consumes: Task 3's vendored entry path `tools/polaris/scripts/code_intelligence_mcp.py`. +- Produces: `project_mcp_target(repo, adapter) -> Path`, `merge_project_mcp(repo, adapter, source_text=None) -> str`, and `validate_project_mcp(repo, adapter) -> None`. + +- [ ] **Step 1: Write failing adapter-v3 manifest tests** + +Require each manifest to contain: + +```json +"project_mcp": { + "server_id": "polaris-codegraph", + "format": "codex-toml", + "target": ".codex/config.toml", + "command": "python3", + "args": ["tools/polaris/scripts/code_intelligence_mcp.py", "--repo", "."] +} +``` + +Claude differs only by `format: "claude-json"` and `target: ".mcp.json"`. Update synthetic adapters in tests to version 3 and provide a unique safe target. Reject unknown format, wrong server ID, absolute/parent paths, a launcher outside `tools/polaris`, missing `--repo .`, duplicate registration targets, and overlap with `skill_target`/`files`. + +- [ ] **Step 2: Run manifest tests and confirm RED** + +Run: `python3 -m unittest tests.test_core.PolarisCoreTests.test_host_adapter_contract_rejects_invalid_or_conflicting_manifests -v` + +Expected: schema rejects v3 or missing `project_mcp`. + +- [ ] **Step 3: Implement adapter-v3 schema and semantic validation** + +Set `adapter_version.const` to 3 and make `project_mcp` required. Validate `server_id`, `format`, target confinement, exact project-relative launcher, argument order, and target overlap in `load_host_adapters`. Keep host discovery declarative—no host ID branches in vendoring or validation. + +- [ ] **Step 4: Write failing non-destructive merge tests** + +For Codex, start with: + +```toml +model = "gpt-5" +[mcp_servers.other] +command = "other" +``` + +Assert the original bytes remain and one marked Polaris block is appended: + +```toml +# POLARIS_MCP_START polaris-codegraph +[mcp_servers.polaris-codegraph] +command = "python3" +args = ["tools/polaris/scripts/code_intelligence_mcp.py", "--repo", "."] +cwd = "." +enabled = true +required = false +enabled_tools = ["polaris_codegraph_explore"] +# POLARIS_MCP_END polaris-codegraph +``` + +For Claude, preserve unrelated top-level fields and `mcpServers.other`, then insert exactly: + +```json +"polaris-codegraph": { + "type": "stdio", + "command": "python3", + "args": ["tools/polaris/scripts/code_intelligence_mcp.py", "--repo", "."], + "env": {} +} +``` + +Assert idempotent rerendering, exact managed-block replacement, malformed TOML/JSON rejection, conflicting unmanaged same-name rejection, symlink rejection, and unrelated-content preservation. + +- [ ] **Step 5: Implement format-specific merge/validation in one focused module** + +Use `tomllib.loads` when available to validate full TOML before and after replacing the uniquely marked block. On Python 3.10, use the vendored standard-library TOML parser compatibility package for equivalent full-document validation; never rewrite unrelated TOML bytes. Use `json.loads` plus four-space `json.dumps(..., ensure_ascii=False, indent=4) + "\n"` for Claude. A same-name Claude entry is accepted only if it exactly equals the managed definition; otherwise raise `RuleFailure`. + +- [ ] **Step 6: Integrate registration into the vendor transaction** + +During `_stage_install`, read the target host config if present, render its merged staged form, and list it as a preserved path. Add each registration target to `_polaris_destinations`/affected paths so rollback backs it up. `init_project.initialize` merges registrations only when `protocol_root(repo) == repo / "tools/polaris"`; source-tree initialization without a vendored runtime must not create a dangling registration. + +- [ ] **Step 7: Validate exact vendored runtime ownership** + +In `validate_project`, parse each registration, require only the declared tool, ensure its launcher is a regular file under the vendored root, reject symlink hops, and require `--repo .`. Include the registration config in the install manifest's preserved paths, never managed files. + +- [ ] **Step 8: Run host/vendoring tests** + +Run: `python3 -m unittest tests.test_core.PolarisCoreTests -k host -v` + +Run: `python3 -m unittest tests.test_core.PolarisCoreTests.test_vendor_rolls_back_after_partial_apply_failure tests.test_core.PolarisCoreTests.test_force_vendor_preserves_unrelated_claude_configuration tests.test_core.PolarisCoreTests.test_vendored_target_is_self_contained -v` + +Expected: all PASS; injected apply failure restores both host configs byte-for-byte. + +- [ ] **Step 9: Commit Task 5** + +```bash +git add scripts/internal/project_mcp_registration.py schemas/host-adapter.schema.json hosts/codex/adapter.json hosts/claude-code/adapter.json scripts/internal/host_adapters.py scripts/vendor_project.py scripts/init_project.py scripts/validate_project.py tests/test_core.py +git commit -m "feat: register the project CodeGraph proxy" +``` + +--- + +### Task 6: Protocol 0.1.21 migration and frozen v2 inventory + +**Files:** +- Modify: `VERSION` +- Modify: `pyproject.toml` +- Modify: `templates/project.json` +- Modify: `templates/task/state.json` +- Modify: `templates/task-sources/state.json` +- Modify: `workflow/migrations.json` +- Modify: `scripts/internal/migration_protocol.py` +- Modify: `README.md` +- Modify: `README.zh-CN.md` +- Modify: `docs/USAGE.md` +- Modify: `plan.md` +- Test: `tests/test_codegraph.py` +- Test: `tests/test_core.py` + +**Interfaces:** +- Consumes: Task 4's `validate_historical_v2_record_value` and Task 5 adapter v3 registration. +- Produces: adjacent migration `0.1.20-to-0.1.21`, protocol/package version `0.1.21`, unchanged workflow `0.1.3`. + +- [ ] **Step 1: Write failing version-route and frozen-v2 migration tests** + +Create a valid v2 record in both current and prior revision directories, freeze their bytes, migrate, then assert: + +```python +self.assertEqual(result["from"], "0.1.20") +self.assertEqual(result["to"], "0.1.21") +self.assertEqual(read_json(self.repo / ".polaris/project.json")["workflow_version"], "0.1.3") +self.assertEqual(v2_path.read_bytes(), frozen_v2_bytes) +self.assertIn( + {"task_id": "TASK-0001", "path": "code-intelligence/r001/planning.json", + "sha256": hashlib.sha256(frozen_v2_bytes).hexdigest()}, + migration["retired_code_intelligence_records"], +) +``` + +Add rejection tests for noncanonical v2 path, mutated v2 bytes on resumed migration, cross-root/dangling symlink, skipped `0.1.20 -> 0.1.22`, and workflow version change. + +- [ ] **Step 2: Run migration tests and confirm RED** + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests -k migration -v` + +Expected: no `0.1.20-to-0.1.21` route and/or v2 not inventoried. + +- [ ] **Step 3: Add the adjacent migration and version replacements** + +Append exactly: + +```json +{ + "migration_id": "0.1.20-to-0.1.21", + "from_polaris_version": "0.1.20", + "to_polaris_version": "0.1.21", + "from_workflow_version": "0.1.3", + "to_workflow_version": "0.1.3", + "project_strategy": "replace_version", + "task_strategy": "append_version_event" +} +``` + +Change only protocol/package/template version literals to `0.1.21`; leave every workflow version at `0.1.3`. + +- [ ] **Step 4: Inventory v2 without weakening historical validation** + +Pass the migration step into `_retired_code_intelligence_records`. For `0.1.20-to-0.1.21`, include canonical validated v2 records (and retain already supported v1 inventory) with task-relative path and exact SHA-256. Other historical migration records remain valid and byte-identical. On resume, compare the recomputed inventory with the frozen migration record before appending events. + +- [ ] **Step 5: Update authority/version documentation** + +Update the current-version lines and migration ledger in both READMEs, `docs/USAGE.md`, and `plan.md`. State that `0.1.21` adds the project-scoped proxy, host adapter v3, and v3 records; state explicitly that Workflow remains `0.1.3`, CodeGraph remains optional/non-gating, and v1/v2 are historical only. + +- [ ] **Step 6: Run version and migration tests** + +Run: `python3 -m unittest tests.test_core.PolarisCoreTests -k migration -v` + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests -k migration -v` + +Expected: all PASS; route is adjacent and frozen v2 bytes do not change. + +- [ ] **Step 7: Commit Task 6** + +```bash +git add VERSION pyproject.toml templates/project.json templates/task/state.json templates/task-sources/state.json workflow/migrations.json scripts/internal/migration_protocol.py README.md README.zh-CN.md docs/USAGE.md plan.md tests/test_codegraph.py tests/test_core.py +git commit -m "feat: migrate CodeGraph evidence to protocol 0.1.21" +``` + +--- + +### Task 7: Stage Skills, host renderings, and user-facing proxy contract + +**Files:** +- Modify: `skills/code-intelligence/SKILL.md` +- Modify: `skills/architecture-planning/SKILL.md` +- Modify: `skills/implementation/SKILL.md` +- Modify: `skills/documentation-sync/SKILL.md` +- Modify: `skills/adversarial-review/SKILL.md` +- Modify: `templates/AGENTS.md` +- Modify: relevant `hosts/*/skill-appendices/*.md` +- Modify: `README.md` +- Modify: `README.zh-CN.md` +- Modify: `docs/USAGE.md` +- Test: `tests/test_codegraph.py` + +**Interfaces:** +- Consumes: MCP tool name/arguments from Task 3, bundle-to-record CLI from Task 4, and registration behavior from Task 5. +- Produces: one consistent Agent/human contract across canonical, rendered, vendored, and localized surfaces. + +- [ ] **Step 1: Invoke the required Skill-writing workflow** + +Before changing any Skill, read and follow `superpowers:writing-skills`. Record the required baseline adversarial evaluation using these four prompts: skip the proxy, trust a clean graph with pending changes, reuse an old Implementation envelope, and treat `UNKNOWN` as current. + +- [ ] **Step 2: Write failing semantic contract tests** + +Render every Skill through every adapter and assert the relevant surfaces contain all exact anchors: + +```python +anchors = ( + "polaris_codegraph_explore", + "freshness envelope", + "NON_AUTHORITATIVE_CONTEXT", + "NAVIGATION_ONLY", + "source/Git fallback", + "raw `codegraph_explore`", + "cannot back `CURRENT` Polaris evidence", +) +``` + +Assert Documentation Sync also contains `sync_if_needed: true`, changed source paths/documented symbols, and “no separate status/sync MCP tool”. Assert Validation surfaces contain none of `polaris_codegraph_explore`, `codegraph status`, or `codegraph sync`. Mutation-check removal of the proxy tool name and fallback branch from one rendered surface. + +- [ ] **Step 3: Run semantic tests and confirm RED** + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests.test_all_agent_surfaces_require_proxy_provenance tests.test_codegraph.CodeGraphTests.test_documentation_sync_uses_one_proxy_query tests.test_codegraph.CodeGraphTests.test_validation_remains_graph_free -v` + +Expected: existing direct status/sync instructions violate the new anchors. + +- [ ] **Step 4: Rewrite canonical stage behavior** + +Use the exact lifecycle: + +1. If Code Intelligence is disabled or `.codegraph/` is absent, skip proxy and use source/Git. +2. Call only `polaris_codegraph_explore` for Polaris graph evidence. +3. Read the envelope before graph content. +4. Treat `CURRENT` as non-authoritative context. +5. For `STALE`/`UNKNOWN`, complete every named fallback before using affected conclusions. +6. Record via `record_code_intelligence.py --bundle ... --annotations ...` only if a proxy operation occurred. +7. Raw Provider MCP/shell output remains allowed but is always out-of-band and cannot support `CURRENT`. + +Implementation must require a fresh call after edits. Documentation Sync must perform one bounded changed-path/symbol query with `sync_if_needed: true` only when supported source changed. Review must independently call the proxy and never inherit the implementer's envelope. Validation must remain graph-free. + +- [ ] **Step 5: Update human docs without obsolete paths** + +Document project trust/first-use approval, `.codex/config.toml`, `.mcp.json`, exact envelope meanings, source/Git fallback, optional/non-gating status, and user ownership of CodeGraph install/init/config/watchers. Remove stage instructions that tell users to separately choose `status` or `sync-if-needed` before raw MCP explore. + +- [ ] **Step 6: Regenerate/render and run adversarial evaluation** + +Run: `python3 scripts/materialize_task_layout.py` + +Run: `python3 -m unittest tests.test_codegraph.CodeGraphTests.test_all_agent_surfaces_require_proxy_provenance tests.test_codegraph.CodeGraphTests.test_documentation_sync_uses_one_proxy_query tests.test_codegraph.CodeGraphTests.test_validation_remains_graph_free -v` + +Repeat the four baseline prompts through the Skill evaluation method required by `writing-skills`. The changed behavior must refuse every shortcut and state the required fallback. + +- [ ] **Step 7: Commit Task 7** + +```bash +git add skills hosts templates/AGENTS.md README.md README.zh-CN.md docs/USAGE.md tests/test_codegraph.py +git commit -m "docs: route Polaris stages through the CodeGraph proxy" +``` + +--- + +### Task 8: End-to-end acceptance, packaging, and PR readiness + +**Files:** +- Modify if tests expose gaps: only files already named in Tasks 1-7 +- Test: `tests/test_codegraph.py` +- Test: `tests/test_core.py` + +**Interfaces:** +- Consumes: all prior tasks. +- Produces: a clean, reviewable feature branch ready for draft PR and CI monitoring. + +- [ ] **Step 1: Add a fake-CLI end-to-end MCP test** + +Create an executable temporary `codegraph` fixture that records cwd/argv and returns scripted status/sync/explore outputs. Vendor Polaris into a disposable Git repo, initialize the task and `.codegraph/`, launch the registered MCP command, call the tool, record its bundle into v3, and run `validate_project.py`. Assert envelope order, exact call maxima, runtime ignore, record hash binding, and no graph calls during project validation. + +- [ ] **Step 2: Add optional real-CLI smoke coverage** + +Extend the existing disposable-repo guard. If `codegraph` is absent, skip. If present, create the temporary repo outside the Polaris workspace, never run init/configuration, and run only when the fixture already has an explicitly created safe marker/index. Assert all CLI commands use the disposable repo cwd and clean up through `TemporaryDirectory`. + +- [ ] **Step 3: Run focused suites** + +Run: `python3 -m unittest tests.test_codegraph -v` + +Run: `python3 -m unittest tests.test_core -v` + +Expected: all PASS; the optional real-CLI test may be the only skip. + +- [ ] **Step 4: Run repository-wide verification** + +Run: `python3 tests/run_tests.py` + +Run: `python3 -m compileall -q polaris_cli.py scripts tests` + +Run: `python3 scripts/materialize_task_layout.py && git diff --exit-code` + +Run: `git diff --check dev...HEAD` + +Run: `rg -n "record_version.?[:=].?2|0\.1\.20|adapter_version.?[:=].?2|status.*or.*sync-if-needed" README.md README.zh-CN.md docs plan.md skills hosts templates scripts schemas tests --glob '!schemas/code-intelligence-record-v2.schema.json' --glob '!schemas/code-intelligence-record-v1.schema.json'` + +Expected: full suite PASS; compile/materialize/diff checks exit 0; the final scan contains only deliberate historical/migration assertions. + +- [ ] **Step 5: Review the final diff against all 17 acceptance tests** + +For each numbered item in the design's Testing and Evaluation section, point to one passing test name. Confirm no implementation added a dependency, daemon, watcher, retry loop, Validation graph call, alternate Provider, global MCP registration, or raw-tool restriction. + +Use this coverage map during review: + +| Spec test | Plan coverage | +| --- | --- | +| 1-3 pre/post pending and clean currentness | Task 2 query-window table tests | +| 4 failed/malformed status | Task 2 no-explore `UNKNOWN` tests | +| 5-6 exact, stale, wrapped, and suspicious banners | Task 1 classifier tests and Task 2 delivery tests | +| 7 shared cwd | Task 1 call-sequence test and Task 8 fake CLI | +| 8 project mismatch | Task 2 discard test | +| 9 envelope ordering | Task 3 tool result test and Task 8 transcript | +| 10 v3 pending/post-query rejection | Task 4 mutation tests | +| 11 historical v1/v2 bytes/readability | Tasks 4 and 6 migration tests | +| 12 stage/vendored/host provenance | Task 7 rendered semantic contract | +| 13 graph-free Validation | Task 7 negative surface test and Task 8 validation call log | +| 14 suite without CodeGraph | Task 8 full suite | +| 15 disposable real CLI | Task 8 guarded smoke test | +| 16 non-destructive host registration | Task 5 merge/rollback tests | +| 17 single Documentation Sync proxy | Task 7 semantic test | + +- [ ] **Step 6: Commit any verification-only test corrections** + +```bash +git add tests README.md README.zh-CN.md docs plan.md +git commit -m "test: verify the CodeGraph MCP proxy end to end" +``` + +Skip this commit if Step 4 required no changes. + +- [ ] **Step 7: Hand off for branch finishing and CI** + +After a clean verification run, use `superpowers:requesting-code-review`, then `superpowers:finishing-a-development-branch`. Publish a draft PR targeting `dev` with the GitHub publishing workflow. Monitor every GitHub Actions check; for any failure, use `github:gh-fix-ci`, reproduce the failing check locally, add or tighten a regression test, push the fix, and continue until all required checks pass. diff --git a/docs/superpowers/specs/2026-08-18-codegraph-freshness-design.md b/docs/superpowers/specs/2026-08-18-codegraph-freshness-design.md index 5f63fc2..979cbaa 100644 --- a/docs/superpowers/specs/2026-08-18-codegraph-freshness-design.md +++ b/docs/superpowers/specs/2026-08-18-codegraph-freshness-design.md @@ -3,11 +3,13 @@ ## 状态 - 日期:2026-08-18 -- 状态:已在对话中确认,等待书面规格复核 +- 状态:历史设计,已由 `2026-08-18-codegraph-polaris-mcp-proxy-design.md` 取代;不得作为当前操作指南 - 适用版本:Polaris v0.1 的下一协议版本 - 产品 authority:`plan.md` - 唯一正式 CodeGraph Provider:[`colbymchenry/codegraph`](https://github.com/colbymchenry/codegraph) +> 当前阶段必须只调用 Polaris 项目代理 `polaris_codegraph_explore`。本文以下对阶段直接编排 `status`、`sync-if-needed` 或 raw `codegraph_explore` 的描述仅用于解释 v2 历史协议,不能支持新的 Polaris `CURRENT` 证据。 + ## 背景 Polaris 已有可选 Code Intelligence Provider 协议,但当前 `codegraph` descriptor 指向另一个同名且协议不兼容的产品。目标 Provider 实际应为 `colbymchenry/codegraph`。它默认向 MCP 暴露单一高价值入口 `codegraph_explore`,并提供 `codegraph explore`、`codegraph status --json` 和 `codegraph sync` CLI。 diff --git a/docs/superpowers/specs/2026-08-18-codegraph-polaris-mcp-proxy-design.md b/docs/superpowers/specs/2026-08-18-codegraph-polaris-mcp-proxy-design.md index 30a3cff..b3d9918 100644 --- a/docs/superpowers/specs/2026-08-18-codegraph-polaris-mcp-proxy-design.md +++ b/docs/superpowers/specs/2026-08-18-codegraph-polaris-mcp-proxy-design.md @@ -3,10 +3,22 @@ ## Status - Date: 2026-08-18 -- Status: proposed for implementation +- Status: approved for implementation on 2026-08-19 - Scope: Polaris Code Intelligence protocol and stage behavior - Provider: `colbymchenry/codegraph` +## Version Boundary + +- Polaris protocol and package version: `0.1.20` to `0.1.21`. +- Workflow graph version: remains `0.1.3`; no state or transition changes. +- Host adapter manifest version: v2 to v3. +- Code Intelligence records written by new stages: v3. +- Code Intelligence v1 and v2 records: immutable historical read support only. + +The adjacent `0.1.20` to `0.1.21` migration inventories immutable v2 records +without rewriting them. It upgrades the host registration and protocol files, +but does not alter the workflow graph. + ## Problem Polaris currently lets Planning, Implementation, and Review choose either a @@ -204,10 +216,27 @@ an arbitrary project path from the tool call. It rejects a missing, moved, or symlinked project root. Removing or disabling the project-local Polaris MCP registration disables the proxy without affecting CodeGraph itself. -The adapter contract must represent the project-scoped MCP registration for -each supported host rather than embedding host-specific configuration writes -in the Code Intelligence adapter. Vendoring and project validation verify that -the registration launches only the repository's vendored Polaris runtime. +Host adapter v3 adds one required declarative `project_mcp` registration. It +identifies the fixed `polaris-codegraph` server ID, the host-native project +configuration target and format, and the project-relative vendored launcher. +For the supported hosts, the targets are `.codex/config.toml` for Codex and +`.mcp.json` for Claude Code. The host renderer owns these syntax differences; +the Code Intelligence adapter remains host-neutral. + +Initialization and vendoring merge only the named `polaris-codegraph` entry +and preserve unrelated user servers and settings. They refuse malformed host +configuration, path/symlink escape, or an existing same-name registration with +a different definition instead of silently overwriting it. Upgrade removes or +replaces only the previously managed Polaris entry. Project validation parses +the resulting host configuration and proves that this entry launches only +`tools/polaris/scripts/code_intelligence_mcp.py`, fixes the repository argument +to the project root, and exposes no user-selected repository parameter. + +The launcher and server independently resolve and compare the configured root +with the actual project root. A host starting the process from an unexpected +working directory therefore fails closed rather than querying another +repository. Host-native trust or first-use approval remains a user decision; +Polaris does not bypass it. ## Envelope @@ -260,8 +289,8 @@ mismatch, response-hash mismatch, or a `CURRENT` claim with any non-zero pending count. Migration inventories immutable v2 records without rewriting them. The -Polaris protocol version increments; the workflow graph version does not -change because no workflow state or transition changes. +Polaris protocol and package version become `0.1.21`; the workflow graph stays +at `0.1.3` because no workflow state or transition changes. ## Stage Behavior @@ -271,8 +300,13 @@ change because no workflow state or transition changes. `CURRENT` stage record. - Implementation may query after edits only through a fresh proxy invocation for Polaris evidence; it does not reuse the entry envelope. -- Documentation Sync uses the same proxy/status machinery for its final - bounded sync evidence when supported source changed. +- When supported source changed, Documentation Sync makes one bounded + `polaris_codegraph_explore` call with `stage: DOCUMENTATION_SYNC`, + `sync_if_needed: true`, and a query limited to the changed source paths and + documented symbols. Its post-query status is the final sync observation; + there is no second status/sync MCP tool. When no supported source changed, + it creates no Code Intelligence record. `STALE` or `UNKNOWN` results require + the same source/Git fallback before documentation conclusions are used. - Validation remains graph-free. - On `STALE` or `UNKNOWN`, stage conclusions concerning returned files or relationships require the envelope's source/Git fallback before use. @@ -319,6 +353,12 @@ Deterministic unit and integration tests must cover: 13. Validation remains graph-free; 14. the full Polaris suite passes without requiring CodeGraph; 15. an optional real-CLI smoke test uses only a disposable temporary repo. +16. host registration merges into existing Codex TOML and Claude JSON without + changing unrelated servers or settings, rejects conflicting same-name + entries and unsafe paths, and validates the exact vendored launcher/root; +17. Documentation Sync uses the single explore proxy for its final bounded + query, skips graph evidence when supported source did not change, and + never calls a separate sync/status MCP tool. Skill evaluation must include pressure cases where an Agent is asked to skip the proxy, trust a clean-looking graph despite pending changes, reuse an old diff --git a/docs/superpowers/specs/2026-08-18-codegraph-polaris-mcp-proxy-design.zh-CN.md b/docs/superpowers/specs/2026-08-18-codegraph-polaris-mcp-proxy-design.zh-CN.md index 543a9f8..015f747 100644 --- a/docs/superpowers/specs/2026-08-18-codegraph-polaris-mcp-proxy-design.zh-CN.md +++ b/docs/superpowers/specs/2026-08-18-codegraph-polaris-mcp-proxy-design.zh-CN.md @@ -3,10 +3,21 @@ ## 状态 - 日期:2026-08-18 -- 状态:提议实施 +- 状态:已于 2026-08-19 批准实施 - 范围:Polaris Code Intelligence 协议与阶段行为 - Provider:`colbymchenry/codegraph` +## 版本边界 + +- Polaris 协议与包版本:`0.1.20` 升到 `0.1.21`。 +- Workflow graph 版本:保持 `0.1.3`;不改变 state 或 transition。 +- 宿主 adapter manifest 版本:v2 升到 v3。 +- 新阶段写入的 Code Intelligence record:v3。 +- Code Intelligence v1 和 v2 record:仅保留不可变的历史读取支持。 + +相邻的 `0.1.20` 到 `0.1.21` migration 会盘点不可变 v2 record,但不改写它们。 +它会升级宿主注册和协议文件,但不修改 workflow graph。 + ## 问题 目前,Polaris 允许 Planning、Implementation 和 Review 在独立的 @@ -183,9 +194,22 @@ server 进程在启动时接收仓库根目录,工具调用本身不接受任 缺失、已移动或为 symlink 时,server 必须拒绝。移除或禁用项目本地的 Polaris MCP 注册,只会禁用代理,不影响 CodeGraph 本身。 -adapter 契约必须为每个受支持宿主表示项目级 MCP 注册,而不能把宿主专属配置写入 -Code Intelligence adapter。Vendoring 和项目校验会验证该注册只启动仓库中的 -vendored Polaris runtime。 +宿主 adapter v3 新增一个必需的声明式 `project_mcp` 注册。它声明固定的 +`polaris-codegraph` server ID、宿主原生项目配置的目标与格式,以及项目相对路径 +下的 vendored launcher。当前受支持宿主中,Codex 目标为 +`.codex/config.toml`,Claude Code 目标为 `.mcp.json`。宿主 renderer 负责这些 +语法差异;Code Intelligence adapter 保持宿主无关。 + +初始化和 vendoring 只合并名为 `polaris-codegraph` 的条目,并保留用户其他 server +与设置。若宿主配置畸形、路径或 symlink 逃逸,或者存在同名但定义不同的注册, +系统必须拒绝,而不是静默覆盖。升级只移除或替换先前由 Polaris 管理的条目。项目 +校验会解析最终宿主配置,证明该条目只能启动 +`tools/polaris/scripts/code_intelligence_mcp.py`,仓库参数被固定为项目根目录,且 +不存在用户可选的仓库参数。 + +launcher 与 server 都会独立解析并比对配置根目录和实际项目根目录。因此,即使 +宿主从意外工作目录启动进程,也会 fail closed,而不会查询另一个仓库。宿主原生 +的项目信任或首次使用批准仍由用户决定;Polaris 不绕过它。 ## Envelope @@ -233,8 +257,9 @@ v3 查询证据包含: validator 会拒绝观察结果缺失、状态自相矛盾、项目不匹配、响应 hash 不匹配,或在 任何 pending 计数非零时声明 `CURRENT`。 -迁移过程会盘点不可变 v2 record,但不改写它们。Polaris 协议版本递增;workflow -graph 版本不变,因为 workflow 状态和 transition 都没有变化。 +迁移过程会盘点不可变 v2 record,但不改写它们。Polaris 协议与包版本升到 +`0.1.21`;workflow graph 保持 `0.1.3`,因为 workflow 状态和 transition 都没有 +变化。 ## 阶段行为 @@ -243,8 +268,12 @@ graph 版本不变,因为 workflow 状态和 transition 都没有变化。 Polaris 而言始终是未验证的,不能支撑 `CURRENT` 阶段 record。 - Implementation 在编辑后若要为 Polaris 生成证据,只能发起一次新的代理调用; 不能复用阶段入口的 envelope。 -- Documentation Sync 在受支持源码发生变化时,使用相同的代理/status 机制生成 - 最终有界 sync 证据。 +- 当受支持源码发生变化时,Documentation Sync 发起一次有界的 + `polaris_codegraph_explore` 调用,使用 `stage: DOCUMENTATION_SYNC`、 + `sync_if_needed: true`,并把查询限制为已变更源码路径和文档涉及的 symbol。查询 + 后 status 就是最终 sync 观察结果;不增加第二个 status/sync MCP 工具。如果受 + 支持源码没有变化,则不创建 Code Intelligence record。`STALE` 或 `UNKNOWN` + 结果必须在使用文档结论前完成同样的源码/Git 回退。 - Validation 仍然不使用 graph。 - 当状态为 `STALE` 或 `UNKNOWN` 时,任何涉及返回文件或关系的阶段结论,都必须先 完成 envelope 要求的源码/Git 回退。 @@ -285,6 +314,11 @@ Vendored `AGENTS.md`、宿主 overlay 和 canonical Skill 必须共享此契约 13. Validation 仍然不使用 graph; 14. 完整 Polaris 测试套件不依赖 CodeGraph 即可通过; 15. 可选的真实 CLI smoke test 只使用一次性临时仓库。 +16. 宿主注册会把配置合并到现有 Codex TOML 和 Claude JSON 中,且不修改无关 + server 或设置;拒绝冲突的同名条目和不安全路径,并验证精确的 vendored + launcher/root; +17. Documentation Sync 使用唯一的 explore 代理完成最终有界查询;受支持源码未 + 变化时跳过 graph 证据;并且绝不调用单独的 sync/status MCP 工具。 Skill 评估必须包含以下压力场景:要求 Agent 跳过代理;在存在 pending changes 时 仍信任看似干净的 graph;复用旧 Implementation envelope;或者把 `UNKNOWN` 当成 diff --git a/hosts/claude-code/adapter.json b/hosts/claude-code/adapter.json index 760de14..f328e98 100644 --- a/hosts/claude-code/adapter.json +++ b/hosts/claude-code/adapter.json @@ -1,5 +1,5 @@ { - "adapter_version": 2, + "adapter_version": 3, "host_id": "claude-code", "display_name": "Claude Code", "skill_target": ".claude/skills", @@ -17,6 +17,17 @@ ], "skill_overlay_root": null, "skill_appendix_root": "skill-appendices", + "project_mcp": { + "server_id": "polaris-codegraph", + "format": "claude-json", + "target": ".mcp.json", + "command": "python3", + "args": [ + "tools/polaris/scripts/code_intelligence_mcp.py", + "--repo", + "." + ] + }, "files": [ { "source": "agents/polaris-implementer.md", diff --git a/hosts/codex/adapter.json b/hosts/codex/adapter.json index 3ebf39c..d87c0ee 100644 --- a/hosts/codex/adapter.json +++ b/hosts/codex/adapter.json @@ -1,5 +1,5 @@ { - "adapter_version": 2, + "adapter_version": 3, "host_id": "codex", "display_name": "Codex", "skill_target": ".agents/skills", @@ -15,5 +15,16 @@ "entry_frontmatter": [], "skill_overlay_root": "skill-overlays", "skill_appendix_root": "skill-appendices", + "project_mcp": { + "server_id": "polaris-codegraph", + "format": "codex-toml", + "target": ".codex/config.toml", + "command": "python3", + "args": [ + "tools/polaris/scripts/code_intelligence_mcp.py", + "--repo", + "." + ] + }, "files": [] } diff --git a/plan.md b/plan.md index 1abe8f9..d383d9c 100644 --- a/plan.md +++ b/plan.md @@ -2,7 +2,7 @@ > 状态:Implementation underway > 目标版本:v0.1 -> 当前协议:`0.1.20`;Workflow:`0.1.3` +> 当前协议:`0.1.21`;Workflow:`0.1.3` > 产品形态:Repo-native Skill System > 宿主 Runtime:声明式可扩展;v0.1 内置 Codex、Claude Code > @@ -288,9 +288,9 @@ target-repo/ ### 宿主适配契约 -每个宿主占用独立、平级的 `hosts//`,并提供由 `host-adapter.schema.json` 校验的 `adapter.json`。清单版本 `adapter_version=2` 声明 Skill 目标目录、调用前缀、入口 Skill、宿主能力、额外 frontmatter、可选 metadata overlay、执行附录和宿主专用文件。能力至少包括结构化用户输入、worker 创建、状态查询、续接和稳定身份;依赖关系必须机械自洽。共享 Skill 只使用 `{{skill:}}` 占位符和宿主无关 worker 语义;vendoring 时再渲染调用语法并追加宿主执行机制。 +每个宿主占用独立、平级的 `hosts//`,并提供由 `host-adapter.schema.json` 校验的 `adapter.json`。清单版本 `adapter_version=3` 除 Skill 目标、调用前缀、入口 Skill、宿主能力、frontmatter、overlay、appendix 与专用文件外,还声明唯一项目级 `polaris-codegraph` MCP 注册。能力依赖必须机械自洽;MCP 注册固定启动器、vendored server、`--repo .` 与宿主配置目标。共享 Skill 只使用 `{{skill:}}` 占位符和宿主无关 worker 语义;vendoring 时再渲染调用语法、执行机制与宿主 MCP 格式。 -`vendor_project.py`、`init_project.py` 和 `validate_project.py` 必须通过 `scripts/internal/host_adapters.py` 发现 canonical Skills 与所有清单,不得按宿主 ID 编写条件分支。入口必须指向实际 Skill;overlay 只能向已知 Skill 增加 canonical 源中不存在的普通文件,不能替换 `SKILL.md` 或其他源内容;adapter 源树与全部目标路径禁止 symlink。新增满足 v2 文件型契约的宿主只增加目录和资产;目标路径冲突、越界路径、缺失源文件、能力矛盾和未知清单版本都必须机械拒绝。需要超出 v2 表达能力的新机制时,先升级 adapter schema/version,再保持旧版本迁移边界,不把宿主差异写回共享 Workflow 或 Authority schema。 +`vendor_project.py`、`init_project.py` 和 `validate_project.py` 必须通过 `scripts/internal/host_adapters.py` 发现 canonical Skills 与所有清单,不得按宿主 ID 编写条件分支。入口必须指向实际 Skill;overlay 只能向已知 Skill 增加 canonical 源中不存在的普通文件,不能替换 `SKILL.md` 或其他源内容;adapter 源树、MCP 启动器与全部目标路径禁止 symlink。新增满足 v3 契约的宿主只增加目录和资产;目标路径冲突、越界路径、缺失源文件、能力矛盾、同名 MCP 冲突和未知清单版本都必须机械拒绝。需要超出 v3 表达能力的新机制时,先升级 adapter schema/version,再保持旧版本迁移边界,不把宿主差异写回共享 Workflow 或 Authority schema。 JSON 文件是机械门禁的权威输入。结构化 artifact 不生成同名 Markdown 副本;用户可直接查看四格缩进 JSON,主任务也可按需格式化展示。旧 revision 和旧 attempt 文件不可覆盖,`state.json` 仅保存当前有效 artifact 的指针。 @@ -452,7 +452,7 @@ v0.1 不设置 `FAILED`:可修复失败通过治理回路处理,外部阻塞 `.polaris/workflow.json` 保存当前项目实际使用且版本锁定的节点、边、依赖和门禁 ID;`tools/polaris/workflow/default-workflow.json` 只用于初始化。`transition_task.py` 只接受图中边并先运行对应 validators,Skill 不直接编辑 `state` 字段。v0.1 遇到 `polaris_version` 或 `workflow_version` 不匹配时拒绝正常执行,不做隐式迁移。 -版本升级必须先 vendoring 目标协议,再显式运行 vendored `migrate_project.py`。`workflow/migrations.json` 是迁移路径唯一且 append-only 的注册表,一次只执行一个从当前版本到目标版本的相邻步骤;历史步骤必须保留以校验已提交记录。Migration protocol v2 保留 `replace_version` / `append_version_event`,并增加 `replace_version_and_workflow` / `append_mapped_workflow_event`。`0.1.19 → 0.1.20` 原子替换冻结 workflow 为 `0.1.3`,追加带源/目标状态及旧版本字段的迁移事件;旧 `IMPLEMENTED`、`DOCS_SYNCED` 映射到 `IMPLEMENTING`,旧 `REVIEWED` 映射到 `VALIDATING`,旧 R0/R1 `VERIFIED` 映射到 `VALIDATING` 以重新提交 `PASS_AND_CLOSE`,仅 R2 保持 `VERIFIED`。迁移以 `.polaris/migrations/MIG-*.json` 记录 `IN_PROGRESS/COMPLETED`、各任务 sequence 和状态映射;重跑必须可恢复且不得重复事件。未知路径、跨版本跳跃、未声明的 workflow 变化、任务集合并发变化和不完整记录都必须机械拒绝。 +版本升级必须先 vendoring 目标协议,再显式运行 vendored `migrate_project.py`。`workflow/migrations.json` 是迁移路径唯一且 append-only 的注册表,一次只执行一个从当前版本到目标版本的相邻步骤;历史步骤必须保留以校验已提交记录。Migration protocol v2 保留 `replace_version` / `append_version_event`,并增加 `replace_version_and_workflow` / `append_mapped_workflow_event`。`0.1.19 → 0.1.20` 原子替换冻结 workflow 为 `0.1.3`,追加带源/目标状态及旧版本字段的迁移事件;旧 `IMPLEMENTED`、`DOCS_SYNCED` 映射到 `IMPLEMENTING`,旧 `REVIEWED` 映射到 `VALIDATING`,旧 R0/R1 `VERIFIED` 映射到 `VALIDATING` 以重新提交 `PASS_AND_CLOSE`,仅 R2 保持 `VERIFIED`。`0.1.20 → 0.1.21` 保持 Workflow `0.1.3`,新增项目级 CodeGraph 代理、Host Adapter v3 与 record v3,并把 canonical v1/v2 record 作为仅可读取的不可变历史证据按路径和 SHA-256 清点;迁移恢复前必须重算并比对清单。迁移以 `.polaris/migrations/MIG-*.json` 记录 `IN_PROGRESS/COMPLETED`、各任务 sequence 和状态映射;重跑必须可恢复且不得重复事件。未知路径、跨版本跳跃、未声明的 workflow 变化、任务集合并发变化和不完整记录都必须机械拒绝。 迁移占用任务转换锁时必须写入结构化 owner:迁移 ID、任务 ID、主机名、PID 和创建时间。重跑只允许接管同一迁移在同一主机上、且原 PID 已确认不存在的锁;活跃 PID、其他迁移、其他主机、空锁或损坏锁一律拒绝。这样既能从进程崩溃或机器重启恢复,又不把真实并发误判为遗留锁。 @@ -486,13 +486,12 @@ AGENTS.md ### 可选 Code Intelligence 协议 -- v0.1 的唯一正式 Provider 是 [colbymchenry/codegraph](https://github.com/colbymchenry/codegraph)。`providers/code-intelligence/codegraph.json` 声明其 MCP `codegraph_explore` 和 CLI `status`、`explore`、`sync` 能力;核心 record 使用 Provider-neutral 的新鲜度和回退字段。 -- `.codegraph/` 由用户创建和维护。Polaris 允许用户显式运行 `polaris code-intelligence add codegraph --repo .`,但绝不安装、初始化、启动或配置 Provider、watcher、daemon、锁或 MCP;缺少 marker 或策略禁用时直接回退源码,不生成新的阶段 record。 -- Provider 原生 watcher 与连接时 reconciliation 是保持索引接近工作树的主机制。Polaris 只在阶段入口、已知索引冻结或最终 Documentation Sync 的有界点读取 status;仅在 status 表示 pending 时至多执行一次 `codegraph sync`,随后至多复查一次,绝不等待或轮询。 -- 记录的结论限定为检查时:`CURRENT_AT_CHECK`、`PARTIAL_STALE`、`INDEX_STALE`、`NOT_VERIFIED` 或 `UNAVAILABLE`,不得宣称与某个 Git commit 严格一致。逐文件 stale point 必须记录路径和原因;文件仍存在时 Agent 直接读取源码并记录 `READ_SOURCE`,已删除时检查注册 subject 的 Git diff 并记录 `INSPECT_GIT_DIFF`;索引级失效使用 `SEARCH_SOURCE` 和 Git 证据。 -- Planning、Implementation 与 Reviewer 只在冻结范围内使用图关系;返回路径必须经源码确认才可进入 Working Set,Reviewer 必须独立查询。响应的局部 stale 不会丢弃其余图线索,但 stale 路径不能直接作为编辑或 Review 结论。 -- 只有阶段实际执行 Provider `status`、`sync` 或 `explore` 操作时才写耐久 record,并准确记录成功、失败和新鲜度;未执行操作时省略 artifact 引用。图不扩展 scope,也不是 Workflow gate;Validation 完全不调用 CodeGraph,仍只依赖源码、Git、构建、测试、静态检查和 Human Check。 -- Git 只保存绑定 Provider、阶段、subject、目的、有限摘要、响应哈希、新鲜度、stale point 与源码回退证据;原始响应只进入 ignored runtime。已提交 v1 record 是不可变历史证据,迁移后标为 `retired_provider_evidence`,不能支持新的新鲜度结论。 +- v0.1 的唯一正式 Provider 是 [colbymchenry/codegraph](https://github.com/colbymchenry/codegraph)。项目级 `polaris-codegraph` MCP 只暴露 `polaris_codegraph_explore`;raw `codegraph_explore` 与 shell 仍可带外使用,但不能支持 Polaris `CURRENT` 证据。 +- `.codegraph/` 与 CodeGraph 安装、初始化、配置、raw MCP、watcher 和 daemon 由用户拥有。Polaris 只非破坏地管理自身项目代理注册;缺少 marker 或策略禁用时直接回退源码,不生成阶段 record。 +- 代理在一个有界窗口内完成 pre-status、可选一次 `codegraph sync`、一次 explore 和 post-status,并先返回 freshness envelope。阶段不分别选择 status/sync;不等待、轮询或重试。 +- envelope 状态为 `CURRENT / NON_AUTHORITATIVE_CONTEXT`、`STALE / NAVIGATION_ONLY`、`UNKNOWN / NAVIGATION_ONLY` 或 `UNAVAILABLE / NO_GRAPH`。`STALE`/`UNKNOWN` 必须先完成具名 `READ_SOURCE`、删除路径 `INSPECT_GIT_DIFF` 或索引级 `SEARCH_SOURCE` 回退;`UNKNOWN` 不得提升为 current。 +- Planning、Implementation 与 Reviewer 只在冻结范围内使用图关系;Implementation 修改关系后必须 fresh proxy call,Reviewer 必须独立调用且不得继承 Implementer envelope。Documentation Sync 仅在 supported source 改变时,以 `sync_if_needed: true` 对 changed paths/symbols执行一次查询。Validation 完全不调用 CodeGraph。 +- 代理 bundle 与原始响应只进入 ignored runtime。Agent 完成 fallback 后提供 annotations,由 `record_code_intelligence.py --bundle ... --annotations ...` 投影不可变 v3 record;没有代理调用就省略 record。v1/v2 仅作为不可变历史证据读取。 ## 9. 确定性脚本 @@ -510,8 +509,9 @@ AGENTS.md | `materialize_task_layout.py` | 从 `internal/task_layout.py` 生成模板样例树和真实任务目录,并校验生成物与平铺模板正文一致 | | `update_implementation_progress.py` | 通过明确事件原子更新 ignored 的线性步骤进度;拒绝 session 接管、跳步、回退、未知验收 ID 和非法 blocker | | `doctor_project.py` | 只读聚合环境、协议、Authority、清单、迁移、索引、任务与操作残留诊断,输出版本化报告、证据和人工动作 | -| `record_code_intelligence.py` | 写入不可变的精简 Code Intelligence Record;只接受已检查的 v2 新鲜度、stale point 和源码回退证据 | -| `code_intelligence_runtime.py` | 内部阶段工具:读取一次 status、按需至多 sync 一次并复查一次,或分类 explore 响应;不暴露为用户 CLI 命令 | +| `record_code_intelligence.py` | 从保留的代理 bundle 与受 Schema 校验的 annotations 投影不可变 v3 Code Intelligence Record;拒绝手写 v1/v2/v3 输入 | +| `code_intelligence_mcp.py` | 项目级 stdio MCP,只暴露一个有界 `polaris_codegraph_explore` 工具并保证 freshness envelope 先于图内容 | +| `code_intelligence_runtime.py` | 保留的底层适配入口;阶段 Skill 不直接编排它,而由项目代理统一执行状态、同步和响应分类窗口 | | `configure_code_intelligence.py` | 启用并优先一个已配置 Provider,保留现有索引范围,不安装或运行 Provider | | `validate_project.py` | 检查目录、ID、结构化索引、活动任务、dangling refs、graph schema | | `validate_task.py` | 检查 revision、artifact JSON、commit/diff hash、finding、AC evidence、docs delta 和 closure eligibility | @@ -630,7 +630,7 @@ Work Item 的 `risk_flags` 用于机械计算最低 rigor:任意 risk flag 为 - [x] 建源仓库目录、JSON artifact 模板、必要 Markdown 上下文模板和八个 Skill - [x] 建立版本化 `hosts/*/adapter.json` 契约,从宿主无关 Skills 生成 Codex/Claude Code 目录与 worker 文件,并将适配器、脚本、Schema、模板和 Workflow vendoring 到 `tools/polaris/` -- [x] 将 Adapter 升级到 v2,校验真实入口、overlay 新增边界、symlink confinement 与宿主能力依赖 +- [x] 将 Adapter 升级到 v3,校验真实入口、overlay 新增边界、symlink confinement、宿主能力依赖与非破坏项目 MCP 注册 - [x] 用安装清单登记 vendored 文件归属、跨平台文本哈希/严格字节哈希,并以预生成、备份、回滚和崩溃恢复事务执行强制升级 - [x] 建立显式相邻迁移注册表、可恢复迁移记录与 append-only 任务版本事件 - [ ] 用最小 fixture 验证当前 Codex 宿主能够发现仓库内 Skills diff --git a/pyproject.toml b/pyproject.toml index d5c11c6..e8fb9c0 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "corona-polaris" -version = "0.1.20" +version = "0.1.21" description = "Repo-native AI engineering workflow command dispatcher" requires-python = ">=3.10" dependencies = [] diff --git a/schemas/code-intelligence-record-annotations.schema.json b/schemas/code-intelligence-record-annotations.schema.json new file mode 100644 index 0000000..6a77c92 --- /dev/null +++ b/schemas/code-intelligence-record-annotations.schema.json @@ -0,0 +1,56 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "title": "Polaris Code Intelligence v3 record annotations", + "type": "object", + "required": ["summary", "symbols", "source_fallbacks"], + "additionalProperties": false, + "properties": { + "summary": {"type": "string"}, + "symbols": { + "type": "array", + "items": { + "type": "object", + "required": ["path", "line", "name"], + "additionalProperties": false, + "properties": { + "path": {"type": "string", "minLength": 1}, + "line": {"type": ["integer", "null"], "minimum": 1}, + "name": {"type": "string", "minLength": 1} + } + } + }, + "source_fallbacks": { + "type": "array", + "items": { + "type": "object", + "required": [ + "action", "path", "observed_sha256", "base_commit", + "head_commit", "diff_hash", "purpose", "result_paths" + ], + "additionalProperties": false, + "properties": { + "action": {"type": "string", "enum": ["READ_SOURCE", "INSPECT_GIT_DIFF", "SEARCH_SOURCE"]}, + "path": {"type": ["string", "null"]}, + "observed_sha256": {"type": ["string", "null"], "pattern": "^[0-9a-f]{64}$"}, + "base_commit": {"type": ["string", "null"], "pattern": "^[0-9a-f]{40}$"}, + "head_commit": {"type": ["string", "null"], "pattern": "^[0-9a-f]{40}$"}, + "diff_hash": {"type": ["string", "null"], "pattern": "^[0-9a-f]{64}$"}, + "purpose": {"type": "string", "minLength": 1}, + "result_paths": { + "type": "array", + "maxItems": 100, + "items": { + "type": "object", + "required": ["path", "observed_sha256"], + "additionalProperties": false, + "properties": { + "path": {"type": "string", "minLength": 1}, + "observed_sha256": {"type": "string", "pattern": "^[0-9a-f]{64}$"} + } + } + } + } + } + } + } +} diff --git a/schemas/code-intelligence-record-v2.schema.json b/schemas/code-intelligence-record-v2.schema.json new file mode 100644 index 0000000..14ca0f6 --- /dev/null +++ b/schemas/code-intelligence-record-v2.schema.json @@ -0,0 +1,496 @@ +{ + "$schema": "https://json-schema.org/draft/2020-12/schema", + "title": "Polaris Code Intelligence record v2", + "type": "object", + "required": [ + "record_version", + "task_id", + "work_item_revision", + "stage", + "artifact_attempt", + "reviewer_slot", + "provider", + "target", + "status", + "queries", + "status_check", + "sync", + "freshness", + "source_fallbacks", + "recorded_at" + ], + "additionalProperties": false, + "properties": { + "record_version": { + "const": 2 + }, + "task_id": { + "type": "string", + "pattern": "^TASK-[0-9]{4}$" + }, + "work_item_revision": { + "type": "integer", + "minimum": 1 + }, + "stage": { + "type": "string", + "enum": [ + "PLANNING", + "IMPLEMENTATION", + "DOCUMENTATION_SYNC", + "REVIEW" + ] + }, + "artifact_attempt": { + "type": [ + "integer", + "null" + ], + "minimum": 1 + }, + "reviewer_slot": { + "type": [ + "integer", + "null" + ], + "minimum": 1 + }, + "provider": { + "type": [ + "object", + "null" + ], + "required": [ + "id", + "descriptor_version", + "transport", + "available_operations" + ], + "additionalProperties": false, + "properties": { + "id": { + "type": "string", + "pattern": "^[a-z][a-z0-9-]*$" + }, + "descriptor_version": { + "const": 2 + }, + "transport": { + "const": "mcp" + }, + "available_operations": { + "type": "array", + "uniqueItems": true, + "items": { + "type": "string", + "enum": [ + "explore", + "status", + "sync" + ] + } + } + } + }, + "target": { + "type": "object", + "required": [ + "base_commit", + "head_commit", + "diff_hash" + ], + "additionalProperties": false, + "properties": { + "base_commit": { + "type": "string", + "pattern": "^[0-9a-f]{40}$" + }, + "head_commit": { + "type": [ + "string", + "null" + ], + "pattern": "^[0-9a-f]{40}$" + }, + "diff_hash": { + "type": [ + "string", + "null" + ], + "pattern": "^[0-9a-f]{64}$" + } + } + }, + "status": { + "type": "string", + "enum": [ + "USED", + "UNAVAILABLE", + "FAILED", + "SKIPPED" + ] + }, + "queries": { + "type": "array", + "items": { + "type": "object", + "required": [ + "id", + "operation", + "purpose", + "status", + "summary", + "symbols", + "response_sha256", + "error" + ], + "additionalProperties": false, + "properties": { + "id": { + "type": "string", + "pattern": "^CIQ-[0-9]{3}$" + }, + "operation": { + "const": "explore" + }, + "purpose": { + "type": "string", + "minLength": 1 + }, + "status": { + "type": "string", + "enum": [ + "SUCCESS", + "EMPTY", + "FAILED", + "UNAVAILABLE" + ] + }, + "summary": { + "type": "string" + }, + "symbols": { + "type": "array", + "items": { + "type": "object", + "required": [ + "path", + "line", + "name" + ], + "additionalProperties": false, + "properties": { + "path": { + "type": "string", + "minLength": 1 + }, + "line": { + "type": [ + "integer", + "null" + ], + "minimum": 1 + }, + "name": { + "type": "string", + "minLength": 1 + } + } + } + }, + "response_sha256": { + "type": [ + "string", + "null" + ], + "pattern": "^[0-9a-f]{64}$" + }, + "error": { + "type": [ + "string", + "null" + ] + } + } + } + }, + "status_check": { + "type": [ + "object", + "null" + ], + "required": [ + "status", + "phase", + "response_sha256", + "error" + ], + "additionalProperties": false, + "properties": { + "status": { + "type": "string", + "enum": [ + "SUCCESS", + "FAILED", + "SKIPPED", + "UNAVAILABLE" + ] + }, + "phase": { + "type": "string", + "enum": [ + "STAGE_ENTRY", + "POST_SYNC" + ] + }, + "response_sha256": { + "type": [ + "string", + "null" + ], + "pattern": "^[0-9a-f]{64}$" + }, + "error": { + "type": [ + "string", + "null" + ] + } + } + }, + "sync": { + "type": [ + "object", + "null" + ], + "required": [ + "status", + "response_sha256", + "error" + ], + "additionalProperties": false, + "properties": { + "status": { + "type": "string", + "enum": [ + "SUCCESS", + "FAILED", + "SKIPPED", + "UNAVAILABLE" + ] + }, + "response_sha256": { + "type": [ + "string", + "null" + ], + "pattern": "^[0-9a-f]{64}$" + }, + "error": { + "type": [ + "string", + "null" + ] + } + } + }, + "freshness": { + "type": "object", + "required": [ + "status", + "checked_at", + "basis", + "response_sha256", + "stale_points" + ], + "additionalProperties": false, + "properties": { + "status": { + "type": "string", + "enum": [ + "CURRENT_AT_CHECK", + "PARTIAL_STALE", + "INDEX_STALE", + "NOT_VERIFIED", + "UNAVAILABLE" + ] + }, + "checked_at": { + "type": "string", + "minLength": 1 + }, + "response_sha256": { + "type": [ + "string", + "null" + ], + "pattern": "^[0-9a-f]{64}$" + }, + "basis": { + "type": "array", + "minItems": 1, + "uniqueItems": true, + "items": { + "type": "string", + "enum": [ + "CONNECT_RECONCILIATION", + "STATUS_JSON", + "SYNC_ACKNOWLEDGED", + "RESPONSE_BANNER", + "NONE" + ] + } + }, + "stale_points": { + "type": "array", + "items": { + "type": "object", + "required": [ + "scope", + "path", + "reason", + "fallback", + "observed_sha256" + ], + "additionalProperties": false, + "properties": { + "scope": { + "type": "string", + "enum": [ + "FILE", + "INDEX" + ] + }, + "path": { + "type": [ + "string", + "null" + ] + }, + "reason": { + "type": "string", + "enum": [ + "PENDING_SYNC", + "AUTO_SYNC_DISABLED", + "WORKTREE_MISMATCH", + "INDEX_PARTIAL", + "INDEX_INDEXING", + "INDEX_FAILED", + "PENDING_REFERENCES", + "REINDEX_RECOMMENDED", + "SYNC_FAILED", + "STATUS_UNREADABLE" + ] + }, + "fallback": { + "type": "string", + "enum": [ + "READ_SOURCE", + "INSPECT_GIT_DIFF", + "SEARCH_SOURCE" + ] + }, + "observed_sha256": { + "type": [ + "string", + "null" + ], + "pattern": "^[0-9a-f]{64}$" + } + } + } + } + } + }, + "source_fallbacks": { + "type": "array", + "items": { + "type": "object", + "required": [ + "action", + "path", + "observed_sha256", + "base_commit", + "head_commit", + "diff_hash", + "purpose", + "result_paths" + ], + "additionalProperties": false, + "properties": { + "action": { + "type": "string", + "enum": [ + "READ_SOURCE", + "INSPECT_GIT_DIFF", + "SEARCH_SOURCE" + ] + }, + "path": { + "type": [ + "string", + "null" + ] + }, + "observed_sha256": { + "type": [ + "string", + "null" + ], + "pattern": "^[0-9a-f]{64}$" + }, + "base_commit": { + "type": [ + "string", + "null" + ], + "pattern": "^[0-9a-f]{40}$" + }, + "head_commit": { + "type": [ + "string", + "null" + ], + "pattern": "^[0-9a-f]{40}$" + }, + "diff_hash": { + "type": [ + "string", + "null" + ], + "pattern": "^[0-9a-f]{64}$" + }, + "purpose": { + "type": "string" + }, + "result_paths": { + "type": "array", + "maxItems": 100, + "items": { + "type": "object", + "required": [ + "path", + "observed_sha256" + ], + "additionalProperties": false, + "properties": { + "path": { + "type": "string", + "minLength": 1 + }, + "observed_sha256": { + "type": "string", + "pattern": "^[0-9a-f]{64}$" + } + } + } + } + } + } + }, + "recorded_at": { + "type": "string", + "minLength": 1 + } + } +} diff --git a/schemas/code-intelligence-record.schema.json b/schemas/code-intelligence-record.schema.json index 14ca0f6..8199c11 100644 --- a/schemas/code-intelligence-record.schema.json +++ b/schemas/code-intelligence-record.schema.json @@ -1,6 +1,6 @@ { "$schema": "https://json-schema.org/draft/2020-12/schema", - "title": "Polaris Code Intelligence record v2", + "title": "Polaris Code Intelligence record v3", "type": "object", "required": [ "record_version", @@ -10,396 +10,146 @@ "artifact_attempt", "reviewer_slot", "provider", + "repository", "target", "status", - "queries", - "status_check", - "sync", - "freshness", + "proxy", + "query", + "query_window", + "delivery", "source_fallbacks", "recorded_at" ], "additionalProperties": false, "properties": { - "record_version": { - "const": 2 - }, - "task_id": { - "type": "string", - "pattern": "^TASK-[0-9]{4}$" - }, - "work_item_revision": { - "type": "integer", - "minimum": 1 - }, + "record_version": {"const": 3}, + "task_id": {"type": "string", "pattern": "^TASK-[0-9]{4}$"}, + "work_item_revision": {"type": "integer", "minimum": 1}, "stage": { "type": "string", - "enum": [ - "PLANNING", - "IMPLEMENTATION", - "DOCUMENTATION_SYNC", - "REVIEW" - ] - }, - "artifact_attempt": { - "type": [ - "integer", - "null" - ], - "minimum": 1 - }, - "reviewer_slot": { - "type": [ - "integer", - "null" - ], - "minimum": 1 + "enum": ["PLANNING", "IMPLEMENTATION", "DOCUMENTATION_SYNC", "REVIEW"] }, + "artifact_attempt": {"type": ["integer", "null"], "minimum": 1}, + "reviewer_slot": {"type": ["integer", "null"], "minimum": 1}, "provider": { - "type": [ - "object", - "null" - ], - "required": [ - "id", - "descriptor_version", - "transport", - "available_operations" - ], + "type": "object", + "required": ["id", "descriptor_version"], "additionalProperties": false, "properties": { - "id": { - "type": "string", - "pattern": "^[a-z][a-z0-9-]*$" - }, - "descriptor_version": { - "const": 2 - }, - "transport": { - "const": "mcp" - }, - "available_operations": { - "type": "array", - "uniqueItems": true, - "items": { - "type": "string", - "enum": [ - "explore", - "status", - "sync" - ] - } - } + "id": {"const": "codegraph"}, + "descriptor_version": {"const": 2} } }, - "target": { + "repository": { "type": "object", - "required": [ - "base_commit", - "head_commit", - "diff_hash" - ], + "required": ["project_id", "root_sha256"], "additionalProperties": false, "properties": { - "base_commit": { - "type": "string", - "pattern": "^[0-9a-f]{40}$" - }, - "head_commit": { - "type": [ - "string", - "null" - ], - "pattern": "^[0-9a-f]{40}$" - }, - "diff_hash": { - "type": [ - "string", - "null" - ], - "pattern": "^[0-9a-f]{64}$" - } + "project_id": {"type": "string", "minLength": 1}, + "root_sha256": {"type": "string", "pattern": "^[0-9a-f]{64}$"} } }, - "status": { - "type": "string", - "enum": [ - "USED", - "UNAVAILABLE", - "FAILED", - "SKIPPED" - ] + "target": { + "type": "object", + "required": ["base_commit", "head_commit", "diff_hash"], + "additionalProperties": false, + "properties": { + "base_commit": {"type": "string", "pattern": "^[0-9a-f]{40}$"}, + "head_commit": {"type": ["string", "null"], "pattern": "^[0-9a-f]{40}$"}, + "diff_hash": {"type": ["string", "null"], "pattern": "^[0-9a-f]{64}$"} + } }, - "queries": { - "type": "array", - "items": { - "type": "object", - "required": [ - "id", - "operation", - "purpose", - "status", - "summary", - "symbols", - "response_sha256", - "error" - ], - "additionalProperties": false, - "properties": { - "id": { - "type": "string", - "pattern": "^CIQ-[0-9]{3}$" - }, - "operation": { - "const": "explore" - }, - "purpose": { - "type": "string", - "minLength": 1 - }, - "status": { - "type": "string", - "enum": [ - "SUCCESS", - "EMPTY", - "FAILED", - "UNAVAILABLE" - ] - }, - "summary": { - "type": "string" - }, - "symbols": { - "type": "array", - "items": { - "type": "object", - "required": [ - "path", - "line", - "name" - ], - "additionalProperties": false, - "properties": { - "path": { - "type": "string", - "minLength": 1 - }, - "line": { - "type": [ - "integer", - "null" - ], - "minimum": 1 - }, - "name": { - "type": "string", - "minLength": 1 - } - } - } - }, - "response_sha256": { - "type": [ - "string", - "null" - ], - "pattern": "^[0-9a-f]{64}$" - }, - "error": { - "type": [ - "string", - "null" - ] - } - } + "status": {"type": "string", "enum": ["USED", "FAILED", "UNAVAILABLE"]}, + "proxy": { + "type": "object", + "required": ["server_id", "tool", "evidence_bundle_sha256"], + "additionalProperties": false, + "properties": { + "server_id": {"const": "polaris-codegraph"}, + "tool": {"const": "polaris_codegraph_explore"}, + "evidence_bundle_sha256": {"type": "string", "pattern": "^[0-9a-f]{64}$"} } }, - "status_check": { - "type": [ - "object", - "null" - ], + "query": { + "type": "object", "required": [ - "status", - "phase", - "response_sha256", - "error" + "id", "purpose", "text", "status", "summary", "symbols", + "response_sha256", "error" ], "additionalProperties": false, "properties": { - "status": { - "type": "string", - "enum": [ - "SUCCESS", - "FAILED", - "SKIPPED", - "UNAVAILABLE" - ] - }, - "phase": { - "type": "string", - "enum": [ - "STAGE_ENTRY", - "POST_SYNC" - ] - }, - "response_sha256": { - "type": [ - "string", - "null" - ], - "pattern": "^[0-9a-f]{64}$" + "id": {"type": "string", "pattern": "^CIQ-[0-9]{3}$"}, + "purpose": {"type": "string", "minLength": 1}, + "text": {"type": "string", "minLength": 1}, + "status": {"type": "string", "enum": ["SUCCESS", "FAILED", "UNAVAILABLE"]}, + "summary": {"type": "string"}, + "symbols": { + "type": "array", + "items": { + "type": "object", + "required": ["path", "line", "name"], + "additionalProperties": false, + "properties": { + "path": {"type": "string", "minLength": 1}, + "line": {"type": ["integer", "null"], "minimum": 1}, + "name": {"type": "string", "minLength": 1} + } + } }, - "error": { - "type": [ - "string", - "null" - ] - } + "response_sha256": {"type": ["string", "null"], "pattern": "^[0-9a-f]{64}$"}, + "error": {"type": ["string", "null"]} } }, - "sync": { - "type": [ - "object", - "null" - ], + "query_window": { + "type": "object", "required": [ - "status", - "response_sha256", - "error" + "pre_status", "sync", "post_sync_status", + "response_classification", "post_query_status" ], "additionalProperties": false, "properties": { - "status": { - "type": "string", - "enum": [ - "SUCCESS", - "FAILED", - "SKIPPED", - "UNAVAILABLE" - ] - }, - "response_sha256": { - "type": [ - "string", - "null" - ], - "pattern": "^[0-9a-f]{64}$" - }, - "error": { - "type": [ - "string", - "null" - ] - } + "pre_status": {"type": "object"}, + "sync": {"type": ["object", "null"]}, + "post_sync_status": {"type": ["object", "null"]}, + "response_classification": {"type": ["object", "null"]}, + "post_query_status": {"type": ["object", "null"]} } }, - "freshness": { + "delivery": { "type": "object", "required": [ - "status", - "checked_at", - "basis", - "response_sha256", - "stale_points" + "state", "record_status", "reason", "checked_at", "usage", + "required_fallback", "stale_points", "pending_changes", "error" ], "additionalProperties": false, "properties": { - "status": { + "state": {"type": "string", "enum": ["CURRENT", "STALE", "UNKNOWN", "UNAVAILABLE"]}, + "record_status": { "type": "string", - "enum": [ - "CURRENT_AT_CHECK", - "PARTIAL_STALE", - "INDEX_STALE", - "NOT_VERIFIED", - "UNAVAILABLE" - ] + "enum": ["CURRENT_AT_CHECK", "PARTIAL_STALE", "INDEX_STALE", "NOT_VERIFIED", "UNAVAILABLE"] }, - "checked_at": { + "reason": {"type": "string", "minLength": 1}, + "checked_at": {"type": "string", "minLength": 1}, + "usage": { "type": "string", - "minLength": 1 + "enum": ["NON_AUTHORITATIVE_CONTEXT", "NAVIGATION_ONLY", "NO_GRAPH"] }, - "response_sha256": { - "type": [ - "string", - "null" - ], - "pattern": "^[0-9a-f]{64}$" + "required_fallback": { + "type": "string", + "enum": ["NONE", "READ_SOURCE", "INSPECT_GIT_DIFF", "SEARCH_SOURCE"] }, - "basis": { - "type": "array", - "minItems": 1, - "uniqueItems": true, - "items": { - "type": "string", - "enum": [ - "CONNECT_RECONCILIATION", - "STATUS_JSON", - "SYNC_ACKNOWLEDGED", - "RESPONSE_BANNER", - "NONE" - ] + "stale_points": {"type": "array", "items": {"type": "object"}}, + "pending_changes": { + "type": "object", + "required": ["added", "modified", "removed"], + "additionalProperties": false, + "properties": { + "added": {"type": "integer", "minimum": 0}, + "modified": {"type": "integer", "minimum": 0}, + "removed": {"type": "integer", "minimum": 0} } }, - "stale_points": { - "type": "array", - "items": { - "type": "object", - "required": [ - "scope", - "path", - "reason", - "fallback", - "observed_sha256" - ], - "additionalProperties": false, - "properties": { - "scope": { - "type": "string", - "enum": [ - "FILE", - "INDEX" - ] - }, - "path": { - "type": [ - "string", - "null" - ] - }, - "reason": { - "type": "string", - "enum": [ - "PENDING_SYNC", - "AUTO_SYNC_DISABLED", - "WORKTREE_MISMATCH", - "INDEX_PARTIAL", - "INDEX_INDEXING", - "INDEX_FAILED", - "PENDING_REFERENCES", - "REINDEX_RECOMMENDED", - "SYNC_FAILED", - "STATUS_UNREADABLE" - ] - }, - "fallback": { - "type": "string", - "enum": [ - "READ_SOURCE", - "INSPECT_GIT_DIFF", - "SEARCH_SOURCE" - ] - }, - "observed_sha256": { - "type": [ - "string", - "null" - ], - "pattern": "^[0-9a-f]{64}$" - } - } - } - } + "error": {"type": ["string", "null"]} } }, "source_fallbacks": { @@ -407,90 +157,34 @@ "items": { "type": "object", "required": [ - "action", - "path", - "observed_sha256", - "base_commit", - "head_commit", - "diff_hash", - "purpose", - "result_paths" + "action", "path", "observed_sha256", "base_commit", + "head_commit", "diff_hash", "purpose", "result_paths" ], "additionalProperties": false, "properties": { - "action": { - "type": "string", - "enum": [ - "READ_SOURCE", - "INSPECT_GIT_DIFF", - "SEARCH_SOURCE" - ] - }, - "path": { - "type": [ - "string", - "null" - ] - }, - "observed_sha256": { - "type": [ - "string", - "null" - ], - "pattern": "^[0-9a-f]{64}$" - }, - "base_commit": { - "type": [ - "string", - "null" - ], - "pattern": "^[0-9a-f]{40}$" - }, - "head_commit": { - "type": [ - "string", - "null" - ], - "pattern": "^[0-9a-f]{40}$" - }, - "diff_hash": { - "type": [ - "string", - "null" - ], - "pattern": "^[0-9a-f]{64}$" - }, - "purpose": { - "type": "string" - }, + "action": {"type": "string", "enum": ["READ_SOURCE", "INSPECT_GIT_DIFF", "SEARCH_SOURCE"]}, + "path": {"type": ["string", "null"]}, + "observed_sha256": {"type": ["string", "null"], "pattern": "^[0-9a-f]{64}$"}, + "base_commit": {"type": ["string", "null"], "pattern": "^[0-9a-f]{40}$"}, + "head_commit": {"type": ["string", "null"], "pattern": "^[0-9a-f]{40}$"}, + "diff_hash": {"type": ["string", "null"], "pattern": "^[0-9a-f]{64}$"}, + "purpose": {"type": "string", "minLength": 1}, "result_paths": { "type": "array", "maxItems": 100, "items": { "type": "object", - "required": [ - "path", - "observed_sha256" - ], + "required": ["path", "observed_sha256"], "additionalProperties": false, "properties": { - "path": { - "type": "string", - "minLength": 1 - }, - "observed_sha256": { - "type": "string", - "pattern": "^[0-9a-f]{64}$" - } + "path": {"type": "string", "minLength": 1}, + "observed_sha256": {"type": "string", "pattern": "^[0-9a-f]{64}$"} } } } } } }, - "recorded_at": { - "type": "string", - "minLength": 1 - } + "recorded_at": {"type": "string", "minLength": 1} } } diff --git a/schemas/host-adapter.schema.json b/schemas/host-adapter.schema.json index 75a206f..3c88c8a 100644 --- a/schemas/host-adapter.schema.json +++ b/schemas/host-adapter.schema.json @@ -13,12 +13,13 @@ "entry_frontmatter", "skill_overlay_root", "skill_appendix_root", + "project_mcp", "files" ], "additionalProperties": false, "properties": { "adapter_version": { - "const": 2 + "const": 3 }, "host_id": { "type": "string", @@ -86,6 +87,42 @@ "null" ] }, + "project_mcp": { + "type": "object", + "required": [ + "server_id", + "format", + "target", + "command", + "args" + ], + "additionalProperties": false, + "properties": { + "server_id": { + "const": "polaris-codegraph" + }, + "format": { + "enum": [ + "codex-toml", + "claude-json" + ] + }, + "target": { + "type": "string", + "minLength": 1 + }, + "command": { + "const": "python3" + }, + "args": { + "const": [ + "tools/polaris/scripts/code_intelligence_mcp.py", + "--repo", + "." + ] + } + } + }, "files": { "type": "array", "items": { diff --git a/scripts/code_intelligence_mcp.py b/scripts/code_intelligence_mcp.py new file mode 100644 index 0000000..919afb7 --- /dev/null +++ b/scripts/code_intelligence_mcp.py @@ -0,0 +1,275 @@ +#!/usr/bin/env python3 +"""Project-scoped stdio MCP server for bounded Polaris CodeGraph queries.""" + +from __future__ import annotations + +import argparse +import json +import re +import sys +from pathlib import Path +from typing import Any + +from internal.code_intelligence_proxy import execute_proxy_query +from internal.path_security import require_regular_file +from internal.polaris_core import ( + InputFailure, + RuleFailure, + protocol_root, +) + + +PROTOCOL_VERSION = "2025-11-25" +TOOL_NAME = "polaris_codegraph_explore" +SERVER_ROOT = Path(__file__).resolve().parent.parent +TOOL = { + "name": TOOL_NAME, + "description": "Run one bounded Polaris CodeGraph freshness window.", + "inputSchema": { + "type": "object", + "required": [ + "task_id", + "stage", + "query_id", + "purpose", + "query", + "sync_if_needed", + ], + "additionalProperties": False, + "properties": { + "task_id": {"type": "string", "pattern": r"^TASK-[0-9]{4}$"}, + "stage": { + "type": "string", + "enum": [ + "PLANNING", + "IMPLEMENTATION", + "DOCUMENTATION_SYNC", + "REVIEW", + ], + }, + "query_id": {"type": "string", "pattern": r"^CIQ-[0-9]{3}$"}, + "purpose": {"type": "string", "minLength": 1, "maxLength": 240}, + "query": {"type": "string", "minLength": 1, "maxLength": 8000}, + "sync_if_needed": {"type": "boolean"}, + }, + }, +} + + +def _error(request_id: Any, code: int, message: str) -> dict[str, Any]: + return { + "jsonrpc": "2.0", + "id": request_id, + "error": {"code": code, "message": " ".join(message.split())[:240]}, + } + + +def _result(request_id: Any, value: dict[str, Any]) -> dict[str, Any]: + return {"jsonrpc": "2.0", "id": request_id, "result": value} + + +def _tool_error(message: str) -> dict[str, Any]: + return { + "content": [ + {"type": "text", "text": " ".join(message.split())[:240]} + ], + "isError": True, + } + + +def _valid_request_id(value: Any) -> bool: + return value is None or ( + not isinstance(value, bool) and isinstance(value, (int, str)) + ) + + +def _validate_arguments(value: Any) -> list[str]: + if not isinstance(value, dict): + return ["tool arguments must be an object"] + expected = { + "task_id", + "stage", + "query_id", + "purpose", + "query", + "sync_if_needed", + } + errors: list[str] = [] + missing = expected - set(value) + extra = set(value) - expected + if missing: + errors.append("missing arguments: " + ", ".join(sorted(missing))) + if extra: + errors.append("unknown arguments: " + ", ".join(sorted(extra))) + task_id = value.get("task_id") + if not isinstance(task_id, str) or re.fullmatch(r"TASK-[0-9]{4}", task_id) is None: + errors.append("task_id must match TASK-0000") + if value.get("stage") not in { + "PLANNING", + "IMPLEMENTATION", + "DOCUMENTATION_SYNC", + "REVIEW", + }: + errors.append("stage is invalid") + query_id = value.get("query_id") + if not isinstance(query_id, str) or re.fullmatch(r"CIQ-[0-9]{3}", query_id) is None: + errors.append("query_id must match CIQ-000") + for key, maximum in (("purpose", 240), ("query", 8000)): + item = value.get(key) + if not isinstance(item, str) or not item.strip() or len(item) > maximum: + errors.append(f"{key} must contain 1 to {maximum} characters") + if not isinstance(value.get("sync_if_needed"), bool): + errors.append("sync_if_needed must be a boolean") + return errors + + +class McpServer: + """Small stateful MCP dispatcher with no dependencies beyond the stdlib.""" + + def __init__(self, repo: Path) -> None: + raw_repo = repo.absolute() + if raw_repo.is_symlink() or not raw_repo.is_dir(): + raise InputFailure("MCP repository root must be a fixed real directory") + self.repo = raw_repo.resolve() + if protocol_root(self.repo).resolve() != SERVER_ROOT: + raise RuleFailure( + "MCP repository does not match the executing vendored protocol root" + ) + require_regular_file( + self.repo / ".polaris/project.json", "Polaris project configuration" + ) + self.initialized = False + self.ready = False + + def _initialize(self, request_id: Any, params: dict[str, Any]) -> dict[str, Any]: + if self.initialized: + return _error(request_id, -32600, "MCP server is already initialized") + if ( + params.get("protocolVersion") != PROTOCOL_VERSION + or not isinstance(params.get("capabilities"), dict) + or not isinstance(params.get("clientInfo"), dict) + ): + return _error(request_id, -32602, "unsupported or incomplete initialize params") + self.initialized = True + version = (protocol_root(self.repo) / "VERSION").read_text( + encoding="utf-8" + ).strip() + return _result( + request_id, + { + "protocolVersion": PROTOCOL_VERSION, + "capabilities": {"tools": {"listChanged": False}}, + "serverInfo": {"name": "polaris-codegraph", "version": version}, + }, + ) + + def _call_tool(self, request_id: Any, params: dict[str, Any]) -> dict[str, Any]: + if params.get("name") != TOOL_NAME: + return _error(request_id, -32602, "unknown MCP tool name") + arguments = params.get("arguments") + errors = _validate_arguments(arguments) + if errors: + return _result(request_id, _tool_error("; ".join(errors))) + assert isinstance(arguments, dict) + try: + proxy = execute_proxy_query( + self.repo, + arguments["task_id"], + arguments["stage"], + arguments["query_id"], + arguments["purpose"], + arguments["query"], + arguments["sync_if_needed"], + ) + content = [{"type": "text", "text": proxy["envelope"]}] + if proxy["response"] is not None: + content.append({"type": "text", "text": proxy["response"]}) + return _result( + request_id, + { + "content": content, + "structuredContent": {"bundle": proxy["bundle"]}, + "isError": False, + }, + ) + except (InputFailure, RuleFailure, OSError, ValueError) as error: + return _result(request_id, _tool_error(str(error))) + except Exception as error: # keep the long-lived stdio server usable + return _result( + request_id, + _tool_error(f"CodeGraph proxy execution failed: {type(error).__name__}"), + ) + + def handle(self, message: Any) -> dict[str, Any] | None: + """Handle one decoded JSON-RPC message; notifications return ``None``.""" + if not isinstance(message, dict) or message.get("jsonrpc") != "2.0": + return _error(None, -32600, "invalid JSON-RPC request") + method = message.get("method") + if not isinstance(method, str): + return _error(message.get("id"), -32600, "request method must be a string") + has_id = "id" in message + request_id = message.get("id") + if has_id and not _valid_request_id(request_id): + return _error(None, -32600, "invalid JSON-RPC request id") + params = message.get("params", {}) + if not isinstance(params, dict): + return None if not has_id else _error(request_id, -32602, "params must be an object") + + if method == "initialize": + if not has_id: + return None + return self._initialize(request_id, params) + if method == "notifications/initialized": + if not has_id and self.initialized: + self.ready = True + return None + known_methods = {"ping", "tools/list", "tools/call"} + if method not in known_methods: + return None if not has_id else _error(request_id, -32601, "method not found") + if not self.ready: + return None if not has_id else _error(request_id, -32600, "MCP server is not initialized") + if not has_id: + return None + if method == "ping": + return _result(request_id, {}) + if method == "tools/list": + return _result(request_id, {"tools": [TOOL]}) + return self._call_tool(request_id, params) + + +def _write_message(value: dict[str, Any]) -> None: + sys.stdout.write( + json.dumps(value, ensure_ascii=False, separators=(",", ":")) + "\n" + ) + sys.stdout.flush() + + +def serve(repo: Path) -> int: + if repo.absolute().resolve() != Path.cwd().resolve(): + raise InputFailure("MCP repository must match the process working directory") + server = McpServer(repo) + for line in sys.stdin: + try: + message = json.loads(line) + except (json.JSONDecodeError, UnicodeError): + _write_message(_error(None, -32700, "parse error")) + continue + response = server.handle(message) + if response is not None: + _write_message(response) + return 0 + + +def main() -> int: + parser = argparse.ArgumentParser() + parser.add_argument("--repo", type=Path, required=True) + args = parser.parse_args() + try: + return serve(args.repo) + except (InputFailure, RuleFailure, OSError) as error: + print(" ".join(str(error).split())[:240], file=sys.stderr) + return 2 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/scripts/init_project.py b/scripts/init_project.py index ccd5c93..b735523 100644 --- a/scripts/init_project.py +++ b/scripts/init_project.py @@ -9,6 +9,7 @@ from pathlib import Path from internal.host_adapters import adapter_file_target, load_host_adapters +from internal.project_mcp_registration import merge_project_mcp, project_mcp_target from internal.polaris_core import ( InputFailure, ensure_gitignore_rule, @@ -16,6 +17,7 @@ read_json, run_main, write_json_atomic, + write_text_atomic, ) from internal.task_layout import ( ARCHIVED_RUNTIME_IGNORE_PATTERN, @@ -55,6 +57,9 @@ def initialize(repo: Path, project_id: str | None = None) -> dict[str, str]: if not destination.exists(): destination.parent.mkdir(parents=True, exist_ok=True) shutil.copyfile(adapter["adapter_root"] / item["source"], destination) + if root == repo / "tools" / "polaris": + destination = project_mcp_target(repo, adapter) + write_text_atomic(destination, merge_project_mcp(repo, adapter)) ensure_gitignore_rule(repo, RUNTIME_IGNORE_PATTERN) ensure_gitignore_rule(repo, ARCHIVED_RUNTIME_IGNORE_PATTERN) return {"message": f"initialized Polaris project {project_id}"} diff --git a/scripts/internal/_tomllib_compat/__init__.py b/scripts/internal/_tomllib_compat/__init__.py new file mode 100644 index 0000000..f13f93d --- /dev/null +++ b/scripts/internal/_tomllib_compat/__init__.py @@ -0,0 +1,11 @@ +# SPDX-License-Identifier: MIT +# SPDX-FileCopyrightText: 2021 Taneli Hukkinen +# Licensed to PSF under a Contributor Agreement. + +__all__ = ("loads", "load", "TOMLDecodeError") + +from ._parser import TOMLDecodeError, load, loads + +# Pretend this exception was created here. +TOMLDecodeError.__module__ = __name__ + diff --git a/scripts/internal/_tomllib_compat/_parser.py b/scripts/internal/_tomllib_compat/_parser.py new file mode 100644 index 0000000..2719bb5 --- /dev/null +++ b/scripts/internal/_tomllib_compat/_parser.py @@ -0,0 +1,692 @@ +# SPDX-License-Identifier: MIT +# SPDX-FileCopyrightText: 2021 Taneli Hukkinen +# Licensed to PSF under a Contributor Agreement. + +from __future__ import annotations + +from collections.abc import Iterable +import string +from types import MappingProxyType +from typing import Any, BinaryIO, NamedTuple + +from ._re import ( + RE_DATETIME, + RE_LOCALTIME, + RE_NUMBER, + match_to_datetime, + match_to_localtime, + match_to_number, +) +from ._types import Key, ParseFloat, Pos + +ASCII_CTRL = frozenset(chr(i) for i in range(32)) | frozenset(chr(127)) + +# Neither of these sets include quotation mark or backslash. They are +# currently handled as separate cases in the parser functions. +ILLEGAL_BASIC_STR_CHARS = ASCII_CTRL - frozenset("\t") +ILLEGAL_MULTILINE_BASIC_STR_CHARS = ASCII_CTRL - frozenset("\t\n") + +ILLEGAL_LITERAL_STR_CHARS = ILLEGAL_BASIC_STR_CHARS +ILLEGAL_MULTILINE_LITERAL_STR_CHARS = ILLEGAL_MULTILINE_BASIC_STR_CHARS + +ILLEGAL_COMMENT_CHARS = ILLEGAL_BASIC_STR_CHARS + +TOML_WS = frozenset(" \t") +TOML_WS_AND_NEWLINE = TOML_WS | frozenset("\n") +BARE_KEY_CHARS = frozenset(string.ascii_letters + string.digits + "-_") +KEY_INITIAL_CHARS = BARE_KEY_CHARS | frozenset("\"'") +HEXDIGIT_CHARS = frozenset(string.hexdigits) + +BASIC_STR_ESCAPE_REPLACEMENTS = MappingProxyType( + { + "\\b": "\u0008", # backspace + "\\t": "\u0009", # tab + "\\n": "\u000A", # linefeed + "\\f": "\u000C", # form feed + "\\r": "\u000D", # carriage return + '\\"': "\u0022", # quote + "\\\\": "\u005C", # backslash + } +) + + +class TOMLDecodeError(ValueError): + """An error raised if a document is not valid TOML.""" + + +def load(fp: BinaryIO, /, *, parse_float: ParseFloat = float) -> dict[str, Any]: + """Parse TOML from a binary file object.""" + b = fp.read() + try: + s = b.decode() + except AttributeError: + raise TypeError( + "File must be opened in binary mode, e.g. use `open('foo.toml', 'rb')`" + ) from None + return loads(s, parse_float=parse_float) + + +def loads(s: str, /, *, parse_float: ParseFloat = float) -> dict[str, Any]: # noqa: C901 + """Parse TOML from a string.""" + + # The spec allows converting "\r\n" to "\n", even in string + # literals. Let's do so to simplify parsing. + src = s.replace("\r\n", "\n") + pos = 0 + out = Output(NestedDict(), Flags()) + header: Key = () + parse_float = make_safe_parse_float(parse_float) + + # Parse one statement at a time + # (typically means one line in TOML source) + while True: + # 1. Skip line leading whitespace + pos = skip_chars(src, pos, TOML_WS) + + # 2. Parse rules. Expect one of the following: + # - end of file + # - end of line + # - comment + # - key/value pair + # - append dict to list (and move to its namespace) + # - create dict (and move to its namespace) + # Skip trailing whitespace when applicable. + try: + char = src[pos] + except IndexError: + break + if char == "\n": + pos += 1 + continue + if char in KEY_INITIAL_CHARS: + pos = key_value_rule(src, pos, out, header, parse_float) + pos = skip_chars(src, pos, TOML_WS) + elif char == "[": + try: + second_char: str | None = src[pos + 1] + except IndexError: + second_char = None + out.flags.finalize_pending() + if second_char == "[": + pos, header = create_list_rule(src, pos, out) + else: + pos, header = create_dict_rule(src, pos, out) + pos = skip_chars(src, pos, TOML_WS) + elif char != "#": + raise suffixed_err(src, pos, "Invalid statement") + + # 3. Skip comment + pos = skip_comment(src, pos) + + # 4. Expect end of line or end of file + try: + char = src[pos] + except IndexError: + break + if char != "\n": + raise suffixed_err( + src, pos, "Expected newline or end of document after a statement" + ) + pos += 1 + + return out.data.dict + + +class Flags: + """Flags that map to parsed keys/namespaces.""" + + # Marks an immutable namespace (inline array or inline table). + FROZEN = 0 + # Marks a nest that has been explicitly created and can no longer + # be opened using the "[table]" syntax. + EXPLICIT_NEST = 1 + + def __init__(self) -> None: + self._flags: dict[str, dict[Any, Any]] = {} + self._pending_flags: set[tuple[Key, int]] = set() + + def add_pending(self, key: Key, flag: int) -> None: + self._pending_flags.add((key, flag)) + + def finalize_pending(self) -> None: + for key, flag in self._pending_flags: + self.set(key, flag, recursive=False) + self._pending_flags.clear() + + def unset_all(self, key: Key) -> None: + cont = self._flags + for k in key[:-1]: + if k not in cont: + return + cont = cont[k]["nested"] + cont.pop(key[-1], None) + + def set(self, key: Key, flag: int, *, recursive: bool) -> None: # noqa: A003 + cont = self._flags + key_parent, key_stem = key[:-1], key[-1] + for k in key_parent: + if k not in cont: + cont[k] = {"flags": set(), "recursive_flags": set(), "nested": {}} + cont = cont[k]["nested"] + if key_stem not in cont: + cont[key_stem] = {"flags": set(), "recursive_flags": set(), "nested": {}} + cont[key_stem]["recursive_flags" if recursive else "flags"].add(flag) + + def is_(self, key: Key, flag: int) -> bool: + if not key: + return False # document root has no flags + cont = self._flags + for k in key[:-1]: + if k not in cont: + return False + inner_cont = cont[k] + if flag in inner_cont["recursive_flags"]: + return True + cont = inner_cont["nested"] + key_stem = key[-1] + if key_stem in cont: + cont = cont[key_stem] + return flag in cont["flags"] or flag in cont["recursive_flags"] + return False + + +class NestedDict: + def __init__(self) -> None: + # The parsed content of the TOML document + self.dict: dict[str, Any] = {} + + def get_or_create_nest( + self, + key: Key, + *, + access_lists: bool = True, + ) -> dict[str, Any]: + cont: Any = self.dict + for k in key: + if k not in cont: + cont[k] = {} + cont = cont[k] + if access_lists and isinstance(cont, list): + cont = cont[-1] + if not isinstance(cont, dict): + raise KeyError("There is no nest behind this key") + return cont # type: ignore[no-any-return] + + def append_nest_to_list(self, key: Key) -> None: + cont = self.get_or_create_nest(key[:-1]) + last_key = key[-1] + if last_key in cont: + list_ = cont[last_key] + if not isinstance(list_, list): + raise KeyError("An object other than list found behind this key") + list_.append({}) + else: + cont[last_key] = [{}] + + +class Output(NamedTuple): + data: NestedDict + flags: Flags + + +def skip_chars(src: str, pos: Pos, chars: Iterable[str]) -> Pos: + try: + while src[pos] in chars: + pos += 1 + except IndexError: + pass + return pos + + +def skip_until( + src: str, + pos: Pos, + expect: str, + *, + error_on: frozenset[str], + error_on_eof: bool, +) -> Pos: + try: + new_pos = src.index(expect, pos) + except ValueError: + new_pos = len(src) + if error_on_eof: + raise suffixed_err(src, new_pos, f"Expected {expect!r}") from None + + if not error_on.isdisjoint(src[pos:new_pos]): + while src[pos] not in error_on: + pos += 1 + raise suffixed_err(src, pos, f"Found invalid character {src[pos]!r}") + return new_pos + + +def skip_comment(src: str, pos: Pos) -> Pos: + try: + char: str | None = src[pos] + except IndexError: + char = None + if char == "#": + return skip_until( + src, pos + 1, "\n", error_on=ILLEGAL_COMMENT_CHARS, error_on_eof=False + ) + return pos + + +def skip_comments_and_array_ws(src: str, pos: Pos) -> Pos: + while True: + pos_before_skip = pos + pos = skip_chars(src, pos, TOML_WS_AND_NEWLINE) + pos = skip_comment(src, pos) + if pos == pos_before_skip: + return pos + + +def create_dict_rule(src: str, pos: Pos, out: Output) -> tuple[Pos, Key]: + pos += 1 # Skip "[" + pos = skip_chars(src, pos, TOML_WS) + pos, key = parse_key(src, pos) + + if out.flags.is_(key, Flags.EXPLICIT_NEST) or out.flags.is_(key, Flags.FROZEN): + raise suffixed_err(src, pos, f"Cannot declare {key} twice") + out.flags.set(key, Flags.EXPLICIT_NEST, recursive=False) + try: + out.data.get_or_create_nest(key) + except KeyError: + raise suffixed_err(src, pos, "Cannot overwrite a value") from None + + if not src.startswith("]", pos): + raise suffixed_err(src, pos, "Expected ']' at the end of a table declaration") + return pos + 1, key + + +def create_list_rule(src: str, pos: Pos, out: Output) -> tuple[Pos, Key]: + pos += 2 # Skip "[[" + pos = skip_chars(src, pos, TOML_WS) + pos, key = parse_key(src, pos) + + if out.flags.is_(key, Flags.FROZEN): + raise suffixed_err(src, pos, f"Cannot mutate immutable namespace {key}") + # Free the namespace now that it points to another empty list item... + out.flags.unset_all(key) + # ...but this key precisely is still prohibited from table declaration + out.flags.set(key, Flags.EXPLICIT_NEST, recursive=False) + try: + out.data.append_nest_to_list(key) + except KeyError: + raise suffixed_err(src, pos, "Cannot overwrite a value") from None + + if not src.startswith("]]", pos): + raise suffixed_err(src, pos, "Expected ']]' at the end of an array declaration") + return pos + 2, key + + +def key_value_rule( + src: str, pos: Pos, out: Output, header: Key, parse_float: ParseFloat +) -> Pos: + pos, key, value = parse_key_value_pair(src, pos, parse_float) + key_parent, key_stem = key[:-1], key[-1] + abs_key_parent = header + key_parent + + relative_path_cont_keys = (header + key[:i] for i in range(1, len(key))) + for cont_key in relative_path_cont_keys: + # Check that dotted key syntax does not redefine an existing table + if out.flags.is_(cont_key, Flags.EXPLICIT_NEST): + raise suffixed_err(src, pos, f"Cannot redefine namespace {cont_key}") + # Containers in the relative path can't be opened with the table syntax or + # dotted key/value syntax in following table sections. + out.flags.add_pending(cont_key, Flags.EXPLICIT_NEST) + + if out.flags.is_(abs_key_parent, Flags.FROZEN): + raise suffixed_err( + src, pos, f"Cannot mutate immutable namespace {abs_key_parent}" + ) + + try: + nest = out.data.get_or_create_nest(abs_key_parent) + except KeyError: + raise suffixed_err(src, pos, "Cannot overwrite a value") from None + if key_stem in nest: + raise suffixed_err(src, pos, "Cannot overwrite a value") + # Mark inline table and array namespaces recursively immutable + if isinstance(value, (dict, list)): + out.flags.set(header + key, Flags.FROZEN, recursive=True) + nest[key_stem] = value + return pos + + +def parse_key_value_pair( + src: str, pos: Pos, parse_float: ParseFloat +) -> tuple[Pos, Key, Any]: + pos, key = parse_key(src, pos) + try: + char: str | None = src[pos] + except IndexError: + char = None + if char != "=": + raise suffixed_err(src, pos, "Expected '=' after a key in a key/value pair") + pos += 1 + pos = skip_chars(src, pos, TOML_WS) + pos, value = parse_value(src, pos, parse_float) + return pos, key, value + + +def parse_key(src: str, pos: Pos) -> tuple[Pos, Key]: + pos, key_part = parse_key_part(src, pos) + key: Key = (key_part,) + pos = skip_chars(src, pos, TOML_WS) + while True: + try: + char: str | None = src[pos] + except IndexError: + char = None + if char != ".": + return pos, key + pos += 1 + pos = skip_chars(src, pos, TOML_WS) + pos, key_part = parse_key_part(src, pos) + key += (key_part,) + pos = skip_chars(src, pos, TOML_WS) + + +def parse_key_part(src: str, pos: Pos) -> tuple[Pos, str]: + try: + char: str | None = src[pos] + except IndexError: + char = None + if char in BARE_KEY_CHARS: + start_pos = pos + pos = skip_chars(src, pos, BARE_KEY_CHARS) + return pos, src[start_pos:pos] + if char == "'": + return parse_literal_str(src, pos) + if char == '"': + return parse_one_line_basic_str(src, pos) + raise suffixed_err(src, pos, "Invalid initial character for a key part") + + +def parse_one_line_basic_str(src: str, pos: Pos) -> tuple[Pos, str]: + pos += 1 + return parse_basic_str(src, pos, multiline=False) + + +def parse_array(src: str, pos: Pos, parse_float: ParseFloat) -> tuple[Pos, list[Any]]: + pos += 1 + array: list[Any] = [] + + pos = skip_comments_and_array_ws(src, pos) + if src.startswith("]", pos): + return pos + 1, array + while True: + pos, val = parse_value(src, pos, parse_float) + array.append(val) + pos = skip_comments_and_array_ws(src, pos) + + c = src[pos : pos + 1] + if c == "]": + return pos + 1, array + if c != ",": + raise suffixed_err(src, pos, "Unclosed array") + pos += 1 + + pos = skip_comments_and_array_ws(src, pos) + if src.startswith("]", pos): + return pos + 1, array + + +def parse_inline_table(src: str, pos: Pos, parse_float: ParseFloat) -> tuple[Pos, dict[str, Any]]: + pos += 1 + nested_dict = NestedDict() + flags = Flags() + + pos = skip_chars(src, pos, TOML_WS) + if src.startswith("}", pos): + return pos + 1, nested_dict.dict + while True: + pos, key, value = parse_key_value_pair(src, pos, parse_float) + key_parent, key_stem = key[:-1], key[-1] + if flags.is_(key, Flags.FROZEN): + raise suffixed_err(src, pos, f"Cannot mutate immutable namespace {key}") + try: + nest = nested_dict.get_or_create_nest(key_parent, access_lists=False) + except KeyError: + raise suffixed_err(src, pos, "Cannot overwrite a value") from None + if key_stem in nest: + raise suffixed_err(src, pos, f"Duplicate inline table key {key_stem!r}") + nest[key_stem] = value + pos = skip_chars(src, pos, TOML_WS) + c = src[pos : pos + 1] + if c == "}": + return pos + 1, nested_dict.dict + if c != ",": + raise suffixed_err(src, pos, "Unclosed inline table") + if isinstance(value, (dict, list)): + flags.set(key, Flags.FROZEN, recursive=True) + pos += 1 + pos = skip_chars(src, pos, TOML_WS) + + +def parse_basic_str_escape( + src: str, pos: Pos, *, multiline: bool = False +) -> tuple[Pos, str]: + escape_id = src[pos : pos + 2] + pos += 2 + if multiline and escape_id in {"\\ ", "\\\t", "\\\n"}: + # Skip whitespace until next non-whitespace character or end of + # the doc. Error if non-whitespace is found before newline. + if escape_id != "\\\n": + pos = skip_chars(src, pos, TOML_WS) + try: + char = src[pos] + except IndexError: + return pos, "" + if char != "\n": + raise suffixed_err(src, pos, "Unescaped '\\' in a string") + pos += 1 + pos = skip_chars(src, pos, TOML_WS_AND_NEWLINE) + return pos, "" + if escape_id == "\\u": + return parse_hex_char(src, pos, 4) + if escape_id == "\\U": + return parse_hex_char(src, pos, 8) + try: + return pos, BASIC_STR_ESCAPE_REPLACEMENTS[escape_id] + except KeyError: + raise suffixed_err(src, pos, "Unescaped '\\' in a string") from None + + +def parse_basic_str_escape_multiline(src: str, pos: Pos) -> tuple[Pos, str]: + return parse_basic_str_escape(src, pos, multiline=True) + + +def parse_hex_char(src: str, pos: Pos, hex_len: int) -> tuple[Pos, str]: + hex_str = src[pos : pos + hex_len] + if len(hex_str) != hex_len or not HEXDIGIT_CHARS.issuperset(hex_str): + raise suffixed_err(src, pos, "Invalid hex value") + pos += hex_len + hex_int = int(hex_str, 16) + if not is_unicode_scalar_value(hex_int): + raise suffixed_err(src, pos, "Escaped character is not a Unicode scalar value") + return pos, chr(hex_int) + + +def parse_literal_str(src: str, pos: Pos) -> tuple[Pos, str]: + pos += 1 # Skip starting apostrophe + start_pos = pos + pos = skip_until( + src, pos, "'", error_on=ILLEGAL_LITERAL_STR_CHARS, error_on_eof=True + ) + return pos + 1, src[start_pos:pos] # Skip ending apostrophe + + +def parse_multiline_str(src: str, pos: Pos, *, literal: bool) -> tuple[Pos, str]: + pos += 3 + if src.startswith("\n", pos): + pos += 1 + + if literal: + delim = "'" + end_pos = skip_until( + src, + pos, + "'''", + error_on=ILLEGAL_MULTILINE_LITERAL_STR_CHARS, + error_on_eof=True, + ) + result = src[pos:end_pos] + pos = end_pos + 3 + else: + delim = '"' + pos, result = parse_basic_str(src, pos, multiline=True) + + # Add at maximum two extra apostrophes/quotes if the end sequence + # is 4 or 5 chars long instead of just 3. + if not src.startswith(delim, pos): + return pos, result + pos += 1 + if not src.startswith(delim, pos): + return pos, result + delim + pos += 1 + return pos, result + (delim * 2) + + +def parse_basic_str(src: str, pos: Pos, *, multiline: bool) -> tuple[Pos, str]: + if multiline: + error_on = ILLEGAL_MULTILINE_BASIC_STR_CHARS + parse_escapes = parse_basic_str_escape_multiline + else: + error_on = ILLEGAL_BASIC_STR_CHARS + parse_escapes = parse_basic_str_escape + result = "" + start_pos = pos + while True: + try: + char = src[pos] + except IndexError: + raise suffixed_err(src, pos, "Unterminated string") from None + if char == '"': + if not multiline: + return pos + 1, result + src[start_pos:pos] + if src.startswith('"""', pos): + return pos + 3, result + src[start_pos:pos] + pos += 1 + continue + if char == "\\": + result += src[start_pos:pos] + pos, parsed_escape = parse_escapes(src, pos) + result += parsed_escape + start_pos = pos + continue + if char in error_on: + raise suffixed_err(src, pos, f"Illegal character {char!r}") + pos += 1 + + +def parse_value( # noqa: C901 + src: str, pos: Pos, parse_float: ParseFloat +) -> tuple[Pos, Any]: + try: + char: str | None = src[pos] + except IndexError: + char = None + + # IMPORTANT: order conditions based on speed of checking and likelihood + + # Basic strings + if char == '"': + if src.startswith('"""', pos): + return parse_multiline_str(src, pos, literal=False) + return parse_one_line_basic_str(src, pos) + + # Literal strings + if char == "'": + if src.startswith("'''", pos): + return parse_multiline_str(src, pos, literal=True) + return parse_literal_str(src, pos) + + # Booleans + if char == "t": + if src.startswith("true", pos): + return pos + 4, True + if char == "f": + if src.startswith("false", pos): + return pos + 5, False + + # Arrays + if char == "[": + return parse_array(src, pos, parse_float) + + # Inline tables + if char == "{": + return parse_inline_table(src, pos, parse_float) + + # Dates and times + datetime_match = RE_DATETIME.match(src, pos) + if datetime_match: + try: + datetime_obj = match_to_datetime(datetime_match) + except ValueError as e: + raise suffixed_err(src, pos, "Invalid date or datetime") from e + return datetime_match.end(), datetime_obj + localtime_match = RE_LOCALTIME.match(src, pos) + if localtime_match: + return localtime_match.end(), match_to_localtime(localtime_match) + + # Integers and "normal" floats. + # The regex will greedily match any type starting with a decimal + # char, so needs to be located after handling of dates and times. + number_match = RE_NUMBER.match(src, pos) + if number_match: + return number_match.end(), match_to_number(number_match, parse_float) + + # Special floats + first_three = src[pos : pos + 3] + if first_three in {"inf", "nan"}: + return pos + 3, parse_float(first_three) + first_four = src[pos : pos + 4] + if first_four in {"-inf", "+inf", "-nan", "+nan"}: + return pos + 4, parse_float(first_four) + + raise suffixed_err(src, pos, "Invalid value") + + +def suffixed_err(src: str, pos: Pos, msg: str) -> TOMLDecodeError: + """Return a `TOMLDecodeError` where error message is suffixed with + coordinates in source.""" + + def coord_repr(src: str, pos: Pos) -> str: + if pos >= len(src): + return "end of document" + line = src.count("\n", 0, pos) + 1 + if line == 1: + column = pos + 1 + else: + column = pos - src.rindex("\n", 0, pos) + return f"line {line}, column {column}" + + return TOMLDecodeError(f"{msg} (at {coord_repr(src, pos)})") + + +def is_unicode_scalar_value(codepoint: int) -> bool: + return (0 <= codepoint <= 55295) or (57344 <= codepoint <= 1114111) + + +def make_safe_parse_float(parse_float: ParseFloat) -> ParseFloat: + """A decorator to make `parse_float` safe. + + `parse_float` must not return dicts or lists, because these types + would be mixed with parsed TOML tables and arrays, thus confusing + the parser. The returned decorated callable raises `ValueError` + instead of returning illegal types. + """ + # The default `float` callable never returns illegal types. Optimize it. + if parse_float is float: + return float + + def safe_parse_float(float_str: str) -> Any: + float_value = parse_float(float_str) + if isinstance(float_value, (dict, list)): + raise ValueError("parse_float must not return dicts or lists") + return float_value + + return safe_parse_float + diff --git a/scripts/internal/_tomllib_compat/_re.py b/scripts/internal/_tomllib_compat/_re.py new file mode 100644 index 0000000..330de92 --- /dev/null +++ b/scripts/internal/_tomllib_compat/_re.py @@ -0,0 +1,108 @@ +# SPDX-License-Identifier: MIT +# SPDX-FileCopyrightText: 2021 Taneli Hukkinen +# Licensed to PSF under a Contributor Agreement. + +from __future__ import annotations + +from datetime import date, datetime, time, timedelta, timezone, tzinfo +from functools import lru_cache +import re +from typing import Any + +from ._types import ParseFloat + +# E.g. +# - 00:32:00.999999 +# - 00:32:00 +_TIME_RE_STR = r"([01][0-9]|2[0-3]):([0-5][0-9]):([0-5][0-9])(?:\.([0-9]{1,6})[0-9]*)?" + +RE_NUMBER = re.compile( + r""" +0 +(?: + x[0-9A-Fa-f](?:_?[0-9A-Fa-f])* # hex + | + b[01](?:_?[01])* # bin + | + o[0-7](?:_?[0-7])* # oct +) +| +[+-]?(?:0|[1-9](?:_?[0-9])*) # dec, integer part +(?P + (?:\.[0-9](?:_?[0-9])*)? # optional fractional part + (?:[eE][+-]?[0-9](?:_?[0-9])*)? # optional exponent part +) +""", + flags=re.VERBOSE, +) +RE_LOCALTIME = re.compile(_TIME_RE_STR) +RE_DATETIME = re.compile( + rf""" +([0-9]{{4}})-(0[1-9]|1[0-2])-(0[1-9]|[12][0-9]|3[01]) # date, e.g. 1988-10-27 +(?: + [Tt ] + {_TIME_RE_STR} + (?:([Zz])|([+-])([01][0-9]|2[0-3]):([0-5][0-9]))? # optional time offset +)? +""", + flags=re.VERBOSE, +) + + +def match_to_datetime(match: re.Match[str]) -> datetime | date: + """Convert a `RE_DATETIME` match to `datetime.datetime` or `datetime.date`. + + Raises ValueError if the match does not correspond to a valid date + or datetime. + """ + ( + year_str, + month_str, + day_str, + hour_str, + minute_str, + sec_str, + micros_str, + zulu_time, + offset_sign_str, + offset_hour_str, + offset_minute_str, + ) = match.groups() + year, month, day = int(year_str), int(month_str), int(day_str) + if hour_str is None: + return date(year, month, day) + hour, minute, sec = int(hour_str), int(minute_str), int(sec_str) + micros = int(micros_str.ljust(6, "0")) if micros_str else 0 + if offset_sign_str: + tz: tzinfo | None = cached_tz( + offset_hour_str, offset_minute_str, offset_sign_str + ) + elif zulu_time: + tz = timezone.utc + else: # local date-time + tz = None + return datetime(year, month, day, hour, minute, sec, micros, tzinfo=tz) + + +@lru_cache(maxsize=None) +def cached_tz(hour_str: str, minute_str: str, sign_str: str) -> timezone: + sign = 1 if sign_str == "+" else -1 + return timezone( + timedelta( + hours=sign * int(hour_str), + minutes=sign * int(minute_str), + ) + ) + + +def match_to_localtime(match: re.Match[str]) -> time: + hour_str, minute_str, sec_str, micros_str = match.groups() + micros = int(micros_str.ljust(6, "0")) if micros_str else 0 + return time(int(hour_str), int(minute_str), int(sec_str), micros) + + +def match_to_number(match: re.Match[str], parse_float: ParseFloat) -> Any: + if match.group("floatpart"): + return parse_float(match.group()) + return int(match.group(), 0) + diff --git a/scripts/internal/_tomllib_compat/_types.py b/scripts/internal/_tomllib_compat/_types.py new file mode 100644 index 0000000..cb8380e --- /dev/null +++ b/scripts/internal/_tomllib_compat/_types.py @@ -0,0 +1,11 @@ +# SPDX-License-Identifier: MIT +# SPDX-FileCopyrightText: 2021 Taneli Hukkinen +# Licensed to PSF under a Contributor Agreement. + +from typing import Any, Callable, Tuple + +# Type annotations +ParseFloat = Callable[[str], Any] +Key = Tuple[str, ...] +Pos = int + diff --git a/scripts/internal/code_intelligence_protocol.py b/scripts/internal/code_intelligence_protocol.py index 5ae7ec6..93296a1 100644 --- a/scripts/internal/code_intelligence_protocol.py +++ b/scripts/internal/code_intelligence_protocol.py @@ -2,6 +2,8 @@ from __future__ import annotations +import hashlib +import re from pathlib import Path from typing import Any, Iterable @@ -15,12 +17,18 @@ read_json, subject_diff_hash, task_dir, + utc_now, validate_json_file, validate_schema, write_json_atomic, ) from .task_location_protocol import resolve_repo_reference -from .task_layout import code_intelligence_record_path, state_path +from .task_layout import ( + code_intelligence_record_path, + code_intelligence_runtime_dir, + state_path, + task_relative_path, +) CONFIG_PATH = Path(".polaris/code-intelligence.json") @@ -364,11 +372,18 @@ def validate_historical_legacy_record_value( def _validate_record_identity( - repo: Path, task_id: str, value: dict[str, Any] + repo: Path, + task_id: str, + value: dict[str, Any], + *, + require_current_revision: bool = True, ) -> tuple[Path, str, str | None]: directory = task_dir(repo, task_id) state = read_json(state_path(directory)) - if value["task_id"] != task_id or value["work_item_revision"] != state["current_revision"]: + if value["task_id"] != task_id or ( + require_current_revision + and value["work_item_revision"] != state["current_revision"] + ): raise RuleFailure("Code Intelligence record targets the wrong task revision") _record_name(value) target = value["target"] @@ -532,14 +547,24 @@ def _validate_v2_freshness( def _validate_v2_record_value( - repo: Path, task_id: str, value: dict[str, Any], root: Path + repo: Path, + task_id: str, + value: dict[str, Any], + root: Path, + *, + require_current_revision: bool = True, ) -> dict[str, Any]: errors = validate_schema( - value, read_json(root / "schemas" / "code-intelligence-record.schema.json") + value, read_json(root / "schemas" / "code-intelligence-record-v2.schema.json") ) if errors: raise RuleFailure("Code Intelligence record failed schema validation:\n- " + "\n- ".join(errors)) - _, base, head = _validate_record_identity(repo, task_id, value) + _, base, head = _validate_record_identity( + repo, + task_id, + value, + require_current_revision=require_current_revision, + ) provider = value["provider"] if value["status"] in {"USED", "FAILED"} and provider is None: raise RuleFailure("used or failed Code Intelligence record requires a provider") @@ -689,13 +714,463 @@ def _validate_v2_record_value( return value +def validate_historical_v2_record_value( + repo: Path, + task_id: str, + path: Path, + value: dict[str, Any], + root: Path | None = None, +) -> dict[str, Any]: + """Validate immutable v2 evidence at its canonical historical location.""" + root = protocol_root(repo) if root is None else root + value = _validate_v2_record_value( + repo, + task_id, + value, + root, + require_current_revision=False, + ) + expected = code_intelligence_record_path( + task_dir(repo, task_id), value["work_item_revision"], _record_name(value) + ) + if path != expected: + raise RuleFailure("Code Intelligence record reference uses a non-canonical path") + return value + + +def _require_exact_keys(value: Any, keys: set[str], label: str) -> dict[str, Any]: + if not isinstance(value, dict) or set(value) != keys: + raise RuleFailure(f"{label} has an invalid field set") + return value + + +def _zero_pending(value: Any) -> bool: + return isinstance(value, dict) and value == { + "added": 0, + "modified": 0, + "removed": 0, + } + + +def _is_sha256(value: Any) -> bool: + return isinstance(value, str) and re.fullmatch(r"[0-9a-f]{64}", value) is not None + + +def _validate_v3_stale_point(repo: Path, point: Any) -> dict[str, Any]: + point = _require_exact_keys( + point, + {"scope", "path", "reason", "fallback", "observed_sha256"}, + "Code Intelligence stale point", + ) + if point["scope"] == "FILE": + if ( + not isinstance(point["path"], str) + or point["reason"] != "PENDING_SYNC" + or point["fallback"] not in {"READ_SOURCE", "INSPECT_GIT_DIFF"} + ): + raise RuleFailure("v3 file stale point is invalid") + resolved = resolve_repo_reference(repo, point["path"]) + if point["fallback"] == "READ_SOURCE": + if ( + not resolved.is_file() + or point["observed_sha256"] != file_sha256(resolved) + ): + raise RuleFailure("v3 READ_SOURCE stale point hash is stale") + elif point["observed_sha256"] is not None or resolved.is_file(): + raise RuleFailure("v3 INSPECT_GIT_DIFF stale point must name a missing path") + elif point["scope"] == "INDEX": + if ( + point["path"] is not None + or point["fallback"] != "SEARCH_SOURCE" + or point["observed_sha256"] is not None + or point["reason"] not in { + "PENDING_CHANGES", + "AUTO_SYNC_DISABLED", + "WORKTREE_MISMATCH", + "INDEX_PARTIAL", + "INDEX_INDEXING", + "INDEX_FAILED", + "PENDING_REFERENCES", + "REINDEX_RECOMMENDED", + "SYNC_FAILED", + "STATUS_UNREADABLE", + } + ): + raise RuleFailure("v3 index stale point is invalid") + else: + raise RuleFailure("v3 stale point scope is invalid") + return point + + +def _validate_v3_status_observation( + repo: Path, observation: Any, label: str +) -> dict[str, Any]: + observation = _require_exact_keys( + observation, + { + "status", + "checked_at", + "basis", + "stale_points", + "status_response_sha256", + "error", + "needs_sync", + "pending_changes", + }, + label, + ) + if ( + observation["status"] + not in {"CURRENT_AT_CHECK", "INDEX_STALE", "NOT_VERIFIED", "UNAVAILABLE"} + or not isinstance(observation["checked_at"], str) + or not observation["checked_at"] + or not isinstance(observation["basis"], list) + or not isinstance(observation["needs_sync"], bool) + or not isinstance(observation["stale_points"], list) + ): + raise RuleFailure(f"{label} has invalid status fields") + for point in observation["stale_points"]: + _validate_v3_stale_point(repo, point) + status = observation["status"] + if status in {"CURRENT_AT_CHECK", "INDEX_STALE"}: + if ( + not _is_sha256(observation["status_response_sha256"]) + or not isinstance(observation["pending_changes"], dict) + or set(observation["pending_changes"]) + != {"added", "modified", "removed"} + or any( + isinstance(value, bool) or not isinstance(value, int) or value < 0 + for value in observation["pending_changes"].values() + ) + or observation["error"] is not None + or "STATUS_JSON" not in observation["basis"] + or len(observation["basis"]) != len(set(observation["basis"])) + or not set(observation["basis"]).issubset( + {"STATUS_JSON", "SYNC_ACKNOWLEDGED"} + ) + ): + raise RuleFailure(f"{label} successful status evidence is incomplete") + if status == "CURRENT_AT_CHECK" and observation["stale_points"]: + raise RuleFailure(f"{label} current status cannot contain stale points") + if status == "INDEX_STALE" and not observation["stale_points"]: + raise RuleFailure(f"{label} stale status requires stale points") + elif status == "NOT_VERIFIED": + if ( + observation["pending_changes"] is not None + or not observation["error"] + or observation["basis"] != ["STATUS_JSON"] + or ( + observation["status_response_sha256"] is not None + and not _is_sha256(observation["status_response_sha256"]) + ) + ): + raise RuleFailure(f"{label} NOT_VERIFIED evidence is incomplete") + elif ( + observation["status_response_sha256"] is not None + or observation["pending_changes"] is not None + or not observation["error"] + or observation["basis"] != ["NONE"] + or observation["stale_points"] + ): + raise RuleFailure(f"{label} UNAVAILABLE evidence is invalid") + return observation + + +def _validate_v3_sync(value: Any) -> dict[str, Any] | None: + if value is None: + return None + value = _require_exact_keys( + value, + {"status", "response_sha256", "error"}, + "Code Intelligence v3 sync", + ) + if value["status"] not in {"SUCCESS", "FAILED"}: + raise RuleFailure("v3 sync must describe exactly one attempted command") + if value["status"] == "SUCCESS" and ( + not _is_sha256(value["response_sha256"]) or value["error"] is not None + ): + raise RuleFailure("successful v3 sync requires a response hash") + if value["status"] == "FAILED" and ( + not value["error"] + or ( + value["response_sha256"] is not None + and not _is_sha256(value["response_sha256"]) + ) + ): + raise RuleFailure("failed v3 sync requires an error") + return value + + +def _validate_v3_response(repo: Path, value: Any) -> dict[str, Any] | None: + if value is None: + return None + value = _require_exact_keys( + value, + { + "classification", + "checked_at", + "basis", + "stale_points", + "response_sha256", + "error", + }, + "Code Intelligence v3 response classification", + ) + if ( + value["classification"] + not in {"NONE", "PARTIAL_STALE", "INDEX_STALE", "NOT_VERIFIED"} + or not _is_sha256(value["response_sha256"]) + or value["basis"] != ["RESPONSE_BANNER"] + ): + raise RuleFailure("v3 response classification is invalid") + for point in value["stale_points"]: + _validate_v3_stale_point(repo, point) + if value["classification"] == "NONE" and ( + value["stale_points"] or value["error"] is not None + ): + raise RuleFailure("neutral v3 response cannot contain stale evidence") + if value["classification"] in {"PARTIAL_STALE", "INDEX_STALE"} and ( + not value["stale_points"] or value["error"] is not None + ): + raise RuleFailure("stale v3 response requires explicit stale points") + if value["classification"] == "NOT_VERIFIED" and not value["error"]: + raise RuleFailure("unverified v3 response requires an error") + return value + + +def _validate_v3_fallback_matches(value: dict[str, Any]) -> None: + fallbacks = value["source_fallbacks"] + points = value["delivery"]["stale_points"] + for point in points: + if point["reason"] == "STATUS_UNREADABLE": + continue + if _matching_fallback( + fallbacks, + point["fallback"], + point["path"], + point["observed_sha256"], + ) is None: + raise RuleFailure("v3 stale point requires exact source fallback evidence") + required = value["delivery"]["required_fallback"] + if required == "NONE" and fallbacks: + raise RuleFailure("CURRENT v3 evidence cannot contain source fallbacks") + if required != "NONE" and not fallbacks: + raise RuleFailure("non-current v3 evidence requires source fallback evidence") + if required == "SEARCH_SOURCE" and not any( + fallback["action"] == "SEARCH_SOURCE" for fallback in fallbacks + ): + raise RuleFailure("v3 evidence requires a SEARCH_SOURCE fallback") + + +def _validate_v3_record_value( + repo: Path, task_id: str, value: dict[str, Any], root: Path +) -> dict[str, Any]: + errors = validate_schema( + value, read_json(root / "schemas/code-intelligence-record.schema.json") + ) + if errors: + raise RuleFailure( + "Code Intelligence record failed schema validation:\n- " + + "\n- ".join(errors) + ) + _, base, head = _validate_record_identity(repo, task_id, value) + project = read_json(repo / ".polaris/project.json") + expected_repository = { + "project_id": project["project_id"], + "root_sha256": hashlib.sha256( + str(repo.resolve()).encode("utf-8") + ).hexdigest(), + } + if value["repository"] != expected_repository: + raise RuleFailure("Code Intelligence v3 repository identity does not match") + if value["provider"] != {"id": "codegraph", "descriptor_version": 2}: + raise RuleFailure("Code Intelligence v3 requires the official CodeGraph provider") + query = value["query"] + if query["status"] == "SUCCESS": + if query["response_sha256"] is None or query["error"] is not None: + raise RuleFailure("successful v3 query requires a response hash and no error") + elif not query["error"] or query["response_sha256"] is not None: + raise RuleFailure("unsuccessful v3 query requires only a finite error") + for symbol in query["symbols"]: + if not resolve_repo_reference(repo, symbol["path"]).is_file(): + raise RuleFailure(f"v3 symbol path is not a current file: {symbol['path']}") + + window = value["query_window"] + pre = _validate_v3_status_observation(repo, window["pre_status"], "pre-query status") + sync = _validate_v3_sync(window["sync"]) + post_sync = ( + _validate_v3_status_observation(repo, window["post_sync_status"], "post-sync status") + if window["post_sync_status"] is not None + else None + ) + response = _validate_v3_response(repo, window["response_classification"]) + post = ( + _validate_v3_status_observation(repo, window["post_query_status"], "post-query status") + if window["post_query_status"] is not None + else None + ) + if sync is None and post_sync is not None: + raise RuleFailure("post-sync status requires one sync attempt") + if sync is not None and sync["status"] == "SUCCESS" and post_sync is None: + raise RuleFailure("successful sync requires post-sync status evidence") + if sync is not None and sync["status"] == "SUCCESS" and ( + post_sync["status"] != "CURRENT_AT_CHECK" or post_sync["needs_sync"] + ): + raise RuleFailure("successful sync requires a current post-sync status") + if query["status"] == "SUCCESS" and ( + response is None + or response["response_sha256"] != query["response_sha256"] + or post is None + ): + raise RuleFailure("successful v3 query requires matching response and post-status evidence") + if query["status"] != "SUCCESS" and (response is not None or post is not None): + raise RuleFailure("unsuccessful v3 query cannot claim response or post-status evidence") + + delivery = value["delivery"] + for point in delivery["stale_points"]: + _validate_v3_stale_point(repo, point) + effective = post_sync if post_sync is not None else pre + observed_points: list[dict[str, Any]] = [] + for observation in (effective, post): + if observation is None: + continue + observed_points.extend(observation["stale_points"]) + if ( + observation is effective + and sync is not None + and sync["status"] == "FAILED" + ): + observed_points.append({ + "scope": "INDEX", + "path": None, + "reason": "SYNC_FAILED", + "fallback": "SEARCH_SOURCE", + "observed_sha256": None, + }) + pending = observation["pending_changes"] + if isinstance(pending, dict) and any(pending.values()): + pending_point = { + "scope": "INDEX", + "path": None, + "reason": "PENDING_CHANGES", + "fallback": "SEARCH_SOURCE", + "observed_sha256": None, + } + if pending_point not in observed_points: + observed_points.append(pending_point) + if response is not None: + observed_points.extend(response["stale_points"]) + unique_points: list[dict[str, Any]] = [] + for point in observed_points: + if point not in unique_points: + unique_points.append(point) + if delivery["stale_points"] != unique_points: + raise RuleFailure("v3 delivery stale points do not match the observed query window") + successful_pending = [ + observation["pending_changes"] + for observation in (effective, post) + if observation is not None and isinstance(observation["pending_changes"], dict) + ] + expected_pending = { + key: max((pending[key] for pending in successful_pending), default=0) + for key in ("added", "modified", "removed") + } + if delivery["pending_changes"] != expected_pending: + raise RuleFailure("v3 delivery pending counts do not match the query window") + state = delivery["state"] + expected_record_status = { + "CURRENT": "CURRENT_AT_CHECK", + "STALE": ( + "INDEX_STALE" + if any(point["scope"] == "INDEX" for point in delivery["stale_points"]) + else "PARTIAL_STALE" + ), + "UNKNOWN": "NOT_VERIFIED", + "UNAVAILABLE": "UNAVAILABLE", + }[state] + if delivery["record_status"] != expected_record_status: + raise RuleFailure("v3 delivery state contradicts its record freshness status") + if value["status"] != { + "CURRENT": "USED", + "STALE": "USED", + "UNKNOWN": "FAILED", + "UNAVAILABLE": "UNAVAILABLE", + }[state]: + raise RuleFailure("v3 record status contradicts proxy delivery") + if state == "CURRENT": + if ( + query["status"] != "SUCCESS" + or response["classification"] != "NONE" + or effective["status"] != "CURRENT_AT_CHECK" + or effective["needs_sync"] + or not _zero_pending(effective["pending_changes"]) + or post["status"] != "CURRENT_AT_CHECK" + or post["needs_sync"] + or not _zero_pending(post["pending_changes"]) + or delivery["usage"] != "NON_AUTHORITATIVE_CONTEXT" + or delivery["required_fallback"] != "NONE" + or delivery["stale_points"] + or not _zero_pending(delivery["pending_changes"]) + or delivery["error"] is not None + ): + raise RuleFailure("CURRENT v3 evidence lacks a complete zero-pending window") + elif state == "STALE": + if ( + query["status"] != "SUCCESS" + or delivery["usage"] != "NAVIGATION_ONLY" + or delivery["required_fallback"] == "NONE" + or not any( + point["reason"] != "STATUS_UNREADABLE" + for point in delivery["stale_points"] + ) + ): + raise RuleFailure("STALE v3 evidence lacks an explicit stale reason") + elif state == "UNKNOWN": + if ( + delivery["usage"] != "NAVIGATION_ONLY" + or delivery["required_fallback"] != "SEARCH_SOURCE" + or not delivery["error"] + ): + raise RuleFailure("UNKNOWN v3 evidence lacks verification failure evidence") + elif ( + query["status"] != "UNAVAILABLE" + or pre["status"] != "UNAVAILABLE" + or any(item is not None for item in (sync, post_sync, response, post)) + or delivery["usage"] != "NO_GRAPH" + or delivery["stale_points"] + ): + raise RuleFailure("UNAVAILABLE v3 evidence contains an attempted operation") + _validate_source_fallbacks(repo, value, base, head) + if state in {"STALE", "UNKNOWN"}: + confirmed_paths: set[str] = set() + for fallback in value["source_fallbacks"]: + if fallback["action"] == "READ_SOURCE": + confirmed_paths.add(fallback["path"]) + elif fallback["action"] == "SEARCH_SOURCE": + confirmed_paths.update( + result["path"] for result in fallback["result_paths"] + ) + for symbol in query["symbols"]: + if symbol["path"] not in confirmed_paths: + raise RuleFailure( + "non-current v3 symbol requires current source fallback evidence" + ) + _validate_v3_fallback_matches(value) + return value + + def validate_record_value( repo: Path, task_id: str, value: dict[str, Any], root: Path | None = None ) -> dict[str, Any]: root = protocol_root(repo) if root is None else root - if value.get("record_version") == 1: + version = value.get("record_version") + if version == 1: return validate_legacy_record_value(repo, task_id, value, root) - return _validate_v2_record_value(repo, task_id, value, root) + if version == 2: + return _validate_v2_record_value(repo, task_id, value, root) + if version == 3: + return _validate_v3_record_value(repo, task_id, value, root) + raise RuleFailure("unsupported Code Intelligence record version") def record( @@ -705,8 +1180,8 @@ def record( root: Path | None = None, ) -> dict[str, Any]: root = protocol_root(repo) if root is None else root - if value.get("record_version") == 1: - raise InputFailure("new Code Intelligence records must use record_version 2") + if value.get("record_version") != 3: + raise InputFailure("new Code Intelligence records must use record_version 3") value = validate_record_value(repo, task_id, value, root) directory = task_dir(repo, task_id) destination = code_intelligence_record_path( @@ -725,6 +1200,172 @@ def record( } +def record_proxy_bundle( + repo: Path, + task_id: str, + bundle_path: Path, + annotations: dict[str, Any], + root: Path | None = None, +) -> dict[str, Any]: + """Project one immutable proxy bundle into the only writable v3 record shape.""" + repo = repo.resolve() + root = protocol_root(repo) if root is None else root + directory = task_dir(repo, task_id) + runtime = code_intelligence_runtime_dir(directory) + candidate = bundle_path if bundle_path.is_absolute() else repo / bundle_path + candidate = confined_target(runtime, candidate, "CodeGraph proxy bundle") + require_regular_file(candidate, "CodeGraph proxy bundle") + bundle_digest = file_sha256(candidate) + bundle = read_json(candidate) + _require_exact_keys( + bundle, + { + "bundle_version", + "proxy", + "provider", + "repository", + "task_context", + "query", + "pre_status", + "sync", + "post_sync_status", + "response_classification", + "post_query_status", + "delivery", + "response_path", + }, + "CodeGraph proxy bundle", + ) + if bundle["bundle_version"] != 1 or bundle["proxy"] != { + "server_id": "polaris-codegraph", + "tool": "polaris_codegraph_explore", + }: + raise RuleFailure("CodeGraph proxy bundle has an unsupported identity") + if bundle["provider"] != {"id": "codegraph", "descriptor_version": 2}: + raise RuleFailure("CodeGraph proxy bundle does not use the official provider") + context = bundle["task_context"] + _require_exact_keys( + context, + { + "task_id", + "work_item_revision", + "stage", + "artifact_attempt", + "reviewer_slot", + "record_name", + "target", + }, + "CodeGraph proxy task context", + ) + if context["task_id"] != task_id: + raise RuleFailure("CodeGraph proxy bundle targets the wrong task") + from .code_intelligence_proxy import resolve_stage_context + + if context != resolve_stage_context(repo, task_id, context["stage"]): + raise RuleFailure("CodeGraph proxy bundle stage context is no longer current") + query = _require_exact_keys( + bundle["query"], + {"id", "purpose", "text", "status", "response_sha256", "error"}, + "CodeGraph proxy query", + ) + query_match = re.fullmatch(r"CIQ-([0-9]{3})", str(query["id"])) + if query_match is None or query_match.group(1) == "000": + raise RuleFailure("CodeGraph proxy bundle has an invalid query ID") + expected_path = directory / task_relative_path( + "code_intelligence_proxy_bundle", + record_name=context["record_name"], + query_id=query["id"], + ) + if candidate != expected_path: + raise RuleFailure("CodeGraph proxy bundle uses a non-canonical runtime path") + query_number = int(query_match.group(1)) + for number in range(1, query_number + 1): + prior = directory / task_relative_path( + "code_intelligence_proxy_bundle", + record_name=context["record_name"], + query_id=f"CIQ-{number:03d}", + ) + require_regular_file(prior, "sequential CodeGraph proxy bundle") + project = read_json(repo / ".polaris/project.json") + expected_repository = { + "project_id": project["project_id"], + "root_sha256": hashlib.sha256( + str(repo.resolve()).encode("utf-8") + ).hexdigest(), + } + if bundle["repository"] != expected_repository: + raise RuleFailure("CodeGraph proxy bundle repository identity does not match") + response_path = bundle["response_path"] + if response_path is None: + if query["status"] == "SUCCESS" and bundle["delivery"]["state"] != "UNKNOWN": + raise RuleFailure("successful CodeGraph bundle lost its response evidence") + else: + if not isinstance(response_path, str): + raise RuleFailure("CodeGraph proxy response path is invalid") + response_file = confined_target( + directory, directory / response_path, "CodeGraph proxy response" + ) + require_regular_file(response_file, "CodeGraph proxy response") + if query["response_sha256"] != file_sha256(response_file): + raise RuleFailure("CodeGraph proxy response hash does not match") + annotation_errors = validate_schema( + annotations, + read_json(root / "schemas/code-intelligence-record-annotations.schema.json"), + ) + if annotation_errors: + raise RuleFailure( + "Code Intelligence annotations failed schema validation:\n- " + + "\n- ".join(annotation_errors) + ) + if response_path is None and annotations["symbols"]: + raise RuleFailure("discarded CodeGraph output cannot annotate graph symbols") + delivery = bundle.get("delivery") + if not isinstance(delivery, dict) or delivery.get("state") not in { + "CURRENT", + "STALE", + "UNKNOWN", + "UNAVAILABLE", + }: + raise RuleFailure("CodeGraph proxy bundle has an invalid delivery state") + record_value = { + "record_version": 3, + "task_id": task_id, + "work_item_revision": context["work_item_revision"], + "stage": context["stage"], + "artifact_attempt": context["artifact_attempt"], + "reviewer_slot": context["reviewer_slot"], + "provider": bundle["provider"], + "repository": bundle["repository"], + "target": context["target"], + "status": { + "CURRENT": "USED", + "STALE": "USED", + "UNKNOWN": "FAILED", + "UNAVAILABLE": "UNAVAILABLE", + }[delivery["state"]], + "proxy": { + **bundle["proxy"], + "evidence_bundle_sha256": bundle_digest, + }, + "query": { + **query, + "summary": annotations["summary"], + "symbols": annotations["symbols"], + }, + "query_window": { + "pre_status": bundle["pre_status"], + "sync": bundle["sync"], + "post_sync_status": bundle["post_sync_status"], + "response_classification": bundle["response_classification"], + "post_query_status": bundle["post_query_status"], + }, + "delivery": delivery, + "source_fallbacks": annotations["source_fallbacks"], + "recorded_at": utc_now(), + } + return record(repo, task_id, record_value, root) + + def record_reference(repo: Path, task_id: str, reference: Any) -> dict[str, Any]: if reference is None: return {} diff --git a/scripts/internal/code_intelligence_proxy.py b/scripts/internal/code_intelligence_proxy.py new file mode 100644 index 0000000..24f2e45 --- /dev/null +++ b/scripts/internal/code_intelligence_proxy.py @@ -0,0 +1,619 @@ +"""Bound one CodeGraph explore call to an auditable Polaris freshness window.""" + +from __future__ import annotations + +import hashlib +import re +import shutil +import subprocess +from pathlib import Path +from typing import Any + +from .code_intelligence_protocol import ( + _project_marker_path, + _record_name, + load_config, + load_providers, +) +from .codegraph_adapter import ( + classify_response, + inspect_status, + run_explore, + synchronize_observed_status, +) +from .implementation_protocol import validate_handoff as validate_implementation_handoff +from .path_security import confined_target +from .polaris_core import ( + InputFailure, + RuleFailure, + file_sha256, + full_commit, + protocol_root, + read_json, + require_protocol_compatible, + subject_diff_hash, + task_dir, + utc_now, + validate_json_file, + write_json_atomic, + write_text_atomic, +) +from .review_handoff_protocol import validate_handoff as validate_review_handoff +from .task_layout import state_path, task_relative_path, work_item_path + + +QUERY_ID_PATTERN = re.compile(r"^CIQ-(?P[0-9]{3})$") +STAGE_STATUSES = { + "PLANNING": {"QUALIFIED"}, + "IMPLEMENTATION": {"IMPLEMENTING"}, + "DOCUMENTATION_SYNC": {"IMPLEMENTING"}, + "REVIEW": {"REVIEWING"}, +} +_INDEX_FALLBACK = { + "scope": "INDEX", + "path": None, + "reason": "STATUS_UNREADABLE", + "fallback": "SEARCH_SOURCE", + "observed_sha256": None, +} + + +def _target(base: str, head: str | None, diff_hash: str | None) -> dict[str, str | None]: + return {"base_commit": base, "head_commit": head, "diff_hash": diff_hash} + + +def _validated_subject(repo: Path, subject: Any, fallback_base: str) -> dict[str, str | None]: + if subject is None: + base = full_commit(repo, fallback_base) + head = full_commit(repo) + return _target(base, head, subject_diff_hash(repo, base, head)) + if not isinstance(subject, dict): + raise RuleFailure("CodeGraph stage context has an invalid subject") + base = full_commit(repo, subject.get("base_commit", "")) + head = full_commit(repo, subject.get("head_commit", "")) + digest = subject_diff_hash(repo, base, head) + if subject.get("diff_hash") != digest: + raise RuleFailure("CodeGraph stage context subject diff hash is stale") + return _target(base, head, digest) + + +def resolve_stage_context(repo: Path, task_id: str, stage: str) -> dict[str, Any]: + """Resolve one stage identity from validated frozen task artifacts.""" + if stage not in STAGE_STATUSES: + raise InputFailure(f"invalid CodeGraph stage: {stage}") + root = protocol_root(repo) + directory = task_dir(repo, task_id) + state = validate_json_file(state_path(directory), root / "schemas/task-state.schema.json") + require_protocol_compatible(repo, state) + if state["task_id"] != task_id: + raise RuleFailure("CodeGraph stage context targets the wrong task") + if state["status"] not in STAGE_STATUSES[stage]: + raise RuleFailure( + f"CodeGraph stage {stage} is inconsistent with task status {state['status']}" + ) + revision = state["current_revision"] + work_item = validate_json_file( + work_item_path(directory, revision), root / "schemas/work-item.schema.json" + ) + if work_item["id"] != task_id or work_item["revision"] != revision: + raise RuleFailure("CodeGraph stage context has the wrong frozen Work Item") + + attempt: int | None = None + reviewer_slot: int | None = None + if stage == "PLANNING": + target = _target(full_commit(repo, work_item["base_commit"]), None, None) + elif stage in {"IMPLEMENTATION", "DOCUMENTATION_SYNC"}: + handoff, _reference = validate_implementation_handoff( + repo, root, directory, state + ) + attempt = handoff["artifact_attempt"] + target = _validated_subject(repo, state.get("subject"), handoff["subject_base_commit"]) + else: + handoff = validate_review_handoff(repo, root, directory, state) + attempt = handoff["artifact_attempt"] + target = _target( + full_commit(repo, handoff["subject_base_commit"]), + full_commit(repo, handoff["subject_head_commit"]), + handoff["subject_diff_hash"], + ) + if target["diff_hash"] != subject_diff_hash( + repo, str(target["base_commit"]), str(target["head_commit"]) + ): + raise RuleFailure("CodeGraph Review handoff diff hash is stale") + if state["artifacts"].get("review_2") is not None: + raise RuleFailure("CodeGraph Review already has both reviewer slots") + reviewer_slot = 2 if state["artifacts"].get("review") is not None else 1 + + identity = { + "stage": stage, + "artifact_attempt": attempt, + "reviewer_slot": reviewer_slot, + } + return { + "task_id": task_id, + "work_item_revision": revision, + "stage": stage, + "artifact_attempt": attempt, + "reviewer_slot": reviewer_slot, + "record_name": _record_name(identity), + "target": target, + } + + +def _validated_query_id(query_id: str) -> int: + match = QUERY_ID_PATTERN.fullmatch(query_id) if isinstance(query_id, str) else None + if match is None or match["number"] == "000": + raise InputFailure(f"invalid CodeGraph query id: {query_id}") + return int(match["number"]) + + +def _proxy_path( + repo: Path, + task_id: str, + context: dict[str, Any], + query_id: str, + artifact: str, +) -> Path: + _validated_query_id(query_id) + if context.get("task_id") != task_id: + raise RuleFailure("CodeGraph proxy context targets the wrong task") + record_name = context.get("record_name") + if not isinstance(record_name, str) or re.fullmatch( + r"(?:planning|implementation-[0-9]{3}|documentation-sync-[0-9]{3}|review-[0-9]{3}-slot-[12])", + record_name, + ) is None: + raise RuleFailure("CodeGraph proxy context has an invalid record name") + directory = task_dir(repo, task_id) + relative = task_relative_path( + artifact, record_name=record_name, query_id=query_id + ) + return confined_target(directory, directory / relative, "CodeGraph proxy evidence") + + +def proxy_bundle_path( + repo: Path, task_id: str, context: dict[str, Any], query_id: str +) -> Path: + """Return the next immutable bundle path for this stage record.""" + number = _validated_query_id(query_id) + destination = _proxy_path( + repo, task_id, context, query_id, "code_intelligence_proxy_bundle" + ) + parent = destination.parent + confined_target(task_dir(repo, task_id), parent, "CodeGraph proxy runtime directory") + existing_numbers: list[int] = [] + if parent.exists(): + if not parent.is_dir(): + raise RuleFailure("CodeGraph proxy runtime path is not a directory") + for path in parent.glob("CIQ-*.json"): + match = QUERY_ID_PATTERN.fullmatch(path.stem) + if match is None or path.is_symlink() or not path.is_file(): + raise RuleFailure(f"invalid CodeGraph proxy bundle path: {path}") + existing_numbers.append(int(match["number"])) + expected_numbers = list(range(1, len(existing_numbers) + 1)) + if sorted(existing_numbers) != expected_numbers: + raise RuleFailure("existing CodeGraph proxy query IDs are not sequential") + expected = len(existing_numbers) + 1 + if expected > 999: + raise InputFailure("CodeGraph proxy query limit exceeded for this stage") + if number != expected: + raise InputFailure(f"CodeGraph query id must be the next sequential ID CIQ-{expected:03d}") + if destination.exists() or destination.is_symlink(): + raise InputFailure(f"CodeGraph proxy bundle is immutable: {destination}") + return destination + + +def _unavailable_status(reason: str) -> dict[str, Any]: + return { + "status": "UNAVAILABLE", + "checked_at": utc_now(), + "basis": ["NONE"], + "stale_points": [], + "status_response_sha256": None, + "error": reason[:240], + "needs_sync": False, + "pending_changes": None, + } + + +def _pending_point() -> dict[str, Any]: + return {**_INDEX_FALLBACK, "reason": "PENDING_CHANGES"} + + +def _observation_points(observation: dict[str, Any] | None) -> list[dict[str, Any]]: + if observation is None: + return [] + points = list(observation.get("stale_points", [])) + pending = observation.get("pending_changes") + if isinstance(pending, dict) and any(pending.get(key, 0) for key in ("added", "modified", "removed")): + if _pending_point() not in points: + points.append(_pending_point()) + return points + + +def _is_unknown(observation: dict[str, Any] | None) -> bool: + return observation is not None and observation.get("status") == "NOT_VERIFIED" + + +def _successful_statuses(*observations: dict[str, Any] | None) -> list[dict[str, Any]]: + return [ + item + for item in observations + if item is not None and isinstance(item.get("pending_changes"), dict) + ] + + +def _pending_counts(*observations: dict[str, Any] | None) -> dict[str, int]: + values = _successful_statuses(*observations) + return { + key: max((item["pending_changes"][key] for item in values), default=0) + for key in ("added", "modified", "removed") + } + + +def _deduplicate(items: list[dict[str, Any]]) -> list[dict[str, Any]]: + result: list[dict[str, Any]] = [] + for item in items: + if item not in result: + result.append(item) + return result + + +def _unsafe_response(classification: dict[str, Any]) -> bool: + error = str(classification.get("error") or "").lower() + return classification.get("classification") == "NOT_VERIFIED" and any( + token in error + for token in ( + "path", + "symlink", + "regular file", + "escapes", + "repository reference", + ) + ) + + +def _delivery( + effective_pre: dict[str, Any], + query_result: dict[str, Any], + classification: dict[str, Any] | None, + post_status: dict[str, Any] | None, + *, + forced_unknown: str | None = None, +) -> dict[str, Any]: + checked_at = ( + (post_status or {}).get("checked_at") + or (classification or {}).get("checked_at") + or query_result.get("checked_at") + or effective_pre.get("checked_at") + or utc_now() + ) + points = _deduplicate([ + *_observation_points(effective_pre), + *((classification or {}).get("stale_points", [])), + *_observation_points(post_status), + ]) + known_stale = any( + point.get("reason") != "STATUS_UNREADABLE" for point in points + ) or effective_pre.get("status") in {"PARTIAL_STALE", "INDEX_STALE"} or ( + post_status is not None + and post_status.get("status") in {"PARTIAL_STALE", "INDEX_STALE"} + ) or (classification or {}).get("classification") in {"PARTIAL_STALE", "INDEX_STALE"} + unknown = ( + forced_unknown is not None + or query_result.get("status") != "SUCCESS" + or _is_unknown(effective_pre) + or post_status is None + or _is_unknown(post_status) + or (classification or {}).get("classification") == "NOT_VERIFIED" + ) + errors = [ + forced_unknown, + effective_pre.get("error"), + query_result.get("error"), + (classification or {}).get("error"), + (post_status or {}).get("error"), + ] + error = next((str(item)[:240] for item in errors if item), None) + if known_stale: + index_points = [point for point in points if point.get("scope") == "INDEX"] + state = "STALE" + record_status = "INDEX_STALE" if index_points else "PARTIAL_STALE" + reason = next( + ( + str(point["reason"]) + for point in points + if point.get("reason") != "STATUS_UNREADABLE" + ), + "INDEX_STALE", + ) + actions = {point.get("fallback") for point in points} + required_fallback = ( + "SEARCH_SOURCE" + if index_points or len(actions) != 1 + else str(next(iter(actions))) + ) + usage = "NAVIGATION_ONLY" + elif unknown: + state = "UNKNOWN" + record_status = "NOT_VERIFIED" + if forced_unknown: + reason = "RESPONSE_INTEGRITY_UNVERIFIED" + elif _is_unknown(effective_pre): + reason = ( + "PROJECT_MISMATCH" + if "different project" in str(effective_pre.get("error", "")).lower() + else "STATUS_UNREADABLE" + ) + elif query_result.get("status") != "SUCCESS": + reason = "EXPLORE_FAILED" + elif (classification or {}).get("classification") == "NOT_VERIFIED": + reason = "RESPONSE_NOT_VERIFIED" + else: + reason = "POST_STATUS_UNREADABLE" + required_fallback = "SEARCH_SOURCE" + usage = "NAVIGATION_ONLY" + else: + state = "CURRENT" + record_status = "CURRENT_AT_CHECK" + reason = "VERIFIED_WINDOW" + required_fallback = "NONE" + usage = "NON_AUTHORITATIVE_CONTEXT" + return { + "state": state, + "record_status": record_status, + "reason": reason, + "checked_at": checked_at, + "usage": usage, + "required_fallback": required_fallback, + "stale_points": points, + "pending_changes": _pending_counts(effective_pre, post_status), + "error": error, + } + + +def _bundle_base( + repo: Path, + context: dict[str, Any], + query_id: str, + purpose: str, + query: str, + descriptor: dict[str, Any], +) -> dict[str, Any]: + project = read_json(repo / ".polaris/project.json") + return { + "bundle_version": 1, + "proxy": { + "server_id": "polaris-codegraph", + "tool": "polaris_codegraph_explore", + }, + "provider": { + "id": descriptor["provider_id"], + "descriptor_version": descriptor["provider_version"], + }, + "repository": { + "project_id": project["project_id"], + "root_sha256": hashlib.sha256(str(repo.resolve()).encode("utf-8")).hexdigest(), + }, + "task_context": context, + "query": { + "id": query_id, + "purpose": purpose, + "text": query, + "status": "UNAVAILABLE", + "response_sha256": None, + "error": None, + }, + "pre_status": None, + "sync": None, + "post_sync_status": None, + "response_classification": None, + "post_query_status": None, + "delivery": None, + "response_path": None, + } + + +def _write_bundle(path: Path, bundle: dict[str, Any]) -> None: + if path.exists() or path.is_symlink(): + raise InputFailure(f"CodeGraph proxy bundle is immutable: {path}") + path.parent.mkdir(parents=True, exist_ok=True) + confined_target(path.parents[3], path, "CodeGraph proxy evidence") + write_json_atomic(path, bundle) + + +def execute_proxy_query( + repo: Path, + task_id: str, + stage: str, + query_id: str, + purpose: str, + query: str, + sync_if_needed: bool, + *, + runner: Any = subprocess.run, +) -> dict[str, Any]: + """Execute one immutable CodeGraph query window and persist its evidence.""" + if not isinstance(purpose, str) or not purpose.strip() or len(purpose) > 240: + raise InputFailure("CodeGraph query purpose must contain 1 to 240 characters") + if not isinstance(query, str) or not query.strip() or len(query) > 8000: + raise InputFailure("CodeGraph query must contain 1 to 8000 characters") + if not isinstance(sync_if_needed, bool): + raise InputFailure("sync_if_needed must be a boolean") + repo = repo.absolute() + if repo.is_symlink() or not repo.is_dir(): + raise RuleFailure("CodeGraph proxy repository root must be a fixed real directory") + repo = repo.resolve() + context = resolve_stage_context(repo, task_id, stage) + bundle_path = proxy_bundle_path(repo, task_id, context, query_id) + response_path = _proxy_path( + repo, task_id, context, query_id, "code_intelligence_proxy_response" + ) + if response_path.exists() or response_path.is_symlink(): + raise InputFailure(f"CodeGraph proxy response is immutable: {response_path}") + root = protocol_root(repo) + descriptor = load_providers(root)["codegraph"] + bundle = _bundle_base(repo, context, query_id, purpose.strip(), query, descriptor) + + config = load_config(repo, root) + marker = _project_marker_path(repo, descriptor["project_marker"]) + unavailable_reason: str | None = None + if config["mode"] == "disabled": + unavailable_reason = "POLICY_DISABLED" + elif not marker.is_dir() or marker.is_symlink(): + unavailable_reason = "MARKER_UNAVAILABLE" + elif shutil.which(descriptor["cli"]["executable"]) is None: + unavailable_reason = "CLI_UNAVAILABLE" + if unavailable_reason is not None: + pre_status = _unavailable_status(unavailable_reason) + bundle["pre_status"] = pre_status + bundle["query"]["error"] = unavailable_reason + bundle["delivery"] = { + "state": "UNAVAILABLE", + "record_status": "UNAVAILABLE", + "reason": unavailable_reason, + "checked_at": pre_status["checked_at"], + "usage": "NO_GRAPH", + "required_fallback": "SEARCH_SOURCE", + "stale_points": [], + "pending_changes": {"added": 0, "modified": 0, "removed": 0}, + "error": unavailable_reason, + } + _write_bundle(bundle_path, bundle) + return { + "bundle": bundle, + "bundle_path": bundle_path, + "response": None, + "envelope": render_freshness_envelope(bundle), + } + + pre_status = inspect_status(repo, descriptor, runner=runner) + bundle["pre_status"] = pre_status + effective_pre = pre_status + if sync_if_needed and pre_status.get("needs_sync"): + synchronized = synchronize_observed_status( + repo, descriptor, pre_status, runner=runner + ) + bundle["sync"] = synchronized["sync"] + effective_pre = synchronized["freshness"] + bundle["post_sync_status"] = synchronized["post_sync_status"] + + if effective_pre["status"] in {"UNAVAILABLE", "NOT_VERIFIED"}: + bundle["query"]["status"] = ( + "UNAVAILABLE" if effective_pre["status"] == "UNAVAILABLE" else "FAILED" + ) + bundle["query"]["error"] = effective_pre.get("error") + if effective_pre["status"] == "UNAVAILABLE": + bundle["delivery"] = { + "state": "UNAVAILABLE", + "record_status": "UNAVAILABLE", + "reason": "PROVIDER_UNAVAILABLE", + "checked_at": effective_pre["checked_at"], + "usage": "NO_GRAPH", + "required_fallback": "SEARCH_SOURCE", + "stale_points": [], + "pending_changes": {"added": 0, "modified": 0, "removed": 0}, + "error": effective_pre.get("error"), + } + else: + bundle["delivery"] = _delivery( + effective_pre, + bundle["query"], + None, + None, + ) + _write_bundle(bundle_path, bundle) + return { + "bundle": bundle, + "bundle_path": bundle_path, + "response": None, + "envelope": render_freshness_envelope(bundle), + } + + query_result = run_explore(repo, descriptor, query, runner=runner) + bundle["query"].update({ + "status": query_result["status"], + "response_sha256": ( + query_result["response_sha256"] + if query_result["status"] == "SUCCESS" + else None + ), + "error": query_result["error"], + }) + response: str | None = query_result.get("response") + classification: dict[str, Any] | None = None + post_status: dict[str, Any] | None = None + forced_unknown: str | None = None + if query_result["status"] == "SUCCESS" and response is not None: + actual_digest = hashlib.sha256(response.encode("utf-8")).hexdigest() + if actual_digest != query_result["response_sha256"]: + forced_unknown = "CodeGraph response digest mismatch" + response = None + else: + classification = classify_response(repo, response) + bundle["response_classification"] = classification + if classification.get("response_sha256") != actual_digest: + forced_unknown = "CodeGraph classification digest mismatch" + response = None + elif _unsafe_response(classification): + forced_unknown = "CodeGraph response contains an unsafe repository path" + response = None + else: + response_path.parent.mkdir(parents=True, exist_ok=True) + confined_target(task_dir(repo, task_id), response_path, "CodeGraph response") + write_text_atomic(response_path, response) + if file_sha256(response_path) != actual_digest: + forced_unknown = "persisted CodeGraph response digest mismatch" + response_path.unlink() + response = None + else: + bundle["response_path"] = response_path.relative_to( + task_dir(repo, task_id) + ).as_posix() + post_status = inspect_status(repo, descriptor, runner=runner) + bundle["post_query_status"] = post_status + + bundle["delivery"] = _delivery( + effective_pre, + bundle["query"], + classification, + post_status, + forced_unknown=forced_unknown, + ) + if response is None: + bundle["response_path"] = None + _write_bundle(bundle_path, bundle) + return { + "bundle": bundle, + "bundle_path": bundle_path, + "response": response, + "envelope": render_freshness_envelope(bundle), + } + + +def render_freshness_envelope(bundle: dict[str, Any]) -> str: + """Render the finite freshness block that must precede graph content.""" + delivery = bundle["delivery"] + pending = delivery.get("pending_changes") or {} + error = " ".join(str(delivery.get("error") or "").split())[:240] + bundle_path = task_relative_path( + "code_intelligence_proxy_bundle", + record_name=bundle["task_context"]["record_name"], + query_id=bundle["query"]["id"], + ).as_posix() + lines = [ + "[POLARIS_CODEGRAPH_FRESHNESS]", + f"state: {delivery['state']}", + f"record_status: {delivery['record_status']}", + f"reason: {delivery['reason']}", + f"checked_at: {delivery['checked_at']}", + f"pending_added: {pending.get('added', 0)}", + f"pending_modified: {pending.get('modified', 0)}", + f"pending_removed: {pending.get('removed', 0)}", + f"usage: {delivery['usage']}", + f"required_fallback: {delivery['required_fallback']}", + f"evidence_bundle: {bundle_path}", + ] + if error: + lines.append(f"error: {error}") + lines.append("[/POLARIS_CODEGRAPH_FRESHNESS]") + return "\n".join(lines) + "\n" diff --git a/scripts/internal/codegraph_adapter.py b/scripts/internal/codegraph_adapter.py index 5e586b0..0d51a1d 100644 --- a/scripts/internal/codegraph_adapter.py +++ b/scripts/internal/codegraph_adapter.py @@ -43,6 +43,10 @@ r"^ - (?P.+) \(edited [^\n()]+, pending sync\)$" ) _DISABLED_BANNER = "⚠️ CodeGraph auto-sync is DISABLED — the index is frozen." +_SUSPICIOUS_FRESHNESS_SIGNAL = re.compile( + r"(?:⚠|\bwarning\b|\bstale\b|\bpending(?:[- ]sync)?\b|\bout[- ]of[- ]date\b)", + re.IGNORECASE, +) def _checked_at() -> str: @@ -204,6 +208,12 @@ def classify_response( return result if not normalized.startswith(_PARTIAL_BANNER_HEADER): + if _SUSPICIOUS_FRESHNESS_SIGNAL.search(normalized): + result = _response_not_verified( + checked_at, "unrecognized CodeGraph freshness warning" + ) + result["response_sha256"] = response_sha256 + return result result = _response_result("NONE", checked_at, stale_points=[]) result["response_sha256"] = response_sha256 return result @@ -328,9 +338,14 @@ def _run_cli( args_key: str, timeout_seconds: float, runner: Runner, + extra_args: list[str] | None = None, ) -> subprocess.CompletedProcess[str]: timeout = _validated_timeout(timeout_seconds) - command = [descriptor["cli"]["executable"], *descriptor["cli"][args_key]] + command = [ + descriptor["cli"]["executable"], + *descriptor["cli"][args_key], + *(extra_args or []), + ] return runner( command, cwd=repo, @@ -423,7 +438,7 @@ def _status_result( stale_points=[_index_point(reason) for reason in stale_reasons], status_response_sha256=response_sha256, error=None, - needs_sync=False, + needs_sync=any(pending.values()), pending_changes=pending, ) @@ -476,6 +491,66 @@ def inspect_status( return _not_verified(checked_at, error, response_sha256) +def run_explore( + repo: Path, + descriptor: dict[str, Any], + query: str, + *, + runner: Runner = subprocess.run, + timeout_seconds: float = 60, +) -> dict[str, Any]: + """Run exactly one bounded CodeGraph explore command in ``repo``.""" + checked_at = _checked_at() + if not isinstance(query, str) or not query.strip(): + return { + "status": "FAILED", + "checked_at": checked_at, + "response": None, + "response_sha256": None, + "error": "CodeGraph query must not be blank", + } + try: + completed = _run_cli( + repo, + descriptor, + "explore_args", + timeout_seconds, + runner, + extra_args=[query], + ) + raw, response_sha256 = _stdout_and_hash(completed) + except ( + KeyError, + OSError, + TypeError, + UnicodeError, + ValueError, + subprocess.TimeoutExpired, + ) as error: + return { + "status": "FAILED", + "checked_at": checked_at, + "response": None, + "response_sha256": None, + "error": _error_summary(error), + } + if completed.returncode != 0: + return { + "status": "FAILED", + "checked_at": checked_at, + "response": None, + "response_sha256": response_sha256, + "error": f"CodeGraph explore exited with {completed.returncode}", + } + return { + "status": "SUCCESS", + "checked_at": checked_at, + "response": raw, + "response_sha256": response_sha256, + "error": None, + } + + def _sync_result(status: str, response_sha256: str | None, error: str | None) -> dict[str, Any]: return { "status": status, @@ -487,6 +562,8 @@ def _sync_result(status: str, response_sha256: str | None, error: str | None) -> def _sync_failed( freshness: dict[str, Any], sync: dict[str, Any], + *, + post_sync_status: dict[str, Any] | None = None, ) -> dict[str, Any]: points = [*freshness["stale_points"], _index_point("SYNC_FAILED")] return { @@ -498,33 +575,36 @@ def _sync_failed( "error": sync["error"] or freshness["error"], }, "sync": sync, + "post_sync_status": post_sync_status, } -def sync_if_needed( +def synchronize_observed_status( repo: Path, descriptor: dict[str, Any], + initial: dict[str, Any], *, runner: Runner = subprocess.run, status_timeout_seconds: float = 15, sync_timeout_seconds: float = 120, ) -> dict[str, Any]: - """Synchronize at most once, then inspect status at most once more.""" + """Synchronize one already-observed status at most once, then recheck once.""" try: status_timeout = _validated_timeout(status_timeout_seconds) sync_timeout = _validated_timeout(sync_timeout_seconds) except ValueError as error: freshness = _not_verified(_checked_at(), error) - return {"freshness": freshness, "sync": _sync_result("SKIPPED", None, None)} - initial = inspect_status( - repo, descriptor, runner=runner, timeout_seconds=status_timeout - ) + return { + "freshness": freshness, + "sync": _sync_result("SKIPPED", None, None), + "post_sync_status": None, + } skipped = _sync_result("SKIPPED", None, None) unavailable = _sync_result("UNAVAILABLE", None, None) if initial["status"] == "UNAVAILABLE" or _marker_path(repo, descriptor) is None: - return {"freshness": initial, "sync": unavailable} + return {"freshness": initial, "sync": unavailable, "post_sync_status": None} if not initial["needs_sync"]: - return {"freshness": initial, "sync": skipped} + return {"freshness": initial, "sync": skipped, "post_sync_status": None} try: completed = _run_cli(repo, descriptor, "sync_args", sync_timeout, runner) @@ -555,6 +635,36 @@ def sync_if_needed( response_sha256, "CodeGraph post-sync status is not current", ), + post_sync_status=rechecked, ) rechecked["basis"] = [*rechecked["basis"], "SYNC_ACKNOWLEDGED"] - return {"freshness": rechecked, "sync": sync} + return {"freshness": rechecked, "sync": sync, "post_sync_status": rechecked} + + +def sync_if_needed( + repo: Path, + descriptor: dict[str, Any], + *, + runner: Runner = subprocess.run, + status_timeout_seconds: float = 15, + sync_timeout_seconds: float = 120, +) -> dict[str, Any]: + """Inspect once, synchronize at most once, then inspect at most once more.""" + try: + status_timeout = _validated_timeout(status_timeout_seconds) + sync_timeout = _validated_timeout(sync_timeout_seconds) + except ValueError as error: + freshness = _not_verified(_checked_at(), error) + return {"freshness": freshness, "sync": _sync_result("SKIPPED", None, None)} + initial = inspect_status( + repo, descriptor, runner=runner, timeout_seconds=status_timeout + ) + result = synchronize_observed_status( + repo, + descriptor, + initial, + runner=runner, + status_timeout_seconds=status_timeout, + sync_timeout_seconds=sync_timeout, + ) + return {"freshness": result["freshness"], "sync": result["sync"]} diff --git a/scripts/internal/host_adapters.py b/scripts/internal/host_adapters.py index 7ce6a4f..238a88e 100644 --- a/scripts/internal/host_adapters.py +++ b/scripts/internal/host_adapters.py @@ -172,6 +172,23 @@ def load_host_adapters(root: Path) -> list[dict[str, Any]]: adapter["skill_target"], "skill_target" ) current_targets = [(skill_target, f"{host_id} skill_target")] + project_mcp = adapter["project_mcp"] + project_mcp_target = _relative_path( + project_mcp["target"], "project_mcp.target" + ) + if project_mcp["server_id"] != "polaris-codegraph": + raise RuleFailure(f"host adapter has an invalid MCP server ID: {path}") + if project_mcp["format"] not in {"codex-toml", "claude-json"}: + raise RuleFailure(f"host adapter has an invalid MCP format: {path}") + if project_mcp["command"] != "python3": + raise RuleFailure(f"host adapter has an invalid MCP command: {path}") + if project_mcp["args"] != [ + "tools/polaris/scripts/code_intelligence_mcp.py", + "--repo", + ".", + ]: + raise RuleFailure(f"host adapter has invalid MCP arguments: {path}") + current_targets.append((project_mcp_target, f"{host_id} project_mcp")) overlay = adapter["skill_overlay_root"] if overlay is not None: _validate_overlay(path.parent, overlay, root, available_skills) diff --git a/scripts/internal/migration_protocol.py b/scripts/internal/migration_protocol.py index deb7a7a..1477666 100644 --- a/scripts/internal/migration_protocol.py +++ b/scripts/internal/migration_protocol.py @@ -8,7 +8,7 @@ from .code_intelligence_protocol import ( _record_name, validate_historical_legacy_record_value, - validate_record_value, + validate_historical_v2_record_value, ) from .polaris_core import ( InputFailure, @@ -272,9 +272,13 @@ def _step_for_record( def _retired_code_intelligence_records( - repo: Path, task_id: str, directory: Path, protocol_root: Path + repo: Path, + task_id: str, + directory: Path, + protocol_root: Path, + step: dict[str, Any], ) -> list[dict[str, str]]: - """Inventory immutable v1 records at their canonical task-local locations.""" + """Inventory immutable historical records at canonical task-local locations.""" records_root = directory / "code-intelligence" if records_root.is_symlink(): raise RuleFailure( @@ -297,9 +301,15 @@ def _retired_code_intelligence_records( repo, task_id, path, value, protocol_root ) elif value.get("record_version") == 2: - value = validate_record_value(repo, task_id, value, protocol_root) + value = validate_historical_v2_record_value( + repo, task_id, path, value, protocol_root + ) else: raise RuleFailure(f"Code Intelligence record path is non-canonical: {path}") + should_inventory = value["record_version"] == 1 or ( + value["record_version"] == 2 + and step["migration_id"] == "0.1.20-to-0.1.21" + ) if value["record_version"] != 1: expected = code_intelligence_record_path( directory, value["work_item_revision"], _record_name(value) @@ -308,6 +318,7 @@ def _retired_code_intelligence_records( raise RuleFailure( f"Code Intelligence record path is non-canonical: {path}" ) + if not should_inventory: continue retired.append( { @@ -349,7 +360,9 @@ def _new_record( } ) retired_code_intelligence_records.extend( - _retired_code_intelligence_records(repo, task_id, directory, protocol_root) + _retired_code_intelligence_records( + repo, task_id, directory, protocol_root, step + ) ) return { "record_version": 2, @@ -437,6 +450,23 @@ def migrate_project(repo: Path, protocol_root: Path) -> dict[str, Any]: item["task_id"] for item in record["tasks"] }: raise RuleFailure("project task list changed during migration") + if incomplete is not None: + current_inventory: list[dict[str, str]] = [] + for item in record["tasks"]: + directory = task_dir(repo, item["task_id"]) + current_inventory.extend( + _retired_code_intelligence_records( + repo, + item["task_id"], + directory, + protocol_root, + step, + ) + ) + if record.get("retired_code_intelligence_records", []) != current_inventory: + raise RuleFailure( + "retired Code Intelligence record inventory changed during migration" + ) locks: list[tuple[Path, int]] = [] try: diff --git a/scripts/internal/project_mcp_registration.py b/scripts/internal/project_mcp_registration.py new file mode 100644 index 0000000..5ec4f32 --- /dev/null +++ b/scripts/internal/project_mcp_registration.py @@ -0,0 +1,232 @@ +"""Render and validate project-local Polaris MCP registrations.""" + +from __future__ import annotations + +import copy +import json +import math +import re +from pathlib import Path +from typing import Any + +try: + import tomllib +except ModuleNotFoundError: # Python 3.10 compatibility + from . import _tomllib_compat as tomllib + +from .path_security import confined_target, require_regular_file +from .polaris_core import InputFailure, RuleFailure + + +SERVER_ID = "polaris-codegraph" +TOOL_NAME = "polaris_codegraph_explore" +LAUNCHER = "tools/polaris/scripts/code_intelligence_mcp.py" +ARGS = [LAUNCHER, "--repo", "."] +CODEX_START = f"# POLARIS_MCP_START {SERVER_ID}" +CODEX_END = f"# POLARIS_MCP_END {SERVER_ID}" + + +def project_mcp_target(repo: Path, adapter: dict[str, Any]) -> Path: + value = adapter["project_mcp"]["target"] + relative = Path(value) + if relative.is_absolute() or not relative.parts or ".." in relative.parts: + raise RuleFailure(f"project MCP target must be a safe relative path: {value}") + return confined_target(repo, repo / relative, "project MCP target") + + +def _definition(adapter: dict[str, Any]) -> dict[str, Any]: + registration = adapter["project_mcp"] + if registration["server_id"] != SERVER_ID: + raise RuleFailure("project MCP registration has an invalid server ID") + if registration["command"] != "python3" or registration["args"] != ARGS: + raise RuleFailure("project MCP registration has an invalid launcher") + return registration + + +def _codex_definition(adapter: dict[str, Any]) -> dict[str, Any]: + registration = _definition(adapter) + return { + "command": registration["command"], + "args": registration["args"], + "cwd": ".", + "enabled": True, + "required": False, + "enabled_tools": [TOOL_NAME], + } + + +def _codex_block(adapter: dict[str, Any]) -> str: + definition = _codex_definition(adapter) + args = json.dumps(definition["args"], ensure_ascii=False) + tools = json.dumps(definition["enabled_tools"], ensure_ascii=False) + return ( + f"{CODEX_START}\n" + f"[mcp_servers.{SERVER_ID}]\n" + f'command = "{definition["command"]}"\n' + f"args = {args}\n" + f'cwd = "{definition["cwd"]}"\n' + "enabled = true\n" + "required = false\n" + f"enabled_tools = {tools}\n" + f"{CODEX_END}\n" + ) + + +def _parse_toml(source: str) -> dict[str, Any]: + try: + value = tomllib.loads(source) + except (tomllib.TOMLDecodeError, UnicodeDecodeError) as exc: + raise RuleFailure(f"project MCP TOML is invalid: {exc}") from exc + if not isinstance(value, dict): + raise RuleFailure("project MCP TOML root must be a table") + return value + + +def _without_managed_server(value: dict[str, Any]) -> dict[str, Any]: + cleaned = copy.deepcopy(value) + servers = cleaned.get("mcp_servers") + if isinstance(servers, dict): + servers.pop(SERVER_ID, None) + if not servers: + cleaned.pop("mcp_servers", None) + return cleaned + + +def _toml_values_equal(left: Any, right: Any) -> bool: + if type(left) is not type(right): + return False + if isinstance(left, float) and math.isnan(left) and math.isnan(right): + return True + if isinstance(left, dict): + return left.keys() == right.keys() and all( + _toml_values_equal(left[key], right[key]) for key in left + ) + if isinstance(left, list): + return len(left) == len(right) and all( + _toml_values_equal(left_item, right_item) + for left_item, right_item in zip(left, right) + ) + return left == right + + +def _merge_codex(adapter: dict[str, Any], source: str) -> str: + parsed = _parse_toml(source) + starts = source.count(CODEX_START) + ends = source.count(CODEX_END) + if starts != ends or starts > 1: + raise RuleFailure("project MCP TOML has malformed or duplicate managed markers") + block = _codex_block(adapter) + if starts == 1: + pattern = re.compile( + rf"(?m)^{re.escape(CODEX_START)}\n.*?^{re.escape(CODEX_END)}(?:\n|$)", + re.DOTALL, + ) + if not pattern.search(source): + raise RuleFailure("project MCP TOML has malformed managed markers") + rendered = pattern.sub(block, source, count=1) + else: + servers = parsed.get("mcp_servers", {}) + if not isinstance(servers, dict): + raise RuleFailure("project MCP TOML mcp_servers must be a table") + if SERVER_ID in servers: + raise RuleFailure( + f"project MCP TOML has a conflicting unmanaged {SERVER_ID} entry" + ) + separator = "" if not source else ("\n" if source.endswith("\n") else "\n\n") + rendered = source + separator + block + final = _parse_toml(rendered) + if not _toml_values_equal( + _without_managed_server(parsed), _without_managed_server(final) + ): + raise RuleFailure("project MCP markers would rewrite unrelated TOML") + servers = final.get("mcp_servers") + if not isinstance(servers, dict) or servers.get(SERVER_ID) != _codex_definition( + adapter + ): + raise RuleFailure("project MCP TOML did not render the exact Polaris entry") + return rendered + + +def _claude_definition(adapter: dict[str, Any]) -> dict[str, Any]: + registration = _definition(adapter) + return { + "type": "stdio", + "command": registration["command"], + "args": registration["args"], + "env": {}, + } + + +def _parse_json(source: str) -> dict[str, Any]: + try: + value = json.loads(source) + except (json.JSONDecodeError, UnicodeDecodeError) as exc: + raise RuleFailure(f"project MCP JSON is invalid: {exc}") from exc + if not isinstance(value, dict): + raise RuleFailure("project MCP JSON root must be an object") + return value + + +def _merge_claude(adapter: dict[str, Any], source: str) -> str: + value = _parse_json(source or "{}") + servers = value.setdefault("mcpServers", {}) + if not isinstance(servers, dict): + raise RuleFailure("project MCP JSON mcpServers must be an object") + expected = _claude_definition(adapter) + existing = servers.get(SERVER_ID) + if existing is not None and existing != expected: + raise RuleFailure(f"project MCP JSON has a conflicting {SERVER_ID} entry") + servers[SERVER_ID] = expected + return json.dumps(value, ensure_ascii=False, indent=4) + "\n" + + +def merge_project_mcp( + repo: Path, + adapter: dict[str, Any], + source_text: str | None = None, +) -> str: + registration = _definition(adapter) + target = project_mcp_target(repo, adapter) + if source_text is None: + if target.exists(): + require_regular_file(target, "project MCP configuration") + try: + source_text = target.read_text(encoding="utf-8") + except UnicodeDecodeError as exc: + raise InputFailure( + f"project MCP configuration is not UTF-8: {target}" + ) from exc + else: + source_text = "" + if registration["format"] == "codex-toml": + return _merge_codex(adapter, source_text) + if registration["format"] == "claude-json": + return _merge_claude(adapter, source_text) + raise RuleFailure(f"unsupported project MCP format: {registration['format']}") + + +def validate_project_mcp(repo: Path, adapter: dict[str, Any]) -> None: + registration = _definition(adapter) + target = project_mcp_target(repo, adapter) + require_regular_file(target, "project MCP configuration") + source = target.read_text(encoding="utf-8") + if registration["format"] == "codex-toml": + if source.count(CODEX_START) != 1 or source.count(CODEX_END) != 1: + raise RuleFailure("project MCP TOML lacks the unique managed block") + parsed = _parse_toml(source) + servers = parsed.get("mcp_servers") + if not isinstance(servers, dict) or servers.get(SERVER_ID) != _codex_definition( + adapter + ): + raise RuleFailure("project MCP TOML Polaris entry is invalid") + elif registration["format"] == "claude-json": + parsed = _parse_json(source) + servers = parsed.get("mcpServers") + if not isinstance(servers, dict) or servers.get(SERVER_ID) != _claude_definition( + adapter + ): + raise RuleFailure("project MCP JSON Polaris entry is invalid") + else: + raise RuleFailure(f"unsupported project MCP format: {registration['format']}") + launcher = confined_target(repo, repo / LAUNCHER, "project MCP launcher") + require_regular_file(launcher, "project MCP launcher") diff --git a/scripts/internal/task_layout.py b/scripts/internal/task_layout.py index 10e55be..d125423 100644 --- a/scripts/internal/task_layout.py +++ b/scripts/internal/task_layout.py @@ -13,6 +13,12 @@ "working_set": "working-set.json", "progress": "runtime/progress.json", "code_intelligence_runtime": "runtime/code-intelligence", + "code_intelligence_proxy_bundle": ( + "runtime/code-intelligence/{record_name}/{query_id}.json" + ), + "code_intelligence_proxy_response": ( + "runtime/code-intelligence/{record_name}/{query_id}.response.txt" + ), "code_intelligence_revision": "code-intelligence/r{revision:03d}", "code_intelligence_record": "code-intelligence/r{revision:03d}/{record_name}.json", "work_item": "revisions/work-item-r{revision:03d}.json", @@ -48,6 +54,7 @@ def task_relative_path( reviewer: int = 1, exploration_id: str = "EXP-0001", record_name: str = "planning", + query_id: str = "CIQ-001", ) -> Path: pattern = TASK_PATH_PATTERNS[artifact] reviewer_suffix = "" if reviewer == 1 else f"-{reviewer}" @@ -58,6 +65,7 @@ def task_relative_path( reviewer_suffix=reviewer_suffix, exploration_id=exploration_id, record_name=record_name, + query_id=query_id, ) ) diff --git a/scripts/record_code_intelligence.py b/scripts/record_code_intelligence.py index 28afb60..a259cc4 100644 --- a/scripts/record_code_intelligence.py +++ b/scripts/record_code_intelligence.py @@ -8,7 +8,7 @@ from pathlib import Path from internal.code_intelligence_protocol import ( - record, + record_proxy_bundle, select_provider, ) from internal.polaris_core import ( @@ -23,7 +23,8 @@ def main() -> int: parser = argparse.ArgumentParser() parser.add_argument("task_id", nargs="?") parser.add_argument("--repo", type=Path, default=Path.cwd()) - parser.add_argument("--input", type=Path) + parser.add_argument("--bundle", type=Path) + parser.add_argument("--annotations", type=Path) parser.add_argument("--available-tool", action="append", default=[]) parser.add_argument("--available-executable", action="append", default=[]) parser.add_argument("--select-provider", action="store_true") @@ -40,10 +41,22 @@ def execute() -> dict[str, object]: available_executables=args.available_executable, ) return {"selected": selected, "available": selected is not None} - if args.task_id is None or args.input is None: - raise InputFailure("recording requires task_id and --input") - input_path = args.input if args.input.is_absolute() else repo / args.input - return record(repo, args.task_id, read_json(input_path)) + if args.task_id is None or args.bundle is None or args.annotations is None: + raise InputFailure( + "recording requires task_id, --bundle, and --annotations" + ) + bundle_path = args.bundle if args.bundle.is_absolute() else repo / args.bundle + annotations_path = ( + args.annotations + if args.annotations.is_absolute() + else repo / args.annotations + ) + return record_proxy_bundle( + repo, + args.task_id, + bundle_path, + read_json(annotations_path), + ) return run_main(execute, args.json) diff --git a/scripts/validate_project.py b/scripts/validate_project.py index 4f7820f..3ddf08b 100644 --- a/scripts/validate_project.py +++ b/scripts/validate_project.py @@ -14,6 +14,7 @@ load_host_adapters, ) from internal.install_manifest import validate_install_manifest +from internal.project_mcp_registration import project_mcp_target, validate_project_mcp from internal.code_intelligence_protocol import validate_static_configuration from internal.migration_protocol import validate_completed_migrations from internal.polaris_core import RuleFailure, protocol_root, read_json, run_main, validate_json_file @@ -126,6 +127,14 @@ def validate(repo: Path) -> dict[str, object]: f"{adapter['display_name']} adapter file is not {ownership}: " f"{relative}" ) + registration = project_mcp_target(repo, adapter) + relative = registration.relative_to(repo).as_posix() + if relative not in preserved_paths: + raise RuleFailure( + f"{adapter['display_name']} project MCP configuration is not preserved: " + f"{relative}" + ) + validate_project_mcp(repo, adapter) listed = set(project["active_tasks"]) validate_task_locations(repo, listed) diff --git a/scripts/vendor_project.py b/scripts/vendor_project.py index 185d67a..8387078 100644 --- a/scripts/vendor_project.py +++ b/scripts/vendor_project.py @@ -28,6 +28,7 @@ validate_install_manifest, write_install_manifest, ) +from internal.project_mcp_registration import merge_project_mcp, project_mcp_target from internal.polaris_core import ( InputFailure, RuleFailure, @@ -69,6 +70,7 @@ def _polaris_destinations( for item in adapter["files"] if item["overwrite"] ) + destinations.append(project_mcp_target(target, adapter)) return destinations @@ -280,6 +282,10 @@ def _stage_install( else: preserved_paths.append(destination) + registration = project_mcp_target(stage, adapter) + write_text_atomic(registration, merge_project_mcp(target, adapter)) + preserved_paths.append(registration) + tools_target = confined_target( stage, stage / "tools" / "polaris", "staged vendored protocol target" ) @@ -407,8 +413,13 @@ def vendor( previous_manifest = ( read_install_manifest(target, source) if manifest_path.is_file() else None ) + registration_targets = { + project_mcp_target(target, adapter) for adapter in adapters + } if not force and any( - path.exists() for path in _polaris_destinations(target, adapters, skills) + path.exists() + for path in _polaris_destinations(target, adapters, skills) + if path not in registration_targets ): raise InputFailure("vendored Polaris files already exist; use --force to update") if force and previous_manifest is not None and not discard_managed_changes: diff --git a/skills/adversarial-review/SKILL.md b/skills/adversarial-review/SKILL.md index 3d8a611..421befd 100644 --- a/skills/adversarial-review/SKILL.md +++ b/skills/adversarial-review/SKILL.md @@ -10,16 +10,16 @@ For R1/R2, run only in the fresh Reviewer context defined by the active host ada 1. Run `recover_task.py --repo . --json` and require state `REVIEWING`. 2. Require an explicit Reviewer slot and registered `review_handoff` path from the dispatcher. Load only that handoff and its package paths. Do not use implementer explanations, prior chat, another Reviewer's artifact, or an expected verdict. 3. Verify handoff hashes, task revision, Review attempt, exact subject commits/diff hash, and the required isolation mode. -4. Assign a reviewer session ID distinct from the implementer for R1/R2. Attest truthfully to isolation and chat-history inheritance; do not fabricate independence. At the Review boundary invoke `{{skill:code-intelligence}}` for optional independent registered-subject relationships, first using `status` or `sync-if-needed`. Use CodeGraph only with an existing `.codegraph/` directory: prefer `codegraph_explore`, with `codegraph explore` as the non-MCP fallback. A bounded `codegraph sync` is non-blocking. Do not reuse Implementer query conclusions. Missing or failing Provider output immediately falls back to the original frozen-package and source review. +4. Assign a reviewer session ID distinct from the implementer for R1/R2. Attest truthfully to isolation and chat-history inheritance; do not fabricate independence. At the Review boundary invoke `{{skill:code-intelligence}}` for optional independent registered-subject relationships. With enabled policy and an existing `.codegraph/`, make a fresh `polaris_codegraph_explore` call and read its freshness envelope before graph content. Never inherit or reuse an Implementer envelope, bundle, or graph conclusion; ignore any pre-existing Review-stage bundle and use the bundle created by this Reviewer call. Missing, stale, unknown, or failing output follows the required source/Git fallback and never blocks the frozen-package/source review. 5. Check specification compliance first: correct problem, scope, exclusions, constraints, and every acceptance criterion. 6. Check engineering quality second: correctness, failure paths, lifetime, concurrency, security, performance, compatibility, maintainability, test gaps, and counterexamples. 7. Preserve every prior Finding ID in a follow-up Review. Read the registered author response, recheck the entire new patch, and record a concrete `reviewer_resolution` for each carried Finding. 8. Give new Findings monotonic IDs and mark critical/high, acceptance failures, and scope violations as blocking. -9. If this stage actually performed a Provider status, sync, or explore operation, finalize an immutable v2 Review Code Intelligence record. If no Provider operation ran, omit the Code Intelligence record and its optional artifact reference. Resolve the output with `task_layout.review_path` from the handoff revision, attempt, and Reviewer slot. Write a new immutable Review JSON bound to the handoff, slot, session attestation, and optional record. Never assemble the path independently or overwrite an existing artifact. Reject while any blocking Finding remains open; Code Intelligence cannot determine the verdict. +9. If this stage ran the proxy, create annotations and run `record_code_intelligence.py --repo . --bundle --annotations ` to project the immutable v3 Review record. If no proxy operation ran, omit the Code Intelligence record and optional artifact reference. Resolve the output with `task_layout.review_path` from the handoff revision, attempt, and Reviewer slot. Write a new immutable Review JSON bound to the handoff, slot, session attestation, and optional record. Never assemble the path independently or overwrite an existing artifact. Reject while any blocking Finding remains open; Code Intelligence cannot determine the verdict. 10. Return the verdict and exact Review path to the dispatching `{{skill:engineering-task}}` context. Do not run `ACCEPT_REVIEW` or `REJECT_REVIEW`; the dispatcher validates and registers all required Review artifacts before applying the graph transition. Never modify implementation code or start another Reviewer task during Review. Return a concise structured result to the dispatcher with verdict, Review attempt, Reviewer slot, reviewer session ID, subject commits/diff hash, every Finding ID and status, and the immutable Review path. Do not emit a Polaris checkpoint marker from the child task. The dispatching context emits `[POLARIS:REVIEW_ACCEPTED]` or `[POLARIS:REVIEW_REJECTED]` with the nine fixed fields only after the corresponding transition succeeds. If isolation or handoff validation prevents review, do not write a Review; report the exact required fresh-session or handoff action to the dispatcher. Only the Reviewer context may write `ACCEPT`. -CodeGraph fallback contract: never run `codegraph init` or manage the Provider. Save and classify each response; when `RESPONSE_BANNER` is present, persist its successful explore response hash as `freshness.response_sha256`. For `PARTIAL_STALE`, if a named path is a current confined regular file, directly read it and record `READ_SOURCE` with its current SHA-256; if a safe path is missing/deleted, inspect the registered subject Git diff and record `INSPECT_GIT_DIFF` with null observed SHA-256 and bound base/head/diff evidence; for unsafe paths, record `NOT_VERIFIED` and use source search. For `INDEX_STALE` or `NOT_VERIFIED`, use source search and Git evidence, then stop graph calls for this stage. Each `SEARCH_SOURCE` fallback records `result_paths`: zero or at most 100 unique POSIX paths, each a current confined regular file with its current SHA-256; non-`SEARCH_SOURCE` fallbacks use empty `result_paths`. Graph evidence cannot determine the Review verdict. +Proxy evidence contract: `CURRENT` is `NON_AUTHORITATIVE_CONTEXT`; `STALE` and `UNKNOWN` are `NAVIGATION_ONLY`. `NAVIGATION_ONLY` never substantiates an edit or conclusion, even after fallback; only completed current source/Git fallback evidence does, and index-wide uncertainty affects the entire graph response. Use no separate status/sync MCP tool and do not retry, poll, wait, or run another query after fallback is required. A raw `codegraph_explore` or `codegraph explore` result is out-of-band and cannot back `CURRENT` Polaris evidence. Never run `codegraph init` or manage the Provider. Graph evidence cannot determine the Review verdict. diff --git a/skills/architecture-planning/SKILL.md b/skills/architecture-planning/SKILL.md index bbb2bee..e89fe2f 100644 --- a/skills/architecture-planning/SKILL.md +++ b/skills/architecture-planning/SKILL.md @@ -6,7 +6,7 @@ description: Internal Polaris stage for an explicitly started `{{skill:engineeri # Architecture Planning 1. Read the frozen Work Item and project rules. -2. Refresh `working-set.json` with `build_working_set.py`. At the Planning boundary use `{{skill:code-intelligence}}` only when optional frozen-task relationship discovery is useful and a Provider operation can run. Use CodeGraph only when `.codegraph/` already exists: prefer `codegraph_explore`, with `codegraph explore` as the non-MCP fallback. A bounded `codegraph sync` is non-blocking. When a status, sync, or explore operation runs, write a compact v2 Planning record; otherwise omit the Code Intelligence record. Confirm every returned path from repository source before adding it with the query ID as `discovered_from`; provider failure immediately falls back to the original repository search path and never blocks Planning. Include `.polaris/code-intelligence.json` and the finalized Planning record in the Working Set only when they exist. Record every entry as section, path, reason, and discovery source; add explicit entries only for concrete dependencies. Do not parse or create a duplicate Markdown Working Set. +2. Refresh `working-set.json` with `build_working_set.py`. At the Planning boundary use `{{skill:code-intelligence}}` only when frozen-task relationship discovery is useful. If policy is enabled and `.codegraph/` exists, call `polaris_codegraph_explore` with stage `PLANNING` and a bounded frozen-scope query. Read its freshness envelope before graph content. Confirm every safe current returned path from repository source before adding it with the query ID as `discovered_from`; confirm a safe missing/deleted path through the registered subject Git diff. Proxy failure immediately uses the original source/Git fallback and never blocks Planning. Include `.polaris/code-intelligence.json` and the projected Planning record in the Working Set only when they exist. Record every entry as section, path, reason, and discovery source; add explicit entries only for concrete dependencies. Do not parse or create a duplicate Markdown Working Set. 3. Investigate only paths justified by the task or a discovered dependency. Provider observations cannot expand frozen scope. 4. Write `PLAN.md` as a delta from `base_commit`, including alternatives, risks, affected invariants, and expected documentation changes. Keep rationale in Markdown; do not use it as decision authority. 5. Map every acceptance criterion to a planned validation command or Human check. Code Intelligence observations are not acceptance evidence. @@ -21,4 +21,4 @@ After the transition succeeds, reload state and emit `[POLARIS:PLAN_READY]` with Do not modify the frozen Work Item or start implementation from this stage. -CodeGraph fallback contract: never run `codegraph init` or manage the Provider. Save and classify each response; when `RESPONSE_BANNER` is present, persist its successful explore response hash as `freshness.response_sha256`. For `PARTIAL_STALE`, if a named path is a current confined regular file, directly read it and record `READ_SOURCE` with its current SHA-256; if a safe path is missing/deleted, inspect the registered subject Git diff and record `INSPECT_GIT_DIFF` with null observed SHA-256 and bound base/head/diff evidence; for unsafe paths, record `NOT_VERIFIED` and use source search. For `INDEX_STALE` or `NOT_VERIFIED`, use source search and Git evidence, then stop graph calls for this stage. Each `SEARCH_SOURCE` fallback records `result_paths`: zero or at most 100 unique POSIX paths, each a current confined regular file with its current SHA-256; non-`SEARCH_SOURCE` fallbacks use empty `result_paths`. Graph evidence cannot expand scope or act as a gate. +Proxy evidence contract: `CURRENT` is `NON_AUTHORITATIVE_CONTEXT`; `STALE` and `UNKNOWN` are `NAVIGATION_ONLY`. `NAVIGATION_ONLY` never substantiates an edit or conclusion, even after fallback; only completed current source/Git fallback evidence does, and index-wide uncertainty affects the entire graph response. Use no separate status/sync MCP tool and do not retry, poll, wait, or run another query after fallback is required. A raw `codegraph_explore` or `codegraph explore` result is out-of-band and cannot back `CURRENT` Polaris evidence. After a proxy operation, write annotations and run `record_code_intelligence.py --repo . --bundle --annotations ` to project v3; without a proxy operation, omit the Code Intelligence record. Never run `codegraph init` or manage the Provider. diff --git a/skills/code-intelligence/SKILL.md b/skills/code-intelligence/SKILL.md index cbe7a39..e584d19 100644 --- a/skills/code-intelligence/SKILL.md +++ b/skills/code-intelligence/SKILL.md @@ -5,21 +5,23 @@ description: Internal optional Polaris stage support for bounded CodeGraph relat # Code Intelligence -Treat Code Intelligence as read-only, best-effort evidence. Source, Git, builds, tests, and frozen Polaris artifacts remain authority. +CodeGraph is optional navigation context. Source, Git, builds, tests, frozen artifacts, Review, Validation, and Human decisions remain authority. -1. Load `.polaris/code-intelligence.json` when present and project rules. If policy disables Code Intelligence or the repository has no `.codegraph/` directory, use the stage's source path and omit the Code Intelligence record because no Provider operation ran. Stop CodeGraph calls for this project for the session and tell the user they may choose to initialize it; never run `codegraph init`. -2. At the calling stage's declared boundary, run `code_intelligence_runtime.py status` or `sync-if-needed`. The latter may run one bounded `codegraph sync` only when status reports pending changes; it never loops, waits for a watcher, or treats a successful command as a gate. -3. For an allowed frozen-scope relationship query, use only `codegraph_explore` when MCP exposes it. If MCP is unavailable and the executable is available, use `codegraph explore` as the non-MCP fallback. Do not select retired narrow operations. Bound the query to the Work Item, Working Set, registered subject, or a confirmed dependency; graph output cannot expand frozen scope, authorize change, satisfy acceptance, or determine a Review verdict. -4. Save each raw explore response only below the task's ignored `runtime/code-intelligence/` directory, then run `code_intelligence_runtime.py classify-response` for it. When `RESPONSE_BANNER` is a freshness basis, persist that successful explore response hash as `freshness.response_sha256`; final records contain the response hash and finite summary, never the response itself. -5. On `PARTIAL_STALE`, process every named path by its current safe state. If it is a current confined regular file, directly read it and record `READ_SOURCE` with its current SHA-256. If a safe path is missing/deleted, inspect the registered subject Git diff and record `INSPECT_GIT_DIFF` with null observed SHA-256 and bound base/head/diff evidence. For unsafe paths, record `NOT_VERIFIED` and use source search. The remaining graph response may still be navigation evidence, but never a conclusion about a stale path. -6. On `INDEX_STALE` or `NOT_VERIFIED`, use repository source search and Git evidence, record the `SEARCH_SOURCE` fallback, and stop repeated graph calls for that stage. Every `SEARCH_SOURCE` fallback records `result_paths`: zero or at most 100 unique POSIX paths, each a current confined regular file with its current SHA-256; non-`SEARCH_SOURCE` fallbacks use empty `result_paths`. On malformed, missing, or unavailable Provider output, continue the same source fallback without blocking the stage. -7. Never initialize, install, start, authenticate, or reconfigure CodeGraph. Do not manage its watcher, daemon, lock, or host MCP settings. -8. Finalize an immutable v2 Code Intelligence record only after an actual Provider status, sync, or explore operation, including its real freshness, stale points, and source fallbacks. If no operation ran, omit the Code Intelligence record. Code Intelligence is never a workflow gate. +1. Load `.polaris/code-intelligence.json` and project rules. If policy disables Code Intelligence or the repository root lacks `.codegraph/`, skip the proxy, use source/Git, and omit the Code Intelligence record because no proxy operation ran. Never run `codegraph init`. +2. For Polaris graph evidence call only `polaris_codegraph_explore`, using the active task ID, legal stage, next `CIQ-NNN`, bounded purpose/query, and the stage's declared `sync_if_needed` value. The project registration fixes the repository root. Use no separate status/sync MCP tool and do not retry, poll, wait, or run another query after an envelope requires the stage fallback. +3. Read the `freshness envelope` before any graph content: + - `CURRENT` with `usage: NON_AUTHORITATIVE_CONTEXT` permits the graph only as non-authoritative context. + - `STALE` or `UNKNOWN` with `usage: NAVIGATION_ONLY` requires every named source/Git fallback. `NAVIGATION_ONLY` never substantiates an edit or conclusion, even after fallback; only the resulting current source/Git evidence does. Index-wide uncertainty affects the entire graph response. `UNKNOWN` is never current. + - `UNAVAILABLE` with `usage: NO_GRAPH` means use source/Git and do not expect graph content. +4. Complete fallbacks exactly. For a safe named current regular file, read it and record `READ_SOURCE` with its current SHA-256. For a safe missing/deleted path, inspect the registered subject diff and record `INSPECT_GIT_DIFF` with null observed SHA-256 and bound base/head/diff hashes. For an unsafe path or index-wide stale/unknown result, perform an actual bounded repository search and record `SEARCH_SOURCE` with zero to 100 unique confined POSIX `result_paths`, each a current regular file and current SHA-256; an empty result is valid only when that search found no current file. In `STALE`/`UNKNOWN`, annotate a symbol only when its current path is covered by `READ_SOURCE` or a hashed `SEARCH_SOURCE` result. +5. A raw `codegraph_explore` MCP call or `codegraph explore` shell command remains user-accessible out-of-band, but its output is always unverified for Polaris and cannot back `CURRENT` Polaris evidence. Never project raw Provider output into a Polaris record. +6. If the proxy ran, use the envelope's `evidence_bundle` path, write an annotations JSON containing only `summary`, confirmed `symbols`, and completed `source_fallbacks`, then run `record_code_intelligence.py --repo . --bundle --annotations `. This projects an immutable v3 record; do not hand-author records. If the proxy did not run, omit the Code Intelligence record and optional artifact reference. +7. Never install, initialize, start, authenticate, configure, reconfigure, or manage CodeGraph, its watcher, daemon, lock, raw MCP registration, or project index. Proxy failure and every non-current state are non-gating. Stage policy: -- Planning: at the Planning boundary, request only frozen-task relationship discovery needed to justify Working Set entries; confirm every returned path in repository source and record its query ID as `discovered_from`. -- Implementation: before editing, request only handoff-scoped edit relationships. Query again mid-stage only when a later declared implementation step depends on relationships changed by the current subject. -- Documentation Sync: run `sync-if-needed` once only when the final subject changed supported source files and the Provider is available; otherwise omit the Code Intelligence record. -- Review: independently request only registered-subject impact relationships. Do not reuse Implementer query conclusions. -- Validation: do not invoke this Skill; use builds, tests, static checks, and Human Checks as the acceptance evidence. +- Planning: query only frozen-task relationships needed to justify Working Set entries. Confirm safe current returned paths in current source before recording the query ID as `discovered_from`; confirm a safe missing/deleted path through the registered subject Git diff instead. +- Implementation: make a bounded handoff-scoped call before editing when useful. Any conclusion needed after edits requires a fresh `polaris_codegraph_explore` call; never reuse the entry freshness envelope. If an earlier non-current envelope ended graph use for the stage, use source/Git only rather than making that post-edit call. +- Documentation Sync: only when supported source changed, make one query over changed source paths and documented symbols with `sync_if_needed: true`; there is no separate status/sync MCP tool. +- Review: independently query only registered-subject impact relationships. Never inherit or reuse the Implementer's envelope, bundle, or conclusions. +- Validation: do not invoke this Skill. Validation remains graph-free. diff --git a/skills/documentation-sync/SKILL.md b/skills/documentation-sync/SKILL.md index b906cc5..c9793b4 100644 --- a/skills/documentation-sync/SKILL.md +++ b/skills/documentation-sync/SKILL.md @@ -12,7 +12,7 @@ description: Internal Polaris worker stage for an explicitly started `{{skill:en 5. Record failed attempts with `record_exploration.py`. Keep task-only conclusions in the task; promote reusable, evidence-backed conclusions to `.polaris/explorations/` with the same script. 6. Leave no unresolved `STALE` entry. 7. Create the final subject checkpoint and recompute the subject diff hash. -8. When the final subject includes supported source changes and the Provider is available, invoke `{{skill:code-intelligence}}` once at the Documentation Sync boundary with `sync-if-needed`. Use CodeGraph only with an existing `.codegraph/` directory: prefer `codegraph_explore`, with `codegraph explore` as the non-MCP fallback. A bounded `codegraph sync` is non-blocking. If a Provider operation ran, reference its immutable v2 record from the Knowledge Delta; otherwise omit the Code Intelligence record and its optional artifact reference. Never claim commit-exact freshness. +8. When the final subject includes supported source changes, policy is enabled, and `.codegraph/` exists, invoke `{{skill:code-intelligence}}` once at the Documentation Sync boundary. Call `polaris_codegraph_explore` with stage `DOCUMENTATION_SYNC`, a query limited to changed source paths and documented symbols, and `sync_if_needed: true`; there is no separate status/sync MCP tool. Read the freshness envelope first and finish every required source/Git fallback. Then create annotations and run `record_code_intelligence.py --repo . --bundle --annotations ` to project the immutable v3 record and reference it from the Knowledge Delta. Otherwise omit the Code Intelligence record and optional artifact reference. 9. Refresh the Working Set if a promoted exploration, documentation change, or confirmed Code Intelligence dependency alters the next stage's justified inputs. 10. Run `check_docs.py` with the final subject base/head. When live telemetry exists, append its result with `ADD_CHECK`, then use `SET_PHASE` to enter `COMPLETED` with no blocker. Return the Knowledge Delta path, final subject base/head, diff hash, changed documentation, promoted explorations, Code Intelligence refresh status, and check result. @@ -20,4 +20,4 @@ Do not run workflow transitions or emit a Polaris checkpoint marker. The main `{ Do not edit Review, Validation, Result, event, or state artifacts directly. -CodeGraph fallback contract: never run `codegraph init` or manage the Provider. Save and classify each response; when `RESPONSE_BANNER` is present, persist its successful explore response hash as `freshness.response_sha256`. For `PARTIAL_STALE`, if a named path is a current confined regular file, directly read it and record `READ_SOURCE` with its current SHA-256; if a safe path is missing/deleted, inspect the registered subject Git diff and record `INSPECT_GIT_DIFF` with null observed SHA-256 and bound base/head/diff evidence; for unsafe paths, record `NOT_VERIFIED` and use source search. For `INDEX_STALE` or `NOT_VERIFIED`, use source search and Git evidence, then stop graph calls for this stage. Each `SEARCH_SOURCE` fallback records `result_paths`: zero or at most 100 unique POSIX paths, each a current confined regular file with its current SHA-256; non-`SEARCH_SOURCE` fallbacks use empty `result_paths`. Graph evidence never gates documentation checks or state changes. +Proxy evidence contract: `CURRENT` is `NON_AUTHORITATIVE_CONTEXT`; `STALE` and `UNKNOWN` are `NAVIGATION_ONLY`. `NAVIGATION_ONLY` never substantiates an edit or conclusion, even after fallback; only completed current source/Git fallback evidence does, and index-wide uncertainty affects the entire graph response. Use no separate status/sync MCP tool and do not retry, poll, wait, or run another query after fallback is required. A raw `codegraph_explore` or `codegraph explore` result is out-of-band and cannot back `CURRENT` Polaris evidence. Never run `codegraph init` or manage the Provider. Graph evidence never gates documentation checks or state changes. diff --git a/skills/implementation/SKILL.md b/skills/implementation/SKILL.md index 308b586..95240af 100644 --- a/skills/implementation/SKILL.md +++ b/skills/implementation/SKILL.md @@ -6,7 +6,7 @@ description: Internal Polaris worker stage for an explicitly started `{{skill:en # Implementation 1. Require the task ID and registered Implementation handoff path returned by the main task. Load only that handoff and its package as task context; read `state.json` only to verify registration. Use paths carried by the handoff or resolved by `task_layout.py`; never reconstruct them from prose. Do not read the main conversation or infer unstated requirements. -2. Confirm state is `IMPLEMENTING`, the handoff hash matches `state.json`, and `artifact_attempt`, revision, base commit, output path, and progress paths are current. At the Implementation boundary invoke `{{skill:code-intelligence}}` before editing for optional handoff-scoped edit relationships, first using `status` or `sync-if-needed`. Use CodeGraph only with an existing `.codegraph/` directory: prefer `codegraph_explore`, with `codegraph explore` as the non-MCP fallback. A bounded `codegraph sync` is non-blocking. Missing or failing Provider output immediately falls back to direct source reading. Query again during Implementation only when a later declared step depends on relationships from newly changed code. +2. Confirm state is `IMPLEMENTING`, the handoff hash matches `state.json`, and `artifact_attempt`, revision, base commit, output path, and progress paths are current. At the Implementation boundary invoke `{{skill:code-intelligence}}` before editing when handoff-scoped relationships are useful. With enabled policy and an existing `.codegraph/`, call only `polaris_codegraph_explore` and read its freshness envelope before graph content. Missing, failing, stale, or unknown output immediately follows the required source/Git fallback. A conclusion about relationships changed by current edits requires a fresh proxy call after edits and never reuses the entry envelope, except that an earlier non-current envelope ends graph use for this stage and requires source/Git only. 3. Generate one stable Implementer session ID for this conversation. Before changing code, create a non-empty ordered `implementation_steps` list. Every step receives the next `STEP-NNN` ID and must reference one or more acceptance IDs from the frozen Work Item. If an ignored live snapshot was initialized, mirror the list through its `DEFINE_STEPS` event. 4. Execute steps linearly. When live telemetry exists, use `START_STEP`, then `COMPLETE_STEP`, `BLOCK_STEP`, or `RESUME_STEP`; use `SKIP_STEP` only with an explicit reason. Existing step identity, title, order, and acceptance bindings are immutable. Newly discovered work may only be added at the end, using `APPEND_STEP` when telemetry exists. Never edit `progress.json` directly or create it as a durable prerequisite. 5. Change only declared subject paths and protect unrelated user changes. Work in small build/test/fix loops. @@ -14,9 +14,9 @@ description: Internal Polaris worker stage for an explicitly started `{{skill:en 7. Record Plan deviations and reasons. After Review rejection, load the handoff's prior Review, answer every open Finding once in an immutable Review Response, and bind it to the new subject. 8. Run planned local checks and record reproducible evidence in the immutable Implementation artifact. When live telemetry exists, append reproducible evidence to the snapshot with `ADD_CHECK`; without a snapshot, do not initialize one merely to report checks. Never report a made-up percentage; derive completed, current, and remaining work from the ordered steps. 9. Complete or explicitly skip every step, then create a subject checkpoint commit containing scoped code, tests, build configuration, and relevant project docs only. -10. If this stage actually performed a Provider status, sync, or explore operation, finalize an immutable v2 Implementation Code Intelligence record and reference it. If no Provider operation ran, omit the Code Intelligence record and its optional artifact reference. Write the immutable Implementation JSON at the handoff's `output_path`, bind the handoff, subject, session, deviations, and checks, and copy the exact terminal `id`, `status`, and `result` projection into `step_results`. Code Intelligence evidence is never a gate. +10. If this stage ran the proxy, create its annotations and run `record_code_intelligence.py --repo . --bundle --annotations ` to project the immutable v3 Implementation record and reference it. If no proxy operation ran, omit the Code Intelligence record and its optional artifact reference. Write the immutable Implementation JSON at the handoff's `output_path`, bind the handoff, subject, session, deviations, and checks, and copy the exact terminal `id`, `status`, and `result` projection into `step_results`. Code Intelligence evidence is never a gate. 11. After every step is `COMPLETED` or `SKIPPED`, use `SET_PHASE` to enter `CHECKPOINTING` only when live telemetry exists. Return the artifact path, session ID, subject base/head, diff hash, step results, checks, deviations, Review Response path when present, and remaining Documentation Sync work. Do not run workflow transitions, Review, Validation, or task closure. Do not emit a Polaris checkpoint marker; the main `{{skill:engineering-task}}` validates the artifact and continues this same task for `{{skill:documentation-sync}}` while authority remains `IMPLEMENTING`. -CodeGraph fallback contract: never run `codegraph init` or manage the Provider. Save and classify each response; when `RESPONSE_BANNER` is present, persist its successful explore response hash as `freshness.response_sha256`. For `PARTIAL_STALE`, if a named path is a current confined regular file, directly read it and record `READ_SOURCE` with its current SHA-256; if a safe path is missing/deleted, inspect the registered subject Git diff and record `INSPECT_GIT_DIFF` with null observed SHA-256 and bound base/head/diff evidence; for unsafe paths, record `NOT_VERIFIED` and use source search. For `INDEX_STALE` or `NOT_VERIFIED`, use source search and Git evidence, then stop graph calls for this stage. Each `SEARCH_SOURCE` fallback records `result_paths`: zero or at most 100 unique POSIX paths, each a current confined regular file with its current SHA-256; non-`SEARCH_SOURCE` fallbacks use empty `result_paths`. Graph evidence never gates implementation. +Proxy evidence contract: `CURRENT` is `NON_AUTHORITATIVE_CONTEXT`; `STALE` and `UNKNOWN` are `NAVIGATION_ONLY`. `NAVIGATION_ONLY` never substantiates an edit or conclusion, even after fallback; only completed current source/Git fallback evidence does, and index-wide uncertainty affects the entire graph response. Use no separate status/sync MCP tool and do not retry, poll, wait, or run another query after fallback is required. A raw `codegraph_explore` or `codegraph explore` result is out-of-band and cannot back `CURRENT` Polaris evidence. Never run `codegraph init` or manage the Provider. diff --git a/templates/AGENTS.md b/templates/AGENTS.md index a4d44b8..30726bc 100644 --- a/templates/AGENTS.md +++ b/templates/AGENTS.md @@ -15,8 +15,9 @@ ## Optional CodeGraph rules -- Use CodeGraph only when the repository root already contains `.codegraph/`. When it is absent, stop CodeGraph calls for this session and use repository source and Git; a user may choose to initialize CodeGraph, but agents must never run `codegraph init`. -- Prefer MCP `codegraph_explore`; when MCP is unavailable, use `codegraph explore` as the CLI fallback. A bounded `codegraph sync` may run only through the Polaris stage boundary procedure and never gates a task. -- Save and classify every graph response in task runtime; when `RESPONSE_BANNER` is present, persist its successful explore response hash as `freshness.response_sha256`. For `PARTIAL_STALE`, if a named path is a current confined regular file, directly read it and record `READ_SOURCE` with its current SHA-256; if a safe path is missing/deleted, inspect the registered subject Git diff and record `INSPECT_GIT_DIFF` with null observed SHA-256 and bound base/head/diff evidence; for unsafe paths, record `NOT_VERIFIED` and use source search. For `INDEX_STALE` or `NOT_VERIFIED`, treat graph output only as a lead, use source search and Git evidence, and stop repeated graph calls for that stage. Each `SEARCH_SOURCE` fallback records `result_paths`: zero or at most 100 unique POSIX paths, each a current confined regular file with its current SHA-256; non-`SEARCH_SOURCE` fallbacks use empty `result_paths`. -- Never install, start, authenticate, reconfigure, or manage CodeGraph, its watcher, daemon, lock, or MCP settings. CodeGraph cannot expand frozen scope or replace source, Git, builds, tests, Review, Validation, or Human gates. +- Use CodeGraph only when project policy permits it and the repository root already contains `.codegraph/`. Otherwise skip the proxy, use source/Git, and omit the Code Intelligence record; agents never run `codegraph init`. +- For Polaris evidence call only `polaris_codegraph_explore` and read its freshness envelope before graph content. `CURRENT` is `NON_AUTHORITATIVE_CONTEXT`; `STALE` and `UNKNOWN` are `NAVIGATION_ONLY`. `NAVIGATION_ONLY` never substantiates an edit or conclusion, even after fallback; only completed current source/Git fallback evidence does, and index-wide uncertainty affects the entire graph response. Use no separate status/sync MCP tool and do not retry, poll, wait, or run another query after fallback is required. `UNAVAILABLE` means no graph. +- Complete fallbacks exactly: a safe current regular file uses `READ_SOURCE` with current SHA-256; a safe missing/deleted path uses `INSPECT_GIT_DIFF` with null observed SHA-256 and bound base/head/diff hashes; unsafe or index-wide stale/unknown results use `SEARCH_SOURCE` with finite confined POSIX result paths and current hashes. +- A raw `codegraph_explore` or `codegraph explore` result is out-of-band and cannot back `CURRENT` Polaris evidence. If the proxy ran, write annotations and run `record_code_intelligence.py --repo . --bundle --annotations ` to project v3; do not hand-author a record. +- Never install, initialize, start, authenticate, configure, reconfigure, or manage CodeGraph, its watcher, daemon, lock, raw MCP registration, or index. CodeGraph cannot expand frozen scope or replace source, Git, builds, tests, Review, Validation, or Human gates. - Preserve any installer-managed marker block exactly as owned by that installer; Polaris does not add, edit, or remove installer marker fences. diff --git a/templates/project.json b/templates/project.json index 4bfc7aa..e914021 100644 --- a/templates/project.json +++ b/templates/project.json @@ -1,6 +1,6 @@ { "project_id": "PROJECT_ID", - "polaris_version": "0.1.20", + "polaris_version": "0.1.21", "workflow_version": "0.1.3", "active_tasks": [] } diff --git a/templates/task-sources/code-intelligence-record.json b/templates/task-sources/code-intelligence-record.json index 9e17567..419a9c3 100644 --- a/templates/task-sources/code-intelligence-record.json +++ b/templates/task-sources/code-intelligence-record.json @@ -1,27 +1,83 @@ { - "record_version": 2, + "record_version": 3, "task_id": "TASK-0001", "work_item_revision": 1, "stage": "PLANNING", "artifact_attempt": null, "reviewer_slot": null, - "provider": null, + "provider": { + "id": "codegraph", + "descriptor_version": 2 + }, + "repository": { + "project_id": "PROJECT_ID", + "root_sha256": "0000000000000000000000000000000000000000000000000000000000000000" + }, "target": { "base_commit": "0000000000000000000000000000000000000000", "head_commit": null, "diff_hash": null }, "status": "UNAVAILABLE", - "queries": [], - "status_check": null, - "sync": null, - "freshness": { + "proxy": { + "server_id": "polaris-codegraph", + "tool": "polaris_codegraph_explore", + "evidence_bundle_sha256": "0000000000000000000000000000000000000000000000000000000000000000" + }, + "query": { + "id": "CIQ-001", + "purpose": "Record unavailable CodeGraph context", + "text": "No CodeGraph query was executed", "status": "UNAVAILABLE", - "checked_at": "1970-01-01T00:00:00Z", - "basis": ["NONE"], + "summary": "", + "symbols": [], "response_sha256": null, - "stale_points": [] + "error": "CodeGraph provider unavailable" + }, + "query_window": { + "pre_status": { + "status": "UNAVAILABLE", + "checked_at": "1970-01-01T00:00:00Z", + "basis": [ + "NONE" + ], + "stale_points": [], + "status_response_sha256": null, + "error": "CodeGraph provider unavailable", + "needs_sync": false, + "pending_changes": null + }, + "sync": null, + "post_sync_status": null, + "response_classification": null, + "post_query_status": null + }, + "delivery": { + "state": "UNAVAILABLE", + "record_status": "UNAVAILABLE", + "reason": "PROVIDER_UNAVAILABLE", + "checked_at": "1970-01-01T00:00:00Z", + "usage": "NO_GRAPH", + "required_fallback": "SEARCH_SOURCE", + "stale_points": [], + "pending_changes": { + "added": 0, + "modified": 0, + "removed": 0 + }, + "error": "CodeGraph provider unavailable" }, - "source_fallbacks": [], + "source_fallbacks": [ + { + "action": "SEARCH_SOURCE", + "path": null, + "observed_sha256": null, + "base_commit": null, + "head_commit": null, + "diff_hash": null, + "purpose": "Use repository source because CodeGraph is unavailable", + "result_paths": [] + } + ], "recorded_at": "1970-01-01T00:00:00Z" } diff --git a/templates/task-sources/state.json b/templates/task-sources/state.json index a623be4..7f63d16 100644 --- a/templates/task-sources/state.json +++ b/templates/task-sources/state.json @@ -1,6 +1,6 @@ { "task_id": "TASK-0001", - "polaris_version": "0.1.20", + "polaris_version": "0.1.21", "workflow_version": "0.1.3", "current_revision": 1, "status": "DRAFT", diff --git a/templates/task/code-intelligence/r001/planning.json b/templates/task/code-intelligence/r001/planning.json index 9e17567..419a9c3 100644 --- a/templates/task/code-intelligence/r001/planning.json +++ b/templates/task/code-intelligence/r001/planning.json @@ -1,27 +1,83 @@ { - "record_version": 2, + "record_version": 3, "task_id": "TASK-0001", "work_item_revision": 1, "stage": "PLANNING", "artifact_attempt": null, "reviewer_slot": null, - "provider": null, + "provider": { + "id": "codegraph", + "descriptor_version": 2 + }, + "repository": { + "project_id": "PROJECT_ID", + "root_sha256": "0000000000000000000000000000000000000000000000000000000000000000" + }, "target": { "base_commit": "0000000000000000000000000000000000000000", "head_commit": null, "diff_hash": null }, "status": "UNAVAILABLE", - "queries": [], - "status_check": null, - "sync": null, - "freshness": { + "proxy": { + "server_id": "polaris-codegraph", + "tool": "polaris_codegraph_explore", + "evidence_bundle_sha256": "0000000000000000000000000000000000000000000000000000000000000000" + }, + "query": { + "id": "CIQ-001", + "purpose": "Record unavailable CodeGraph context", + "text": "No CodeGraph query was executed", "status": "UNAVAILABLE", - "checked_at": "1970-01-01T00:00:00Z", - "basis": ["NONE"], + "summary": "", + "symbols": [], "response_sha256": null, - "stale_points": [] + "error": "CodeGraph provider unavailable" + }, + "query_window": { + "pre_status": { + "status": "UNAVAILABLE", + "checked_at": "1970-01-01T00:00:00Z", + "basis": [ + "NONE" + ], + "stale_points": [], + "status_response_sha256": null, + "error": "CodeGraph provider unavailable", + "needs_sync": false, + "pending_changes": null + }, + "sync": null, + "post_sync_status": null, + "response_classification": null, + "post_query_status": null + }, + "delivery": { + "state": "UNAVAILABLE", + "record_status": "UNAVAILABLE", + "reason": "PROVIDER_UNAVAILABLE", + "checked_at": "1970-01-01T00:00:00Z", + "usage": "NO_GRAPH", + "required_fallback": "SEARCH_SOURCE", + "stale_points": [], + "pending_changes": { + "added": 0, + "modified": 0, + "removed": 0 + }, + "error": "CodeGraph provider unavailable" }, - "source_fallbacks": [], + "source_fallbacks": [ + { + "action": "SEARCH_SOURCE", + "path": null, + "observed_sha256": null, + "base_commit": null, + "head_commit": null, + "diff_hash": null, + "purpose": "Use repository source because CodeGraph is unavailable", + "result_paths": [] + } + ], "recorded_at": "1970-01-01T00:00:00Z" } diff --git a/templates/task/state.json b/templates/task/state.json index a623be4..7f63d16 100644 --- a/templates/task/state.json +++ b/templates/task/state.json @@ -1,6 +1,6 @@ { "task_id": "TASK-0001", - "polaris_version": "0.1.20", + "polaris_version": "0.1.21", "workflow_version": "0.1.3", "current_revision": 1, "status": "DRAFT", diff --git a/tests/fixtures/code-intelligence-record-v2.json b/tests/fixtures/code-intelligence-record-v2.json new file mode 100644 index 0000000..099cd5d --- /dev/null +++ b/tests/fixtures/code-intelligence-record-v2.json @@ -0,0 +1,29 @@ +{ + "record_version": 2, + "task_id": "TASK-0001", + "work_item_revision": 1, + "stage": "PLANNING", + "artifact_attempt": null, + "reviewer_slot": null, + "provider": null, + "target": { + "base_commit": "0000000000000000000000000000000000000000", + "head_commit": null, + "diff_hash": null + }, + "status": "UNAVAILABLE", + "queries": [], + "status_check": null, + "sync": null, + "freshness": { + "status": "UNAVAILABLE", + "checked_at": "1970-01-01T00:00:00Z", + "basis": [ + "NONE" + ], + "response_sha256": null, + "stale_points": [] + }, + "source_fallbacks": [], + "recorded_at": "1970-01-01T00:00:00Z" +} diff --git a/tests/test_codegraph.py b/tests/test_codegraph.py index 2dc9938..0b76da9 100644 --- a/tests/test_codegraph.py +++ b/tests/test_codegraph.py @@ -1,14 +1,17 @@ from __future__ import annotations +import copy import importlib +import hashlib import io import json +import os import shutil import subprocess import sys import tempfile import unittest -from contextlib import redirect_stdout +from contextlib import contextmanager, redirect_stdout from pathlib import Path from unittest import mock @@ -20,6 +23,7 @@ from init_task import initialize as init_task # noqa: E402 from migrate_project import migrate as migrate_project # noqa: E402 from new_revision import create as new_revision # noqa: E402 +from transition_task import transition # noqa: E402 from internal.code_intelligence_protocol import ( # noqa: E402 _project_marker_path, load_providers, @@ -36,6 +40,7 @@ write_json_atomic, write_text_atomic, ) +from internal.recovery_protocol import refresh_project_index # noqa: E402 from vendor_project import vendor # noqa: E402 @@ -45,6 +50,20 @@ def completed( return subprocess.CompletedProcess([], returncode, stdout, stderr) +@contextmanager +def protocol_source_at(version: str): + """Materialize a historical protocol target for adjacent migration tests.""" + with tempfile.TemporaryDirectory(prefix="polaris-codegraph-protocol-") as temp: + source = Path(temp) / "source" + shutil.copytree( + ROOT, + source, + ignore=shutil.ignore_patterns(".git", "__pycache__", "*.pyc"), + ) + (source / "VERSION").write_text(version + "\n", encoding="utf-8") + yield source + + def healthy_status(project: Path) -> str: return json.dumps( { @@ -158,7 +177,7 @@ def test_managed_surfaces_only_name_the_official_codegraph(self) -> None: self.assertIn(official, path.read_text(encoding="utf-8"), path.relative_to(ROOT).as_posix()) for path in [ROOT / "README.md", ROOT / "README.zh-CN.md"]: text = path.read_text(encoding="utf-8") - self.assertIn("0.1.20", text, path.relative_to(ROOT).as_posix()) + self.assertIn("0.1.21", text, path.relative_to(ROOT).as_posix()) self.assertIn("0.1.3", text, path.relative_to(ROOT).as_posix()) def test_authority_surfaces_publish_workflow_013(self) -> None: @@ -169,7 +188,7 @@ def test_authority_surfaces_publish_workflow_013(self) -> None: ROOT / "plan.md", ]: text = path.read_text(encoding="utf-8") - self.assertIn("0.1.20", text, path.relative_to(ROOT).as_posix()) + self.assertIn("0.1.21", text, path.relative_to(ROOT).as_posix()) self.assertIn("0.1.3", text, path.relative_to(ROOT).as_posix()) def test_readmes_keep_codegraph_operational_boundaries(self) -> None: @@ -243,7 +262,7 @@ def classify_response( def v2_record(self) -> dict[str, object]: value = json.loads( - (ROOT / "templates" / "task-sources" / "code-intelligence-record.json").read_text( + (ROOT / "tests/fixtures/code-intelligence-record-v2.json").read_text( encoding="utf-8" ) ) @@ -281,6 +300,1352 @@ def add_explore_response_evidence(value: dict[str, object]) -> None: def initialize_task(self) -> None: init_task(self.repo, "TASK-0001", "R1") + def qualify_task(self) -> None: + self.initialize_task() + path = self.repo / ".polaris/tasks/TASK-0001/revisions/work-item-r001.json" + value = json.loads(path.read_text(encoding="utf-8")) + value.update({ + "title": "CodeGraph proxy task", + "goal": "Exercise the bounded proxy", + "motivation": "Keep graph evidence freshness-aware", + }) + value["scope"]["in"] = ["scripts"] + value["acceptance"][0].update({ + "statement": "The proxy emits a freshness envelope", + "evidence": "proxy bundle", + }) + value["implementation_dispatch"]["authorized"] = True + value["review_dispatch"]["authorized"] = True + write_json_atomic(path, value) + transition( + self.repo, + "TASK-0001", + "QUALIFY", + [], + None, + None, + None, + None, + None, + None, + ) + + def proxy_module(self) -> object: + return importlib.import_module("internal.code_intelligence_proxy") + + def record_current_v3_fixture(self) -> tuple[dict[str, object], dict[str, object]]: + self.qualify_task() + (self.repo / ".codegraph").mkdir() + source = self.repo / "src/a.py" + source.parent.mkdir() + source.write_text("class A:\n pass\n", encoding="utf-8") + proxy = self.proxy_module() + responses = [ + completed(healthy_status(self.repo)), + completed("A is defined in src/a.py\n"), + completed(healthy_status(self.repo)), + ] + + def runner(command: list[str], **_kwargs: object) -> subprocess.CompletedProcess[str]: + return responses.pop(0) + + with mock.patch("internal.code_intelligence_proxy.shutil.which", return_value="/bin/codegraph"): + query = proxy.execute_proxy_query( + self.repo, + "TASK-0001", + "PLANNING", + "CIQ-001", + "locate A", + "symbol A", + False, + runner=runner, + ) + protocol = importlib.import_module("internal.code_intelligence_protocol") + result = protocol.record_proxy_bundle( + self.repo, + "TASK-0001", + query["bundle_path"], + { + "summary": "Located the affected symbol.", + "symbols": [{"path": "src/a.py", "line": 1, "name": "A"}], + "source_fallbacks": [], + }, + ROOT, + ) + recorded = json.loads(Path(result["path"]).read_text(encoding="utf-8")) + return recorded, query + + def test_proxy_stage_context_uses_record_name_and_sequential_query_ids(self) -> None: + self.qualify_task() + proxy = self.proxy_module() + + context = proxy.resolve_stage_context(self.repo, "TASK-0001", "PLANNING") + + self.assertEqual(context["work_item_revision"], 1) + self.assertEqual(context["artifact_attempt"], None) + self.assertEqual(context["reviewer_slot"], None) + self.assertEqual(context["record_name"], "planning") + expected = ( + self.repo + / ".polaris/tasks/TASK-0001/runtime/code-intelligence/planning/CIQ-001.json" + ) + self.assertEqual( + proxy.proxy_bundle_path(self.repo, "TASK-0001", context, "CIQ-001"), + expected, + ) + with self.assertRaisesRegex(InputFailure, "next sequential"): + proxy.proxy_bundle_path(self.repo, "TASK-0001", context, "CIQ-002") + with self.assertRaisesRegex(InputFailure, "invalid CodeGraph query id"): + proxy.proxy_bundle_path(self.repo, "TASK-0001", context, "CIQ-000") + + def test_proxy_stage_context_binds_attempt_subject_and_review_slot(self) -> None: + self.qualify_task() + proxy = self.proxy_module() + task = self.repo / ".polaris/tasks/TASK-0001" + state_path = task / "state.json" + state = json.loads(state_path.read_text(encoding="utf-8")) + base = subprocess.run( + ["git", "rev-parse", "HEAD"], + cwd=self.repo, + capture_output=True, + check=True, + text=True, + encoding="utf-8", + ).stdout.strip() + implementation_handoff = { + "artifact_attempt": 2, + "subject_base_commit": base, + } + state["status"] = "IMPLEMENTING" + write_json_atomic(state_path, state) + with mock.patch( + "internal.code_intelligence_proxy.validate_implementation_handoff", + return_value=(implementation_handoff, {"path": "handoff.json", "sha256": "0" * 64}), + ): + implementation = proxy.resolve_stage_context( + self.repo, "TASK-0001", "IMPLEMENTATION" + ) + documentation = proxy.resolve_stage_context( + self.repo, "TASK-0001", "DOCUMENTATION_SYNC" + ) + self.assertEqual(implementation["record_name"], "implementation-002") + self.assertEqual(documentation["record_name"], "documentation-sync-002") + self.assertEqual(implementation["target"], { + "base_commit": base, + "head_commit": base, + "diff_hash": subject_diff_hash(self.repo, base, base), + }) + + state["status"] = "REVIEWING" + write_json_atomic(state_path, state) + review_handoff = { + "artifact_attempt": 2, + "subject_base_commit": base, + "subject_head_commit": base, + "subject_diff_hash": subject_diff_hash(self.repo, base, base), + } + with mock.patch( + "internal.code_intelligence_proxy.validate_review_handoff", + return_value=review_handoff, + ): + review = proxy.resolve_stage_context(self.repo, "TASK-0001", "REVIEW") + self.assertEqual(review["record_name"], "review-002-slot-1") + state["artifacts"]["review"] = {"path": "review.json", "sha256": "0" * 64} + write_json_atomic(state_path, state) + with mock.patch( + "internal.code_intelligence_proxy.validate_review_handoff", + return_value=review_handoff, + ): + review_2 = proxy.resolve_stage_context(self.repo, "TASK-0001", "REVIEW") + self.assertEqual(review_2["record_name"], "review-002-slot-2") + + with self.assertRaisesRegex(RuleFailure, "inconsistent with task status"): + proxy.resolve_stage_context(self.repo, "TASK-0001", "PLANNING") + + def test_proxy_window_requires_clean_pre_and_post_status_for_current(self) -> None: + self.qualify_task() + (self.repo / ".codegraph").mkdir() + proxy = self.proxy_module() + calls: list[tuple[list[str], dict[str, object]]] = [] + responses = [ + completed(healthy_status(self.repo)), + completed("graph bytes\n"), + completed(healthy_status(self.repo)), + ] + + def runner(command: list[str], **kwargs: object) -> subprocess.CompletedProcess[str]: + calls.append((command, kwargs)) + if not responses: + raise AssertionError(f"unexpected extra CodeGraph call: {command}") + return responses.pop(0) + + with mock.patch("internal.code_intelligence_proxy.shutil.which", return_value="/bin/codegraph"): + result = proxy.execute_proxy_query( + self.repo, + "TASK-0001", + "PLANNING", + "CIQ-001", + "locate affected symbols", + "symbol A", + False, + runner=runner, + ) + + bundle = result["bundle"] + self.assertEqual(bundle["delivery"]["state"], "CURRENT") + self.assertEqual(bundle["delivery"]["usage"], "NON_AUTHORITATIVE_CONTEXT") + self.assertEqual(bundle["delivery"]["record_status"], "CURRENT_AT_CHECK") + self.assertEqual([call[0][1] for call in calls], ["status", "explore", "status"]) + self.assertTrue(all(Path(call[1]["cwd"]).resolve() == self.repo.resolve() for call in calls)) + self.assertEqual(result["response"], "graph bytes\n") + self.assertTrue(result["envelope"].startswith("[POLARIS_CODEGRAPH_FRESHNESS]\n")) + self.assertEqual( + hashlib.sha256( + (self.repo / ".polaris/tasks/TASK-0001/runtime/code-intelligence/planning/CIQ-001.response.txt").read_bytes() + ).hexdigest(), + bundle["query"]["response_sha256"], + ) + + def test_proxy_window_downgrades_pending_unknown_and_unavailable_states(self) -> None: + cases = [ + ("pending", "STALE", 3), + ("malformed", "UNKNOWN", 1), + ("missing_marker", "UNAVAILABLE", 0), + ] + for index, (case, expected_state, expected_calls) in enumerate(cases, start=1): + with self.subTest(case=case): + if index > 1: + self.tearDown() + self.setUp() + self.qualify_task() + proxy = self.proxy_module() + query_id = "CIQ-001" + calls: list[list[str]] = [] + status = json.loads(healthy_status(self.repo)) + if case == "pending": + status["pendingChanges"]["modified"] = 1 + responses = [ + completed(json.dumps(status)), + completed("graph bytes\n"), + completed(json.dumps(status)), + ] + else: + responses = [completed("not-json\n")] + if case != "missing_marker": + (self.repo / ".codegraph").mkdir() + + def runner(command: list[str], **_kwargs: object) -> subprocess.CompletedProcess[str]: + calls.append(command) + return responses.pop(0) + + with mock.patch("internal.code_intelligence_proxy.shutil.which", return_value="/bin/codegraph"): + result = proxy.execute_proxy_query( + self.repo, + "TASK-0001", + "PLANNING", + query_id, + "inspect freshness", + "symbol A", + False, + runner=runner, + ) + self.assertEqual(result["bundle"]["delivery"]["state"], expected_state) + self.assertEqual(len(calls), expected_calls) + if expected_state == "UNAVAILABLE": + self.assertIsNone(result["response"]) + + def test_proxy_window_syncs_once_and_rechecks_after_the_query(self) -> None: + self.qualify_task() + (self.repo / ".codegraph").mkdir() + proxy = self.proxy_module() + pending = json.loads(healthy_status(self.repo)) + pending["pendingChanges"]["modified"] = 1 + responses = [ + completed(json.dumps(pending)), + completed("synced\n"), + completed(healthy_status(self.repo)), + completed("graph bytes\n"), + completed(healthy_status(self.repo)), + ] + calls: list[list[str]] = [] + + def runner(command: list[str], **_kwargs: object) -> subprocess.CompletedProcess[str]: + calls.append(command) + return responses.pop(0) + + with mock.patch("internal.code_intelligence_proxy.shutil.which", return_value="/bin/codegraph"): + result = proxy.execute_proxy_query( + self.repo, + "TASK-0001", + "PLANNING", + "CIQ-001", + "refresh one query window", + "symbol A", + True, + runner=runner, + ) + + self.assertEqual( + [command[1] for command in calls], + ["status", "sync", "status", "explore", "status"], + ) + self.assertEqual(result["bundle"]["sync"]["status"], "SUCCESS") + self.assertEqual(result["bundle"]["delivery"]["state"], "CURRENT") + self.assertEqual( + result["bundle"]["delivery"]["pending_changes"], + {"added": 0, "modified": 0, "removed": 0}, + ) + + def test_proxy_window_never_promotes_failed_or_post_stale_queries(self) -> None: + cases = [ + ("explore_failed", "UNKNOWN", ["status", "explore"]), + ("post_pending", "STALE", ["status", "explore", "status"]), + ] + for index, (case, expected_state, expected_calls) in enumerate(cases, start=1): + with self.subTest(case=case): + if index > 1: + self.tearDown() + self.setUp() + self.qualify_task() + (self.repo / ".codegraph").mkdir() + proxy = self.proxy_module() + post = json.loads(healthy_status(self.repo)) + post["pendingChanges"]["added"] = 1 + responses = ( + [completed(healthy_status(self.repo)), completed("", 1, "failed")] + if case == "explore_failed" + else [ + completed(healthy_status(self.repo)), + completed("graph bytes\n"), + completed(json.dumps(post)), + ] + ) + calls: list[list[str]] = [] + + def runner(command: list[str], **_kwargs: object) -> subprocess.CompletedProcess[str]: + calls.append(command) + return responses.pop(0) + + with mock.patch("internal.code_intelligence_proxy.shutil.which", return_value="/bin/codegraph"): + result = proxy.execute_proxy_query( + self.repo, + "TASK-0001", + "PLANNING", + "CIQ-001", + "verify failure handling", + "symbol A", + False, + runner=runner, + ) + self.assertEqual(result["bundle"]["delivery"]["state"], expected_state) + self.assertEqual([command[1] for command in calls], expected_calls) + if case == "explore_failed": + self.assertIsNone(result["response"]) + else: + self.assertEqual( + result["bundle"]["delivery"]["reason"], "PENDING_CHANGES" + ) + + def test_proxy_window_classifies_banners_and_discards_unsafe_paths(self) -> None: + partial = """⚠️ Some files referenced below were edited since the last index sync — +their codegraph entries may be stale: + - src/widget.py (edited 800ms ago, pending sync) +For accurate content of those specific files, Read them directly. +""" + cases = [ + ("partial", partial, "STALE", True), + ("suspicious", "warning: graph may be stale\n", "UNKNOWN", True), + ("unsafe", partial.replace("src/widget.py", "../escape.py"), "UNKNOWN", False), + ] + for index, (case, response, expected_state, retained) in enumerate(cases, start=1): + with self.subTest(case=case): + if index > 1: + self.tearDown() + self.setUp() + self.qualify_task() + (self.repo / ".codegraph").mkdir() + source = self.repo / "src/widget.py" + source.parent.mkdir() + source.write_text("value = 1\n", encoding="utf-8") + proxy = self.proxy_module() + responses = [ + completed(healthy_status(self.repo)), + completed(response), + completed(healthy_status(self.repo)), + ] + + def runner(command: list[str], **_kwargs: object) -> subprocess.CompletedProcess[str]: + return responses.pop(0) + + with mock.patch("internal.code_intelligence_proxy.shutil.which", return_value="/bin/codegraph"): + result = proxy.execute_proxy_query( + self.repo, + "TASK-0001", + "PLANNING", + "CIQ-001", + "classify response", + "symbol A", + False, + runner=runner, + ) + self.assertEqual(result["bundle"]["delivery"]["state"], expected_state) + self.assertEqual(result["response"] is not None, retained) + self.assertEqual(result["bundle"]["response_path"] is not None, retained) + + def test_proxy_window_rejects_cross_project_and_runtime_symlink_evidence(self) -> None: + self.qualify_task() + (self.repo / ".codegraph").mkdir() + proxy = self.proxy_module() + other = self.repo / "other" + other.mkdir() + wrong = json.loads(healthy_status(self.repo)) + wrong["projectPath"] = str(other) + calls: list[list[str]] = [] + + def runner(command: list[str], **_kwargs: object) -> subprocess.CompletedProcess[str]: + calls.append(command) + return completed(json.dumps(wrong)) + + with mock.patch("internal.code_intelligence_proxy.shutil.which", return_value="/bin/codegraph"): + result = proxy.execute_proxy_query( + self.repo, + "TASK-0001", + "PLANNING", + "CIQ-001", + "reject cross-project status", + "symbol A", + False, + runner=runner, + ) + self.assertEqual([command[1] for command in calls], ["status"]) + self.assertEqual(result["bundle"]["delivery"]["state"], "UNKNOWN") + self.assertEqual(result["bundle"]["delivery"]["reason"], "PROJECT_MISMATCH") + + self.tearDown() + self.setUp() + self.qualify_task() + runtime = self.repo / ".polaris/tasks/TASK-0001/runtime" + shutil.rmtree(runtime) + runtime.symlink_to(self.repo / "escaped-runtime", target_is_directory=True) + context = proxy.resolve_stage_context(self.repo, "TASK-0001", "PLANNING") + with self.assertRaisesRegex(RuleFailure, "crosses a symlink"): + proxy.proxy_bundle_path(self.repo, "TASK-0001", context, "CIQ-001") + + def test_proxy_window_disabled_policy_or_missing_cli_never_calls_provider(self) -> None: + for index, case in enumerate(("disabled", "missing_cli"), start=1): + with self.subTest(case=case): + if index > 1: + self.tearDown() + self.setUp() + self.qualify_task() + (self.repo / ".codegraph").mkdir() + if case == "disabled": + config_path = self.repo / ".polaris/code-intelligence.json" + config = json.loads( + (ROOT / "templates/code-intelligence.json").read_text( + encoding="utf-8" + ) + ) + config["mode"] = "disabled" + write_json_atomic(config_path, config) + proxy = self.proxy_module() + + def runner(*_args: object, **_kwargs: object) -> subprocess.CompletedProcess[str]: + raise AssertionError("disabled or unavailable CodeGraph must not run") + + executable = "/bin/codegraph" if case == "disabled" else None + with mock.patch( + "internal.code_intelligence_proxy.shutil.which", + return_value=executable, + ): + result = proxy.execute_proxy_query( + self.repo, + "TASK-0001", + "PLANNING", + "CIQ-001", + "verify activation gate", + "symbol A", + True, + runner=runner, + ) + self.assertEqual(result["bundle"]["delivery"]["state"], "UNAVAILABLE") + self.assertEqual(result["bundle"]["query"]["status"], "UNAVAILABLE") + self.assertIsNone(result["response"]) + + def test_proxy_envelope_is_finite_and_truncates_diagnostics(self) -> None: + self.qualify_task() + proxy = self.proxy_module() + context = proxy.resolve_stage_context(self.repo, "TASK-0001", "PLANNING") + bundle = { + "task_context": context, + "query": {"id": "CIQ-001"}, + "delivery": { + "state": "UNKNOWN", + "record_status": "NOT_VERIFIED", + "reason": "STATUS_UNREADABLE", + "checked_at": "2026-08-19T00:00:00Z", + "pending_changes": {"added": 0, "modified": 0, "removed": 0}, + "usage": "NAVIGATION_ONLY", + "required_fallback": "SEARCH_SOURCE", + "error": "x" * 300, + }, + } + + envelope = proxy.render_freshness_envelope(bundle) + + self.assertTrue(envelope.startswith("[POLARIS_CODEGRAPH_FRESHNESS]\n")) + self.assertTrue(envelope.endswith("[/POLARIS_CODEGRAPH_FRESHNESS]\n")) + error_line = next(line for line in envelope.splitlines() if line.startswith("error: ")) + self.assertEqual(len(error_line.removeprefix("error: ")), 240) + + def test_mcp_server_initializes_and_lists_one_proxy_tool(self) -> None: + messages = [ + { + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-11-25", + "capabilities": {}, + "clientInfo": {"name": "test", "version": "1"}, + }, + }, + {"jsonrpc": "2.0", "method": "notifications/initialized"}, + {"jsonrpc": "2.0", "id": 2, "method": "tools/list", "params": {}}, + ] + completed_process = subprocess.run( + [ + sys.executable, + SCRIPTS / "code_intelligence_mcp.py", + "--repo", + ".", + ], + cwd=self.repo, + input="".join(json.dumps(item) + "\n" for item in messages), + text=True, + capture_output=True, + check=False, + ) + + self.assertEqual(completed_process.returncode, 0, completed_process.stderr) + responses = [json.loads(line) for line in completed_process.stdout.splitlines()] + self.assertEqual(len(responses), 2) + self.assertEqual(responses[0]["result"]["protocolVersion"], "2025-11-25") + self.assertEqual( + responses[0]["result"]["capabilities"], + {"tools": {"listChanged": False}}, + ) + tools = responses[1]["result"]["tools"] + self.assertEqual([item["name"] for item in tools], ["polaris_codegraph_explore"]) + self.assertNotIn("repository", tools[0]["inputSchema"]["properties"]) + self.assertEqual(completed_process.stderr, "") + + def test_mcp_server_returns_envelope_before_graph_and_preserves_bundle(self) -> None: + module = importlib.import_module("code_intelligence_mcp") + server = module.McpServer(self.repo) + initialized = server.handle({ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-11-25", + "capabilities": {}, + "clientInfo": {"name": "test", "version": "1"}, + }, + }) + self.assertIn("result", initialized) + self.assertIsNone(server.handle({ + "jsonrpc": "2.0", "method": "notifications/initialized" + })) + bundle = { + "delivery": {"state": "STALE"}, + "query": {"id": "CIQ-001"}, + } + proxy_result = { + "bundle": bundle, + "bundle_path": self.repo / "bundle.json", + "response": "graph bytes\n", + "envelope": "[POLARIS_CODEGRAPH_FRESHNESS]\nstate: STALE\n[/POLARIS_CODEGRAPH_FRESHNESS]\n", + } + request = { + "jsonrpc": "2.0", + "id": 2, + "method": "tools/call", + "params": { + "name": "polaris_codegraph_explore", + "arguments": { + "task_id": "TASK-0001", + "stage": "PLANNING", + "query_id": "CIQ-001", + "purpose": "locate symbols", + "query": "symbol A", + "sync_if_needed": False, + }, + }, + } + with mock.patch( + "code_intelligence_mcp.execute_proxy_query", return_value=proxy_result + ): + response = server.handle(request) + + result = response["result"] + self.assertFalse(result["isError"]) + self.assertEqual(result["content"][0]["text"], proxy_result["envelope"]) + self.assertEqual(result["content"][1]["text"], "graph bytes\n") + self.assertEqual(result["structuredContent"], {"bundle": bundle}) + + proxy_result["response"] = None + with mock.patch( + "code_intelligence_mcp.execute_proxy_query", return_value=proxy_result + ): + no_graph = server.handle({**request, "id": 3}) + self.assertFalse(no_graph["result"]["isError"]) + self.assertEqual(len(no_graph["result"]["content"]), 1) + + def test_mcp_server_rejects_lifecycle_tool_and_input_errors(self) -> None: + module = importlib.import_module("code_intelligence_mcp") + server = module.McpServer(self.repo) + before = server.handle({ + "jsonrpc": "2.0", "id": 1, "method": "tools/list", "params": {} + }) + self.assertEqual(before["error"]["code"], -32600) + server.handle({ + "jsonrpc": "2.0", + "id": 2, + "method": "initialize", + "params": { + "protocolVersion": "2025-11-25", + "capabilities": {}, + "clientInfo": {"name": "test", "version": "1"}, + }, + }) + server.handle({"jsonrpc": "2.0", "method": "notifications/initialized"}) + unknown = server.handle({ + "jsonrpc": "2.0", + "id": 3, + "method": "tools/call", + "params": {"name": "codegraph_explore", "arguments": {}}, + }) + self.assertEqual(unknown["error"]["code"], -32602) + invalid = server.handle({ + "jsonrpc": "2.0", + "id": 4, + "method": "tools/call", + "params": { + "name": "polaris_codegraph_explore", + "arguments": { + "task_id": "TASK-0001", + "stage": "INVALID", + "query_id": "CIQ-000", + "purpose": "locate symbols", + "query": "symbol A", + "sync_if_needed": False, + }, + }, + }) + self.assertTrue(invalid["result"]["isError"]) + self.assertNotIn("structuredContent", invalid["result"]) + + def test_mcp_server_emits_jsonrpc_parse_and_method_errors_one_per_line(self) -> None: + transcript = "{bad json\n" + "\n".join([ + json.dumps({"jsonrpc": "2.0", "id": 1, "method": "initialize", "params": {}}), + json.dumps({"jsonrpc": "2.0", "id": 2, "method": "unknown", "params": {}}), + ]) + "\n" + completed_process = subprocess.run( + [sys.executable, SCRIPTS / "code_intelligence_mcp.py", "--repo", "."], + cwd=self.repo, + input=transcript, + text=True, + capture_output=True, + check=False, + ) + + responses = [json.loads(line) for line in completed_process.stdout.splitlines()] + self.assertEqual([item["error"]["code"] for item in responses], [-32700, -32602, -32601]) + self.assertTrue(all("\n" not in line for line in completed_process.stdout.splitlines())) + + def test_mcp_server_rejects_cwd_and_vendored_launcher_project_mismatch(self) -> None: + mismatched_cwd = subprocess.run( + [sys.executable, SCRIPTS / "code_intelligence_mcp.py", "--repo", self.repo], + cwd=ROOT, + input="", + text=True, + capture_output=True, + check=False, + ) + self.assertEqual(mismatched_cwd.returncode, 2) + self.assertIn("working directory", mismatched_cwd.stderr) + + with tempfile.TemporaryDirectory(prefix="polaris-codegraph-launcher-") as temporary: + fixture = Path(temporary) + projects = [fixture / "project-a", fixture / "project-b"] + for index, project in enumerate(projects, start=1): + project.mkdir() + subprocess.run(["git", "init", "-q"], cwd=project, check=True) + vendor(ROOT, project, False) + init_project(project, f"launcher-{index}") + foreign_launcher = ( + projects[0] / "tools/polaris/scripts/code_intelligence_mcp.py" + ) + mismatched_launcher = subprocess.run( + [sys.executable, foreign_launcher, "--repo", "."], + cwd=projects[1], + input="", + text=True, + capture_output=True, + check=False, + ) + self.assertEqual(mismatched_launcher.returncode, 2) + self.assertIn("vendored protocol root", mismatched_launcher.stderr) + + def test_vendored_mcp_proxy_runs_one_auditable_fake_cli_window(self) -> None: + """Registered MCP, fake CLI, v3 projection, and Validation compose end to end.""" + with tempfile.TemporaryDirectory(prefix="polaris-codegraph-e2e-") as temporary: + fixture_root = Path(temporary) + repo = fixture_root / "repository" + fake_bin = fixture_root / "bin" + repo.mkdir() + fake_bin.mkdir() + subprocess.run(["git", "init", "-q"], cwd=repo, check=True) + subprocess.run( + ["git", "config", "user.email", "polaris@test.local"], + cwd=repo, + check=True, + ) + subprocess.run( + ["git", "config", "user.name", "Polaris Test"], + cwd=repo, + check=True, + ) + vendor(ROOT, repo, False) + init_project(repo, "codegraph-e2e") + source = repo / "src/a.py" + source.parent.mkdir() + source.write_text("class A:\n pass\n", encoding="utf-8") + (repo / ".codegraph").mkdir() + subprocess.run(["git", "add", "."], cwd=repo, check=True) + subprocess.run( + ["git", "commit", "-q", "-m", "initialize fixture"], + cwd=repo, + check=True, + ) + init_task(repo, "TASK-0001", "R1") + work_item_path = ( + repo + / ".polaris/tasks/TASK-0001/revisions/work-item-r001.json" + ) + work_item = json.loads(work_item_path.read_text(encoding="utf-8")) + work_item.update({ + "title": "Exercise the registered proxy", + "goal": "Prove one bounded CodeGraph window", + "motivation": "Keep graph evidence auditable", + }) + work_item["scope"]["in"] = ["src/a.py"] + work_item["acceptance"][0].update({ + "statement": "The registered proxy emits a current envelope", + "evidence": "v3 Code Intelligence record", + }) + work_item["implementation_dispatch"]["authorized"] = True + work_item["review_dispatch"]["authorized"] = True + write_json_atomic(work_item_path, work_item) + transition( + repo, + "TASK-0001", + "QUALIFY", + [], + None, + None, + None, + None, + None, + None, + ) + + call_log = fixture_root / "codegraph-calls.jsonl" + executable = fake_bin / "codegraph.py" + write_text_atomic( + executable, + """#!/usr/bin/env python3 +import json +import os +import sys +from pathlib import Path + +log = Path(os.environ["POLARIS_FAKE_CODEGRAPH_LOG"]) +entry = {"cwd": str(Path.cwd().resolve()), "argv": sys.argv[1:]} +with log.open("a", encoding="utf-8") as stream: + stream.write(json.dumps(entry, separators=(",", ":")) + "\\n") +entries = [json.loads(line) for line in log.read_text(encoding="utf-8").splitlines()] +args = sys.argv[1:] +if args == ["status", "--json"]: + status_count = sum(item["argv"] == ["status", "--json"] for item in entries) + pending = 1 if status_count == 1 else 0 + print(json.dumps({ + "initialized": True, + "projectPath": str(Path.cwd().resolve()), + "pendingChanges": {"added": 0, "modified": pending, "removed": 0}, + "worktreeMismatch": None, + "index": {"state": "complete", "pendingRefs": 0, "reindexRecommended": False}, + })) +elif args == ["sync", "--quiet"]: + print("synchronized") +elif len(args) == 2 and args[0] == "explore": + print("A is defined in src/a.py") +else: + print("unexpected fake CodeGraph arguments", file=sys.stderr) + raise SystemExit(2) +""", + ) + descriptor_path = ( + repo + / "tools/polaris/providers/code-intelligence/codegraph.json" + ) + descriptor_bytes = descriptor_path.read_bytes() + runtime_descriptor = json.loads(descriptor_bytes) + runtime_descriptor["cli"] = { + "executable": sys.executable, + "explore_args": [str(executable), "explore"], + "status_args": [str(executable), "status", "--json"], + "sync_args": [str(executable), "sync", "--quiet"], + } + write_json_atomic(descriptor_path, runtime_descriptor) + environment = os.environ.copy() + environment["POLARIS_FAKE_CODEGRAPH_LOG"] = str(call_log) + + registration = json.loads((repo / ".mcp.json").read_text(encoding="utf-8"))[ + "mcpServers" + ]["polaris-codegraph"] + transcript = "\n".join([ + json.dumps({ + "jsonrpc": "2.0", + "id": 1, + "method": "initialize", + "params": { + "protocolVersion": "2025-11-25", + "capabilities": {}, + "clientInfo": {"name": "Polaris test", "version": "1"}, + }, + }), + json.dumps({ + "jsonrpc": "2.0", + "method": "notifications/initialized", + }), + json.dumps({ + "jsonrpc": "2.0", + "id": 2, + "method": "tools/call", + "params": { + "name": "polaris_codegraph_explore", + "arguments": { + "task_id": "TASK-0001", + "stage": "PLANNING", + "query_id": "CIQ-001", + "purpose": "locate A", + "query": "symbol A", + "sync_if_needed": True, + }, + }, + }), + ]) + "\n" + mcp = subprocess.run( + [registration["command"], *registration["args"]], + cwd=repo, + env=environment, + input=transcript, + text=True, + capture_output=True, + check=False, + ) + descriptor_path.write_bytes(descriptor_bytes) + self.assertEqual(mcp.returncode, 0, mcp.stderr) + responses = [json.loads(line) for line in mcp.stdout.splitlines()] + self.assertEqual([response["id"] for response in responses], [1, 2]) + tool_result = responses[1]["result"] + self.assertFalse(tool_result["isError"]) + self.assertTrue( + tool_result["content"][0]["text"].startswith( + "[POLARIS_CODEGRAPH_FRESHNESS]\nstate: CURRENT\n" + ) + ) + self.assertEqual( + tool_result["content"][1]["text"], "A is defined in src/a.py\n" + ) + bundle = tool_result["structuredContent"]["bundle"] + envelope = tool_result["content"][0]["text"] + bundle_relative = next( + line.split(": ", 1)[1] + for line in envelope.splitlines() + if line.startswith("evidence_bundle: ") + ) + task_relative = Path(".polaris/tasks/TASK-0001") + bundle_repo_relative = task_relative / bundle_relative + bundle_path = repo / bundle_repo_relative + response_relative = task_relative / bundle["response_path"] + for ignored in (bundle_repo_relative, response_relative): + ignored_result = subprocess.run( + ["git", "check-ignore", "-q", ignored.as_posix()], + cwd=repo, + check=False, + ) + self.assertEqual(ignored_result.returncode, 0, ignored.as_posix()) + + calls = [ + json.loads(line) + for line in call_log.read_text(encoding="utf-8").splitlines() + ] + self.assertEqual( + [entry["argv"] for entry in calls], + [ + ["status", "--json"], + ["sync", "--quiet"], + ["status", "--json"], + ["explore", "symbol A"], + ["status", "--json"], + ], + ) + self.assertTrue(all(entry["cwd"] == str(repo.resolve()) for entry in calls)) + + annotations_path = fixture_root / "annotations.json" + write_json_atomic( + annotations_path, + { + "summary": "Located A through the bounded proxy.", + "symbols": [{"path": "src/a.py", "line": 1, "name": "A"}], + "source_fallbacks": [], + }, + ) + recorder = subprocess.run( + [ + sys.executable, + "tools/polaris/scripts/record_code_intelligence.py", + "TASK-0001", + "--repo", + ".", + "--bundle", + bundle_repo_relative.as_posix(), + "--annotations", + str(annotations_path), + "--json", + ], + cwd=repo, + text=True, + capture_output=True, + check=False, + ) + self.assertEqual(recorder.returncode, 0, recorder.stderr) + record_result = json.loads(recorder.stdout) + record_value = json.loads( + Path(record_result["path"]).read_text(encoding="utf-8") + ) + self.assertEqual( + record_value["proxy"]["evidence_bundle_sha256"], + file_sha256(bundle_path), + ) + + calls_before_validation = call_log.read_bytes() + validation = subprocess.run( + [ + sys.executable, + "tools/polaris/scripts/validate_project.py", + "--repo", + ".", + "--json", + ], + cwd=repo, + env=environment, + text=True, + capture_output=True, + check=False, + ) + self.assertEqual(validation.returncode, 0, validation.stderr) + self.assertEqual(call_log.read_bytes(), calls_before_validation) + + def test_v3_record_projects_exact_proxy_bundle(self) -> None: + recorded, query = self.record_current_v3_fixture() + self.assertEqual(recorded["record_version"], 3) + self.assertEqual(recorded["proxy"]["server_id"], "polaris-codegraph") + self.assertEqual( + recorded["query_window"]["pre_status"]["pending_changes"], + {"added": 0, "modified": 0, "removed": 0}, + ) + self.assertEqual(recorded["delivery"]["state"], "CURRENT") + self.assertEqual( + recorded["proxy"]["evidence_bundle_sha256"], + file_sha256(query["bundle_path"]), + ) + self.assertEqual(recorded["query"]["symbols"][0]["path"], "src/a.py") + + def test_v3_record_rejects_mutated_window_identity_and_fallbacks(self) -> None: + recorded, _query = self.record_current_v3_fixture() + protocol = importlib.import_module("internal.code_intelligence_protocol") + + def current_to_stale(value: dict[str, object]) -> None: + pending = { + "scope": "INDEX", + "path": None, + "reason": "PENDING_CHANGES", + "fallback": "SEARCH_SOURCE", + "observed_sha256": None, + } + value["query_window"]["post_query_status"]["pending_changes"]["added"] = 1 + value["query_window"]["post_query_status"]["needs_sync"] = True + value["delivery"].update({ + "state": "STALE", + "record_status": "INDEX_STALE", + "reason": "PENDING_CHANGES", + "usage": "NAVIGATION_ONLY", + "required_fallback": "SEARCH_SOURCE", + "stale_points": [pending], + "pending_changes": {"added": 1, "modified": 0, "removed": 0}, + }) + + mutations: list[tuple[str, object]] = [ + ( + "current pending", + lambda value: value["delivery"]["pending_changes"].update( + {"modified": 1} + ), + ), + ( + "missing post status", + lambda value: value["query_window"].update( + {"post_query_status": None} + ), + ), + ( + "response hash mismatch", + lambda value: value["query_window"]["response_classification"].update( + {"response_sha256": "1" * 64} + ), + ), + ( + "bundle hash shape", + lambda value: value["proxy"].update( + {"evidence_bundle_sha256": "short"} + ), + ), + ( + "repository mismatch", + lambda value: value["repository"].update( + {"project_id": "another-project"} + ), + ), + ( + "delivery usage", + lambda value: value["delivery"].update( + {"usage": "NAVIGATION_ONLY"} + ), + ), + ( + "sync without post-sync", + lambda value: value["query_window"].update({ + "sync": { + "status": "SUCCESS", + "response_sha256": "2" * 64, + "error": None, + }, + "post_sync_status": None, + }), + ), + ] + for name, mutation in mutations: + with self.subTest(name=name): + value = copy.deepcopy(recorded) + mutation(value) + with self.assertRaises(RuleFailure): + protocol.validate_record_value(self.repo, "TASK-0001", value, ROOT) + + stale_without_fallback = copy.deepcopy(recorded) + current_to_stale(stale_without_fallback) + with self.assertRaisesRegex(RuleFailure, "source fallback"): + protocol.validate_record_value( + self.repo, "TASK-0001", stale_without_fallback, ROOT + ) + + unconfirmed_symbol = copy.deepcopy(stale_without_fallback) + unconfirmed_symbol["source_fallbacks"] = [{ + "action": "SEARCH_SOURCE", + "path": None, + "observed_sha256": None, + "base_commit": None, + "head_commit": None, + "diff_hash": None, + "purpose": "inspect current repository source", + "result_paths": [], + }] + with self.assertRaisesRegex(RuleFailure, "symbol.*source fallback"): + protocol.validate_record_value( + self.repo, "TASK-0001", unconfirmed_symbol, ROOT + ) + + contradictory_status = copy.deepcopy(stale_without_fallback) + contradictory_status["source_fallbacks"] = [{ + "action": "SEARCH_SOURCE", + "path": None, + "observed_sha256": None, + "base_commit": None, + "head_commit": None, + "diff_hash": None, + "purpose": "inspect current repository source", + "result_paths": [{ + "path": "src/a.py", + "observed_sha256": file_sha256(self.repo / "src/a.py"), + }], + }] + contradictory_status["delivery"]["record_status"] = "CURRENT_AT_CHECK" + with self.assertRaisesRegex(RuleFailure, "record freshness status"): + protocol.validate_record_value( + self.repo, "TASK-0001", contradictory_status, ROOT + ) + + unsafe_fallback = copy.deepcopy(stale_without_fallback) + unsafe_fallback["source_fallbacks"] = [{ + "action": "SEARCH_SOURCE", + "path": None, + "observed_sha256": None, + "base_commit": None, + "head_commit": None, + "diff_hash": None, + "purpose": "inspect current repository source", + "result_paths": [{ + "path": "../outside.py", + "observed_sha256": "0" * 64, + }], + }] + with self.assertRaises(RuleFailure): + protocol.validate_record_value( + self.repo, "TASK-0001", unsafe_fallback, ROOT + ) + + with self.assertRaisesRegex(InputFailure, "record_version 3"): + record(self.repo, "TASK-0001", self.v2_record(), ROOT) + + def test_failed_explore_proxy_bundle_projects_to_unknown_v3(self) -> None: + self.qualify_task() + (self.repo / ".codegraph").mkdir() + source = self.repo / "src/a.py" + source.parent.mkdir() + source.write_text("class A:\n pass\n", encoding="utf-8") + responses = [ + completed(healthy_status(self.repo)), + completed("failed explore output\n", returncode=1), + ] + + def runner(command: list[str], **_kwargs: object) -> subprocess.CompletedProcess[str]: + return responses.pop(0) + + proxy = self.proxy_module() + with mock.patch( + "internal.code_intelligence_proxy.shutil.which", + return_value="/bin/codegraph", + ): + query = proxy.execute_proxy_query( + self.repo, + "TASK-0001", + "PLANNING", + "CIQ-001", + "locate A", + "symbol A", + False, + runner=runner, + ) + protocol = importlib.import_module("internal.code_intelligence_protocol") + result = protocol.record_proxy_bundle( + self.repo, + "TASK-0001", + query["bundle_path"], + { + "summary": "Explore failed; verified current source instead.", + "symbols": [], + "source_fallbacks": [{ + "action": "SEARCH_SOURCE", + "path": None, + "observed_sha256": None, + "base_commit": None, + "head_commit": None, + "diff_hash": None, + "purpose": "locate A in current source", + "result_paths": [{ + "path": "src/a.py", + "observed_sha256": file_sha256(source), + }], + }], + }, + ROOT, + ) + recorded = json.loads(Path(result["path"]).read_text(encoding="utf-8")) + self.assertEqual(recorded["delivery"]["state"], "UNKNOWN") + self.assertEqual(recorded["query"]["response_sha256"], None) + self.assertEqual(recorded["delivery"]["stale_points"], []) + + def test_failed_sync_proxy_bundle_preserves_only_observed_post_status(self) -> None: + self.qualify_task() + (self.repo / ".codegraph").mkdir() + source = self.repo / "src/a.py" + source.parent.mkdir() + source.write_text("class A:\n pass\n", encoding="utf-8") + pending = json.loads(healthy_status(self.repo)) + pending["pendingChanges"]["modified"] = 1 + responses = [ + completed(json.dumps(pending)), + completed("sync failed\n", returncode=1), + completed("A is defined in src/a.py\n"), + completed(healthy_status(self.repo)), + ] + + def runner(command: list[str], **_kwargs: object) -> subprocess.CompletedProcess[str]: + return responses.pop(0) + + proxy = self.proxy_module() + with mock.patch( + "internal.code_intelligence_proxy.shutil.which", + return_value="/bin/codegraph", + ): + query = proxy.execute_proxy_query( + self.repo, + "TASK-0001", + "PLANNING", + "CIQ-001", + "locate A after one sync attempt", + "symbol A", + True, + runner=runner, + ) + self.assertIsNone(query["bundle"]["post_sync_status"]) + protocol = importlib.import_module("internal.code_intelligence_protocol") + result = protocol.record_proxy_bundle( + self.repo, + "TASK-0001", + query["bundle_path"], + { + "summary": "Sync failed; verified current source instead.", + "symbols": [{"path": "src/a.py", "line": 1, "name": "A"}], + "source_fallbacks": [{ + "action": "SEARCH_SOURCE", + "path": None, + "observed_sha256": None, + "base_commit": None, + "head_commit": None, + "diff_hash": None, + "purpose": "verify A in current source", + "result_paths": [{ + "path": "src/a.py", + "observed_sha256": file_sha256(source), + }], + }], + }, + ROOT, + ) + recorded = json.loads(Path(result["path"]).read_text(encoding="utf-8")) + self.assertEqual(recorded["delivery"]["state"], "STALE") + self.assertIn( + "SYNC_FAILED", + [point["reason"] for point in recorded["delivery"]["stale_points"]], + ) + + def test_v3_record_preserves_stale_unknown_and_unavailable_restrictions(self) -> None: + cases = [ + ("stale", "STALE", "USED"), + ("unknown", "UNKNOWN", "FAILED"), + ("unavailable", "UNAVAILABLE", "UNAVAILABLE"), + ] + for index, (case, delivery_state, record_status) in enumerate(cases, start=1): + with self.subTest(case=case): + if index > 1: + self.tearDown() + self.setUp() + self.qualify_task() + source = self.repo / "src/a.py" + source.parent.mkdir() + source.write_text("class A:\n pass\n", encoding="utf-8") + if case != "unavailable": + (self.repo / ".codegraph").mkdir() + if case == "stale": + pending = json.loads(healthy_status(self.repo)) + pending["pendingChanges"]["modified"] = 1 + responses = [ + completed(json.dumps(pending)), + completed("A is defined in src/a.py\n"), + completed(json.dumps(pending)), + ] + elif case == "unknown": + responses = [completed("not-json\n")] + else: + responses = [] + + def runner(command: list[str], **_kwargs: object) -> subprocess.CompletedProcess[str]: + if not responses: + raise AssertionError(f"unexpected provider call: {command}") + return responses.pop(0) + + proxy = self.proxy_module() + with mock.patch( + "internal.code_intelligence_proxy.shutil.which", + return_value="/bin/codegraph", + ): + query = proxy.execute_proxy_query( + self.repo, + "TASK-0001", + "PLANNING", + "CIQ-001", + "locate A conservatively", + "symbol A", + False, + runner=runner, + ) + fallback = { + "action": "SEARCH_SOURCE", + "path": None, + "observed_sha256": None, + "base_commit": None, + "head_commit": None, + "diff_hash": None, + "purpose": "verify the graph conclusion in current source", + "result_paths": [{ + "path": "src/a.py", + "observed_sha256": file_sha256(source), + }], + } + protocol = importlib.import_module("internal.code_intelligence_protocol") + result = protocol.record_proxy_bundle( + self.repo, + "TASK-0001", + query["bundle_path"], + { + "summary": "Used source fallback.", + "symbols": [], + "source_fallbacks": [fallback], + }, + ROOT, + ) + recorded = json.loads(Path(result["path"]).read_text(encoding="utf-8")) + self.assertEqual(recorded["delivery"]["state"], delivery_state) + self.assertEqual(recorded["status"], record_status) + self.assertEqual(recorded["delivery"]["usage"], ( + "NO_GRAPH" if case == "unavailable" else "NAVIGATION_ONLY" + )) + self.assertEqual( + recorded["source_fallbacks"][0]["action"], "SEARCH_SOURCE" + ) + + def test_v2_schema_is_frozen_for_historical_reads(self) -> None: + frozen = ROOT / "schemas/code-intelligence-record-v2.schema.json" + self.assertTrue(frozen.is_file()) + self.assertEqual(json.loads(frozen.read_text(encoding="utf-8"))["title"], + "Polaris Code Intelligence record v2") + + self.initialize_task() + value = self.v2_record() + path = self.repo / ".polaris/tasks/TASK-0001/code-intelligence/r001/planning.json" + write_json_atomic(path, value) + original = path.read_bytes() + protocol = importlib.import_module("internal.code_intelligence_protocol") + validated = protocol.validate_historical_v2_record_value( + self.repo, "TASK-0001", path, value, ROOT + ) + self.assertEqual(validated["record_version"], 2) + self.assertEqual(path.read_bytes(), original) + def set_protocol_version(self, version: str) -> None: project_path = self.repo / ".polaris/project.json" project = json.loads(project_path.read_text(encoding="utf-8")) @@ -322,6 +1687,98 @@ def set_workflow_version(self, version: str) -> None: "".join(json.dumps(event, separators=(",", ":")) + "\n" for event in events), ) + def prepare_v2_migration_records(self) -> list[tuple[Path, bytes]]: + """Create immutable v2 records in the current and prior revision slots.""" + self.initialize_task() + new_revision(self.repo, "TASK-0001") + task = self.repo / ".polaris/tasks/TASK-0001" + state_path = task / "state.json" + state = json.loads(state_path.read_text(encoding="utf-8")) + state["current_revision"] = 2 + write_json_atomic(state_path, state) + event_path = task / "events.jsonl" + events = [ + json.loads(line) + for line in event_path.read_text(encoding="utf-8").splitlines() + ] + events[0]["current_revision"] = 2 + write_text_atomic( + event_path, + "".join( + json.dumps(event, separators=(",", ":")) + "\n" + for event in events + ), + ) + frozen: list[tuple[Path, bytes]] = [] + for revision in (1, 2): + value = self.v2_record() + value["work_item_revision"] = revision + path = ( + task + / "code-intelligence" + / f"r{revision:03d}" + / "planning.json" + ) + write_json_atomic(path, value) + frozen.append((path, path.read_bytes())) + self.set_protocol_version("0.1.20") + self.set_workflow_version("0.1.3") + refresh_project_index(self.repo) + return frozen + + def test_migration_inventories_frozen_v2_records_without_rewriting_them( + self, + ) -> None: + """0.1.21 inventories current/prior v2 evidence and preserves Workflow 0.1.3.""" + frozen = self.prepare_v2_migration_records() + vendor(ROOT, self.repo, False) + + result = migrate_project(self.repo) + + self.assertEqual(result["from"], "0.1.20") + self.assertEqual(result["to"], "0.1.21") + project = json.loads( + (self.repo / ".polaris/project.json").read_text(encoding="utf-8") + ) + self.assertEqual(project["workflow_version"], "0.1.3") + migration = json.loads( + ( + self.repo + / ".polaris/migrations/MIG-0.1.20-to-0.1.21.json" + ).read_text(encoding="utf-8") + ) + self.assertEqual( + migration["retired_code_intelligence_records"], + [ + { + "task_id": "TASK-0001", + "path": f"code-intelligence/r{revision:03d}/planning.json", + "sha256": hashlib.sha256(content).hexdigest(), + } + for revision, (_path, content) in enumerate(frozen, start=1) + ], + ) + for path, content in frozen: + self.assertEqual(path.read_bytes(), content) + + def test_migration_resume_rejects_mutated_frozen_v2_inventory(self) -> None: + """中断迁移重跑前会重算 v2 清单,拒绝已经变化的历史证据。""" + frozen = self.prepare_v2_migration_records() + vendor(ROOT, self.repo, False) + with mock.patch( + "internal.migration_protocol.append_jsonl", + side_effect=OSError("injected migration interruption"), + ): + with self.assertRaisesRegex(OSError, "injected migration interruption"): + migrate_project(self.repo) + + path, _content = frozen[0] + value = json.loads(path.read_text(encoding="utf-8")) + value["recorded_at"] = "2026-08-19T00:00:00Z" + write_json_atomic(path, value) + with self.assertRaisesRegex(RuleFailure, "inventory changed"): + migrate_project(self.repo) + def test_legacy_v1_records_remain_readable_but_cannot_be_written(self) -> None: self.initialize_task() protocol = importlib.import_module("internal.code_intelligence_protocol") @@ -360,7 +1817,7 @@ def test_legacy_v1_records_remain_readable_but_cannot_be_written(self) -> None: } self.assertEqual(legacy_validator(self.repo, "TASK-0001", value, ROOT)["record_version"], 1) with self.assertRaisesRegex( - InputFailure, "new Code Intelligence records must use record_version 2" + InputFailure, "new Code Intelligence records must use record_version 3" ): record(self.repo, "TASK-0001", value, ROOT) @@ -401,7 +1858,8 @@ def test_migration_retires_v1_records_without_rewriting_them(self) -> None: } write_json_atomic(legacy_path, legacy) legacy_bytes = legacy_path.read_bytes() - vendor(ROOT, self.repo, False) + with protocol_source_at("0.1.20") as source: + vendor(source, self.repo, False) result = migrate_project(self.repo) @@ -430,11 +1888,10 @@ def test_migration_retires_v1_records_without_rewriting_them(self) -> None: current = self.v2_record() current.update({"stage": "IMPLEMENTATION", "artifact_attempt": 1}) - result = record(self.repo, "TASK-0001", current, ROOT) - self.assertEqual( - Path(result["path"]).parts[-3:], - ("code-intelligence", "r001", "implementation-001.json"), - ) + with self.assertRaisesRegex( + InputFailure, "new Code Intelligence records must use record_version 3" + ): + record(self.repo, "TASK-0001", current, ROOT) def test_migration_rejects_noncanonical_v2_record_paths(self) -> None: """Migration scans only the canonical Code Intelligence record layout.""" @@ -446,7 +1903,8 @@ def test_migration_rejects_noncanonical_v2_record_paths(self) -> None: / ".polaris/tasks/TASK-0001/code-intelligence/r001/not-a-stage.json" ) write_json_atomic(noncanonical, self.v2_record()) - vendor(ROOT, self.repo, False) + with protocol_source_at("0.1.20") as source: + vendor(source, self.repo, False) with self.assertRaisesRegex(RuleFailure, "non-canonical"): migrate_project(self.repo) @@ -503,7 +1961,8 @@ def test_migration_inventories_v1_records_from_prior_revisions(self) -> None: RuleFailure, "targets the wrong task revision" ): validate_record_value(self.repo, "TASK-0001", legacy, ROOT) - vendor(ROOT, self.repo, False) + with protocol_source_at("0.1.20") as source: + vendor(source, self.repo, False) try: migrate_project(self.repo) @@ -534,7 +1993,8 @@ def test_migration_rejects_a_dangling_code_intelligence_symlink(self) -> None: records_root = self.repo / ".polaris/tasks/TASK-0001/code-intelligence" shutil.rmtree(records_root) records_root.symlink_to(self.repo / "missing-code-intelligence") - vendor(ROOT, self.repo, False) + with protocol_source_at("0.1.20") as source: + vendor(source, self.repo, False) with self.assertRaisesRegex(RuleFailure, "must not be a symlink"): migrate_project(self.repo) @@ -1515,45 +2975,19 @@ def test_official_descriptor_uses_explore_status_and_sync(self) -> None: ) self.assertEqual(descriptor["cli"]["sync_args"], ["sync", "--quiet"]) - def test_all_agent_surfaces_share_codegraph_fallback_rules(self) -> None: - """Stage instructions keep CodeGraph stale-data fallbacks identical per host.""" - required_fragments = ( - ".codegraph/", - "codegraph_explore", - "codegraph explore", - "codegraph sync", - "PARTIAL_STALE", - "INDEX_STALE", - "directly read", - "never run `codegraph init`", - ) - partial_stale_branches = ( - "current confined regular file", - "READ_SOURCE", - "current SHA-256", - "missing/deleted", - "INSPECT_GIT_DIFF", - "null observed SHA-256", - "base/head/diff evidence", - "unsafe paths", - "NOT_VERIFIED", - "source search", - ) - audit_binding_fragments = ( - "freshness.response_sha256", - "successful explore response", - "result_paths", - "at most 100", - "POSIX", - "current confined regular file", - "empty `result_paths`", - ) - retired_operations = ( - "symbol" + "_search", - "call" + "_graph", - "review" + "_context", - "refresh" + "_files", - "refresh" + "_workspace", + def test_all_agent_surfaces_require_proxy_provenance(self) -> None: + """Every CodeGraph-capable stage requires the proxy envelope and fallbacks.""" + anchors = ( + "polaris_codegraph_explore", + "freshness envelope", + "NON_AUTHORITATIVE_CONTEXT", + "NAVIGATION_ONLY", + "never substantiates", + "source/Git fallback", + "no separate status/sync MCP tool", + "do not retry", + "raw `codegraph_explore`", + "cannot back `CURRENT` Polaris evidence", ) stage_skills = ( "code-intelligence", @@ -1563,6 +2997,15 @@ def test_all_agent_surfaces_share_codegraph_fallback_rules(self) -> None: "documentation-sync", ) available_skills = set(discover_skills(ROOT)) + + def assert_contract(text: str, label: str) -> None: + for anchor in anchors: + self.assertIn(anchor, text, f"{label}: {anchor}") + self.assertIn("record_code_intelligence.py", text, label) + self.assertIn("--bundle", text, label) + self.assertIn("--annotations", text, label) + self.assertIn("v3", text, label) + for adapter in load_host_adapters(ROOT): for skill_name in stage_skills: source = (ROOT / "skills" / skill_name / "SKILL.md").read_text( @@ -1571,30 +3014,59 @@ def test_all_agent_surfaces_share_codegraph_fallback_rules(self) -> None: rendered = render_skill( source, skill_name, adapter, available_skills ) - for fragment in required_fragments: - self.assertIn(fragment, rendered, f"{adapter['host_id']}:{skill_name}") - for fragment in partial_stale_branches: - self.assertIn(fragment, rendered, f"{adapter['host_id']}:{skill_name}") - for fragment in audit_binding_fragments: - self.assertIn(fragment, rendered, f"{adapter['host_id']}:{skill_name}") - self.assertIn("v2", rendered, f"{adapter['host_id']}:{skill_name}") - self.assertNotIn("directly read every listed stale file", rendered) - for retired in retired_operations: - self.assertNotIn(retired, rendered, f"{adapter['host_id']}:{skill_name}") - - validation = (ROOT / "skills" / "validation" / "SKILL.md").read_text( - encoding="utf-8" - ) - self.assertIn("Do not invoke Code Intelligence", validation) + assert_contract(rendered, f"{adapter['host_id']}:{skill_name}") agents = (ROOT / "templates" / "AGENTS.md").read_text(encoding="utf-8") - self.assertIn("stop CodeGraph calls for this session", agents) - self.assertIn("installer-managed marker block", agents) - for fragment in partial_stale_branches: - self.assertIn(fragment, agents) - for fragment in audit_binding_fragments: - self.assertIn(fragment, agents) - self.assertNotIn("directly read every listed stale file", agents) + assert_contract(agents, "templates/AGENTS.md") + + rendered = render_skill( + (ROOT / "skills/implementation/SKILL.md").read_text(encoding="utf-8"), + "implementation", + load_host_adapters(ROOT)[0], + available_skills, + ) + for mutation in ( + rendered.replace("polaris_codegraph_explore", "missing_proxy"), + rendered.replace("source/Git fallback", "missing fallback"), + ): + with self.assertRaises(AssertionError): + assert_contract(mutation, "mutated implementation") + + def test_documentation_sync_uses_one_proxy_query(self) -> None: + """Documentation Sync uses one bounded changed-path/symbol proxy query.""" + source = (ROOT / "skills/documentation-sync/SKILL.md").read_text( + encoding="utf-8" + ) + for adapter in load_host_adapters(ROOT): + rendered = render_skill( + source, + "documentation-sync", + adapter, + set(discover_skills(ROOT)), + ) + for anchor in ( + "polaris_codegraph_explore", + "sync_if_needed: true", + "changed source paths", + "documented symbols", + "no separate status/sync MCP tool", + ): + self.assertIn(anchor, rendered, f"{adapter['host_id']}: {anchor}") + + def test_validation_remains_graph_free(self) -> None: + """Validation never invokes the proxy or raw CodeGraph lifecycle commands.""" + source = (ROOT / "skills/validation/SKILL.md").read_text(encoding="utf-8") + self.assertIn("Do not invoke Code Intelligence", source) + for adapter in load_host_adapters(ROOT): + rendered = render_skill( + source, "validation", adapter, set(discover_skills(ROOT)) + ) + for forbidden in ( + "polaris_codegraph_explore", + "codegraph status", + "codegraph sync", + ): + self.assertNotIn(forbidden, rendered, adapter["host_id"]) def test_provider_requires_marker_and_accepts_mcp_or_cli(self) -> None: self.assertIsNone( @@ -1620,7 +3092,7 @@ def test_project_marker_rejects_unsafe_paths_and_symlinks(self) -> None: with self.assertRaisesRegex(RuleFailure, "must not be a symlink"): _project_marker_path(self.repo, ".codegraph") - def test_record_cli_requires_task_id_and_input(self) -> None: + def test_record_cli_requires_task_id_bundle_and_annotations(self) -> None: completed = subprocess.run( [ sys.executable, @@ -1638,7 +3110,10 @@ def test_record_cli_requires_task_id_and_input(self) -> None: self.assertEqual(completed.returncode, 2) payload = json.loads(completed.stdout) self.assertEqual(payload["status"], "ERROR") - self.assertEqual(payload["message"], "recording requires task_id and --input") + self.assertEqual( + payload["message"], + "recording requires task_id, --bundle, and --annotations", + ) def test_old_product_tool_names_are_absent_from_descriptor(self) -> None: text = (ROOT / "providers/code-intelligence/codegraph.json").read_text( @@ -1694,6 +3169,78 @@ def runner( self.assertEqual(result["freshness"]["status"], "CURRENT_AT_CHECK") self.assertIn("SYNC_ACKNOWLEDGED", result["freshness"]["basis"]) + def test_pending_changes_still_sync_once_with_an_index_stale_reason(self) -> None: + _, sync_if_needed = self.adapter_functions() + (self.repo / ".codegraph").mkdir() + pending = json.loads(healthy_status(self.repo)) + pending["pendingChanges"]["modified"] = 1 + pending["index"]["state"] = "partial" + responses = iter([ + completed(json.dumps(pending)), + completed("Synced 1 changed file\n"), + completed(healthy_status(self.repo)), + ]) + calls: list[list[str]] = [] + + def runner( + command: list[str], **_kwargs: object + ) -> subprocess.CompletedProcess[str]: + calls.append(command) + return next(responses) + + result = sync_if_needed( + self.repo, load_providers(ROOT)["codegraph"], runner=runner + ) + + self.assertEqual([call[1] for call in calls], ["status", "sync", "status"]) + self.assertEqual(result["sync"]["status"], "SUCCESS") + self.assertEqual(result["freshness"]["status"], "CURRENT_AT_CHECK") + + def test_explore_and_observed_sync_are_bounded_to_one_repo(self) -> None: + adapter = self.adapter_module() + (self.repo / ".codegraph").mkdir() + calls: list[tuple[list[str], Path]] = [] + + def runner( + command: list[str], **kwargs: object + ) -> subprocess.CompletedProcess[str]: + cwd = kwargs["cwd"] + self.assertIsInstance(cwd, Path) + calls.append((command, cwd)) + if command[1:3] == ["status", "--json"]: + return completed(healthy_status(self.repo)) + if command[1:] == ["sync", "--quiet"]: + return completed("synced\n") + return completed("graph response\n") + + pending = json.loads(healthy_status(self.repo)) + pending["pendingChanges"]["modified"] = 1 + initial = adapter._status_result( + self.repo, pending, "2026-08-19T00:00:00Z", "a" * 64 + ) + synchronized = adapter.synchronize_observed_status( + self.repo, + load_providers(ROOT)["codegraph"], + initial, + runner=runner, + ) + explored = adapter.run_explore( + self.repo, + load_providers(ROOT)["codegraph"], + "find symbol A", + runner=runner, + ) + + self.assertEqual(synchronized["sync"]["status"], "SUCCESS") + self.assertEqual(explored["status"], "SUCCESS") + self.assertEqual( + explored["response_sha256"], + hashlib.sha256(b"graph response\n").hexdigest(), + ) + self.assertTrue(all(cwd == self.repo for _command, cwd in calls)) + self.assertEqual(sum(command[1] == "sync" for command, _cwd in calls), 1) + self.assertEqual(sum(command[1] == "explore" for command, _cwd in calls), 1) + def test_index_wide_stale_reasons_do_not_sync(self) -> None: inspect_status, sync_if_needed = self.adapter_functions() (self.repo / ".codegraph").mkdir() @@ -2101,29 +3648,35 @@ def test_response_banner_rejects_unsafe_windows_style_paths(self) -> None: result["stale_points"][0]["reason"], "STATUS_UNREADABLE" ) - def test_arbitrary_warning_is_not_a_codegraph_banner(self) -> None: - result = self.classify_response("⚠️ maybe stale: src/widget.py\n") - - self.assertEqual(result["classification"], "NONE") - self.assertEqual(result["stale_points"], []) - - def test_prefixed_or_quoted_official_banner_is_not_recognized(self) -> None: + def test_suspicious_or_wrapped_freshness_warnings_are_not_verified(self) -> None: banner = """⚠️ Some files referenced below were edited since the last index sync — their codegraph entries may be stale: - src/deleted.py (edited 800ms ago, pending sync) For accurate content of those specific files, Read them directly. """ - for response in ( + samples = ( + "warning: graph may be stale\n", + "⚠️ maybe stale: src/widget.py\n", + "quoted: ⚠️ CodeGraph auto-sync is DISABLED — the index is frozen.\n", + "\ufeff⚠️ CodeGraph auto-sync is DISABLED — the index is frozen.\n", + " pending-sync required\n", f"context before banner\n{banner}", f"> {banner}", f"quoted response: {banner}", f" {banner}", f"\ufeff{banner}", - ): - with self.subTest(response=response[:20]): + ) + for response in samples: + with self.subTest(response=response[:30]): result = self.classify_response(response) - self.assertEqual(result["classification"], "NONE") - self.assertEqual(result["stale_points"], []) + self.assertEqual(result["classification"], "NOT_VERIFIED") + self.assertEqual( + result["stale_points"][0]["reason"], "STATUS_UNREADABLE" + ) + self.assertEqual( + result["response_sha256"], + hashlib.sha256(response.encode("utf-8")).hexdigest(), + ) def test_merge_freshness_uses_conservative_status_and_ordered_evidence(self) -> None: merger = getattr(self.adapter_module(), "merge_freshness", None) diff --git a/tests/test_core.py b/tests/test_core.py index e00dfd4..0303c29 100644 --- a/tests/test_core.py +++ b/tests/test_core.py @@ -29,10 +29,12 @@ add_provider, load_config, record as record_code_intelligence, + record_proxy_bundle, select_provider, validate_record_value, validate_static_configuration, ) +from internal.code_intelligence_proxy import execute_proxy_query # noqa: E402 from internal.host_adapters import ( # noqa: E402 discover_skills, load_host_adapters, @@ -130,6 +132,20 @@ def repository_file_snapshot(repo: Path) -> dict[str, bytes]: } +@contextmanager +def protocol_source_at(version: str) -> Iterator[Path]: + """Materialize a historical protocol target for adjacent migration tests.""" + with tempfile.TemporaryDirectory(prefix="polaris-protocol-source-") as temp: + source = Path(temp) / "source" + shutil.copytree( + ROOT, + source, + ignore=shutil.ignore_patterns(".git", "__pycache__", "*.pyc"), + ) + (source / "VERSION").write_text(version + "\n", encoding="utf-8") + yield source + + @contextmanager def simulated_symlinks(*paths: Path) -> Iterator[None]: """Report selected paths as symlinks without requiring filesystem support.""" @@ -346,6 +362,53 @@ def enter_implementing(self) -> None: self.enter_planned() self.register_implementation_handoff() + def record_current_proxy_intelligence(self, stage: str) -> dict[str, object]: + """Create genuine current v3 evidence for an active implementation stage.""" + (self.repo / ".codegraph").mkdir(exist_ok=True) + status = json.dumps({ + "initialized": True, + "projectPath": str(self.repo.resolve()), + "pendingChanges": {"added": 0, "modified": 0, "removed": 0}, + "worktreeMismatch": None, + "index": { + "state": "complete", + "pendingRefs": 0, + "reindexRecommended": False, + }, + }) + responses = [ + subprocess.CompletedProcess([], 0, status, ""), + subprocess.CompletedProcess([], 0, "graph context\n", ""), + subprocess.CompletedProcess([], 0, status, ""), + ] + + def runner( + _command: list[str], **_kwargs: object + ) -> subprocess.CompletedProcess[str]: + return responses.pop(0) + + with mock.patch( + "internal.code_intelligence_proxy.shutil.which", + return_value="/bin/codegraph", + ): + query = execute_proxy_query( + self.repo, + "TASK-0001", + stage, + "CIQ-001", + "bind final subject", + "final subject symbols", + False, + runner=runner, + ) + return record_proxy_bundle( + self.repo, + "TASK-0001", + query["bundle_path"], + {"summary": "Current graph context.", "symbols": [], "source_fallbacks": []}, + ROOT, + ) + def register_implementation_handoff(self) -> dict[str, object]: """Register the next deterministic handoff for initial work or rework.""" handoff = build_implementation_handoff(self.repo, "TASK-0001") @@ -2139,7 +2202,8 @@ def test_explicit_migration_appends_task_event_and_records_completion(self) -> N """相邻版本迁移追加审计事件,不改写任务历史,并留下完成记录。""" self.set_protocol_version("0.1.19") self.set_workflow_version("0.1.2") - vendor(ROOT, self.repo, False) + with protocol_source_at("0.1.20") as source: + vendor(source, self.repo, False) result = migrate_project(self.repo) @@ -2291,7 +2355,8 @@ def test_migration_replaces_frozen_workflow_and_maps_tasks(self) -> None: for event in events ), ) - vendor(ROOT, self.repo, False) + with protocol_source_at("0.1.20") as source: + vendor(source, self.repo, False) result = migrate_project(self.repo) @@ -2368,7 +2433,8 @@ def test_migrated_r1_verified_task_can_close(self) -> None: ) self.set_protocol_version("0.1.19") self.set_workflow_version("0.1.2") - vendor(ROOT, self.repo, False) + with protocol_source_at("0.1.20") as source: + vendor(source, self.repo, False) migrate_project(self.repo) @@ -2406,7 +2472,8 @@ def test_migration_resumes_after_event_append_without_duplication(self) -> None: """中断后重跑会采用已追加的迁移事件并完成投影,不重复写事件。""" self.set_protocol_version("0.1.19") self.set_workflow_version("0.1.2") - vendor(ROOT, self.repo, False) + with protocol_source_at("0.1.20") as source: + vendor(source, self.repo, False) state = read_json(self.task / "state.json") started_at = "2026-08-15T00:00:00Z" record = { @@ -2470,7 +2537,8 @@ def test_migration_reclaims_only_its_own_dead_process_lock(self) -> None: """迁移可接管同一迁移的崩溃锁,但不能抢占仍存活的进程。""" self.set_protocol_version("0.1.19") self.set_workflow_version("0.1.2") - vendor(ROOT, self.repo, False) + with protocol_source_at("0.1.20") as source: + vendor(source, self.repo, False) lock_path = self.task / ".transition.lock" write_json_atomic( lock_path, @@ -2510,8 +2578,9 @@ def test_migration_reclaims_only_its_own_dead_process_lock(self) -> None: def test_migration_rejects_an_undeclared_version_jump(self) -> None: """没有注册的跨版本路径机械拒绝,且不创建部分迁移记录。""" - self.set_protocol_version("0.1.10") - vendor(ROOT, self.repo, False) + self.set_protocol_version("0.1.20") + with protocol_source_at("0.1.22") as source: + vendor(source, self.repo, False) with self.assertRaisesRegex(RuleFailure, "no explicit adjacent migration"): migrate_project(self.repo) @@ -2519,9 +2588,24 @@ def test_migration_rejects_an_undeclared_version_jump(self) -> None: self.assertFalse((self.repo / ".polaris" / "migrations").exists()) self.assertEqual( read_json(self.repo / ".polaris" / "project.json")["polaris_version"], - "0.1.10", + "0.1.20", ) + def test_version_only_migration_rejects_a_workflow_version_change(self) -> None: + """0.1.20→0.1.21 路由不能暗中改变已冻结的 Workflow 0.1.3。""" + self.set_protocol_version("0.1.20") + with protocol_source_at("0.1.21") as source: + migrations_path = source / "workflow" / "migrations.json" + migrations = read_json(migrations_path) + migrations["steps"][-1]["to_workflow_version"] = "0.1.4" + write_json_atomic(migrations_path, migrations) + vendor(source, self.repo, False) + + with self.assertRaisesRegex( + RuleFailure, "workflow migration requires replacement" + ): + migrate_project(self.repo) + def test_code_intelligence_auto_detects_available_operations_and_can_be_disabled(self) -> None: """已初始化的可选代码情报按 MCP 工具能力发现;缺失或禁用时不产生硬依赖。""" (self.repo / ".codegraph").mkdir() @@ -2554,7 +2638,7 @@ def test_code_intelligence_v1_record_is_read_only_historical_evidence(self) -> N """v1 精简记录升级后仍可读取,并绑定任务、提交和安全路径。""" base = run_git(self.repo, "rev-parse", "HEAD") value = read_json( - ROOT / "templates" / "task-sources" / "code-intelligence-record.json" + ROOT / "tests" / "fixtures" / "code-intelligence-record-v2.json" ) value["record_version"] = 1 value.pop("sync") @@ -2567,7 +2651,7 @@ def test_code_intelligence_v1_record_is_read_only_historical_evidence(self) -> N validate_record_value(self.repo, "TASK-0001", value, ROOT)["status"], "UNAVAILABLE", ) - with self.assertRaisesRegex(InputFailure, "record_version 2"): + with self.assertRaisesRegex(InputFailure, "record_version 3"): record_code_intelligence(self.repo, "TASK-0001", value, ROOT) invalid = copy.deepcopy(value) @@ -2692,6 +2776,8 @@ def test_risk_flag_requires_r2(self) -> None: def test_vendored_target_is_self_contained(self) -> None: """目标仓库 vendoring 后同时包含 Codex、Claude Code 与机械协议。""" + self.assertFalse((self.repo / ".codex" / "config.toml").exists()) + self.assertFalse((self.repo / ".mcp.json").exists()) vendor(ROOT, self.repo, False) for adapter in load_host_adapters(ROOT): skill_root = self.repo / str(adapter["skill_target"]) @@ -2723,6 +2809,46 @@ def test_vendored_target_is_self_contained(self) -> None: result = validate_project(self.repo) self.assertEqual(result["active_tasks"], 1) + def test_vendor_preserves_and_validates_project_mcp_configuration(self) -> None: + """vendoring 注册项目代理,把宿主配置列为保留文件并校验启动边界。""" + from internal.project_mcp_registration import validate_project_mcp + + codex_path = self.repo / ".codex" / "config.toml" + codex_path.parent.mkdir() + codex_path.write_text( + 'model = "gpt-5"\n[mcp_servers.other]\ncommand = "other"\n', + encoding="utf-8", + ) + claude_path = self.repo / ".mcp.json" + write_json_atomic( + claude_path, + { + "permissions": {"allow": ["Read"]}, + "mcpServers": {"other": {"command": "other", "args": []}}, + }, + ) + + vendor(ROOT, self.repo, False) + + adapters = load_host_adapters(self.repo / "tools" / "polaris") + for adapter in adapters: + validate_project_mcp(self.repo, adapter) + self.assertIn('model = "gpt-5"', codex_path.read_text(encoding="utf-8")) + claude = read_json(claude_path) + self.assertEqual(claude["permissions"], {"allow": ["Read"]}) + self.assertIn("other", claude["mcpServers"]) + manifest = read_json( + self.repo / "tools" / "polaris" / "install-manifest.json" + ) + self.assertIn(".codex/config.toml", manifest["preserved_files"]) + self.assertIn(".mcp.json", manifest["preserved_files"]) + self.assertEqual(validate_project(self.repo)["active_tasks"], 1) + + claude["mcpServers"]["polaris-codegraph"]["args"][-1] = "../other" + write_json_atomic(claude_path, claude) + with self.assertRaisesRegex(RuleFailure, "Polaris entry is invalid"): + validate_project(self.repo) + def test_doctor_reports_a_healthy_vendored_project_without_writing(self) -> None: """Doctor 聚合健康检查并通过报告 Schema,且诊断前后项目文件完全不变。""" vendor(ROOT, self.repo, False) @@ -2967,8 +3093,12 @@ def test_vendor_rolls_back_after_partial_apply_failure(self) -> None: vendor(ROOT, self.repo, False) manifest_path = self.repo / "tools" / "polaris" / "install-manifest.json" skill_path = self.repo / ".agents" / "skills" / "engineering-task" / "SKILL.md" + codex_path = self.repo / ".codex" / "config.toml" + claude_mcp_path = self.repo / ".mcp.json" original_manifest = manifest_path.read_bytes() original_skill = skill_path.read_bytes() + original_codex = codex_path.read_bytes() + original_claude_mcp = claude_mcp_path.read_bytes() with tempfile.TemporaryDirectory(prefix="polaris-vendor-source-") as temp: source = Path(temp) / "source" shutil.copytree( @@ -2999,6 +3129,8 @@ def fail_after_copy(staged: Path, destination: Path) -> None: self.assertEqual(manifest_path.read_bytes(), original_manifest) self.assertEqual(skill_path.read_bytes(), original_skill) + self.assertEqual(codex_path.read_bytes(), original_codex) + self.assertEqual(claude_mcp_path.read_bytes(), original_claude_mcp) self.assertEqual(validate_project(self.repo)["active_tasks"], 1) self.assertEqual( list( @@ -3183,6 +3315,36 @@ def test_host_adapters_render_from_one_host_neutral_skill_source(self) -> None: adapters = {item["host_id"]: item for item in load_host_adapters(ROOT)} self.assertEqual(set(adapters), {"codex", "claude-code"}) + self.assertIn("project_mcp", adapters["codex"]) + self.assertIn("project_mcp", adapters["claude-code"]) + self.assertEqual( + adapters["codex"]["project_mcp"], + { + "server_id": "polaris-codegraph", + "format": "codex-toml", + "target": ".codex/config.toml", + "command": "python3", + "args": [ + "tools/polaris/scripts/code_intelligence_mcp.py", + "--repo", + ".", + ], + }, + ) + self.assertEqual( + adapters["claude-code"]["project_mcp"], + { + "server_id": "polaris-codegraph", + "format": "claude-json", + "target": ".mcp.json", + "command": "python3", + "args": [ + "tools/polaris/scripts/code_intelligence_mcp.py", + "--repo", + ".", + ], + }, + ) self.assertFalse((ROOT / "hosts" / "codex" / "skills").exists()) codex = render_skill(source, "engineering-task", adapters["codex"]) claude = render_skill(source, "engineering-task", adapters["claude-code"]) @@ -3217,7 +3379,7 @@ def test_host_adapter_contract_rejects_invalid_or_conflicting_manifests(self) -> def adapter(host_id: str) -> dict[str, object]: return { - "adapter_version": 2, + "adapter_version": 3, "host_id": host_id, "display_name": host_id, "skill_target": f".{host_id}/skills", @@ -3233,6 +3395,17 @@ def adapter(host_id: str) -> dict[str, object]: "entry_frontmatter": [], "skill_overlay_root": None, "skill_appendix_root": None, + "project_mcp": { + "server_id": "polaris-codegraph", + "format": "claude-json", + "target": f".{host_id}/mcp.json", + "command": "python3", + "args": [ + "tools/polaris/scripts/code_intelligence_mcp.py", + "--repo", + ".", + ], + }, "files": [ { "source": "bridge.md", @@ -3244,7 +3417,7 @@ def adapter(host_id: str) -> dict[str, object]: cases = { "unknown version": lambda first, _second: first.update( - {"adapter_version": 3} + {"adapter_version": 4} ), "blank prefix": lambda first, _second: first.update( {"invocation_prefix": ""} @@ -3255,6 +3428,27 @@ def adapter(host_id: str) -> dict[str, object]: "overlapping target": lambda first, second: second["files"][0].update( {"target": first["skill_target"]} ), + "unknown MCP format": lambda first, _second: first["project_mcp"].update( + {"format": "yaml"} + ), + "wrong MCP server": lambda first, _second: first["project_mcp"].update( + {"server_id": "other"} + ), + "unsafe MCP target": lambda first, _second: first["project_mcp"].update( + {"target": "../config.json"} + ), + "wrong MCP launcher": lambda first, _second: first["project_mcp"].update( + {"args": ["scripts/code_intelligence_mcp.py", "--repo", "."]} + ), + "missing fixed repo": lambda first, _second: first["project_mcp"].update( + {"args": ["tools/polaris/scripts/code_intelligence_mcp.py"]} + ), + "duplicate MCP target": lambda first, second: second["project_mcp"].update( + {"target": first["project_mcp"]["target"]} + ), + "MCP overlaps skill": lambda first, _second: first["project_mcp"].update( + {"target": first["skill_target"]} + ), } for name, mutate in cases.items(): with self.subTest(case=name), tempfile.TemporaryDirectory( @@ -3280,6 +3474,221 @@ def adapter(host_id: str) -> dict[str, object]: with self.assertRaises(RuleFailure): load_host_adapters(root) + def test_project_mcp_registration_preserves_unrelated_host_configuration( + self, + ) -> None: + """项目 MCP 合并只管理 Polaris 条目,并且重复执行保持稳定。""" + from internal.project_mcp_registration import merge_project_mcp + + adapters = {item["host_id"]: item for item in load_host_adapters(ROOT)} + codex_source = 'model = "gpt-5"\n[mcp_servers.other]\ncommand = "other"\n' + codex_expected = codex_source + """ +# POLARIS_MCP_START polaris-codegraph +[mcp_servers.polaris-codegraph] +command = "python3" +args = ["tools/polaris/scripts/code_intelligence_mcp.py", "--repo", "."] +cwd = "." +enabled = true +required = false +enabled_tools = ["polaris_codegraph_explore"] +# POLARIS_MCP_END polaris-codegraph +""" + codex = merge_project_mcp( + self.repo, adapters["codex"], source_text=codex_source + ) + self.assertEqual(codex, codex_expected) + self.assertEqual( + merge_project_mcp(self.repo, adapters["codex"], source_text=codex), + codex_expected, + ) + stale_managed_block = codex.replace("enabled = true", "enabled = false") + self.assertEqual( + merge_project_mcp( + self.repo, + adapters["codex"], + source_text=stale_managed_block, + ), + codex_expected, + ) + + claude_source = json.dumps( + { + "permissions": {"allow": ["Read"]}, + "mcpServers": {"other": {"command": "other", "args": []}}, + } + ) + claude = merge_project_mcp( + self.repo, adapters["claude-code"], source_text=claude_source + ) + claude_value = json.loads(claude) + self.assertEqual(claude_value["permissions"], {"allow": ["Read"]}) + self.assertEqual( + claude_value["mcpServers"]["other"], + {"command": "other", "args": []}, + ) + self.assertEqual( + claude_value["mcpServers"]["polaris-codegraph"], + { + "type": "stdio", + "command": "python3", + "args": [ + "tools/polaris/scripts/code_intelligence_mcp.py", + "--repo", + ".", + ], + "env": {}, + }, + ) + self.assertEqual( + merge_project_mcp(self.repo, adapters["claude-code"], source_text=claude), + claude, + ) + + def test_project_mcp_registration_imports_without_python_311_tomllib( + self, + ) -> None: + """Python 3.10 can register the managed Codex block without a dependency.""" + valid_source = ''' +[[profiles]] +name = "first" + +[[profiles]] +name = "second" +description = """ +[mcp_servers.not-a-real-table] +""" + +[mcp_servers.polaris-codegraph] +command = "python3" +args = ["tools/polaris/scripts/code_intelligence_mcp.py", "--repo", "."] +cwd = "." +enabled = true +required = false +enabled_tools = ["polaris_codegraph_explore"] +''' + script = f""" +import builtins +import sys + +real_import = builtins.__import__ + +def import_without_tomllib(name, *args, **kwargs): + if name == "tomllib": + raise ModuleNotFoundError("simulated Python 3.10") + return real_import(name, *args, **kwargs) + +builtins.__import__ = import_without_tomllib +sys.path.insert(0, {str(SCRIPTS)!r}) +from internal.project_mcp_registration import _parse_toml +from internal.polaris_core import RuleFailure + +value = _parse_toml({valid_source!r}) +assert [profile["name"] for profile in value["profiles"]] == ["first", "second"] +assert value["profiles"][1]["description"].strip() == "[mcp_servers.not-a-real-table]" +assert value["mcp_servers"]["polaris-codegraph"]["command"] == "python3" +try: + _parse_toml("this is not TOML\\n") +except RuleFailure: + pass +else: + raise AssertionError("malformed TOML was accepted") +""" + completed_process = subprocess.run( + [sys.executable, "-c", script], + cwd=ROOT, + text=True, + capture_output=True, + check=False, + ) + self.assertEqual(completed_process.returncode, 0, completed_process.stderr) + + def test_project_mcp_registration_rejects_unsafe_or_conflicting_configuration( + self, + ) -> None: + """项目 MCP 拒绝损坏配置、非受管同名项与 symlink 目标。""" + from internal.project_mcp_registration import merge_project_mcp + + adapters = {item["host_id"]: item for item in load_host_adapters(ROOT)} + cases = ( + ( + adapters["codex"], + '[mcp_servers.other\ncommand = "broken"\n', + "TOML", + ), + ( + adapters["codex"], + '[mcp_servers.polaris-codegraph]\ncommand = "other"\n', + "conflicting unmanaged", + ), + (adapters["claude-code"], "{broken", "JSON"), + ( + adapters["claude-code"], + json.dumps( + {"mcpServers": {"polaris-codegraph": {"command": "other"}}} + ), + "conflicting", + ), + ) + for adapter, source, message in cases: + with self.subTest(format=adapter["project_mcp"]["format"]): + with self.assertRaisesRegex(RuleFailure, message): + merge_project_mcp(self.repo, adapter, source_text=source) + + target = self.repo / ".mcp.json" + outside = self.repo / "outside-mcp.json" + outside.write_text("{}\n", encoding="utf-8") + try: + target.symlink_to(outside) + except (NotImplementedError, OSError) as exc: + self.skipTest(f"file symlink creation is unavailable: {exc}") + with self.assertRaisesRegex(RuleFailure, "symlink"): + merge_project_mcp(self.repo, adapters["claude-code"]) + + def test_project_mcp_registration_rejects_markers_inside_toml_strings( + self, + ) -> None: + """Managed markers cannot claim or rewrite unrelated multiline strings.""" + from internal.project_mcp_registration import merge_project_mcp + + adapter = {item["host_id"]: item for item in load_host_adapters(ROOT)}[ + "codex" + ] + source = ''' +description = """ +# POLARIS_MCP_START polaris-codegraph +This text belongs to the user. +# POLARIS_MCP_END polaris-codegraph +""" + +[mcp_servers.polaris-codegraph] +command = "python3" +args = ["tools/polaris/scripts/code_intelligence_mcp.py", "--repo", "."] +cwd = "." +enabled = true +required = false +enabled_tools = ["polaris_codegraph_explore"] +''' + with self.assertRaisesRegex(RuleFailure, "unrelated TOML"): + merge_project_mcp(self.repo, adapter, source_text=source) + + def test_project_mcp_registration_preserves_unrelated_toml_nan_values( + self, + ) -> None: + """TOML NaN values remain equivalent across insert and managed updates.""" + from internal.project_mcp_registration import merge_project_mcp + + adapter = {item["host_id"]: item for item in load_host_adapters(ROOT)}[ + "codex" + ] + source = 'metric = nan\n[mcp_servers.other]\ncommand = "other"\n' + inserted = merge_project_mcp(self.repo, adapter, source_text=source) + self.assertTrue(inserted.startswith(source)) + stale = inserted.replace("enabled = true", "enabled = false") + self.assertEqual( + merge_project_mcp(self.repo, adapter, source_text=stale), + inserted, + ) + def test_host_adapter_hardening_rejects_entry_overlay_and_capability_errors(self) -> None: """入口必须存在,overlay 不得覆写 Skill,worker 能力依赖必须自洽。""" cases = { @@ -3422,7 +3831,7 @@ def test_vendor_and_validator_discover_a_third_host_without_code_changes(self) - write_json_atomic( synthetic_root / "adapter.json", { - "adapter_version": 2, + "adapter_version": 3, "host_id": "synthetic", "display_name": "Synthetic Host", "skill_target": ".synthetic/skills", @@ -3438,6 +3847,17 @@ def test_vendor_and_validator_discover_a_third_host_without_code_changes(self) - "entry_frontmatter": [], "skill_overlay_root": None, "skill_appendix_root": None, + "project_mcp": { + "server_id": "polaris-codegraph", + "format": "claude-json", + "target": ".synthetic/mcp.json", + "command": "python3", + "args": [ + "tools/polaris/scripts/code_intelligence_mcp.py", + "--repo", + ".", + ], + }, "files": [], }, ) @@ -3461,6 +3881,16 @@ def test_force_vendor_preserves_unrelated_claude_configuration(self) -> None: (unrelated_skill / "SKILL.md").write_text("# Keep me\n", encoding="utf-8") unrelated_agent.write_text("# Keep me\n", encoding="utf-8") (self.repo / "CLAUDE.md").write_text("# Project-owned Claude rules\n", encoding="utf-8") + codex_config = self.repo / ".codex" / "config.toml" + codex_config.parent.mkdir() + codex_config.write_text('model = "gpt-5"\n', encoding="utf-8") + write_json_atomic( + self.repo / ".mcp.json", + { + "permissions": {"allow": ["Read"]}, + "mcpServers": {"other": {"command": "other", "args": []}}, + }, + ) vendor(ROOT, self.repo, False) vendor(ROOT, self.repo, True) self.assertEqual( @@ -3471,6 +3901,11 @@ def test_force_vendor_preserves_unrelated_claude_configuration(self) -> None: (self.repo / "CLAUDE.md").read_text(encoding="utf-8"), "# Project-owned Claude rules\n", ) + self.assertIn('model = "gpt-5"', codex_config.read_text(encoding="utf-8")) + claude_mcp = read_json(self.repo / ".mcp.json") + self.assertEqual(claude_mcp["permissions"], {"allow": ["Read"]}) + self.assertIn("other", claude_mcp["mcpServers"]) + self.assertIn("polaris-codegraph", claude_mcp["mcpServers"]) def test_validate_project_requires_complete_claude_adapter(self) -> None: """vendored 项目缺少 Claude Skill 或 worker 定义时机械拒绝。""" @@ -4351,22 +4786,8 @@ def test_implementation_and_final_documentation_subjects_are_bound(self) -> None final_head = run_git(self.repo, "rev-parse", "HEAD") final_diff_hash = subject_diff_hash(self.repo, base, final_head) - implementation_intelligence = read_json( - ROOT / "templates" / "task-sources" / "code-intelligence-record.json" - ) - implementation_intelligence.update( - { - "stage": "IMPLEMENTATION", - "artifact_attempt": 1, - "target": { - "base_commit": base, - "head_commit": final_head, - "diff_hash": final_diff_hash, - }, - } - ) - implementation_intelligence_result = record_code_intelligence( - self.repo, "TASK-0001", implementation_intelligence, ROOT + implementation_intelligence_result = self.record_current_proxy_intelligence( + "IMPLEMENTATION" ) implementation_intelligence_path = Path( implementation_intelligence_result["path"] @@ -4374,26 +4795,14 @@ def test_implementation_and_final_documentation_subjects_are_bound(self) -> None implementation_path = self.task / "implementations" / "r001" / "attempt-001.json" implementation = self.implementation_value(base, final_head, "impl-session") implementation["code_intelligence"] = { - "path": implementation_intelligence_path.relative_to(self.task).as_posix(), + "path": implementation_intelligence_path.relative_to( + self.task.resolve() + ).as_posix(), "sha256": file_sha256(implementation_intelligence_path), } write_json_atomic(implementation_path, implementation) - documentation_intelligence = read_json( - ROOT / "templates" / "task-sources" / "code-intelligence-record.json" - ) - documentation_intelligence.update( - { - "stage": "DOCUMENTATION_SYNC", - "artifact_attempt": 1, - "target": { - "base_commit": base, - "head_commit": final_head, - "diff_hash": final_diff_hash, - }, - } - ) - documentation_intelligence_result = record_code_intelligence( - self.repo, "TASK-0001", documentation_intelligence, ROOT + documentation_intelligence_result = self.record_current_proxy_intelligence( + "DOCUMENTATION_SYNC" ) documentation_intelligence_path = Path( documentation_intelligence_result["path"] @@ -4401,7 +4810,9 @@ def test_implementation_and_final_documentation_subjects_are_bound(self) -> None knowledge_path = self.task / "knowledge" / "r001" / "knowledge-delta-001.json" knowledge = self.knowledge_value(1, base, final_head) knowledge["code_intelligence"] = { - "path": documentation_intelligence_path.relative_to(self.task).as_posix(), + "path": documentation_intelligence_path.relative_to( + self.task.resolve() + ).as_posix(), "sha256": file_sha256(documentation_intelligence_path), } knowledge["entries"][0].update( diff --git a/workflow/migrations.json b/workflow/migrations.json index 3582a1b..ab69d72 100644 --- a/workflow/migrations.json +++ b/workflow/migrations.json @@ -99,6 +99,15 @@ "to_workflow_version": "0.1.3", "project_strategy": "replace_version_and_workflow", "task_strategy": "append_mapped_workflow_event" + }, + { + "migration_id": "0.1.20-to-0.1.21", + "from_polaris_version": "0.1.20", + "to_polaris_version": "0.1.21", + "from_workflow_version": "0.1.3", + "to_workflow_version": "0.1.3", + "project_strategy": "replace_version", + "task_strategy": "append_version_event" } ] }