Repository navigation
feat(skill): run the pipeline as graphify pipeline subcommands, no inline Python (#197) - #4275
nothariharan wants to merge 6 commits into
Conversation
Graphify-Labs#197) The agent skill carried each pipeline step as an inline `$(cat graphify-out/.graphify_python) -c "..."` block. That form trips Claude Code's command_substitution and "newline followed by #" approval prompts once per step and cannot be allowlisted, so a run cost ~30 manual confirmations. Add `graphify pipeline <step>` so every step is a plain console command: no inline script, no command substitution, and `Bash(graphify *)` is allowlistable. The step bodies live in graphify.pipeline and are transcriptions of the inline blocks they replace, so a run produces the same graph.json. Regenerate the skill artifacts from the updated core/shell fragments, and move the drift, cross-shell parity, and Graphify-Labs#2490 curated-label guards to follow the flow to the subcommands. The hand-authored aider/devin monoliths keep their inline pipelines for now.
) Follow-up fixes to the graphify pipeline steps, found in review: - detect: pass the directory that CONTAINS graphify-out as cache_root, not graphify-out itself, so the stat cache and Google-Workspace conversions are not written to a nested graphify-out/graphify-out. - merge-chunks: write an empty .graphify_semantic_new.json on zero chunks instead of erroring, so the all-cached path (which routes through B3) still produces the sidecar Part C reads unconditionally (Graphify-Labs#1392). - empty-semantic: return fresh lists per call rather than sharing module state. - cli: honor --no-gitignore (was silently dropped) and thread --force to build/label so the documented Graphify-Labs#479 remediation actually overrides the guard. - cli: reject non-integer --labels keys with a clean error instead of a traceback, and honor --help anywhere after the step. - out dir: honor the GRAPHIFY_OUT env override like the rest of the CLI (Graphify-Labs#686). - restore the @@INTERP_GUARD@@ block so the on-demand reference flows that still read .graphify_python keep their re-resolve safety net. Regenerate the skill artifacts.
…phify-Labs#197) The core pipeline was converted in the prior commit, but the reference flows it loads on demand (update, query, add/watch, transcribe, exports) still ran $(cat graphify-out/.graphify_python) -c "...", so /graphify update, /graphify query, /graphify add, transcription, and the MCP export still triggered the same Claude Code approval prompts. Convert them too, so the whole Claude Code skill set is free of inline Python and interpreter substitution. New pipeline steps: vocab, transcribe, detect-incremental, code-only-check, empty-extract, update-merge, graph-diff. Reference rewrites: - query: vocab step, graphify query/path/explain, graphify save-result; the obsolete inline NetworkX traversal fallback (which duplicated the CLI) is removed, and the query stub no longer promises it. - update: detect-incremental / code-only-check / empty-extract / update-merge / graph-diff replace the six inline blocks. - add-watch: graphify add, graphify watch (which now parses --debounce and defaults to the .graphify_root marker instead of the cwd). - transcribe: the transcribe step. - exports: graphify-mcp console script instead of -m graphify.serve. The drift guard now also covers graphify/skills/*/references/*.md, so a future inline block in any split host fails the test.
…abs#197) - cli: boolean pipeline flags (--directed/--force/--no-gitignore/ --google-workspace) no longer consume the next token, so detect --no-gitignore <path> scans the corpus instead of the cwd. - cli: watch --debounce with no value errors instead of silently defaulting. - query reference: drop stale prose about the removed inline fallback. - exports reference: note the absolute-interpreter fallback for a GUI Claude Desktop that lacks graphify-mcp on PATH. - tests: cover the boolean-flag ordering and the watch debounce error.
|
Thanks for the pull request, @nothariharan. A maintainer will review it soon. Want to talk it through while it is in review? Come join us on our Discord server. For longer-form discussion there is also GitHub Discussions. A couple of things that speed up review: make sure the test suite passes on Python 3.10 and 3.13, and that the change keeps extraction deterministic. |
There was a problem hiding this comment.
Graphify reviewed this change.
Formal verification could not match the changed code to the code graph, so it may not have checked the code this PR changed.
Worth a look — the grounded gate found no coupling regressions or blocking issues, but 1 advisory finding(s) below merit a look before merge.
Not checked: tests were not run; formal verification proved 0 of 1 changed function(s) (1 not verified).
Formal verification. PR-changed functions: 0/1 verified (0 proven, 0 may-equivalent, 0 distinguished) · 1 not verified (1 vacuous).
Not verified on this run: dispatch\_command (vacuous: never exercised).
Graphify review — findings
Adds a graphify pipeline <step> command that exposes each stage of the agent skill as a plain CLI step: detect, AST and semantic extraction with cache hits/misses, merging, build, the read-only diagnose health gate, labeling, manifest/cost tracking, cleanup, plus the query, transcribe and incremental-update steps. The generated skill files for every agent call these steps instead of inline python -c blocks, so they no longer need interpreter path substitution. Step sidecars default to the cwd, and boolean flags like --no-gitignore and --directed never swallow the following path.
Worth a look
- step_label writes the report and labels before the shrink-guard can refuse —
graphify/pipeline.py· Escalate · medium- agreed by 2 of 2 members but NOT verified (no proof, no reproducing execution) — consensus is not a verdict; needs human review
Review partial — this diff was larger than one review pass covers, so later files were not reviewed; some findings may be missing.
Analysis details — impact, health, verification
Impact & health
Graphify review
Impact — 2073 functions depend on the 1834 functions this change touches.
Health — this change adds coupling hotspots:
- worse:
dispatch_command()— 2 callers, 131 callees - new:
_cmd_pipeline()— 1 callers, 23 callees - new:
step_build()— 1 callers, 13 callees - new:
step_label()— 1 callers, 9 callees - new:
step_extract_ast()— 1 callers, 7 callees - new:
step_save_manifest()— 1 callers, 7 callees - new:
step_update_merge()— 1 callers, 7 callees - new:
step_graph_diff()— 1 callers, 6 callees - …and 1 more — each is listed as a finding
Verification — 2073 functions in the blast radius were not formally verified this run (proofs are advisory here).
Gate & verification
graphify gate
PASS — objectively clean (no health regressions, tests not run — proofs not run this pass (advisory)). Grounded, not self-assessed.
Advisory (not blocking):
- verification_scope: 2005 function(s) in the blast radius were not formally verified this run
Test selection
Test selection
357 of 357 test file(s) selected (100%) via static blast radius.
Escalated to a full run for safety — the selection is not trustworthy on its own (see below). CI should run the whole suite.
tests/test_affected_cli.py— impact, full-run-safetytests/test_affected_member_seed.py— full-run-safetytests/test_agents_platform.py— impact, full-run-safetytests/test_analyze.py— full-run-safetytests/test_anthropic_custom_endpoint.py— full-run-safetytests/test_antigravity_install.py— full-run-safetytests/test_apm_fallback_version.py— full-run-safetytests/test_architecture_doc.py— full-run-safetytests/test_astro_extraction.py— full-run-safetytests/test_astro_import_ids.py— full-run-safetytests/test_atomic_canvas_export.py— full-run-safetytests/test_atomic_version_stamp.py— full-run-safetytests/test_atomic_writes.py— full-run-safetytests/test_backend_env_isolation.py— full-run-safetytests/test_backend_extras.py— full-run-safetytests/test_benchmark.py— full-run-safetytests/test_benchmark_raw_graph.py— full-run-safetytests/test_blade_extractor.py— full-run-safetytests/test_build.py— full-run-safetytests/test_build_located_semantic_identity.py— full-run-safetytests/test_build_merge_dedup_scope.py— full-run-safetytests/test_build_merge_hyperedges_and_prune.py— full-run-safetytests/test_build_merge_shrink_guard.py— full-run-safetytests/test_builtin_global_type_refs.py— full-run-safetytests/test_cache.py— full-run-safetytests/test_cache_stale_import_target.py— full-run-safetytests/test_callflow_html.py— full-run-safetytests/test_cargo_introspect.py— full-run-safetytests/test_cargo_missing_manifest.py— full-run-safetytests/test_carried_hyperedge_remap.py— full-run-safetytests/test_case_sensitive_resolution.py— full-run-safetytests/test_charmap_encoding.py— full-run-safetytests/test_chunking.py— full-run-safetytests/test_cjs_module_extension.py— full-run-safetytests/test_claude_cli_backend.py— full-run-safetytests/test_claude_md.py— full-run-safetytests/test_cli_broken_pipe.py— full-run-safetytests/test_cli_export.py— full-run-safetytests/test_cli_help.py— full-run-safetytests/test_cloud_cta.py— impact, full-run-safetytests/test_cluster.py— full-run-safetytests/test_cluster_ambiguous_scale.py— full-run-safetytests/test_cluster_exclude_hubs.py— full-run-safetytests/test_cobol_extractor.py— full-run-safetytests/test_codebuddy.py— impact, full-run-safetytests/test_community_hub_labels.py— full-run-safetytests/test_community_labels_skill.py— impact, changed-test, full-run-safetytests/test_confidence.py— full-run-safetytests/test_corrupt_graph_json.py— full-run-safetytests/test_cpp_method_declarations.py— full-run-safety- … and 307 more
non-code file(s) changed (
graphify/skill-agents.md,graphify/skill-amp.md,graphify/skill-claw.md,graphify/skill-codex.md,graphify/skill-copilot.md…) → running the full suite for safety (a code graph can't see config/fixture/data deps)
changed code file(s) with no mapped test (
graphify/skill-agents.md,graphify/skill-amp.md,graphify/skill-claw.md,graphify/skill-codex.md,graphify/skill-copilot.md…) — a coverage gap or a missing link — running the full suite rather than only the selected tests
Selection is safe under the controlled-regression assumption; always-run tests + a periodic full run are the backstops. Advisory — it never changes the check verdict.
Docs that may be stale (advisory)
CHANGELOG.md§ 0.9.44 (2026-08-15) (lines 519-531): references changed symbolscorpusCHANGELOG.md§ 0.9.33 (2026-08-05) (lines 641-647): references changed symbolscorpusCHANGELOG.md§ 0.9.18 (2026-07-17) (lines 805-818): references changed symbolscorpusCHANGELOG.md§ 0.9.17 (2026-07-16) (lines 819-848): references changed symbolscorpusCHANGELOG.md§ 0.9.12 (2026-07-10) (lines 929-948): references changed symbolscorpusCHANGELOG.md§ 0.8.1 (2026-05-15) (lines 1592-1603): references changed symbolscorpusdocs/how-it-works.md§ Token benchmark (lines 53-70): references changed symbolscorpusdocs/superpowers/plans/2026-05-04-incremental-updates-dedup.md§ Task 4: Incremental updates — semantic cache + manifest in __main__.py (lines 603-911): references changed symbolscorpusdocs/translations/README.ja-JP.md§ 実例 (lines 193-202): references changed symbolscorpusdocs/translations/README.ko-KR.md§ 실전 예제 (lines 233-242): references changed symbolscorpus
…and 10 more.
Formal verification
Could not verify: Could not verify dispatch\_command.
The verifier did not have enough to check dispatch\_command, so it is saying so rather than guessing. No false assurance is the whole point.
Guarantee: No guarantee either way, this is an honest abstention, not a pass.
Note: Reason: no capturable inputs from the test suite; property tier: not verifiable: all 40 sampled inputs raised on both versions — the function never executed, so 'no divergence' would be vacuous (mostly SystemExit — names the real obstacle, not a sampling gap)
· 9 grounded finding(s) anchored inline below.
| return "\n".join(lines) | ||
|
|
||
|
|
||
| def _cmd_pipeline(argv: list[str]) -> None: |
There was a problem hiding this comment.
_cmd_pipeline()
fans out to 23 callees (efferent coupling).
Grounded coupling-delta finding (deterministic), not an LLM guess.
| sys.exit(1) | ||
|
|
||
|
|
||
| def dispatch_command(cmd: str) -> None: |
There was a problem hiding this comment.
dispatch_command()
fans out to 131 callees (efferent coupling).
Grounded coupling-delta finding (deterministic), not an LLM guess.
| # -------------------------------------------------------------------------- | ||
|
|
||
|
|
||
| def step_extract_ast( |
There was a problem hiding this comment.
step_extract_ast()
fans out to 7 callees (efferent coupling).
Grounded coupling-delta finding (deterministic), not an LLM guess.
| # -------------------------------------------------------------------------- | ||
|
|
||
|
|
||
| def step_merge_chunks( |
There was a problem hiding this comment.
step_merge_chunks()
fans out to 6 callees (efferent coupling).
Grounded coupling-delta finding (deterministic), not an LLM guess.
| # -------------------------------------------------------------------------- | ||
|
|
||
|
|
||
| def step_build( |
There was a problem hiding this comment.
step_build()
fans out to 13 callees (efferent coupling).
Grounded coupling-delta finding (deterministic), not an LLM guess.
| # -------------------------------------------------------------------------- | ||
|
|
||
|
|
||
| def step_label( |
There was a problem hiding this comment.
step_label()
fans out to 9 callees (efferent coupling).
Grounded coupling-delta finding (deterministic), not an LLM guess.
| # -------------------------------------------------------------------------- | ||
|
|
||
|
|
||
| def step_save_manifest( |
There was a problem hiding this comment.
step_save_manifest()
fans out to 7 callees (efferent coupling).
Grounded coupling-delta finding (deterministic), not an LLM guess.
| return {"created": True} | ||
|
|
||
|
|
||
| def step_update_merge( |
There was a problem hiding this comment.
step_update_merge()
fans out to 7 callees (efferent coupling).
Grounded coupling-delta finding (deterministic), not an LLM guess.
| } | ||
|
|
||
|
|
||
| def step_graph_diff( |
There was a problem hiding this comment.
step_graph_diff()
fans out to 6 callees (efferent coupling).
Grounded coupling-delta finding (deterministic), not an LLM guess.
…mands # Conflicts: # graphify/skill-agents.md # graphify/skill-amp.md # graphify/skill-claw.md # graphify/skill-codex.md # graphify/skill-copilot.md # graphify/skill-droid.md # graphify/skill-kilo.md # graphify/skill-kiro.md # graphify/skill-opencode.md # graphify/skill-pi.md # graphify/skill-trae.md # graphify/skill-vscode.md # graphify/skill-windows.md # graphify/skill.md # graphify/skills/agents/references/update.md # graphify/skills/amp/references/update.md # graphify/skills/claude/references/update.md # graphify/skills/claw/references/update.md # graphify/skills/codex/references/update.md # graphify/skills/copilot/references/update.md # graphify/skills/droid/references/update.md # graphify/skills/kilo/references/update.md # graphify/skills/kiro/references/update.md # graphify/skills/opencode/references/update.md # graphify/skills/pi/references/update.md # graphify/skills/trae/references/update.md # graphify/skills/vscode/references/update.md # graphify/skills/windows/references/update.md # tools/skillgen/expected/graphify__skill-agents.md # tools/skillgen/expected/graphify__skill-amp.md # tools/skillgen/expected/graphify__skill-claw.md # tools/skillgen/expected/graphify__skill-codex.md # tools/skillgen/expected/graphify__skill-copilot.md # tools/skillgen/expected/graphify__skill-droid.md # tools/skillgen/expected/graphify__skill-kilo.md # tools/skillgen/expected/graphify__skill-kiro.md # tools/skillgen/expected/graphify__skill-opencode.md # tools/skillgen/expected/graphify__skill-pi.md # tools/skillgen/expected/graphify__skill-trae.md # tools/skillgen/expected/graphify__skill-vscode.md # tools/skillgen/expected/graphify__skill-windows.md # tools/skillgen/expected/graphify__skill.md # tools/skillgen/expected/graphify__skills__agents__references__update.md # tools/skillgen/expected/graphify__skills__amp__references__update.md # tools/skillgen/expected/graphify__skills__claude__references__update.md # tools/skillgen/expected/graphify__skills__claw__references__update.md # tools/skillgen/expected/graphify__skills__codex__references__update.md # tools/skillgen/expected/graphify__skills__copilot__references__update.md # tools/skillgen/expected/graphify__skills__droid__references__update.md # tools/skillgen/expected/graphify__skills__kilo__references__update.md # tools/skillgen/expected/graphify__skills__kiro__references__update.md # tools/skillgen/expected/graphify__skills__opencode__references__update.md # tools/skillgen/expected/graphify__skills__pi__references__update.md # tools/skillgen/expected/graphify__skills__trae__references__update.md # tools/skillgen/expected/graphify__skills__vscode__references__update.md # tools/skillgen/expected/graphify__skills__windows__references__update.md # tools/skillgen/fragments/core/core.md # tools/skillgen/fragments/references/shared/update.md
…Graphify-Labs#197, Graphify-Labs#4240) The split hosts no longer carry an inline detect block, so the Graphify-Labs#4240 guard now exercises graphify.pipeline.step_detect / step_detect_incremental, and the monoliths' remaining inline blocks are still executed. Same contract, new path.
What this fixes
Closes the remaining work in #197.
The Graphify agent skill delivered every pipeline step — and, on demand, every reference flow — as inline Python inside a bash string:
In Claude Code that trips two content-based approval heuristics once per step:
command_substitution— from$(cat graphify-out/.graphify_python).#inside a quoted argument" — from the multi-line Python comments.Neither can be satisfied by permission allow-rules, so a single run cost roughly 30 manual "Yes" clicks and the skill was close to unusable without hand-patching it.
How it's fixed
Every step is now a plain console command. The step bodies live in a new
graphify/pipeline.pyand are exposed asgraphify pipeline <step>:detect(...)+ write sidecargraphify pipeline detect <path>collect_files+extractgraphify pipeline extract-ast <path>.graphify_semantic.jsongraphify pipeline empty-semanticcheck_semantic_cache(...)graphify pipeline cache-check <path> --prompt-file SPECgraphify pipeline merge-chunkssave_semantic_cache(...)graphify pipeline save-cache <path> --prompt-file SPECgraphify pipeline merge-semanticgraphify pipeline merge-extractiongraphify pipeline build <path> [--directed]graphify pipeline diagnose <path>graphify pipeline label <path> --labels '<json>'graphify pipeline save-manifest <path>graphify pipeline cleanupThe on-demand reference flows are converted the same way, with seven more steps plus existing commands:
graphify pipeline vocab,graphify query,graphify path,graphify explain,graphify save-resultgraphify pipeline detect-incremental | code-only-check | empty-extract | update-merge | graph-diffgraphify add,graphify watch(now parses--debounceand defaults to the.graphify_rootmarker)graphify pipeline transcribegraphify-mcpconsole script instead of-m graphify.serveBecause every step is a console-script call,
Bash(graphify *)becomes an allowlistable prefix, thecat-to-interpreter indirection disappears, and the security surface shrinks: arguments go through argparse instead of being interpolated into a Python string.The skill is generated — the fragment is the source of truth
graphify/skill*.mdandgraphify/skills/*/references/*.mdare rendered fromtools/skillgen/fragments/bypython -m tools.skillgen. The edits are in the fragments (core/core.md,references/**,query-stub/default.md); the ~200 generated artifacts and theexpected/snapshots are regenerated and blessed (python -m tools.skillgen --bless).Verification
.graphify_detect.json,.graphify_ast.json,.graphify_semantic.json,.graphify_extract.json,.graphify_analysis.json,graph.json,manifest.json, andGRAPH_REPORT.mdare all identical (report modulo its timestamp). Same graph, different delivery.v8; the full suite is 4 failures, all pre-existing onv8(3 intest_hooks.py, 1 intest_unmapped_at_alias_resolution.py) — none introduced by this change. ~4,200 targeted skill/pipeline tests pass.tests/test_pipeline_cli.py(pipeline steps end to end, including the incremental-update flow, the#479shrink guard +--force,--no-gitignore,GRAPHIFY_OUT, and the boolean-flag/positional parsing) andtests/test_skill_no_inline_python.py(a drift guard that fails if any split host or reference reintroduces an inline-cblock).graphify-outfrom a wrongcache_root,merge-chunkserroring on the all-cached path,--no-gitignore/--forcesilently dropped, a boolean flag swallowing the following path, non-numeric label keys crashing,GRAPHIFY_OUTignored, stale fallback prose, and the MCP config caveat).Known blocker / deliberately out of scope: the aider & devin monoliths
skill-aider.mdandskill-devin.mdstill contain inline Python (37 and 39-cblocks). This is not a capability gap — it's a deliberately isolated, high-blast-radius change, and it is not what #197 reports (those are Aider and Devin, not Claude Code).Why they're separate:
They aren't generated from
core.md+ references.gen.render()returns_read_fragment("core/{aider|devin}.md")verbatim — each is a hand-maintained 1,300–1,400-line single file that duplicates the entire pipeline and every reference flow inline. Converting them is the same work as everything in this PR, repeated twice.They're frozen against a pinned v8 blob.
monolith_roundtrip(tools/skillgen/gen.py:1232) diffs the rendered monolith againstroundtrip_ref(a pinned v8 commit) and requires every non-blank changed line to match a predicate in_SANCTIONED_MONOLITH_DIFFS(19 predicates). So a conversion there also has to add a new sanctioned-diff predicate and re-validate the round-trip — the guard exists precisely to stop arbitrary drift, and touching it is a wide, cross-cutting change.Different hosts. Aider and Devin have their own delivery mechanism; the Claude Code heuristics in Skill should use CLI subcommands instead of inline python -c #197 don't apply to them.
I chose to keep this PR scoped, correct, and low-risk rather than fold a 76-block rewrite of two frozen files into it. The monolith conversion is a clean follow-up PR on its own.
Scope summary
skill.md+skills/claude/references/*): zero inline Python, zero$(cat graphify-out/.graphify_python). Only Step 1's one-time interpreter detection remains (inherent to installing/detecting graphify, and it is a single block, not per-step).-c.Ref #114 (which fixed only the
save_query_resultblock).