Skip to content

Agent trace recipe for broken call chains + tpk eval harness (#17 step 4, #35) - #91

Merged
gangtao merged 2 commits into
mainfrom
feat/agent-trace-strategy-and-eval
Sep 21, 2026
Merged

gangtao merged 2 commits into
mainfrom
feat/agent-trace-strategy-and-eval

Conversation

@gangtao

@gangtao gangtao commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Step 4 ("D") of the #17 plan, and the eval harness #35 asked for. Does not close #17; closes the tooling half of #35.

Why

Agent — a trace recipe (agent.py, new rule 7)

For "trace / how does X reach Y" questions: (a) find both endpoints — methods are named Class::method since #88; (b) path_between; mode: "calls" → done; (c) otherwise walk neighbors(id, rels=["calls"], direction="out") a hop at a time (and "in" from the far end); (d) at a dead end, read_source the function, find the interface call (x->write(...)), look up implementations with search_entities("::write", kinds=["function"]), pick by context, continue; (e) answer with every hop marked [graph] or [code] and say what is still unconfirmed. Traces may use 18 tool calls instead of 12.

tpk eval (src/tpk/evals.py, src/tpk/eval_questions.toml)

Runs a bundled 12-question set — 3 trace, 3 relationship, 3 architecture, 3 docs/deployment — through the real agent on the live graph and records, per question: tool-call sequence, whether a graph-edge tool was used, call count vs budget, a keyword sanity check on the answer, latency, errors. Trace/relationship questions require neighbors or path_between.

tpk eval --yes -o before.json
tpk eval --yes -o after.json --compare before.json      # passed / graph_tool_rate / avg_tool_calls, fixed + regressed ids
tpk eval --only trace --limit 2                          # narrow a run;  --questions my.toml for a custom set

Output: JSON report + a markdown table by category. An agent/gateway error fails that question and is recorded; it never aborts the run. It is a behaviour probe, not a correctness benchmark — expect_answer_any only catches an answer that isn't about the thing asked. Costs real LLM tokens (confirmation prompt unless --yes). The question file ships inside the package (verified in the built wheel), so it works in the container.

Testing

  • tests/test_evals.py (infra-free, fake LangGraph agent): tool-sequence/answer extraction incl. Anthropic block content and parallel tool calls; scoring (tools / answer / budget); error isolation; per-category summary; compare (metric deltas, fixed / regressed); the bundled set is well-formed; report round-trip; the CLI end-to-end with --only --limit --compare.
  • tests/test_agent.py: the prompt carries the recipe.
  • 58 infra-free tests pass. The DB-backed suite was not re-run; this PR touches no DB code.

Live results (2026-09-21) — local stack, both proton entries re-ingested with #88, agent = Opus 4.8

Same image, same graph, same 12 questions; only agent.py differs (main's prompt vs this PR's).

main prompt + trace recipe
passed 11 / 12 11 / 12
used a graph-edge tool 67% (trace 100%, relationship 100%, architecture 67%, docs 0%) 67% (identical split)
avg tool calls 6.92 7.83
errors 0 0

Honest reading:

  • The recipe does not move the numbers. One run per arm, so ±1–2 calls per question is noise, but there is no measurable gain in pass rate or graph usage, and a small cost (+0.9 calls/question).
  • The big change had already happened — from ingest: derive code kinds and class-qualified names from graph structure (#17, step 1) #88 (kinds / Class::method names, after re-ingest) and path_between: directed call chains (#17, step 2) — re-land of #89 onto main #90 (path_between + its prompt text). Production chats audited before those: 1 of 13 turns touched a graph edge (~8%). Baseline here: 67%, and 100% on trace and relationship questions. "Which classes inherit from IInterpreter?" is answered in 2 calls.
  • The Deep call-path queries (e.g. "trace HTTP insert → nativelog write") return no result — sparse C++ call graph + agent doesn't use path_between #17 question is now answered. "Trace the call path from HTTP insert to nativelog write" → HTTPHandler::processQuery → InterpreterInsertQuery::execute → StorageStream::write → StreamSink → shard store → NativeLog, stitched across the interface breaks by reading source. It takes 21–22 tool calls; it "failed" only the harness's 20-call budget, which was my arbitrary number — that one question now has max_tool_calls = 24 (still under the recursion limit).
  • What the recipe does buy is provenance. With it, that answer marks each hop [graph] / [code] (2 / 4) and states that path_between found only a related chain, not a call chain. Without it the agent presented the whole path as "the confirmed chain". That is the reason to keep it; it is not a performance change.

So: the UI's suggested trace question can stay. Re-ingest of the proton entries took ~4 min each (app peak ≈ 1.6 GiB); node count for proton-enterprise@v3.3.1 went 163,924 → 168,419 because overloads no longer collide.

Caveats: n = 1 run per arm; the answer check is a keyword sanity check, not a correctness judgement — I read the hard-trace answers, not all 24.

🤖 Generated with Claude Code

gangtao and others added 2 commits September 21, 2026 13:02
…#17, #35)

#35's verification showed the 'lean on the graph' nudge did nothing: 13 audited
turns = search 61 / read_source 31 / neighbors 1 / path_between 0. And #17's
measurements showed why more nudging can't work: the extracted C++ call graph
breaks at every virtual or pointer call (77% of member calls can't be typed), so
a chain cannot be completed from edges alone.

Agent: a TRACE RECIPE rule -- find endpoints by Class::method name, try
path_between, otherwise walk neighbors(rels=['calls']) a hop at a time, and at a
dead end read_source the function, find the interface call, look up the
implementations with search_entities('::method', kinds=['function']) and
continue; answer with each hop marked [graph] or [code]. Traces get 18 calls.

Eval: 'tpk eval' runs a bundled 12-question set (trace / relationship /
architecture / docs) through the REAL agent and reports tool sequence,
graph-tool usage, call count and sanity checks per question and per category;
JSON report, markdown summary, --compare against a previous run, --only/--limit/
--questions. An agent or gateway error fails that question without aborting.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ed 21-22)

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Deep call-path queries (e.g. "trace HTTP insert → nativelog write") return no result — sparse C++ call graph + agent doesn't use path_between

1 participant