Agent trace recipe for broken call chains + tpk eval harness (#17 step 4, #35) - #91
Merged
Merged
Conversation
…#17, #35) #35's verification showed the 'lean on the graph' nudge did nothing: 13 audited turns = search 61 / read_source 31 / neighbors 1 / path_between 0. And #17's measurements showed why more nudging can't work: the extracted C++ call graph breaks at every virtual or pointer call (77% of member calls can't be typed), so a chain cannot be completed from edges alone. Agent: a TRACE RECIPE rule -- find endpoints by Class::method name, try path_between, otherwise walk neighbors(rels=['calls']) a hop at a time, and at a dead end read_source the function, find the interface call, look up the implementations with search_entities('::method', kinds=['function']) and continue; answer with each hop marked [graph] or [code]. Traces get 18 calls. Eval: 'tpk eval' runs a bundled 12-question set (trace / relationship / architecture / docs) through the REAL agent and reports tool sequence, graph-tool usage, call count and sanity checks per question and per category; JSON report, markdown summary, --compare against a previous run, --only/--limit/ --questions. An agent or gateway error fails that question without aborting. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ed 21-22) Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Step 4 ("D") of the #17 plan, and the eval harness #35 asked for. Does not close #17; closes the tooling half of #35.
Why
search_entities61 /read_source31 /neighbors1 /path_between0. The "lean on the graph" prompt nudge (feat(agent): encourage graph traversal beyond explicit call-path questions #34) did nothing.IInterpreter::executehas 63 overrides and 0 callers). A chain cannot be completed from edges alone — the agent has to read the code at the break, which is what a person does.Agent — a trace recipe (
agent.py, new rule 7)For "trace / how does X reach Y" questions: (a) find both endpoints — methods are named
Class::methodsince #88; (b)path_between;mode: "calls"→ done; (c) otherwise walkneighbors(id, rels=["calls"], direction="out")a hop at a time (and"in"from the far end); (d) at a dead end,read_sourcethe function, find the interface call (x->write(...)), look up implementations withsearch_entities("::write", kinds=["function"]), pick by context, continue; (e) answer with every hop marked[graph]or[code]and say what is still unconfirmed. Traces may use 18 tool calls instead of 12.tpk eval(src/tpk/evals.py,src/tpk/eval_questions.toml)Runs a bundled 12-question set — 3 trace, 3 relationship, 3 architecture, 3 docs/deployment — through the real agent on the live graph and records, per question: tool-call sequence, whether a graph-edge tool was used, call count vs budget, a keyword sanity check on the answer, latency, errors. Trace/relationship questions require
neighborsorpath_between.Output: JSON report + a markdown table by category. An agent/gateway error fails that question and is recorded; it never aborts the run. It is a behaviour probe, not a correctness benchmark —
expect_answer_anyonly catches an answer that isn't about the thing asked. Costs real LLM tokens (confirmation prompt unless--yes). The question file ships inside the package (verified in the built wheel), so it works in the container.Testing
tests/test_evals.py(infra-free, fake LangGraph agent): tool-sequence/answer extraction incl. Anthropic block content and parallel tool calls; scoring (tools / answer / budget); error isolation; per-category summary; compare (metric deltas, fixed / regressed); the bundled set is well-formed; report round-trip; the CLI end-to-end with--only --limit --compare.tests/test_agent.py: the prompt carries the recipe.Live results (2026-09-21) — local stack, both proton entries re-ingested with #88, agent = Opus 4.8
Same image, same graph, same 12 questions; only
agent.pydiffers (main's prompt vs this PR's).Honest reading:
Class::methodnames, after re-ingest) and path_between: directed call chains (#17, step 2) — re-land of #89 onto main #90 (path_between+ its prompt text). Production chats audited before those: 1 of 13 turns touched a graph edge (~8%). Baseline here: 67%, and 100% on trace and relationship questions. "Which classes inherit from IInterpreter?" is answered in 2 calls.HTTPHandler::processQuery→InterpreterInsertQuery::execute→StorageStream::write→StreamSink→ shard store → NativeLog, stitched across the interface breaks by reading source. It takes 21–22 tool calls; it "failed" only the harness's 20-call budget, which was my arbitrary number — that one question now hasmax_tool_calls = 24(still under the recursion limit).[graph]/[code](2 / 4) and states thatpath_betweenfound only a related chain, not a call chain. Without it the agent presented the whole path as "the confirmed chain". That is the reason to keep it; it is not a performance change.So: the UI's suggested trace question can stay. Re-ingest of the proton entries took ~4 min each (app peak ≈ 1.6 GiB); node count for
proton-enterprise@v3.3.1went 163,924 → 168,419 because overloads no longer collide.Caveats: n = 1 run per arm; the answer check is a keyword sanity check, not a correctness judgement — I read the hard-trace answers, not all 24.
🤖 Generated with Claude Code