Add ainative-router submission (AINative Router, content-only routing) - #213
agent-studio-kr wants to merge 12 commits into
Conversation
Task-group routing policy fit only on an external 2,179-query calibration set (same public sources, all RouterArena items removed), hash-frozen before routing RouterArena. Registers openai/gpt-6-luna and deepseek/deepseek-v4.1-flash (OpenRouter list prices). Local official evaluation: Arena 76.32, accuracy 78.72%, $0.379 per 1K queries, robustness 90.24.
|
/evaluate |
|
Leaderboard display request (if accepted):
Thanks! |
Router Evaluation ResultsRouter: RouterArena Metrics
Optimality Metrics
Evaluation completed by RouterArena automated workflow |
|
@yl231 merge please |
Ko-Agent Router v2. Same method as v1 (task-group policy fit only on an external 2,179-query calibration set with all RouterArena items removed); adds deepseek/deepseek-v4-flash-0731 to the pool (registered at its OpenRouter list price, $0.021 / $0.32 per M tokens). The change was accepted on a fresh 2,177-query held-out calibration extension (same construction, disjoint from RouterArena and the fit set) before routing RouterArena: dArena +0.50, 95% CI [+0.27, +0.75] vs v1. Policy and signals re-frozen before routing. Local official evaluation: Arena 76.64 (v1 76.32), accuracy 79.13%, $0.385 per 1K queries, robustness 88.81 (v1 90.24). Renames the submission ID arena-router -> ko-agent-router (config, adapter KoAgentRouter, prediction files).
|
Updated this PR to v2 and renamed the submission ID to |
|
/evaluate |
1 similar comment
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Optimality Metrics
Evaluation completed by RouterArena automated workflow |
|
Leaderboard display request (updated for v2; supersedes the earlier one):
Thanks! |
No change to the policy, predictions, or generations; config, adapter class (AINativeRouter), and prediction files are renamed only.
|
Renamed the submission to AINative Router ( Leaderboard display request (supersedes the earlier ones):
Thanks! |
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Optimality Metrics
Evaluation completed by RouterArena automated workflow |
|
Thanks for the thorough documentation, @agent-studio-kr. The method notes, cross-validation and freeze records made this straightforward to review, and the model prices check out. We can't accept the router as it stands, though, because its main routing signal comes from RouterArena itself. From
We ruled on this approach in #140: routing that matches the prompt-template strings the RouterArena harness adds to each dataset is dataset-specific routing, and routing must depend on the query's own content. It falls under the README's evaluation-only rule:
Your per-group model choices were calibrated on external items, which is the right way to fit a policy. The problem is the grouping itself: keying on RouterArena's exact templates measures the benchmark's structure rather than the router. To be reconsidered: route on the question itself without RouterArena's template strings or config files. For example, strip the instruction header and classify the question body with a signal fitted on external data. Please also drop the re-weighting to RouterArena's per-config proportions, since it depends on the same dataset identity. Then regenerate the predictions, and we're happy to re-review. |
Routing now uses the question content only: the question body (first/last paragraph dropped, field labels and option letters stripped generically) is classified into a content category by a classifier trained only on external calibration data. RouterArena template strings, config files and per-config proportions are no longer used anywhere. Predictions regenerated. Local official evaluation: Arena 75.48, accuracy 78.39%, $0.547 per 1K, robustness 85.00.
|
Thanks for the careful review, @yl231 — agreed, and fixed. The router now routes on the question content only:
Local official scripts: Arena 0.7548, robustness 0.8500. Details in the updated PR description. Happy to adjust anything else. |
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Optimality Metrics
Evaluation completed by RouterArena automated workflow |
Six regular rows had empty responses (provider-side errors / whitespace-only answer); they are re-requested. Routing is unchanged.
|
Regenerated the 6 rows flagged by the evaluation bot (empty responses: provider-side errors, one whitespace-only answer, one runaway generation). Routing and everything else are unchanged; all regular rows now have valid generations. |
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Optimality Metrics
Evaluation completed by RouterArena automated workflow |
The policy can also choose a reasoning setting (effort low / reasoning off) for a model; chosen per content category on external calibration data only. The prediction field keeps the base model name, generated_result.request_params records the reasoning setting sent. Predictions regenerated.
|
Update: the policy can now also choose a lower reasoning setting per content category (effort low / reasoning off), selected on external calibration data only (same Arena on calibration CV, 22% lower cost). |
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Optimality Metrics
Evaluation completed by RouterArena automated workflow |
Final policy is the per-category majority vote over 25 group-bootstrap refits on the external calibration set, which stabilises near-tied options. Multiple-choice moves to deepseek-v4.1-flash with reasoning effort low (same accuracy within noise, ~31% lower cost on calibration). Predictions regenerated; all rows valid.
|
Update: the policy is now the per-category majority vote over 25 bootstrap refits on the external calibration set (bagging), which stabilises near-tied choices. Multiple-choice moves to |
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Optimality Metrics
Evaluation completed by RouterArena automated workflow |
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Optimality Metrics
Evaluation completed by RouterArena automated workflow |
…> deepseek-v4.1-flash, reasoning effort low)
|
/evaluate |
1 similar comment
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Optimality Metrics
Evaluation completed by RouterArena automated workflow |
…soning effort low)
|
Hi @yl231, thanks again for the earlier review. Since then the router routes on the question body only, and the per-config re-weighting you flagged has been removed (the policy now weights each external source equally). For transparency, one detail about our external calibration set: when we collected it, the per-source target counts were initialized from RouterArena's per-source query counts (with a floor of 20 items per source). So the raw composition of the set partly mirrors the benchmark's source mix. The current policy does not use that composition (source-uniform weights), but the category classifier is trained on the rows as collected. Some accepted submissions describe a similar collection protocol (starting the calibration mix from the benchmark's source proportions) and fit their policy on the calibration rows as collected, without a separate re-weighting step. Could you clarify where the boundary is?
We'll follow your guidance. Until then the policy stays source-uniform. |
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Optimality Metrics
Evaluation completed by RouterArena automated workflow |
…on -> deepseek-v4-flash-0731 (reasoning effort low)
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Optimality Metrics
Evaluation completed by RouterArena automated workflow |
…> deepseek-v4.1-flash, reasoning effort low)
|
/evaluate |
Router Evaluation ResultsRouter: RouterArena Metrics
Optimality Metrics
Evaluation completed by RouterArena automated workflow |
Summary
Leaderboard submission for AINative Router (
ainative-router, Agent Studio). Revised after review: routing now uses the question content only.The question body is extracted structurally (first and last paragraph dropped; line-leading field labels, option letters and
Noneplaceholders stripped with generic patterns), embedded with all-MiniLM-L6-v2, and classified into one of 11 content categories by a logistic regression trained only on an external calibration set. Each category maps to the option (model, optionally with a lower reasoning setting) with the best accuracy–cost trade-off on that set. One OpenRouter call per query, no cascades.What changed after review
tests/signals/test_content_signal.py). Format cues are stripped on purpose: with them the category classifier reached 94% CV accuracy, without them 85%, so the 9-point gap is formatting that the router no longer uses.template-v2of the repo above.Files
router_inference/config/ainative-router.jsonrouter_inference/predictions/ainative-router.json(8,400 regular + 4,045 optimality rows,generated_resultpopulated)router_inference/predictions/ainative-router-robustness.json(420 rows)router_inference/router/ainative_router.py(AINativeRouter) +__init__.pyexport (adapter; loads the frozen router from the repo above viaARENA_ROUTER_ROOT)openai/gpt-6-luna($0.10 / $0.50 per M),deepseek/deepseek-v4.1-flash($0.03 / $0.60 per M),deepseek/deepseek-v4-flash-0731($0.021 / $0.32 per M) inuniversal_model_names.py,model_cost/model_cost.json,llm_inference/model_inference.py(OpenRouter provider mapping). No existing entries are changed.Reasoning settings as options
Besides the six models at provider defaults, the policy can choose
deepseek-v4.1-flashwith reasoning effort low,gemini-3-flash-previewwith reasoning off, anddeepseek-v4-flash-0731with reasoning effort low (all measured on every calibration query). Thepredictionfield keeps the base model name (priced bymodel_cost.json);generated_result.request_paramsrecords the reasoning setting sent, and token usage (incl. reasoning tokens) is the provider-reported usage.deepseek-v4.1-flashwith reasoning effort low (18 of 25 refits).check_config_prediction_files.py ainative-router full --check-generated-result: all checks passed;tools/audit_token_accounting.py --strict: 0 unaccounted reasoning tokens; all 8,400 regular rows have valid generations.Compliance
artifacts/FREEZE.json) before routing RouterArena and verified at load. Every row recordsrequested_model,model_used,upstream_provider,request_id,invoked_at, and provider-reported token usage.Disclosures
arena-router,ko-agent-router, CI 0.7625 / 0.7657), which are withdrawn.gpt-6-lunamedium 0.7491; multiple-choice/ethics →deepseek-v4-flashlow 0.7550; a question-only reasoning-length rule for multiple-choice plus translation →deepseek-v4-flash-0731low 0.7559 (code at ainative-router commit 3904e84). The submitted version is the bagged policy selected by calibration CV (predictions identical to commit 9428f51, CI 0.7583).