Skip to content

Add ainative-router submission (AINative Router, content-only routing) - #213

Open
agent-studio-kr wants to merge 12 commits into
RouteWorks:mainfrom
agent-studio-kr:arena-router
Open

agent-studio-kr wants to merge 12 commits into
RouteWorks:mainfrom
agent-studio-kr:arena-router

Conversation

@agent-studio-kr

@agent-studio-kr agent-studio-kr commented Sep 28, 2026 •

Copy link
Copy Markdown

Summary

Leaderboard submission for AINative Router (ainative-router, Agent Studio). Revised after review: routing now uses the question content only.

The question body is extracted structurally (first and last paragraph dropped; line-leading field labels, option letters and None placeholders stripped with generic patterns), embedded with all-MiniLM-L6-v2, and classified into one of 11 content categories by a logistic regression trained only on an external calibration set. Each category maps to the option (model, optionally with a lower reasoning setting) with the best accuracy–cost trade-off on that set. One OpenRouter call per query, no cascades.

What changed after review

  • Removed the grouping by RouterArena eval-config instruction heads and the kNN over instruction prefixes. The router reads no RouterArena files (config, templates, datasets); a test checks that no router module references benchmark files.
  • Removed the re-weighting to RouterArena per-config proportions; calibration sources are weighted equally when fitting the policy.
  • Routing is invariant to rewriting the instruction / answer-format paragraphs and to relabelling fields or option letters (tests in tests/signals/test_content_signal.py). Format cues are stripped on purpose: with them the category classifier reached 94% CV accuracy, without them 85%, so the 9-point gap is formatting that the router no longer uses.
  • The earlier template-based code is kept only under the tag template-v2 of the repo above.

Files

  • router_inference/config/ainative-router.json
  • router_inference/predictions/ainative-router.json (8,400 regular + 4,045 optimality rows, generated_result populated)
  • router_inference/predictions/ainative-router-robustness.json (420 rows)
  • router_inference/router/ainative_router.py (AINativeRouter) + __init__.py export (adapter; loads the frozen router from the repo above via ARENA_ROUTER_ROOT)
  • Model registration for three OpenRouter models at OpenRouter list prices: openai/gpt-6-luna ($0.10 / $0.50 per M), deepseek/deepseek-v4.1-flash ($0.03 / $0.60 per M), deepseek/deepseek-v4-flash-0731 ($0.021 / $0.32 per M) in universal_model_names.py, model_cost/model_cost.json, llm_inference/model_inference.py (OpenRouter provider mapping). No existing entries are changed.

Reasoning settings as options

Besides the six models at provider defaults, the policy can choose deepseek-v4.1-flash with reasoning effort low, gemini-3-flash-preview with reasoning off, and deepseek-v4-flash-0731 with reasoning effort low (all measured on every calibration query). The prediction field keeps the base model name (priced by model_cost.json); generated_result.request_params records the reasoning setting sent, and token usage (incl. reasoning tokens) is the provider-reported usage.

  • Calibration CV: router vs best single option ΔArena +2.07 [95% CI +1.26, +2.86].
  • The final policy is the per-category majority vote over 25 group-bootstrap refits on the calibration set (bagging), which stabilises near-tied options; multiple-choice gets deepseek-v4.1-flash with reasoning effort low (18 of 25 refits).
  • check_config_prediction_files.py ainative-router full --check-generated-result: all checks passed; tools/audit_token_accounting.py --strict: 0 unaccounted reasoning tokens; all 8,400 regular rows have valid generations.
  • Robustness (local): 0.8310. Official CI: see the bot comment below.

Compliance

  • Routing depends on the query content only. No RouterArena template string, config file, dataset identifier, query, answer, label, or per-config proportion is used by, or was used to fit, any router component. RouterArena data was used only as an exclusion list when building the calibration set, and for this evaluation.
  • Calibration set: 6,409 queries from the same public source benchmarks with every RouterArena item removed (normalized exact match, MiniLM cosine ≥ 0.85 on question + context, and source-ID checks). Rebuilding from public sources reproduces the published manifests (prompt hashes).
  • Category labels are our own 11-category taxonomy of the external sources (e.g. multiple-choice knowledge, math, code, translation, reading comprehension); they are used only as training labels.
  • Validation: group 5-fold CV on the calibration set with the classifier, the policy and the best single model all refit inside each fold; ΔArena vs best single model +2.91 [95% CI +2.19, +3.61].
  • Classifier weights, policy and code are hash-frozen (artifacts/FREEZE.json) before routing RouterArena and verified at load. Every row records requested_model, model_used, upstream_provider, request_id, invoked_at, and provider-reported token usage.

Disclosures

  1. The calibration set's per-source sample sizes were set from RouterArena's per-config counts when it was built (before this revision). They only determine how many external items were drawn per source; the router does not use them (sources are weighted equally).
  2. The model pool was shortlisted with the help of public leaderboard model pools and, in early exploration, aggregate per-model statistics from other submissions' public prediction files. Model selection per category uses only the external calibration set.
  3. This PR previously carried template-based versions (arena-router, ko-agent-router, CI 0.7625 / 0.7657), which are withdrawn.
  4. Cost-first variants of this router were also evaluated on this PR and are withdrawn (bot comments above): multiple-choice → gpt-6-luna medium 0.7491; multiple-choice/ethics → deepseek-v4-flash low 0.7550; a question-only reasoning-length rule for multiple-choice plus translation → deepseek-v4-flash-0731 low 0.7559 (code at ainative-router commit 3904e84). The submitted version is the bagged policy selected by calibration CV (predictions identical to commit 9428f51, CI 0.7583).

Task-group routing policy fit only on an external 2,179-query calibration
set (same public sources, all RouterArena items removed), hash-frozen
before routing RouterArena. Registers openai/gpt-6-luna and
deepseek/deepseek-v4.1-flash (OpenRouter list prices).

Local official evaluation: Arena 76.32, accuracy 78.72%,
$0.379 per 1K queries, robustness 90.24.
@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

@agent-studio-kr

Copy link
Copy Markdown
Author

Leaderboard display request (if accepted):

  • Router name: Ko-Agent Router (submission ID / file names stay arena-router)
  • Affiliation: Agent Studio (@agent-studio-kr)
  • Links: [Code]
  • Type: open-source (method, code, calibration manifest and freeze record are public)

Thanks!

@github-actions

Copy link
Copy Markdown

Router Evaluation Results

Router: arena-router
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7625
Accuracy 78.64%
Total Cost $3.182717
Avg Cost per Query $0.000379
Avg Cost per 1K Queries $0.3789
Number of Queries 8400
Abnormal Entries 0
Robustness Score 0.9024

Optimality Metrics

Metric Value
Opt.Sel (Optimal Selection) 0.0751
Opt.Cost (Cost Efficiency) 0.1633
Opt.Acc (Accuracy vs Optimal) 0.9293

Evaluation completed by RouterArena automated workflow

@agent-studio-kr

Copy link
Copy Markdown
Author

@yl231 merge please

Ko-Agent Router v2. Same method as v1 (task-group policy fit only on an
external 2,179-query calibration set with all RouterArena items removed);
adds deepseek/deepseek-v4-flash-0731 to the pool (registered at its
OpenRouter list price, $0.021 / $0.32 per M tokens).

The change was accepted on a fresh 2,177-query held-out calibration
extension (same construction, disjoint from RouterArena and the fit set)
before routing RouterArena: dArena +0.50, 95% CI [+0.27, +0.75] vs v1.
Policy and signals re-frozen before routing.

Local official evaluation: Arena 76.64 (v1 76.32), accuracy 79.13%,
$0.385 per 1K queries, robustness 88.81 (v1 90.24).

Renames the submission ID arena-router -> ko-agent-router (config,
adapter KoAgentRouter, prediction files).
@agent-studio-kr agent-studio-kr changed the title Add arena-router submission Add ko-agent-router submission (Ko-Agent Router) Sep 28, 2026
@agent-studio-kr

Copy link
Copy Markdown
Author

Updated this PR to v2 and renamed the submission ID to ko-agent-router (display name: Ko-Agent Router). v2 adds one model (deepseek/deepseek-v4-flash-0731); it was accepted on a fresh held-out calibration extension (ΔArena +0.50 [+0.27, +0.75] vs v1) before RouterArena was routed. Local official scripts: Arena 0.7664 (v1 0.7632). Details and disclosures are in the updated PR description. The old arena-router files are removed.

@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

1 similar comment
@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

@github-actions

Copy link
Copy Markdown

Router Evaluation Results

Router: ko-agent-router
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7657
Accuracy 79.04%
Total Cost $3.230867
Avg Cost per Query $0.000385
Avg Cost per 1K Queries $0.3846
Number of Queries 8400
Abnormal Entries 0
Robustness Score 0.8881

Optimality Metrics

Metric Value
Opt.Sel (Optimal Selection) 0.0643
Opt.Cost (Cost Efficiency) 0.1580
Opt.Acc (Accuracy vs Optimal) 0.9240

Evaluation completed by RouterArena automated workflow

@agent-studio-kr

Copy link
Copy Markdown
Author

Leaderboard display request (updated for v2; supersedes the earlier one):

  • Router name: Ko-Agent Router (submission ID / file names: ko-agent-router)
  • Affiliation: Agent Studio (@agent-studio-kr)
  • Links: [Code]
  • Type: open-source (method, code, calibration manifests, held-out report and freeze record are public)
  • CI: Arena 0.7657, accuracy 79.04%, $0.3846 / 1K, robustness 0.8881 (the first /evaluate run was aborted by a runner shutdown; the re-run completed)

Thanks!

No change to the policy, predictions, or generations; config, adapter
class (AINativeRouter), and prediction files are renamed only.
@agent-studio-kr agent-studio-kr changed the title Add ko-agent-router submission (Ko-Agent Router) Add ainative-router submission (AINative Router) Sep 28, 2026
@agent-studio-kr

agent-studio-kr commented Sep 28, 2026 •

Copy link
Copy Markdown
Author

Renamed the submission to AINative Router (ainative-router); policy, predictions and generations are unchanged from the ko-agent-router v2 run (CI 0.7657). A pre-registered second held-out test did not justify further changes, so v2 stands.

Leaderboard display request (supersedes the earlier ones):

  • Router name: AINative Router (submission ID / file names: ainative-router)
  • Affiliation: Agent Studio (@agent-studio-kr)
  • Links: [Code]
  • Type: open-source (method, code, calibration manifests, held-out and pre-registration reports, freeze record are public)

Thanks!

@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

@github-actions

Copy link
Copy Markdown

Router Evaluation Results

Router: ainative-router
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7657
Accuracy 79.04%
Total Cost $3.230867
Avg Cost per Query $0.000385
Avg Cost per 1K Queries $0.3846
Number of Queries 8400
Abnormal Entries 0
Robustness Score 0.8881

Optimality Metrics

Metric Value
Opt.Sel (Optimal Selection) 0.0643
Opt.Cost (Cost Efficiency) 0.1580
Opt.Acc (Accuracy vs Optimal) 0.9240

Evaluation completed by RouterArena automated workflow

@yl231

yl231 commented Sep 28, 2026

Copy link
Copy Markdown
Contributor

Thanks for the thorough documentation, @agent-studio-kr. The method notes, cross-validation and freeze records made this straightforward to review, and the model prices check out. We can't accept the router as it stands, though, because its main routing signal comes from RouterArena itself.

From docs/method.md and arena_router/signals.py:

  1. template_heads() reads RouterArena's own eval configs (third_party/RouterArena/config/eval_config/zero-shot/*.json) and maps each config's fixed instruction header to the RouterArena source configs that use it (26 groups over 41 configs).
  2. The router's primary signal, task_by_template(), matches each prompt against those headers, so the group it routes on is effectively the RouterArena dataset the prompt came from. The kNN fallback is built to recover the same groups when the header is reworded.

We ruled on this approach in #140: routing that matches the prompt-template strings the RouterArena harness adds to each dataset is dataset-specific routing, and routing must depend on the query's own content. It falls under the README's evaluation-only rule:

Submissions that train, fit, or tune any router component on RouterArena data (including the label files) will be rejected.

Your per-group model choices were calibrated on external items, which is the right way to fit a policy. The problem is the grouping itself: keying on RouterArena's exact templates measures the benchmark's structure rather than the router.

To be reconsidered: route on the question itself without RouterArena's template strings or config files. For example, strip the instruction header and classify the question body with a signal fitted on external data. Please also drop the re-weighting to RouterArena's per-config proportions, since it depends on the same dataset identity. Then regenerate the predictions, and we're happy to re-review.

Routing now uses the question content only: the question body (first/last
paragraph dropped, field labels and option letters stripped generically) is
classified into a content category by a classifier trained only on external
calibration data. RouterArena template strings, config files and per-config
proportions are no longer used anywhere. Predictions regenerated.

Local official evaluation: Arena 75.48, accuracy 78.39%, $0.547 per 1K,
robustness 85.00.
@agent-studio-kr agent-studio-kr changed the title Add ainative-router submission (AINative Router) Add ainative-router submission (AINative Router, content-only routing) Sep 28, 2026
@agent-studio-kr

Copy link
Copy Markdown
Author

Thanks for the careful review, @yl231 — agreed, and fixed. The router now routes on the question content only:

  1. No RouterArena templates or config files. Template grouping and the kNN over instruction prefixes are removed. The question body is extracted structurally (first/last paragraph dropped), then line-leading field labels, option letters and None placeholders are stripped with generic patterns, so formatting cannot act as a dataset fingerprint (classifier CV accuracy drops from 94% to 85% without them; that gap is exactly the formatting we no longer use). The body is classified into 11 content categories by a MiniLM + logistic-regression model trained only on our external calibration set.
  2. No RouterArena proportions. The re-weighting is removed; external sources are weighted equally when fitting the policy.
  3. Tests check routing invariance to rewritten instruction / format paragraphs and relabelled fields, and that no router module references benchmark files. Predictions are regenerated; the old code is kept only under a tag in our repo.

Local official scripts: Arena 0.7548, robustness 0.8500. Details in the updated PR description. Happy to adjust anything else.

@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

@github-actions

Copy link
Copy Markdown

Router Evaluation Results

Router: ainative-router
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7541
Accuracy 78.31%
Total Cost $4.596102
Avg Cost per Query $0.000547
Avg Cost per 1K Queries $0.5472
Number of Queries 8400
Abnormal Entries 6
Robustness Score 0.8500

⚠️ 6 of 8400 queries (0.1%) had no valid generation (inference failed / empty answer) and were scored as incorrect (0). These queries still count toward the denominator, so accuracy and cost reflect the full query set. Please regenerate predictions for these queries and resubmit for a complete evaluation.

Optimality Metrics

Metric Value
Opt.Sel (Optimal Selection) 0.0424
Opt.Cost (Cost Efficiency) 0.2212
Opt.Acc (Accuracy vs Optimal) 0.8950

Evaluation completed by RouterArena automated workflow

Six regular rows had empty responses (provider-side errors / whitespace-only
answer); they are re-requested. Routing is unchanged.
@agent-studio-kr

Copy link
Copy Markdown
Author

Regenerated the 6 rows flagged by the evaluation bot (empty responses: provider-side errors, one whitespace-only answer, one runaway generation). Routing and everything else are unchanged; all regular rows now have valid generations.

@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

@github-actions

Copy link
Copy Markdown

Router Evaluation Results

Router: ainative-router
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7543
Accuracy 78.34%
Total Cost $4.601825
Avg Cost per Query $0.000548
Avg Cost per 1K Queries $0.5478
Number of Queries 8400
Abnormal Entries 0
Robustness Score 0.8500

Optimality Metrics

Metric Value
Opt.Sel (Optimal Selection) 0.0424
Opt.Cost (Cost Efficiency) 0.2212
Opt.Acc (Accuracy vs Optimal) 0.8950

Evaluation completed by RouterArena automated workflow

The policy can also choose a reasoning setting (effort low / reasoning off) for a
model; chosen per content category on external calibration data only. The
prediction field keeps the base model name, generated_result.request_params
records the reasoning setting sent. Predictions regenerated.
@agent-studio-kr

Copy link
Copy Markdown
Author

Update: the policy can now also choose a lower reasoning setting per content category (effort low / reasoning off), selected on external calibration data only (same Arena on calibration CV, 22% lower cost). prediction keeps the base model name; request_params records the setting per row. Predictions regenerated; all checks pass.

@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

@github-actions

Copy link
Copy Markdown

Router Evaluation Results

Router: ainative-router
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7542
Accuracy 78.28%
Total Cost $4.490355
Avg Cost per Query $0.000535
Avg Cost per 1K Queries $0.5346
Number of Queries 8400
Abnormal Entries 0
Robustness Score 0.8310

Optimality Metrics

Metric Value
Opt.Sel (Optimal Selection) 0.0439
Opt.Cost (Cost Efficiency) 0.2254
Opt.Acc (Accuracy vs Optimal) 0.8962

Evaluation completed by RouterArena automated workflow

Final policy is the per-category majority vote over 25 group-bootstrap refits on
the external calibration set, which stabilises near-tied options. Multiple-choice
moves to deepseek-v4.1-flash with reasoning effort low (same accuracy within noise,
~31% lower cost on calibration). Predictions regenerated; all rows valid.
@agent-studio-kr

Copy link
Copy Markdown
Author

Update: the policy is now the per-category majority vote over 25 bootstrap refits on the external calibration set (bagging), which stabilises near-tied choices. Multiple-choice moves to deepseek-v4.1-flash with reasoning effort low. Predictions regenerated; all 8,400 regular rows have valid generations and all checks pass.

@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

@github-actions

Copy link
Copy Markdown

Router Evaluation Results

Router: ainative-router
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7583
Accuracy 78.09%
Total Cost $3.080138
Avg Cost per Query $0.000367
Avg Cost per 1K Queries $0.3667
Number of Queries 8400
Abnormal Entries 0
Robustness Score 0.8310

Optimality Metrics

Metric Value
Opt.Sel (Optimal Selection) 0.0394
Opt.Cost (Cost Efficiency) 0.2913
Opt.Acc (Accuracy vs Optimal) 0.8832

Evaluation completed by RouterArena automated workflow

@jh941213 jh941213 mentioned this pull request Sep 29, 2026
@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

@github-actions

Copy link
Copy Markdown

Router Evaluation Results

Router: ainative-router
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7491
Accuracy 76.03%
Total Cost $1.522593
Avg Cost per Query $0.000181
Avg Cost per 1K Queries $0.1813
Number of Queries 8400
Abnormal Entries 0
Robustness Score 0.8143

Optimality Metrics

Metric Value
Opt.Sel (Optimal Selection) 0.0656
Opt.Cost (Cost Efficiency) 0.3723
Opt.Acc (Accuracy vs Optimal) 0.8746

Evaluation completed by RouterArena automated workflow

…> deepseek-v4.1-flash, reasoning effort low)
@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

1 similar comment
@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

@github-actions

Copy link
Copy Markdown

Router Evaluation Results

Router: ainative-router
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7583
Accuracy 78.09%
Total Cost $3.080138
Avg Cost per Query $0.000367
Avg Cost per 1K Queries $0.3667
Number of Queries 8400
Abnormal Entries 0
Robustness Score 0.8310

Optimality Metrics

Metric Value
Opt.Sel (Optimal Selection) 0.0394
Opt.Cost (Cost Efficiency) 0.2913
Opt.Acc (Accuracy vs Optimal) 0.8832

Evaluation completed by RouterArena automated workflow

@agent-studio-kr

Copy link
Copy Markdown
Author

Hi @yl231, thanks again for the earlier review. Since then the router routes on the question body only, and the per-config re-weighting you flagged has been removed (the policy now weights each external source equally).

For transparency, one detail about our external calibration set: when we collected it, the per-source target counts were initialized from RouterArena's per-source query counts (with a floor of 20 items per source). So the raw composition of the set partly mirrors the benchmark's source mix. The current policy does not use that composition (source-uniform weights), but the category classifier is trained on the rows as collected.

Some accepted submissions describe a similar collection protocol (starting the calibration mix from the benchmark's source proportions) and fit their policy on the calibration rows as collected, without a separate re-weighting step. Could you clarify where the boundary is?

  1. Is initializing an external calibration collection plan from RouterArena's per-source counts acceptable, or should we rebuild it with a benchmark-independent allocation?
  2. If it is acceptable, may the policy be fit on those rows as collected (per-query weighting), or should it stay source-uniform?
  3. Does the classifier training on the collected rows need any change?

We'll follow your guidance. Until then the policy stays source-uniform.

@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

@github-actions

Copy link
Copy Markdown

Router Evaluation Results

Router: ainative-router
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7550
Accuracy 77.58%
Total Cost $2.821043
Avg Cost per Query $0.000336
Avg Cost per 1K Queries $0.3358
Number of Queries 8400
Abnormal Entries 0
Robustness Score 0.8310

Optimality Metrics

Metric Value
Opt.Sel (Optimal Selection) 0.0465
Opt.Cost (Cost Efficiency) 0.2699
Opt.Acc (Accuracy vs Optimal) 0.8852

Evaluation completed by RouterArena automated workflow

…on -> deepseek-v4-flash-0731 (reasoning effort low)
@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

@github-actions

Copy link
Copy Markdown

Router Evaluation Results

Router: ainative-router
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7559
Accuracy 77.75%
Total Cost $2.958265
Avg Cost per Query $0.000352
Avg Cost per 1K Queries $0.3522
Number of Queries 8400
Abnormal Entries 0
Robustness Score 0.8024

Optimality Metrics

Metric Value
Opt.Sel (Optimal Selection) 0.0525
Opt.Cost (Cost Efficiency) 0.2672
Opt.Acc (Accuracy vs Optimal) 0.8790

Evaluation completed by RouterArena automated workflow

…> deepseek-v4.1-flash, reasoning effort low)
@agent-studio-kr

Copy link
Copy Markdown
Author

/evaluate

@github-actions

Copy link
Copy Markdown

Router Evaluation Results

Router: ainative-router
Dataset Split: full

RouterArena Metrics

Metric Value
RouterArena Score 0.7583
Accuracy 78.09%
Total Cost $3.080138
Avg Cost per Query $0.000367
Avg Cost per 1K Queries $0.3667
Number of Queries 8400
Abnormal Entries 0
Robustness Score 0.8310

Optimality Metrics

Metric Value
Opt.Sel (Optimal Selection) 0.0394
Opt.Cost (Cost Efficiency) 0.2913
Opt.Acc (Accuracy vs Optimal) 0.8832

Evaluation completed by RouterArena automated workflow

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants