Hello RouterArena maintainers,
We are integrating TITAN, Theus's commercial model router, and completed a preregistered, cache-only routing pilot on the 809 shared sub_10 cases. The pool uses the exact three historical model IDs in your official caches. Selection received only each question; we froze the policy/state and sealed every prediction before accessing scores. No component was trained or tuned on RouterArena, and no inference, embedding or model-judge calls were made.
Our evaluator stopped on one cached value outside our preregistered [0,1] quality domain:
- Repository revision:
cff9659dd6a3a07815f4842fcdadea5a597a1df1
- Cache:
claude-3-haiku-20240307.jsonl, SHA-256 9f5bed666a524d74e629091b57ef2fdf644181dcc1039ef44b44460f254d150a
global_index: NarrativeQA_4102
evaluation_result.metric: meteor_score
evaluation_result.score: -0.0330188679245283
We retained all 809 routing decisions and have published no aggregate quality/cost for this failed run. We did not clip the value, drop the case, or impute a replacement.
Evidence: public report and downloadable artifacts, preregistered protocols and prediction seals.
Could you clarify:
- Is this negative cached METEOR value intentional, and how should the current official aggregator treat it?
- For a future commercial-router submission, may the complete 8400-case and 420-case robustness predictions reuse exact public cached generations with fixed, disclosed provenance? Which artifacts and redistribution permissions are required when the router implementation and operational snapshot remain private?
We understand that this 809-case pilot is not eligible as a full leaderboard result and are not requesting that it be listed as one. We want to establish a valid zero-new-inference path to a complete submission.
Thank you.
Hello RouterArena maintainers,
We are integrating TITAN, Theus's commercial model router, and completed a preregistered, cache-only routing pilot on the 809 shared
sub_10cases. The pool uses the exact three historical model IDs in your official caches. Selection received only each question; we froze the policy/state and sealed every prediction before accessing scores. No component was trained or tuned on RouterArena, and no inference, embedding or model-judge calls were made.Our evaluator stopped on one cached value outside our preregistered [0,1] quality domain:
cff9659dd6a3a07815f4842fcdadea5a597a1df1claude-3-haiku-20240307.jsonl, SHA-2569f5bed666a524d74e629091b57ef2fdf644181dcc1039ef44b44460f254d150aglobal_index:NarrativeQA_4102evaluation_result.metric:meteor_scoreevaluation_result.score:-0.0330188679245283We retained all 809 routing decisions and have published no aggregate quality/cost for this failed run. We did not clip the value, drop the case, or impute a replacement.
Evidence: public report and downloadable artifacts, preregistered protocols and prediction seals.
Could you clarify:
We understand that this 809-case pilot is not eligible as a full leaderboard result and are not requesting that it be listed as one. We want to establish a valid zero-new-inference path to a complete submission.
Thank you.