Skip to content

Model price profiles: exact price matching, max price for unmatched models - #218

Open
yl231 wants to merge 2 commits into
mainfrom
pricing/model-profiles
Open

yl231 wants to merge 2 commits into
mainfrom
pricing/model-profiles

Conversation

@yl231

@yl231 yl231 commented Sep 29, 2026

Copy link
Copy Markdown
Contributor

Introduces model price profiles as the single source of truth for prices, makes price lookup exact, and charges models with no price profile the highest price in the table instead of $0. No score on the leaderboard changes.

What changes

  • model_cost/model_profiles.yaml (new): one profile per model — price (USD per 1M input/output tokens), aliases (other spellings of the same model), as_of, source, note. 86 profiles, 19 aliases. Prices are seeded from today's model_cost.json unchanged; a market sync can update them later.
  • scripts/pricing/build_model_cost.py (new): generates model_cost.json from the profiles (each id and alias becomes an entry at the profile's price), so the evaluator, /evaluate bot, and other tools keep reading the same file. --check fails if the JSON was edited by hand; wired in as a pre-commit hook, so it runs in CI.
  • Evaluator (evaluate_models.py): price lookup is exact id/alias only. The substring fallback — which let a name pick up whichever key happened to be a substring of it — is removed. A model with no price profile is charged the table's highest input and output price (currently $15/$75), warned once per name, and summarized at the end of the run.
  • run.py: a row whose model name is not in universal_model_names.py is now graded and charged the maximum price instead of being dropped from the results.
  • Submission check: a model without a price profile is now a warning (with the max-price consequence spelled out), not a failure, so /evaluate runs and applies the rule above.

Keeping it score-neutral

  • Duplicate spellings are folded into aliases only where that moves no score. Three models keep two profiles each (qwen3-235b-a22b-2507, claude-haiku-4.5, qwen3-30b-a3b-instruct-2507) because different board entries are charged different prices for them today; they carry a conflict note to resolve when prices are re-synced.
  • Eight spellings that were priced only through the substring fallback are added as explicit aliases (e.g. Lynkr's openai/gpt-oss-120b, glm-4-air, and model_used spellings like google/gemini-3-flash-preview).
  • Three duplicate entries that no row is charged at get their profile's price (deepseek-v3.2, openai_gpt-oss-120b, qwen_qwen3-235b-a22b-2507).

Testing

Check Result
New unit tests (tests/test_pricing.py, 7) pass
pre-commit on changed files (ruff, format, codespell, yaml/json, mypy) pass
Re-price all 265,568 rows in router_inference/predictions/ with old vs new code (same path as run.py, incl. model_used) identical
Full run.py lynkr full --force on old vs new code 10,018 rows, 0 accuracy / 0 cost differences; Arena 0.6770 both
Synthetic file with an unregistered model 100/100 rows graded, each charged exactly tokens × $15/$75
Hook negative test (hand-edited model_cost.json) fails as intended

Notes for reviewers

  • Open PRs that add prices by editing model_cost.json directly will need their rows moved into model_profiles.yaml when they are merged.
  • Pre-existing issues seen while testing, not addressed here: a checkpoint-save race in parallel evaluation ("dictionary changed size during iteration"), and code grading where sys.stdin has no .buffer, so solutions reading sys.stdin.buffer are marked wrong.

🤖 Generated with Claude Code

Louie Lu and others added 2 commits September 27, 2026 21:52
…ls the max price

- model_cost/model_profiles.yaml is now the source of truth for prices: one
  profile per model with aliases; model_cost.json is generated from it by
  scripts/pricing/build_model_cost.py (--check verifies it is in sync).
- Duplicate spellings are folded into aliases where that moves no score.
  Three models stay double-profiled with a conflict note (qwen3-235b,
  claude-haiku-4.5, qwen3-30b) because different board entries are charged
  different prices for them today.
- Evaluator: price lookup is exact id/alias only (the substring fallback is
  removed). Spellings that relied on it are added as explicit aliases.
- A model with no price profile is charged the highest input/output price in
  the table instead of $0, and an unknown model name is graded and charged
  rather than dropped. The submission check warns instead of failing.
- Re-pricing all 265,568 rows in router_inference/predictions gives identical
  costs before and after.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Runs scripts/pricing/build_model_cost.py --check whenever model_cost/ or
scripts/pricing/ changes, locally and in the Pre-commit CI job.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant