Skip to content

Move --llm polish to claude-sonnet-5-5 and handle incomplete responses - #1

Merged
MaxGhenis merged 1 commit into
mainfrom
chore/model-pins-2026-09-28
Sep 29, 2026
Merged

MaxGhenis merged 1 commit into
mainfrom
chore/model-pins-2026-09-28

Conversation

@MaxGhenis

@MaxGhenis MaxGhenis commented Sep 29, 2026 •

Copy link
Copy Markdown
Contributor

Summary

The optional --llm transcript polish defaulted to claude-haiku-4-5-20251001. That model is the main quality model for published transcript prose on a low-volume path (one call per speaker turn), not a bulk tier. This PR moves the default to claude-sonnet-5-5 and updates the call so it is valid and safe on that model.

File Old -> new Role
transcript_tools/cli.py claude-haiku-4-5-20251001 -> claude-sonnet-5-5 --model default for --llm
pyproject.toml [llm] anthropic>=0.40 -> anthropic>=1.0; anthropic>=1.0 added to [dev] SDK floor with output_config.effort; CI runs the real-SDK test

API changes (transcript_tools/llm.py polish_turn)

  • output_config={"effort": "low"}. This is per-turn cleanup, and at low Sonnet 5.5 skips thinking on most requests. No thinking param is sent (adaptive is the default).
  • max_tokens 4096 -> 16000, because thinking counts toward max_tokens.
  • The response is used only when stop_reason == "end_turn". A refusal, a max_tokens cut-off or pause_turn raises. A truncated or refused turn with no numbers after the cut would have passed the number guard and quietly dropped the rest of the turn. Now polish_paragraphs keeps the deterministic text and reports error: incomplete response (stop_reason=...).
  • Text is still read by block type, so thinking blocks are ignored.

Invariants (tested)

  • No call site sends temperature, top_p, top_k, thinking, tool_choice or an assistant prefill to the model. This is checked on the kwargs and on the JSON body the real SDK serializes.
  • A polished turn is used only when stop_reason == "end_turn". This is checked for every value of anthropic.types.StopReason (7 cases).
  • The number guard is unchanged: a polished turn never changes a number, dollar amount or percentage.

Tests run

  • uv run --extra dev pytest -q -rs: 28 passed, 0 skipped (anthropic 1.9.0). The baseline was 16.
  • The same suite in a clean venv with anthropic==1.0.0 (the new floor): 28 passed.
  • The real-SDK test uses an httpx2.MockTransport, so there is no network access and no API key.
  • ruff format (88 columns) is clean on tests/test_llm.py, tests/test_cli.py and transcript_tools/llm.py. transcript_tools/cli.py has one formatting difference at line 30, which this PR does not touch and which is identical on main. ruff check has one finding on an unchanged line (BLE001 at llm.py:103, from the initial commit). Ruff is not a CI gate.

Env overrides to check

None. The model comes only from the --model CLI flag.

Left alone on purpose

  • glossary.yaml GPT-5.x regexes, the GPT-5.5 strings in tests/test_core.py and tests/test_llm.py, and the README glossary example. These normalize spoken product names in transcripts; they don't select a model.
  • The server-side refusal fallbacks beta. It isn't needed: a refused turn already falls back to the deterministic text, and --model can name models where it may not apply.

Notes

  • --model still accepts any id, but it must now be a model that takes output_config.effort. Haiku 4.5 and Sonnet 4.5 reject it, so a run with those models would keep every turn unpolished. The help text and README say this.
  • Sonnet 5.5 vs Opus 5.5 for this default is open for Max. Opus costs about 2x, still cents per webinar.

Part of the 2026-09-28 cross-repo model-pin audit (inventory: ~/reviews/model-pin-audit-2026-09-28/REPORT.md). Pins classified as benchmark, historical record, fixture or needing Max's call were deliberately left unchanged.

🤖 Generated with Claude Code

- transcript_tools/cli.py: --model default claude-haiku-4-5-20251001 ->
  claude-sonnet-5-5 (the --llm polish is the primary quality model for
  published transcripts on a low-volume path, not a bulk tier). Help text
  notes the model must accept the effort parameter.

API-shape changes in transcript_tools/llm.py polish_turn:
- output_config={"effort": "low"}: per-turn cleanup; low effort skips
  thinking on most turns. No thinking param is sent (adaptive by default).
- max_tokens 4096 -> 16000: thinking counts toward max_tokens.
- Raise on any stop_reason other than end_turn (refusal, max_tokens,
  pause_turn, ...). A partial reply could pass the number guard, so the
  turn now falls back to the deterministic text with reason
  "error: incomplete response (stop_reason=...)".
- Text is still read by block type, so thinking blocks are skipped.
- No temperature/top_p/top_k, tool_choice, or assistant prefill was ever
  sent; none is added.

Dependencies: [llm] extra anthropic>=0.40 -> anthropic>=1.0 (first release
line whose messages.create takes output_config with effort). anthropic
added to the dev extra so CI exercises the real SDK request shape.

Tests: request shape (fake client), every StopReason value, the real SDK
over an httpx2 MockTransport (no network), and the CLI default.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MaxGhenis

Copy link
Copy Markdown
Contributor Author

Merge audit (model-pin audit 2026-09-28): independent review by GPT-6 Astra via Subfleet (job 20260928-214012-mpr-policyengine-transcript-tools) at head 49fd1f6. It re-ran the suite (28 passed on anthropic 1.9.0 and on the 1.0.0 floor) and requested one change: correct the PR body's Ruff-format claim. The body now scopes that claim to the three clean files and discloses the pre-existing cli.py:30 difference. Code is unchanged; CI test is green; mergeable. Main session (Opus 5.5) reviewed the diff: explicit effort low, max_tokens 16000 for thinking headroom, refusal/max_tokens stop_reason guard, text read by block type.

🤖 Generated with Claude Code

@MaxGhenis
MaxGhenis merged commit 0d6698c into main Sep 29, 2026
1 check passed
@MaxGhenis
MaxGhenis deleted the chore/model-pins-2026-09-28 branch September 29, 2026 02:04
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant