Move --llm polish to claude-sonnet-5-5 and handle incomplete responses - #1
Merged
Merged
Conversation
- transcript_tools/cli.py: --model default claude-haiku-4-5-20251001 ->
claude-sonnet-5-5 (the --llm polish is the primary quality model for
published transcripts on a low-volume path, not a bulk tier). Help text
notes the model must accept the effort parameter.
API-shape changes in transcript_tools/llm.py polish_turn:
- output_config={"effort": "low"}: per-turn cleanup; low effort skips
thinking on most turns. No thinking param is sent (adaptive by default).
- max_tokens 4096 -> 16000: thinking counts toward max_tokens.
- Raise on any stop_reason other than end_turn (refusal, max_tokens,
pause_turn, ...). A partial reply could pass the number guard, so the
turn now falls back to the deterministic text with reason
"error: incomplete response (stop_reason=...)".
- Text is still read by block type, so thinking blocks are skipped.
- No temperature/top_p/top_k, tool_choice, or assistant prefill was ever
sent; none is added.
Dependencies: [llm] extra anthropic>=0.40 -> anthropic>=1.0 (first release
line whose messages.create takes output_config with effort). anthropic
added to the dev extra so CI exercises the real SDK request shape.
Tests: request shape (fake client), every StopReason value, the real SDK
over an httpx2 MockTransport (no network), and the CLI default.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Contributor
Author
|
Merge audit (model-pin audit 2026-09-28): independent review by GPT-6 Astra via Subfleet (job 20260928-214012-mpr-policyengine-transcript-tools) at head 49fd1f6. It re-ran the suite (28 passed on anthropic 1.9.0 and on the 1.0.0 floor) and requested one change: correct the PR body's Ruff-format claim. The body now scopes that claim to the three clean files and discloses the pre-existing cli.py:30 difference. Code is unchanged; CI 🤖 Generated with Claude Code |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
The optional
--llmtranscript polish defaulted toclaude-haiku-4-5-20251001. That model is the main quality model for published transcript prose on a low-volume path (one call per speaker turn), not a bulk tier. This PR moves the default toclaude-sonnet-5-5and updates the call so it is valid and safe on that model.transcript_tools/cli.pyclaude-haiku-4-5-20251001->claude-sonnet-5-5--modeldefault for--llmpyproject.toml[llm]anthropic>=0.40->anthropic>=1.0;anthropic>=1.0added to[dev]output_config.effort; CI runs the real-SDK testAPI changes (
transcript_tools/llm.pypolish_turn)output_config={"effort": "low"}. This is per-turn cleanup, and atlowSonnet 5.5 skips thinking on most requests. Nothinkingparam is sent (adaptive is the default).max_tokens4096 -> 16000, because thinking counts towardmax_tokens.stop_reason == "end_turn". A refusal, amax_tokenscut-off orpause_turnraises. A truncated or refused turn with no numbers after the cut would have passed the number guard and quietly dropped the rest of the turn. Nowpolish_paragraphskeeps the deterministic text and reportserror: incomplete response (stop_reason=...).Invariants (tested)
temperature,top_p,top_k,thinking,tool_choiceor an assistant prefill to the model. This is checked on the kwargs and on the JSON body the real SDK serializes.stop_reason == "end_turn". This is checked for every value ofanthropic.types.StopReason(7 cases).Tests run
uv run --extra dev pytest -q -rs: 28 passed, 0 skipped (anthropic 1.9.0). The baseline was 16.anthropic==1.0.0(the new floor): 28 passed.httpx2.MockTransport, so there is no network access and no API key.ruff format(88 columns) is clean ontests/test_llm.py,tests/test_cli.pyandtranscript_tools/llm.py.transcript_tools/cli.pyhas one formatting difference at line 30, which this PR does not touch and which is identical onmain.ruff checkhas one finding on an unchanged line (BLE001 atllm.py:103, from the initial commit). Ruff is not a CI gate.Env overrides to check
None. The model comes only from the
--modelCLI flag.Left alone on purpose
glossary.yamlGPT-5.x regexes, theGPT-5.5strings intests/test_core.pyandtests/test_llm.py, and the README glossary example. These normalize spoken product names in transcripts; they don't select a model.fallbacksbeta. It isn't needed: a refused turn already falls back to the deterministic text, and--modelcan name models where it may not apply.Notes
--modelstill accepts any id, but it must now be a model that takesoutput_config.effort. Haiku 4.5 and Sonnet 4.5 reject it, so a run with those models would keep every turn unpolished. The help text and README say this.Part of the 2026-09-28 cross-repo model-pin audit (inventory:
~/reviews/model-pin-audit-2026-09-28/REPORT.md). Pins classified as benchmark, historical record, fixture or needing Max's call were deliberately left unchanged.🤖 Generated with Claude Code