Same root cause as #252 (the unguarded json.loads), but this is the sharper end of it: what happens when the accumulator is not just empty but truncated mid-string by a cut stream, and what it does to the chat. Filing separately because the blast radius is different from the "empty/short JSON" case there.
What happened, on cptr 0.9.21 behind an OpenAI-compatible proxy:
The proxy cut the SSE stream mid-argument (broken pipe on its side, logged at the same second as the crash). cptr had accumulated arguments_json by concatenating deltas, and the accumulation ended mid-string: {"content": "\"\"\"Regression tests for ... and then nothing. When a later chunk carried finish_reason: tool_calls, cptr ran json.loads on that truncated string and raised:
json.decoder.JSONDecodeError: Unterminated string starting at: line 1 column 13 (char 12)
at json.loads(tc["arguments_json"]) inside stream_openai_completions, which propagates straight out of the async generator into run_chat_task's generic except Exception. So the whole chat task dies, not just the tool call.
The consequences that make this worse than a one-off 500:
- Every turn after it fails identically until the model happens to produce a fully-contained tool call. In the affected chat, eleven consecutive assistant turns ended in the exact same
Unterminated string error block. The user sees the model "reply" with a wall of the same error and cannot get work done; each failed turn consumes a full model call including all the reasoning tokens.
- Each crash persists an error turn into the chat tree, so the history accumulates a dozen
> **Error:** Unterminated string... branches that then ride along in the context for every subsequent attempt, burning tokens and burying the actual work.
- It is invisible from the outside. There is no signal distinguishing "the provider truncated a tool call" from "the model crashed". The user cannot tell whether retrying will help.
Two things would fix this cleanly:
- Guard the parse: strict
json.loads, fall back to whatever lenient extraction exists, and on total failure fail that tool call with an error the model can see and retry, instead of raising out of the stream loop and killing the task.
- Treat a stream that ends before the accumulated arguments parse as an incomplete stream (retry it), rather than parsing whatever partial bytes landed.
The today-workaround for anyone hitting this: nothing on the cptr side, it needs the guard. If your provider is a proxy, the cuts that produce truncated arguments show up as broken pipes or read timeouts in the proxy log, and a proxy that retries internally or keeps the connection alive makes it rare. That is mitigation, not a fix.
Same root cause as #252 (the unguarded
json.loads), but this is the sharper end of it: what happens when the accumulator is not just empty but truncated mid-string by a cut stream, and what it does to the chat. Filing separately because the blast radius is different from the "empty/short JSON" case there.What happened, on cptr 0.9.21 behind an OpenAI-compatible proxy:
The proxy cut the SSE stream mid-argument (broken pipe on its side, logged at the same second as the crash). cptr had accumulated
arguments_jsonby concatenating deltas, and the accumulation ended mid-string:{"content": "\"\"\"Regression tests for ...and then nothing. When a later chunk carriedfinish_reason: tool_calls, cptr ranjson.loadson that truncated string and raised:at
json.loads(tc["arguments_json"])insidestream_openai_completions, which propagates straight out of the async generator intorun_chat_task's genericexcept Exception. So the whole chat task dies, not just the tool call.The consequences that make this worse than a one-off 500:
Unterminated stringerror block. The user sees the model "reply" with a wall of the same error and cannot get work done; each failed turn consumes a full model call including all the reasoning tokens.> **Error:** Unterminated string...branches that then ride along in the context for every subsequent attempt, burning tokens and burying the actual work.Two things would fix this cleanly:
json.loads, fall back to whatever lenient extraction exists, and on total failure fail that tool call with an error the model can see and retry, instead of raising out of the stream loop and killing the task.The today-workaround for anyone hitting this: nothing on the cptr side, it needs the guard. If your provider is a proxy, the cuts that produce truncated arguments show up as broken pipes or read timeouts in the proxy log, and a proxy that retries internally or keeps the connection alive makes it rare. That is mitigation, not a fix.