Skip to content

Free-lane turns end at max-tokens with zero output: max_tokens is hard-coded to 32768 and shared with reasoning #16

Description

@codeOct

Summary

On the anonymous Zen free lane a turn can end with finish_reason: length after spending the entire 32768-token output budget on reasoning, returning no answer text and no tool call. Reproduced twice in a row with mimo-v2.6-flash-free at thinking level high.

Two causes, both in packages/plugin/src/adapter/zen-adapter.ts:

  1. The per-model output cap is a constant. Every model is advertised with maxTokens: DEFAULT_MAX_TOKENS (32768) and defaultMaxTokens: DEFAULT_MAX_TOKENS, and the catalog never reads the model's real output/context limits — so a model that supports more than 32768 is still capped at 32768.
  2. reasoning_effort is forwarded while max_tokens stays at 32768. Reasoning and the answer share that single budget on this lane, so high lets the model consume all of it before emitting anything at all.

The reasoning channel itself is fine — the reasoning arrives as a clean, separate type: "reasoning" block with no encoding damage. This is purely a budget-allocation problem.

Environment

  • @opencode2dsh/dsh-plugin@0.3.3 (latest; master does not change any of this)
  • opencode2dsh/mimo-v2.6-flash-free, reasoningEffort: high
  • DSH harness on Windows

Observed

Two consecutive turns, taken from the DSH session log (decompressed session.v3.jsonl):

turn 1: finish.reason.kind = "max-tokens" | usage: inputTokens 15069, outputTokens 32768
        content: [ { type: "reasoning", ... } ]          <- reasoning only

turn 2: finish.reason.kind = "max-tokens" | usage: inputTokens 35, cacheReadTokens 15040, outputTokens 32768
        content: [ { type: "reasoning", ... } ]

outputTokens is exactly 32768 in both turns — the declared cap, to the token. Neither turn contains a text block or a tool call.

Impact in DSH: the turn surfaces as "Output token limit reached", and on max-tokens DSH drops tool-call blocks from the truncated response (they cannot be executed safely), so the turn produces nothing usable. Sending "continue" cannot recover a dropped tool call — the user has to re-issue the request.

Source references (master)

packages/plugin/src/adapter/zen-adapter.ts

Line Code
39–40 const DEFAULT_CONTEXT_WINDOW = 262144 / const DEFAULT_MAX_TOKENS = 32768
152, 167–168 toPiModel() → contextWindow: DEFAULT_CONTEXT_WINDOW, maxTokens: DEFAULT_MAX_TOKENS
247, 261–262 resolveModel() → context: { contextWindow: DEFAULT_CONTEXT_WINDOW }, defaultMaxTokens: DEFAULT_MAX_TOKENS
485, 506 reasoning_effort: effortWire is injected into the payload while maxTokens: options.maxTokens is forwarded unchanged
49, 52, 69 REASONING_EFFORT_LADDER, DEFAULT_EFFORT_LADDER = ['off','minimal','low','medium','high'], reasoningEfforts()

packages/plugin/src/adapter/catalog.ts

Line Code
88 decodeModelsDev()
122–126 stores only cost.input, cost.output, deprecated, reasoning, reasoning_options

No limit fields are read anywhere: a grep over the shipped lib/ finds zero occurrences of limit., max_output, maxOutputTokens or maxOutputTokens. Because decodeModelsDev() also never reads limit.context, DEFAULT_CONTEXT_WINDOW is a guess for every model as well.

Related: reasoningEfforts() falls back to DEFAULT_EFFORT_LADDER whenever the metadata declares no ladder, so high is offered for models where — given a shared 32768 budget — it is effectively a "produce nothing" setting.

Reproduction

  1. Select mimo-v2.6-flash-free and set the thinking level to High.
  2. Ask for something that invites a long plan, e.g. "write a single-file HTML + SVG animation of …".
  3. The turn ends at exactly 32768 output tokens with max-tokens, containing only a reasoning block.

Suggested fix

  • Prefer the model's real cap. Read limit.output (and limit.context) in decodeModelsDev(), carry it on the catalog entry, and use it for maxTokens / defaultMaxTokens, keeping 32768 only as the fallback when metadata cannot speak.
  • Give thinking its own budget. Do not hand the entire output cap to the reasoning phase: either reserve headroom for the answer, or stop advertising high / max (as a default, at least) when a thinking-only turn would consume the whole cap.
  • Optional safety net. Warn — or mark the turn distinctly — when a turn ends at the cap having emitted no content block at all.

Workaround (no plugin change needed)

Set the thinking level to Off. Per the 0.3.3 changelog, reasoningEffort: off maps to wire reasoning_effort: "none", which is the only spelling that actually stops the always-think free models. Switching provider also avoids the pattern.

Caveat

I did not measure the true output cap of mimo-v2.6-flash-free on this lane. If 32768 happens to be its real limit, then the constant is numerically correct and what remains is the missing budget isolation between thinking and answering — the truncation itself would then be the model's own behaviour. The per-model-metadata gap (cause 1) stands either way.


中文摘要:mimo-v2.6-flash-free 在思考等级 High 下,会把全部 32768 输出预算用在 reasoning 上,正文和工具调用一个字都没输出,上游以 finish_reason: length 结束(实测 outputTokens 正好 32768)。原因有两点,都在 packages/plugin/src/adapter/zen-adapter.ts:其一,每个模型的输出上限被写死为常量 32768,catalog 从不读取 models.dev 的 limit.output / limit.context;其二,reasoning_effort 直接透传而上限不变,而该通道的思考与正文共用同一预算。建议按模型真实上限取值,并为思考单独预留预算;用户侧可先把思考等级设为 Off(线上发 reasoning_effort: "none")规避。

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions