Derive ReAct text completion from the task output - #272
Conversation
deepfates
left a comment
There was a problem hiding this comment.
Independent review of cded163. The central change is coherent: task-field descriptions drive the live step, loop history retains its durable vocabulary, and no prose classifier turns nonempty output into silence. I ran the affected completion/step-contract/round-trip/native-reduction/JSON-fallback tests: 59 tests, 0 failures. No provider calls.
One medium contract gap remains before I would call generic output-name history support complete: a host-supplied answer-only turn for a task whose output is tool_calls silently disappears on replay. step_turn/2 (lib/imp/predict/react_v2.ex:1493–1500) treats any tool_calls key as durable call storage; Chat then treats the string as a native-call turn (lib/imp/adapter/chat.ex:1200), leaving no assistant message. A static-model probe with Imp.react("intent -> tool_calls", [], lm: lm) and Imp.History.new([%{intent: "before", tool_calls: "PRIOR AUTHORED WORDS"}]) captured no assistant history; the otherwise identical answer task retained the words. The new generic-name tests only replay the loop's own normalized history, so do not catch this boundary. The guard predates this patch; this is an uncovered limitation of the broader compatibility claim, not evidence that ordinary answer tasks regressed.
Please either handle the unambiguous binary task answer before classifying the reserved storage key, with a host-history regression across native/written adapters, or explicitly bound the supported host-history contract. Keep existing real durable call lists unchanged. No repository edits made during review.
|
Addressed the independent review's supplied-history collision in c23f2d8. When the task text output is named The captured-prompt regression covers native/textual modes and atom/string historical keys. It fails before the repair (27 tests, 1 failure) and passes after; the broader completion/history/adapter/restore/last-action selection passes all 99 tests. No models or production calls. |
deepfates
left a comment
There was a problem hiding this comment.
Follow-up independent review of c23f2d8: the supplied-history collision is resolved. I reran the original disposable static-model probe: the tool_calls text task now preserves PRIOR AUTHORED WORDS as the assistant message, matching the ordinary answer case. The binary guard leaves actual call collections on the existing durable-history path, and FieldMap.put avoids parallel atom/string aliases.
Rechecked the repair diff and ran task-completion, written-chat-history and round-trip restoration tests: 30 tests, 0 failures. The new regression covers native/written modes, atom/string historical keys, and unchanged retained history. No remaining blockers found in this reviewed head. This establishes the declaration/replay contract, not improved live model behavior. No repository changes or provider calls during review.
|
CI repair in 98cafc4 changes only the existing narrow Chat Dialyzer filter from Before: |
A one-text-output ReAct task declared
answerbut asked the model to finish innext_thought. In Dwell's observed turns, that mismatch accompanied nonempty "no reply" summaries that the host correctly published. This change derives the step's text field and description from the task output, so the request and completion share one declaration.The existing typed/constrained
submitpath stays intact. Native and textual tools, JSON fallback, last requests, saved agents and legacy demonstrations use the derived roles. Durable history still usesnext_thought/tool_calls; boundary projection preserves existing conversations. A task namedtool_callsusesagent_tool_callsfor its written calls. Scripted model responses/custom step renderers must use the task output name.Validation:
mix checkpassed (59 doctests, 9 properties, 3588 tests, 0 failures; 13 skipped, 221 excluded). A final focused run including the additional textual-call collision test passed all 72 tests. Coverage includes captured requests/adapters, history and demo collisions, blank and nonempty completion, and existing terminal/last-request behavior. The new regression suite fails against unchanged main (20 failures in the original 25-test set) and passes with the implementation. No models, production services or social publishing were invoked. This repairs the contract; it does not prove improved model behavior. Nonempty silence prose still remains output, and malformed-wrapper detection is outside this change.Closes #270. Context: https://github.com/GroveResearch/dwell/issues/184#issuecomment-5942331037