Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions crates/xgeny-cli/src/composition.rs
Original file line number Diff line number Diff line change
Expand Up @@ -473,6 +473,7 @@ const fn model_rejection_code(reason: ModelCallRejectionReason) -> &'static str
"model_rejected.planner_invalid_response"
}
ModelCallRejectionReason::ProviderLimit => "model_rejected.provider_limit",
ModelCallRejectionReason::OutputTruncated => "model_rejected.output_truncated",
ModelCallRejectionReason::ProviderRejected => "model_rejected.provider_rejected",
ModelCallRejectionReason::ProposalRejected => "model_rejected.proposal_rejected",
ModelCallRejectionReason::MaterializationFailed => "model_rejected.materialization_failed",
Expand Down Expand Up @@ -1672,6 +1673,7 @@ fn map_planner_unavailable(run_id: String, failure: PlannerPortFailure) -> Local
}
PlannerPortFailure::InvalidResponse => ModelCallRejectionReason::PlannerInvalidResponse,
PlannerPortFailure::ProviderLimit => ModelCallRejectionReason::ProviderLimit,
PlannerPortFailure::OutputTruncated => ModelCallRejectionReason::OutputTruncated,
PlannerPortFailure::ProviderRejected => ModelCallRejectionReason::ProviderRejected,
};
LocalCommandResult::Rejected {
Expand Down Expand Up @@ -2689,6 +2691,10 @@ mod tests {
ModelCallRejectionReason::ProviderLimit,
"model_rejected.provider_limit",
),
(
ModelCallRejectionReason::OutputTruncated,
"model_rejected.output_truncated",
),
(
ModelCallRejectionReason::ProviderRejected,
"model_rejected.provider_rejected",
Expand Down
3 changes: 3 additions & 0 deletions crates/xgeny-cli/tests/live_go50902_public.rs
Original file line number Diff line number Diff line change
Expand Up @@ -1342,6 +1342,9 @@ fn require_workspace_completion(output: &Output, state_root: &Path) {
ModelCallRejectionReason::ProviderLimit => {
"live workspace provider response exceeded a limit"
}
ModelCallRejectionReason::OutputTruncated => {
"live workspace provider output was truncated by the output token budget"
}
ModelCallRejectionReason::ProviderRejected => {
"live workspace provider rejected the request"
}
Expand Down
5 changes: 3 additions & 2 deletions crates/xgeny-provider-openai/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -960,6 +960,7 @@ const fn map_compatibility_transport_failure(
PlannerPortFailure::Unavailable => OpenAiCompatibilityCheckFailure::Unavailable,
PlannerPortFailure::InvalidResponse => OpenAiCompatibilityCheckFailure::InvalidResponse,
PlannerPortFailure::ProviderLimit => OpenAiCompatibilityCheckFailure::ProviderLimit,
PlannerPortFailure::OutputTruncated => OpenAiCompatibilityCheckFailure::OutputTruncated,
PlannerPortFailure::ProviderRejected => OpenAiCompatibilityCheckFailure::RequestRejected,
}
}
Expand Down Expand Up @@ -1059,7 +1060,7 @@ fn decode_chat_response(
return Err(PlannerPortFailure::InvalidResponse);
}
if choice.finish_reason == "length" {
return Err(PlannerPortFailure::ProviderLimit);
return Err(PlannerPortFailure::OutputTruncated);
}
if choice.finish_reason != "stop"
|| choice.message.refusal.is_some()
Expand Down Expand Up @@ -1768,7 +1769,7 @@ mod tests {
}
assert_eq!(
decode_chat_response(&response(&valid_plan(), "length"), MODEL, 256 * 1024, 64,),
Err(PlannerPortFailure::ProviderLimit)
Err(PlannerPortFailure::OutputTruncated)
);
assert_eq!(
decode_chat_response(&response(&valid_plan(), "stop"), MODEL, 8, 64),
Expand Down
68 changes: 68 additions & 0 deletions crates/xgeny-provider-openai/tests/http_contract.rs
Original file line number Diff line number Diff line change
Expand Up @@ -804,6 +804,74 @@ fn deterministic_provider_rejection_is_closed_without_raw_error_body() {
assert!(!durable.contains(RAW_RESPONSE_SENTINEL));
}

#[test]
fn truncated_planner_output_is_closed_as_output_truncated_not_provider_limit() {
// ADR-0039: a 200 whose choice ended at the output token budget is a budget problem the user
// fixes with --max-output-tokens, not a rate limit. The journal must keep that distinction.
let body = serde_json::to_vec(&json!({
"id": RAW_RESPONSE_SENTINEL,
"model": "qwen3.8-27b",
"choices": [{
"index": 0,
"message": {"role": "assistant", "content": "{\"objective\":\"partial"},
"finish_reason": "length"
}]
}))
.unwrap();
let server = TestServer::spawn("200 OK", body);
let mut planner = planner(&server.base_url);
let mut store = seed_store();
let loop_runtime = configured_loop(&mut store, &mut planner);
let mut events = DeterministicEvents;
let mut materializer = EphemeralMaterializer;
let tick = loop_runtime
.tick(
&mut store,
&mut events,
&FixedLease,
&synthetic_registry(),
&IdentityResolver::default(),
&mut planner,
&mut materializer,
)
.expect("truncation should settle as a rejection");
assert!(matches!(
tick,
AgentLoopTick::PlannerUnavailable {
failure: PlannerPortFailure::OutputTruncated,
..
}
));
let request = server.finish();
assert_eq!(
std::str::from_utf8(&request)
.unwrap()
.matches("POST /v1/chat/completions")
.count(),
1,
"a truncated response must not be retried"
);
let snapshot = store.load().unwrap().unwrap();
let last = snapshot.records.last().expect("settlement should exist");
assert!(matches!(
&last.event.body,
RunEventBody::ModelCallSettled {
settlement: ModelCallSettlement::Rejected {
reason: ModelCallRejectionReason::OutputTruncated
},
..
}
));
let durable = format!(
"{}{}",
serde_json::to_string(&snapshot.records).unwrap(),
serde_json::to_string(&snapshot.state).unwrap()
);
assert!(durable.contains("\"output_truncated\""));
assert!(!durable.contains(RAW_RESPONSE_SENTINEL));
assert!(!durable.contains("partial"));
}

/// A provider that accepts the request and then stalls longer than the planner budget.
fn spawn_stalling_server(stall: Duration) -> (String, thread::JoinHandle<()>) {
let listener = TcpListener::bind("127.0.0.1:0").expect("test listener should bind");
Expand Down
13 changes: 13 additions & 0 deletions crates/xgeny-runtime/src/agent_loop.rs
Original file line number Diff line number Diff line change
Expand Up @@ -685,6 +685,9 @@ pub enum PlannerPortFailure {
InvalidResponse,
#[error("planner request exceeded provider limits")]
ProviderLimit,
/// The provider stopped at the output token budget before the proposal was complete.
#[error("planner output was truncated by the output token budget")]
OutputTruncated,
#[error("planner provider rejected the request")]
ProviderRejected,
}
Expand Down Expand Up @@ -1706,6 +1709,16 @@ impl AgentLoop {
),
ModelCallConflictIntent::RejectStale,
),
PlannerPortFailure::OutputTruncated => (
append_model_call_rejection(
store,
events,
reserved_state,
call_id,
ModelCallRejectionReason::OutputTruncated,
),
ModelCallConflictIntent::RejectStale,
),
PlannerPortFailure::ProviderRejected => (
append_model_call_rejection(
store,
Expand Down
2 changes: 2 additions & 0 deletions crates/xgeny-workgraph/src/lib.rs
Original file line number Diff line number Diff line change
Expand Up @@ -1462,6 +1462,8 @@ pub enum ModelCallUnknownReason {
pub enum ModelCallRejectionReason {
PlannerInvalidResponse,
ProviderLimit,
/// The provider stopped at the output token budget before the proposal was complete.
OutputTruncated,
ProviderRejected,
ProposalRejected,
MaterializationFailed,
Expand Down
86 changes: 86 additions & 0 deletions docs/adr/0039-planner-output-truncation-rejection-class.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,86 @@
# ADR-0039: planner 출력 잘림을 provider limit과 다른 rejection class로 기록한다

- 상태: Accepted
- 날짜: 2026-09-14
- 관련: ADR-0016 durable model call lifecycle, ADR-0017 OpenAI-compatible provider adapter, ADR-0035 model profile inference limits, ADR-0037 portable proposal schema and compact output

## 배경

Planner 호출이 `finish_reason: "length"`로 끝나면 provider adapter는 `PlannerPortFailure::ProviderLimit`을
돌려주고, Core는 journal에 `ModelCallRejectionReason::ProviderLimit`(`provider_limit`)으로 settlement를
기록하며, CLI는 `XGENY_REJECTED … reason=model_rejected.provider_limit`을 출력한다. HTTP 413과 429도
같은 class다. 즉 사용자는 세 가지 다른 원인을 하나의 문자열로 본다.

| 원인 | 사용자가 취해야 할 조치 |
| --- | --- |
| 출력 token 예산에서 잘림 | `--max-output-tokens`를 올리거나 model의 thinking을 줄인다 |
| HTTP 429 rate limit | 기다리거나 동시 요청을 줄인다 |
| HTTP 413 요청 크기 초과 | workspace/goal을 줄인다 |

Compatibility probe(PR #53, ADR-0032 §4)는 이미 `provider_output_truncated`와 `rate_limited`를 구분한다.
Production planner 경로만 구분하지 않아 진단이 비대칭이다.

실측:

- llama.cpp(Ollama 번들 `llama-server`, Qwen3.8 27B GGUF, `--jinja`)에서 쓰기 Run이 출력 예산 4096
token에서 `model_rejected.provider_limit`으로 닫혔다. 응답의 `reasoning_content`가 10,148자였고
`finish_reason`은 `length`였다. Rate limit이 아니었지만 결과 코드로는 알 수 없었다.
- 2026-09-14 재현: Ollama 0.33 Qwen3.8 27B 읽기 Run에 `XGENY_OPENAI_MAX_OUTPUT_TOKENS=64`를 주면 25초
만에 `model_rejected.provider_limit`으로 닫히고 journal에 `"reason":"provider_limit"`이 남는다.

## 결정

### 1. Provider port에 `OutputTruncated` failure를 추가한다

`xgeny_runtime::PlannerPortFailure`에 `OutputTruncated` variant를 더한다. OpenAI-compatible adapter는
choice의 `finish_reason == "length"`일 때 이 값을 돌려준다. HTTP 413/429와 요청 크기 상한 초과는 그대로
`ProviderLimit`이다.

### 2. Journal에 `output_truncated` rejection class를 추가한다

`xgeny_workgraph::ModelCallRejectionReason`에 `OutputTruncated`를 더한다. serde 표현은 다른 variant와
같은 규칙으로 `output_truncated`다. AgentLoop는 `PlannerPortFailure::OutputTruncated`를 이 class로
settle하고 conflict intent는 `ProviderLimit`과 같은 `RejectStale`이다. 재시도하지 않는다는 ADR-0016의
의미는 그대로다.

### 3. 공개 결과 코드 `model_rejected.output_truncated`를 추가한다

CLI는 `model_rejected.output_truncated`를 출력한다. `model_rejected.provider_limit`의 의미는 "출력
예산·요청 크기 초과"에서 "요청 크기 초과와 429"로 좁아진다. Getting-started troubleshooting 표에 새
코드와 조치(`--max-output-tokens`, `XGENY_OPENAI_MAX_OUTPUT_TOKENS`, thinking 설정)를 적는다.

### 4. 호환성 경계를 그대로 받아들인다

- Request profile digest(ADR-0017)의 입력이 아니다. 진행 중 Run의 resume은 이 변경으로 깨지지 않는다.
- Journal event JSON에 새 enum 값이 생긴다. 이 변경 이후 binary가 기록한 `output_truncated` settlement는
이전 binary가 역직렬화하지 못한다. RC3 Run을 RC2로 여는 것은 이미 비지원이고 `v0.1.0-rc.3` tag 전이므로
store schema version은 올리지 않는다.
- 이전 binary가 기록한 journal은 새 binary가 그대로 읽는다. `provider_limit`은 계속 유효한 값이다.

## 결과

- 사용자가 `output_truncated`를 보면 예산을 올리면 되고, `provider_limit`을 보면 기다리거나 요청을 줄이면
된다. 두 조치를 섞어 시도할 필요가 없다.
- Probe와 planner의 진단 어휘가 같은 사실을 같은 방식으로 말한다.
- Journal settlement가 provider 응답의 `finish_reason` 사실을 보존하므로 이후 분석에서 잘림 빈도를 셀 수
있다.

## 대안

- `provider_limit`을 유지하고 settlement event에 부가 필드를 넣는다: journal schema 변경은 어차피
발생하고, 사용자가 grep하는 것은 class 문자열이라 부가 필드는 진단에 쓰이지 않는다.
- `planner_invalid_response`로 분류한다: 잘림은 형식 위반이 아니라 예산 문제다. 같은 class에 넣으면 strict
schema 미준수 provider와 구분할 수 없어 ADR-0037의 진단이 흐려진다.
- 잘리면 예산을 올려 자동 재시도한다: planner 호출당 HTTP POST 1회와 불확정 호출 무재시도(ADR-0016)
원칙을 깬다. 예산은 request profile digest 입력이라 호출 중간에 바꿀 수도 없다.

## 검증

- 단위 테스트: `decode_chat_response`가 `finish_reason: "length"`에 `OutputTruncated`를 돌려주고
413/429는 `ProviderLimit`으로 남는다. `RejectionReason::ModelRejected(OutputTruncated).code()`가
`model_rejected.output_truncated`다.
- 계약 테스트: HTTP 200에 `finish_reason: "length"`를 돌려주는 서버로 AgentLoop tick을 돌리면 journal
마지막 event가 `ModelCallSettled { Rejected { OutputTruncated } }`이고 raw body는 저장되지 않는다.
- 실측: 위 재현 절차(27B, 예산 64)가 `XGENY_REJECTED … reason=model_rejected.output_truncated`를 내고,
같은 입력에 예산 1024를 주면 `XGENY_COMPLETED`다. 닫힌 포트와 429 응답은 계속 `provider_limit`이 아닌
각자의 class(`model_call_unknown.transport_unavailable`, `model_rejected.provider_limit`)를 낸다.
1 change: 1 addition & 0 deletions docs/development/durable-model-call-lifecycle.md
Original file line number Diff line number Diff line change
Expand Up @@ -169,6 +169,7 @@ Port 안에서 retry, fallback model, secondary endpoint 또는 response-repair
| valid Plan + Core/materialization 성공 | `PlanAccepted`와 모든 input sidecar가 success settlement | +1 |
| valid Completion + Core gate 성공 | `CompletionCandidateRecorded`가 success settlement | +1 |
| `PlannerPortFailure::ProviderLimit` | closed `ProviderLimit` settlement | 변화 없음 |
| `PlannerPortFailure::OutputTruncated` | closed `OutputTruncated` settlement (ADR-0039) | 변화 없음 |
| `PlannerPortFailure::ProviderRejected` | closed `ProviderRejected` settlement | 변화 없음 |
| invalid response/decode | closed rejection settlement | 변화 없음 |
| proposal/context/graph/Core validation 실패 | closed rejection settlement | 변화 없음 |
Expand Down
2 changes: 1 addition & 1 deletion docs/getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -437,7 +437,7 @@ profile 저장소를 격리하려면 `XGENY_CONFIG_HOME`을 따로 설정한다.
| `model setup`/`model check`의 `timeout` | Inference probe는 production planner와 같은 wall-clock 예산을 쓴다. 로컬 model은 첫 로드가 크기와 disk cache 상태에 따라 수십 초 걸릴 수 있으므로 endpoint에서 model을 한 번 warm-up한 뒤 재시도한다. |
| `provider_output_truncated` | Probe나 planner 응답이 출력 token 예산에서 잘렸다. Reasoning을 많이 쓰는 model은 최종 JSON 전에 예산을 소진할 수 있으므로 model의 thinking 설정이나 profile의 출력 예산을 조정한다. Rate limit이 아니므로 재시도로 해결되지 않는다. |
| `proposal_rejected.*` | 뒤의 class가 Core가 제안을 거부한 이유다. `capability_unavailable`/`capability_unsupported`는 허용하지 않은 capability 선택, `invocation_invalid`는 scope 밖 인자나 스키마 위반, `tool_call_budget_exhausted`는 예산 소진이다. Class는 Core 판정이며 model 출력 원문이 아니다. |
| `model_rejected.*` | 뒤의 class는 journal의 model call settlement와 같은 값이다. `planner_invalid_response`는 provider가 strict JSON Schema를 지키지 않은 응답(문법 미지원·미적용), `provider_limit`은 출력 예산·요청 크기 초과, `provider_rejected`는 4xx 거부다. Class는 Core 판정이며 model 출력 원문이 아니다. |
| `model_rejected.*` | 뒤의 class는 journal의 model call settlement와 같은 값이다. `planner_invalid_response`는 provider가 strict JSON Schema를 지키지 않은 응답(문법 미지원·미적용), `output_truncated`는 provider가 출력 token 예산에서 멈춰 proposal이 완성되지 않은 것(`--max-output-tokens` 또는 `XGENY_OPENAI_MAX_OUTPUT_TOKENS`를 올리거나 model의 thinking을 줄인다; 재시도로 해결되지 않는다), `provider_limit`은 요청 크기 초과(413)나 rate limit(429), `provider_rejected`는 그 밖의 4xx 거부다. Class는 Core 판정이며 model 출력 원문이 아니다. |
| `configuration_mismatch` | 원래 workspace, file/directory scope, executable와 model profile binding(inference timeout·출력 예산 포함)으로 resume한다. 자동 대체하지 말고 필요하면 새 Run을 시작한다. |
| `model_call_unknown.timeout`이 planner 호출마다 반복 | 프로필의 inference timeout이 model·hardware에 비해 짧다. 로컬 27B는 호출당 60초 안팎이 걸리므로 `--inference-timeout`을 올린다. |
| `model_call_unknown.transport_unavailable` / `.interrupted` | 요청이 전송됐을 수 있으나 결과를 못 받았다(연결 끊김) 또는 process가 응답 전에 끝났다. 자동 replay하지 않으므로 endpoint 상태를 확인한 뒤 `resume`한다. |
Expand Down