Skip to content

feat(daemon): capture the full Claude Code and Codex OTEL spec - #320

Merged
AnnatarHe merged 2 commits into
mainfrom
claude/determined-turing-85hoiv
Oct 8, 2026
Merged

AnnatarHe merged 2 commits into
mainfrom
claude/determined-turing-85hoiv

Conversation

@AnnatarHe

Copy link
Copy Markdown
Contributor

Summary

This aligns the AICodeOtel receiver with its two sources of truth: the Claude Code monitoring doc and Codex codex-rs/otel/src. It is part of a cross-repo change; the server, web and iOS PRs follow.

Data that was lost or wrong before

  • success=false and zero values never reached the server. The event struct used non-pointer fields with omitempty, so false and 0 were dropped and failed tool calls always counted as 0. Optional scalars are now pointers.
  • Timestamps were truncated to seconds. A new timestampMs field is sent; timestamp (seconds) is kept for older servers.
  • Exporter retries were stored twice. Every receipt got a random UUID. Ids are now stable ot1: hashes of the resource, scope and record, so a retry produces the same id.
  • Every attribute outside a fixed switch was dropped. Unmapped attributes now go into an attributes catch-all (64 keys, 2 KB per value).
  • New first-class fields: promptId (prompt.id), sequence (event.sequence), requestId, speed, querySource, response / responseLength (assistant_response, Codex agent_response), toolInput (Claude tool_input, Codex arguments, raw), entrypoint (app.entrypoint / Codex originator / service.name).
  • Spec mappings that were missing:
    • tool_use_id → callId
    • Codex output → toolOutput
    • cache_write_token_count → cacheCreationTokens
    • cost_usd_micros used as a fallback for cost
    • effort / model_reasoning_effort → reasoningEffort
    • mcp_servers sent as a comma-joined string is now parsed
    • host.name used as the machine-name fallback
    • Codex tool_token_count (which is really total_tokens) no longer lands in toolTokens
  • Event names: read from event.name, then LogRecord.EventName, then the body, with a generic prefix strip. The opt-in raw API bodies and system_prompt are dropped.
  • Metrics:
    • type is routed per metric; active_time user/cli no longer lands in tokenType.
    • tool_name is accepted on code_edit_tool.decision.
    • The invented codex.* metric names are removed (Codex exports no such metrics).
    • A warning is logged once if temporality is cumulative.
  • Size limits: the gRPC receive limit is raised to 32 MB, and outgoing requests are split at 8 MB.

Install and config

  • shelltime cc install adds OTEL_LOG_TOOL_DETAILS=1, OTEL_LOG_ASSISTANT_RESPONSES=1, OTEL_METRICS_INCLUDE_ENTRYPOINT=true, OTEL_METRICS_INCLUDE_REPOSITORY=true and OTEL_EXPORTER_OTLP_METRICS_TEMPORALITY_PREFERENCE=delta.
  • shelltime codex install merges into [otel] instead of replacing the table. It keeps the user's other keys and adds log_agent_responses = true. Uninstall removes only the keys it manages.
  • doctor warns when a config is missing the new keys and offers the existing auto-fix.
  • Privacy: both installs print a note about what is now sent (bash commands, file paths, truncated tool input, tool errors, assistant responses).

Compatibility

The payload changes are additive. Older servers ignore the new keys and already decode optional fields as pointers. The server PR stores the new fields.

Test plan

  • go vet ./...
  • go test -timeout 3m ./... passes: commands, daemon, model, stloader.
  • New daemon/aicode_otel_processor_v2_test.go: table tests built from the spec's record shapes (15 Claude and 10 Codex cases), plus the drop list, name fallbacks, stable ids, metric routing, attribute caps and chunked sending.
  • gRPC round-trip test asserts the raw JSON carries "success":false, zero cost and duration, timestampMs and ot1: ids; a 6 MB export is accepted.
  • Codex config tests: merge keeps user keys, reinstall is idempotent, uninstall removes only the managed keys.

🤖 Generated with Claude Code

https://claude.ai/code/session_01EsMo3p9YUBzdsjFHxGWoCL


Generated by Claude Code

Align the AICodeOtel receiver with the Claude Code monitoring doc and
Codex codex-rs/otel so no attribute is silently dropped:

- Optional scalars are pointers, so success=false and zero values reach
  the server (failed tool calls were always counted as 0).
- Payload v2 fields: timestampMs, promptId, sequence, requestId, speed,
  querySource, response, responseLength, toolInput, entrypoint and an
  attributes catch-all for every unmapped attribute (capped).
- Stable ot1: event and metric ids so exporter retries de-duplicate.
- Event names: event.name, then LogRecord.EventName, then Body; generic
  claude_code./codex. prefix strip; raw API bodies and system_prompt are
  dropped.
- Spec mappings: tool_use_id, tool_input/arguments, Codex output,
  cache_write_token_count, cost_usd_micros, effort, string mcp_servers,
  host.name and originator fallbacks; tool_token_count (total tokens) no
  longer lands in toolTokens.
- Metrics: type routed per metric, tool_name accepted, fabricated codex.*
  names removed, cumulative temporality warning.
- gRPC receive limit 32 MB; outgoing requests split at 8 MB.
- cc install adds OTEL_LOG_TOOL_DETAILS, OTEL_LOG_ASSISTANT_RESPONSES,
  entrypoint/repository attributes and delta temporality; codex install
  merges into [otel] and enables log_agent_responses; doctor flags stale
  configs; installs print a privacy note.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EsMo3p9YUBzdsjFHxGWoCL
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@codecov

codecov Bot commented Oct 8, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 90.58615% with 53 lines in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
daemon/aicode_otel_processor.go 88.20% 46 Missing ⚠️
model/codex_otel_config.go 95.00% 3 Missing ⚠️
model/api_aicode_otel.go 93.93% 2 Missing ⚠️
model/aicode_backfill_claude.go 90.00% 1 Missing ⚠️
model/aicode_otel_claude_settings.go 94.44% 1 Missing ⚠️
Flag Coverage Δ
unittests 85.85% <90.58%> (?)

Flags with carried forward coverage won't be shown. Click here to find out more.

Files with missing lines Coverage Δ
commands/cc.go 81.81% <100.00%> (+3.24%) ⬆️
commands/codex.go 100.00% <100.00%> (ø)
commands/doctor_checks.go 95.12% <100.00%> (+0.13%) ⬆️
daemon/aicode_otel_server.go 95.23% <100.00%> (ø)
model/aicode_backfill_codex.go 85.76% <100.00%> (ø)
model/aicode_backfill_common.go 91.97% <ø> (+0.24%) ⬆️
model/aicode_otel_types.go 100.00% <100.00%> (ø)
model/aicode_backfill_claude.go 89.43% <90.00%> (ø)
model/aicode_otel_claude_settings.go 87.16% <94.44%> (+0.77%) ⬆️
model/api_aicode_otel.go 95.83% <93.93%> (-4.17%) ⬇️
... and 2 more

... and 3 files with indirect coverage changes

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@github-actions

github-actions Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

⏱️ shelltime track performance

✅ No significant change in shelltime track latency

Comparing main @ d82ea6b → #320 @ 4ffd4b3. Each scenario execs the binary the way the shell hooks do; values are the median of 10 interleaved rounds. Lower is better.

Scenario base head Δ p
Startup/floor 487 µs ±1% 488 µs ±2% ~ 0.971 A/A
Startup/version 3.58 ms ±1% 3.55 ms ±1% ~ 0.165 ✅
Track/daemon/pre 3.92 ms ±1% 3.88 ms ±1% -1.21% 0.001 ✅
Track/daemon/post 3.88 ms ±1% 3.85 ms ±1% ~ 0.075 ✅
Track/direct/pre 3.85 ms ±1% 3.82 ms ±1% -0.67% 0.023 ✅
Track/direct/post 7.42 ms ±1% 7.33 ms ±1% -1.16% 0.003 ✅
Track/direct/post-sync 9.04 ms ±2% 8.99 ms ±1% ~ 0.052 ✅

Binary size: 19.1 MiB → 19.1 MiB (+0.0%).

CPU, tail latency and memory

Per exec, base → head (Δ when significant).

Scenario p95 wall user CPU sys CPU peak RSS
Startup/floor 548 µs → 544 µs 276 µs → 278 µs 159 µs → 161 µs 12.2 MiB → 12.2 MiB
Startup/version 3.89 ms → 3.91 ms 1.71 ms → 1.63 ms 2.35 ms → 2.34 ms 18.0 MiB → 18.1 MiB
Track/daemon/pre 4.28 ms → 4.20 ms (-2.03%) 1.96 ms → 1.94 ms 2.48 ms → 2.42 ms 18.9 MiB → 19.0 MiB
Track/daemon/post 4.21 ms → 4.23 ms 1.93 ms → 1.92 ms 2.45 ms → 2.41 ms 18.8 MiB → 18.8 MiB
Track/direct/pre 4.18 ms → 4.12 ms (-1.46%) 1.84 ms → 1.89 ms 2.47 ms → 2.41 ms 18.7 MiB → 18.9 MiB (+0.97%)
Track/direct/post 8.11 ms → 7.98 ms (-1.57%) 5.21 ms → 5.19 ms 3.44 ms → 3.41 ms 21.1 MiB → 21.1 MiB
Track/direct/post-sync 9.83 ms → 9.87 ms 6.98 ms → 6.90 ms 4.63 ms → 4.49 ms 23.4 MiB → 23.4 MiB
benchstat
goos: linux
goarch: amd64
pkg: github.com/malamtime/cli/perf
cpu: AMD EPYC 7763 64-Core Processor                
                         │    base     │                head                │
                         │   sec/op    │   sec/op     vs base               │
Startup/floor-4            487.2µ ± 1%   488.0µ ± 2%       ~ (p=0.971 n=10)
Startup/version-4          3.577m ± 1%   3.549m ± 1%       ~ (p=0.165 n=10)
Track/daemon/pre-4         3.923m ± 1%   3.875m ± 1%  -1.21% (p=0.001 n=10)
Track/daemon/post-4        3.884m ± 1%   3.854m ± 1%       ~ (p=0.075 n=10)
Track/direct/pre-4         3.848m ± 1%   3.822m ± 1%  -0.67% (p=0.023 n=10)
Track/direct/post-4        7.415m ± 1%   7.330m ± 1%  -1.16% (p=0.003 n=10)
Track/direct/post-sync-4   9.043m ± 2%   8.989m ± 1%       ~ (p=0.052 n=10)
geomean                    3.532m        3.506m       -0.72%

                         │    base     │                head                │
                         │ p50-sec/op  │ p50-sec/op   vs base               │
Startup/floor-4            478.8µ ± 1%   477.8µ ± 1%       ~ (p=0.912 n=10)
Startup/version-4          3.542m ± 2%   3.507m ± 2%  -1.00% (p=0.043 n=10)
Track/daemon/pre-4         3.910m ± 1%   3.860m ± 1%  -1.29% (p=0.000 n=10)
Track/daemon/post-4        3.859m ± 2%   3.814m ± 1%       ~ (p=0.063 n=10)
Track/direct/pre-4         3.840m ± 1%   3.825m ± 2%       ~ (p=0.143 n=10)
Track/direct/post-4        7.368m ± 1%   7.287m ± 2%  -1.10% (p=0.029 n=10)
Track/direct/post-sync-4   9.018m ± 1%   8.935m ± 1%  -0.92% (p=0.001 n=10)
geomean                    3.507m        3.477m       -0.87%

                         │    base     │                head                │
                         │ p95-sec/op  │ p95-sec/op   vs base               │
Startup/floor-4            548.5µ ± 3%   544.0µ ± 2%       ~ (p=0.436 n=10)
Startup/version-4          3.893m ± 2%   3.911m ± 3%       ~ (p=0.579 n=10)
Track/daemon/pre-4         4.282m ± 2%   4.196m ± 1%  -2.03% (p=0.043 n=10)
Track/daemon/post-4        4.208m ± 2%   4.225m ± 3%       ~ (p=0.971 n=10)
Track/direct/pre-4         4.184m ± 1%   4.123m ± 1%  -1.46% (p=0.011 n=10)
Track/direct/post-4        8.107m ± 3%   7.980m ± 2%  -1.57% (p=0.035 n=10)
Track/direct/post-sync-4   9.832m ± 2%   9.865m ± 2%       ~ (p=0.579 n=10)
geomean                    3.863m        3.837m       -0.67%

                         │     base     │                head                 │
                         │  peak-rss-B  │  peak-rss-B   vs base               │
Startup/floor-4            12.23Mi ± 2%   12.22Mi ± 1%       ~ (p=0.928 n=10)
Startup/version-4          17.97Mi ± 1%   18.06Mi ± 1%       ~ (p=0.218 n=10)
Track/daemon/pre-4         18.94Mi ± 1%   18.99Mi ± 1%       ~ (p=0.247 n=10)
Track/daemon/post-4        18.80Mi ± 1%   18.84Mi ± 1%       ~ (p=0.529 n=10)
Track/direct/pre-4         18.72Mi ± 1%   18.90Mi ± 1%  +0.97% (p=0.001 n=10)
Track/direct/post-4        21.11Mi ± 1%   21.13Mi ± 1%       ~ (p=0.631 n=10)
Track/direct/post-sync-4   23.39Mi ± 1%   23.37Mi ± 0%       ~ (p=0.529 n=10)
geomean                    18.43Mi        18.48Mi       +0.27%

                         │    base     │                head                 │
                         │ sys-sec/op  │  sys-sec/op   vs base               │
Startup/floor-4            159.1µ ± 8%   161.4µ ± 15%       ~ (p=0.631 n=10)
Startup/version-4          2.354m ± 5%   2.342m ±  4%       ~ (p=0.912 n=10)
Track/daemon/pre-4         2.476m ± 4%   2.416m ±  4%       ~ (p=0.190 n=10)
Track/daemon/post-4        2.449m ± 5%   2.410m ±  8%       ~ (p=0.739 n=10)
Track/direct/pre-4         2.469m ± 6%   2.414m ±  5%       ~ (p=0.218 n=10)
Track/direct/post-4        3.443m ± 2%   3.407m ±  4%       ~ (p=0.684 n=10)
Track/direct/post-sync-4   4.629m ± 5%   4.490m ±  5%       ~ (p=0.247 n=10)
geomean                    1.900m        1.874m        -1.35%

                         │    base     │                head                │
                         │ user-sec/op │ user-sec/op  vs base               │
Startup/floor-4            275.6µ ± 6%   277.8µ ± 5%       ~ (p=0.912 n=10)
Startup/version-4          1.708m ± 6%   1.633m ± 6%       ~ (p=0.247 n=10)
Track/daemon/pre-4         1.962m ± 5%   1.939m ± 6%       ~ (p=0.912 n=10)
Track/daemon/post-4        1.930m ± 7%   1.921m ± 7%       ~ (p=0.529 n=10)
Track/direct/pre-4         1.844m ± 9%   1.895m ± 4%       ~ (p=0.579 n=10)
Track/direct/post-4        5.209m ± 3%   5.186m ± 6%       ~ (p=0.529 n=10)
Track/direct/post-sync-4   6.978m ± 3%   6.897m ± 5%       ~ (p=0.353 n=10)
geomean                    1.980m        1.968m       -0.60%

Δ compares medians; ~ means no significant difference (Mann-Whitney U, p ≥ 0.05). A scenario is flagged when it is more than 10% and more than 250µs slower with p < 0.05. Startup/floor runs true, not shelltime: it is an A/A check of runner noise. Runner: linux/amd64, AMD EPYC 7763 64-Core Processor . go1.27.1, 10 rounds × 100 execs per scenario and binary.

@claude

claude Bot commented Oct 8, 2026

Copy link
Copy Markdown
Contributor

Code review

No issues found. Checked for bugs and CLAUDE.md compliance.

Copy link
Copy Markdown
Contributor Author

/claude review


Generated by Claude Code

@AnnatarHe

Copy link
Copy Markdown
Contributor Author

@claude review

@claude

claude Bot commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @AnnatarHe's task in 42s —— View job


Review of #320: capture the full Claude Code and Codex OTEL spec

I read the diff for the processor, the request splitter, the Claude and Codex config services and the gRPC server. I did not re-run the tests. The PR reports go vet and go test passing, and the earlier automated review found nothing. The change is well structured and the table-driven tests are thorough. I found no blocking bugs, but a few points are worth deciding on before merge.

Worth addressing

  1. Existing Codex users start sending tool output without a new opt-in. daemon/aicode_otel_processor.go now maps Codex output to toolOutput (capped at AICodeOtelMaxTextBytes). Before, only tool_output / toolOutput were read, and Codex never emits those names, so tool output was dropped. Command output and file contents can include secrets. The privacy note only prints on shelltime codex install / cc install. Anyone who installed earlier and just upgrades the daemon starts sending this data with no notice. Options:

    • Gate toolOutput on a config flag.
    • Have doctor or the daemon log a one-time notice about the new fields.
    • Confirm that it is acceptable for the server to receive this.
  2. The attributes catch-all forwards every unmapped attribute. The default: branches in applyEventAttribute and extractResourceAttributes send everything, within 64 keys and 2 KB per value. This is deliberate and capped. But it means any field a future Claude or Codex version adds is sent automatically, including fields that are sensitive by default. An allow/deny list for known-sensitive keys, or a note in the PR description, would make the data scope explicit.

  3. A chunked send can outlast the exporter's timeout. send posts chunks one after another from inside the gRPC handler. The receive limit is now 32 MB and chunks are 8 MB, so one export can mean several sequential HTTP calls. If the total exceeds the exporter's timeout (the OTLP default is 10 s), the exporter retries the whole batch. The stable ot1: ids make that safe if the server dedupes on id. Sending chunks concurrently, or running send after returning the response, would remove the risk. Please confirm the server-side dedupe covers this.

Minor

  • Check() can report a false "installed". In model/codex_otel_config.go, Check() returns true whenever an [otel] table exists. Install now merges instead of replacing. A user with only environment = ... set therefore looks installed. Checking for the ShellTime exporter endpoint, as Endpoint() does, would be more precise.
  • Uninstall can remove a pre-existing log_user_prompt. Uninstall deletes log_user_prompt unconditionally. If the user set it before installing, their setting is removed. This matches "removes only managed keys", but it is a small surprise.
  • Misplaced comment. In extractResourceAttributes, "Kept as fields and also forwarded in the attributes." sits under the project / project.path case, but it describes the service.name case below it.
  • Identical records collide on id. stableMessageID hashes the full LogRecord, so two records that are byte-identical (same nanosecond timestamp, same attributes) get the same id and the server will treat them as one. Claude's event.sequence avoids this. For Codex it should be rare. It is worth stating as an accepted limitation.
  • Chunk failures drop data. A failed chunk is logged and skipped, and the other chunks are still sent. That fits passthrough mode, but the server can then see partial batches.

Looks good

  • Pointer fields for success, cost, duration and the token counts fix the dropped false and 0 values.
  • timestampMs is added while timestamp stays for older servers.
  • The length-prefixed hash in stableOtelID prevents collisions between different splits.
  • Codex tool_token_count is correctly treated as total tokens.
  • Cumulative-temporality warning is logged once through sync.Once.
  • MissingManagedKeys respects a user's explicit opt-out value, such as OTEL_LOG_TOOL_DETAILS=0.

Install now merges into the user's [otel] table, so Check() returning
true for any [otel] table made a config with only `environment` (or an
exporter pointing at another collector) look installed. Check() now
requires the exporter to point at the ShellTime daemon.

Also move a misplaced comment in extractResourceAttributes to the
service.name case it describes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EsMo3p9YUBzdsjFHxGWoCL

Copy link
Copy Markdown
Contributor Author

Thanks for the review. Changes in 4ffd4b3:

  • Check() false "installed". Fixed. It now returns true only when the [otel] exporter points at the ShellTime daemon. Before, any [otel] table counted, including one with only environment or an exporter pointing at another collector. Covered by a new test.
  • Misplaced comment. Moved to the service.name case it describes.

Not changed, with reasons:

  • Codex output and the attributes catch-all sent without a new opt-in. This is the agreed collection level: tool details plus responses. The install privacy note already lists "truncated tool output". The catch-all implements the "nothing silently dropped" goal and is capped at 64 keys × 2 KB.
  • Chunked send outlasting the exporter timeout. A retried batch carries the same ot1: ids. The server namespaces them per user and inserts with ON CONFLICT(event_id) DO NOTHING, so a whole-batch retry is absorbed (covered by TestOtelIDsAreNamespacedPerUser in shelltime/server#517).
  • Uninstall removing a pre-set log_user_prompt. This is the documented behaviour of removing the keys ShellTime manages.
  • Byte-identical records colliding. This needs the same nanosecond timestamp and identical attributes. I'm treating it as an accepted limitation; Claude events also carry event.sequence.
  • A failed chunk dropping data. This matches the existing passthrough design. Each chunk is independent, and the server accepts partial batches.

go vet passes and go test is green for model, daemon and commands.


Generated by Claude Code

@AnnatarHe
AnnatarHe merged commit cb27a72 into main Oct 8, 2026
5 checks passed
@AnnatarHe
AnnatarHe deleted the claude/determined-turing-85hoiv branch October 8, 2026 10:49
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants