Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
160 changes: 160 additions & 0 deletions .github/workflows/perf.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,160 @@
name: Track Perf

# `shelltime track` runs in the foreground of every shell prompt, so its latency
# is user-facing. This builds the base and head binaries and times the real
# process in interleaved rounds on one runner (perf/compare.sh), then reports
# per-scenario changes in the job summary and, on PRs, in a sticky comment.
# Regressions are warnings (annotations + report), never a failed check.

on:
pull_request:
branches:
- main
types:
- opened
- synchronize
- reopened
push:
branches:
- main
workflow_dispatch:
inputs:
base:
description: Commit to compare against (default HEAD^1)
required: false
default: ""
rounds:
description: Interleaved rounds per binary
required: false
default: "10"

permissions:
contents: read

concurrency:
group: track-perf-${{ github.event.pull_request.number || github.ref }}
# Superseded PR pushes are cancelled; every main commit keeps its report.
cancel-in-progress: ${{ github.event_name == 'pull_request' }}

jobs:
track-perf:
name: shelltime track latency
runs-on: ubuntu-24.04
timeout-minutes: 30
permissions:
contents: read
pull-requests: write # sticky comment; read-only on fork PRs, which get the summary only
env:
# Build both sides with head's toolchain, so a change shows the code's
# effect, not the compiler's.
GOTOOLCHAIN: local
ROUNDS: ${{ inputs.rounds || '10' }}
ITERS: "100"
THRESHOLD: "10"
steps:
- name: Checkout code
uses: actions/checkout@v5
with:
fetch-depth: 0

- name: Setup Go
uses: actions/setup-go@v5
with:
go-version-file: go.mod
cache-dependency-path: |
go.sum
perf/go.sum

- name: Test the perf tooling
run: |
go -C perf vet ./...
go -C perf test ./...

- name: Pick the base commit
id: base
env:
EVENT: ${{ github.event_name }}
BEFORE: ${{ github.event.before }}
INPUT_BASE: ${{ inputs.base }}
PR_NUMBER: ${{ github.event.pull_request.number }}
PR_HEAD: ${{ github.event.pull_request.head.sha }}
run: |
case "$EVENT" in
push) base=$BEFORE ;;
workflow_dispatch) base=$INPUT_BASE ;;
# pull_request checks out the merge commit: its first parent
# is exactly the base the PR is being merged onto.
*) base= ;;
esac
if [[ -z $base ]] || ! git cat-file -e "$base^{commit}" 2>/dev/null; then
base=$(git rev-parse HEAD^1)
fi
base=$(git rev-parse "$base^{commit}")

# Harness changes need a full run even when the binary is unchanged.
force=0
git diff --quiet "$base" HEAD -- perf .github/workflows/perf.yaml || force=1

if [[ $EVENT == pull_request ]]; then
base_label="\`$GITHUB_BASE_REF\` @ \`${base::7}\`"
head_label="#$PR_NUMBER @ \`${PR_HEAD::7}\`"
else
base_label="\`${base::7}\`"
head_label="\`$(git rev-parse --short=7 HEAD)\`"
fi
{
echo "sha=$base"
echo "force=$force"
echo "base_label=$base_label"
echo "head_label=$head_label"
} >> "$GITHUB_OUTPUT"

- name: Benchmark base vs head
id: perf
env:
BASE_SHA: ${{ steps.base.outputs.sha }}
FORCE: ${{ steps.base.outputs.force }}
BASE_LABEL: ${{ steps.base.outputs.base_label }}
HEAD_LABEL: ${{ steps.base.outputs.head_label }}
run: perf/compare.sh "$BASE_SHA" WORKTREE "$RUNNER_TEMP/perf"

- name: Job summary
if: always()
run: |
report=$RUNNER_TEMP/perf/report.md
if [[ -f $report ]]; then
cat "$report" >> "$GITHUB_STEP_SUMMARY"
fi

- name: Comment on the PR
if: >-
always() &&
github.event_name == 'pull_request' &&
github.event.pull_request.head.repo.full_name == github.repository
env:
GH_TOKEN: ${{ github.token }}
PR_NUMBER: ${{ github.event.pull_request.number }}
SKIPPED: ${{ steps.perf.outputs.skipped }}
run: |
report=$RUNNER_TEMP/perf/report.md
if [[ ! -f $report ]]; then
exit 0
fi
# An unchanged binary refreshes an existing report but never
# starts a new comment thread.
if [[ $SKIPPED == true ]]; then
perf/comment.sh "$PR_NUMBER" "$report" --update-only
else
perf/comment.sh "$PR_NUMBER" "$report"
fi

- name: Upload results
if: always()
uses: actions/upload-artifact@v4
with:
name: track-perf
path: |
${{ runner.temp }}/perf/*.txt
${{ runner.temp }}/perf/*.md
${{ runner.temp }}/perf/*.log
if-no-files-found: ignore
3 changes: 3 additions & 0 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,6 +15,7 @@ This is a Go monorepo for the ShellTime CLI and daemon.
- `model/`: config, API clients, shell integrations, crypto, and shared domain logic
- `docs/`: user-facing docs such as `CONFIG.md` and `CC_STATUSLINE.md`
- `fixtures/`: reusable test fixtures
- `perf/`: separate Go module that benchmarks the real `shelltime track` process, plus the CI comparison tooling

Keep new code inside the existing package boundary. Do not mix CLI wiring, daemon internals, and model logic in the same package.

Expand All @@ -31,6 +32,8 @@ Keep new code inside the existing package boundary. Do not mix CLI wiring, daemo
- `go vet ./...`: run static analysis
- `mockery`: regenerate mocks when interfaces change
- `pp g`: regenerate PromptPal-generated artifacts when relevant
- `go -C perf test -run '^$' -bench . -benchtime 50x -count 5`: benchmark `shelltime track` latency on the current tree (`perf/` is its own module; see `perf/README.md`)
- `perf/compare.sh origin/main`: compare `track` latency of a base commit and the working tree, as the `Track Perf` workflow does on every PR and push to `main`

Use Go 1.27.1, as declared in `go.mod`.

Expand Down
11 changes: 11 additions & 0 deletions CLAUDE.md
Original file line number Diff line number Diff line change
Expand Up @@ -40,6 +40,17 @@ go test -run TestHandlerName ./daemon/

Tests use **testify** (assertions + suites). Suite-based tests use `suite.Suite` with `SetupTest`/`TearDownTest` lifecycle hooks (see `daemon/cc_info_handler_test.go` for example). Simple functions use table-driven tests.

### Performance (`shelltime track` latency)
`perf/` is a separate Go module (root `./...` skips it). It benchmarks the real binary the way the shell hooks run it: daemon and direct paths, pre and post, plus a sync. See `perf/README.md`.
```bash
# Benchmark the current tree
go -C perf test -run '^$' -bench . -benchtime 50x -count 5

# Compare a base commit with the working tree, as CI does
perf/compare.sh origin/main
```
The `Track Perf` workflow (`.github/workflows/perf.yaml`) runs this on every PR and push to `main`. It posts the report as a PR comment and in the job summary. A scenario that gets more than 10% slower (and more than 0.25 ms, with p < 0.05) raises a warning, never a failure. `Track/*` scenarios skip while a real daemon owns `/tmp/shelltime.sock`.

### Code Generation
```bash
# Generate mocks (uses .mockery.yml configuration, Mockery v3)
Expand Down
91 changes: 91 additions & 0 deletions perf/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,91 @@
# `shelltime track` performance tests

The shell hooks (`model/hooks/`) run `shelltime track` in the foreground before
and after **every** command, so its latency sits directly between the user and
their next prompt. The benchmarks here exec the real binary, with the same flags
the zsh hook passes, in an isolated `$HOME`. That measures what users feel:
exec, Go runtime and package init, `main`'s config read and telemetry setup,
then the track logic.

This directory is its own Go module, so `golang.org/x/perf` never takes part in
version selection for the shipped binaries, and the root `go test ./...` skips
it.

## Scenarios

| Benchmark | What it measures |
|---|---|
| `Startup/floor` | Runs `true`, not shelltime: the runner's fork/exec cost. Base and head run the same thing, so this row is an A/A check of runner noise |
| `Startup/version` | `shelltime --version`: process start, init, config read, no command logic |
| `Track/daemon/pre`, `/post` | The default install. A fake daemon listens on `/tmp/shelltime.sock`; track sends it the event and returns |
| `Track/direct/pre` | No daemon. track appends to `~/.shelltime/commands/pre.txt` |
| `Track/direct/post` | No daemon, below the flush threshold. track appends to `post.txt` and reads the whole store (500 synced pairs) to decide whether to sync |
| `Track/direct/post-sync` | The post that reaches `flushCount` and sends the batch to a local fake API (once every 10 commands in direct mode) |

Every scenario checks that the binary really took its path: the fake daemon got
one event per run, the store grew, the API got one batch per run, and
`log.log` holds no errors. `track` always exits 0, so without these checks a
broken run would just look fast. A failed check fails that scenario only, and
the report shows it as ❌.

Besides wall time (`sec/op`), each scenario reports per-exec user and sys CPU,
p50/p95 wall time and peak RSS of the child process.

On Linux the direct post path execs `lsb_release`. On Ubuntu that is a Python
script that adds tens of noisy milliseconds, so the harness puts a shell stub
with the same output first in `PATH`. Set `SHELLTIME_BENCH_REAL_OSINFO=1` to
measure the real one.

## Running locally

```sh
# Benchmark the current tree (builds ./cmd/cli once)
go -C perf test -run '^$' -bench . -benchtime 50x -count 5

# Benchmark a specific binary
SHELLTIME_BENCH_BIN=/path/to/shelltime go -C perf test -run '^$' -bench . -benchtime 50x

# Reproduce the CI comparison: base commit vs the working tree, uncommitted changes included
perf/compare.sh origin/main
ROUNDS=4 ITERS=30 perf/compare.sh HEAD~1 HEAD /tmp/perf # quicker, explicit head and output dir
```

The `Track/*` scenarios are skipped while a real shelltime daemon owns
`/tmp/shelltime.sock`, which is the usual state of a developer machine: the
daemon would receive the fake commands. Stop the daemon to run them. The socket
path is fixed in the CLI, so it cannot be redirected.

## In CI

`.github/workflows/perf.yaml` runs on every pull request to `main`, every push
to `main` and on demand. It compares:

- a pull request against the commit it is merged onto (`HEAD^1` of the merge commit),
- a push to `main` against the previous `main` tip (`github.event.before`).

`compare.sh` builds both binaries with identical flags (`CGO_ENABLED=0
-trimpath -buildvcs=false -ldflags "-s -w"`). If they come out byte-identical,
for example on a docs-only change, it skips the run, unless the harness itself
changed. Otherwise it compiles the harness once from head and runs it against
both binaries in 10 rounds of 100 execs per scenario. The rounds alternate in
ABBA order on the same runner, so drift and noise hit both sides alike.

The report goes to the job summary and, on pull requests from this repository,
to a single comment that later runs update. Fork pull requests get a read-only
token, so they only get the summary.

### Reading the report

Values are medians across rounds, with a 95% confidence interval. A scenario is
flagged ⚠️ when it is **more than 10% slower, more than 0.25 ms slower and
significant (Mann-Whitney U, p < 0.05)**. The absolute floor keeps
sub-millisecond scenarios from flagging on scheduler jitter. ~ means no
significant difference. 🚀 is the same rule in the other direction.

If `Startup/floor` itself moved significantly, the headline says the runner was
noisy. Treat small changes in that run with care.

A regression never fails the check. It shows up as ⚠️ in the report and as a
warning annotation on the run. Rerun the job if you suspect noise, or reproduce
it locally with `compare.sh`. To tune the gate, change `THRESHOLD` in the
workflow, or `-min-delta` in `cmd/perfreport`.
Loading
Loading