Skip to content

feat(config): add per-command continue-on-error and always policies - #53

Merged
quike merged 1 commit into
mainfrom
feat/command-failure-policy
Sep 16, 2026
Merged

quike merged 1 commit into
mainfrom
feat/command-failure-policy

Conversation

@quike

@quike quike commented Sep 16, 2026 •

Copy link
Copy Markdown
Owner

Implements item 2 of #37. Item 3 (the composition decision) is the only part left.

Two independent, composable keys on map-form commands: entries:

groups:
  - name: test-with-teardown
    commands:
      - ./setup.sh                                               # starts a database
      - { command: ./optional-lint.sh, continue-on-error: true }
      - { command: go, params: [test, "./..."] }
      - { command: ./teardown.sh, always: true }                 # runs even if tests fail

continue-on-error tolerates an entry's failure: warn-logged, sequence continues, group still succeeds with exitCode 0. always guarantees the entry runs after an earlier failure but does not forgive it — the group still fails with the original error.

Why this closes a real gap

Splitting the sequence into separate groups does not work as a substitute: a failed group aborts the flow, so a downstream cleanup group never runs. And because runGroup returns the error before e.outputs.Set, a failed group stores no result, so downstream when: cannot react to it either. Before this change there was no way to guarantee cleanup at all.

always covers cancellation, not just failure

A hung test that trips timeout: is exactly when teardown matters, so pending always entries also run when the context is cancelled — detached via context.WithoutCancel and bounded by a 30-second grace period. Their own failures are logged rather than returned, since the cancellation is the real cause.

This matters for the name: in GitHub Actions if: always() runs on cancellation too, so anyone arriving with that mental model would have been surprised by an always that quietly skipped on timeout.

Demonstrated on a group with timeout: 1s whose second command sleeps for 10:

before:  setup                                   → teardown never ran
         Error: step 1: context deadline exceeded

after:   setup
         teardown-ran                            ← runs despite the timeout
         Error: step 1: command 2 of 3: run "hangs": signal: killed

The error text is also more specific now — first-error-wins names the failing command rather than reporting a bare deadline.

Semantics

  • First failure wins, over both a later always failure and a subsequent cancellation, so the diagnostic keeps pointing at the real cause.
  • Fully tolerated. A continue-on-error failure is not recorded in ExitCode/Status — a group reported as successful must not carry a non-zero exit code.
  • Per-entry only. No group-level on-error:; per-entry is strictly more expressive and a group default can be layered on later.
  • Entries between a failure and an always entry are skipped, not run.

Cache

Both keys fold into the fingerprint — unlike silent, they change whether the group succeeds, so a cached "ok" produced under continue-on-error must not be replayed after the flag is removed. Distinct markers keep the two from aliasing. Emitted only when set, so configs not using them keep their fingerprints.

Two things found while implementing

  1. A tolerated-only sequence left Status: "". RunResult's doc reserves the empty status for groups that never ran, so a group whose only entry was tolerated would have published a meaningless status. Fixed by folding StatusOK without disturbing an earlier non-ok status.
  2. RunResult.ExitCode's doc was stale. It said the field was "laid down for a future soft-fail / continue-on-error feature". That feature is this one, and with fully-tolerated semantics the field still stays 0 in stored results — corrected rather than left pointing at work that no longer implies it.

Verification

gofmt, go vet, golangci-lint, and go test -race ./... all clean.

Unit tests cover each policy and their interactions: teardown after failure, tolerated failure continuing, both keys combined, an always entry failing after a prior failure, a real failure after a tolerated one, entries being skipped, and all three cancellation paths. The test-resources/config-commands-valid.yml fixture gained a policy group, which the existing end-to-end test executes for real.

Both scenarios were also driven through the built binary rather than only through tests — the failure path returns exit 1 with the original error after running teardown, and the timeout path now does the same.

Remaining limitation

retries: replays the whole sequence, so an always teardown runs once per attempt. Pre-existing behaviour, documented in docs/CONFIG.md; keep teardowns idempotent.

@codecov

codecov Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 96.61017% with 2 lines in your changes missing coverage. Please review.
✅ Project coverage is 92.97%. Comparing base (0e2c6aa) to head (9b59575).

Files with missing lines Patch % Lines
internal/engine/engine.go 95.91% 2 Missing ⚠️
Additional details and impacted files
@@            Coverage Diff             @@
##             main      #53      +/-   ##
==========================================
+ Coverage   92.84%   92.97%   +0.12%     
==========================================
  Files          27       27              
  Lines        1524     1565      +41     
==========================================
+ Hits         1415     1455      +40     
- Misses        106      107       +1     
  Partials        3        3              
Files with missing lines Coverage Δ
internal/cache/cache.go 92.98% <100.00%> (+0.52%) ⬆️
internal/config/config.go 95.72% <100.00%> (+0.09%) ⬆️
internal/engine/engine.go 96.32% <95.91%> (+0.10%) ⬆️
🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

@quike
quike force-pushed the feat/command-failure-policy branch from c3832c4 to 9b59575 Compare September 16, 2026 20:10
@quike
quike merged commit d8653eb into main Sep 16, 2026
4 checks passed
@quike
quike deleted the feat/command-failure-policy branch September 16, 2026 20:19
@quike

quike commented Sep 16, 2026

Copy link
Copy Markdown
Owner Author

🎉 This PR is included in version 1.31.0 🎉

The release is available on GitHub release

Your semantic-release bot 📦🚀

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant