Skip to content

perf(spmc): specialize single-producer coordination - #346

Closed
tisonkun wants to merge 1 commit into
mainfrom
codex/spmc-single-producer
Closed

tisonkun wants to merge 1 commit into
mainfrom
codex/spmc-single-producer

Conversation

@tisonkun

@tisonkun tisonkun commented Oct 3, 2026

Copy link
Copy Markdown
Member

Summary

SPMC used a producer waiter queue even though send(&mut self) permits only one pending send. Replace that queue with one waker slot, and carry the exclusive producer capability into the implementation. Add cancellation, notification-race, reentrant-callback, and payload-destruction tests, plus SPMC primitive benchmarks.

Validation: cargo x test, cargo x check, cargo x lint, and cargo x miri. All 28 SPMC tests pass normally; Miri passes 26, with the two existing Tokio runtime tests excluded.

Design Notes

  • The bounded sender owns its capacity limit. Consumers only remove values: once they free space, no competing producer can claim it. Sending needs no waiter ID, node allocation, capacity grant, or cancellation handoff.
  • Keep the immediate try_send fast path. Pending sends reuse the single waker slot; receiving checks the queue and registers interest in one critical section. Unbounded publication has no capacity check or Full path.
  • Keep the shared mutex and receiver waiter queue: consumers still compete for values and must transfer unconsumed notifications on cancellation. Wake callbacks and waker/payload destruction remain outside the lock. No new unsafe code is introduced.
  • Bounded storage is preallocated, moving allocation to channel creation instead of growing the buffer while sending. This requires the full requested buffer allocation upfront; the documentation and changelog record the change.

Local measurements on Apple M4 Max, rustc 1.99.0-nightly (3d6c19bb9), against f998f35 with the same benchmark cases. Values are medians of five run medians, alternating old/new execution order.

Primitive cycle Before After Time change
Cancel pending send 19.92 ns 16.67 ns -16.3%
Wake blocked sender 29.69 ns 24.65 ns -17.0%
Wake receiver, bounded 32.46 ns 24.64 ns -24.1%
Wake receiver, unbounded 32.94 ns 24.97 ns -24.2%
Ready send/receive, bounded 14.88 ns 15.37 ns +3.3%

For the existing 16,384-message task batch with four runtime workers:

Consumers Bounded before → after Unbounded before → after
1 391.4 → 397.3 µs (+1.5%) 235.0 → 252.9 µs (+7.6%)
2 748.3 → 652.8 µs (-12.8%) 538.0 → 435.6 µs (-19.0%)
4 925.2 → 821.8 µs (-11.2%) 827.6 → 749.0 µs (-9.5%)
8 1064 → 981.4 µs (-7.8%) 817.0 → 779.5 µs (-4.6%)

The gains are concentrated in waiting and competing-consumer workloads; this is not a universal throughput improvement. In particular, the unbounded single-consumer task batch regressed in these local measurements. Current-thread task-batch changes ranged from -4.7% to +2.0%.

cargo x bench --bench primitives -- 'spmc::' --sample-count 2000
cargo x bench --bench ecosystem -- 'spmc::.*Spmc' --sample-count 100

@tisonkun tisonkun closed this Oct 5, 2026
@tisonkun
tisonkun deleted the codex/spmc-single-producer branch October 5, 2026 00:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant