Skip to content

PipeWire and ALSA raw: port lifecycle fixes, and one meaning for the observer transport filters - #256

Merged
jcelerier merged 5 commits into
masterfrom
pipewire-alsa-lifecycle-fixes
Sep 13, 2026
Merged

jcelerier merged 5 commits into
masterfrom
pipewire-alsa-lifecycle-fixes

Conversation

@jcelerier

Copy link
Copy Markdown
Member

Three independent backend fixes, one per commit, found while building a MIDI device tree in ossia score against real hardware (M-Audio Keystation Pro 88) on PipeWire, ALSA sequencer and ALSA raw.

They can be reviewed and taken separately.

1. PipeWire: a per-object error is not the connection dying

The visible symptom: create a port, destroy it, create another — the second open fails, on any port, for the rest of the process. It only happens when an observer is alive across the cycle, which in an application is always, because the port picker holds one.

Closing a port races the daemon over who destroys the link. pw_filter_new_simple() calls pw_context_new() rather than reusing ours, so the filter's node and its link live on two different sockets; the daemon serves them independently, and when it destroys the node first that cascades into the link's global. Our already-queued core.destroy() then names a resource that is gone, and the daemon answers -ENOENT, "unknown resource N op:7" — raised via pw_resource_errorf(client->core_resource, …) and therefore carrying id == PW_ID_CORE.

on_core_error read that as connection loss. Grounding the classification in PipeWire's own source: a genuine socket failure has exactly one origin, on_remote_data in module-protocol-native.c, which reports id = 0, res = -EPIPE. res is the whole discriminator, and it is the one every first-party client uses — pw-cli, pw-dump, pw-link and pw-mon quit on id == PW_ID_CORE && res == -EPIPE and ignore the rest; module-loopback.c logs -ENOENT at info level right beside that test.

Three changes follow:

  • connection_lost(id, res) states the rule. A per-object complaint also stops recording a sync error — finalize_sync() turns any recorded error into broken, so recording one ended a round-trip the daemon was still answering.
  • shared_context() returned the cached context without checking it, so once broken it was handed to every later port for as long as any holder kept it alive. That is what made the failure permanent. It now rebuilds in place, preserving that holder's subscriptions, and drops the cache when it cannot.
  • The filter is created on our own core, so node and links share one connection and the daemon sees their teardown in the order it was sent. pw_filter_new() does not take the events, so the listener is attached explicitly; the old entry point stays as a fallback where the symbol is unavailable.

tests/integration/pipewire_context_object_errors.cpp provokes the race deterministically with no MIDI hardware, and fails against the previous handler. It is separate from pipewire_context_error_scope on purpose: that one covers id != PW_ID_CORE, while the error that causes this arrives on the core id and would pass it.

2. alsa_raw: opening a busy output port blocked instead of failing

A raw MIDI device is exclusive, and snd_rawmidi_open() on one already held waits in the kernel instead of returning EBUSY. midi_out asked for SND_RAWMIDI_SYNC alone, so a second open hung the caller outright with no way to report the port was in use. midi_in already opened non-blocking, which is why only output showed it.

It now opens non-blocking and clears the flag immediately with snd_rawmidi_nonblock(handle, 0): the mode is wanted only for the open, since writes must still block until the device takes the bytes or a burst is silently truncated. The new symbol goes through the existing loader and the call is guarded, so an older libasound keeps today's behaviour.

3. observer: one meaning for the transport filters on every backend

track_hardware / track_virtual / track_network were interpreted by whichever backend implemented them, so the same configuration selected different ports depending on what was compiled in — ALSA sequencer counted anything it could not classify as hardware, ALSA raw dropped it, CoreMIDI had no notion of a virtual port, JACK and WinMM disagreed again.

observer_configuration::accepts(transport_type) states the rule once, and each backend answers only what it can actually answer: which transport a port is on. A port the platform genuinely cannot classify now belongs to no group rather than being forced into one, so a filter no longer silently gains or loses ports.

Testing

  • ctest in this repo: 18/18, including all four PipeWire integration tests plus the new one.
  • Against real hardware through ossia score, all three backends open/destroy/reopen cleanly with an observer alive: PipeWire (was failing 4 runs in 5), ALSA sequencer, ALSA raw (was hanging).
  • CoreMIDI behaviour for (3) was validated earlier on an M1 Mac against an Arturia keyboard and virtual/network ports.

🤖 Generated with Claude Code

https://claude.ai/code/session_01UH29eo7sTWNSkDx8XSfPpt

@jcelerier
jcelerier force-pushed the pipewire-alsa-lifecycle-fixes branch from b471f22 to 8865eca Compare September 13, 2026 02:38
jcelerier and others added 3 commits September 13, 2026 00:39
Opening a port, closing it and opening another failed for the rest of the
process, on any port, as soon as one observer stayed alive across the
cycle.

Closing a port races the daemon over who destroys the link. The filter's
node and the link live on two different sockets -- pw_filter_new_simple()
calls pw_context_new() rather than reusing ours -- and the daemon serves
them independently, so destroying the node first cascades into the link's
global and our own core.destroy() then names a resource that is already
gone. The daemon answers -ENOENT, "unknown resource N op:7", raised on
the core resource and therefore carrying id == PW_ID_CORE.

Three things follow, and this fixes each:

- on_core_error treated that as connection loss. res is the whole
  discriminator: a real socket failure is uniquely -EPIPE from
  on_remote_data, and pw-cli, pw-dump, pw-link and pw-mon all quit on
  that alone. connection_lost() states the rule, and a per-object
  complaint no longer records a sync error either -- finalize_sync()
  turns any recorded error into `broken`, so recording one ended the
  round-trip the daemon was still answering.

- shared_context() handed out the cached context without checking it.
  Once broken it was given to every later port for as long as any holder
  kept it alive, which is what made the failure permanent. It now
  rebuilds in place, which keeps that holder's subscriptions, and drops
  the cache if it cannot.

- the filter is created on our own core, so the node and its links share
  one connection and the daemon sees their teardown in the order it was
  sent. pw_filter_new() does not take the events, so the listener is
  added explicitly; the old entry point remains as a fallback.

The new test provokes the race deterministically and needs no MIDI
hardware: ask for a factory that does not exist, then destroy the proxy
before the loop delivers remove_id. It fails against the previous
handler. pipewire_context_error_scope covers the neighbouring case of
id != PW_ID_CORE, which is why this one is separate: the error that
matters here arrives *on* the core id.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UH29eo7sTWNSkDx8XSfPpt
track_hardware, track_virtual and track_network were each interpreted by
the backend that implemented them, so the same configuration selected
different ports depending on which one was compiled in. ALSA sequencer
counted anything it could not classify as hardware; ALSA raw dropped it;
CoreMIDI had no notion of a virtual port at all; JACK, its UMP twin and
WinMM disagreed again.

observer_configuration::accepts(transport_type) states the rule once, and
each backend now answers only the question it can actually answer: which
transport a given port is on. Ports a platform genuinely cannot classify
belong to no group rather than being forced into one, so a filter no
longer silently gains or loses them.

track_network defaults on. These flags exist to exclude, so a default
observer has to keep listing what it listed before one rule replaced
several: an RTP session on CoreMIDI is device-backed and was admitted as
hardware, and libremidi::client and the C API expose no way to ask for it
back.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UH29eo7sTWNSkDx8XSfPpt
A raw MIDI device is exclusive, and snd_rawmidi_open() on one that is
already held waits in the kernel rather than returning EBUSY. midi_out
asked for SND_RAWMIDI_SYNC alone, so a second open of the same device
hung the caller with no way to report that the port was in use. midi_in
already opened non-blocking, which is why only output showed it. The UMP
backend opens the same device nodes and needed the same treatment.

Opening with SND_RAWMIDI_NONBLOCK turns that case into an error that can
be returned. The non-blocking mode is wanted only for the open -- writes
must still block until the device has taken the bytes, or a burst is
silently truncated -- so it is cleared immediately afterwards.

A short write is now an error rather than a success: with a 4 KiB ring
and a larger SysEx dump, the device would otherwise receive a truncated
message it cannot tell from a complete one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UH29eo7sTWNSkDx8XSfPpt
jcelerier and others added 2 commits September 13, 2026 15:23
Network ports came on by default, so an observer configured for hardware
or for virtual ports listed them too and a host offering all three groups
showed every network port three times.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UH29eo7sTWNSkDx8XSfPpt
Reverts the previous commit, which had it backwards. A default observer
has to keep listing what it listed before one rule replaced several: an
RTP session on CoreMIDI is device-backed and was admitted as hardware,
and libremidi::client and the C API expose no way to ask for it back.

What went wrong was on the other side of the interface -- a host wanting
hardware, software and network in three disjoint groups has to say so on
each observer, rather than relying on a default that is there to avoid
losing ports.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01UH29eo7sTWNSkDx8XSfPpt
@jcelerier
jcelerier merged commit 4d39cfc into master Sep 13, 2026
39 of 90 checks passed
@jcelerier
jcelerier deleted the pipewire-alsa-lifecycle-fixes branch September 13, 2026 21:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant