Skip to content

Recover registration after outages and stop on hard failures - #19

Merged
calebtt merged 3 commits into
masterfrom
fix/issue-13-registration-recovery
Sep 26, 2026
Merged

calebtt merged 3 commits into
masterfrom
fix/issue-13-registration-recovery

Conversation

@calebtt

@calebtt calebtt commented Sep 26, 2026

Copy link
Copy Markdown
Owner

Summary

Closes #13.

A failed registration refresh scheduled a reconnect that called SIPRegistrationUserAgent.Start() on an agent that was still running. That threw "already running" and ended the client's retry chain after one attempt. OnRegistrationTemporaryFailure also never cleared IsRegistered, so status said registered for the whole outage. Recovery came only from SIPSorcery's own retry, 300 s after the failure. The issue has the reproduction timeline.

Changes

SipClient

  • Flag and state on failure. Any registration failure clears IsRegistered and raises RegistrationStatusChanged, with or without auto-reconnect.
  • Restarts. Every start stops the current agent (Stop(sendZeroExpiryRegister: false)), detaches it, and starts a fresh one. "Already running" can't happen, and late events from an old agent are ignored. StartRegistration() is safe to call again at any time, including after a hard failure.
  • Default retry is unchanged. After a temporary failure the client makes up to 5 attempts at 2 s × n. After the last one the agent is left running, so SIPSorcery's own retry still recovers the process. SipBotOpen and homeline keep today's recovery.
  • Optional extended retry. SipConfig.RegistrationRetry.Extended, for always-on hosts. The client owns the retries: the delay doubles from 30 s to at most 5 min, and retries never stop.
  • Hard failures stop. A 401/407 after authentication, or a 402, 403, or 404, classified by SIP status rather than by which SIPSorcery event fired, stops the agent. No further REGISTER goes out until StartRegistration() is called. Without this stop, a rejected password on the first REGISTER keeps being re-sent by Start()'s repeating timer every (Expires − 5) s.
  • Health check. It does nothing before registration has been started, after a hard failure, or while an attempt is in flight. With auto-reconnect off, it restarts registration through the same stop-then-start path.
  • New properties. RegistrationState (NotStarted, Registering, Registered, TemporaryFailure, HardFailure) and LastRegistrationError.
  • Loopback. A loopback registrar given as host:port is treated as loopback, so no STUN runs.

sipbot serve

  • --extended-retry / SIP_EXTENDED_RETRY=1 turns extended retry on.
  • --sip-trace / SIP_TRACE=1 logs every SIP message sent and received to stderr. It is for lab use only, because the messages include headers such as digest Authorization.
  • Both are off by default and read by sipbot only, not by the library's settings loader. The startup line reports extendedRetry= and sipTrace=.
  • JSONL event and field names are unchanged. status.registered is now false while registration is failing. AGENTS.md documents all of this.

Tests

A scripted loopback registrar (FakeRegistrar) answers REGISTER with 200, 503, 402/403/404, a 401 challenge, or nothing. It counts transactions, not retransmissions. 26 new tests:

Area Tests
Hard failures 403, 404, and 402 on the first REGISTER, and 403, 404, 402, and 401 on the authenticated REGISTER, each in default and extended mode (14 cases). Registration stops, LastRegistrationError is set, and no REGISTER follows
Repeating timer After a 403 on the first REGISTER, no REGISTER for a full 25 s refresh period
Re-register StartRegistration() after a hard failure sends one REGISTER while the 403 persists, then recovers once the account is accepted (both modes)
Default schedule Five attempts at 100 ms × n in the tests, then the agent's own retry
Temporary failure A timeout after success clears the flag and reports the change, and recovery works with and without auto-reconnect
Extended retry Backoff doubles to the cap and never stops, and restarts after a success. The health check adds no REGISTERs and never interrupts an attempt in flight (both modes)
Dispose Disposing with a reconnect pending sends nothing more

dotnet test: 86 passed (60 existing + 26 new). Four full runs were all green. Build warnings unchanged (17).

Live check against the lab PBX (VitalPBX / Asterisk 20.14)

sipbot serve --extended-retry registered as the subject extension. Then, on the PBX, SIP UDP to and from that client's address and port only was dropped for 4 minutes.

Time (CDT) Event
20:07:26 Registered (Expires 120); block applied
20:09:52 Refresh REGISTER timed out: temporary failure. status reports registered: false from here on; attempt 1 is scheduled in 30 s
20:10:22 Attempt 1, dropped. It reports its own failure at 20:10:54; attempt 2 is scheduled in 60 s
20:11:35 Block lifted (22 inbound and 5 outbound packets dropped)
20:11:54 Attempt 2 registers (20:11:54.9), 19.4 s after the lift: the 19 s still left on the wait, plus one round trip

No "already running" appeared. status said registered: false in all 12 polls during the outage. For comparison, master stayed true for 43 polls and recovered only through SIPSorcery's 300 s retry.

An earlier run of the first commit found the health check starting the next attempt 3 s into the current one. Commit 7a784ab fixes that and adds a test. In the run above, each attempt reports its own failure.

Compatibility

  • Library consumers. SipBotOpen and homeline use the defaults and call StartRegistration() once. They keep the same retry schedule and recovery, with two differences. IsRegistered is now false during an outage. And a rejected account (for example a changed password) now stops registering instead of re-sending the rejected credentials, so restart after fixing it.
  • Health check. It no longer starts registration on its own if the host never called StartRegistration().
  • API. Additive: RegistrationState, LastRegistrationError, and SipConfig.RegistrationRetry. The previous constants became option defaults.

🤖 Generated with Claude Code

calebtt and others added 3 commits September 25, 2026 19:58
A failed registration refresh scheduled a reconnect that called
SIPRegistrationUserAgent.Start() on an agent that was still running, which
threw "already running" and ended the client's retry chain after one
attempt. OnRegistrationTemporaryFailure also never cleared IsRegistered, so
the client reported registered for the whole outage. Recovery came only from
SIPSorcery's own retry, 300 s after the failure.

SipClient now:
- clears IsRegistered and raises RegistrationStatusChanged on any failure,
  with or without auto-reconnect;
- stops the current registration agent (no unregister) and starts a fresh
  one on every start, so "already running" cannot happen and late events
  from an old agent are ignored; StartRegistration() is safe to call again;
- keeps the default retry schedule (5 attempts at 2 s x n, then the agent's
  own retry stays armed), and adds optional extended retry
  (RegistrationRetryOptions.Extended: 30 s doubling to 5 min, never stops);
- classifies hard failures by SIP status (401/407 after authentication,
  402, 403, 404) and stops the agent, so rejected credentials are not sent
  again by SIPSorcery's repeating timer; the health check does not restart
  registration after a hard failure or while an attempt is in flight;
- exposes RegistrationState and LastRegistrationError;
- treats a loopback registrar given as host:port as loopback (no STUN).

sipbot serve gains --extended-retry / SIP_EXTENDED_RETRY and --sip-trace /
SIP_TRACE (both off, read by sipbot only) and reports both on its startup
line. JSONL event and field names are unchanged; status.registered is now
false while registration is failing.

Tests: a scripted loopback registrar covers hard failures (first and
authenticated REGISTER, both retry modes), re-registering after a hard
failure, SIPSorcery's repeating-timer case, the default schedule, extended
backoff and its reset, recovery with and without auto-reconnect, the health
check, and Dispose with a reconnect pending.

Closes #13

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Found in a live outage test against the lab PBX with --extended-retry: with
REGISTERs being dropped, the health check fired 3 s into a reconnect
attempt and scheduled the next one. In extended mode that stopped the
in-flight agent, so its own failure was never reported and the backoff
skipped a step; in default mode the scheduled attempt cut the in-flight one
short.

The health check now does nothing while an attempt is in flight
(RegistrationState.Registering), in both modes. New test: every REGISTER
dropped, 2 s per attempt, health check every 100 ms; without the fix
attempts were cut short after 176-600 ms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ck test

On the CI runner a REGISTER can leave late (SIPSorcery sends it via a
blocking call on a thread-pool thread) while its attempt timeout is already
running, so gaps between REGISTERs came out short even with the fix. Assert
the bug's own signature instead: every reconnect attempt must follow a
reported temporary failure, taken from the client's metrics. Without the
fix both cases still fail.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@calebtt

calebtt commented Sep 26, 2026

Copy link
Copy Markdown
Owner Author

The first CI run failed only in Health_check_does_not_interrupt_an_attempt_in_flight: one short gap between REGISTERs in each mode. On the runner, SIPSorcery can send a REGISTER late (it uses a blocking call on a thread-pool thread) while that attempt's timeout is already running, so gap length isn't a reliable signal there.

Commit 02d9d4d now checks the bug's own signature instead: every reconnect attempt must follow a reported temporary failure, counted from SipClient.Metrics. It passed 3 of 3 local runs. Both cases still fail on the commit without the fix. CI is green, and #20 is rebased onto it.

🤖 Generated with Claude Code

@calebtt
calebtt added this pull request to stack #21 September 26, 2026 01:23
@calebtt
calebtt merged commit a598942 into master Sep 26, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Registration does not recover after an outage: reconnect throws "already running" and IsRegistered stays true

1 participant