fix(connectors): bound source forwarding channel with backpressure - #3795
fix(connectors): bound source forwarding channel with backpressure#3795mlevkov wants to merge 4 commits into
Conversation
|
Thanks for the PR. It is labeled Slash commands (own line, regular comment) move it around the queue:
See CONTRIBUTING.md for details. |
Codecov Report❌ Patch coverage is Additional details and impacted files@@ Coverage Diff @@
## master #3795 +/- ##
=============================================
- Coverage 83.14% 21.01% -62.13%
Complexity 1340 1340
=============================================
Files 1219 1218 -1
Lines 166113 138190 -27923
Branches 134282 106488 -27794
=============================================
- Hits 138108 29039 -109069
- Misses 24348 108446 +84098
+ Partials 3657 705 -2952
🚀 New features to boost your workflow:
|
e5ecd55 to
f03fe1f
Compare
aa2a33f to
d7580d5
Compare
|
/request-review @hubcio |
Iggy has no way to receive a webhook. Every provider that pushes events over HTTP needs something in front of it, and today that means running a separate service whose only job is to accept a POST and republish it. This connector removes that hop: it runs an embedded HTTP server, accepts authenticated POST bodies, and produces them to the instance's stream and topic as raw bytes. One plugin .so is loaded once no matter how many source entries reference it, so the listener cannot live on any single instance. It lives in a process-global registry keyed by listen address: the first open binds the public and admin ports, later opens validate their body limit, admin address, management token and instance name against the running listener before joining, and the last close releases both ports. Mismatches fail that instance's open rather than silently handing it a listener its configuration does not describe. A single port can therefore serve many providers, each routed to its own topic. Requests resolve against an ArcSwap route table that is rebuilt whole on every control-plane change, so one atomic load yields both the endpoint's auth rules and the destination bridge. Secret paths carry 128 bits in the URL itself, on the model of a Slack webhook, with optional bearer or HMAC on top; HMAC is verified over the raw body in constant time. Revoked endpoints answer 404 alongside paths that never existed, so a leaked URL cannot be used to confirm it was once live. Endpoints can be registered, re-keyed and revoked at runtime through a token-guarded API on the admin listener, because revoking a compromised endpoint is time-critical and provisioning one per tenant is inherently programmatic. Those endpoints ride the SDK's ConnectorState, and state is attached only to an empty batch: the runtime saves state solely on the success branch of the Iggy send, and an empty send always succeeds, so a mutation cannot be lost to an unrelated send failure. Revocation writes a tombstone that outranks TOML on restore, so a stale config file cannot resurrect an endpoint an operator revoked. Delivery is best-effort in both directions and the README says so first, before anything else: HTTP 200 means accepted into an in-memory buffer, and both the loss and duplicate windows are enumerated with what mitigates each. A full bridge answers 429 with Retry-After rather than blocking, since holding the connection open would turn a slow Iggy into a retry storm. Gateway metrics on the admin listener cover accept-to-200 latency, which the runtime's own stage histograms begin too late to see. Part of the webhook gateway design accepted in apache#3039. The backpressure chain is only complete once the bounded runtime forwarding channel from apache#3795 lands; until then a full bridge signals an arrival burst rather than a slow Iggy, which the README documents. Co-authored-by: Claude <noreply@anthropic.com>
bd19da2 to
399e46c
Compare
Iggy has no way to receive a webhook. Every provider that pushes events over HTTP needs something in front of it, and today that means running a separate service whose only job is to accept a POST and republish it. This connector removes that hop: it runs an embedded HTTP server, accepts authenticated POST bodies, and produces them to the instance's stream and topic as raw bytes. One plugin .so is loaded once no matter how many source entries reference it, so the listener cannot live on any single instance. It lives in a process-global registry keyed by listen address: the first open binds the public and admin ports, later opens validate their body limit, admin address, management token and instance name against the running listener before joining, and the last close releases both ports. Mismatches fail that instance's open rather than silently handing it a listener its configuration does not describe. A single port can therefore serve many providers, each routed to its own topic. Requests resolve against an ArcSwap route table that is rebuilt whole on every control-plane change, so one atomic load yields both the endpoint's auth rules and the destination bridge. Secret paths carry 128 bits in the URL itself, on the model of a Slack webhook, with optional bearer or HMAC on top; HMAC is verified over the raw body in constant time. Revoked endpoints answer 404 alongside paths that never existed, so a leaked URL cannot be used to confirm it was once live. Endpoints can be registered, re-keyed and revoked at runtime through a token-guarded API on the admin listener, because revoking a compromised endpoint is time-critical and provisioning one per tenant is inherently programmatic. Those endpoints ride the SDK's ConnectorState, and state is attached only to an empty batch: the runtime saves state solely on the success branch of the Iggy send, and an empty send always succeeds, so a mutation cannot be lost to an unrelated send failure. Revocation writes a tombstone that outranks TOML on restore, so a stale config file cannot resurrect an endpoint an operator revoked. Delivery is best-effort in both directions and the README says so first, before anything else: HTTP 200 means accepted into an in-memory buffer, and both the loss and duplicate windows are enumerated with what mitigates each. A full bridge answers 429 with Retry-After rather than blocking, since holding the connection open would turn a slow Iggy into a retry storm. Gateway metrics on the admin listener cover accept-to-200 latency, which the runtime's own stage histograms begin too late to see. Part of the webhook gateway design accepted in apache#3039. The backpressure chain is only complete once the bounded runtime forwarding channel from apache#3795 lands; until then a full bridge signals an arrival burst rather than a slow Iggy, which the README documents. Co-authored-by: Claude <noreply@anthropic.com>
hubcio
left a comment
There was a problem hiding this comment.
a few things that don't fit a diff line:
manager/source.rs:164- the 5s task-await is per handle (there are two), fixed, andchannel_capacitynow scales the post-cleanup_senderdrain against it (buffered batches still drain after the sender drops, up to capacity send+save iterations). abort mid-drain is a clean tail truncation since the state save is atomic, so it costs lost position, not corruption - but the interaction deserves a line in the readme: raising capacity lengthens shutdown.source.rs:627-iggy_source_handle's i32 is discarded too, same family as the close-code comment; only reachable in a start-then-stop race today, so hygiene.- elasticsearch_source with
[state] enabled = truepersists its own cursor at close and overrides the runtime state at open, so a runtime-side latch can't cover it. opt-in and off by default; follow-up issue. - the pre-existing
producer.send()err path has the same cursor-supersession shape (skip save, continue, next batch persists) - the latch here doesn't close that one; separate follow-up. - worth a regression test once the latch lands: saturate the channel, stop the connector, restart, assert no row gap -
state_persists_across_connector_restartin the postgres integration suite is a natural template. spawn_source_handleris at 11 params; passing&SourceConfiginstead needs the resolved-path/version split onSourceConnectorPluginfirst, so follow-up sized.
|
Thanks - the shutdown-drop finding is a real bug and I had it characterised wrong in the PR description. Pushed The silent supersession. Traced it and you are right: the drop returns, the forwarding loop keeps draining so a slot frees, the next batch enqueues and ships, and the save that follows carries a cursor covering the batch that was dropped. Calling that tail truncation was wrong; it is mid-stream loss, and it fires on SIGTERM and on API restart, not on some exotic path. Implemented the per-instance latch as you described. First drop sets Mutation-checked: removing the latch check fails the new test with "a batch after the gap must be dropped even with capacity free, or the persisted cursor would advance past the batch that was lost". Grace round before the first drop, as suggested - one bounded Dedicated counter. Capacity default 1024 -> 64, landed together with the fix rather than before it, per your sequencing note. READMEs and the skill doc follow. Tests. Both backoff tests move to capacity 16, since capacity 1 routes to Also taken: close return code checked in Left for follow-ups, as you framed them: the elasticsearch_source cursor override, the The saturate-stop-restart-assert-no-gap regression test you suggested is the one I would most like to add, but it belongs in the postgres integration suite next to Note on the branch: you had updated it with a merge from master, so I rebased my commit on top of that rather than force-pushing, and this was a normal fast-forward push - your merge commit is intact. Gate: fmt, sort, clippy |
|
/ready |
The channel between a source plugin's send callback and the runtime's forwarding loop was flume::unbounded(), so a slow or hung Iggy meant batches accumulated without bound instead of propagating backpressure into the plugin's polling loop. Swap it for a bounded crossfire channel (the shard and server-ng standard), sized by an optional SourceConfig channel_capacity counted in batches, defaulting to 1024. The FFI callback retries with send_timeout while re-reading a shutdown flag, set by the manager before iggy_source_close and for every instance ahead of the sequential process-shutdown stops, since same-library instances share one plugin runtime and a wedged sibling would otherwise hold a worker an earlier close needs. A unit test pins that buffered batches drain after the senders drop, which shutdown relies on and crossfire's docs do not promise. This drops flume from the runtime. Requested in the HTTP source discussion (apache#3039). Co-authored-by: Claude <noreply@anthropic.com>
Review found the shutdown drop is silent mid-stream loss, not the tail truncation this PR claimed. Drop batch N with the channel full, and the forwarding loop keeps draining, so a slot frees, batch N+1 enqueues, ships, and persists its state. Sources snapshot their cursor into every batch, so the saved position now covers N and a restart resumes past the hole. It fires on routine paths: SIGTERM arms every source for the whole sequential-stop window, and connector restart through the API does the same. A per-instance latch closes it. The first drop sets `dropped` and every later batch from that instance is dropped too, so the cursor can never pass the gap. Queued batches still flush, since the ring is FIFO and they all predate the drop. The resume point becomes the last delivered batch's state, which the README now states as the guarantee rather than describing an error count. Dropping also gets one grace `send_timeout` round first. The forwarding loop drains until `cleanup_sender` and the parked sender wakes the moment a slot frees, so the wait is usually the drain, and it is the difference between losing an in-flight batch and delivering it. Loss now has its own counter. Folding it into `iggy_connector_errors_total` put permanent data loss in the same series as decode and send failures that get retried; `iggy_connector_messages_dropped_total` counts messages rather than batches. The error log latches with the flag so a wedged instance emits one line, not one per poll, while the counter keeps moving. Default capacity drops 1024 -> 64. In batches, against postgres' 1000-row default, four figures admitted millions of messages before backpressure engaged. Also from review: the close return code is checked in `stop_connector` the way `init` already checks it, since the new ordering's safety argument depends on close having actually stopped the callbacks; `SourceSenderEntry` moves behind one `Arc` in the map so the callback clones once and `send_with_backpressure` takes three parameters instead of six; the unsafe block narrows to `from_raw_parts`; and the two capacity-1 backoff tests move to 16, since crossfire routes capacity 1 to a different queue and backoff regime than the shipped default. The two shutdown-signal tests merge into one, because `signal_shutdown_all` sets every entry in the process-global map and could have covered for `signal_shutdown` being a no-op.
Same family as the close-code check, and the last of the discarded FFI return codes in this path. The SDK returns non-zero when it could not register the send handler, which leaves an instance reporting Running while producing nothing.
8c4100e to
5de9283
Compare
Iggy has no way to receive a webhook. Every provider that pushes events over HTTP needs something in front of it, and today that means running a separate service whose only job is to accept a POST and republish it. This connector removes that hop: it runs an embedded HTTP server, accepts authenticated POST bodies, and produces them to the instance's stream and topic as raw bytes. One plugin .so is loaded once no matter how many source entries reference it, so the listener cannot live on any single instance. It lives in a process-global registry keyed by listen address: the first open binds the public and admin ports, later opens validate their body limit, admin address, management token and instance name against the running listener before joining, and the last close releases both ports. Mismatches fail that instance's open rather than silently handing it a listener its configuration does not describe. A single port can therefore serve many providers, each routed to its own topic. Requests resolve against an ArcSwap route table that is rebuilt whole on every control-plane change, so one atomic load yields both the endpoint's auth rules and the destination bridge. Secret paths carry 128 bits in the URL itself, on the model of a Slack webhook, with optional bearer or HMAC on top; HMAC is verified over the raw body in constant time. Revoked endpoints answer 404 alongside paths that never existed, so a leaked URL cannot be used to confirm it was once live. Endpoints can be registered, re-keyed and revoked at runtime through a token-guarded API on the admin listener, because revoking a compromised endpoint is time-critical and provisioning one per tenant is inherently programmatic. Those endpoints ride the SDK's ConnectorState, and state is attached only to an empty batch: the runtime saves state solely on the success branch of the Iggy send, and an empty send always succeeds, so a mutation cannot be lost to an unrelated send failure. Revocation writes a tombstone that outranks TOML on restore, so a stale config file cannot resurrect an endpoint an operator revoked. Delivery is best-effort in both directions and the README says so first, before anything else: HTTP 200 means accepted into an in-memory buffer, and both the loss and duplicate windows are enumerated with what mitigates each. A full bridge answers 429 with Retry-After rather than blocking, since holding the connection open would turn a slow Iggy into a retry storm. Gateway metrics on the admin listener cover accept-to-200 latency, which the runtime's own stage histograms begin too late to see. Part of the webhook gateway design accepted in apache#3039. The backpressure chain is only complete once the bounded runtime forwarding channel from apache#3795 lands; until then a full bridge signals an arrival burst rather than a slow Iggy, which the README documents. Co-authored-by: Claude <noreply@anthropic.com>
|
Follow-ups from your review are filed so they survive this PR merging:
One I took here instead of filing: the discarded On the regression test - the saturate, stop, restart, assert-no-row-gap one against Also rebased onto current master per your note, along with my other three. Gate green on all four; this branch is 134 passing. |
| pub verbose: bool, | ||
| #[serde(default)] | ||
| pub benchmark: bool, | ||
| /// Forwarding channel capacity in batches; defaults to 1024. |
There was a problem hiding this comment.
still says 1024 - DEFAULT_CHANNEL_CAPACITY is 64 and the readme says 64. same at line 183.
| use super::*; | ||
| use iggy_connector_sdk::ProducedMessage; | ||
|
|
||
| // Prod default (1024) is crossfire's `ArrayMpsc` with `large = true`. |
There was a problem hiding this comment.
prod default is 64 now, not 1024. the substance holds - 64 is still ArrayMpsc with large = true - but the number is stale.
Summary
The channel between a source plugin's send callback and the runtime's forwarding loop was
flume::unbounded(), so a slow or hung Iggy meant batches accumulated in memory without bound instead of propagating backpressure into the plugin's polling loop. This is the prerequisite runtime fix requested in the HTTP source discussion (#3039), and it applies to every source connector, including the source PRs currently in flight.What changed
crossfire::mpsc::bounded_blocking_async), the same shapeshardandserver-nguse. flume is no longer a runtime dependency.SourceConfigfield,channel_capacity, counted in batches (a single batch can be megabytes), defaulting to 1024 and clamped to [1, 65536] since crossfire eagerly allocates the ring and asserts capacity < 2^31. The existingConfigEnvderive providesIGGY_CONNECTORS_SOURCE_<KEY>_CHANNEL_CAPACITY; configs without the field behave as before apart from the bound.try_sendfast path and asend_timeout(10ms)retry loop that re-reads a per-instance shutdown flag between waits. The manager sets that flag beforeiggy_source_closeso a hung Iggy cannot deadlock the close. Process shutdown sets every instance's flag (signal_shutdown_all) before the sequential stops, because instances loaded from one plugin library share a single tokio runtime and a wedged sibling would otherwise hold a worker an earlier close needs.warn!per backpressure episode (latched, cleared on genuine recovery). A batch that still cannot be enqueued after the stop signal is dropped and counted iniggy_connector_errors_total.One correction to the discussion notes
@hubcio the spec assumed crossfire's blocking sender has no
send_timeout. It does:blocking_tx.rs:288onTx, reachable fromMTxviaDeref. The loop is built on it instead oftry_sendplus sleep, so the sender wakes as soon as capacity frees while shutdown latency stays bounded by the retry interval.Known limitation
Stopping a single connector via the runtime API while enough same-library sibling instances are saturated can delay that close until the siblings drain, because the callback parks a worker of the shared plugin runtime. The code comment and the connector skill document this. The complete fix is an SDK-side worker handoff (
tokio::task::block_in_placearound the callback invocation); happy to file it as a follow-up issue.Test plan
cargo clippy -p iggy-connectors --all-targets -- -D warningscleancargo test -p iggy-connectors: 128 passed, including the new channel and shutdown testscargo build -p iggy_connector_stdout_sink -p iggy_connector_random_sourcecargo test -p integration -- connectors::runtime::could not run on this machine (hwlocality-sysneedspkg-config); relying on CI for the integration suite