Skip to content

feat(config): make startup bounds configurable - #412

Open
EnRaiha wants to merge 3 commits into
NodeDB-Lab:mainfrom
EnRaiha:drill/issue354
Open

EnRaiha wants to merge 3 commits into
NodeDB-Lab:mainfrom
EnRaiha:drill/issue354

Conversation

@EnRaiha

@EnRaiha EnRaiha commented Oct 6, 2026 •

Copy link
Copy Markdown
Contributor

feat: startup readiness bounds, WAL format naming

Closes #354

Why

A store written by another build failed every boot with a message that called its WAL segments
corrupt. The reader knew the format version and threw it away, which is the root cause. This branch
surfaces that version in every reader that decides a record header, so startup names what it found
and what it reads. It also makes the two startup readiness bounds configurable, with a range check on
each one.

What changed

Commit Unit What it does
7034193ee feat(config) Reads raft_ready_timeout_ms and data_group_recovery_timeout_ms from [tuning.startup] as newtypes, range-checks both at load, and refuses a bound written under [server] while naming the path that is read
41bbc0620 fix(wal) Returns the format version from the three readers that decide a record header, judges the version once per segment at its first record, attaches the segment path where the open fails, and keeps startup validation to a single pass
4adf4a87e fix(wal) Starts over a newest segment that is non-empty and holds no parseable record, as its own commit, with the version check still ahead of it

torn_tail is deliberately not one of the three readers. A version mismatch past a corruption stop
is damage there, never a resync point, and its own test pins that.

Behaviour changes, stated plainly

Change From To
tuning.startup.raft_ready_timeout_ms default hard-coded 30 s 300 s, configurable in 1 to 86400000 ms
tuning.startup.data_group_recovery_timeout_ms default hard-coded 60 s 600 s, configurable in the same range
A store in another WAL format "contains no valid WAL records, the segment appears to be corrupted" "holds records in WAL format version 1; this build reads version 3. Start the store with the build that wrote that segment, or re-create the data directory and load the data again"
A non-empty newest segment with no parseable record refused at startup accepted, because the next boot resumes that same file

The wait for the first authorization lease keeps its own constant. It is a hard deadline on one grant
and does not reset on progress, so it is not the metadata stall bound under another name.

Bounds near u64::MAX do not overflow on Linux: Instant::now() + Duration::from_millis(u64::MAX)
computes (tv_sec: 18446744073719120). They are still refused, because a day is the longest window
that still bounds a boot, and a platform whose Instant is a u64 nanosecond counter does overflow.

How it was verified

Arm Result Log
Focused suites exit 0: 155 nodedb, 304 nodedb-wal, 835 nodedb-types 20261009-green-final.log
Full library suite exit 0: 9,137 tests, 0 failures 20261009-fullsuite-final.log
Mutation, version error removed exit 100: 24 of 27 pass, exactly the three version-message tests fail 20261009-mutA-final.log
Mutation, exemption removed exit 100: 23 of 25 pass, exactly the two record-less tests fail 20261009-mutB-final.log
End to end exit 0: a real version-1 production copy is refused with both versions and an action, no "corrupted", and a store this build writes boots, reboots and keeps its rows 20261009-e2e-final.log
Preflight exit 1, one violation, not this branch's 20261009-preflight-final.log

Every arm ran against 4adf4a87e0b97eddb765415948a99cdc800d31d5, whose tree is
e4bff056a296ce0590edb17a9e6abd0e57be71bc.

Both mutation arms run -p nodedb --lib -E 'test(/wal::manager/)', so they falsify the startup
validation tests and the three reader tests through their caller. They do not falsify the
nodedb-wal unit tests added for the readers, which ran only in the focused suites above.

Preflight: one violation, and it is main's

Preflight stops on two clippy::nonminimal_bool errors under clippy 1.96, at
nodedb/src/control/planner/sql_plan_convert/dml/vector_primary.rs:171 and
nodedb/src/data/executor/handlers/control/calvin_reply/stage.rs:95. Both files are byte-identical
to main on this branch:

git diff --quiet origin/main -- \
  nodedb/src/control/planner/sql_plan_convert/dml/vector_primary.rs \
  nodedb/src/data/executor/handlers/control/calvin_reply/stage.rs

The earlier revision of this branch carried the two rewrites, and preflight passed there. The review
asked for those fixes to be split into their own pull request, so this branch no longer carries them.
That leaves the gate red on code this branch does not touch. The two ways forward are a one-shot
PREFLIGHT_SKIP=1 with this paragraph as the reason, or a separate pull request that fixes the two
sites on main.

Review

Review 2 verdict: PASS (0 blockers) on 4adf4a87e0b9. The reviewer re-read the three readers,
the startup validation pass, both bound checks, the [server] path guard, and the exemption commit,
and reproduced the byte-identity of the two clippy sites against main.

Scope

Two units share this pull request, one commit each: the startup bounds under [tuning.startup], and
the WAL format diagnosis with the record-less newest segment as its own commit. The clippy fixes the
review asked to split out are not here.

Copilot AI balanced review requested due to automatic review settings October 6, 2026 13:11

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@EnRaiha
EnRaiha marked this pull request as draft October 6, 2026 14:19
@EnRaiha EnRaiha changed the title feat(config): make startup bounds configurable feat(config): make startup bounds configurable Oct 7, 2026
@EnRaiha
EnRaiha marked this pull request as ready for review October 7, 2026 09:29

@farhan-syah farhan-syah left a comment •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Request changes. The head commit does not compile its own test target. The WAL fix is at the right layer, but it covers one of four places that read record headers.

Rebase first

The branch is 25 commits behind main. It merges cleanly today, but rebase onto current main before the next push. Re-run the checks on the rebased head and post the output from that head.

Claims checked

Claim Result
cargo fmt --all -- --check passes Holds.
Full library suite passes on this head Does not hold. cargo nextest run -p nodedb-wal fails to compile at c620167 with E0382 at nodedb-wal/src/reader.rs:595.
nodedb-types startup tests Hold: 3/3 pass.

Every pass claim in the description needs re-running on the rebased head.

Layer

The startup bounds are at the right layer and match what #354 asks for.

The WAL change is at the right layer: the reader that both recovery and the startup gate share. It needs four changes before it fixes the bug class instead of one instance:

  1. Fix every reader. nodedb-wal/src/lazy_reader.rs:153 and nodedb-wal/src/mmap_reader/reader.rs:299 still turn UnsupportedVersion into a Corruption stop. nodedb-wal/src/torn_tail.rs:156 (intact_record_lsn) still treats a non-current version as "not a record". After this PR, the sentence the description writes about the old reader is still true of three of them.
  2. Decide the version per segment, not per record. header.validate checks the version before the CRC, so the reader declares a format gap from two unverified bytes. The version-0 carve-out patches one case of that. A writer never mixes format versions inside one segment. A mismatch on the segment's first record (after any preamble) is a format gap. A mismatch after an intact current-version record is damage, and torn_tail must classify it.
  3. Attach the segment path where the open fails. The pre-open validate_segments_for_startup call runs a second full scan of every segment. It exists only because the writer's open returns the version error without a path. recovery::recover(path) and SegmentedWal::open already know the path. Attach it there, and keep one validation pass.
  4. Name an action that exists. No WAL migration path exists, so "needs a migration" sends the operator looking for a tool that is not there.

Scope

  • Split the clippy changes out. vector_primary.rs and calvin_reply/stage.rs keep the logic the same but are unrelated to #354. They belong in their own PR.
  • Split the newest-segment relaxation out. Accepting a record-less newest segment changes startup behavior, separate from version diagnosis. It needs its own commit and an accurate justification (see the inline note).

What landed correctly

The [tuning.startup] section follows the existing tuning pattern, and the bounds are wired from main.rs to both gates. The [server] stale-key error does what #354 asks for.

Comment thread nodedb-wal/src/reader.rs
// An empty stream is also what a silent skip or a clean end would
// return, so the reason has to be asserted for this to prove a stop.
assert!(
matches!(reader.stop_reason(), Some(StopReason::Corruption { .. })),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocker: this does not compile. records(self) moves reader at line 585, and this line borrows it after the move (E0382). Keep the reader and iterate with next_record() in a loop, then read stop_reason(). Re-run the full suite on the rebased head after the fix.

Comment thread nodedb-wal/src/reader.rs Outdated
offset: header_offset,
});
}
Err(error @ WalError::UnsupportedVersion { .. }) => return Err(error),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocker: this reader now returns the version error, but lazy_reader.rs:153, mmap_reader/reader.rs:299, and torn_tail.rs:156 still turn the same UnsupportedVersion into corruption or "not a record". Fix all four in this PR.

This arm also fires before the CRC check and at any offset. A mismatch after an intact current-version record in the same segment cannot be a format gap, because the writer never mixes versions in one segment. Return the version error only for the segment's first record (after any preamble). Route a later mismatch through the corruption stop, so torn_tail classifies it. The version-0 arm above then needs no special case.

};
let status_fn = Arc::clone(status_fn);
let deadline = Instant::now() + DATA_GROUP_RECOVERY_TIMEOUT;
let deadline = Instant::now() + timeout;

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Blocker: a configured value overflows Instant and panics the boot. For example, data_group_recovery_timeout_ms = 18446744073709551615 panics here, and admitted_within in auth_lease/status.rs does the same. 0 fails every boot at once. Check both bounds at config load against a stated range, and return a config error that names the key and its range.

/// resets the clock, so a large replay finishes; only a stuck group fails.
/// Default: 300_000 (5 minutes).
#[serde(default = "default_raft_ready_timeout_ms")]
pub raft_ready_timeout_ms: u64,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Nothing checks either bound. 0 and u64::MAX both deserialize, and both break the boot (see the note on data_group_recovery.rs:185). Add a range check at config load that names the key, the value, and the valid range.

pub async fn await_planning_admitted(state: &SharedState, timeout: Duration) -> crate::Result<()> {
pub async fn await_planning_admitted(
state: &SharedState,
timeout: RaftReadyTimeout,

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Planning admission is a different wait from the metadata apply stall. It is a hard deadline, and the stall bound resets on progress. Typing it RaftReadyTimeout couples them, and it raises this wait's default from 30 s to 300 s without saying so. startup.rs says different waits get different types. Give this wait its own type and bound, or keep its own constant, and state the default in the description.

Comment thread nodedb/src/wal/manager/replay.rs Outdated
"the message must name the version it found, and it named none: {text}"
);
assert!(
text.contains("requires 3"),

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

"requires 3" breaks on the next format bump. Build the expected text from nodedb_wal::record::WAL_FORMAT_VERSION.

Comment thread nodedb/src/wal/manager/replay.rs Outdated
// torn tail must not stop a store from opening.
let info = match recovered {
Ok(info) => info,
// A store written by another build used to be reported as

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Comments describe what the code does now, not what an earlier version did. "used to be reported as corruption" inverts once this merges. The same applies to line 510 ("The production store failed every boot"), reader.rs:548 ("The reader used to file..."), and the comment at reader.rs:233. Rewrite each to state the current rule.

Comment thread nodedb-wal/src/segmented.rs Outdated
/// write rather than a format this build cannot read. Opening over it must
/// succeed, because the writer's open is the first thing a boot does.
///
/// A review of an earlier version of the reader change found this missing:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove the review history from the comment. State what the test pins: a newest segment with the record magic and a zeroed version opens.

Comment thread nodedb-wal/src/segmented.rs Outdated
/// rolled would refuse to open. This uses the real writer to produce the
/// state, rather than a file crafted by hand.
#[test]
fn open_accepts_a_segment_the_writer_rolled_and_never_wrote() {

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This passes on main too. The writer creates the rolled segment with 0 bytes, and every path already skips a 0-byte segment. It pins no behavior this PR adds. Remove it, or make it cover the state the relaxation exists for.

Comment thread nodedb/src/config/server/config.rs Outdated
/// names the path that is read instead. `deny_unknown_fields` catches the key
/// anyway, but its message lists every valid field instead of the replacement.
///
/// Neither key was ever a `[server]` field on `main`. They appear at that path

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Remove the branch history ("Neither key was ever a [server] field on main... never landed"). The guard's purpose is enough: it names the path that is read for a misplaced key.

The metadata stall bound and the data-group recovery bound were constants in the bootstrap
code, so an operator with a long WAL tail could not raise them. Both are newtypes now, read
from [tuning.startup], which also stops a call site from transposing two adjacent Duration
arguments.

The shipped defaults change with them: the metadata stall bound moves from 30 s to 300 s and
the data-group recovery bound from 60 s to 600 s. Those are the values production runs, and
the backlog they cover takes minutes to replay, not seconds.

The server config path already rejected such a key through `deny_unknown_fields`, but that
message lists every valid field instead of the path that is read. A guard names the
replacement path, so a misplaced key says where to put it.

Both bounds are range-checked where the config is loaded, and the message names the key, the
value and the range. Zero fails every boot on its first poll, and a bound near u64::MAX stops
being a bound at all.

The wait for the first authorization lease keeps its own constant: it is a
hard deadline on one grant, with no progress reset, so it is not the metadata stall bound
under another name.
A store written by another build failed every boot with a message calling its segments
corrupt. The reader knew the format version and flattened it, along with every other header
validation failure, into StopReason::Corruption, so the operator was sent looking for damage
that was not there.

The version error now travels out of the three readers that decide a record header. It is
judged once per segment, at its first record: one writer never mixes formats inside one
segment, so a mismatch there is another build's format, and a mismatch anywhere later is
damage that torn_tail classifies. A zeroed version is neither, because no build writes format
zero, so it stays a torn-write stop.

torn_tail is deliberately not one of the three. A version mismatch past a corruption stop is
damage there, never a resync point, and its own test pins that.

The callers that know the segment path attach it, so the refusal names the file and both
versions. Startup validation stays a single pass after the open: the open recovers the newest
segment and fails there with the path attached, which is the same refusal without reading
every segment twice.
A crash can leave the newest segment non-empty and holding nothing that parses, such as
zero-filled pages, and the next boot resumes that same file. Refusing it leaves the store
permanently unbootable, because every later boot reads the same bytes and refuses them again.

The newest segment is now the one exemption from the empty-segment check. A record-less
segment anywhere else stays fatal, since nothing resumes it and replay would silently skip
whatever it held. Only segments without a preamble reach that check at all: a segment that
opens with one reports its end at the end of the preamble, never at zero. The version check
runs first, so a segment written by another build is still refused by name rather than
accepted as an empty tail.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

feat(config): make the startup stall and data-group recovery bounds configurable

3 participants