Repository navigation
feat: run Redis as a primary and two replicas behind three Sentinels - #28
Merged
Merged
Conversation
The single Redis node becomes three t3.small nodes, each running Redis and a Sentinel on the org profile: tcnaw-redis01 and tcnaw-redis03 in us-east-1c, tcnaw-redis02 in us-east-1a, placed as the Gitaly and Praefect nodes are. Each Redis node admits 6379 and 26379 from gitlab-node and may call both on its peers; the Rails nodes gain 26379 egress. Only the Rails and Redis nodes have egress to either port and only the Redis nodes admit them, so no Gitaly or Praefect node reaches Redis or a Sentinel. Each Redis node's nftables opens 26379 beside 6379. No IAM or standing resource changes. The role's redis node runs Redis and its Sentinel. The node redis.primary names starts with redis_master_role and the others with redis_replica_role, every node naming primary_host as the primary it starts with; from then on Sentinel decides which node is primary. The Sentinels monitor gitlab-redis with a quorum of two, require their own password, a sixth run secret, and keep the package's down-after and failover timeouts. A rails node no longer names a Redis host: Puma, Sidekiq and Workhorse ask the Sentinels for the primary, and ask again when a connection to it fails other than by timing out, or it answers as a replica. Validation requires the new inputs, and a Redis node must listen on both ports. The secrets origin is the first Redis node by name, tcnaw-redis01, never the first by inventory order, everywhere it is read; the run secrets and the shared gitlab-secrets.json are created and read on it alone. The Redis plays run in order before Gitaly, Praefect and Rails: the first node as primary, the read of its shared secrets, the other two as its replicas with those secrets, then a Redis check play. The check play requires every Sentinel to report three usable Sentinels, all three to name the same Redis node as primary, and every node to replicate as its Sentinel says: two online replicas on the primary, a live link on each replica. It reads roles, never names, so the second converge passes it after a failover. The proof gains a final section. It reads the primary the Sentinels name and stops every GitLab service on that node, Redis and its Sentinel together. The two remaining Sentinels must promote another node within 36 reads 5 seconds apart, a bound that holds one retry after a split vote, since each candidate then waits twice the failover timeout. The promoted node must serve as primary, the session from before the stop must still be root's, a new sign-in must give a session, and a push over HTTP must succeed and appear in the project's push events, which Sidekiq records from a job it takes from Redis. GitLab is started on the stopped node whatever happened, with a named cleanup assert on gitlab-ctl and both ports. The proof then waits for that node to rejoin as an online replica and for every Sentinel to reach its quorum again, so the second converge's Redis check finds what it requires. Every Sentinel read carries the Sentinel password on stdin only, under no_log, judged by the assert after it. A reconfigure after a failover keeps the primary Sentinel recorded but renders redis.conf's replicaof from gitlab.rb again and restarts Redis. On a local Redis 7.2.16 run, the node gitlab.rb names primary came back a primary and was demoted, and a replica Sentinel had promoted came back a replica, leaving no primary for about 40 seconds until the Sentinels failed over. The template comment and the role README say so; nothing the role is given depends on which node is primary, so a failover alone triggers no reconfigure. The workflow describes eleven nodes, estimates the converge near 95 minutes inside its 120 from the 81 nine nodes took, and gives the proof step 45 minutes. The role, inventory and Terraform READMEs describe the three Redis nodes, their Sentinels and the rules.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
The single Redis node becomes three Redis nodes, each running a Sentinel (quorum 2): GitLab's documented minimal
Redis HA. Rails, Sidekiq and Workhorse find the primary through the Sentinels. Losing the primary's whole node fails
over to a replica, and GitLab keeps serving.
tcnaw-redis01andtcnaw-redis03in us-east-1c,tcnaw-redis02in us-east-1a, all t3.small. Eachadmits 6379 and 26379 only from the Rails and Redis nodes. No Gitaly or Praefect node can reach either port.
redis.primarynames starts as primary and the others as its replicas; Sentinel decides afterthat. The Sentinels have their own password, a sixth run secret. Rails asks the Sentinels again when a connection
fails other than by timing out, or the node answers as a replica (redis-client
sentinel_config.rb:109-111, 118-133).tcnaw-redis01, the first Redis node by name, everywhere it is read.two online replicas. It reads roles, never names, so it passes after a failover.
sentinel.c:4769, 4980-4981);recorded;
always, with a named cleanup;Measured locally on real Redis 7.2.16 (three nodes, omnibus 19.4.1's own config templates): eight clean
failovers, with promotion in 10.5–13.6 s and no split votes. The rejoin took about 10 s. The rehearsal also corrected
a review assumption about a reconfigure after a failover; the template comment and README record what it found.
Review
delegate_tobeforewhen, so thedelegated cleanups'
is definedguards error (ignored) rather than skip. There is no run-time impact.Proof
Local:
ruby -c(all-in-one byte-equal to R0);nft -c;Live: this merge is followed by a 240-minute hold dispatch, the first live run of three Redis nodes and the
Sentinel failover.