Skip to content

Optimize Sqids encoding, decoding, and blocklist matching - #12

Open
rbviz wants to merge 3 commits into
sqids:mainfrom
rbviz:performance-tweaks
Open

rbviz wants to merge 3 commits into
sqids:mainfrom
rbviz:performance-tweaks

Conversation

@rbviz

@rbviz rbviz commented Oct 6, 2026 •

Copy link
Copy Markdown

Reduce encoding, decoding, and blocklist overhead while preserving the public API and generated IDs. Interpreter encode/decode time falls 32.9–93.1% across 14 workloads. Default Sqids.new takes 11.6 µs without JIT and 8.4 µs with YJIT.

  • Use byte buffers and cached alphabet lookups to avoid repeated string conversion, slicing, and searches; small alphabets use sparse maps.
  • Share exact default-blocklist lookups for IDs up to five characters and prefix-factored regexes for longer IDs.
  • Index custom blocklists in frozen sets grouped by length and matching rule; share an empty index and avoid per-character filtering allocations.

Ruby 4.0.5, macOS ARM64, baseline 9ea13c38c3fa13e32cc9eeec12bac6a00de555ce. Throughput and JIT timings are warmed medians with GC enabled; gem loading and JIT startup are excluded. Memory results distinguish Ruby heap estimates from process RSS. A default instance retains 2,912 bytes versus 200 bytes originally; shared matchers and exact sets add 27,796 bytes once per process, excluding existing blocklist words. Constructor and some JIT gains are smaller than the interpreter encode/decode gains.

Encode/decode throughput — 14 workloads
Ruby 4.0.5; 5 alternating rounds; baseline 9ea13c3
Workload                    Before us   After us   Speedup Time saved
single: encode                 36.349      2.517    14.44x      93.1%
single: decode                  7.562      1.145     6.60x      84.9%
three numbers: encode          85.891     16.108     5.33x      81.2%
three numbers: decode          32.422     11.432     2.84x      64.7%
100 numbers: encode          1200.203    478.313     2.51x      60.1%
100 numbers: decode          1189.828    453.125     2.63x      61.9%
padding (255): encode         151.457     56.285     2.69x      62.8%
padding (255): decode         101.334     15.629     6.48x      84.6%
small alphabet: encode          3.821      2.562     1.49x      32.9%
small alphabet: decode          4.606      2.100     2.19x      54.4%
no blocklist: encode           14.906      6.273     2.38x      57.9%
no blocklist: decode           19.406      5.934     3.27x      69.4%
blocklist retries: encode      91.194     37.069     2.46x      59.4%
blocklist retries: decode      21.240      6.395     3.32x      69.9%
Allocations, heap growth, and retained memory
Ruby 4.0.5; baseline 9ea13c3; 3 alternating rounds; 100 batches/round
Objects/op: GC total_allocated_objects delta. Bytes/op: uncollected Ruby heap growth with GC disabled.
Includes returned values and temporary objects. Bytes are ObjectSpace estimates, not cumulative malloc or process RSS.
Workload                   Before obj After obj Obj saved Before B/op  After B/op   B saved
single: encode                  85.8      13.8     84.0%      4606.0      1248.0     72.9%
single: decode                 155.8       6.5     95.8%      8170.0       860.0     89.5%
three numbers: encode          363.0      24.0     93.4%     19152.0      1802.7     90.6%
three numbers: decode          438.0      14.0     96.8%     22640.0      1160.0     94.9%
100 numbers: encode          13847.0     509.0     96.3%    721800.0     25616.0     96.5%
100 numbers: decode          14187.0     403.0     97.2%    731216.0     18024.0     97.5%
padding (255): encode          498.0      27.0     94.6%     27756.0      4612.0     83.4%
padding (255): decode          616.5      17.0     97.2%     31560.0      1560.0     95.1%
small alphabet: encode          47.0      19.0     59.6%      1920.0       988.0     48.5%
small alphabet: decode          72.5      10.5     85.5%      3100.0       440.0     85.8%
no blocklist: encode           226.5      19.0     91.6%     11816.0      1512.0     87.2%
no blocklist: decode           299.5      11.0     96.3%     15520.0      1040.0     93.3%
blocklist retries: encode     1317.0     105.0     92.0%     71344.0      9808.0     86.3%
blocklist retries: decode      302.0      10.0     96.7%     15760.0      1000.0     93.7%

Construction allocations (per new instance, before collection):
single                         192.0      71.0     63.0%      9560.0      7312.0     23.5%
three numbers                  192.0      71.0     63.0%      9560.0      7312.0     23.5%
100 numbers                    192.0      71.0     63.0%      9560.0      7312.0     23.5%
padding (255)                  192.0      71.0     63.0%      9560.0      7312.0     23.5%
small alphabet                  24.0      12.0     50.0%      1024.0       560.0     45.3%
no blocklist                   260.0      71.0     72.7%     13224.0      7312.0     44.7%
blocklist retries              317.0     109.1     65.6%     15970.0     10753.6     32.7%

Retained heap per instance (200 live instances, full GC, shared constants excluded):
Configuration              Before bytes  After bytes  Extra bytes
single                            200.0       2912.0       2712.0
three numbers                     200.0       2912.0       2712.0
100 numbers                       200.0       2912.0       2712.0
padding (255)                     200.0       2912.0       2712.0
small alphabet                    224.0        320.0         96.0
no blocklist                      344.0       2912.0       2568.0
blocklist retries                 409.0       3913.0       3504.0

Default matchers and exact sets reachable heap: 35236 bytes once per process (includes referenced words).
Additional matcher heap excluding objects already in DEFAULT_BLOCKLIST: 27796 bytes once per process.
Large custom blocklists — 1,000 to 100,000 words
Ruby 4.0.5; baseline 9ea13c3; 3 alternating timing rounds
24 inputs per encode batch; 8 canonical IDs inserted to force blocklist hits/retries.
Retained bytes include words reachable from each instance; temporary construction heap is uncollected growth.
Profile           Words Impl      Build ms   Build MiB    Alloc obj   Retain MiB   Encode us   Pad255 us    Checks
random_letters     1000 before        2.23        0.82        17538         0.14      448.25      497.54      PASS
random_letters     1000 after         0.46        0.10         1170         0.15       55.21      276.92      PASS
random_letters    10000 before       20.81        8.10       170902         1.50     2325.12     3539.00      PASS
random_letters    10000 after         4.03        1.10        10136         1.55       41.33      167.04      PASS
random_letters   100000 before      208.26       80.51      1700474        14.87    26193.54    35669.96      PASS
random_letters   100000 after        42.06       10.06       100136        14.69       53.17      200.21      PASS
digit_edges        1000 before        2.37        0.94        19797         0.15      215.04      271.25      PASS
digit_edges        1000 after         0.46        0.10         1135         0.16       27.42       57.75      PASS
digit_edges       10000 before       22.95        9.20       192889         1.58     1742.13     1919.08      PASS
digit_edges       10000 after         4.16        1.19        10135         1.63       26.58       57.50      PASS
digit_edges      100000 before      248.65       91.49      1921685        15.71    21177.33    23053.17      PASS
digit_edges      100000 after        48.34       10.91       100135        15.53       29.33       58.75      PASS
shared_prefix      1000 before        2.20        1.39        26523         0.17      253.88      387.92      PASS
shared_prefix      1000 after         0.45        0.13         1111         0.18       26.62       68.58      PASS
shared_prefix     10000 before       22.89       13.77       260523         1.84     2104.17     3048.79      PASS
shared_prefix     10000 after         4.27        1.40        10111         1.84       27.75       68.42      PASS
shared_prefix    100000 before      250.12      137.30      2600523        18.26    21008.46    30577.58      PASS
shared_prefix    100000 after        45.30       13.64       100111        18.26       30.04       78.12      PASS
Long-word index: 100001 words including a 10,000-character word: PASS
1780 compatibility comparisons passed; all constructors and matcher probes completed.
100,000-word construction — normal GC and process RSS
Ruby 4.0.5; baseline 9ea13c3; 100000 unique words; fresh process per measurement; GC enabled
RSS is a post-construction snapshot of the whole process, including runtime and input blocklist; not peak RSS.
Profile         Impl       Build ms    Alloc obj  GC runs      RSS MiB
random_letters  before       229.76      1700290       39        40.84
random_letters  after         41.95       100139        2        43.28
digit_edges     before       270.86      1900290       55        40.45
digit_edges     after         49.06       100139        3        42.94
shared_prefix   before       281.28      3000290       38        53.30
shared_prefix   after         56.35       100139        1        50.78
Interpreter, YJIT, and ZJIT — original versus optimized
Original versus current implementation on the same Ruby build and JIT mode.
Median microseconds per operation; speedup = original time / current time.
Workload                     Mode          Original us   Current us      Speedup
single: encode               interpreter        37.418        2.620       14.28x
single: encode               yjit               28.650        2.151       13.32x
single: encode               zjit               38.795        2.278       17.03x
single: decode               interpreter         7.341        1.125        6.53x
single: decode               yjit                7.027        0.804        8.74x
single: decode               zjit                7.154        0.915        7.82x
three numbers: encode        interpreter        86.823       16.630        5.22x
three numbers: encode        yjit               59.834        9.502        6.30x
three numbers: encode        zjit               68.224       11.661        5.85x
three numbers: decode        interpreter        31.445       10.372        3.03x
three numbers: decode        yjit               23.217        4.637        5.01x
three numbers: decode        zjit               25.381        6.786        3.74x
100 numbers: encode          interpreter      1201.914      490.766        2.45x
100 numbers: encode          yjit              822.500      219.350        3.75x
100 numbers: encode          zjit              948.938      321.023        2.96x
100 numbers: decode          interpreter      1226.445      448.996        2.73x
100 numbers: decode          yjit              809.047      193.143        4.19x
100 numbers: decode          zjit              906.484      304.098        2.98x
padding (255): encode        interpreter       152.934       56.764        2.69x
padding (255): encode        yjit              116.436       43.781        2.66x
padding (255): encode        zjit              137.667       51.480        2.67x
padding (255): decode        interpreter       101.612       16.086        6.32x
padding (255): decode        yjit               90.641        8.964       10.11x
padding (255): decode        zjit               96.417       11.790        8.18x
small alphabet: encode       interpreter         3.971        2.570        1.55x
small alphabet: encode       yjit                2.712        1.661        1.63x
small alphabet: encode       zjit                3.886        2.221        1.75x
small alphabet: decode       interpreter         4.366        2.093        2.09x
small alphabet: decode       yjit                3.382        1.283        2.64x
small alphabet: decode       zjit                3.862        1.665        2.32x
no blocklist: encode         interpreter        15.111        6.471        2.34x
no blocklist: encode         yjit               12.181        3.026        4.03x
no blocklist: encode         zjit               12.133       10.420        1.16x
no blocklist: decode         interpreter        19.237        5.834        3.30x
no blocklist: decode         yjit               14.778        2.880        5.13x
no blocklist: decode         zjit               16.308        4.038        4.04x
blocklist retries: encode    interpreter        89.235       37.517        2.38x
blocklist retries: encode    yjit               62.250       18.251        3.41x
blocklist retries: encode    zjit               73.475       62.487        1.18x
blocklist retries: decode    interpreter        20.819        6.503        3.20x
blocklist retries: decode    yjit               16.500        3.180        5.19x
blocklist retries: decode    zjit               19.745        4.396        4.49x
Default Sqids.new            interpreter        15.017       11.626        1.29x
Default Sqids.new            yjit               10.407        8.392        1.24x
Default Sqids.new            zjit               11.767        9.723        1.21x
Custom Sqids.new             interpreter        25.188       17.474        1.44x
Custom Sqids.new             yjit               20.210       13.589        1.49x
Custom Sqids.new             zjit               21.485       15.796        1.36x
Default blocklist matchers — short IDs and padding

Ruby 4.0.5 + YJIT, reused instance, 100 seeded random single-value inputs from 0–1,000,000. Medians from three rotated fresh-process rounds. µs per encoding; lower is better. This compares the regex matcher alone, scanning sets for all lengths, and the exact-lookup specialization used here.

Configuration Regex matcher Set scanning Exact lookup up to five characters
Default length 5.66 1.48 0.99
Minimum length 12 8.27 6.36 8.31
Minimum length 255 43.17 110.33 42.84

The exact lookups use shared frozen sets with no extra per-instance storage. Longer IDs retain the regex matcher.

Validation: all 30 existing specs pass under the interpreter, YJIT, and ZJIT. 12,000 seeded differential comparisons cover custom alphabets, malformed IDs, padding, and blocklists. The short-ID specialization passed 25,002 encode/decode comparisons and 13,920 matcher comparisons; the custom-list suite passed 1,780 comparisons, including a 100,001-word list with a 10,000-character word.

@david-uhlig

Copy link
Copy Markdown

Thanks for this! I maintain Sqinky, which re-encodes every decoded ID to accept only canonical encodings, so encode speed matters a lot.

Suggestion: use the set index for the default blocklist too. In my measurements the regex matcher for the default list is slower than the set index this PR already builds for custom lists:

   EMPTY_BLOCKLIST_INDEX = [Set.new.freeze, [].freeze, [].freeze].freeze
+  DEFAULT_BLOCKLIST_INDEX = blocklist_index(DEFAULT_BLOCKLIST)
@@
     if default_blocklist
-      @blocklist_patterns = DEFAULT_BLOCKLIST_PATTERNS
+      @blocklist_index = DEFAULT_BLOCKLIST_INDEX

Sqids.new.encode([id]) with random IDs up to 1,000,000, measured with benchmark-ips on Linux x86_64:

Ruby 3.3.10 Ruby 3.3.10 + YJIT Ruby 4.0.5 Ruby 4.0.5 + YJIT
0.2.2 49.7 µs 38.1 µs 53.2 µs 37.5 µs
This PR 13.4 µs 10.8 µs 13.4 µs 13.0 µs
This PR + diff above 3.4 µs 2.4 µs 3.2 µs 2.4 µs

For my use case, decoding an ID and re-encoding it to check that it is canonical, the time drops from 62 µs on 0.2.2 to 14 µs with this PR and 5.3 µs with the diff (Ruby 3.3.10).

Checks with the diff applied, on both Ruby versions:

  • The 30 existing specs pass, with and without YJIT on Ruby 4.0.5.
  • The encodings are byte-for-byte identical to 0.2.2 for 250,001 inputs, each encoded with the default options and with min_length: 12 (500,002 encodings in total). The inputs are IDs 0–200,000 and 50,000 random lists of 1–4 values up to Sqids.max_value. They include the 15 IDs below 200,000 that the default blocklist forces to be regenerated.

With the diff, blocklist_patterns, pattern_source, DEFAULT_BLOCKLIST_PATTERNS, and the regex branch of blocked_id? are no longer used and could be removed, which would also simplify the PR.

@rbviz

rbviz commented Oct 6, 2026

Copy link
Copy Markdown
Author

@david-uhlig

Thanks for the suggestion and thorough testing! I went with shared exact lookups for 4–5 character IDs, keeping the regex matcher for longer IDs.

Local Ruby 4.0.5 + YJIT results, µs per encoding:

ID length Previous PR Your suggestion Updated PR
Typical short IDs 5.66 1.48 0.99
Padded to 12 8.27 6.36 8.31
Padded to 255 43.17 110.33 42.84

Your approach wins at 12 characters, but slows the extreme padding case by about 2.6×. This version gets even faster short IDs while keeping longer-ID performance steady.

All 30 specs and the additional compatibility checks pass. Thanks again - this helped improve the PR!

@rbviz

rbviz commented Oct 6, 2026

Copy link
Copy Markdown
Author

Quick update: I've re-assessed (or realized past my change) what the actual IDs for performance critical data sets look like and expanded the fast path to cover 6-12 characters, using shared sets for whole-ID/edge matches and letter-only sets for interior matches. This avoids checking edge substrings twice. Longer IDs still use regexes.

The update is about 4.4× faster at six characters, 2× at twelve, and 1.7× for padded twelve-character IDs versus the previous PR version.

Full comparison - Ruby 4.0.5 + YJIT, µs per encoding

Local macOS ARM64 measurements: warmed medians from three rotated fresh-process rounds, reused instances, GC enabled, 32 seeded inputs per case. Lower is better.

Workload Previous PR Your set-index suggestion Updated PR
4 characters 0.86 1.27 0.88
5 characters 1.02 1.58 1.02
6 characters 5.72 1.93 1.30
7 characters 5.45 2.36 1.50
8 characters 6.08 2.68 1.81
10 characters 6.74 3.64 2.51
12 characters 6.57 4.73 3.35
Padded to 12 8.64 6.55 4.95
Padded to 16 8.53 8.43 8.86
Padded to 24 9.12 11.56 9.42
Padded to 32 8.17 14.41 8.23
Padded to 64 11.98 27.35 12.21
Padded to 128 21.64 53.72 22.11
Padded to 255 41.41 107.94 41.40
Tenant + product (10–12) 8.50 6.14 4.91
Three large values (18–20) 11.02 11.57 11.13

The 255-character case stays flat; the full sweep also shows small measured regressions of about 3–4% at 16/24 characters. All 30 specs pass under the interpreter, YJIT, and ZJIT, and the additional compatibility checks pass. The new indexes add about 14.5 KiB once per process, with no extra per-instance storage.

@david-uhlig

Copy link
Copy Markdown

Great catch, that's an excellent update. I can confirm your results on Linux x86_64.

Compatibility. Everything passes!

Speed. This is Sqinky's canonical check (decode, re-encode, compare) on Ruby 4.0.5 + YJIT, in µs per check, as the median of 3 fresh-process rounds:

Encodings Length 0.2.2 Previous PR This PR
1 id ≤ 13.8M 5 48.1 2.7 2.9
1 id ≤ 844M 6 57.7 14.5 3.3
1 id ~ 1e12 8 78.9 15.0 4.2
1 id ~ 2^62 12 90.4 16.4 7.1
2 values 7–12 84.9–109.0 19.1–21.6 8.1–11.1

It's also about 1.2x faster than my set-index suggestion at 6–12 characters, so checking the interior separately pays off.

Optional follow-up: custom alphabets and blocklists still use blocklist_index at every length. With a shuffled alphabet, the check gets much faster at normal lengths (5.7 vs. 71.6 µs for 5 characters on Ruby 3.3.10). But the gain shrinks as min_length grows: 46.3 vs. 166.1 µs at 32 characters, and 310.1 vs. 320.1 µs at 255. It's never slower than 0.2.2, so it's fine to leave as is. Since the default blocklist now switches to regexes above 12 characters, a similar cutoff could help custom lists with large min_length, if the regex construction cost is acceptable there.

Hope this lands, it's such a performance improvement. Thanks for putting this together - great work!

@rbviz rbviz changed the title Performance and memory optimizations Optimize Sqids encoding, decoding, and blocklist matching Oct 7, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants