Repository navigation
Conversation
|
Thanks for this! I maintain Sqinky, which re-encodes every decoded ID to accept only canonical encodings, so Suggestion: use the set index for the default blocklist too. In my measurements the regex matcher for the default list is slower than the set index this PR already builds for custom lists: EMPTY_BLOCKLIST_INDEX = [Set.new.freeze, [].freeze, [].freeze].freeze
+ DEFAULT_BLOCKLIST_INDEX = blocklist_index(DEFAULT_BLOCKLIST)
@@
if default_blocklist
- @blocklist_patterns = DEFAULT_BLOCKLIST_PATTERNS
+ @blocklist_index = DEFAULT_BLOCKLIST_INDEX
For my use case, decoding an ID and re-encoding it to check that it is canonical, the time drops from 62 µs on 0.2.2 to 14 µs with this PR and 5.3 µs with the diff (Ruby 3.3.10). Checks with the diff applied, on both Ruby versions:
With the diff, |
|
Thanks for the suggestion and thorough testing! I went with shared exact lookups for 4–5 character IDs, keeping the regex matcher for longer IDs. Local Ruby 4.0.5 + YJIT results, µs per encoding:
Your approach wins at 12 characters, but slows the extreme padding case by about 2.6×. This version gets even faster short IDs while keeping longer-ID performance steady. All 30 specs and the additional compatibility checks pass. Thanks again - this helped improve the PR! |
|
Quick update: I've re-assessed (or realized past my change) what the actual IDs for performance critical data sets look like and expanded the fast path to cover 6-12 characters, using shared sets for whole-ID/edge matches and letter-only sets for interior matches. This avoids checking edge substrings twice. Longer IDs still use regexes. The update is about 4.4× faster at six characters, 2× at twelve, and 1.7× for padded twelve-character IDs versus the previous PR version. Full comparison - Ruby 4.0.5 + YJIT, µs per encodingLocal macOS ARM64 measurements: warmed medians from three rotated fresh-process rounds, reused instances, GC enabled, 32 seeded inputs per case. Lower is better.
The 255-character case stays flat; the full sweep also shows small measured regressions of about 3–4% at 16/24 characters. All 30 specs pass under the interpreter, YJIT, and ZJIT, and the additional compatibility checks pass. The new indexes add about 14.5 KiB once per process, with no extra per-instance storage. |
|
Great catch, that's an excellent update. I can confirm your results on Linux x86_64. Compatibility. Everything passes! Speed. This is Sqinky's canonical check (decode, re-encode, compare) on Ruby 4.0.5 + YJIT, in µs per check, as the median of 3 fresh-process rounds:
It's also about 1.2x faster than my set-index suggestion at 6–12 characters, so checking the interior separately pays off. Optional follow-up: custom alphabets and blocklists still use Hope this lands, it's such a performance improvement. Thanks for putting this together - great work! |
Reduce encoding, decoding, and blocklist overhead while preserving the public API and generated IDs. Interpreter encode/decode time falls 32.9–93.1% across 14 workloads. Default
Sqids.newtakes 11.6 µs without JIT and 8.4 µs with YJIT.Ruby 4.0.5, macOS ARM64, baseline
9ea13c38c3fa13e32cc9eeec12bac6a00de555ce. Throughput and JIT timings are warmed medians with GC enabled; gem loading and JIT startup are excluded. Memory results distinguish Ruby heap estimates from process RSS. A default instance retains 2,912 bytes versus 200 bytes originally; shared matchers and exact sets add 27,796 bytes once per process, excluding existing blocklist words. Constructor and some JIT gains are smaller than the interpreter encode/decode gains.Encode/decode throughput — 14 workloads
Allocations, heap growth, and retained memory
Large custom blocklists — 1,000 to 100,000 words
100,000-word construction — normal GC and process RSS
Interpreter, YJIT, and ZJIT — original versus optimized
Default blocklist matchers — short IDs and padding
Ruby 4.0.5 + YJIT, reused instance, 100 seeded random single-value inputs from 0–1,000,000. Medians from three rotated fresh-process rounds. µs per encoding; lower is better. This compares the regex matcher alone, scanning sets for all lengths, and the exact-lookup specialization used here.
The exact lookups use shared frozen sets with no extra per-instance storage. Longer IDs retain the regex matcher.
Validation: all 30 existing specs pass under the interpreter, YJIT, and ZJIT. 12,000 seeded differential comparisons cover custom alphabets, malformed IDs, padding, and blocklists. The short-ID specialization passed 25,002 encode/decode comparisons and 13,920 matcher comparisons; the custom-list suite passed 1,780 comparisons, including a 100,001-word list with a 10,000-character word.