Repository navigation
Add a portable SIMD block index for forward overlap queries - #27
Merged
Merged
Conversation
sstadick
marked this pull request as ready for review
August 19, 2026 18:23
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
A note on AI 🤖
I used agents extensively to benchmark possible implementations and explore different indexing methods. The two key ideas, the block-based index and the SIMD overlap masks, were, surprisingly, my own.
I'm still conflicted about having agents do any of this, as this library had not previously been touched by AI, and I don't know how its dependents feel about agentically written code. If you have strong feelings and read this, open an issue and let's discuss it.
Summary
This prepares rust-lapper 2.0.0-beta.1 by replacing the original
max_len-bounded forward scan with an always-on 32-interval block index while preserving rust-lapper's borrowed, start-ordered iterator identity.The interval vector stays sorted by
(start, stop). Construction groups it into 32-interval blocks and records each block's minimum end, maximum end, prefix maximum (the largest end seen in that block or any earlier block), and next block with a greater maximum.find()binary-searches those running maxima to enter at the first possible block, whileseek()narrows the same entry point with its caller-owned cursor.The dense route only needs the sorted starts. The mixed route tests both half-open overlap conditions,
interval.stop > query.startandinterval.start < query.stop, and drains the least-significant set bit first so results remain borrowed and start ordered.Compatibility
find()andseek()keep their public call shapes and forward result order.count()remains the two-binary-search BITS implementation; coverage, depth, merging, union, and intersection keep their existing semantics.u8throughu64, their signed forms,usize, andisizeuse specialized SIMD kernels where supported.u128,i128, and customPrimIntimplementations use the scalar mask.The branch adds an
I: 'staticbound for safeTypeId-based primitive dispatch. Stable Rust cannot specialize primitive SIMD kernels while retaining a blanket custom-PrimIntfallback, so the bound is accepted and documented as the intentional major-version API change. It excludes lifetime-carrying coordinate types, not short-livedLappervalues.Performance
The three cases and source data come from the September 2025 polars-bio interval benchmark. That post measures end-to-end Polars operations; the tables here isolate one index build plus one complete query batch after input loading, and compare both rust-lapper versions in the same native harness.
Apple M3/AArch64 NEON medians versus rust-lapper 1.3.0:
1-27-38-7AMD Ryzen 9 3950X/x86-64 AVX2 medians versus rust-lapper 1.3.0:
1-27-38-7Every implementation returned the expected overlap counts: 54,246 for
1-2, 4,408,383 for7-3, and 307,184,634 for8-7. CPU, compiler, flags, methods, raw samples, medians, and pinned competitor revisions are retained in the AArch64 record and AVX2 record.Validation
cargo fmt --all -- --checkcargo test --all-features --lockedcargo clippy --all-targets --all-features --locked -- -D warningscargo publish --dry-run --lockedA physical non-AVX2 or Intel-branded host is not an additional release gate. QEMU executes the complete scalar x86-64 suite with AVX2 hidden, while native AMD hosts execute and benchmark the vendor-neutral AVX2 path.
Todo List
The implementation checks are complete and recorded in
plans/simd/productionization.md.2.0.0-beta.1and verify that it installs from crates.io.2.0.0if no blockers appear, or2.0.0-beta.2if code changes are needed.