Skip to content

Add a portable SIMD block index for forward overlap queries - #27

Merged
sstadick merged 29 commits into
masterfrom
worked/portable-simd-index
Aug 20, 2026
Merged

sstadick merged 29 commits into
masterfrom
worked/portable-simd-index

Conversation

@sstadick

@sstadick sstadick commented Jul 28, 2026 •

Copy link
Copy Markdown
Owner

A note on AI 🤖

I used agents extensively to benchmark possible implementations and explore different indexing methods. The two key ideas, the block-based index and the SIMD overlap masks, were, surprisingly, my own.

I'm still conflicted about having agents do any of this, as this library had not previously been touched by AI, and I don't know how its dependents feel about agentically written code. If you have strong feelings and read this, open an issue and let's discuss it.

Summary

This prepares rust-lapper 2.0.0-beta.1 by replacing the original max_len-bounded forward scan with an always-on 32-interval block index while preserving rust-lapper's borrowed, start-ordered iterator identity.

The interval vector stays sorted by (start, stop). Construction groups it into 32-interval blocks and records each block's minimum end, maximum end, prefix maximum (the largest end seen in that block or any earlier block), and next block with a greater maximum. find() binary-searches those running maxima to enter at the first possible block, while seek() narrows the same entry point with its caller-owned cursor.

query [start, stop)
        |
        v
binary-search block prefix maximum
        |
        v
candidate block
        |-- first_start >= stop -> done
        |-- max_end <= start    -> follow next-greater link
        |-- min_end > start     -> return active start prefix
        `-- mixed ends          -> NEON / AVX2 / scalar overlap mask
                                      |
                                      v
                              drain low bits in start order

The dense route only needs the sorted starts. The mixed route tests both half-open overlap conditions, interval.stop > query.start and interval.start < query.stop, and drains the least-significant set bit first so results remain borrowed and start ordered.

Compatibility

  • find() and seek() keep their public call shapes and forward result order.
  • count() remains the two-binary-search BITS implementation; coverage, depth, merging, union, and intersection keep their existing semantics.
  • Signed coordinates are now supported.
  • u8 through u64, their signed forms, usize, and isize use specialized SIMD kernels where supported.
  • u128, i128, and custom PrimInt implementations use the scalar mask.
  • Constructors, insertion, merging, and deserialization rebuild the derived index.
  • Serde retains the original six-field representation.
  • There is no workload heuristic or user-visible mode.

The branch adds an I: 'static bound for safe TypeId-based primitive dispatch. Stable Rust cannot specialize primitive SIMD kernels while retaining a blanket custom-PrimInt fallback, so the bound is accepted and documented as the intentional major-version API change. It excludes lifetime-carrying coordinate types, not short-lived Lapper values.

Performance

The three cases and source data come from the September 2025 polars-bio interval benchmark. That post measures end-to-end Polars operations; the tables here isolate one index build plus one complete query batch after input loading, and compare both rust-lapper versions in the same native harness.

Apple M3/AArch64 NEON medians versus rust-lapper 1.3.0:

Dataset Original total This branch total Improvement
1-2 8.534 ms 5.729 ms 32.87%
7-3 5,075.695 ms 66.385 ms 98.69% (76.46x)
8-7 920.998 ms 588.766 ms 36.07%

AMD Ryzen 9 3950X/x86-64 AVX2 medians versus rust-lapper 1.3.0:

Dataset Original total This branch total Improvement
1-2 13.715 ms 8.600 ms 37.30%
7-3 8,756.427 ms 97.084 ms 98.89% (90.20x)
8-7 1,192.626 ms 779.564 ms 34.64%

Every implementation returned the expected overlap counts: 54,246 for 1-2, 4,408,383 for 7-3, and 307,184,634 for 8-7. CPU, compiler, flags, methods, raw samples, medians, and pinned competitor revisions are retained in the AArch64 record and AVX2 record.

Validation

  • cargo fmt --all -- --check
  • cargo test --all-features --locked
  • cargo clippy --all-targets --all-features --locked -- -D warnings
  • cargo publish --dry-run --locked
  • AArch64 assembly inspection and native comparative benchmarks
  • Native AArch64 NEON execution on Apple and GitHub ARM64 runners
  • Native AVX2 execution and primitive-mask verification on AMD EPYC 7763 and Ryzen 9 3950X hosts
  • Native Ryzen performance comparison against rust-lapper 1.3.0 and pinned competitors
  • x86-64 scalar dispatch and the full test suite under a QEMU Nehalem CPU model with AVX2 hidden
  • Rust 1.59 MSRV and current-stable all-feature tests
  • Default, serde-only, unstable-sort-only, and all-feature configurations
  • i686, PowerPC64LE, and Wasm scalar-fallback compilation

A physical non-AVX2 or Intel-branded host is not an additional release gate. QEMU executes the complete scalar x86-64 suite with AVX2 hidden, while native AMD hosts execute and benchmark the vendor-neutral AVX2 path.

Todo List

The implementation checks are complete and recorded in plans/simd/productionization.md.

  • Publish 2.0.0-beta.1 and verify that it installs from crates.io.
  • Confirm the beta documentation builds on docs.rs.
  • Ask downstream users to compile and test against the beta.
  • Leave the beta open for one or two weeks for correctness and compatibility reports.
  • Release 2.0.0 if no blockers appear, or 2.0.0-beta.2 if code changes are needed.

@sstadick
sstadick marked this pull request as ready for review August 19, 2026 18:23
@sstadick
sstadick merged commit b3159be into master Aug 20, 2026
12 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant