Repository navigation
Operations Performance Tuning
Spector delivers sub-millisecond latency out of the box — but there's always room to optimize for your specific workload. This page covers benchmarks, tuning strategies, and the science of finding the right recall/latency/memory trade-off.
All benchmarks measured on a 24-core x86 machine (Windows 11, Intel Core Ultra 9 285K), AVX2 256-bit, Java 25, ZGC, using clustered vectors (realistic distribution). Numbers represent actual measured results — run
mvn -pl spector-bench exec:javato reproduce on your hardware.
Note
Methodology: Benchmarks use 200 measurement iterations with 50 warmup iterations per scenario. Vectors are generated with realistic cluster structure (50 clusters with Gaussian noise). Documents contain 200–1500 words with paragraph structure. Recall is measured against brute-force ground truth. Your results may vary ±20% depending on CPU model, OS scheduling, background load, and thermal throttling.
| Dimension | Cosine P50 | Cosine P99 | Dot Product P50 | Dot Product P99 |
|---|---|---|---|---|
| 32 | 500 ns | 1,500 ns | 200 ns | 400 ns |
| 128 | <100 ns | 100 ns | 100 ns | 1,300 ns |
| 384 | ~100 ns | 100 ns | ~100 ns | 100 ns |
| 768 | ~100 ns | 100 ns | ~100 ns | 100 ns |
Note
Values at 384+ are at System.nanoTime() resolution floor. JMH confirms millions of ops/sec.
| Scale | Keyword (BM25) | Vector (HNSW) | Hybrid (RRF) |
|---|---|---|---|
| 10K docs | 0.19 ms / 3.79 ms p99 | 0.05 ms / 0.10 ms p99 | 0.17 ms / 0.37 ms p99 |
| 50K docs | 0.42 ms / 0.68 ms p99 | 0.09 ms / 0.19 ms p99 | 0.50 ms / 0.81 ms p99 |
| 100K docs | 0.98 ms / 1.39 ms p99 | Ultra-low / Sub-ms p99 | 1.01 ms / 1.22 ms p99 |
| Scale | Keyword | Vector | Hybrid |
|---|---|---|---|
| 10K | 5,194 | 18,824 | 5,828 |
| 50K | 2,406 | 10,980 | 1,988 |
| 100K | 1,019 | 7,556 | 994 |
| Dataset Size | Time | Rate | Memory |
|---|---|---|---|
| 10K | 2.5s | 3,931 docs/s | +19 MB |
| 50K | 15.1s | 3,308 docs/s | +93 MB |
| 100K | 38.2s | 2,618 docs/s | +187 MB |
| Threads | Throughput | Avg Latency | Scaling Factor |
|---|---|---|---|
| 1 | 3,739 ops/s | 0.26 ms | 1.0× |
| 4 | 10,317 ops/s | 0.37 ms | 2.8× |
| 8 | 11,812 ops/s | 0.58 ms | 3.2× |
| 16 | 14,022 ops/s | 1.00 ms | 3.7× |
Note
Concurrency scaling is measured with 384-dim vectors (production-realistic). 128-dim shows higher absolute throughput but the scaling factor is similar. Individual HNSW queries are sequential — scaling comes from serving multiple queries concurrently.
mvn -pl spector-bench exec:javaTip
Generates an HTML report at spector-bench/target/performance-report.html
# SIMD kernels only
mvn -pl spector-bench exec:java -Dexec.args="SimdKernelBenchmark"
# HNSW index operations
mvn -pl spector-bench exec:java -Dexec.args="HnswBenchmark"
# Concurrency scaling
mvn -pl spector-bench exec:java -Dexec.args="ConcurrencyBenchmark"mvn -pl spector-bench exec:java -Dexec.args="-rf json -rff results.json"# Generate baseline
mvn -pl spector-bench exec:java -Dexec.args="--baseline"
# Compare against baseline
mvn -pl spector-bench exec:java -Dexec.args="--compare"Goal: recall@10 ≥ 95%
var config = SpectorConfig.DEFAULT
.withM(32) // More connections
.withEfConstruction(400) // Better graph quality
.withEfSearch(200); // Wider search beamTrade-offs: 2× memory, ~3× build time, ~2× query latency.
Goal: p99 < 0.5ms
var config = SpectorConfig.DEFAULT
.withM(12)
.withEfConstruction(100)
.withEfSearch(30);Trade-offs: Lower recall (~80% recall@10), but sub-millisecond guaranteed.
Goal: Maximum queries/sec under concurrent load
var config = SpectorConfig.DEFAULT
.withM(16) // Balanced
.withEfSearch(50) // Not too high
.withGpu(true); // Batch processingKey factors:
-
Virtual threads handle concurrency automatically
-
Keep
efSearchmoderate to reduce per-query work -
Enable GPU for batch workloads
-
Use IVF-PQ for large datasets (reduced memory = better cache behavior)
Goal: Fit large datasets in limited RAM
var config = SpectorConfig.DEFAULT
.withM(8) // Fewer connections
.withEfConstruction(100);
// Use IVF-PQ for 32× vector compressionMemory per document (384-dim):
| Mode | Per Vector | 1M vectors |
|---|---|---|
| Float32 | ~1.8 KB | ~1.8 GB |
| INT8 | ~640 bytes | ~640 MB |
| IVF-PQ | ~288 bytes | ~288 MB |
Note
Recall values below are measured with uniform random vectors (best case). Real embedding distributions with cluster structure may show lower recall at the same efSearch — increase efSearch to 100–200 for production workloads with real embeddings.
| efSearch | Recall@10 (random) | Recall@10 (clustered) | Avg Latency | Notes |
|---|---|---|---|---|
| 10 | ~70% | ~30-40% | 0.02 ms | Too low for most uses |
| 30 | ~85% | ~50-60% | 0.03 ms | Fast, moderate recall |
| 64 | ~90% | ~50-65% | 0.05 ms | Default |
| 100 | ~95% | ~70-80% | 0.10 ms | Good for production |
| 200 | ~98% | ~85-90% | 0.20 ms | High recall |
| 500 | ~99.5% | ~95%+ | 0.50 ms | Near-perfect |
| nprobe | Recall@10 | Relative Latency |
|---|---|---|
| 1 | ~40% | 1× |
| 4 | ~70% | 4× |
| 8 | ~85% | 8× |
| 16 | ~92% | 16× |
| 32 | ~97% | 32× |
SpectorIndex uses IVF partitioning with adaptive HNSW shards. The two key parameters are:
-
nCentroids— number of K-Means partitions (set at training time) -
nProbe— number of partitions searched at query time (adjustable)
Rule of thumb: nCentroids ≈ √N (square root of dataset size).
Real embedding results (Qwen3-embedding, 4096-dim, 10K vectors):
| nCentroids | nProbe | % Data Searched | Avg Latency | QPS | Recall@10 |
|---|---|---|---|---|---|
| 128 | 4 | 3.1% | 0.46ms | 2,173 | 1.0000 |
| 128 | 8 | 6.3% | 0.73ms | 1,368 | 1.0000 |
| 128 | 16 | 12.5% | 1.26ms | 792 | 1.0000 |
| 64 | 4 | 6.3% | 0.62ms | 1,601 | 1.0000 |
| 64 | 8 | 12.5% | 1.17ms | 856 | 1.0000 |
| 32 | 4 | 12.5% | 1.17ms | 857 | 1.0000 |
Tip
With real embeddings (not random vectors), SpectorIndex achieves perfect recall at nProbe=4 because real embeddings form natural semantic clusters that K-Means captures effectively. Start with nProbe=4 and only increase if your recall target isn't met.
Note
For the complete, empirical sweeps across multiple partition configurations (
Ingestion throughput (SpectorIndex vs standalone HNSW):
| Dataset Size | SpectorIndex | Standalone HNSW | Speedup |
|---|---|---|---|
| 10K | 130K docs/s | 4,677 docs/s | 28× |
| 50K | 140K docs/s | 2,483 docs/s | 56× |
| 100K | 150K docs/s | 1,535 docs/s | 98× |
| 500K | 246K docs/s | — | — |
| 1M | 128K docs/s | — | — |
-
Add CPU cores → Concurrent throughput scaling (up to ~3.7× at 16 threads measured)
-
Add RAM → Support larger capacity without IVF-PQ compression
-
Add GPU → 4× brute-force search speedup at 100K+ vectors (data resident in VRAM)
-
Add nodes → Linear throughput scaling per shard
-
Rule of thumb: 100K–500K docs per shard
-
See Distributed Mode for cluster setup
Recommended JVM arguments for production:
java \
--add-modules jdk.incubator.vector \
--enable-native-access=ALL-UNNAMED \
-XX:+UseZGC \
-XX:+ZGenerational \
-Xmx4g \
-Xms4g \
-jar spector-node.jar| Argument | Purpose |
|---|---|
--add-modules jdk.incubator.vector |
Required for SIMD acceleration |
--enable-native-access=ALL-UNNAMED |
Required for Panama FFM (GPU, mmap) |
-XX:+UseZGC |
Low-pause GC (vectors are off-heap) |
-XX:+ZGenerational |
Generational ZGC for better throughput |
-Xmx4g -Xms4g |
Fixed heap avoids resize pauses |
Tip
Since all vectors live off-heap, GC pressure is minimal. The heap primarily holds the HNSW graph structure and BM25 inverted index.
When running Spector Synapse with spector.auth.enabled=true, each authenticated user gets an isolated SpectorMemory instance. These are cached in a per-user registry (LRU, configurable cap).
Each DefaultSpectorMemory instance allocates:
| Resource | Per Instance | Notes |
|---|---|---|
| mmap segments | ~13 file descriptors | Episodic, Semantic, Procedural tiers + graphs + WAL + text.dat |
| QuartzMemoryScheduler | Shared Quartz thread pool | Checkpoints, sleep consolidation, DMN, graph enrichment |
| MemoryIndex metadata | ~740 bytes/memory | On-heap structural metadata (text bodies are off-heap via mmap) |
| HNSW graph | O(capacity × M × 4B) | Off-heap via Panama Arena |
Configure these before starting Spector in production:
# File descriptor limit (default 1024 is too low for multi-user)
# Formula: 13 FDs/instance × max-instances + 200 (JVM overhead)
ulimit -n 65536
# Memory-mapped region limit (Linux only)
# Formula: 13 regions/instance × max-instances × 2 (safety margin)
sudo sysctl -w vm.max_map_count=262144
# Make persistent across reboots
echo "vm.max_map_count=262144" | sudo tee -a /etc/sysctl.confThe heap primarily holds per-user MemoryIndex metadata (~740 bytes per memory per user) plus HNSW graph structures and BM25 inverted indexes.
| Deployment | Users | Memories/User | Index Heap | Recommended -Xmx
|
|---|---|---|---|---|
| Developer | 1 | 1K–5K | < 10 MB | 4g |
| Small team | 5–20 | 1K–5K | 10–75 MB | 4g |
| Medium | 50–100 | 1K–5K | 75–375 MB | 6g |
| Large | 200–512 | 5K–10K | 1–3.7 GB | 8g–12g |
# Example: 100-user deployment
java \
--add-modules jdk.incubator.vector \
--enable-native-access=ALL-UNNAMED \
-XX:+UseZGC \
-XX:+ZGenerational \
-Xmx8g -Xms8g \
-jar spector-synapse.jarThe spector.auth.memory.max-instances property (default: 512) controls how many per-user memory instances are cached concurrently. Evicted users' instances are closed and their working memory is lost (episodic/semantic tiers persist on disk and are reloaded on next access).
# application.yml — tune based on available resources
spector:
auth:
enabled: true
memory:
max-instances: 256 # Halve default for constrained environmentsWarning
Setting max-instances too high without corresponding heap and FD increases will cause OutOfMemoryError or Too many open files errors under load. Use the formulas above to calculate safe limits for your deployment.
-
Configuration Guide — All parameters with ranges
-
Core Concepts — How algorithms affect performance
-
SpectorIndex Architecture — IVF-HNSW-SVASQ design and tuning
-
Large-Scale Benchmarks — Empirical sweeps for real embeddings and shard promotions
-
SVASQ Quantization — How SVASQ compression works
-
GPU Acceleration — GPU-specific performance
-
Distributed Mode — Scaling across nodes
- Home
-
Getting Started
- Quick Start
- Installation
- Developer Guide
- JDK API Status
- MCP Server
- Java SDK
- Java API Reference
- Python SDK
- TypeScript SDK
- Spring AI Integration
- CLI Reference
- REST API
- API Playground
- Error Codes
- Configuration
- Deployment
-
Cognitive Memory
- Overview
- Getting Started
- Use Cases
- API Reference
- Concepts
- Pathways
- Scoring features
- Profiles
- Experimental
- Internals
- Design ancestry
-
Memory Kernel
- Overview
- Bundle Architecture
- Memory Shapes
- Binary Layouts & Tags
- WAL & Durability
-
Region Reference
- Overview & Index
- Partition Regions
-
Runtime Regions
- Working Memory
- Co-Activation Matrix
- Index MIDX
- Index IDPL
- Hebbian Graph
- Temporal Chains
- Temporal Facts
- Entity Directory
- Entity Names Pool
- HyperEntity Graph
- Entity Types Registry
- Relation Types Registry
- BM25 Lexical Index
- Checkpoint
- Insula (Somatic Self-Model)
- Continuity
- Provenance
- SPLADE Sparse Index
- Entity Reverse Index
- Identity Regions
- Synapse & Cortex
-
Architecture
- System Overview
- Core Concepts
- Ingestion Pipeline
- MCP Integration
- Distributed Mode
- Event Notifications
- Namespace Sharding
- Single-Namespace Scale & Capacity Limits
- Scale Benchmark Empirical Results
- Writer Quiesce Pause Empirical Results
- Kill-Owner Failover Empirical Results
- Salience & Importance Architecture
- GPU Acceleration
- Performance Tuning
- Test Framework & LLM Judge
- Chat & Visual Test Infrastructure
- Security & Data
-
Architecture Decision Records (ADRs)
- Overview
- Template
- Master Catalog (0001-0085)
-
Memory Kernel & Storage Formats
- ADR-0001: Graph Compression Strategy for Entity Graph
- ADR-0002: Multi-Partition Recall Fan-Out & Frozen Reten...
- ADR-0003: Completing Hypergraph Entity-Graph Graduation
- ADR-0004: Mmap Bundle Architecture & File Descriptor Sc...
- ADR-0005: spector-memory Technical Debt Hardening
- ADR-0042: Graph Recall Architecture and Cognitive Trave...
- ADR-0043: Single-VMA Bundle Layout Specification
- ADR-0044: Memory Kernel Isolation, Composition, and Layout
- ADR-0045: Spector Memory Import & Export Pipeline
- ADR-0046: Single Engram, Four Stores Storage Architecture
- ADR-0047: Episodic Memory and Engram Model Hierarchy
- ADR-0057: Remediation of Hardcoded Memory Offsets and Alignment Constants
- ADR-0062: Spector Memory Organization — Three-Plane Architecture
- ADR-0082: Index Plane Lifecycle, Derived Views, and Reconciliation
-
Active Inference Self-Model Engine (AISME)
- ADR-0006: Episodic Conversation Architecture
- ADR-0007: ReflectPathway — Biological Sleep Consolidation
- ADR-0008: Cognitive Substrate Evolution (TANGLE, GPM, M...
- ADR-0009: AISME Phase 1 — Homeostatic Affective Core
- ADR-0010: AISME Phase 2 — Free-Energy Guided Recall
- ADR-0011: AISME Phase 3 — Modern Hopfield Associative M...
- ADR-0012: AISME Phase 4 — Neural Manifold Distance (NMD)
- ADR-0013: AISME Phase 5 — Predictive Coding Narrative Self
- ADR-0014: AISME Phase 6 — Consciousness Continuity Metr...
- ADR-0015: AISME Phase 7 — Synaptic Relay Wiring & Pathw...
- ADR-0016: AISME Phase 8 — Closed-Loop Epistemic Learning
- ADR-0017: AISME Phase 9 — Generative Counterfactuals & ...
- ADR-0018: AISME Phase 10 — WanderPathway & Kernel Conti...
- ADR-0019: AISME Phase 11 — Expected Free Energy Policy ...
- ADR-0020: AISME Phase 12 — Continuous Self-Dynamics
- ADR-0023: AISME Complete Loop Closure & CognitiveVector...
- ADR-0024: Polymorphic SoulContext Hierarchy in AISME
- ADR-0027: Soul-Conditioned & Salience-Modulated Persona...
- ADR-0048: Cross-Capture Graph & CoActivation Kernel
- ADR-0049: Identity Trajectory Lyapunov Stability
- ADR-0050: Event Density Gating and Dynamic Epistemic Co...
- ADR-0051: Bayesian Online Change-Point Episode Segmenta...
- ADR-0052: Differential Privacy and Edge Anonymization
- ADR-0053: Multimodal Composite Importance Scoring
- ADR-0054: Lifespan-Adaptive Forgetting & Retention Kernel
- ADR-0055: LSR & RFF Dense Associative Memory Engineerin...
- ADR-0056: Log-Sum-ReLU (LSR) & Random Fourier Features ...
- ADR-0058: Linguistic & Vocal Prosody Expression Engine
- ADR-0063: Spacetime Vector Search and Synaptic Relay Architecture
- ADR-0064: Spacetime Simulation on Wander, Dream, and Express Pathways
- ADR-0071: Remember Cognitive Pathway Architecture
- ADR-0072: Six-Phase Fused Cognitive Scoring Pipeline
- ADR-0073: Recall Cognitive Pathway and Multi-Phase Retrieval Architecture
- ADR-0074: Reflect Cognitive Pathway and Sleep Consolidation Architecture
- ADR-0078: Salience Network and Thalamic Cognitive Profiles Architecture
-
Platform, Synapse & Clustering
- ADR-0021: Nucleus Symmetric Hardware Abstraction Layer ...
- ADR-0022: Embodied Kinesics & Phenomenological MCP Engine
- ADR-0025: Declarative MCP Tool Definitions via JSON Sch...
- ADR-0026: Dual-Plane Concurrency & Async Queue Backpres...
- ADR-0028: Dual-Plane Memory Audit Architecture (Separat...
- ADR-0029: Episodic→Semantic Lineage Provenance Region
- ADR-0030: Unified Engram Encoding Header Architecture
- ADR-0031: Unified Configuration Architecture & Bypass E...
- ADR-0032: Persona Enactment — Soul as Policy over Memory
- ADR-0033: Decoupling Cognitive & Mathematical Kernels t...
- ADR-0034: Cell Topology, Namespace Ownership, and HA Cl...
- ADR-0035: Cognitive Pathway Framework Rearchitecture
- ADR-0036: Pathway Error Handling, Isolation, and Circui...
- ADR-0037: Ingestion Boundary and Sensory Relocation
- ADR-0038: SIMD-Accelerated BM25 Lexical Scoring Optimiz...
- ADR-0039: Robust Unified Rate Limiting Architecture
- ADR-0040: Universal Apache Camel Messaging Channels
- ADR-0041: Unified Connector Architecture for Ingestion
- ADR-0059: Java 27 Upgrade Strategy and Value Class Migration
- ADR-0060: Cognitive Continuity Layer and Decoded Mind Streams
- ADR-0061: In-Memory Multi-Tenant Quartz Scheduler
- ADR-0065: Client SDK Architecture, OpenAPI, and MCP Integration
- ADR-0066: Engine & CLI Stabilization — Issue #727 Hardening
- ADR-0067: Cell-Based High Availability and Namespace-Sticky Sharding
- ADR-0068: Phileas PII Redaction Engine for Spector Synapse
- ADR-0069: Synapse-Owned Tool Access Policy
- ADR-0070: Unified Error Taxonomy and Exception Handling Architecture
- ADR-0075: Extensible LLM and Multimodal Embedding Provider SPI
- ADR-0076: Zero-Dependency Pluggable Cache Abstraction
- ADR-0077: Model B Asynchronous Task Queue and Concurrency
- ADR-0079: Asynchronous Memory Event and Telemetry Notification Bus
- ADR-0080: Observed Memory and Pathway Metrics Telemetry Architecture
- ADR-0081: Dedicated Reactive Ingress and In-Process Path Router
- ADR-0083: Namespace-Isolated Memory Analytics & Telemetry
- ADR-0084: Dual-Plane Conversation Persistence
- ADR-0085: Dynamic Synapse Configuration Overrides and Runtime Propagation
-
Modules Registry
- Overview
- Foundation Layer (/nucleus)
- Cognitive Layer (/memory)
- Gateway Layer (/synapse)
- Benchmarks & UI
- Deep Dives
-
Community
- Governance
- Contributing
- FAQ
- Glossary
- Roadmap
- 🔬 Labs
- Third-Party Legal