Repository navigation
Memory Performance
Spector Memory is engineered for microsecond-scale latency. This page documents the benchmark results and the key performance techniques that make it possible.
Measured on Intel Core Ultra 9 285K, Java 25, AVX2 256-bit (8 float lanes), ZGC:
| Benchmark | Result | Notes |
|---|---|---|
| SIMD L2 Distance (128-dim) | 0.8 µs/vector | 1.2M vectors/sec |
| SIMD L2 Distance (384-dim) | 1.5 µs/vector | 2.6M vectors/sec |
| SIMD L2 Distance (768-dim) | 2.2 µs/vector | 1.4M vectors/sec |
| SIMD L2 Distance (1024-dim) | 3.0 µs/vector | 1.0M vectors/sec |
| Reverse Index Lookup | 180 ns/lookup | O(1) packed-key map |
| Cognitive Scorer (10K × 128-dim) | 2.9 ms total | Full 6-phase pipeline |
| Batch Habituation (1K IDs) | 101 µs total | 100 ns per penalty computation |
| Tier Count Query | 17 ms / 100K calls | 170 ns per call |
| Full Pipeline (1K ingest + 100 recall) | < 50 ms/query | End-to-end latency |
| Real Embedding (qwen3-embedding 4096-dim) | 31 ms/embed | Via Ollama (network bound) |
Memory IDs are resolved in constant time using a packed-key map. The key packs (type, offset) into a single 64-bit long — zero string concatenation, zero hashing overhead.
This yields 180 ns lookups at 50K entries.
Quantized INT8 Euclidean distance uses the Java Vector API for hardware acceleration:
flowchart LR
READ["Read INT8 bytes<br/>from MemorySegment"] --> CAST["Cast INT8 → float32<br/><i>vectorized</i>"]
CAST --> DEQUANT["Affine dequantize<br/><i>float = byte × scale + min</i>"]
DEQUANT --> L2["Fused multiply-add<br/><i>accumulate squared diff</i>"]
L2 --> RESULT["L2 distance<br/><b>2.2 µs/768-dim</b>"]
style READ fill:#4a6fa5,color:white
style L2 fill:#0984e3,color:white
style RESULT fill:#00b894,color:white
The entire computation runs in SIMD registers — no intermediate Java objects are created.
Throughput: 2.2 µs/vector at 768 dimensions (1.4M vectors/sec on AVX2).
The habituation penalty module computes all penalties in a single batch call with amortized map access, processing 1K penalties in 101 µs total.
Scored records capture the cognitive header inline during scoring, eliminating N×8 off-heap re-reads per recall query.
Tier count queries use direct field access to typed store references rather than iteration, completing 100K calls in 17 ms (170 ns/call).
Each memory tier is scanned on a dedicated Virtual Thread:
gantt
title Parallel Recall: 5 concurrent scans
dateFormat X
axisFormat %L ms
section Working (100 records)
Scan :a1, 0, 1
section Episodic P1 (5K records)
Scan :a2, 0, 3
section Episodic P2 (3K records)
Scan :a3, 0, 2
section Semantic (200 headers)
Scan :a4, 0, 1
section Procedural (50 records)
Scan :a5, 0, 1
section Merge + Rank
Top-K :a6, 3, 4
Key insight: Episodic partitions use disjoint memory segments — each partition's mmap is a separate off-heap buffer. This guarantees zero contention between virtual threads, enabling perfect parallel scaling.
Fallback: If parallel scanning fails (e.g., thread pool exhaustion), the pipeline falls back to sequential scanning with identical results.
| Component | Formula | 10K memories (768-dim) |
|---|---|---|
| Episodic partition | 64B header + N × (64B + vecBytes) | 64B + 10K × 832B = 8.1 MB |
| Working memory | capacity × (64B + vecBytes) | 100 × 832B = 81 KB |
| Semantic headers | capacity × 64B | 5K × 64B = 312 KB |
| Procedural store | capacity × (64B + vecBytes) | 500 × 832B = 406 KB |
| Forward index | ~120B per entry | 10K × 120B = 1.2 MB |
| Reverse index | ~60B per entry | 10K × 60B = 600 KB |
| Total | ~10.7 MB |
tip: vs. Python Memory Layers A Python memory system stores each memory as a Python object (~500-800 bytes overhead) plus the vector in NumPy (~3KB for 768-dim float32). Spector stores the same memory in 832 bytes (64B header + 768B INT8 vector) — a 4-8× reduction.
spector-core: 276 tests ✅ (includes 15 SIMD kernel verification tests)
spector-memory: 167 tests ✅ (includes performance benchmarks + index tests)
+ 10 Ollama real embedding E2E tests (gated by OLLAMA_LIVE=true)
Total: 443 tests, 0 failures
# Run all memory tests (includes benchmark assertions)
mvn test -pl spector-memory
# Run only performance benchmarks
mvn test -pl spector-memory -Dtest=PerformanceBenchmarkTest
# Run Ollama real embedding E2E tests
OLLAMA_LIVE=true mvn test -pl spector-memory -Dtest=OllamaRealEmbeddingTest- :material-memory: [[Off-Heap Panama Design|Memory--Panama-Design]] — zero-GC architecture
- :material-lightning-bolt: [[6-Phase Scoring Pipeline|Memory--Scoring-Pipeline]] — the SIMD hot-loop
- :material-brain: [[Architecture|Memory--Architecture]] — system-level design
- Home
-
Getting Started
- Quick Start
- Installation
- Developer Guide
- JDK API Status
- MCP Server
- Java SDK
- Java API Reference
- Python SDK
- TypeScript SDK
- Spring AI Integration
- CLI Reference
- REST API
- API Playground
- Error Codes
- Configuration
- Deployment
-
Cognitive Memory
- Overview
- Getting Started
- Use Cases
- API Reference
- Concepts
- Pathways
- Scoring features
- Profiles
- Experimental
- Internals
- Design ancestry
-
Memory Kernel
- Overview
- Bundle Architecture
- Memory Shapes
- Binary Layouts & Tags
- WAL & Durability
-
Region Reference
- Overview & Index
- Partition Regions
-
Runtime Regions
- Working Memory
- Co-Activation Matrix
- Index MIDX
- Index IDPL
- Hebbian Graph
- Temporal Chains
- Temporal Facts
- Entity Directory
- Entity Names Pool
- HyperEntity Graph
- Entity Types Registry
- Relation Types Registry
- BM25 Lexical Index
- Checkpoint
- Insula (Somatic Self-Model)
- Continuity
- Provenance
- SPLADE Sparse Index
- Entity Reverse Index
- Identity Regions
- Synapse & Cortex
-
Architecture
- System Overview
- Core Concepts
- Ingestion Pipeline
- MCP Integration
- Distributed Mode
- Event Notifications
- Namespace Sharding
- Single-Namespace Scale & Capacity Limits
- Scale Benchmark Empirical Results
- Writer Quiesce Pause Empirical Results
- Kill-Owner Failover Empirical Results
- Salience & Importance Architecture
- GPU Acceleration
- Performance Tuning
- Test Framework & LLM Judge
- Chat & Visual Test Infrastructure
- Security & Data
-
Architecture Decision Records (ADRs)
- Overview
- Template
- Master Catalog (0001-0085)
-
Memory Kernel & Storage Formats
- ADR-0001: Graph Compression Strategy for Entity Graph
- ADR-0002: Multi-Partition Recall Fan-Out & Frozen Reten...
- ADR-0003: Completing Hypergraph Entity-Graph Graduation
- ADR-0004: Mmap Bundle Architecture & File Descriptor Sc...
- ADR-0005: spector-memory Technical Debt Hardening
- ADR-0042: Graph Recall Architecture and Cognitive Trave...
- ADR-0043: Single-VMA Bundle Layout Specification
- ADR-0044: Memory Kernel Isolation, Composition, and Layout
- ADR-0045: Spector Memory Import & Export Pipeline
- ADR-0046: Single Engram, Four Stores Storage Architecture
- ADR-0047: Episodic Memory and Engram Model Hierarchy
- ADR-0057: Remediation of Hardcoded Memory Offsets and Alignment Constants
- ADR-0062: Spector Memory Organization — Three-Plane Architecture
- ADR-0082: Index Plane Lifecycle, Derived Views, and Reconciliation
-
Active Inference Self-Model Engine (AISME)
- ADR-0006: Episodic Conversation Architecture
- ADR-0007: ReflectPathway — Biological Sleep Consolidation
- ADR-0008: Cognitive Substrate Evolution (TANGLE, GPM, M...
- ADR-0009: AISME Phase 1 — Homeostatic Affective Core
- ADR-0010: AISME Phase 2 — Free-Energy Guided Recall
- ADR-0011: AISME Phase 3 — Modern Hopfield Associative M...
- ADR-0012: AISME Phase 4 — Neural Manifold Distance (NMD)
- ADR-0013: AISME Phase 5 — Predictive Coding Narrative Self
- ADR-0014: AISME Phase 6 — Consciousness Continuity Metr...
- ADR-0015: AISME Phase 7 — Synaptic Relay Wiring & Pathw...
- ADR-0016: AISME Phase 8 — Closed-Loop Epistemic Learning
- ADR-0017: AISME Phase 9 — Generative Counterfactuals & ...
- ADR-0018: AISME Phase 10 — WanderPathway & Kernel Conti...
- ADR-0019: AISME Phase 11 — Expected Free Energy Policy ...
- ADR-0020: AISME Phase 12 — Continuous Self-Dynamics
- ADR-0023: AISME Complete Loop Closure & CognitiveVector...
- ADR-0024: Polymorphic SoulContext Hierarchy in AISME
- ADR-0027: Soul-Conditioned & Salience-Modulated Persona...
- ADR-0048: Cross-Capture Graph & CoActivation Kernel
- ADR-0049: Identity Trajectory Lyapunov Stability
- ADR-0050: Event Density Gating and Dynamic Epistemic Co...
- ADR-0051: Bayesian Online Change-Point Episode Segmenta...
- ADR-0052: Differential Privacy and Edge Anonymization
- ADR-0053: Multimodal Composite Importance Scoring
- ADR-0054: Lifespan-Adaptive Forgetting & Retention Kernel
- ADR-0055: LSR & RFF Dense Associative Memory Engineerin...
- ADR-0056: Log-Sum-ReLU (LSR) & Random Fourier Features ...
- ADR-0058: Linguistic & Vocal Prosody Expression Engine
- ADR-0063: Spacetime Vector Search and Synaptic Relay Architecture
- ADR-0064: Spacetime Simulation on Wander, Dream, and Express Pathways
- ADR-0071: Remember Cognitive Pathway Architecture
- ADR-0072: Six-Phase Fused Cognitive Scoring Pipeline
- ADR-0073: Recall Cognitive Pathway and Multi-Phase Retrieval Architecture
- ADR-0074: Reflect Cognitive Pathway and Sleep Consolidation Architecture
- ADR-0078: Salience Network and Thalamic Cognitive Profiles Architecture
-
Platform, Synapse & Clustering
- ADR-0021: Nucleus Symmetric Hardware Abstraction Layer ...
- ADR-0022: Embodied Kinesics & Phenomenological MCP Engine
- ADR-0025: Declarative MCP Tool Definitions via JSON Sch...
- ADR-0026: Dual-Plane Concurrency & Async Queue Backpres...
- ADR-0028: Dual-Plane Memory Audit Architecture (Separat...
- ADR-0029: Episodic→Semantic Lineage Provenance Region
- ADR-0030: Unified Engram Encoding Header Architecture
- ADR-0031: Unified Configuration Architecture & Bypass E...
- ADR-0032: Persona Enactment — Soul as Policy over Memory
- ADR-0033: Decoupling Cognitive & Mathematical Kernels t...
- ADR-0034: Cell Topology, Namespace Ownership, and HA Cl...
- ADR-0035: Cognitive Pathway Framework Rearchitecture
- ADR-0036: Pathway Error Handling, Isolation, and Circui...
- ADR-0037: Ingestion Boundary and Sensory Relocation
- ADR-0038: SIMD-Accelerated BM25 Lexical Scoring Optimiz...
- ADR-0039: Robust Unified Rate Limiting Architecture
- ADR-0040: Universal Apache Camel Messaging Channels
- ADR-0041: Unified Connector Architecture for Ingestion
- ADR-0059: Java 27 Upgrade Strategy and Value Class Migration
- ADR-0060: Cognitive Continuity Layer and Decoded Mind Streams
- ADR-0061: In-Memory Multi-Tenant Quartz Scheduler
- ADR-0065: Client SDK Architecture, OpenAPI, and MCP Integration
- ADR-0066: Engine & CLI Stabilization — Issue #727 Hardening
- ADR-0067: Cell-Based High Availability and Namespace-Sticky Sharding
- ADR-0068: Phileas PII Redaction Engine for Spector Synapse
- ADR-0069: Synapse-Owned Tool Access Policy
- ADR-0070: Unified Error Taxonomy and Exception Handling Architecture
- ADR-0075: Extensible LLM and Multimodal Embedding Provider SPI
- ADR-0076: Zero-Dependency Pluggable Cache Abstraction
- ADR-0077: Model B Asynchronous Task Queue and Concurrency
- ADR-0079: Asynchronous Memory Event and Telemetry Notification Bus
- ADR-0080: Observed Memory and Pathway Metrics Telemetry Architecture
- ADR-0081: Dedicated Reactive Ingress and In-Process Path Router
- ADR-0083: Namespace-Isolated Memory Analytics & Telemetry
- ADR-0084: Dual-Plane Conversation Persistence
- ADR-0085: Dynamic Synapse Configuration Overrides and Runtime Propagation
-
Modules Registry
- Overview
- Foundation Layer (/nucleus)
- Cognitive Layer (/memory)
- Gateway Layer (/synapse)
- Benchmarks & UI
- Deep Dives
-
Community
- Governance
- Contributing
- FAQ
- Glossary
- Roadmap
- 🔬 Labs
- Third-Party Legal