Repository navigation
Why Spector
The short answer: AI memory requires fusing semantic similarity, temporal decay, emotional valence, adaptive importance, and the associations between memories into a single sub-millisecond ranking decision. No existing database — SQL, NoSQL, or vector — can do this. Spector was built from scratch to solve exactly this problem.
When an AI agent needs to recall a relevant memory, it must answer a question that no database was designed to handle:
"Of the millions of things I've ever observed, which ones are most relevant right now, considering how similar they are to the current query, how important they were when stored, how they've decayed over time, what emotional tone they carry, and how often they've been recalled before?"
This is fundamentally different from SELECT * FROM memories WHERE topic = 'X' ORDER BY created_at DESC.
| Dimension | SQL / NoSQL | Spector |
|---|---|---|
| Cache locality | Fields scattered across B-tree pages (4KB+). Each field access causes a cache miss. | 64-byte cache-line-aligned cognitive headers. All scoring fields in one CPU cache line — one cycle. |
| Memory access | JDBC/driver serialization → heap allocation → GC pressure | Zero-copy MemorySegment reads. No heap allocation. 0.01% GC overhead at 1M memories. |
| Cognitive scoring | Impossible in SQL. Power-law decay, Bloom filter tag gating, and SIMD vector distance cannot be expressed in queries or aggregation pipelines. | Natively fused in a single off-heap scan — six sequential phases, each gating before the expensive vector math. |
| Micro-mutations | Every recall-count increment or valence adjustment triggers WAL write amplification and row locking. | Lock-free hardware-level atomic memory swaps. Thousands of concurrent updates with zero contention. |
| Search latency | Low milliseconds (ms) | Microseconds (µs) — 100-1000× faster |
quote: The Key Insight Spector is not a database with vector search bolted on. It's a cognitive scoring engine where every byte of the storage layout, every SIMD instruction, and every decay function is co-designed for a single purpose: signal-complete recall without the truncation trap, evaluating relevance, decay, valence, and associative reach simultaneously.
Vector databases (Pinecone, Weaviate, Qdrant, Milvus, pgvector) solve semantic similarity. But AI memory needs more:
The Truncation Trap: If you retrieve the top-100 nearest vectors and then sort by importance in application code, a critical 6-month-old memory that's slightly less similar than a trivial conversation from 5 minutes ago is irreversibly lost — dropped before your code ever sees it.
Spector eliminates this by fusing all signals into a single scoring pass. Every memory is evaluated against similarity, importance, decay, and valence simultaneously. Nothing is truncated prematurely.
Systems like Mem0, Zep, and Letta add a thin layer over existing databases. They inherit all the database limitations above, plus:
- Network hop tax: Every recall crosses a REST boundary → 1-5ms added
- JSON serialization: GC pressure from parsing/encoding every result
- No hardware co-design: Cannot exploit cache-line alignment, SIMD, or off-heap storage
Spector grounds its architecture in formal cognitive science and the Memory Fundamentals Specification — modeling decay, associative networks, and two-factor retention as computational solutions to retrieval failure modes:
| Cognitive Principle | Spector Implementation | Operational Effect |
|---|---|---|
| Prediction Error (Surprise) | Adaptive surprise detection (z-score) | Novel memories get higher importance automatically |
| Ebbinghaus forgetting curve | Power-law temporal decay (12-bucket table) | Memories fade naturally, frequently-recalled ones persist |
| Bjork Two-Factor theory | Storage strength × retrieval strength | Recalled memories become progressively easier to surface |
| Hebb's rule | Co-activation graph with spreading activation | Related memories cluster and reinforce each other |
| Emotional modulation (Valence) | Affective valence + arousal modulation | High-arousal events resist decay — like flashbulb memories |
| Consolidation & Replay | Sleep consolidation cycles | Background process promotes important memories, prunes weak ones |
| Benchmark | Result | Notes |
|---|---|---|
| Cognitive recall | Ultra-low latency | Hardware-accelerated in-process SIMD |
| Vector search p50 | 88–143µs | 10K–100K docs, HNSW M=16 |
| Peak QPS (16 threads) | 61,011 | Concurrent vectorSearch |
| GC overhead | 0.01% | 1 pause / 100K searches |
| SVASQ-8 compression | 4× smaller | 99.5%+ recall preserved |
Spector needs no Docker, no external database, and no extra services. Reach it from any language over MCP or REST/gRPC, drive it from the Python SDK, or embed it directly:
- Embedded library: Add a single JAR to your Java/Kotlin/Scala application
- Standalone server: REST + gRPC + MCP APIs on a single port
- Clustered mode: gRPC fan-out across multiple nodes with namespace sharding
- Kubernetes: Helm chart with horizontal pod autoscaling
- Spring AI / Micronaut: First-class framework integration
Spector includes a 37+ tool MCP server for AI agent integration — Claude Desktop, Cursor, and custom agents can use cognitive memory, graph recall, and namespace RBAC out of the box. No wrapper libraries needed.
Every tenant gets physically separate files with independent encryption keys:
- AES-256-GCM text and WAL encryption (per-tenant keys)
- HMAC blind indexing for tag search over encrypted data
- BYOK — users can supply their own encryption keys; server operators cannot decrypt BYOK data
-
File-level isolation — no shared database tables, no
WHERE tenant_id = ?leaks
| Feature | Spector | Pinecone | Weaviate | Qdrant | Milvus | ChromaDB | pgvector |
|---|---|---|---|---|---|---|---|
| Deployment | Embedded / Standalone / Clustered | Cloud SaaS | Self-hosted / Cloud | Self-hosted / Cloud | Self-hosted / Cloud | Embedded / Server | PostgreSQL extension |
| Language | Java 25 | Managed | Go | Rust | Go/C++ | Python | C |
| Dependencies | Zero (JDK only) | N/A (SaaS) | Docker | Docker | Docker + etcd + MinIO | Python packages | PostgreSQL |
| SIMD acceleration | ✅ AVX2/AVX-512/NEON | ✅ (internal) | ✅ | ✅ | ✅ | ❌ | ✅ (pgvector 0.5+) |
| Off-heap / Zero GC | ✅ Panama FFM & V4 Bundles | N/A | Partial | ✅ (Rust) | Partial | ❌ | N/A |
| Fused cognitive scoring | ✅ 6-Phase SIMD Recall Pathway | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
| Hybrid search | ✅ HNSW + BM25 + RRF | ✅ | ✅ | ✅ | ✅ | ❌ | ❌ |
| Built-in MCP server | ✅ 37+ tools | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
| Cognitive memory | ✅ 4-tier with decay & consolidation | ❌ | ❌ | ❌ | ❌ | ❌ | ❌ |
| Quantization | SVASQ-8/4, IVF-PQ | ✅ | ✅ BQ | ✅ SQ/PQ | ✅ IVF-PQ/SQ | ❌ | ❌ |
| GPU acceleration | ✅ CUDA via Panama | ✅ | ❌ | ❌ | ✅ | ❌ | ❌ |
| License | Apache 2.0 | Proprietary | BSD-3 | Apache 2.0 | Apache 2.0 | Apache 2.0 | PostgreSQL |
| Feature | Spector Memory | Mem0 | Letta (MemGPT) | Zep | Stanford Generative Agents |
|---|---|---|---|---|---|
| Benchmark Accuracy | MindSpan: 100% LongMemEval: 94% LoCoMo: 85% |
LoCoMo: 62–68% LongMemEval: 62–68% |
LoCoMo: ~65% | LoCoMo: 75–80% LongMemEval: 75–80% |
N/A (Research) |
| Recall latency | 3–12ms (in-process SIMD) | 50–200ms (API) / 657ms (Graph) | 100ms+ | 50–150ms (API) / 632ms (Graph) | N/A |
| Context injected | 1,257–1,731 tokens (compact) | 1,764 tokens | 2,000+ tokens | 3,911 tokens | Unbounded |
| Temporal decay | ✅ Power-law (configurable) | ❌ None | ❌ Agent-managed | ✅ Limited | ✅ Exponential |
| Scoring model | ACT-R inspired | Vector similarity | Agent-managed | Hybrid | Additive |
| Two-Factor strengthening | ✅ Bjork model (Strength Region) | ❌ | ❌ | ❌ | ❌ |
| Emotional valence | ✅ Amygdala model | ❌ | ❌ | ❌ | ❌ |
| Salience profiles | ✅ Persona + interest-based | ❌ | ❌ | ❌ | ❌ |
| Sleep consolidation | ✅ Hippocampus model | ❌ | ❌ | ❌ | ❌ |
| Hebbian associations | ✅ Co-activation graph | ❌ | ❌ | ❌ | ❌ |
| Entity knowledge graph | ✅ HyperEntityGraph | ❌ | ❌ | ❌ | ❌ |
| GC pressure | 0.01% (off-heap) | High (Python) | High (Python) | Moderate | N/A |
| MCP integration | ✅ Built-in (37+ tools) | ❌ | ❌ | ❌ | ❌ |
| Infrastructure | Zero (embedded JVM) | Redis + API | PostgreSQL + API | PostgreSQL + API | Research code |
- You need a managed cloud service with zero ops → Pinecone
- You want a pure-Python library with no JVM in the process at all → ChromaDB
- You already have PostgreSQL and just want to add basic vector search → pgvector
- You need multi-modal search (images, video, audio) → Weaviate, Milvus
📖 Full Benchmark Report → · Performance Tuning → · Cognitive Memory →
Corrections welcome — if any comparison is inaccurate, please open an issue.
- Home
-
Getting Started
- Quick Start
- Installation
- Developer Guide
- JDK API Status
- MCP Server
- Java SDK
- Java API Reference
- Python SDK
- TypeScript SDK
- Spring AI Integration
- CLI Reference
- REST API
- API Playground
- Error Codes
- Configuration
- Deployment
-
Cognitive Memory
- Overview
- Getting Started
- Use Cases
- API Reference
- Concepts
- Pathways
- Scoring features
- Profiles
- Experimental
- Internals
- Design ancestry
-
Memory Kernel
- Overview
- Bundle Architecture
- Memory Shapes
- Binary Layouts & Tags
- WAL & Durability
-
Region Reference
- Overview & Index
- Partition Regions
-
Runtime Regions
- Working Memory
- Co-Activation Matrix
- Index MIDX
- Index IDPL
- Hebbian Graph
- Temporal Chains
- Temporal Facts
- Entity Directory
- Entity Names Pool
- HyperEntity Graph
- Entity Types Registry
- Relation Types Registry
- BM25 Lexical Index
- Checkpoint
- Insula (Somatic Self-Model)
- Continuity
- Provenance
- SPLADE Sparse Index
- Entity Reverse Index
- Identity Regions
- Synapse & Cortex
-
Architecture
- System Overview
- Core Concepts
- Ingestion Pipeline
- MCP Integration
- Distributed Mode
- Event Notifications
- Namespace Sharding
- Single-Namespace Scale & Capacity Limits
- Scale Benchmark Empirical Results
- Writer Quiesce Pause Empirical Results
- Kill-Owner Failover Empirical Results
- Salience & Importance Architecture
- GPU Acceleration
- Performance Tuning
- Test Framework & LLM Judge
- Chat & Visual Test Infrastructure
- Security & Data
-
Architecture Decision Records (ADRs)
- Overview
- Template
- Master Catalog (0001-0085)
-
Memory Kernel & Storage Formats
- ADR-0001: Graph Compression Strategy for Entity Graph
- ADR-0002: Multi-Partition Recall Fan-Out & Frozen Reten...
- ADR-0003: Completing Hypergraph Entity-Graph Graduation
- ADR-0004: Mmap Bundle Architecture & File Descriptor Sc...
- ADR-0005: spector-memory Technical Debt Hardening
- ADR-0042: Graph Recall Architecture and Cognitive Trave...
- ADR-0043: Single-VMA Bundle Layout Specification
- ADR-0044: Memory Kernel Isolation, Composition, and Layout
- ADR-0045: Spector Memory Import & Export Pipeline
- ADR-0046: Single Engram, Four Stores Storage Architecture
- ADR-0047: Episodic Memory and Engram Model Hierarchy
- ADR-0057: Remediation of Hardcoded Memory Offsets and Alignment Constants
- ADR-0062: Spector Memory Organization — Three-Plane Architecture
- ADR-0082: Index Plane Lifecycle, Derived Views, and Reconciliation
-
Active Inference Self-Model Engine (AISME)
- ADR-0006: Episodic Conversation Architecture
- ADR-0007: ReflectPathway — Biological Sleep Consolidation
- ADR-0008: Cognitive Substrate Evolution (TANGLE, GPM, M...
- ADR-0009: AISME Phase 1 — Homeostatic Affective Core
- ADR-0010: AISME Phase 2 — Free-Energy Guided Recall
- ADR-0011: AISME Phase 3 — Modern Hopfield Associative M...
- ADR-0012: AISME Phase 4 — Neural Manifold Distance (NMD)
- ADR-0013: AISME Phase 5 — Predictive Coding Narrative Self
- ADR-0014: AISME Phase 6 — Consciousness Continuity Metr...
- ADR-0015: AISME Phase 7 — Synaptic Relay Wiring & Pathw...
- ADR-0016: AISME Phase 8 — Closed-Loop Epistemic Learning
- ADR-0017: AISME Phase 9 — Generative Counterfactuals & ...
- ADR-0018: AISME Phase 10 — WanderPathway & Kernel Conti...
- ADR-0019: AISME Phase 11 — Expected Free Energy Policy ...
- ADR-0020: AISME Phase 12 — Continuous Self-Dynamics
- ADR-0023: AISME Complete Loop Closure & CognitiveVector...
- ADR-0024: Polymorphic SoulContext Hierarchy in AISME
- ADR-0027: Soul-Conditioned & Salience-Modulated Persona...
- ADR-0048: Cross-Capture Graph & CoActivation Kernel
- ADR-0049: Identity Trajectory Lyapunov Stability
- ADR-0050: Event Density Gating and Dynamic Epistemic Co...
- ADR-0051: Bayesian Online Change-Point Episode Segmenta...
- ADR-0052: Differential Privacy and Edge Anonymization
- ADR-0053: Multimodal Composite Importance Scoring
- ADR-0054: Lifespan-Adaptive Forgetting & Retention Kernel
- ADR-0055: LSR & RFF Dense Associative Memory Engineerin...
- ADR-0056: Log-Sum-ReLU (LSR) & Random Fourier Features ...
- ADR-0058: Linguistic & Vocal Prosody Expression Engine
- ADR-0063: Spacetime Vector Search and Synaptic Relay Architecture
- ADR-0064: Spacetime Simulation on Wander, Dream, and Express Pathways
- ADR-0071: Remember Cognitive Pathway Architecture
- ADR-0072: Six-Phase Fused Cognitive Scoring Pipeline
- ADR-0073: Recall Cognitive Pathway and Multi-Phase Retrieval Architecture
- ADR-0074: Reflect Cognitive Pathway and Sleep Consolidation Architecture
- ADR-0078: Salience Network and Thalamic Cognitive Profiles Architecture
-
Platform, Synapse & Clustering
- ADR-0021: Nucleus Symmetric Hardware Abstraction Layer ...
- ADR-0022: Embodied Kinesics & Phenomenological MCP Engine
- ADR-0025: Declarative MCP Tool Definitions via JSON Sch...
- ADR-0026: Dual-Plane Concurrency & Async Queue Backpres...
- ADR-0028: Dual-Plane Memory Audit Architecture (Separat...
- ADR-0029: Episodic→Semantic Lineage Provenance Region
- ADR-0030: Unified Engram Encoding Header Architecture
- ADR-0031: Unified Configuration Architecture & Bypass E...
- ADR-0032: Persona Enactment — Soul as Policy over Memory
- ADR-0033: Decoupling Cognitive & Mathematical Kernels t...
- ADR-0034: Cell Topology, Namespace Ownership, and HA Cl...
- ADR-0035: Cognitive Pathway Framework Rearchitecture
- ADR-0036: Pathway Error Handling, Isolation, and Circui...
- ADR-0037: Ingestion Boundary and Sensory Relocation
- ADR-0038: SIMD-Accelerated BM25 Lexical Scoring Optimiz...
- ADR-0039: Robust Unified Rate Limiting Architecture
- ADR-0040: Universal Apache Camel Messaging Channels
- ADR-0041: Unified Connector Architecture for Ingestion
- ADR-0059: Java 27 Upgrade Strategy and Value Class Migration
- ADR-0060: Cognitive Continuity Layer and Decoded Mind Streams
- ADR-0061: In-Memory Multi-Tenant Quartz Scheduler
- ADR-0065: Client SDK Architecture, OpenAPI, and MCP Integration
- ADR-0066: Engine & CLI Stabilization — Issue #727 Hardening
- ADR-0067: Cell-Based High Availability and Namespace-Sticky Sharding
- ADR-0068: Phileas PII Redaction Engine for Spector Synapse
- ADR-0069: Synapse-Owned Tool Access Policy
- ADR-0070: Unified Error Taxonomy and Exception Handling Architecture
- ADR-0075: Extensible LLM and Multimodal Embedding Provider SPI
- ADR-0076: Zero-Dependency Pluggable Cache Abstraction
- ADR-0077: Model B Asynchronous Task Queue and Concurrency
- ADR-0079: Asynchronous Memory Event and Telemetry Notification Bus
- ADR-0080: Observed Memory and Pathway Metrics Telemetry Architecture
- ADR-0081: Dedicated Reactive Ingress and In-Process Path Router
- ADR-0083: Namespace-Isolated Memory Analytics & Telemetry
- ADR-0084: Dual-Plane Conversation Persistence
- ADR-0085: Dynamic Synapse Configuration Overrides and Runtime Propagation
-
Modules Registry
- Overview
- Foundation Layer (/nucleus)
- Cognitive Layer (/memory)
- Gateway Layer (/synapse)
- Benchmarks & UI
- Deep Dives
-
Community
- Governance
- Contributing
- FAQ
- Glossary
- Roadmap
- 🔬 Labs
- Third-Party Legal