Repository navigation
Architecture Overview
Spector is a SIMD-accelerated AI memory backbone with built-in MCP server, hybrid search, and multi-tier cognitive memory. This page covers the system architecture, data flows, threading model, and memory architecture that make sub-millisecond, agent-native search possible.
graph TB
subgraph Clients["Client Interfaces & SDKs"]
claude["🤖 Claude Desktop"]
cursor["✏️ Cursor / Windsurf"]
agents["🦾 Autonomous Agents"]
py["🐍 Python SDK"]
ts["🔷 TypeScript SDK"]
sdk["☕ Java Client SDK"]
spring["🌱 Spring AI"]
cli["🖥️ spector CLI"]
rest["🌐 REST / gRPC"]
end
subgraph Transport["Synapse Application Layer"]
mcp["MCP Server<br/><i>stdio · Streamable HTTP · 37+ tools</i>"]
armeria["Armeria Gateway :7070<br/><i>REST + gRPC + SSE streaming</i>"]
persona["Persona Enactment<br/><i>Dual-process cognitive appraisal</i>"]
end
subgraph Engine["Spector Memory Engine"]
subgraph Pathways["Cognitive Pathways"]
rem["Remember Pathway<br/><i>Surprise · Flashbulb · Dedup</i>"]
rec["Recall Pathway<br/><i>6-Phase SIMD Fused Scoring</i>"]
ref["Reflect Pathway<br/><i>Sleep Consolidation · Replay</i>"]
drm["Dream Pathway<br/><i>Counterfactual Simulation</i>"]
dec["Decide Pathway<br/><i>Active Inference Policy</i>"]
exp["Express Pathway<br/><i>Prosody · Stylometry · Kinesics</i>"]
wan["Wander Pathway<br/><i>DMN Associative Synthesis</i>"]
end
subgraph Memory["4-Tier Cortex & Graphs"]
cortex["4-Tier Cortex<br/><i>Working · Episodic · Semantic · Procedural</i>"]
hebbian["Cognitive Graphs<br/><i>Hebbian · Temporal · HyperEntity</i>"]
decay["Memory Decay<br/><i>Power-law forgetting · Bjork strength</i>"]
end
subgraph Search["Search & Retrieval Stack"]
hybrid["Hybrid Retrieval<br/><i>BM25 + Dense + SPLADE</i>"]
hnsw["HNSW Graph Index<br/><i>M=16, ef=200</i>"]
colbert["ColBERT v2 Reranking<br/><i>Late-interaction MaxSim</i>"]
end
end
subgraph Kernel["⚡ Memory Kernel (spector-kernel — Zero GC)"]
direction TB
ns["NamespaceKernel Facade"]
bundles["V4 Single-VMA Bundles<br/><i>runtime.bundle · partition.bundle · identity.bundle</i>"]
shapes["8 Sealed Memory Shapes<br/><i>Record · Append · Graph · Chain · Hash · Insula</i>"]
panama["Panama FFM Storage<br/><i>Shared Arena · MemorySegment · mmap</i>"]
simd["Hardware SIMD Acceleration<br/><i>Vector API · AVX2 / AVX-512 / NEON</i>"]
gpu["GPU Acceleration<br/><i>CUDA via Panama FFM</i>"]
end
subgraph Observe["Observability & Telemetry"]
events["TelemetryBus<br/><i>Event streams</i>"]
metrics["Micrometer<br/><i>Prometheus export</i>"]
sse["SSE Event Stream<br/><i>Cortex 3D Galaxy</i>"]
end
claude & cursor & agents --> mcp
py & ts & sdk & spring --> armeria
cli & rest --> armeria
mcp & armeria & persona --> Pathways
Pathways --> Memory & Search
Memory & Search --> ns
ns --> bundles --> shapes --> panama
panama --> simd
gpu -.->|batch compute| simd
Engine --> events
events --> metrics & sse
style Clients fill:#5b6abf,stroke:#e94560,color:#fff
style Transport fill:#4a6fa5,stroke:#3b82f6,color:#fff
style Engine fill:#3b82f6,stroke:#7c3aed,color:#fff
style Kernel fill:#1e293b,stroke:#0f172a,color:#fff
style Pathways fill:#2563eb,stroke:#1d4ed8,color:#fff
style Memory fill:#1d4ed8,stroke:#1e40af,color:#fff
style Search fill:#1e40af,stroke:#1e3a8a,color:#fff
style Observe fill:#5b6abf,stroke:#7c3aed,color:#fff
graph LR
subgraph Ingest["Ingest"]
docs["📄 Documents"]
files["📁 Files"]
api["🌐 API Data"]
end
subgraph Process["Process"]
chunk["✂️ Chunk"]
embed["🧬 Embed"]
quantize["🗜️ Quantize"]
end
subgraph Store["Store"]
vectors["📊 Vector Index<br/><i>HNSW · IVF-PQ</i>"]
text["📝 Text Index<br/><i>BM25</i>"]
memory["🧠 Cognitive Store<br/><i>4-tier cortex</i>"]
end
subgraph Query["Query"]
search["🔍 Hybrid Search"]
recall["💭 Memory Recall"]
rag["🤖 RAG Pipeline"]
end
docs & files & api --> chunk --> embed --> quantize
quantize --> vectors & text & memory
vectors & text --> search --> rag
memory --> recall --> rag
style Ingest fill:#5b6abf,stroke:#e94560,color:#fff
style Process fill:#4a6fa5,stroke:#3b82f6,color:#fff
style Store fill:#3b82f6,stroke:#7c3aed,color:#fff
style Query fill:#7c3aed,stroke:#e94560,color:#fff
graph LR
subgraph Embedded["Embedded Mode"]
lib["SpectorMemory API<br/><i>In-process · zero-network · drop-in JAR</i>"]
end
subgraph Standalone["Standalone Mode"]
jar["java -jar spector.jar<br/><i>Engine + MCP + REST/gRPC + SSE</i>"]
end
subgraph Distributed["Distributed Mode"]
coord["Coordinator<br/><i>Query routing · fan-out</i>"]
s1["Shard 1"] & s2["Shard 2"] & s3["Shard N"]
coord --> s1 & s2 & s3
end
style Embedded fill:#4a6fa5,stroke:#3b82f6,color:#fff
style Standalone fill:#3b82f6,stroke:#7c3aed,color:#fff
style Distributed fill:#7c3aed,stroke:#e94560,color:#fff
Spector's MCP server runs in-process — the agent's tool calls go directly into SIMD kernels with zero network hops, zero serialization, and zero GC pressure. This is the architectural advantage over adapters that wrap a database behind an HTTP API.
graph TB
subgraph Agents["AI Agents"]
claude["🤖 Claude Desktop"]
cursor["✏️ Cursor / Windsurf"]
cline["🔧 Cline / Aider"]
custom["🦾 Autonomous Multi-Agents"]
end
subgraph MCP["MCP Server — Dual Transport · JSON-RPC 2.0"]
transport["Transport Layer<br/><i>stdio (stdin/stdout) for CLI agents<br/>Streamable HTTP (/mcp) for remote agents</i>"]
registry["SpectorToolRegistry<br/><i>37+ tools · dynamic route dispatch</i>"]
handler["McpToolHandler<br/><i>Base class · thread-safe · virtual threads</i>"]
subgraph Mem["1. Memory Tier Operations (16 Tools)"]
m1["memory_remember — Store with importance & tags"]
m2["memory_recall — Fused SIMD scoring recall"]
m3["memory_scratchpad — Working-memory scratchpad"]
m4["memory_reinforce — Outcome feedback (+/-)"]
m5["memory_forget — Tombstone intentional forgetting"]
m6["memory_status — Per-tier statistics & health"]
m7["memory_introspect — Metamemory self-reflection"]
m8["memory_suppress — Temporary recall suppression"]
m9["memory_resolve — Mark resolved/unresolved"]
m10["memory_reminder — Proactive intent triggers"]
m11["memory_why_not — Explain recall misses"]
m12["memory_compute_importance — Pre-ingest scoring"]
m13["memory_inspect — Full cognitive X-ray"]
m14["memory_export — Bulk JSON memory export"]
m15["memory_browse — Browse by tag/tier filter"]
m16["memory_salience — Inspect & tune salience profile"]
end
subgraph GraphContext["2. Graph & Multi-Evidence Retrieval (7 Tools)"]
g1["memory_graph_recall — Spreading activation graph walk"]
g2["memory_context_pack — Assembled agent prompt pack"]
g3["memory_fact_history — Temporal chain evolution"]
g4["memory_persona_context — Soul-aligned contextual injection"]
g5["memory_multi_evidence_recall — Multi-vector consensus"]
g6["vector_search — Pure vector cosine similarity"]
g7["memory_express — Natural language memory synthesis"]
end
subgraph NamespaceRBAC["3. Namespace & Multi-Tenancy (9 Tools)"]
n1["namespace_create — Provision isolated namespace"]
n2["namespace_list — Enumerate active namespaces"]
n3["namespace_info — Inspect V4 bundle layout & size"]
n4["namespace_switch — Set active session namespace"]
n5["namespace_set_default — Update default namespace"]
n6["namespace_delete — Safely purge namespace files"]
n7["namespace_grant — RBAC access delegation"]
n8["namespace_revoke — Revoke access permissions"]
n9["namespace_list_grants — Audit security grants"]
end
subgraph SoulPolicy["4. Agent Soul & Persona Enactment (5 Tools)"]
s1["update_agent_soul — Mutate agent persona & dogmas"]
s2["persona_enact — Dual-process cognitive appraisal"]
s3["account_introspect — Account & tenant introspection"]
s4["invoke_connector_route — External data connector dispatch"]
s5["send_notification — Dispatch proactive agent alerts"]
end
end
subgraph Core["In-Process Engine — Zero Network Overhead"]
pathways["Cognitive Pathways<br/><i>Remember · Recall · Reflect · Dream · Decide · Express · Wander</i>"]
kernel["Sealed Memory Kernel<br/><i>V4 Bundles · 8 Shapes · Panama FFM</i>"]
end
Agents -->|stdio / HTTP| transport --> registry --> handler
handler --> Mem & GraphContext & NamespaceRBAC & SoulPolicy
Mem & GraphContext & NamespaceRBAC & SoulPolicy --> pathways --> kernel
style Agents fill:#5b6abf,stroke:#e94560,color:#fff
style MCP fill:#4a6fa5,stroke:#3b82f6,color:#fff
style Mem fill:#3b82f6,stroke:#2563eb,color:#fff
style GraphContext fill:#2563eb,stroke:#1d4ed8,color:#fff
style NamespaceRBAC fill:#1d4ed8,stroke:#1e40af,color:#fff
style SoulPolicy fill:#1e40af,stroke:#1e3a8a,color:#fff
style Core fill:#1e293b,stroke:#0f172a,color:#fff
sequenceDiagram
participant Agent as 🤖 AI Agent
participant MCP as 📡 MCP Server
participant Tools as 🔧 ToolRegistry
participant Memory as 🧠 SpectorMemory
participant SIMD as 🔬 SIMD (off-heap)
Note over Agent,SIMD: Single JVM process — no HTTP, no gRPC, no serialization
Agent->>MCP: tools/call {"name": "memory_remember", ...}
MCP->>Tools: Route → MemoryRememberTool
Tools->>Memory: remember(text, tags, importance)
Memory->>SIMD: Embed → HNSW insert → tier assign
SIMD-->>Agent: ✅ memoryId + tier (~1ms)
Agent->>MCP: tools/call {"name": "memory_recall", ...}
MCP->>Tools: Route → MemoryRecallTool
Tools->>Memory: recall(query, topK)
Memory->>SIMD: Fused scoring: sim × importance × decay
SIMD-->>Agent: 📋 Ranked memories (ultra-fast)
Agent->>MCP: tools/call {"name": "memory_introspect", ...}
MCP->>Tools: Route → MemoryIntrospectTool
Tools->>Runtime: memory().introspect(topic)
Runtime->>SIMD: Confidence + knowledge-gap analysis over tiers
SIMD-->>Agent: 🔍 Knowledge report (~0.2ms)
| Metric | Spector (in-process) | Typical MCP adapter |
|---|---|---|
| Architecture | Engine + MCP in one JVM | Python → HTTP → DB → HTTP → agent |
| Search latency | 88µs (SIMD) | 5–50ms (network round-trip) |
| Memory recall | Ultra-low latency (fused scoring) | 50–200ms (Mem0/Letta/Zep) |
| Tools | 16 (cognitive memory tools) | 3–5 basic CRUD |
| GC pressure | Zero (Panama off-heap) | Full GC overhead |
| Deployment | java -jar spector.jar |
Python + pip + DB + config |
Tip
For full MCP integration details, tool schemas, and Claude Desktop configuration, see the dedicated MCP Integration page.
graph LR
subgraph "🔬 Foundation & Acceleration (nucleus/)"
core["spector-core<br/><i>Compute SPIs & Quantization</i>"]
cpu["spector-cpu<br/><i>Java 25 SIMD Kernels</i>"]
gpu["spector-gpu<br/><i>Panama FFM + CUDA GPU</i>"]
hdc["spector-hdc<br/><i>Hyperdimensional vectors</i>"]
index["spector-index<br/><i>HNSW + SpectorIndex + BM25</i>"]
commons["spector-commons<br/><i>Error codes & concurrency</i>"]
config["spector-config<br/><i>SpectorProperties & YAML</i>"]
events["spector-events<br/><i>Telemetry event bus</i>"]
testsupport["spector-test-support<br/><i>Harnesses & mocks</i>"]
end
subgraph "🧠 Cognitive Memory Layer (memory/)"
kernel["spector-kernel<br/><i>Sealed Off-Heap Bundle Kernel</i>"]
memory["spector-memory<br/><i>4-Tier Memory, Cognitive Pathways & Daemons</i>"]
providerapi["spector-provider-api<br/><i>Provider SPI</i>"]
providers["spector-providers<br/><i>AI Providers (Ollama, OpenAI, ONNX)</i>"]
ingestion["spector-ingestion<br/><i>Sensory & file ingest pipeline</i>"]
inspect["spector-inspect<br/><i>Bundle inspection CLI</i>"]
metrics["spector-metrics<br/><i>Micrometer + Prometheus</i>"]
end
subgraph "⚡ Nervous System & Gateways (synapse/)"
synapse["spector-synapse<br/><i>Spring Boot 4 REST/SSE & Chat Graph</i>"]
gateway["spector-gateway<br/><i>API Gateway & routing</i>"]
connector["spector-connector<br/><i>Apache Camel connectors</i>"]
mcp["spector-mcp<br/><i>MCP Server — Agent-native</i>"]
cli["spector-cli<br/><i>spectorctl CLI & standalone spector.jar</i>"]
spring["spector-spring<br/><i>Spring AI VectorStore</i>"]
batch["spector-batch<br/><i>Batch migration engine</i>"]
end
subgraph "🌐 Distributed (cluster/)"
cluster["spector-cluster<br/><i>Multi-node coordination</i>"]
end
subgraph "📦 Client SDKs (sdks/)"
javaclient["spector-client<br/><i>Java client SDK</i>"]
end
subgraph "📈 Performance & Validation (bench/)"
bench["spector-bench<br/><i>JMH benchmarks & cognitive eval</i>"]
end
Note
Index implementations in spector-index: hnsw/ (graph-based ANN, Quantized HNSW), spectrum/ (SpectorIndex, multi-tier sharding), bm25/ (keyword scoring + analyzers), splade/ (sparse neural representations).
graph TD
synapse["🌐 synapse"] --> mcp["🤖 mcp"]
synapse --> connector["🔌 connector"]
synapse --> metrics["📈 metrics"]
synapse --> events["📡 events"]
synapse --> memory["🧠 memory"]
mcp --> memory
mcp --> ingestion["📥 ingestion"]
cli["🖥️ cli"] --> memory
cli --> mcp
cli --> ingestion
memory --> index["📊 index"]
memory --> core["🔬 core"]
memory --> cpu["⚡ cpu"]
memory --> config["⚙️ config"]
memory --> providerapi["🧬 provider-api"]
index --> core
index --> config
index --> commons["📄 commons"]
gpu --> index
gpu --> core
gpu --> commons
cpu --> core
cpu --> commons
metrics --> memory
metrics --> events
connector --> ingestion
connector --> providerapi
spring["🌱 spring"] --> memory
spring --> metrics
bench["🧪 bench"] --> memory
bench --> providers["🤖 providers"]
Legend: Solid arrows = compile dependency. Dotted arrow (
bench) = benchmark execution dependency.
Dependency rules:
| Path | Description |
|---|---|
memory → kernel, index, core, cpu |
Cognitive memory composes kernel bundles, 4-tier engrams, and HNSW/BM25 indexes |
kernel → (JDK only) |
Sealed storage kernel — zero external dependencies, Panama FFM + Vector API only |
cli → memory + mcp + ingestion |
CLI with local batch and remote (client) modes |
synapse → memory + ingestion |
Spring Boot 4 application: REST + SSE + cluster coordination |
mcp → memory + ingestion |
MCP agent entry point (in-process, zero network) |
commons ← ingestion & memory |
Houses IngestionBoundary decoupling sensory ingestion from memory (ADR-0037) |
index → core, config, commons |
HNSW, SpectorIndex, BM25, and SPLADE storage foundations |
!!! important
No circular dependencies. spector-memory contains both vector search and cognitive memory stores, keeping the API gateway (spector-synapse) decoupled from low-level storage.
Spector structures all cognitive processes into a unified, composable pipeline architecture modeled after neurobiological synaptic pathways. Every cognitive operation extends AbstractPathway<S, R> and executes a directed sequence of SynapticRelay<S> stages orchestrated by PathwayEngine<S>.
Two types are easy to confuse, so it is worth stating the split explicitly:
-
Pathway<I, O>is the public operation — typed input to typed output. It owns scope entry/exit, the conduction outcome, metrics, and the projection of a conducted signal into a report. -
PathwayEngine<S>is the relay conductor — signal in, same signal out. It runs the stage list, applies each stage'sErrorPolicy, and records traces. It is not aPathwayand deliberately does not implement it.
AbstractPathway bridges the two by composition: it holds an engine and conducts it, then projects. Stage lists are declared in a PathwayRecipe and assembled by PathwayComposer, which is the only supported authoring API — it is where the build-time safety checks live (rejecting a retry on a non-idempotent relay, a timeout on a non-interruptible one, or ABORT inside a divergent branch).
| Pathway | Input Signal | Result / Report | Biological Analog & Key Functions |
|---|---|---|---|
| Remember | RememberSignal |
RememberResult |
Encoding & Consolidation: Dopamine-modulated surprise gating, perceptual dedup, transactional cortical commit, Hebbian graph co-activation linking, and knowledge graph hyper-edge enrichment. |
| Recall | RecallSignal |
List<CognitiveResult> |
Retrieval & Reconstruction: 6-phase fused hybrid search (lexical BM25, dense HNSW, SPLADE), spreading activation across Hebbian manifolds, lateral inhibition, and ColBERT v2 token MaxSim reranking. |
| Reflect | ReflectSignal |
ReflectReport |
Circadian Sleep Consolidation: Non-REM/REM sleep cycle simulation, synaptic homeostasis & power-law decay, soul-drift re-fusion, episodic clustering, and off-heap partition compaction. |
| Dream | DreamSignal |
DreamReport |
Counterfactual Replay & Discovery: Generative self-model replay, Langevin stochastic discovery, synthetic episode generation, and Expected Free Energy (EFE) triage. |
| Decide | DecideSignal |
DecideReport |
Active Inference Policy Selection: Free Energy Principle (FEP) optimization, homeostatic goal appraisal, policy rollouts, and action selection. |
| Express | ExpressSignal |
ExpressReport |
Embodied Stylometry & Prosody: Affective somatic feedback, vocal prosody vector calculation, idiolect stylometric matching, SSML tagging, and 3D blendshape parameter synthesis. |
| Wander | WanderSignal |
WanderReport |
Default Mode Network (DMN): Idle-state autobiographical sampling, Hopfield continuous attractor energy minimization, spontaneous associative synthesis, and longitudinal continuity checkpointing. |
Each relay in a cognitive pathway can be declaratively wrapped with resilience and observability decorators via PathwayComposer or builder composition. To guarantee deterministic behavior across all pathways, decorators execute in a strict onion order:
Pathway Conduction Invocation
│
▼
┌─────────────────────────────────────────────────────────────┐
│ 1. Tracing & Metrics Interceptor │
│ Measures relay latency; emits spector.pathway.relay │
├─────────────────────────────────────────────────────────────┤
│ 2. TimeoutRelay │
│ Enforces per-relay deadlines; emits onTimeout │
├─────────────────────────────────────────────────────────────┤
│ 3. RetryRelay │
│ Retries transient failures with backoff; emits onRetry │
├─────────────────────────────────────────────────────────────┤
│ 4. CircuitBreakerRelay │
│ Guards downstream services; emits onCircuitEvent │
├─────────────────────────────────────────────────────────────┤
│ 5. BulkheadRelay │
│ Isolates concurrency per relay; emits onBulkheadReject │
├─────────────────────────────────────────────────────────────┤
│ 6. DegradationPolicy │
│ Gracefully degrades or bypasses stage; emits onDegraded │
├─────────────────────────────────────────────────────────────┤
│ 7. Underlying SynapticRelay │
│ Executes core domain logic (pure cognitive step) │
└─────────────────────────────────────────────────────────────┘
To prevent cascading failures when external providers (LLMs, embedding APIs, graph engines) experience outages, relays support circuit breaker isolation:
-
States:
CLOSED(normal operation),OPEN(tripped after consecutive failures; fast-fails requests), andHALF_OPEN(probe execution testing upstream health). -
Event Lifecycle: Emits
TRIP,PROBE,CLOSE, andREJECTevents throughCircuitBreaker.EventListener. -
Fault Kinds: Exceptions are classified into
TIMEOUT,PROVIDER,CAPACITY,TRANSIENT,DATA, orCONTROLviaFaultClassifier, driving degradation decisions.
The pathway execution subsystem provides zero-dependency instrumentation in spector-commons via PathwayObservationHook and exports 8 standardized Micrometer meters in spector-metrics:
| Meter Name | Type | Tags | Description |
|---|---|---|---|
spector.pathway.conduct |
Timer |
pathway, finish
|
End-to-end execution duration of the full pathway cycle |
spector.pathway.relay |
Timer |
pathway, relay, status
|
Latency of an individual relay (status: success, degraded, bypassed, failed, short_circuited) |
spector.pathway.degraded |
Counter |
pathway, relay, kind
|
Count of degraded relay executions classified by FaultKind
|
spector.pathway.circuit |
Counter |
circuit, state_transition
|
Circuit breaker state transitions (trip, probe, close, reject) |
spector.pathway.bulkhead.reject |
Counter | bulkhead |
Count of relay invocations rejected due to exhausted concurrency permits |
spector.pathway.timeout |
Counter |
pathway, relay
|
Count of relay executions aborted due to timeout expiration |
spector.pathway.retry |
Counter |
pathway, relay
|
Count of retry attempts executed after initial transient failure |
spector.pathway.nested |
Timer |
from, to
|
Latency and frequency of nested pathway invocations |
All domain reports (DreamReport, DecideReport, ExpressReport, WanderReport, ReflectReport, RememberResult) encapsulate an immutable ConductionOutcome. Consumers can inspect:
-
finish(): Terminal disposition (COMPLETED,SHORT_CIRCUITED,FAILED). -
degraded()&isDegraded(scope): Whether any stage completed in a degraded fallback state. -
bypassed()&isBypassed(scope): Whether any stages were skipped due to gate predicates or circuit state. -
traces(): Microsecond-precision per-relay execution trace timeline.
sequenceDiagram
participant Client as 👤 Client (CLI/MCP/REST)
participant Pipeline as 🔄 IngestionPipeline
participant Embed as 🧠 ParallelEmbeddingPipeline
participant Target as 💾 IngestionTarget
participant Store as 💾 Storage (mmap)
Client->>Pipeline: pipeline.ingest(file)
Pipeline->>Embed: generateEmbeddings()
Embed-->>Pipeline: dense + sparse vectors
Pipeline->>Target: target.store(chunk)
Target->>Store: write to off-heap MemorySegment
loop Each chunk
Pipeline->>Pipeline: TextChunker.chunk(content)
Pipeline->>Embed: embed(chunkTexts) via virtual threads
Embed-->>Pipeline: List<vector>
Pipeline->>Target: target.ingest(id, text, vector)
Target->>Store: VectorStore + VectorIndex
end
Store-->>Client: ✅ Indexed
-
Client calls
pipeline.ingest()— unified across CLI, MCP, and application code - IngestionPipeline handles chunking (from config) and parallel embedding
-
IngestionTarget receives pre-embedded chunks — storing directly in
SpectorMemory - Downstream storage writes to off-heap memory and indexes with HNSW/BM25
Tip
FileDiscoveryService can be used independently for file discovery without any engine dependency.
sequenceDiagram
participant Client as 👤 Client
participant Memory as 🧠 SpectorMemory
participant Pipeline as ⚙️ RecallPipeline
participant BM25 as 📝 BM25 Search
participant HNSW as 🧠 Dense HNSW
participant Sparse as 📈 Sparse (SPLADE)
participant RRF as 🧬 RRF Fusion
participant Rerank as 🚀 ColBERT Rerank
participant Graph as 🔗 Graph Expansion
Client->>Memory: recall(query, options)
Memory->>Pipeline: execute(query, options)
par Parallel first-stage retrieval on virtual threads
Pipeline->>BM25: exact term matching
Pipeline->>HNSW: dense semantic search
Pipeline->>Sparse: learned sparse search
end
BM25 & HNSW & Sparse->>RRF: Rank merge
RRF->>Rerank: Token-level late interaction MaxSim
Rerank->>Graph: Multi-hop graph expansion & gating
Graph-->>Client: ✨ Final cognitive memories
-
Recall Pathway receives options (
TextSearchMode,RecallMode, etc.) - Dense Vector, BM25, and Sparse (SPLADE) searches run in parallel on virtual threads
- RRF Fusion merges the ranked lists using reciprocal rank scores
- ColBERT v2 Reranking scores the top candidates using SIMD MaxSim operations
- Graph Expansion traverses Hebbian/Entity/Temporal edges for neighbor expansion
sequenceDiagram
participant Agent as 🤖 AI Agent (Claude/Cursor)
participant MCP as 📡 MCP Transport (stdio / Streamable HTTP)
participant Handler as 🔧 McpToolHandler
participant Memory as 🧠 SpectorMemory
participant SIMD as 🔬 SIMD Kernels
Agent->>MCP: tools/call {"name": "memory_recall", "arguments": {"query": "..."}}
MCP->>Handler: MemoryRecallTool.execute(args)
Handler->>Memory: recall(query, options)
Memory->>SIMD: 6-phase scoring + Panama off-heap reads
SIMD-->>Memory: CognitiveResult[] (~130µs)
Memory-->>Handler: List<CognitiveResult>
Handler-->>MCP: CallToolResult
MCP-->>Agent: JSON-RPC response with recalled memories
The MCP path operates directly against SpectorMemory. The MCP server wraps tool handler calls with JSON-RPC transport. There is zero network overhead because everything runs in the same JVM process.
Tip
For full MCP architecture details and tool schemas, see the dedicated MCP Integration page.
Spector is designed from the ground up for Java virtual threads:
Tip
No synchronized blocks anywhere in the codebase. All coordination uses ReentrantLock to avoid virtual thread pinning.
| Operation | Threading Strategy |
|---|---|
| REST request handling | One virtual thread per request |
| Hybrid search | Parallel BM25 + HNSW via StructuredTaskScope
|
| Bulk ingest | Virtual thread per document |
| Embedding generation | Batched across virtual threads |
| HNSW construction (>10K) | Virtual threads per core for parallel insertion |
| Distributed fan-out | Virtual thread per shard query |
At 50K docs with hybrid search (384-dim, production-realistic):
| Virtual Threads | Throughput | Scaling |
|---|---|---|
| 1 | 3,739 ops/s | 1.0× |
| 4 | 10,317 ops/s | 2.8× |
| 8 | 11,812 ops/s | 3.2× |
| 16 | 14,022 ops/s | 3.7× |
Note
Scaling depends on vector dimensions and workload type. 384-dim shows ~3.7× at 16 threads due to higher per-query memory bandwidth. Individual HNSW queries are inherently sequential (graph traversal data dependencies) — scaling comes from concurrent queries sharing CPU cores.
All vector data lives off-heap using the Panama Foreign Function & Memory API:
graph TB
subgraph "☕ JVM Heap (minimal)"
HG["HNSW Graph<br/>(adjacency lists)"]
BM["BM25 Index<br/>(inverted index)"]
ES["Engine State<br/>(config, lifecycle)"]
end
subgraph "🧊 Off-Heap (Panama MemorySegment)"
VS["Vector Store<br/>Contiguous float32, SIMD-aligned<br/>Zero-copy reads, no GC pressure"]
QS["Quantized Store<br/>INT8 or PQ codes"]
GM["GPU Device Memory<br/>CUDA via FFM"]
end
HG -.-> VS
BM -.-> VS
ES -.-> QS
ES -.-> GM
Benefits:
-
✅ Zero GC pressure — Vectors never touch the garbage collector
-
✅ Instant startup — Memory-mapped files load via
mmapsyscall, no deserialization -
✅ SIMD-friendly layout — Contiguous float32 arrays ready for Vector API operations
-
✅ Explicit lifecycle —
Arena-scoped memory with deterministic cleanup -
✅ Memory efficiency — Store billions of vectors limited only by disk/address space
| Store | Location | Use Case |
|---|---|---|
InMemoryVectorStore |
Off-heap (Arena) | Development, small datasets |
MmapVectorStore |
Memory-mapped file | Production, persistence |
QuantizedVectorStore |
Off-heap (INT8) | Memory-constrained deployments |
IvfPqStore |
Off-heap (PQ codes) | Billion-scale (32× compression) |
graph TD
subgraph "SpectorNode - Armeria Server, single port"
CORS["CorsService decorator"]
Auth["API Key decorator"]
COMPRESS["EncodingService - gzip/brotli"]
subgraph "ApiModule Registration"
SE["🔍 SearchEndpoint"]
IE["📥 IngestEndpoint"]
RE["🤖 RagEndpoint"]
DE["🗑️ DocumentEndpoint"]
STE["📊 StatusEndpoint"]
ESE["📡 EventStreamEndpoint"]
end
gRPC["gRPC Service<br/>inter-node fan-out"]
HEALTH["💚 /health"]
PROM["📊 /metrics"]
end
subgraph "REST Controller Layer"
MC["MemoryController<br/>/api/v1/memory/*"]
SC["SystemController<br/>/api/v1/system/*"]
HC["HealthController<br/>/api/v1/system/*"]
end
subgraph "Service Layer"
MS["MemoryService"]
end
subgraph "Core Engine"
SM["SpectorMemory"]
end
MC & SC & HC --> MS
MS --> SM
Every request runs on its own virtual thread. The Armeria server handles HTTP REST, gRPC, and SSE events on a single port. API endpoints are registered via ApiModule components, enabling straightforward API versioning (/api/v1, /api/v2).
The /api/v1/search/stream endpoint uses Server-Sent Events to emit results progressively. The /api/v1/events endpoint provides a live event stream where clients can subscribe to search, ingest, cluster, MCP, and engine events with optional category filtering.
-
Core Concepts — Algorithms and data structures in detail
-
Distributed Mode — Multi-node clustering architecture
-
GPU Acceleration — CUDA kernel integration via Panama
-
Performance Tuning — Optimizing for your workload
- Home
-
Getting Started
- Quick Start
- Installation
- Developer Guide
- JDK API Status
- MCP Server
- Java SDK
- Java API Reference
- Python SDK
- TypeScript SDK
- Spring AI Integration
- CLI Reference
- REST API
- API Playground
- Error Codes
- Configuration
- Deployment
-
Cognitive Memory
- Overview
- Getting Started
- Use Cases
- API Reference
- Concepts
- Pathways
- Scoring features
- Profiles
- Experimental
- Internals
- Design ancestry
-
Memory Kernel
- Overview
- Bundle Architecture
- Memory Shapes
- Binary Layouts & Tags
- WAL & Durability
-
Region Reference
- Overview & Index
- Partition Regions
-
Runtime Regions
- Working Memory
- Co-Activation Matrix
- Index MIDX
- Index IDPL
- Hebbian Graph
- Temporal Chains
- Temporal Facts
- Entity Directory
- Entity Names Pool
- HyperEntity Graph
- Entity Types Registry
- Relation Types Registry
- BM25 Lexical Index
- Checkpoint
- Insula (Somatic Self-Model)
- Continuity
- Provenance
- SPLADE Sparse Index
- Entity Reverse Index
- Identity Regions
- Synapse & Cortex
-
Architecture
- System Overview
- Core Concepts
- Ingestion Pipeline
- MCP Integration
- Distributed Mode
- Event Notifications
- Namespace Sharding
- Single-Namespace Scale & Capacity Limits
- Scale Benchmark Empirical Results
- Writer Quiesce Pause Empirical Results
- Kill-Owner Failover Empirical Results
- Salience & Importance Architecture
- GPU Acceleration
- Performance Tuning
- Test Framework & LLM Judge
- Chat & Visual Test Infrastructure
- Security & Data
-
Architecture Decision Records (ADRs)
- Overview
- Template
- Master Catalog (0001-0085)
-
Memory Kernel & Storage Formats
- ADR-0001: Graph Compression Strategy for Entity Graph
- ADR-0002: Multi-Partition Recall Fan-Out & Frozen Reten...
- ADR-0003: Completing Hypergraph Entity-Graph Graduation
- ADR-0004: Mmap Bundle Architecture & File Descriptor Sc...
- ADR-0005: spector-memory Technical Debt Hardening
- ADR-0042: Graph Recall Architecture and Cognitive Trave...
- ADR-0043: Single-VMA Bundle Layout Specification
- ADR-0044: Memory Kernel Isolation, Composition, and Layout
- ADR-0045: Spector Memory Import & Export Pipeline
- ADR-0046: Single Engram, Four Stores Storage Architecture
- ADR-0047: Episodic Memory and Engram Model Hierarchy
- ADR-0057: Remediation of Hardcoded Memory Offsets and Alignment Constants
- ADR-0062: Spector Memory Organization — Three-Plane Architecture
- ADR-0082: Index Plane Lifecycle, Derived Views, and Reconciliation
-
Active Inference Self-Model Engine (AISME)
- ADR-0006: Episodic Conversation Architecture
- ADR-0007: ReflectPathway — Biological Sleep Consolidation
- ADR-0008: Cognitive Substrate Evolution (TANGLE, GPM, M...
- ADR-0009: AISME Phase 1 — Homeostatic Affective Core
- ADR-0010: AISME Phase 2 — Free-Energy Guided Recall
- ADR-0011: AISME Phase 3 — Modern Hopfield Associative M...
- ADR-0012: AISME Phase 4 — Neural Manifold Distance (NMD)
- ADR-0013: AISME Phase 5 — Predictive Coding Narrative Self
- ADR-0014: AISME Phase 6 — Consciousness Continuity Metr...
- ADR-0015: AISME Phase 7 — Synaptic Relay Wiring & Pathw...
- ADR-0016: AISME Phase 8 — Closed-Loop Epistemic Learning
- ADR-0017: AISME Phase 9 — Generative Counterfactuals & ...
- ADR-0018: AISME Phase 10 — WanderPathway & Kernel Conti...
- ADR-0019: AISME Phase 11 — Expected Free Energy Policy ...
- ADR-0020: AISME Phase 12 — Continuous Self-Dynamics
- ADR-0023: AISME Complete Loop Closure & CognitiveVector...
- ADR-0024: Polymorphic SoulContext Hierarchy in AISME
- ADR-0027: Soul-Conditioned & Salience-Modulated Persona...
- ADR-0048: Cross-Capture Graph & CoActivation Kernel
- ADR-0049: Identity Trajectory Lyapunov Stability
- ADR-0050: Event Density Gating and Dynamic Epistemic Co...
- ADR-0051: Bayesian Online Change-Point Episode Segmenta...
- ADR-0052: Differential Privacy and Edge Anonymization
- ADR-0053: Multimodal Composite Importance Scoring
- ADR-0054: Lifespan-Adaptive Forgetting & Retention Kernel
- ADR-0055: LSR & RFF Dense Associative Memory Engineerin...
- ADR-0056: Log-Sum-ReLU (LSR) & Random Fourier Features ...
- ADR-0058: Linguistic & Vocal Prosody Expression Engine
- ADR-0063: Spacetime Vector Search and Synaptic Relay Architecture
- ADR-0064: Spacetime Simulation on Wander, Dream, and Express Pathways
- ADR-0071: Remember Cognitive Pathway Architecture
- ADR-0072: Six-Phase Fused Cognitive Scoring Pipeline
- ADR-0073: Recall Cognitive Pathway and Multi-Phase Retrieval Architecture
- ADR-0074: Reflect Cognitive Pathway and Sleep Consolidation Architecture
- ADR-0078: Salience Network and Thalamic Cognitive Profiles Architecture
-
Platform, Synapse & Clustering
- ADR-0021: Nucleus Symmetric Hardware Abstraction Layer ...
- ADR-0022: Embodied Kinesics & Phenomenological MCP Engine
- ADR-0025: Declarative MCP Tool Definitions via JSON Sch...
- ADR-0026: Dual-Plane Concurrency & Async Queue Backpres...
- ADR-0028: Dual-Plane Memory Audit Architecture (Separat...
- ADR-0029: Episodic→Semantic Lineage Provenance Region
- ADR-0030: Unified Engram Encoding Header Architecture
- ADR-0031: Unified Configuration Architecture & Bypass E...
- ADR-0032: Persona Enactment — Soul as Policy over Memory
- ADR-0033: Decoupling Cognitive & Mathematical Kernels t...
- ADR-0034: Cell Topology, Namespace Ownership, and HA Cl...
- ADR-0035: Cognitive Pathway Framework Rearchitecture
- ADR-0036: Pathway Error Handling, Isolation, and Circui...
- ADR-0037: Ingestion Boundary and Sensory Relocation
- ADR-0038: SIMD-Accelerated BM25 Lexical Scoring Optimiz...
- ADR-0039: Robust Unified Rate Limiting Architecture
- ADR-0040: Universal Apache Camel Messaging Channels
- ADR-0041: Unified Connector Architecture for Ingestion
- ADR-0059: Java 27 Upgrade Strategy and Value Class Migration
- ADR-0060: Cognitive Continuity Layer and Decoded Mind Streams
- ADR-0061: In-Memory Multi-Tenant Quartz Scheduler
- ADR-0065: Client SDK Architecture, OpenAPI, and MCP Integration
- ADR-0066: Engine & CLI Stabilization — Issue #727 Hardening
- ADR-0067: Cell-Based High Availability and Namespace-Sticky Sharding
- ADR-0068: Phileas PII Redaction Engine for Spector Synapse
- ADR-0069: Synapse-Owned Tool Access Policy
- ADR-0070: Unified Error Taxonomy and Exception Handling Architecture
- ADR-0075: Extensible LLM and Multimodal Embedding Provider SPI
- ADR-0076: Zero-Dependency Pluggable Cache Abstraction
- ADR-0077: Model B Asynchronous Task Queue and Concurrency
- ADR-0079: Asynchronous Memory Event and Telemetry Notification Bus
- ADR-0080: Observed Memory and Pathway Metrics Telemetry Architecture
- ADR-0081: Dedicated Reactive Ingress and In-Process Path Router
- ADR-0083: Namespace-Isolated Memory Analytics & Telemetry
- ADR-0084: Dual-Plane Conversation Persistence
- ADR-0085: Dynamic Synapse Configuration Overrides and Runtime Propagation
-
Modules Registry
- Overview
- Foundation Layer (/nucleus)
- Cognitive Layer (/memory)
- Gateway Layer (/synapse)
- Benchmarks & UI
- Deep Dives
-
Community
- Governance
- Contributing
- FAQ
- Glossary
- Roadmap
- 🔬 Labs
- Third-Party Legal