Repository navigation
FAQ
title: "Spector FAQ — Frequently Asked Questions" description: "Comprehensive answers to questions about Spector Cognitive Memory: MF-001 recall algebra, Panama FFM zero-GC kernel, 4 tiers, 37 MCP tools, and benchmarks."
Everything you need to know about Spector's architecture, cognitive memory model, Panama FFM kernel, MCP integration, and benchmarks.
Traditional vector databases (Pinecone, Milvus, Qdrant, Weaviate, pgvector) are storage engines optimized for top-$K$ nearest-neighbor vector similarity. They return vectors that are geometrically close to a prompt embedding at write time.
Spector is a formal cognitive memory engine implementing the Memory Fundamentals Specification (MF-001). As formalized in MF-001:
A database returns what was written. A memory engine reconstructs what is reachable from a cue at this moment — under decay, association, and tier physics — without losing a live trace because the first index was the wrong one.
Instead of an undifferentiated flat vector store, Spector organizes memory across four cognitive tiers (Working, Episodic, Semantic, Procedural) and applies a fused 6-phase recall algebra combining vector similarity, power-law temporal decay, affective valence, habituation, and associative graph spreading directly in a single cache-friendly SIMD hot-loop.
When an agent queries a standard vector database with post-filtering:
- The database retrieves the top-$K$ (e.g. 100) nearest vectors by cosine distance.
- The application layer filters or re-ranks candidates by recency, importance, or user tags.
danger: The Truncation Trap If an agent asks "What is the user's architectural guideline for retries?", a vital procedural rule stored 3 months ago might have a cosine similarity of
0.78, while 100 casual discussion snippets from yesterday have similarity0.81. In a standard vector DB, the 3-month-old rule is permanently truncated at step 1 before application filtering ever executes.
Spector eliminates the truncation trap by evaluating semantic similarity, temporal recency, importance, and Bloom-filtered tags simultaneously in a single off-heap pass. The procedural rule is boosted by its tier and importance during the scan, guaranteeing it surfaces in the final recall candidates.
Memory wrappers are orchestration layers built on top of external databases (such as PostgreSQL, Redis, or cloud vector endpoints). While they add cognitive heuristics, they suffer from three structural bottlenecks:
| Dimension | ⚡ Spector Memory Kernel | AI Memory Wrappers |
|---|---|---|
| Execution Substrate | In-process off-heap memory (MemorySegment) |
Network-bound REST / RPC calls |
| Recall Latency (p50) | 1.01 ms (in-process fused scan) | 15–80 ms (multi-query network roundtrips) |
| GC Overhead | Zero-GC (Java 25 Panama FFM) | High heap allocation & JSON serialization |
| Candidate Evaluation | Single-pass SIMD fused scoring | Multi-stage fetch, deserialize, filter, re-rank |
| Graph Association | Real-time Hebbian co-activation & STDP | External graph DB queries (Neo4j / NetworkX) |
MF-001 is an open specification establishing the mathematical and operational foundations of artificial cognitive memory. It formalizes:
- Trace Durability: Distinction between volatile working buffers and consolidated semantic structures.
- Recall Algebra: Unified scoring functions that bind spatial distance, power-law retention decay, and affective valence.
- Associative Spreading: Hebbian co-activation dynamics where recalled engrams prime adjacent concept nodes.
Spector is the reference implementation of the MF-001 standard.
flowchart LR
WM["🧪 Working Memory<br/>(Working)<br/>Volatile Turn Buffer"]
EM["📝 Episodic Memory<br/>(Episodic)<br/>Timestamped Log"]
SE["🧬 Semantic Memory<br/>(Semantic)<br/>Permanent Facts"]
PR["⚙️ Procedural Memory<br/>(Procedural)<br/>Rules & Constraints"]
WM -.->|"Sleep Consolidation"| EM
EM -->|"Dreaming & Pruning"| SE
SE -.->|"Policy Extraction"| PR
classDef t fill:#1e293b,stroke:#6366f1,stroke-width:1.5px,color:#f8fafc
class WM,EM,SE,PR t
Different types of knowledge operate on radically different timescales, access frequencies, and eviction semantics:
- Working Memory: Sub-microsecond circular buffer for active prompt context and turn state. Automatically evicts the oldest items on capacity overflow.
- Episodic Memory: Partitioned by date, backed by memory-mapped files. Preserves autobiographical agent interactions, tool calls, and user queries with chronological fidelity.
- Semantic Memory: Distilled, enduring world knowledge, codebase architecture, and user preferences. Compacted and consolidated across sessions.
- Procedural Memory: High-salience behavioral policies, prompt constraints, and tool protocols that must never be accidentally evicted by casual dialogue.
- Reflect (supported): cluster, promote, prune, rebuild.
- Dream (experimental): stochastic association; see Experimental.
In high-concurrency AI systems, JVM Garbage Collection pauses can degrade recall latency from 1ms to hundreds of milliseconds.
Spector achieves Zero-GC execution using Java 25's Foreign Function & Memory (FFM) API (Project Panama):
- Memory engrams, 128-bit Bloom filters, valence headers, and vector payloads reside strictly off-heap in native memory segments (
java.lang.foreign.MemorySegment). - Search kernels read raw memory addresses directly via hardware SIMD instructions without allocating intermediate Java heap objects.
- High-throughput scans operate with zero garbage collector invocation, maintaining flat
$p99$ latency profiles.
Spector stores memories in V4 Bundles (.seg files) using a page-aligned native binary format:
-
Zero-Copy Loading: Partitions are mapped into the process virtual address space via
mmap. The engine boots in under 50 milliseconds, regardless of whether the index contains 10,000 or 10,000,000 records. - Write-Ahead Log (WAL): Ingestion appends to a memory-mapped journal, ensuring crash consistency and immediate durability.
- Atomic Compaction: Live background compaction rebuilds sparse partitions without blocking ongoing read queries.
When client.memory.recall(...) executes, the CognitiveScorer evaluates off-heap records through six progressive gating filters:
flowchart TD
P1["Phase 1: Tombstone Bit Test (~1 CPU cycle)"] -->|"Live"| P2["Phase 2: 128-bit Bloom Filter Tag Gating (~1 cycle)"]
P2 -->|"Match"| P3["Phase 3: Valence & Affective Range Check (~2 cycles)"]
P3 -->|"In Range"| P4["Phase 4: Precomputed Decay Bucket Lookup (~5 cycles)"]
P4 -->|"Salient"| P5["Phase 5: SIMD L2/Cosine Vector Distance (~200 cycles)"]
P5 --> P6["Phase 6: Fused Cognitive Score & Top-K Heap Insert (~7 cycles)"]
classDef p fill:#0f172a,stroke:#3b82f6,stroke-width:1.5px,color:#f8fafc
class P1,P2,P3,P4,P5,P6 p
- Phase 1 (Tombstone): 1-cycle bit check eliminates deleted or suppressed engrams.
- Phase 2 (Bloom Tag Match): 128-bit Bloom filter test filters out non-matching categorical tags before touching vector data.
- Phase 3 (Valence Filter): Filters memories outside desired affective boundaries (e.g., recalling only error states or only positive feedback).
- Phase 4 (Decay Pre-screen): Array lookup against 12-bucket precomputed power-law decay table; drops traces too weak to enter top-$K$.
- Phase 5 (SIMD Vector Math): Hardware-accelerated distance calculation using AVX-512/AVX2/NEON instructions.
- Phase 6 (Score Fusion): Fuses spatial similarity, adjusted decay, valence, and importance into the final ranking score.
No. Spector is engineered to deliver sub-millisecond search on commodity CPUs:
- On modern x86_64 CPUs, the Java Vector API compiles to AVX2 (256-bit) and AVX-512 (512-bit) vector instructions.
- On Apple Silicon and ARM servers, it compiles to ARM NEON (128-bit) vector operations.
- A GPU (NVIDIA CUDA) is completely optional and primarily beneficial for high-concurrency batch ingestion (
$>32$ concurrent streams).
The Spector MCP server exposes 37 specialized cognitive tools organized into functional clusters:
-
Core Memory Operations:
memory_remember,memory_recall,memory_forget,memory_reinforce,memory_suppress. -
Cognitive Introspection:
memory_introspect,memory_why_not,memory_salience,memory_fact_history,memory_status. -
Associative Graphs:
memory_graph_recall,memory_multi_evidence_recall,vector_search. -
Multi-Tenancy & Governance:
namespace_create,namespace_switch,namespace_grant,namespace_revoke,namespace_list. -
Identity & Affect:
persona_enact,update_agent_soul,memory_persona_context.
Add Spector to your configuration file:
```json title="claude_desktop_config.json"
{
"mcpServers": {
"spector": {
"command": "npx",
"args": ["-y", "@spectrayan/spector", "mcp"]
}
}
}
```
```json title=".cursor/mcp.json"
{
"mcpServers": {
"spector": {
"command": "npx",
"args": ["-y", "@spectrayan/spector", "mcp"]
}
}
}
```
The npx -y @spectrayan/spector mcp launcher automatically connects to your local running Spector daemon (:7070), or launches an embedded in-process memory kernel if no server is running.
Every memory engram belongs to an isolated Namespace:
-
Strict Cryptographic Isolation: Queries in
namespace_Acannot see or scan records innamespace_Bunless explicit cross-namespace grants exist. -
Granular Permissions: Namespaces support read/write delegation via
namespace_grantwith specific access modes (READ,WRITE,ADMIN). - Audit Trails: Ingestion and recall operations log cryptographically verifiable provenance traces (stating author, timestamp, and source).
-
API Key: Configure via
SPECTOR_API_KEYenvironment variable. Clients passX-API-Key: <token>. - Bearer Tokens: Standard JWT / Bearer authentication support for enterprise single-sign-on (SSO) gateways.
- Local Dev Mode: When no key is set, the server accepts local connections for friction-free developer onboarding.
Spector has been evaluated across the industry-standard AI long-term memory benchmarks:
-
:material-bullseye-arrow: LoCoMo Benchmark
Long-Context Mobile & Agentic Memory evaluation.
85% Precision State-of-the-Art
-
:material-timer-sand-complete: LongMemEval Benchmark
Multi-session recall across extended temporal horizons.
94% Recall Accuracy Top Performance
-
:material-brain: MindSpan Benchmark
Complex multi-hop reasoning and associative recall.
100% Accuracy Flawless Resolution
Traditional vector databases drop to 40–60% accuracy on multi-session benchmarks because conversation history accumulates noise that dilutes raw cosine similarity. Spector's fused 6-phase scoring applies temporal reconsolidation (memories recalled in prior sessions gain durability) and associative graph traversal, ensuring that relevant facts are surfaced even when phrasing changes across conversations.
OpenJDK 25 or later is required for running the Spector server or embedded core JAR. This is because Spector leverages:
- Java Vector API (
jdk.incubator.vector) for SIMD acceleration. - Foreign Function & Memory API (
java.lang.foreign) for off-heap Zero-GC storage.
Client SDKs (Python, TypeScript, Node.js) require no Java runtime on client machines.
java \
--add-modules jdk.incubator.vector \
--enable-native-access=ALL-UNNAMED \
-XX:+UseZGC -XX:+ZGenerational \
-Xms4g -Xmx4g \
-jar spector.jar-
--add-modules jdk.incubator.vector: Enables hardware SIMD intrinsics. -
--enable-native-access=ALL-UNNAMED: Permits zero-overhead off-heap FFM access. -
-XX:+UseZGC -XX:+ZGenerational: Generational ZGC guarantees sub-millisecond GC pause times for any minor heap allocations.
-
Python SDK:
pip install spector-client(Documentation) -
TypeScript / Node.js SDK:
npm install @spectrayan/spector-client(Documentation) -
Java Client SDK:
com.spectrayan:spector-client(Documentation) -
Spring AI: First-class
VectorStoreintegration (Documentation)
- 💬 Join the conversation on GitHub Discussions
- 🐛 Report an issue on GitHub Issues
- 📖 Explore the Cognitive Memory Overview
- Home
-
Getting Started
- Quick Start
- Installation
- Developer Guide
- JDK API Status
- MCP Server
- Java SDK
- Java API Reference
- Python SDK
- TypeScript SDK
- Spring AI Integration
- CLI Reference
- REST API
- API Playground
- Error Codes
- Configuration
- Deployment
-
Cognitive Memory
- Overview
- Getting Started
- Use Cases
- API Reference
- Concepts
- Pathways
- Scoring features
- Profiles
- Experimental
- Internals
- Design ancestry
-
Memory Kernel
- Overview
- Bundle Architecture
- Memory Shapes
- Binary Layouts & Tags
- WAL & Durability
-
Region Reference
- Overview & Index
- Partition Regions
-
Runtime Regions
- Working Memory
- Co-Activation Matrix
- Index MIDX
- Index IDPL
- Hebbian Graph
- Temporal Chains
- Temporal Facts
- Entity Directory
- Entity Names Pool
- HyperEntity Graph
- Entity Types Registry
- Relation Types Registry
- BM25 Lexical Index
- Checkpoint
- Insula (Somatic Self-Model)
- Continuity
- Provenance
- SPLADE Sparse Index
- Entity Reverse Index
- Identity Regions
- Synapse & Cortex
-
Architecture
- System Overview
- Core Concepts
- Ingestion Pipeline
- MCP Integration
- Distributed Mode
- Event Notifications
- Namespace Sharding
- Single-Namespace Scale & Capacity Limits
- Scale Benchmark Empirical Results
- Writer Quiesce Pause Empirical Results
- Kill-Owner Failover Empirical Results
- Salience & Importance Architecture
- GPU Acceleration
- Performance Tuning
- Test Framework & LLM Judge
- Chat & Visual Test Infrastructure
- Security & Data
-
Architecture Decision Records (ADRs)
- Overview
- Template
- Master Catalog (0001-0085)
-
Memory Kernel & Storage Formats
- ADR-0001: Graph Compression Strategy for Entity Graph
- ADR-0002: Multi-Partition Recall Fan-Out & Frozen Reten...
- ADR-0003: Completing Hypergraph Entity-Graph Graduation
- ADR-0004: Mmap Bundle Architecture & File Descriptor Sc...
- ADR-0005: spector-memory Technical Debt Hardening
- ADR-0042: Graph Recall Architecture and Cognitive Trave...
- ADR-0043: Single-VMA Bundle Layout Specification
- ADR-0044: Memory Kernel Isolation, Composition, and Layout
- ADR-0045: Spector Memory Import & Export Pipeline
- ADR-0046: Single Engram, Four Stores Storage Architecture
- ADR-0047: Episodic Memory and Engram Model Hierarchy
- ADR-0057: Remediation of Hardcoded Memory Offsets and Alignment Constants
- ADR-0062: Spector Memory Organization — Three-Plane Architecture
- ADR-0082: Index Plane Lifecycle, Derived Views, and Reconciliation
-
Active Inference Self-Model Engine (AISME)
- ADR-0006: Episodic Conversation Architecture
- ADR-0007: ReflectPathway — Biological Sleep Consolidation
- ADR-0008: Cognitive Substrate Evolution (TANGLE, GPM, M...
- ADR-0009: AISME Phase 1 — Homeostatic Affective Core
- ADR-0010: AISME Phase 2 — Free-Energy Guided Recall
- ADR-0011: AISME Phase 3 — Modern Hopfield Associative M...
- ADR-0012: AISME Phase 4 — Neural Manifold Distance (NMD)
- ADR-0013: AISME Phase 5 — Predictive Coding Narrative Self
- ADR-0014: AISME Phase 6 — Consciousness Continuity Metr...
- ADR-0015: AISME Phase 7 — Synaptic Relay Wiring & Pathw...
- ADR-0016: AISME Phase 8 — Closed-Loop Epistemic Learning
- ADR-0017: AISME Phase 9 — Generative Counterfactuals & ...
- ADR-0018: AISME Phase 10 — WanderPathway & Kernel Conti...
- ADR-0019: AISME Phase 11 — Expected Free Energy Policy ...
- ADR-0020: AISME Phase 12 — Continuous Self-Dynamics
- ADR-0023: AISME Complete Loop Closure & CognitiveVector...
- ADR-0024: Polymorphic SoulContext Hierarchy in AISME
- ADR-0027: Soul-Conditioned & Salience-Modulated Persona...
- ADR-0048: Cross-Capture Graph & CoActivation Kernel
- ADR-0049: Identity Trajectory Lyapunov Stability
- ADR-0050: Event Density Gating and Dynamic Epistemic Co...
- ADR-0051: Bayesian Online Change-Point Episode Segmenta...
- ADR-0052: Differential Privacy and Edge Anonymization
- ADR-0053: Multimodal Composite Importance Scoring
- ADR-0054: Lifespan-Adaptive Forgetting & Retention Kernel
- ADR-0055: LSR & RFF Dense Associative Memory Engineerin...
- ADR-0056: Log-Sum-ReLU (LSR) & Random Fourier Features ...
- ADR-0058: Linguistic & Vocal Prosody Expression Engine
- ADR-0063: Spacetime Vector Search and Synaptic Relay Architecture
- ADR-0064: Spacetime Simulation on Wander, Dream, and Express Pathways
- ADR-0071: Remember Cognitive Pathway Architecture
- ADR-0072: Six-Phase Fused Cognitive Scoring Pipeline
- ADR-0073: Recall Cognitive Pathway and Multi-Phase Retrieval Architecture
- ADR-0074: Reflect Cognitive Pathway and Sleep Consolidation Architecture
- ADR-0078: Salience Network and Thalamic Cognitive Profiles Architecture
-
Platform, Synapse & Clustering
- ADR-0021: Nucleus Symmetric Hardware Abstraction Layer ...
- ADR-0022: Embodied Kinesics & Phenomenological MCP Engine
- ADR-0025: Declarative MCP Tool Definitions via JSON Sch...
- ADR-0026: Dual-Plane Concurrency & Async Queue Backpres...
- ADR-0028: Dual-Plane Memory Audit Architecture (Separat...
- ADR-0029: Episodic→Semantic Lineage Provenance Region
- ADR-0030: Unified Engram Encoding Header Architecture
- ADR-0031: Unified Configuration Architecture & Bypass E...
- ADR-0032: Persona Enactment — Soul as Policy over Memory
- ADR-0033: Decoupling Cognitive & Mathematical Kernels t...
- ADR-0034: Cell Topology, Namespace Ownership, and HA Cl...
- ADR-0035: Cognitive Pathway Framework Rearchitecture
- ADR-0036: Pathway Error Handling, Isolation, and Circui...
- ADR-0037: Ingestion Boundary and Sensory Relocation
- ADR-0038: SIMD-Accelerated BM25 Lexical Scoring Optimiz...
- ADR-0039: Robust Unified Rate Limiting Architecture
- ADR-0040: Universal Apache Camel Messaging Channels
- ADR-0041: Unified Connector Architecture for Ingestion
- ADR-0059: Java 27 Upgrade Strategy and Value Class Migration
- ADR-0060: Cognitive Continuity Layer and Decoded Mind Streams
- ADR-0061: In-Memory Multi-Tenant Quartz Scheduler
- ADR-0065: Client SDK Architecture, OpenAPI, and MCP Integration
- ADR-0066: Engine & CLI Stabilization — Issue #727 Hardening
- ADR-0067: Cell-Based High Availability and Namespace-Sticky Sharding
- ADR-0068: Phileas PII Redaction Engine for Spector Synapse
- ADR-0069: Synapse-Owned Tool Access Policy
- ADR-0070: Unified Error Taxonomy and Exception Handling Architecture
- ADR-0075: Extensible LLM and Multimodal Embedding Provider SPI
- ADR-0076: Zero-Dependency Pluggable Cache Abstraction
- ADR-0077: Model B Asynchronous Task Queue and Concurrency
- ADR-0079: Asynchronous Memory Event and Telemetry Notification Bus
- ADR-0080: Observed Memory and Pathway Metrics Telemetry Architecture
- ADR-0081: Dedicated Reactive Ingress and In-Process Path Router
- ADR-0083: Namespace-Isolated Memory Analytics & Telemetry
- ADR-0084: Dual-Plane Conversation Persistence
- ADR-0085: Dynamic Synapse Configuration Overrides and Runtime Propagation
-
Modules Registry
- Overview
- Foundation Layer (/nucleus)
- Cognitive Layer (/memory)
- Gateway Layer (/synapse)
- Benchmarks & UI
- Deep Dives
-
Community
- Governance
- Contributing
- FAQ
- Glossary
- Roadmap
- 🔬 Labs
- Third-Party Legal