Skip to content

Architecture Overview

github-actions[bot] edited this page Sep 26, 2026 · 19 revisions

🏗️ Architecture Overview

Spector is a SIMD-accelerated AI memory backbone with built-in MCP server, hybrid search, and multi-tier cognitive memory. This page covers the system architecture, data flows, threading model, and memory architecture that make sub-millisecond, agent-native search possible.


System Architecture

graph TB
    subgraph Clients["Client Interfaces & SDKs"]
        claude["🤖 Claude Desktop"]
        cursor["✏️ Cursor / Windsurf"]
        agents["🦾 Autonomous Agents"]
        py["🐍 Python SDK"]
        ts["🔷 TypeScript SDK"]
        sdk["☕ Java Client SDK"]
        spring["🌱 Spring AI"]
        cli["🖥️ spector CLI"]
        rest["🌐 REST / gRPC"]
    end

    subgraph Transport["Synapse Application Layer"]
        mcp["MCP Server<br/><i>stdio · Streamable HTTP · 37+ tools</i>"]
        armeria["Armeria Gateway :7070<br/><i>REST + gRPC + SSE streaming</i>"]
        persona["Persona Enactment<br/><i>Dual-process cognitive appraisal</i>"]
    end

    subgraph Engine["Spector Memory Engine"]
        subgraph Pathways["Cognitive Pathways"]
            rem["Remember Pathway<br/><i>Surprise · Flashbulb · Dedup</i>"]
            rec["Recall Pathway<br/><i>6-Phase SIMD Fused Scoring</i>"]
            ref["Reflect Pathway<br/><i>Sleep Consolidation · Replay</i>"]
            drm["Dream Pathway<br/><i>Counterfactual Simulation</i>"]
            dec["Decide Pathway<br/><i>Active Inference Policy</i>"]
            exp["Express Pathway<br/><i>Prosody · Stylometry · Kinesics</i>"]
            wan["Wander Pathway<br/><i>DMN Associative Synthesis</i>"]
        end

        subgraph Memory["4-Tier Cortex & Graphs"]
            cortex["4-Tier Cortex<br/><i>Working · Episodic · Semantic · Procedural</i>"]
            hebbian["Cognitive Graphs<br/><i>Hebbian · Temporal · HyperEntity</i>"]
            decay["Memory Decay<br/><i>Power-law forgetting · Bjork strength</i>"]
        end

        subgraph Search["Search & Retrieval Stack"]
            hybrid["Hybrid Retrieval<br/><i>BM25 + Dense + SPLADE</i>"]
            hnsw["HNSW Graph Index<br/><i>M=16, ef=200</i>"]
            colbert["ColBERT v2 Reranking<br/><i>Late-interaction MaxSim</i>"]
        end
    end

    subgraph Kernel["⚡ Memory Kernel (spector-kernel — Zero GC)"]
        direction TB
        ns["NamespaceKernel Facade"]
        bundles["V4 Single-VMA Bundles<br/><i>runtime.bundle · partition.bundle · identity.bundle</i>"]
        shapes["8 Sealed Memory Shapes<br/><i>Record · Append · Graph · Chain · Hash · Insula</i>"]
        panama["Panama FFM Storage<br/><i>Shared Arena · MemorySegment · mmap</i>"]
        simd["Hardware SIMD Acceleration<br/><i>Vector API · AVX2 / AVX-512 / NEON</i>"]
        gpu["GPU Acceleration<br/><i>CUDA via Panama FFM</i>"]
    end

    subgraph Observe["Observability & Telemetry"]
        events["TelemetryBus<br/><i>Event streams</i>"]
        metrics["Micrometer<br/><i>Prometheus export</i>"]
        sse["SSE Event Stream<br/><i>Cortex 3D Galaxy</i>"]
    end

    claude & cursor & agents --> mcp
    py & ts & sdk & spring --> armeria
    cli & rest --> armeria
    mcp & armeria & persona --> Pathways

    Pathways --> Memory & Search
    Memory & Search --> ns
    ns --> bundles --> shapes --> panama
    panama --> simd
    gpu -.->|batch compute| simd

    Engine --> events
    events --> metrics & sse

    style Clients fill:#5b6abf,stroke:#e94560,color:#fff
    style Transport fill:#4a6fa5,stroke:#3b82f6,color:#fff
    style Engine fill:#3b82f6,stroke:#7c3aed,color:#fff
    style Kernel fill:#1e293b,stroke:#0f172a,color:#fff
    style Pathways fill:#2563eb,stroke:#1d4ed8,color:#fff
    style Memory fill:#1d4ed8,stroke:#1e40af,color:#fff
    style Search fill:#1e40af,stroke:#1e3a8a,color:#fff
    style Observe fill:#5b6abf,stroke:#7c3aed,color:#fff
Loading

High-Level Data Flow

graph LR
    subgraph Ingest["Ingest"]
        docs["📄 Documents"]
        files["📁 Files"]
        api["🌐 API Data"]
    end

    subgraph Process["Process"]
        chunk["✂️ Chunk"]
        embed["🧬 Embed"]
        quantize["🗜️ Quantize"]
    end

    subgraph Store["Store"]
        vectors["📊 Vector Index<br/><i>HNSW · IVF-PQ</i>"]
        text["📝 Text Index<br/><i>BM25</i>"]
        memory["🧠 Cognitive Store<br/><i>4-tier cortex</i>"]
    end

    subgraph Query["Query"]
        search["🔍 Hybrid Search"]
        recall["💭 Memory Recall"]
        rag["🤖 RAG Pipeline"]
    end

    docs & files & api --> chunk --> embed --> quantize
    quantize --> vectors & text & memory
    vectors & text --> search --> rag
    memory --> recall --> rag

    style Ingest fill:#5b6abf,stroke:#e94560,color:#fff
    style Process fill:#4a6fa5,stroke:#3b82f6,color:#fff
    style Store fill:#3b82f6,stroke:#7c3aed,color:#fff
    style Query fill:#7c3aed,stroke:#e94560,color:#fff
Loading

Deployment Modes

graph LR
    subgraph Embedded["Embedded Mode"]
        lib["SpectorMemory API<br/><i>In-process · zero-network · drop-in JAR</i>"]
    end

    subgraph Standalone["Standalone Mode"]
        jar["java -jar spector.jar<br/><i>Engine + MCP + REST/gRPC + SSE</i>"]
    end

    subgraph Distributed["Distributed Mode"]
        coord["Coordinator<br/><i>Query routing · fan-out</i>"]
        s1["Shard 1"] & s2["Shard 2"] & s3["Shard N"]
        coord --> s1 & s2 & s3
    end

    style Embedded fill:#4a6fa5,stroke:#3b82f6,color:#fff
    style Standalone fill:#3b82f6,stroke:#7c3aed,color:#fff
    style Distributed fill:#7c3aed,stroke:#e94560,color:#fff
Loading

🤖 MCP Architecture — Agent-Native Engine

Spector's MCP server runs in-process — the agent's tool calls go directly into SIMD kernels with zero network hops, zero serialization, and zero GC pressure. This is the architectural advantage over adapters that wrap a database behind an HTTP API.

Tool Registry

graph TB
    subgraph Agents["AI Agents"]
        claude["🤖 Claude Desktop"]
        cursor["✏️ Cursor / Windsurf"]
        cline["🔧 Cline / Aider"]
        custom["🦾 Autonomous Multi-Agents"]
    end

    subgraph MCP["MCP Server — Dual Transport · JSON-RPC 2.0"]
        transport["Transport Layer<br/><i>stdio (stdin/stdout) for CLI agents<br/>Streamable HTTP (/mcp) for remote agents</i>"]
        registry["SpectorToolRegistry<br/><i>37+ tools · dynamic route dispatch</i>"]
        handler["McpToolHandler<br/><i>Base class · thread-safe · virtual threads</i>"]

        subgraph Mem["1. Memory Tier Operations (16 Tools)"]
            m1["memory_remember — Store with importance & tags"]
            m2["memory_recall — Fused SIMD scoring recall"]
            m3["memory_scratchpad — Working-memory scratchpad"]
            m4["memory_reinforce — Outcome feedback (+/-)"]
            m5["memory_forget — Tombstone intentional forgetting"]
            m6["memory_status — Per-tier statistics & health"]
            m7["memory_introspect — Metamemory self-reflection"]
            m8["memory_suppress — Temporary recall suppression"]
            m9["memory_resolve — Mark resolved/unresolved"]
            m10["memory_reminder — Proactive intent triggers"]
            m11["memory_why_not — Explain recall misses"]
            m12["memory_compute_importance — Pre-ingest scoring"]
            m13["memory_inspect — Full cognitive X-ray"]
            m14["memory_export — Bulk JSON memory export"]
            m15["memory_browse — Browse by tag/tier filter"]
            m16["memory_salience — Inspect & tune salience profile"]
        end

        subgraph GraphContext["2. Graph & Multi-Evidence Retrieval (7 Tools)"]
            g1["memory_graph_recall — Spreading activation graph walk"]
            g2["memory_context_pack — Assembled agent prompt pack"]
            g3["memory_fact_history — Temporal chain evolution"]
            g4["memory_persona_context — Soul-aligned contextual injection"]
            g5["memory_multi_evidence_recall — Multi-vector consensus"]
            g6["vector_search — Pure vector cosine similarity"]
            g7["memory_express — Natural language memory synthesis"]
        end

        subgraph NamespaceRBAC["3. Namespace & Multi-Tenancy (9 Tools)"]
            n1["namespace_create — Provision isolated namespace"]
            n2["namespace_list — Enumerate active namespaces"]
            n3["namespace_info — Inspect V4 bundle layout & size"]
            n4["namespace_switch — Set active session namespace"]
            n5["namespace_set_default — Update default namespace"]
            n6["namespace_delete — Safely purge namespace files"]
            n7["namespace_grant — RBAC access delegation"]
            n8["namespace_revoke — Revoke access permissions"]
            n9["namespace_list_grants — Audit security grants"]
        end

        subgraph SoulPolicy["4. Agent Soul & Persona Enactment (5 Tools)"]
            s1["update_agent_soul — Mutate agent persona & dogmas"]
            s2["persona_enact — Dual-process cognitive appraisal"]
            s3["account_introspect — Account & tenant introspection"]
            s4["invoke_connector_route — External data connector dispatch"]
            s5["send_notification — Dispatch proactive agent alerts"]
        end
    end

    subgraph Core["In-Process Engine — Zero Network Overhead"]
        pathways["Cognitive Pathways<br/><i>Remember · Recall · Reflect · Dream · Decide · Express · Wander</i>"]
        kernel["Sealed Memory Kernel<br/><i>V4 Bundles · 8 Shapes · Panama FFM</i>"]
    end

    Agents -->|stdio / HTTP| transport --> registry --> handler
    handler --> Mem & GraphContext & NamespaceRBAC & SoulPolicy
    Mem & GraphContext & NamespaceRBAC & SoulPolicy --> pathways --> kernel

    style Agents fill:#5b6abf,stroke:#e94560,color:#fff
    style MCP fill:#4a6fa5,stroke:#3b82f6,color:#fff
    style Mem fill:#3b82f6,stroke:#2563eb,color:#fff
    style GraphContext fill:#2563eb,stroke:#1d4ed8,color:#fff
    style NamespaceRBAC fill:#1d4ed8,stroke:#1e40af,color:#fff
    style SoulPolicy fill:#1e40af,stroke:#1e3a8a,color:#fff
    style Core fill:#1e293b,stroke:#0f172a,color:#fff
Loading

Agent Interaction Flow

sequenceDiagram
    participant Agent as 🤖 AI Agent
    participant MCP as 📡 MCP Server
    participant Tools as 🔧 ToolRegistry
    participant Memory as 🧠 SpectorMemory
    participant SIMD as 🔬 SIMD (off-heap)

    Note over Agent,SIMD: Single JVM process — no HTTP, no gRPC, no serialization

    Agent->>MCP: tools/call {"name": "memory_remember", ...}
    MCP->>Tools: Route → MemoryRememberTool
    Tools->>Memory: remember(text, tags, importance)
    Memory->>SIMD: Embed → HNSW insert → tier assign
    SIMD-->>Agent: ✅ memoryId + tier (~1ms)

    Agent->>MCP: tools/call {"name": "memory_recall", ...}
    MCP->>Tools: Route → MemoryRecallTool
    Tools->>Memory: recall(query, topK)
    Memory->>SIMD: Fused scoring: sim × importance × decay
    SIMD-->>Agent: 📋 Ranked memories (ultra-fast)

    Agent->>MCP: tools/call {"name": "memory_introspect", ...}
    MCP->>Tools: Route → MemoryIntrospectTool
    Tools->>Runtime: memory().introspect(topic)
    Runtime->>SIMD: Confidence + knowledge-gap analysis over tiers
    SIMD-->>Agent: 🔍 Knowledge report (~0.2ms)
Loading

Performance: In-Process vs. External Adapters

Metric Spector (in-process) Typical MCP adapter
Architecture Engine + MCP in one JVM Python → HTTP → DB → HTTP → agent
Search latency 88µs (SIMD) 5–50ms (network round-trip)
Memory recall Ultra-low latency (fused scoring) 50–200ms (Mem0/Letta/Zep)
Tools 16 (cognitive memory tools) 3–5 basic CRUD
GC pressure Zero (Panama off-heap) Full GC overhead
Deployment java -jar spector.jar Python + pip + DB + config

Tip

For full MCP integration details, tool schemas, and Claude Desktop configuration, see the dedicated MCP Integration page.


📦 Module Diagram

graph LR
    subgraph "🔬 Foundation & Acceleration (nucleus/)"
        core["spector-core<br/><i>Compute SPIs & Quantization</i>"]
        cpu["spector-cpu<br/><i>Java 25 SIMD Kernels</i>"]
        gpu["spector-gpu<br/><i>Panama FFM + CUDA GPU</i>"]
        hdc["spector-hdc<br/><i>Hyperdimensional vectors</i>"]
        index["spector-index<br/><i>HNSW + SpectorIndex + BM25</i>"]
        commons["spector-commons<br/><i>Error codes & concurrency</i>"]
        config["spector-config<br/><i>SpectorProperties & YAML</i>"]
        events["spector-events<br/><i>Telemetry event bus</i>"]
        testsupport["spector-test-support<br/><i>Harnesses & mocks</i>"]
    end

    subgraph "🧠 Cognitive Memory Layer (memory/)"
        kernel["spector-kernel<br/><i>Sealed Off-Heap Bundle Kernel</i>"]
        memory["spector-memory<br/><i>4-Tier Memory, Cognitive Pathways &amp; Daemons</i>"]
        providerapi["spector-provider-api<br/><i>Provider SPI</i>"]
        providers["spector-providers<br/><i>AI Providers (Ollama, OpenAI, ONNX)</i>"]
        ingestion["spector-ingestion<br/><i>Sensory &amp; file ingest pipeline</i>"]
        inspect["spector-inspect<br/><i>Bundle inspection CLI</i>"]
        metrics["spector-metrics<br/><i>Micrometer + Prometheus</i>"]
    end

    subgraph "⚡ Nervous System &amp; Gateways (synapse/)"
        synapse["spector-synapse<br/><i>Spring Boot 4 REST/SSE &amp; Chat Graph</i>"]
        gateway["spector-gateway<br/><i>API Gateway &amp; routing</i>"]
        connector["spector-connector<br/><i>Apache Camel connectors</i>"]
        mcp["spector-mcp<br/><i>MCP Server — Agent-native</i>"]
        cli["spector-cli<br/><i>spectorctl CLI &amp; standalone spector.jar</i>"]
        spring["spector-spring<br/><i>Spring AI VectorStore</i>"]
        batch["spector-batch<br/><i>Batch migration engine</i>"]
    end

    subgraph "🌐 Distributed (cluster/)"
        cluster["spector-cluster<br/><i>Multi-node coordination</i>"]
    end

    subgraph "📦 Client SDKs (sdks/)"
        javaclient["spector-client<br/><i>Java client SDK</i>"]
    end

    subgraph "📈 Performance &amp; Validation (bench/)"
        bench["spector-bench<br/><i>JMH benchmarks &amp; cognitive eval</i>"]
    end
Loading

Note

Index implementations in spector-index: hnsw/ (graph-based ANN, Quantized HNSW), spectrum/ (SpectorIndex, multi-tier sharding), bm25/ (keyword scoring + analyzers), splade/ (sparse neural representations).


🔗 Dependency Graph

graph TD
    synapse["🌐 synapse"] --> mcp["🤖 mcp"]
    synapse --> connector["🔌 connector"]
    synapse --> metrics["📈 metrics"]
    synapse --> events["📡 events"]
    synapse --> memory["🧠 memory"]

    mcp --> memory
    mcp --> ingestion["📥 ingestion"]
    cli["🖥️ cli"] --> memory
    cli --> mcp
    cli --> ingestion

    memory --> index["📊 index"]
    memory --> core["🔬 core"]
    memory --> cpu["⚡ cpu"]
    memory --> config["⚙️ config"]
    memory --> providerapi["🧬 provider-api"]

    index --> core
    index --> config
    index --> commons["📄 commons"]

    gpu --> index
    gpu --> core
    gpu --> commons

    cpu --> core
    cpu --> commons

    metrics --> memory
    metrics --> events

    connector --> ingestion
    connector --> providerapi

    spring["🌱 spring"] --> memory
    spring --> metrics
    bench["🧪 bench"] --> memory
    bench --> providers["🤖 providers"]
Loading

Legend: Solid arrows = compile dependency. Dotted arrow (bench) = benchmark execution dependency.

Dependency rules:

Path Description
memory → kernel, index, core, cpu Cognitive memory composes kernel bundles, 4-tier engrams, and HNSW/BM25 indexes
kernel → (JDK only) Sealed storage kernel — zero external dependencies, Panama FFM + Vector API only
cli → memory + mcp + ingestion CLI with local batch and remote (client) modes
synapse → memory + ingestion Spring Boot 4 application: REST + SSE + cluster coordination
mcp → memory + ingestion MCP agent entry point (in-process, zero network)
commons ← ingestion & memory Houses IngestionBoundary decoupling sensory ingestion from memory (ADR-0037)
index → core, config, commons HNSW, SpectorIndex, BM25, and SPLADE storage foundations

!!! important No circular dependencies. spector-memory contains both vector search and cognitive memory stores, keeping the API gateway (spector-synapse) decoupled from low-level storage.


🧠 Cognitive Pathways Architecture (ADR-0035, ADR-0036, ADR-0037)

Spector structures all cognitive processes into a unified, composable pipeline architecture modeled after neurobiological synaptic pathways. Every cognitive operation extends AbstractPathway<S, R> and executes a directed sequence of SynapticRelay<S> stages orchestrated by PathwayEngine<S>.

Two types are easy to confuse, so it is worth stating the split explicitly:

  • Pathway<I, O> is the public operation — typed input to typed output. It owns scope entry/exit, the conduction outcome, metrics, and the projection of a conducted signal into a report.
  • PathwayEngine<S> is the relay conductor — signal in, same signal out. It runs the stage list, applies each stage's ErrorPolicy, and records traces. It is not a Pathway and deliberately does not implement it.

AbstractPathway bridges the two by composition: it holds an engine and conducts it, then projects. Stage lists are declared in a PathwayRecipe and assembled by PathwayComposer, which is the only supported authoring API — it is where the build-time safety checks live (rejecting a retry on a non-idempotent relay, a timeout on a non-interruptible one, or ABORT inside a divergent branch).

The 7 Unified Cognitive Pathways

Pathway Input Signal Result / Report Biological Analog & Key Functions
Remember RememberSignal RememberResult Encoding & Consolidation: Dopamine-modulated surprise gating, perceptual dedup, transactional cortical commit, Hebbian graph co-activation linking, and knowledge graph hyper-edge enrichment.
Recall RecallSignal List<CognitiveResult> Retrieval & Reconstruction: 6-phase fused hybrid search (lexical BM25, dense HNSW, SPLADE), spreading activation across Hebbian manifolds, lateral inhibition, and ColBERT v2 token MaxSim reranking.
Reflect ReflectSignal ReflectReport Circadian Sleep Consolidation: Non-REM/REM sleep cycle simulation, synaptic homeostasis & power-law decay, soul-drift re-fusion, episodic clustering, and off-heap partition compaction.
Dream DreamSignal DreamReport Counterfactual Replay & Discovery: Generative self-model replay, Langevin stochastic discovery, synthetic episode generation, and Expected Free Energy (EFE) triage.
Decide DecideSignal DecideReport Active Inference Policy Selection: Free Energy Principle (FEP) optimization, homeostatic goal appraisal, policy rollouts, and action selection.
Express ExpressSignal ExpressReport Embodied Stylometry & Prosody: Affective somatic feedback, vocal prosody vector calculation, idiolect stylometric matching, SSML tagging, and 3D blendshape parameter synthesis.
Wander WanderSignal WanderReport Default Mode Network (DMN): Idle-state autobiographical sampling, Hopfield continuous attractor energy minimization, spontaneous associative synthesis, and longitudinal continuity checkpointing.

Relay Decorator Execution Order (ADR-0036)

Each relay in a cognitive pathway can be declaratively wrapped with resilience and observability decorators via PathwayComposer or builder composition. To guarantee deterministic behavior across all pathways, decorators execute in a strict onion order:

Pathway Conduction Invocation
  │
  ▼
┌─────────────────────────────────────────────────────────────┐
│ 1. Tracing & Metrics Interceptor                            │
│    Measures relay latency; emits spector.pathway.relay      │
├─────────────────────────────────────────────────────────────┤
│ 2. TimeoutRelay                                             │
│    Enforces per-relay deadlines; emits onTimeout            │
├─────────────────────────────────────────────────────────────┤
│ 3. RetryRelay                                               │
│    Retries transient failures with backoff; emits onRetry   │
├─────────────────────────────────────────────────────────────┤
│ 4. CircuitBreakerRelay                                      │
│    Guards downstream services; emits onCircuitEvent         │
├─────────────────────────────────────────────────────────────┤
│ 5. BulkheadRelay                                            │
│    Isolates concurrency per relay; emits onBulkheadReject   │
├─────────────────────────────────────────────────────────────┤
│ 6. DegradationPolicy                                        │
│    Gracefully degrades or bypasses stage; emits onDegraded  │
├─────────────────────────────────────────────────────────────┤
│ 7. Underlying SynapticRelay                                 │
│    Executes core domain logic (pure cognitive step)         │
└─────────────────────────────────────────────────────────────┘

Circuit Breakers & Fault Tolerance

To prevent cascading failures when external providers (LLMs, embedding APIs, graph engines) experience outages, relays support circuit breaker isolation:

  • States: CLOSED (normal operation), OPEN (tripped after consecutive failures; fast-fails requests), and HALF_OPEN (probe execution testing upstream health).
  • Event Lifecycle: Emits TRIP, PROBE, CLOSE, and REJECT events through CircuitBreaker.EventListener.
  • Fault Kinds: Exceptions are classified into TIMEOUT, PROVIDER, CAPACITY, TRANSIENT, DATA, or CONTROL via FaultClassifier, driving degradation decisions.

Pathway Observability & Micrometer Metrics (ADR-0036 R7)

The pathway execution subsystem provides zero-dependency instrumentation in spector-commons via PathwayObservationHook and exports 8 standardized Micrometer meters in spector-metrics:

Meter Name Type Tags Description
spector.pathway.conduct Timer pathway, finish End-to-end execution duration of the full pathway cycle
spector.pathway.relay Timer pathway, relay, status Latency of an individual relay (status: success, degraded, bypassed, failed, short_circuited)
spector.pathway.degraded Counter pathway, relay, kind Count of degraded relay executions classified by FaultKind
spector.pathway.circuit Counter circuit, state_transition Circuit breaker state transitions (trip, probe, close, reject)
spector.pathway.bulkhead.reject Counter bulkhead Count of relay invocations rejected due to exhausted concurrency permits
spector.pathway.timeout Counter pathway, relay Count of relay executions aborted due to timeout expiration
spector.pathway.retry Counter pathway, relay Count of retry attempts executed after initial transient failure
spector.pathway.nested Timer from, to Latency and frequency of nested pathway invocations

Domain Telemetry & ConductionOutcome

All domain reports (DreamReport, DecideReport, ExpressReport, WanderReport, ReflectReport, RememberResult) encapsulate an immutable ConductionOutcome. Consumers can inspect:

  • finish(): Terminal disposition (COMPLETED, SHORT_CIRCUITED, FAILED).
  • degraded() & isDegraded(scope): Whether any stage completed in a degraded fallback state.
  • bypassed() & isBypassed(scope): Whether any stages were skipped due to gate predicates or circuit state.
  • traces(): Microsecond-precision per-relay execution trace timeline.

📥 Data Flow: Ingest Path

sequenceDiagram
    participant Client as 👤 Client (CLI/MCP/REST)
    participant Pipeline as 🔄 IngestionPipeline
    participant Embed as 🧠 ParallelEmbeddingPipeline
    participant Target as 💾 IngestionTarget
    participant Store as 💾 Storage (mmap)

    Client->>Pipeline: pipeline.ingest(file)
    Pipeline->>Embed: generateEmbeddings()
    Embed-->>Pipeline: dense + sparse vectors
    Pipeline->>Target: target.store(chunk)
    Target->>Store: write to off-heap MemorySegment
    loop Each chunk
        Pipeline->>Pipeline: TextChunker.chunk(content)
        Pipeline->>Embed: embed(chunkTexts) via virtual threads
        Embed-->>Pipeline: List<vector>
        Pipeline->>Target: target.ingest(id, text, vector)
        Target->>Store: VectorStore + VectorIndex
    end
    Store-->>Client: ✅ Indexed
Loading
  1. Client calls pipeline.ingest() — unified across CLI, MCP, and application code
  2. IngestionPipeline handles chunking (from config) and parallel embedding
  3. IngestionTarget receives pre-embedded chunks — storing directly in SpectorMemory
  4. Downstream storage writes to off-heap memory and indexes with HNSW/BM25

Tip

FileDiscoveryService can be used independently for file discovery without any engine dependency.


🔍 Data Flow: Search Path

sequenceDiagram
    participant Client as 👤 Client
    participant Memory as 🧠 SpectorMemory
    participant Pipeline as ⚙️ RecallPipeline
    participant BM25 as 📝 BM25 Search
    participant HNSW as 🧠 Dense HNSW
    participant Sparse as 📈 Sparse (SPLADE)
    participant RRF as 🧬 RRF Fusion
    participant Rerank as 🚀 ColBERT Rerank
    participant Graph as 🔗 Graph Expansion

    Client->>Memory: recall(query, options)
    Memory->>Pipeline: execute(query, options)
    par Parallel first-stage retrieval on virtual threads
        Pipeline->>BM25: exact term matching
        Pipeline->>HNSW: dense semantic search
        Pipeline->>Sparse: learned sparse search
    end
    BM25 & HNSW & Sparse->>RRF: Rank merge
    RRF->>Rerank: Token-level late interaction MaxSim
    Rerank->>Graph: Multi-hop graph expansion & gating
    Graph-->>Client: ✨ Final cognitive memories
Loading
  1. Recall Pathway receives options (TextSearchMode, RecallMode, etc.)
  2. Dense Vector, BM25, and Sparse (SPLADE) searches run in parallel on virtual threads
  3. RRF Fusion merges the ranked lists using reciprocal rank scores
  4. ColBERT v2 Reranking scores the top candidates using SIMD MaxSim operations
  5. Graph Expansion traverses Hebbian/Entity/Temporal edges for neighbor expansion

🤖 Data Flow: MCP Agent Path

sequenceDiagram
    participant Agent as 🤖 AI Agent (Claude/Cursor)
    participant MCP as 📡 MCP Transport (stdio / Streamable HTTP)
    participant Handler as 🔧 McpToolHandler
    participant Memory as 🧠 SpectorMemory
    participant SIMD as 🔬 SIMD Kernels

    Agent->>MCP: tools/call {"name": "memory_recall", "arguments": {"query": "..."}}
    MCP->>Handler: MemoryRecallTool.execute(args)
    Handler->>Memory: recall(query, options)
    Memory->>SIMD: 6-phase scoring + Panama off-heap reads
    SIMD-->>Memory: CognitiveResult[] (~130µs)
    Memory-->>Handler: List<CognitiveResult>
    Handler-->>MCP: CallToolResult
    MCP-->>Agent: JSON-RPC response with recalled memories
Loading

The MCP path operates directly against SpectorMemory. The MCP server wraps tool handler calls with JSON-RPC transport. There is zero network overhead because everything runs in the same JVM process.

Tip

For full MCP architecture details and tool schemas, see the dedicated MCP Integration page.


🧵 Threading Model: Virtual Threads

Spector is designed from the ground up for Java virtual threads:

Tip

No synchronized blocks anywhere in the codebase. All coordination uses ReentrantLock to avoid virtual thread pinning.

Operation Threading Strategy
REST request handling One virtual thread per request
Hybrid search Parallel BM25 + HNSW via StructuredTaskScope
Bulk ingest Virtual thread per document
Embedding generation Batched across virtual threads
HNSW construction (>10K) Virtual threads per core for parallel insertion
Distributed fan-out Virtual thread per shard query

📈 Scaling Results

At 50K docs with hybrid search (384-dim, production-realistic):

Virtual Threads Throughput Scaling
1 3,739 ops/s 1.0×
4 10,317 ops/s 2.8×
8 11,812 ops/s 3.2×
16 14,022 ops/s 3.7×

Note

Scaling depends on vector dimensions and workload type. 384-dim shows ~3.7× at 16 threads due to higher per-query memory bandwidth. Individual HNSW queries are inherently sequential (graph traversal data dependencies) — scaling comes from concurrent queries sharing CPU cores.


💾 Memory Model: Panama Off-Heap

All vector data lives off-heap using the Panama Foreign Function & Memory API:

graph TB
    subgraph "☕ JVM Heap (minimal)"
        HG["HNSW Graph<br/>(adjacency lists)"]
        BM["BM25 Index<br/>(inverted index)"]
        ES["Engine State<br/>(config, lifecycle)"]
    end

    subgraph "🧊 Off-Heap (Panama MemorySegment)"
        VS["Vector Store<br/>Contiguous float32, SIMD-aligned<br/>Zero-copy reads, no GC pressure"]
        QS["Quantized Store<br/>INT8 or PQ codes"]
        GM["GPU Device Memory<br/>CUDA via FFM"]
    end

    HG -.-> VS
    BM -.-> VS
    ES -.-> QS
    ES -.-> GM
Loading

Benefits:

  • ✅ Zero GC pressure — Vectors never touch the garbage collector

  • ✅ Instant startup — Memory-mapped files load via mmap syscall, no deserialization

  • ✅ SIMD-friendly layout — Contiguous float32 arrays ready for Vector API operations

  • ✅ Explicit lifecycle — Arena-scoped memory with deterministic cleanup

  • ✅ Memory efficiency — Store billions of vectors limited only by disk/address space

📊 Storage Types

Store Location Use Case
InMemoryVectorStore Off-heap (Arena) Development, small datasets
MmapVectorStore Memory-mapped file Production, persistence
QuantizedVectorStore Off-heap (INT8) Memory-constrained deployments
IvfPqStore Off-heap (PQ codes) Billion-scale (32× compression)

🌐 API Layer

graph TD
    subgraph "SpectorNode - Armeria Server, single port"
        CORS["CorsService decorator"]
        Auth["API Key decorator"]
        COMPRESS["EncodingService - gzip/brotli"]
        subgraph "ApiModule Registration"
            SE["🔍 SearchEndpoint"]
            IE["📥 IngestEndpoint"]
            RE["🤖 RagEndpoint"]
            DE["🗑️ DocumentEndpoint"]
            STE["📊 StatusEndpoint"]
            ESE["📡 EventStreamEndpoint"]
        end
        gRPC["gRPC Service<br/>inter-node fan-out"]
        HEALTH["💚 /health"]
        PROM["📊 /metrics"]
    end

    subgraph "REST Controller Layer"
        MC["MemoryController<br/>/api/v1/memory/*"]
        SC["SystemController<br/>/api/v1/system/*"]
        HC["HealthController<br/>/api/v1/system/*"]
    end

    subgraph "Service Layer"
        MS["MemoryService"]
    end

    subgraph "Core Engine"
        SM["SpectorMemory"]
    end

    MC & SC & HC --> MS
    MS --> SM
Loading

Every request runs on its own virtual thread. The Armeria server handles HTTP REST, gRPC, and SSE events on a single port. API endpoints are registered via ApiModule components, enabling straightforward API versioning (/api/v1, /api/v2).

Streaming via SSE

The /api/v1/search/stream endpoint uses Server-Sent Events to emit results progressively. The /api/v1/events endpoint provides a live event stream where clients can subscribe to search, ingest, cluster, MCP, and engine events with optional category filtering.


🔗 See Also

🏠 Home


Clone this wiki locally