+===============================================================================+
| Mercury Agent ♱ v2.1.0 |
| Neuro-Symbolic AI for Autonomous, Multi-Model, Multi-Domain Anomaly Detection |
| |
| 7-Phase Evolution | Hybrid Fusion ML | Production Security |
| Neural + Symbolic | 30 Detection Engines | Post-Quantum Crypto |
| Ethical Governance | Multi-Head Attention | OWASP Validation |
| |
| LAYER 3: Ethics | LAYER 2: ML/AI | LAYER 1: Security |
| ------------------- | ------------------- | ------------------- |
| Benevolence >= 0.99 | Fusion Network | Kyber-1024/ML-DSA-65 |
| Lyapunov Stability | Ensemble Averaging | JWT Authentication |
| Civilization-First | Property Testing | Rate Limiting |
| |
| Archetype for a civilized evolution. |
+===============================================================================+
On the Layer-3 entries: these are runtime-enforced gates, not static guarantees. The benevolence gate defaults to ≥ 0.99 and is configurable no lower than a hard 0.70 floor (enforced in
src/omni_mercury_engine/cognitive/ethical_bounding.py); Lyapunov stability is monitored and reported (is_stable), not proven a priori. Seedocs/MATH_SPEC.md§2.2 and §2.7.3.
Copyright 2025 Steel Security Advisors LLC Author/Inventor: Andrew E. A. Contact: steel.sa.llc@gmail.com License: GNU General Public License v3.0 or later (SPDX: GPL-3.0-or-later) Version: v2.1.0 Date: 2026-07-11 AI Co-Architects: Eris ✠ | Eden ♱ | Devin ⚛︎ | Claude ⊛
Mercury Agent is a comprehensive neuro-symbolic AI Archetype implementing a 7-phase cognitive evolution for multi-domain anomaly detection. The system combines neural pattern recognition with symbolic reasoning to produce explainable, ethically-bounded decisions across security, medical, environmental, humanitarian, and infrastructure domains.
The framework embodies a Civilization-First philosophy, prioritizing ethical AI governance and humanitarian impact. Every action must clear a mandatory benevolence enforcement gate — 0.99 by default, configurable no lower than a hard 0.70 floor — keeping the system in service of human flourishing and civilizational progress.
Project Philosophy: Mercury Agent represents the next evolution in AI systems - one that combines the pattern recognition power of neural networks with the interpretability and reasoning capabilities of symbolic AI. This neuro-symbolic fusion enables the system to not only detect anomalies but explain why they matter and what actions should be taken.
Security Disclosure: This is a research-grade implementation. Production use REQUIRES:
- Independent security review by qualified professionals
- Validation on domain-specific real-world datasets (MIMIC-III, NSL-KDD)
- Clinical validation for any medical applications
- Post-quantum cryptography for Mercury Agent and FINDΩYOU™ is derived from AMA Cryptography
- FINDΩYOU™ is a separate, near-future sibling platform whose humanitarian scope — locating the lost, missing, and abducted to reunite families and help bring perpetrators to justice — is served by Mercury under its Civilization-First mission.
This project is licensed under the GNU General Public License v3.0 or later (SPDX: GPL-3.0-or-later)
Everyone is permitted to copy and distribute verbatim copies of this license document, but changing it is not allowed. Any derivative work, fork, or modification must also be released under GPL v3 — no proprietary versions allowed, ever. This ensures the code and all future improvements remain free and open source forever, even if used by corporations or governments.
Status: Research-grade | Community-tested | Not externally audited Last Updated: 2026-07-11
This block is regenerated by
.github/workflows/benchmark.ymlon every push tomainand committed back to the repo, so the most recent live-data run is always front-and-center — never lost to expiring CI artifacts.
The full result file lives at benchmarks/mercury_benchmark_results.json.
Methodology is documented in docs/BENCHMARKS.md.
A multi-panel visual summary appears in the Current Benchmarks and Visual Proof section below.
| Metric | Current | Previous | Δ |
|---|---|---|---|
| Mean ROC-AUC | 0.8251 | 0.8251 | +0.0000 |
| Median ROC-AUC | 0.8747 | 0.8747 | +0.0000 |
| Mean Oracle F1 | 0.5998 | 0.5998 | +0.0000 |
| Datasets (successful / total) | 66 / 75 | 66 / 75 | +0.0000 |
| Run timestamp (UTC) | 2026-06-21T00:11:10.047740+00:00 | 2026-06-21T00:11:10.047740+00:00 | — |
| Commit | a7a194b |
a7a194b |
— |
Regression gates: ROC-AUC must stay ≥ 0.75 and Mean Oracle F1 ≥ 0.55 (the coarse floors set in .github/workflows/benchmark.yml, held ~9% below the current headline AUC/F1 so they do not flap on normal load-availability variance). Fine-grained, tight regression protection is the deterministic per-dataset guard (anomaly_regression_guard.py), not these coarse floors. CI fails the workflow if either drops below threshold.
Comparability note. The aggregate ROC-AUC above blends two evaluation regimes. Only externally-labeled datasets — ADBench and the standard labeled sets, where labels come from a source independent of the scored signal — are comparable to published baselines; the comparable headline is ADBench Mean AUC 0.8251. Several environmental loaders are self-labeled by thresholding the signal they score, which inflates their AUC toward 1.0 (label leakage) and is reported for pipeline transparency only. See Label provenance and comparability in the benchmarks below for the full split.
Fusion-hardening (transductive). The unsupervised fusion changes in detectors/statistical.py are measured independently of the headline run by a committed, re-runnable 18-set real-ADBench transductive harness (research/omni_equation/harness_adbench.py): Mean AUROC 0.7397 → 0.7634 (+2.37 pts, 14 W / 2 tie / 2 L) from base e118e1f to the hardened detector (full per-set results in research/omni_equation/adbench_results.json). Real data only — the harness fails loudly rather than substituting synthetic. Re-run: python research/omni_equation/harness_adbench.py.
Codebase Scale (measured, not estimated)
These numbers are produced by scripts/measure_codebase_scale.py and gated in CI: --check README.md fails the build if the block below drifts from what is on disk, and --update README.md regenerates it. They reflect what is actually on disk, not what marketing copy assumes.
| Measurement | Value |
|---|---|
Python source files in src/omni_mercury_engine/ |
756 |
| Source lines of code (LOC) | ~398,000 |
Top-level subpackages (true Python packages with __init__.py) |
49 |
Files importing PyTorch (optional [ml] extra) |
175 |
Distinct torch.nn.Module subclasses |
173 |
Detector classes (class *Detector) in detectors/ |
88 |
Data-loader classes (class *Loader) in loaders/ |
21 |
Test modules (test_*.py) / total test LOC |
646 modules / ~183,000 LOC |
| GitHub Actions workflows | 23 |
Generated by python scripts/measure_codebase_scale.py --update README.md and gated in CI (scripts/measure_codebase_scale.py --check README.md); do not hand-edit between the markers.
The neuro-symbolic claim is real, not naming theatre. Concrete evidence in-repo:
src/omni_mercury_engine/cognitive/— ~32,000 LOC, 30 modules (excluding__init__.py; measured at v2.1.x), includingneural_memory_layer.py,symbolic_logic_layer.py,neurosymbolic_fusion.py,differentiable_logic.py(withNeuralPredicateEncoder,DifferentiableRuleModule,NeuralTheoremProver,CounterfactualReasoner— all realnn.Modulesubclasses),cognitive_evolution_engine.py,chain_of_thought.py,multi_hop_reasoner.py,causal_discovery.py,case_based_reasoning.py,predictive_coding.py,formal_verification.py,knowledge_graph.py,multi_agent_coordination.py,reflexion.py,plasticity_engine.py,hierarchical_planning.py,ipb_engine.py.src/omni_mercury_engine/core/neurosymbolic_hub.py— 1,446 LOC; definesKnowledgeGraph,NeuralEncoder,SymbolicRule,NeuroSymbolicHub,ExplainableOutput,FusionMode.- PyTorch is the engine of the ML / neuro-symbolic stack, shipped as the optional
[ml]extra (torch>=2.2.0,torchvision>=0.17.0,pytorch-lightning>=2.0.0) — it is not a core install dependency: the top-levelimport omni_mercury_engineand the default (non-ML) detection path import without torch, while the ML / neuro-symbolic submodules importtorchat module scope and require the[ml]extra. The measured count of torch-importing files andnn.Modulesubclasses is in the table above. Install the full neuro-symbolic path withpip install mercury-agent[ml].
Mercury Agent is a neuro-symbolic AI that exposes anomaly detection as one of its capabilities — it is not "an anomaly detection library that happens to use neural networks."
The claim is operational, not just architectural: the production fusion trainer (OmniMercuryEngine.fit_fusion) co-trains a differentiable Logic Tensor Network with the neural network by default (symbolic_weight="adaptive") — a symbolic satisfaction term enters the loss and its gradient flows into the network's anomaly head. The weight follows a label-scarcity schedule that cleared a pre-registered dominance bar on real ADBench labels (no full-data AUC regression beyond noise; a seed-agreed low-data lift) and decays to the purely-neural path when labels are abundant. This is co-training in the loss, not a post-hoc score blend. Full accounting, ablation, and the evidence-based keep/quarantine verdicts: docs/NEUROSYMBOLIC.md.
7-Phase Neuro-Symbolic Evolution
Mercury Agent implements a 7-phase cognitive architecture that progressively builds from basic neural memory to curiosity-driven exploration. Each phase names a real, CI-gated module in src/omni_mercury_engine/cognitive/ (the "Neuro-Symbolic Tests" required check enforces them):
| Phase | Component | Description | Key Features |
|---|---|---|---|
| Phase 1 | Neural Memory Layer | Memory embeddings and pattern detection | K-means clustering, episodic/semantic memory, pattern recognition |
| Phase 2 | Symbolic Logic Layer | NetworkX-based logic graphs | Explainable decisions, threshold rules, inference chains |
| Phase 3 | Neuro-Symbolic Fusion | Hybrid anomaly scoring | Attention-based fusion, confidence weighting, neural-symbolic integration |
| Phase 4 | Enhanced Anomaly Detection | Memory knowledge graph | Bayesian predictor, HMM predictor, external data integration |
| Phase 5 | Autonomous Agent | OODA loop implementation | Observe-Orient-Decide-Act-Reflect, user synchronization, Mercury/AMA Disconnect |
| Phase 6 | Ethical Bounding | Benevolence scoring (>=0.99) | Harm reduction, equity calculation (Gini), empathy module |
| Phase 7 | Cognitive Evolution Engine | Curiosity-driven exploration | CuriosityEngine: measured novelty scoring (online diagonal-Mahalanobis distance) of detected anomalies, wired into the CognitiveOrchestrator. The earlier self-play / genetic-mutation / theory-of-mind internals were measured decorative and removed. |
The seven phases are the historical evolution spine, not the whole cognitive
layer: cognitive/ has since grown to 30 modules, integrated at runtime by the
CognitiveOrchestrator (knowledge graph, multi-hop reasoner, and uncertainty
quantifier always on; plasticity, causal discovery, IPB, case-based reasoning,
and indicator development on by default; curiosity and enhanced detection
opt-in — see cognitive/__init__.py).
Detailed benchmarks and methodology
Snapshot of the committed mercury_benchmark_results.json run (2026-06-21). The
Latest Benchmark Results block at the top of this
README is the CI-refreshed current headline; the tables here are the detailed
breakdown. All numbers are measured on real data — no synthetic substitution, no
hyperparameter tuning.
MercuryAnomalyDetector was run untuned over 66 of 75 attempted real-world
datasets across 12 domains (the 9 failures were unreachable external sources,
tracked by a two-lane reachability harness, .github/workflows/dataset-reachability.yml).
Results fall into two regimes, and only the first is comparable to published
detectors:
- Externally-labeled (comparable). Labels come from a source independent of the scored features (ADBench, NSL-KDD, BATADAL, SMD/NAB, CWRU/MSDS). The headline is the 47-dataset ADBench result: Mean AUC 0.8251 / Mean F1 0.5975.
- Self-labeled / threshold-derived (not comparable). Some environmental loaders manufacture labels by thresholding a feature they also score, so the detector recovers them almost trivially (0.97–1.00 AUC). That is label leakage, not accuracy — reported below for pipeline transparency only.
The same provenance discipline extends to the live-API loaders: of 20 concrete
loaders, only 2 (network_security, sepsis) produce labels independent of any
scored feature, and the governed-fusion suite gates self-improvement on those 2
alone. CI enforces the split (python -m omni_mercury_engine.loaders.label_provenance --check);
see docs/SELF_IMPROVEMENT_LOOP.md for the
governed self-improvement rollout.
Ensemble — unsupervised, three-component weighted fusion:
| Component | Weight | Method | Mean AUC | Median AUC |
|---|---|---|---|---|
| ResonanceScore | 40% | FFT harmonic spectral profiles | 0.7665 | 0.8294 |
| KinematicScore | 30% | Physics jerk/curvature via np.diff |
0.6009 | 0.6116 |
| InfoGeometryScore | 30% | Fisher-information Mahalanobis OOD | 0.8267 | 0.8760 |
| Ensemble | 100% | Weighted combination | 0.8251 | 0.8747 |
| Aggregate (66 successful datasets) | Value |
|---|---|
| Mean AUC / Median AUC | 0.8251 / 0.8747 |
| Std AUC | 0.1638 |
| Mean / Median Oracle F1 | 0.5998 / 0.6747 |
Oracle F1 is an upper bound (best of a multi-strategy threshold sweep), not operational performance.
These rows use labels independent of the scored features — the numbers to line up against other detectors.
| Domain | Datasets | Mean AUC | Mean F1 |
|---|---|---|---|
| ADBench | 47 | 0.8251 | 0.5975 |
| Industrial (BATADAL) | 1 | 0.9114 | 0.5545 |
| Security (NSL-KDD, ThreatIntel) | 2 | 0.8995 | 0.7423 |
| Space (NASA, Solar) | 2 | 0.8753 | 0.7356 |
| Time Series (SMD, NAB) | 2 | 0.6807 | 0.4333 |
| General (ADRepository) | 1 | 0.7086 | 0.3468 |
| Academic (CWRU, MSDS)† | 2 | 1.0000 | 1.0000 |
† Academic/General rows are genuinely labeled but tiny; the 1.0000 AUC reflects easy separability at that size, not headline accuracy. The representative comparable figure is the 47-dataset ADBench row.
Against PyOD baselines on the full 57-dataset ADBench suite, Mercury's tier
detector places 3rd of 8 by mean rank (3.79) — behind knn (2.40) and
local_outlier_factor (3.69), and ahead of isolation_forest (3.94),
hbos (5.14), copod (5.62), and ecod (6.11). Mercury is strong-to-dominant
on high-dimensional sets where distance methods collapse, and weaker on
low-dimensional local-density sets. Full method-by-method tables:
benchmarks/COMPETITIVE_BENCHMARK.md.
Untuned 5-fold-CV spot check versus standard baselines (regenerated from
license-clean scikit-learn datasets, deterministic at --seed 42, via
python -m benchmarks.empirical_benchmark --readme-subset):
| Detector | breast_cancer | covtype | digits_8 |
|---|---|---|---|
| Mercury-Agent | 0.796 | 0.896 | 0.489 |
| One-Class SVM | 0.662 | 0.901 | 0.767 |
| LOF | 0.544 | 0.667 | 0.571 |
| Elliptic Envelope | 0.888 | 0.899 | 0.503 |
| TranAD (supervised SOTA ref) | 0.940 | 0.892 | 0.742 |
| MAAT (supervised SOTA ref) | 0.946 | — | 0.747 |
AUC; bold = best per column, "—" = not evaluated. Mercury is unsupervised and untuned; TranAD/MAAT (Tuli et al., VLDB 2022; Kang & Kang, 2023) are supervised published references marking the ceiling, not re-run here.
| Validation | Result |
|---|---|
| Threshold calibration (MD-011) | 33/52 datasets improved (71.2%), mean F1 +0.097 |
| Fusion-weight CV (MD-003) | adaptive weights within 0.007 F1 of L-BFGS optimal; 82.7% of datasets within 0.02 F1 |
| Conformal coverage (MD-005, CrossConformal @0.90/0.95) | 69.2% empirical (36/52 datasets) |
Per-domain results from run_all_benchmarks.py:
| Domain | Mean AUC | Domain | Mean AUC |
|---|---|---|---|
| Earthquake | 0.9367 | Network Security | 0.7983 |
| Tsunami | 0.8905 | Hurricane | 0.7233 |
| Tornado | 0.8803 | Flood | 0.6837 |
| Pandemic | 0.8588 | FEMA | 0.6573 |
| Energy | 0.8038 | Marine | 0.3540 (gate fail) |
Wildfire, Volcanic, Landslide, Sepsis, and Financial reported no data in this run (API key required or upstream unavailable).
Labels here are a deterministic threshold on a scored feature, so high AUC is leakage, not accuracy. Listed for pipeline transparency only — never averaged into a headline.
| Domain | Datasets | Mean AUC | Label rule (leaky) |
|---|---|---|---|
| Air Quality (EPA) | 1 | 0.9975 | PM2.5 > 35.4 µg/m³ |
| Climate (NOAA GSOD/StormEvents/ERDDAP) | 3 | 0.9939 | ±3σ threshold |
| Environmental (USGS/NOAA/EPA) | 3 | 0.8856 | ±3σ threshold |
| Ocean (NOAA Buoy) | 1 | 0.8510 | ±3σ threshold |
| Disaster (FEMA) | 1 | 0.9993 | threshold/polarity-derived |
Use Mercury when anomaly decisions must be interpretable, data types are diverse and need adaptive profiling, or no labeled anomaly data is available. Prefer alternatives when labeled data exists (use a supervised classifier) or memory is tightly constrained. Caveats, stated plainly: KinematicScore is near-random on shuffled tabular data (mean AUC 0.60); 6 datasets fall below 0.50 AUC (ensemble inversion on high-dimensional data); no tuning was performed.
The benchmark panels below are generated from the current committed
mercury_benchmark_results.json run (2026-06-21; Mean AUC 0.8251 over 66
successful datasets) by benchmarks/generate_benchmark_visuals.py; the
calibration panel comes from the MD-011/MD-005 calibration-validation run
(52 datasets).
Neuro-symbolic report — AUC/F1 distributions, component boxplots, domain performance, dataset rankings:

Anomaly-detection analysis — per-component AUC breakdown, ensemble-vs-best scatter, threshold usage:

Performance dashboard — timing, AUC by category, dataset size, adaptive weights, precision-recall:

Benchmark summary — per-dataset AUC bar chart with the run mean line:

Calibration and conformal — MD-011 calibration lift and MD-005 coverage:

Adaptive weights — unsupervised weight distribution and mean weights by domain:

Full results: benchmarks/mercury_benchmark_results.json · Methodology: docs/BENCHMARKS.md · Competitive detail: benchmarks/COMPETITIVE_BENCHMARK.md.
Click to expand 3R Mathematical Framework
The 3R mechanism (Recursion-Resonance-Refactoring) is a mathematical method for anomaly detection and optimization that forms the core of Mercury Agent's detection capabilities. This framework combines three complementary engines for pattern recognition and adaptive learning.
The Recursion Engine implements hierarchical multi-scale pattern analysis through recursive feature extraction:
Mathematical Foundation:
R(x, d) = f(x) + α · R(g(x), d-1) for d > 0
R(x, 0) = f(x)
Where f(x) is the feature extraction function, g(x) is the downsampling operator, α is the decay factor, and d is the recursion depth.
Key Capabilities:
- Multi-scale hierarchical pattern detection across temporal and spatial dimensions
- Recursive feature extraction that captures both local and global anomaly signatures
- Self-similar pattern recognition inspired by fractal mathematics
- Adaptive depth adjustment based on data complexity
The Resonance Engine performs FFT-based frequency-domain analysis to detect harmonic anomalies and periodic patterns:
Mathematical Foundation:
H(ω) = |FFT(x)|²
A(x) = Σₙ H(n·ω₀) / Σ H(ω) (Harmonic Ratio)
Where ω₀ is the fundamental frequency and n represents harmonic indices.
Key Capabilities:
- Frequency-domain anomaly detection using Fast Fourier Transform
- Harmonic analysis for periodic pattern recognition (e.g., Schumann resonance at 7.83 Hz)
- Signal amplification for weak anomaly signatures
- Cross-domain frequency correlation (seismic, electromagnetic, biological)
The Refactoring Engine provides dynamic optimization and adaptive model refinement:
Mathematical Foundation:
θ_{t+1} = θ_t - η · ∇L(θ_t) + β · (θ_t - θ_{t-1}) (Momentum-based optimization)
L_total = L_detection + λ₁·L_stability + λ₂·L_ethical
Key Capabilities:
- Dynamic model optimization based on detection performance
- Code complexity metrics for security review and maintainability
- Adaptive threshold adjustment using Bayesian calibration
- Self-healing capabilities inspired by CRISPR mechanisms
The three engines work together to create emergent detection capabilities:
| Synergy | Components | Capability |
|---|---|---|
| Harmonic Analysis | Resonance + Recursion | Multi-scale frequency decomposition |
| Quantum-Inspired Paths | Recursion + Refactoring | Simulated annealing for optimization |
| Ava Equation | All 3R | Unified anomaly scoring: A = R·H·O |
| Asymptotic Horizons | Resonance + Refactoring | Convergence monitored via a Lyapunov-style decay schedule |
The unified 3R scoring function with ethical gating:
A = (w_R·R(x) + w_H·H(ω) + w_O·O(θ)) · η_Ethical^Φ
Where:
R(x)= Recursion score (multi-scale hierarchical analysis)H(ω)= Harmonic/Resonance score (frequency coherence)O(θ)= Optimization/Refactoring score (adaptive theta)η_Ethical= Ethical compliance threshold (default 0.96, medical fallback 0.93)Φ= Golden ratio constant (1.618033988749895)- Golden-ratio fusion weights (canonical Φ:1:1):
w_R = φ/(φ+2) ≈ 0.447,w_H = 1/(φ+2) ≈ 0.276,w_O = 1/(φ+2) ≈ 0.276(sum to 1.0)
Lyapunov Stability (decay-schedule monitor): convergence is monitored against the target condition V̇ ≤ -λV with rate λ = 0.25 (LyapunovConstants.LAMBDA_CONVERGENCE; distinct from the double-helix adaptation rate LAMBDA_DECAY = 0.18 — see Mathematical Foundations) — the empirical trajectory is measured against this schedule, not guaranteed a priori.
The ThreeRAnomalyTransformer provides a differentiable PyTorch implementation of the 3R mechanism for end-to-end training:
from omni_mercury_engine.ml import ThreeRAnomalyTransformer, LyapunovAnomalyLoss
# Initialize model with domain-specific configuration
model = ThreeRAnomalyTransformer(
input_dim=25, # Feature dimension (e.g., NSL-KDD)
d_model=256, # Model dimension
n_heads=8, # Attention heads
num_layers=2, # 3R attention layers
ethical_threshold=0.96, # Domain-specific (0.93 for medical)
)
# Lyapunov-constrained loss for stable predictions
loss_fn = LyapunovAnomalyLoss(
mu_stability=0.1, # Stability constraint weight
alpha=0.25, # Convergence rate (matches CONVERGENCE_RATE_PARAMETER)
)
# Training step
output = model(x) # x: [batch, seq_len, input_dim]
loss = loss_fn(
x=x,
x_recon=output["reconstruction"],
anomaly_scores=output["anomaly_scores"],
labels=labels, # Supervised signal for accuracy
)Key Features:
- Golden-ratio OAE weights: Mathematically grounded fusion (0.447/0.276/0.276)
- Bounded outputs: Sigmoid activation ensures anomaly scores in [0, 1]
- Supervised + unsupervised: BCE loss with labels + reconstruction loss
- Lyapunov stability: Penalizes divergent predictions for safety-critical applications
- Domain configs: See
configs/ablation_3r_lyapunov.yamlfor medical/security/infrastructure presets
The 3R mechanism is integrated throughout Mercury Agent:
- Detectors: All 30 detection engines leverage 3R for feature extraction
- Fusion Network: Multi-head attention combines 3R outputs across domains
- Ethical Governance: Refactoring engine ensures Lyapunov stability constraints
- Self-Healing: CRISPR-inspired adaptation uses recursive pattern learning
- PyTorch Training:
ThreeRAnomalyTrainer(Lightning) for end-to-end training with stability constraints
Click to expand navigation
- Executive Summary
- Key Capabilities
- Use Cases by Sector
- Performance Metrics
- Quick Start
- Reproducible Verification
- Testing and Quality Assurance
- Documentation
- Cross-Platform Support
- Build System
- Mathematical Foundations
- Contributing
- Unique Features
- License
- Contact and Support
- Acknowledgments
- Legal Disclaimer & Attribution
Problem Statement and Solution
Modern anomaly detection faces three critical challenges:
- Domain Fragmentation: Security, medical, environmental, and infrastructure domains require specialized expertise with no unified framework
- Ethical Blind Spots: Most ML systems lack bias detection, fairness metrics, and ethical governance
- Production Gaps: Research models often lack security hardening, input validation, and deployment infrastructure
Mercury Agent addresses all three challenges through:
- Unified Framework: 30 detection engines under a single hybrid fusion architecture covering medical, security, space, infrastructure, and environmental domains
- Ethical Governance: Fairlearn bias detection with demographic parity, equalized odds, and 80% rule enforcement; 180+ ethical scalars with Lyapunov stability
- Production Security: OWASP-compliant input validation, post-quantum cryptography (Kyber-1024 / ML-KEM-1024, ML-DSA-65, SPHINCS+-SHA2-256f-simple) via AMA Cryptography v3.3.0, JWT authentication, rate limiting
- Medical & Healthcare: Sepsis detection, cardiology analysis, pandemic response (requires clinical validation)
- Security & Intelligence: Threat detection, intelligence fusion, cyber fortress monitoring
- Space & Environmental: Solar storm detection, Schumann resonance analysis, disaster precursors
- Infrastructure & Humanitarian: Critical infrastructure monitoring, crisis response, climate resilience
See Use Cases by Sector for detailed scenarios.
Unique Differentiators
Defense-in-depth with 3 independent layers for anomaly detection:
| Layer | Protection | Components |
|---|---|---|
| 1. Core Infrastructure | Security foundation | Kyber-1024 / ML-DSA-65 / SPHINCS+ PQC (AMA Cryptography v3.3.0), JWT auth, OWASP validation |
| 2. ML/AI Pipeline | Detection intelligence | 30 engines, hybrid fusion, multi-head attention |
| 3. Ethical Governance | Fairness assurance | Fairlearn bias audit, 180+ ethical scalars, Lyapunov stability |
The signature innovation providing transparent, auditable AI decision-making:
- Fairlearn Integration: Demographic parity, equalized odds, 80% rule enforcement
- 180+ Ethical Scalars: Omnibenevolent constraints across all operations
- Lyapunov Stability: decay-schedule monitor of system convergence (measured, not guaranteed)
- Civilization-First Philosophy: Humanitarian impact prioritized in all design decisions
Optimized for both accuracy and interpretability:
- Feature Fusion:
torch.cat()across 30 detector outputs - Decision Fusion: Weighted voting with learned importance scores
- Attention Fusion: Multi-head attention (4 heads) for cross-domain correlation
- Final Score:
0.7 * MLP + 0.3 * weighted_voteensemble
Closes the loop from interpret to deter on top of the calibrated detection
certificate, with an explicit, principled "don't-know" gate. On the core
engine it is opt-in via engine.enable_decision_layer(); the served deploy
entrypoints — the /api/v1/detect/flagship HTTP route, the MCP
mercury_detect_fusion tool, and mercury-agent detect -d fusion — enable it
by default, so every served flagship detection carries a decision section
(additive key; library consumers constructing OmniMercuryEngine directly are
unchanged).
- Calibration-grounded abstention — reuses the engine-wide
ThreeStatecontract: the conformal label set is authoritative (singleton → GROUNDED;{0,1}→ UNAVAILABLE, a resolvable don't-know;{}→ UNDECIDABLE, a fail-closed hold). Neuro-symbolic disagreement, drift, or an ethical-gate refusal can only move a verdict toward abstention. - Bounded, non-destructive response —
monitor/alert/recommend_mitigation/recommend_conversion/recommend_restoration/escalate_to_human/request_input/hold. The loop recommends and escalates; it never autonomously executes a destructive action (a test invariant enforces this). The restorative convert verbs (ResponsePolicy(restorative=True), opt-in) recommend non-violent convert-to-benign / restore-to-pre-harm paths and are recommend-only by construction. - Auditable & verifiable — a deterministic, JSON-safe
DecisionRecordwith the calibrated confidence, reasons, caveats and full evidence provenance, plus a one-paragraphexplain()(and afrom_dictinverse for reload). An append-onlyDecisionLedger(bounded ring buffer, O(1) incrementalsummary(), thread-safe, JSON-save/load) and aDecisionLoopadd the verify step over a stream of decisions. Closes into the existing CAP 1.2 alerting and autonomy (AgentAction) channels.
See examples/decision_abstention_response_demo.py and
docs/capability_vs_vision_matrix.md.
Key Achievements
| Achievement | Description |
|---|---|
| Multi-Domain Coverage | 30 detection engines across 12 domains (8 new statistical methods) |
| Ethical Governance | Fairlearn bias detection, 180+ ethical scalars |
| Production Security | OWASP validation, PQC support, JWT authentication |
| Comprehensive Testing | 8,789 tests collected (2026-06-10, full optional-dependency surface) across the test modules counted in the CI-gated Codebase Scale block; property-based testing, security scanning |
| Benchmark Coverage | 66 reproducible datasets (of 75 attempted; 47 ADBench + 28 domain); canonical Mean ROC-AUC 0.8251 / Median 0.8747 (CI-refreshed "Latest Benchmark Results" block above); externally-comparable subset ADBench Mean AUC 0.8251 |
| Cross-Platform | Linux (Ubuntu 22.04+ supported in CI), macOS 13+, Windows 10/11 (via WSL2), Docker, Kubernetes (Helm chart); 8 integrated observability platforms (Prometheus, Elastic/OpenSearch, Splunk, Datadog, Azure Anomaly Detector, Netdata, Grafana, InfluxDB) |
| Mathematical Rigor | Lyapunov stability (λ = 0.25, certified by tools/lyapunov_validator.py), σ_Immutable hard gate (trained-network decision threshold 0.93; GOSNN gating default 0.96), Benevolence ≥ 0.99 |
| Codebase Scale | All structural counts are measured and CI-gated in the Codebase Scale block above (source files, LOC, packages, detector/loader classes, nn.Module subclasses, test modules, workflows) — no hand-typed figures |
Implementation Status Matrix
| Component | Status | Notes |
|---|---|---|
| Hybrid Fusion Network | Complete | Multi-head attention, ensemble averaging |
| Decision / Abstention / Response | Complete | Calibration-grounded ThreeState "don't-know" gate; bounded non-destructive response; enabled by default on the served entrypoints (API /detect/flagship, MCP, CLI fusion), opt-in on the core engine via enable_decision_layer() |
| Bias Detection | Complete | Fairlearn metrics, built-in fallback |
| Input Validation | Complete | OWASP-compliant, SQL/XSS/injection detection |
| JWT Authentication | Complete | Native stdlib omni_mercury_engine.security.native_jwt (HS256/HS384/HS512); all three route through AMA Cryptography v3.3.0's ACVP-validated native HMAC, fail-closed (no stdlib downgrade) |
| Property Testing | Complete | Hypothesis-based test suite |
| Post-Quantum Crypto | Complete | AMA Cryptography (sole PQC backend) |
| Real-Data Validation | Pending | Requires MIMIC-III, NSL-KDD datasets |
Legend:
- Complete: Implemented and tested
- Pending: Requires real-world dataset validation
Note: Core anomaly detection is benchmarked on 66 reproducible real-world datasets (of 75 attempted; see
benchmarks/mercury_benchmark_results.jsonand the reproducibility note in the Benchmarks section). Domain-specific modules may still require validation on their target datasets before production deployment.
Experimental Research Areas: The use cases below represent targeted experimental applications where Mercury Agent multi-domain anomaly detection may provide value. These are research-grade implementations requiring independent validation before deployment in regulated, clinical, or mission-critical environments.
Medical & Healthcare
Unique Value: Multi-modal health monitoring with ethical AI governance (Not approved for clinical deployment without independent validation):
- Sepsis Detection: SOFA/qSOFA scoring per JAMA 2016 Sepsis-3 guidelines with bias-audited predictions
- Cardiology Analysis: ECG rhythm analysis covering 13 arrhythmia types, Framingham risk scoring
- Neurocritical Care: ICP monitoring, stroke detection, TBI assessment with fairness constraints
- Pandemic Response: SEIR modeling, outbreak prediction, mutation tracking with ethical oversight
Validation Required: All medical applications require clinical validation on real patient data (MIMIC-III) before any deployment.
Security & Intelligence
Unique Value: Multi-source threat detection with ethical constraints:
- Threat Detection: SQL injection, XSS, path traversal detection with pattern matching and ML classification
- Intelligence Fusion: 13-source fusion (OSINT, SIGINT, HUMINT, GEOINT) with bias-aware aggregation
- Cyber Fortress: Hash integrity verification, quantum-resistant validation with Kyber-1024 / ML-DSA-65 (AMA Cryptography v3.3.0)
- Traffic Analysis: Encrypted traffic anomaly detection with privacy-preserving techniques
Space & Environmental
Unique Value: Cosmic and terrestrial anomaly detection:
- Solar Storm Detection: CME tracking, geomagnetic storm prediction with calibrated confidence intervals
- Schumann Resonance: ELF spectrum analysis (7.83 Hz fundamental) for environmental monitoring
- Disaster Precursors: Earthquake, tsunami early warning systems with uncertainty quantification
- Geological Hazards: Volcanic, landslide, wildfire detection with multi-sensor fusion
Infrastructure & Humanitarian
Unique Value: Critical infrastructure protection with Civilization-First principles:
- Critical Infrastructure: 55 CISA National Critical Functions monitoring with anomaly detection
- Crisis Response: Essential workers, government facilities tracking with ethical constraints
- Climate Resilience: Climate adaptation, extreme weather pattern detection
- Economic Sectors: 21 ISIC categories, financial crisis detection with fairness auditing
Latency Benchmarks
| Configuration | CPU Latency | GPU Latency (RTX 4090) |
|---|---|---|
| Full (30 engines) | ~500ms | ~50ms |
| Standard | ~250ms | ~25ms |
| Fast (statistical only) | ~100ms | ~10ms |
Historical snapshot: synthetic data, Python 3.12, Ubuntu 22.04; real-world
performance may vary 20-40%. The GPU column was measured on the stated
hardware and is not re-verified per release. Reproduce the CPU-side core
claim on your host with mercury-agent tool detector_profiler, which
emits a provenance-stamped JSON latency/RSS profile (schema
mercury.tools.detector_profiler/v1; most recent sandbox re-run: median
47.7 ms, p95 56.3 ms — the <100ms fast-path claim holds).
Memory Footprint
| Component | Memory |
|---|---|
| Harmonic Encoder | ~10 MB |
| Fusion Network | ~50 MB |
| DeepFace (VGG-Face) | ~200 MB |
| Full Runtime | ~500 MB |
Module Performance
| Operation | Latency | Throughput |
|---|---|---|
| Module Instantiation (1) | 0.020ms | 49,932 ops/sec |
| Module Instantiation (12) | 0.057ms | 17,697 ops/sec |
| Space Exploration | 0.206ms | 2,919,469 samples/sec |
| Cosmic Ray Detection | 0.324ms | 3,081,781 samples/sec |
| Collatz Exploration | 67.07ms | 74,544 cases/sec |
Module performance benchmarks measured on synthetic data (2026-01-27). Results may vary by hardware.
Installation
# Clone repository
git clone https://github.com/Steel-SecAdv-LLC/Mercury-Agent.git
cd Mercury-Agent
# Install core dependencies
pip install -e .
# Install with all features
pip install -e ".[all]"Linux (Ubuntu 22.04+):
# Install system dependencies
sudo apt-get install build-essential python3-dev
# Install package
pip install -e ".[all]"macOS (13+):
# Install via Homebrew if needed
brew install python@3.12
# Install package
pip install -e ".[all]"Windows (10/11):
# WSL2 recommended for full compatibility
wsl --install
# Or install directly (some features may be limited)
pip install -e ".[all]"The package uses the following naming conventions:
- Package Name:
mercury-agent(used in pip install and pyproject.toml) - CLI Command:
mercury-agent(the command-line interface) - Internal Module:
omni_mercury_engine(the Python import path)
This means you install with pip install mercury-agent but import with from omni_mercury_engine import .... The internal module name preserves backward compatibility with the original codebase while the package and CLI names reflect the current project identity.
DeepFace/dlib (Optional): For biometric features, install DeepFace:
# Linux/macOS
pip install deepface
# Windows (use pre-built wheel)
pip install https://github.com/z-mahmud22/Dlib_Windows_Python3.x/releases/download/v19.22.99/dlib-19.22.99-cp312-cp312-win_amd64.whl
pip install deepfaceBasic Usage
from omni_mercury_engine import OmniMercuryEngine
# Initialize engine
engine = OmniMercuryEngine(mode="fusion", device="cuda")
# Detect anomalies
result = engine.detect_with_fusion(data)
print(f"Anomaly Score: {result['anomaly_score']:.3f}")
print(f"Is Anomaly: {result['is_anomaly']}")The engine starts with five general-purpose base detectors (statistical, temporal, spatial, dimensional, directive). Other detectors declared in the detector manifest — for example the trajectory geo_movement detector or the unsupervised kmeans_distance clusterer — are opt-in: enable them per engine instance without disturbing the calibrated default fusion path.
from omni_mercury_engine import OmniMercuryEngine
from omni_mercury_engine.detectors import KMeansDistanceDetector
engine = OmniMercuryEngine(mode="fusion")
# Which manifest detectors exist, and which are active on this engine?
engine.available_detectors() # {"statistical": True, ..., "geo_movement": False}
# Enable a registered detector by name (constructed from the manifest)...
engine.enable_detector("geo_movement")
# ...or register your own BaseDetector instance
engine.register_detector("kmeans", KMeansDistanceDetector(n_clusters=8))Once enabled, a detector contributes its feature group to fit_fusion and detect_with_fusion like any built-in; one that cannot process a given input is skipped gracefully rather than failing the call. A detector added after fit_fusion has trained participates only after a re-fit, since fusion inference is restricted to the feature groups training saw.
# View available commands
mercury-agent --help
# Anomaly detection
mercury-agent detect --input data.csv --detector fusion
# Biometric analysis
mercury-agent biometric --help
# Security analysis
mercury-agent security --helpLive Anomaly Detection Demo
A live demonstration script showcases real-time anomaly detection across security, medical, and environmental domains.
# Run a single-domain demo (security / medical / environmental)
python examples/live_anomaly_demo.py --domain security --samples 30
# Run all domains
python examples/live_anomaly_demo.py --all --samples 20Demo features: real-time streaming with configurable sample rates,
multi-domain detection, ethical-governance benevolence scoring (target 0.99+),
threat classification (LOW/MEDIUM/HIGH/CRITICAL), and JSON output for monitoring
integration. A recorded session is at
assets/live_anomaly_demo.mp4.
Sample output:
[ 1/15] 2026-01-06 02:21:17.325 | ANOMALY | Score: 1.000 | Conf: 0.990 | Benev: 0.990
Threat: HIGH | Bytes: 9246/600
[ 2/15] 2026-01-06 02:21:17.442 | NORMAL | Score: 0.797 | Conf: 0.831 | Benev: 0.990
...
Total Samples: 15
Anomalies Detected: 4
Detection Rate: 26.67%
Avg Confidence: 0.8701
Avg Benevolence: 0.9900
Runtime: 1.94s
| Option | Description | Default |
|---|---|---|
--domain |
Detection domain (security/medical/environmental) | security |
--all |
Run demo for all domains | False |
--samples |
Number of samples to process | 30 |
--delay |
Delay between samples in ms | 150 |
--output |
Output JSON file path | None |
--quiet |
Suppress verbose output | False |
Docker Quick Start
# Build the Docker image
docker build -t mercury-agent:latest .The default entrypoint runs the FastAPI server for production inference:
# Run API server (default mode)
docker run -d \
-e JWT_SECRET_KEY=$(openssl rand -hex 32) \
-e OMNI_RATE_LIMIT_ENABLED=true \
-p 8000:8000 \
mercury-agent:latest
# API documentation available at http://localhost:8000/docs
# Health check: http://localhost:8000/healthOverride the default command to run training jobs:
# Train all disaster detection neural networks
docker run -it \
-v $(pwd)/models:/app/models \
mercury-agent:latest \
python -c "from omni_mercury_engine.detectors.geological.disaster_detectors import train_all_disaster_networks; train_all_disaster_networks()"
# Train with GPU support (requires nvidia-docker)
docker run -it --gpus all \
-v $(pwd)/models:/app/models \
mercury-agent:latest \
python -c "from omni_mercury_engine.detectors.geological.disaster_detectors import train_all_disaster_networks; train_all_disaster_networks(device='cuda')"
# Run benchmarks
docker run -it \
-v $(pwd)/results:/app/results \
mercury-agent:latest \
python benchmarks/empirical_benchmark.pyRun batch inference on data files:
# Run anomaly detection on a data file
docker run -it \
-v $(pwd)/data:/app/data \
-v $(pwd)/output:/app/output \
mercury-agent:latest \
python -c "
from omni_mercury_engine import OmniMercuryEngine
import numpy as np
engine = OmniMercuryEngine(mode='fusion')
data = np.load('/app/data/input.npy')
result = engine.detect_with_fusion(data)
print(f'Anomaly Score: {result[\"anomaly_score\"]:.3f}')
"
# Use CLI for detection
docker run -it \
-v $(pwd)/data:/app/data \
mercury-agent:latest \
mercury-agent detect --input /app/data/input.csv --detector fusion# Start interactive shell for development
docker run -it \
-v $(pwd):/app \
mercury-agent:latest \
/bin/bash
# Run tests inside container
docker run -it \
-v $(pwd):/app \
mercury-agent:latest \
pytest tests/ -v| Variable | Description | Default |
|---|---|---|
JWT_SECRET_KEY |
Shared JWT signing key. Unset in production, JWTAuth derives the key via AMA HD Key Management (api/auth.py) — deterministic fleet-wide when AMA_MASTER_SEED is set, per-process (with a logged warning) otherwise |
None |
AMA_MASTER_SEED |
Hex AMA HD master seed (openssl rand -hex 64); makes HD-derived keys identical across workers/replicas/restarts |
unset |
MERCURY_CACHE_SECRET |
Shared HMAC-SHA256 secret for Redis cache entry signing (RedisCache); tampered entries raise CacheIntegrityError |
unset |
OMNI_RATE_LIMIT_ENABLED |
Enable API rate limiting (api/server.py) |
true |
OMNI_RATE_LIMIT_REQUESTS_PER_MINUTE |
Sustained rate-limit budget | 100 |
OMNI_RATE_LIMIT_BURST |
Token-bucket burst size | 20 |
MERCURY_ENV |
Environment mode (development / production; unknown values raise) |
development |
MERCURY_CORS_ORIGINS |
Explicit CORS origin allow-list; in production, unset means same-origin only (CORS middleware disabled) | unset |
OMNI_ENSEMBLE_CALIBRATION |
Per-detector ensemble score calibration (rank/ecdf/isotonic/platt/none); see docs/DETECTION_MECHANISMS.md |
rank |
OMNI_ENSEMBLE_WARMUP |
Warm-up window the ensemble calibrators train on (int points / float fraction) | whole training series |
OMNI_DETECTOR_NAN_POLICY |
Detector-tier NaN/Inf policy (neutral/impute/flag/raise) |
neutral |
OMNI_DETECTOR_MAX_MAGNITUDE |
Unified safe magnitude cap for detector scores/inputs/metadata | 1e100 |
OMNI_DETECTOR_RIDGE_FACTOR |
Digital-twin scale-relative Tikhonov ridge factor | 1e-6 |
OMNI_DETECTION_CONFIG |
Path to a YAML/JSON file supplying the OMNI_DETECTOR_* / OMNI_ENSEMBLE_* knobs (env overrides file) |
unset |
| Mount Point | Purpose |
|---|---|
/app/data |
Input data files |
/app/models |
Trained model checkpoints |
/app/results |
Benchmark and inference results |
/app/output |
Output files from batch processing |
This architecture supports scalable deployments via Kubernetes/Helm where the API server handles inference requests while training jobs can be run as separate batch workloads.
Kubernetes/Helm
# Install via Helm
helm install mercury-agent ./helm/mercury-agent \
--set image.tag=latest \
--set secrets.jwtSecret=$(openssl rand -hex 32)
# Check deployment status
kubectl get pods -l app=mercury-agentExecutable Lyapunov Certificate
The Lyapunov decay rate λ cited throughout this README and docs/MATH_SPEC.md is not a prose claim -- it is enforced by an executable certificate.
Single source of truth. configs/lyapunov_canonical.yaml declares the
canonical linear surrogate (A, P) of the fusion-trajectory dynamics and
the claimed rate λ = 0.25 (certified with a 2× margin against the
larger computed rate; see below). LyapunovConstants.LAMBDA_CONVERGENCE
in src/omni_mercury_engine/core/centralized_constants.py is the matching
Python constant; the reconciliation test
tests/tools/test_lyapunov_reconciliation.py fails CI the moment they
diverge.
Mathematical kernel. tools/lyapunov_validator.py implements two
modes:
| Mode | Inputs | Method |
|---|---|---|
quadratic |
(A, P) matrices, claimed λ |
Symmetric-definite generalized eigenvalue problem Q v = μ P v with Q = AᵀP + PA, solved via Cholesky + numpy.linalg.eigvalsh. Certifies λ* = −μ_max. |
samples |
[{V, Vdot}, …], claimed λ |
Worst observed ratio infₛ(−Vdotₛ / Vₛ). Suitable for non-linear V and regression gating, not a proof. |
The canonical config certifies λ = 0.5 and the claim 0.25 is therefore satisfied with a 2× margin (and a 1e-8 tolerance for floating-point round-off).
CLI.
# Certify any Lyapunov YAML (top-level A/P/λ, or a nested `lyapunov:` block).
python -m tools.lyapunov_validator configs/lyapunov_canonical.yaml
# JSON output, suitable for piping into jq / CI annotations.
python -m tools.lyapunov_validator configs/ablation_3r_lyapunov.yaml | jq .Exit codes: 0 certified pass · 1 claim does not hold · 2 config error.
Ablation Runner with Lyapunov Pre-Gate
scripts/run_ablation.py is the canonical entry-point for experiments that must not run unless their Lyapunov claim is provably satisfied. The pre-gate is non-negotiable.
# Single-purpose Lyapunov certificate -- gate only.
python scripts/run_ablation.py \
--config configs/lyapunov_canonical.yaml \
--out artifacts/lyapunov_check.json \
--skip-run
# Multi-variant ablation with the certificate carried in a nested `lyapunov:` block.
python scripts/run_ablation.py \
--config configs/ablation_3r_lyapunov.yaml \
--out artifacts/ablation_result.json \
--timeout 1800Exit codes: 0 success · 2 config not found or --timeout invalid (no JSON report; absence of the output file is the explicit "no run attempted" signal for pollers) · 3 Lyapunov gate failed, experiment not launched (JSON report is written with the failed certificate's computed_lambda / claimed_lambda so CI dashboards can render the diagnosis) · 4 no run_command declared (JSON report written) · 124 --timeout exceeded, GNU timeout(1) convention (JSON report written with run_timed_out=true). Process-group isolation is enforced via start_new_session=True + os.killpg so shell-spawned grandchildren are reaped on timeout (POSIX); Windows uses the documented weaker Popen.terminate() fallback.
Universal Equation Optimizer & Runtime Equation Profiles
tools/equation_optimizer.py is a reproducible, safety-gated search over the
mathematical surfaces inventoried from docs/MATH_SPEC.md.
It freezes the original equations as an immutable baseline, searches a
constrained universal candidate family, and only promotes a candidate that
clears hard gates — ethical-compliance invariant, output range [0, 1],
contraction α ≤ 0.999, and Lyapunov λ ≥ 1e-6 — emitting versioned artifacts plus
rollback / continuous-revalidation metadata. When no candidate beats the proven
baseline under the gates, the baseline wins under the documented decision rules. Full
operating guide: docs/EQUATION_RESEARCH_PROTOCOL.md.
# Run the optimizer; emits the artifact set under --output-dir.
python -m tools.equation_optimizer \
--math-spec docs/MATH_SPEC.md \
--output-dir artifacts/equation_optimization
# Governed protocol runner (inventory → freeze → search → gate → ledger).
python scripts/run_equation_research_protocol.py --config configs/equation_research_protocol.yaml
# Compare the frozen baseline against a candidate runtime profile on real data.
python scripts/compare_runtime_equation_profiles.py --seed 3 --n 800 --out artifacts/runtime_equation_compare.jsonExit codes: 0 pipeline ok · 1 a hard gate failed after a completed run ·
2 tool/config/input exception. The winner JSON records every objective
metric, the constraint-detail booleans, and the selected parameters so a
promotion is fully auditable.
Runtime equation profiles (opt-in serve path). The optimizer's frozen OAE
surface is also exposed at inference through
omni_mercury_engine.core.equation_profiles: OmniMercuryEngine(..., equation_profile="baseline_original_v1") (or a per-call
detect_with_fusion(..., equation_profile=...) / score_fusion(..., equation_profile=...)) blends the calibrated neural probability with the
golden-ratio R/H/O equation signal. None (the default) is an exact no-op,
so the legacy serve/benchmark path is byte-for-byte unchanged; the blend
metadata is surfaced under result["equation_profile"].
Three profiles ship: baseline_original_v1 and quiet_horizon_v1 are frozen
(fixed 0.70/0.30 split), while phi_fibring_v1 harmonises the blend with
Mercury's canonical fibring fusion (core/fibring_fusion.py) — a golden-ratio
base split (φ/(1+φ) : 1/(1+φ) ≈ 0.618 : 0.382) plus correlation-aware
decorrelation: when the equation signal is redundant with the neural score
(|Pearson r| ≥ 0.85) the lower-variance stream is shrunk by 1/(1+|r|) and the
weights renormalise, so a duplicated channel cannot double-count.
Tests: tests/tools/test_equation_optimizer.py (pipeline artifacts + CLI smoke),
tests/core/test_equation_profiles.py (bounded/distinct profiles, None
pass-through, unknown-profile rejection, channel→R/H/O mapping),
tests/scripts/test_run_equation_research_protocol.py,
tests/scripts/test_compare_runtime_equation_profiles.py.
Hardware Benchmark Harness
scripts/run_hardware_benchmark.py produces reproducible performance numbers for the validator pipeline, paired with the environment fingerprint that makes the result scientifically comparable. Full operating guide: docs/HARDWARE_HARNESS.md.
python scripts/run_hardware_benchmark.py \
--config configs/lyapunov_canonical.yaml \
--iters 2000 --warmup 200 \
--out artifacts/hwbench.json
# Treat measured throughput as a regression gate:
python scripts/run_hardware_benchmark.py --min-ops-per-sec 1500The JSON report contains the certificate result, the environment fingerprint (Python version, NumPy version, platform, CPU count, CPU affinity, scaling governor), and the timing block (mean / p50 / p95 / p99 / max / total). Throughput is reported as samples / total_s (which equals 1 / mean_s exactly when mean_s is the arithmetic mean of the same samples — statistics.fmean, in this harness — up to floating-point rounding). The harness emits both total_s and ops_per_sec so the assertable invariant ops_per_sec == samples / total_s can be checked directly by downstream tooling without re-deriving the mean; this also keeps the contract stable if the central-tendency estimator is ever swapped for a trimmed or median variant.
Documentation Drift Gate
scripts/check_readme_lyapunov.py blocks any pull request whose README.md or docs/MATH_SPEC.md documents a numeric λ that disagrees with the imported source-of-truth constants. The gate is plural: it carries a registry with one entry per canonical constant referenced in user-facing docs, currently λ_lyapunov = LyapunovConstants.LAMBDA_CONVERGENCE (the fusion-trajectory Lyapunov decay rate) and λ_decay = omni_mercury_engine.core.double_helix_engine.LAMBDA_DECAY (the double-helix evolutionary adaptation rate). Each check imports its canonical at call time (rather than parsing a YAML or constants file with a regex), matches every Greek / LaTeX / English prose form of its symbol, applies anchor-token context filtering so unrelated λs are ignored, and enforces a min_occurrences floor so silently deleting every documented mention cannot achieve a vacuous-green pass.
# Equivalent to the ISO Hardening / Workflow Hardening CI gate.
# Requires ``pip install -e .`` so the importer can resolve numpy
# (which double_helix_engine imports at module load).
python scripts/check_readme_lyapunov.pyThis gate is the second mandatory checkpoint: changing LAMBDA_CONVERGENCE or LAMBDA_DECAY requires updating the canonical config, the documentation, and (for LAMBDA_CONVERGENCE) the reconciliation test in lock-step.
Note: Running the full test suite requires dev dependencies. Install with:
pip install -e ".[dev]"
Test Suite
# Install development dependencies
pip install -e ".[dev]"
# Run full test suite
pytest tests/ -v
# Run with coverage
pytest tests/ --cov=src/omni_mercury_engine --cov-report=html
# Run property-based tests
pytest tests/test_property_based.py -v
# Run security scan
bandit -r src/ -f txtThe test suite includes:
- 8,789 tests collected (
pytest --collect-only -q, 2026-06-10, with the optionaltorch/scikit-learn/hypothesis/fastapidependencies installed) across the test modules counted in the CI-gated Codebase Scale block (the exact module count is measured from disk, never hand-typed). A minimal install collects fewer tests because modules gated behind those optional imports skip at collection time. - Property-based testing with Hypothesis for edge case discovery
- Security scanning with Bandit integrated in CI/CD
- Coverage tracking: the merge gate enforces measured floors
(CORE ≥ 33 %, FULL ≥ 62 %) on every PR;
pyproject.toml [tool.coverage.report] fail_under = 85is the long-term aspirational target. Coverage is regenerated per release viapytest --cov=src/omni_mercury_engine --cov-report=term; do not assert a percentage from this README — re-measure on the head ofmain.
Notable Test Suites (historical, v1.4.x → v1.6.x):
test_enhanced_anomaly_detection.py: 38+ tests for enhanced statistical methods, cross-platform hub, ensemble coordinationtest_cortical_network.py: 40+ tests for 6-layer cortical architecturetest_statistical_real.py: 30+ tests for Z-score, IQR, adaptive detectiontest_temporal_real.py: 26+ tests for time-series pattern detectiontest_resilience_real.py: 27+ tests for circuit breaker state machinetest_enhanced_geological_detectors.py: 60+ tests for Landslide/Wildfire/Volcanic with 3R synapsestest_advanced_optimizers.py: 50+ tests for SyntheticGradient/DTP/AMAV integrationtest_mercury_amacrypto.py: 60+ tests for AMA Cryptography PQC adapter and EWMA timing monitor
v1.7 development-cycle additions:
tests/security/test_native_jwt.py+tests/security/test_native_jwt_ama_routing.py: 50 tests pinning the native JWT module (HS256/HS384/HS512 byte-equivalence with stdlib, RFC 4231 KAT incl. SHA-384, AMA-routing vs stdlib cross-check,alg: nonerejection).tests/security/test_cve_2026_6357_regression.py: regression guard for the pip CVE-2026-6357 floor across every install path in CI.tests/security/test_nist_fips_kat.py: NIST FIPS 203/204/205 ACVP-Server KAT vectors verified bit-for-bit (ML-DSA-65 deterministic sigGen, ML-KEM-1024 decapsulation, SLH-DSA-SHAKE-128s sigGen).tests/tools/test_lyapunov_validator.py+tests/tools/test_lyapunov_reconciliation.py: pin the executable Lyapunov certificate against documentation drift and againstLyapunovConstants.LAMBDA_CONVERGENCE.tests/scripts/test_check_readme_lyapunov.py+tests/scripts/test_run_ablation.py+tests/scripts/test_run_hardware_benchmark.py: lock the ISO Hardening operator-tool surface (drift gate, ablation runner pre-gate, hardware harness throughput math).tests/api/test_server_comprehensive.py::TestLifespanWarmup: 4 tests pinning the API warmup lifespan (wiring, success path, internal-failure propagation under the fail-fast contract, TestClient lifecycle).tests/datasets/test_unreachable_loaders_{offline,network}.py: two-lane reachability harness for the 11 watch-listed datasets whose upstream sources are flaky or not currently fetchable (9 failed in the committed benchmark run; NOAA StormEvents and NOAA ERDDAP recovered).tests/validation/test_synthetic_policy_gate.py: locks theMERCURY_ALLOW_SYNTHETICpolicy gate across every loader that previously exposed ause_synthetickwarg bypass.
Test Suite Stabilization (v1.6.0 Patch — historical):
- Fixed 100+ test failures caused by missing FastAPI dependency, uninitialized federation attributes, and unreachable synthetic data thresholds
- All fixes are minimal and targeted — no unnecessary refactoring
Ethics Scoring (test_ai_ethics.py):
The legacy heuristic ethics tests use keyword-presence scoring against
stringified action parameters — for example, the compassion check awards
a base 0.55 plus a +0.1 boost when create_backup appears in the
serialised params. The dual hard ethical gate contract (Benevolence
- σ_Immutable, raising
EthicalConstraintViolationError(check=…)at every public detect / analyze / predict surface) is the production decision-boundary and is exercised bytests/ethical/test_hard_enforcement.py; the keyword tests are retained for backward compatibility with the v1.x heuristic surface and should not be read as the operative ethics contract.
Continuous Integration
GitHub Actions enforce the following gates on every pull request and push to main/develop:
| Workflow | Job | Blocking | Description |
|---|---|---|---|
ci.yml |
Code Quality | yes | black, flake8, ruff, mypy (strict), pydocstyle |
ci.yml |
Workflow Hardening | yes | actionlint + zizmor + repository workflow invariants |
ci.yml |
Type Checking | yes | mypy on src/ and graduated strict test directories |
ci.yml |
Security Scan | yes | bandit, safety, pip-audit, semgrep against an isolated install |
ci.yml |
Core Tests | yes | pytest, ≥ 33 % combined stmt+branch coverage on the curated core lane |
ci.yml |
Neuro-Symbolic Tests | yes | 7-phase cognitive architecture + safeguards + ethics regression suite |
ci.yml |
Integration Tests | yes | tests/integration/ against mocked external services |
ci.yml |
Performance Benchmark | PR-only | TTLCache / synthetic-gradient regression gate |
ci.yml |
Ethics Audit | yes | benchmarks/run_ethics_audit.py (EthicalAutonomyGovernor, σ_Immutable, OAE) |
ci.yml |
ML Tests | nightly/PR-to-main | Full suite under tests/, ≥ 62 % coverage, real AMA Cryptography build |
ci.yml |
Docker Build + Trivy | yes | Multi-stage runtime image, CRITICAL/HIGH = 0 beyond the enumerated, expiring .trivyignore ledger (ignore-unfixed: false) |
ci.yml |
Docs Build | yes | Sphinx build of the narrative docs |
iso-hardening.yml |
Docs λ Drift Gate | yes | scripts/check_readme_lyapunov.py -- canonical λ = 0.25 across docs |
iso-hardening.yml |
Examples Parity | yes | examples/*.py must run end-to-end and emit known markers |
iso-hardening.yml |
Load Tests | yes | k6 smoke + locust headless against the live API with FastAPI lifespan warmup + CI warmup loop, SLO p(95) < 500 ms on detection endpoints, p(99) < 150 ms on /health |
iso-hardening.yml |
ISO Hardening Success | yes | Rollup of the three gates above (single required status check) |
security.yml |
Container/SAST scan | yes | Trivy + Semgrep with deterministic SARIF categories |
pqc-production-check.yml |
PQC Production Readiness | yes | KAT vectors, NIST FIPS ACVP-Server vectors, real AMA Cryptography build |
benchmark.yml |
Live Benchmark | scheduled | Refreshes benchmarks/mercury_benchmark_results.json + README block |
dataset-reachability.yml |
Loader Reachability | nightly | Offline lane + nightly network lane for the 11 watch-listed loaders |
network-tests.yml |
External Source Probe | nightly | Diagnostic probe of upstream data providers |
docker.yml |
Docker Release | tag-driven | Push runtime image with provenance attestation |
format.yml |
Formatting check | yes | Drift guard for black / ruff format output |
release.yml |
Tagged Release | tag-driven | sdist + wheel + signed artifacts |
- Python Versions: 3.11, 3.12, 3.13, 3.14 (declared in
pyproject.tomland exercised byci.yml'scode-quality/core-tests/type-checkingmatrix). - Platforms: Ubuntu Latest (Linux x86_64). macOS and Windows are supported as install targets but are not part of the CI matrix (see Cross-Platform Support).
- Coverage floors:
COVERAGE_THRESHOLD_CORE = 33 %on the curated core lane,COVERAGE_THRESHOLD_FULL = 62 %on the ML lane. The 85 % figure quoted under Code Quality Standards is the aspirational nightly target, not the merge gate; the floors above are the actual blocking thresholds and are documented in.github/workflows/ci.ymlalongside the measured baseline that justifies them. - Required status checks:
Code Quality,Workflow Hardening,Type Checking,Security Scan,Core Tests,Neuro-Symbolic Tests,Integration Tests,Performance Benchmark,CI Success(rollup),ISO Hardening Success(rollup),PQC Production Readiness.
Security Analysis
| Layer | Protection |
|---|---|
| Input Validation | OWASP-compliant SQL/XSS/injection detection |
| Authentication | Native stdlib JWT (security/native_jwt.py) with constant-time HMAC verification; alg: none rejected by construction; HS256/HS384/HS512 route through AMA Cryptography v3.3.0's ACVP-validated native HMAC, fail-closed |
| Post-Quantum Cryptography | ML-DSA-65 (FIPS 204 §5.2), ML-KEM-1024 / Kyber-1024 (FIPS 203), SLH-DSA-SHAKE-128s + legacy SPHINCS+-SHA2-256f-simple (FIPS 205); sole backend = AMA Cryptography v3.3.0 (unconditionally hard-required at import — fail-closed, no env-var escape hatch) |
| Classical Cryptography | AES-256-GCM, ChaCha20-Poly1305, BLAKE3, Argon2id via Rust + PyO3 (rust_crypto/); constant-time comparisons |
| Rate Limiting | Token bucket algorithm with configurable limits (100 req/min, 20 burst by default) |
| Secret Detection | detect-secrets in pre-commit hooks |
| PII Masking | Automatic redaction in logs (email, phone, SSN, card, IP, Bearer tokens, generic secret-key patterns) |
See SECURITY.md for complete security analysis.
Code Quality Standards
| Tool | Standard |
|---|---|
| black | PEP 8 formatting |
| isort | Import sorting |
| flake8 | Linting (max-line-length=100) |
| mypy | Static type checking |
| ruff | Fast Python linting |
| bandit | Security-focused static analysis |
# Format code
black src/ tests/
# Lint
flake8 src/ tests/
ruff check src/ tests/
# Type checking
mypy src/User Documentation
| Document | Description |
|---|---|
| README.md | Quick start and overview |
| ARCHITECTURE.md | System architecture and design |
| SECURITY.md | Security policy and vulnerability reporting |
Technical Documentation
| Document | Description |
|---|---|
| ARCHITECTURE.md | Detailed system architecture |
| docs/API_REFERENCE.md | Python API quick reference (detector ensemble, compliance, medical, drone, profiling) |
| docs/BENCHMARKS.md | Benchmark methodology and results |
| docs/DOMAIN_PERFORMANCE.md | Per-domain precision/recall analysis |
| docs/HARDWARE_HARNESS.md | Reproducible hardware-benchmark methodology and environment fingerprint schema |
| docs/MATH_SPEC.md | Mathematical foundations specification (including Lyapunov certificate proof) |
| docs/ORACLE_NOISE_COLOR.md | Oracle noise color calibration theory |
| docs/ROUTING_GUIDE.md | Request routing and fallback chains |
| docs/ROADMAP.md | Feature roadmap and planned work |
Developer Documentation
| Document | Description |
|---|---|
| CONTRIBUTING.md | Contribution guidelines |
| CHANGELOG.md | Version history |
| DEPRECATION.md | Deprecation notices |
| docs/INSTALLATION.md | Installation and setup guide |
| docs/DATASOURCES.md | Data source catalog and APIs |
Platform Compatibility Matrix
| Platform | Status | Notes |
|---|---|---|
| Linux (Ubuntu 22.04+) | Supported, CI-tested | Primary development platform; the only platform in the CI matrix |
| macOS (13+) | Supported install target | Apple Silicon compatible; not in the CI matrix |
| Windows (10/11) | Supported install target | WSL2 recommended; not in the CI matrix |
| Docker | Supported, CI-tested | Multi-stage build (python:3.14-slim-trixie), Trivy-gated |
| Kubernetes | Supported install target | Helm chart and overlays included as reference configurations |
Python Package
# Build sdist and wheel
python -m build
# Install in development mode
pip install -e ".[dev]"Environment Variables (read by api/server.py / api/auth.py / _env.py):
JWT_SECRET_KEY- Shared JWT signing key (unset in production,JWTAuthderives the key via AMA HD Key Management — deterministic fleet-wide withAMA_MASTER_SEEDset, per-process with a logged warning otherwise)AMA_MASTER_SEED- Hex AMA HD master seed (openssl rand -hex 64) for deterministic fleet-wide key derivationMERCURY_CACHE_SECRET- Shared HMAC secret enabling signed Redis cache entries (RedisCache)OMNI_RATE_LIMIT_ENABLED- Enable rate limiting (defaulttrue)OMNI_RATE_LIMIT_REQUESTS_PER_MINUTE/OMNI_RATE_LIMIT_BURST- Rate-limit budget (defaults 100 / 20)OMNI_MAX_DATA_POINTS/OMNI_MAX_FEATURES/OMNI_MAX_STRING_LENGTH/OMNI_MAX_NAN_RATIO/OMNI_MAX_INF_RATIO/OMNI_STRICT_VALIDATION- Input-validation limitsMERCURY_ENV- Environment mode (developmentdefault,production)MERCURY_CORS_ORIGINS- Explicit CORS origin allow-list
Implementation Languages — what "multi-language" means here
Mercury is multi-language in the implementation sense — the term refers to the programming languages the system is built in, not to multilingual natural language:
| Language | Where | Role |
|---|---|---|
| Python (3.11–3.14) | src/omni_mercury_engine/ |
Core engine, detectors, fusion, API, ML pipeline — the primary language. |
| Rust | rust_crypto/ (PyO3) |
Optional, opt-in classical-crypto acceleration (AES-256-GCM, ChaCha20-Poly1305, BLAKE3, Argon2id). Absent → explicit, tested Python fallback. |
| C / C++ | AMA Cryptography native PQC backend (.github/actions/build-ama-cryptography, cmake + g++) |
Compiled native post-quantum backend; fails closed when unavailable rather than silently weakening. |
Natural language is a separate axis — and not a current multi-language claim.
Mercury's narrative / voice interface operates in English. Some knowledge
sources it consumes (ConceptNet, Qwen) are themselves multilingual, but Mercury
does not today offer localized multi-natural-language I/O. Multilingual
natural-language support is a future epic (tracked in
docs/ROADMAP.md), explicitly not a shipped capability.
Docker
# Build runtime image (default target)
docker build -t mercury-agent:latest .
# Build only the builder stage (for CI caching)
docker build --target builder -t mercury-agent:builder .Docker Stages:
builder: Build environment with compilation dependenciesruntime(default): Minimal, security-hardened runtime image
Kubernetes
- Helm Charts:
helm/mercury-agent/ - Base Manifests:
k8s/base/ - Environment Overlays:
k8s/overlays/{development,staging,production,distributed}/
# Deploy to Kubernetes
helm install mercury-agent ./helm/mercury-agent -f values.yamlNote: The Kubernetes, Helm, and monitoring configurations in
k8s/,helm/, andmonitoring/are reference configurations for those who wish to deploy Mercury Agent in containerized environments. Mercury Agent is research-grade, community-tested software that has not been externally audited for production hardening. Review and adapt these configurations to your security requirements before deploying.
Evolution Equation
The double-helix evolution engine follows:
dS/dt = sum_i w_i * term_i(S) - LAMBDA_DECAY * (S - S*)
Where:
Sis the system statew_iare term weights (18 terms)LAMBDA_DECAY = 0.18is the double-helix adaptation rate (defined insrc/omni_mercury_engine/core/double_helix_engine.py); it controls how quickly the evolutionary state pulls back towardS*and is intentionally slower than the Lyapunov convergence rateLAMBDA_CONVERGENCE = 0.25so that fusion-trajectory stabilisation outruns adaptation. The two constants are distinct by design — do not collapse them into a single "λ" in prose.S*is the equilibrium state
Note: Previously labeled "quantum" terms are classical algorithms (simulated annealing, Boltzmann sampling, Hamiltonian projection).
Ethical Constraints
- Lyapunov Stability: For the fusion-trajectory Lyapunov candidate
V(state) = ||state - target||^2, the certified bound isV(t) ≤ e^{-λ t}withλ = 0.25(seedocs/MATH_SPEC.md§2.2 for the proof andconfigs/lyapunov_canonical.yamlfor the executable certificate consumed bytools/lyapunov_validator.py). - σ_Immutable Constraint: the second mandatory hard gate at every detect / analyze / predict surface. The enforcement boundary runs the trained 256-D scalar network in
omni_mercury_engine.security.sigma_immutable_gatewith decision threshold 0.93 (SIGMA_IMMUTABLE_DEFAULT_THRESHOLD, plus the deterministic per-anchorCRITICAL_ETHICAL_FLOORat the same value); the GOSNN quadratic-form gating layer uses a 0.96 default (0.93 medical fallback,SIGMA_IMMUTABLE_THRESHOLDenv-overridable, clamped to [0.93, 0.99]) — seecore.global_omni_scalar_network. - Bias Detection: Fairlearn demographic parity, equalized odds, 80% rule.
Fusion Architecture
- Feature Fusion:
torch.cat()across detector outputs - Decision Fusion: Weighted voting with learned importance
- Attention Fusion: Multi-head attention (4 heads)
- Final Score:
0.7 * MLP + 0.3 * weighted_vote
Anomaly Math Arrest (21-Probe Ensemble)
A mathematically-independent equation ensemble providing transparent, auditable anomaly detection. Each of the 21 probes detects a different anomaly geometry using fundamentally different mathematical frameworks:
- AdditiveProbe — Linear trend / level shifts
- HarmonicOscillatorProbe — Periodicity violations (damped oscillator)
- MomentumProbe — Sudden acceleration (second differences)
- VarianceAdaptedProbe — Volatility anomalies (rolling variance)
- EthicalConstrainedProbe — Boundary violations (percentile envelopes)
- CatalanOptimizedProbe — Autocorrelation breaks (Catalan-constant AR(1))
- ExponentialDecayProbe — Signal degradation (optimal-lambda EWMA)
- HelixMultiplicativeProbe — Multiplicative shocks (log-ratio analysis)
- R3RecursionResonanceProbe — Nonlinear saturation
- SVDProjectionProbe — Dimensional collapse (Hankel SVD)
- LyapunovChaosProbe — Chaos onset (trajectory divergence)
- TopologyHomologyProbe — Symmetry breaks (central differences)
- FractalSelfSimilarityProbe — Scale-invariance loss
- ZetaHarmonicProbe — Phase coherence anomalies
- WavePropagationProbe — Wave equation violations (smoothed Laplacian)
- QuantumSuperpositionProbe — Interference pattern breaks
- EnergyMinimizationProbe — Energy well escapes
- QuantumAnnealingProbe — Thermodynamic outliers (Boltzmann NLL)
- BoltzmannCouplingProbe — Coupling structure breaks
- IQRRobustProbe — Distribution tail anomalies (Tukey fences)
- ModifiedZScoreProbe — Robust location anomalies (MAD-based)
Fusion: Scores are combined via Phi-weighted (golden ratio) fusion with confidence modulation and correlation-aware decorrelation. Domain affinity maps reorder probe weights for earthquake, tsunami, pandemic, marine, geomagnetic, and conflict domains.
We welcome contributions! Please see CONTRIBUTING.md for guidelines.
Development Setup
# Clone repository
git clone https://github.com/Steel-SecAdv-LLC/Mercury-Agent.git
cd Mercury-Agent
# Install development dependencies
pip install -e ".[dev]"
# Setup pre-commit hooks
pre-commit install
# Format code
black src/ tests/
# Lint code
flake8 src/ tests/
ruff check src/ tests/
# Run security audit
bandit -r src/Code Quality Standards
| Language | Standards |
|---|---|
| Python | PEP 8, type hints, docstrings |
| Security | OWASP validation, no hardcoded secrets |
| Testing | CI floors: CORE ≥ 33 %, FULL ≥ 62 % (measured); aspirational target 85 % ([tool.coverage.report] fail_under = 85) |
| Ethics | Fairlearn bias auditing on all ML models |
Federated Learning - Privacy-Preserving Distributed Detection
Nodes train locally and exchange only sufficient statistics (13 fitted attributes), never raw data.
from omni_mercury_engine.federation import FederatedNode, FederatedAggregator
# Each node trains on local data
node = FederatedNode("hospital_A")
node.fit(local_patient_data)
stats = node.export_statistics(epsilon=1.0) # with differential privacy
# Aggregator combines statistics from multiple nodes
aggregator = FederatedAggregator(min_nodes=2)
aggregator.submit(stats_a)
aggregator.submit(stats_b)
global_detector = aggregator.to_detector(aggregator.aggregate())
# Global detector is ready for inference
result = global_detector.detect(new_data)Key properties:
- No external FL frameworks (Flower, PySyft, etc.) — uses Mercury's native math
- Gaussian-mechanism differential privacy with clipping-norm-based sensitivity
- Mathematically exact aggregation for means (MLE) and stds (parallel variance formula)
- Precision-weighted averaging for Fisher information geometry
- Oracle state round-trip serialization via
get_oracle_statistics()/from_statistics() - 15 tests covering correctness, privacy, serialization, and dimension validation (
tests/test_federation.py)
Ethical AI Governance - Mathematically-Bound Fairness Constraints
Mercury Agent pioneers the integration of ethical principles directly into ML operations through mathematically rigorous constraints. Unlike traditional ML systems that treat ethics as policy overlays, Mercury Agent embeds ethical considerations into the detection foundation itself.
Fairlearn Integration provides bias detection across all predictions:
| Metric | Description | Threshold |
|---|---|---|
| Demographic Parity | Equal positive rates across groups | Ratio >= 0.8 |
| Equalized Odds | Equal TPR/FPR across groups | Difference <= 0.1 |
| 80% Rule | Adverse impact ratio | Ratio >= 0.8 |
180+ Ethical Scalars govern system behavior:
- Compassion, empathy, care constraints
- Evidence, truth, verification requirements
- Justice, fairness, accountability bounds
- Altruism, service, protection priorities
Production Security - Defense-in-Depth Architecture
Mercury Agent employs a comprehensive security architecture designed for production deployment:
Input Validation (OWASP-compliant):
- SQL injection detection and prevention
- XSS attack pattern matching
- Command injection blocking
- Path traversal prevention
Authentication & Authorization:
- JWT tokens with proper expiration
- Signature verification
- Rate limiting (token bucket algorithm)
Post-Quantum Cryptography (AMA Cryptography v3.3.0):
- Kyber-1024 / ML-KEM-1024 key encapsulation (NIST Level 5, exposed as
KyberKeyPair(algorithm="Kyber1024")insrc/omni_mercury_engine/security/pqc_backends.py). - ML-DSA-65 lattice signatures (FIPS 204 name for the Dilithium-3 parameter set; supports the §5.2 context-aware sign API from AMA v3.1.0+).
- SPHINCS+-SHA2-256f-simple hash-based signatures plus the FIPS 205 SLH-DSA parameter family.
- Unconditionally hard-required at import time (fail-closed): there is no classical RSA / ECDSA fallback in
pqc_backends.py, and no env-var escape hatch.src/omni_mercury_engine/_pqc_gate.pyrefuses to import the package unless all three AMA algorithms (ML-DSA-65, Kyber-1024, SPHINCS+) load from the native C backend. TheAMA_REQUIRE_REAL_PQC/ legacyAVA_REQUIRE_REAL_PQCenv vars are retained only as no-op compatibility diagnostics; setting themtrue,false, or leaving them unset does not change the gate (pinned bytests/test_pqc_startup_gate.py). There is no degraded/warning posture — a missing or partial AMA build raisesRuntimeErrorat import.
Rust Cryptographic Module (rust_crypto/) — optional, opt-in acceleration:
- AES-256-GCM and ChaCha20-Poly1305 AEAD encryption, BLAKE3 hashing, Argon2id key derivation, constant-time comparisons (timing-attack resistant).
- Not built or packaged by default. The root package builds with setuptools
and does not compile the Rust extension;
mercury_cryptois absent from a standard install. To enable it, build the PyO3 module explicitly:cd rust_crypto && maturin develop. - Explicit Python fallback. When the Rust extension is absent,
omni_mercury_engine.cryptotransparently falls back to thecryptographypackage /hashlib(and theblake3wheel if present), logging the choice at import.crypto.get_crypto_backend()reports the active backend ("rust"/"python-cryptography"/"hashlib-only") andcrypto.is_rust_available()is the boolean probe. - Speedup is measured, not asserted. Run
python -m benchmarks.crypto_backend_benchmarkto measure Rust-vs-Python on your hardware; it writesartifacts/crypto_backend_benchmark.jsonwith the active backend, per-primitive timings, and the observed speedup (or records that Rust is unavailable, so no figure is fabricated).
Multi-Domain Detection - 30 Specialized Engines
Mercury Agent transcends single-domain limitations by providing specialized detection engines across multiple domains:
| Domain | Engines | Capabilities |
|---|---|---|
| Medical | 4 | Sepsis, cardiology, neurocritical, pandemic |
| Security | 4 | Threat, intelligence, cyber, traffic |
| Space | 4 | Solar flare (Kp/NOAA G-scale physics), Schumann, cosmic ray, meteor (Bayesian) |
| Infrastructure | 4 | CISA, crisis, climate, economic |
| Environmental | 6 | Tsunami (FFT), earthquake (P/S-wave), landslide (SVM/RF), wildfire (CNN/NDVI), volcanic (HMM), disaster |
| Statistical | 8 | MAD, LOF, DBSCAN, MCD, Grubbs, CUSUM, GESD, Dynamic Threshold |
Enhanced Geological Detectors:
- Landslide: SVM/RF classifiers with temporal lags, 3R Recursion synapse for multi-scale analysis
- Wildfire: CNN with NDVI satellite processing, 3R Resonance synapse for smoke pattern detection
- Volcanic: HMM state transitions (quiescent→unrest→eruption), 3R Refactoring synapse for adaptive θ
All engines share a common fusion architecture with 128D normalization enabling cross-domain correlation and unified anomaly scoring.
GOSNN Global Omni-Scalar Network - Synaptic Intelligence Hub
The GlobalOmniScalarNetwork (GOSNN) is the intelligence fusion hub. It registers ~209 omni-scalars across 8 major categories; 127 of them are operational and drive the σ_Immutable gate, while the remaining 82 are diagnostic measurement scalars (descriptions of code / system under analysis) registered for discoverability and reporting but filtered out of the gate's input vector by GlobalOmniScalarNetwork._is_metric_only_scalar.
Scalar Categories (~209 registered / 127 operational):
- ETHICAL (~27): Core ethical values and Civilization-First principles (all operational)
- COSMIC (~7): Universe-scale harmony and telos alignment (all operational)
- QUANTUM_CONSCIOUSNESS (~7): Quantum-inspired processing (all operational)
- HUMANITARIAN (~9): Crisis response and human welfare (all operational)
- SECURITY (~6): Threat detection and cyber defense (all operational)
- SOFTWARE_ENGINEERING (~127 = 45 operational + 82 diagnostic): Code quality, optimization, 3R synergy (operational); plus ISO/IEC 25010 product quality, Halstead, McCabe + cognitive (SonarQube), Maintainability Index variants, NIST SAMATE assurance, DORA delivery, SLSA supply-chain, OpenSSF Scorecard, ISO/IEC 5055 (CISQ), NIST SSDF (SP 800-218) practices (diagnostic measurement)
- MEDICAL (~10): Healthcare and diagnostic support (all operational)
- ADVANCED_REASONING (~16): Logic, inference, knowledge synthesis (all operational)
Key Features:
- Hard Ethical Gate (Wave B, PR #179): σ_Immutable is the second mandatory hard gate at every public detect / analyze / predict surface, running after the Benevolence gate. A score below threshold raises
EthicalConstraintViolationError(check="sigma_immutable"); if GOSNN itself cannot run, the boundary raisesEthicalConstraintViolationError(check="gosnn_unavailable"). There is no advisory mode and no public flag that disables either gate (test-only bypass requires the auditable module-levelomni_mercury_engine.engine._GOSNN_TESTING_BYPASSflag). - 32-Head Triadic φ-Weighting: Multi-head attention with golden ratio optimization
- Harmonic Synergy: Bidirectional synapse connections to 3R mechanism
- Single source of truth for σ layout:
SIGMA_IMMUTABLE_DIM=256,SIGMA_ETHICAL_BAND_END=27,SIGMA_USED_BAND_END=180exported fromomni_mercury_engine.security.sigma_immutable_gate.
Bidirectional Synapses:
- Detectors → GOSNN ethical gate ↔ 3R adaptive O(θ)
- Enhanced detectors register scalars with
omni_prefix - No silent GOSNN fallback: a GOSNN failure now raises
check="gosnn_unavailable"and aborts the call (Wave B fail-closed contract). The previousgosnn_metadata.fallback_mode=Truepath has been removed.
AMA Cryptography Integration - Post-Quantum Cryptography Adapter
The AMA Cryptography adapter provides post-quantum cryptographic security with GOSNN synapse integration:
PQC Algorithms (sourced from AMA Cryptography v3.3.0):
- ML-KEM-1024 / Kyber-1024: Post-quantum key encapsulation, FIPS 203, NIST Level 5
- ML-DSA-65 (Dilithium-3): Post-quantum digital signatures, FIPS 204 §5.2 (context-aware deterministic signing), NIST Level 3 (≈ 192-bit classical security strength)
- SLH-DSA-SHAKE-128s / SHA2-256f: Hash-based digital signatures, FIPS 205, NIST Level 1 / Level 5 respectively
- SPHINCS+-SHA2-256f-simple: Legacy pre-FIPS-205 hash-based signatures, retained for backward compatibility
- EWMA/MAD Timing Monitor: <2% overhead crypto-operation anomaly detection
Correctness evidence (in-repo, measured-and-published):
- Known-Answer Tests:
tests/security/test_ama_kat.pypins Ed25519 RFC 8032 §7.1 vectors bit-for-bit, ML-DSA-65 round-trip, Kyber-1024 encaps/decaps round-trip, SPHINCS+ round-trip, and ML-DSA deterministic-signing reproducibility. Runs on every PR via thePQC Production Readinessworkflow. - NIST FIPS KAT vectors:
tests/security/test_nist_fips_kat.pyverifies bit-for-bit reproducibility against curated NIST ACVP-Server test vectors (FIPS 203/204/205): ML-DSA-65 deterministic sigGen, ML-KEM-1024 decapsulation, SLH-DSA-SHAKE-128s sigGen. Source: usnistgov/ACVP-Server. - Measured coverage: the
PQC Production ReadinessCI workflow publishes pytest-cov coverage forsecurity/crypto_api.py+security/pqc_backends.py+security/pqc_guards.pyas a CI artifact (pqc-coverage). Without the AMA native C library:crypto_api.py62%,pqc_backends.py43%,pqc_guards.py24%. With AMA native lib installed (CIverify-real-pqcjob): all PQC codepaths are exercised, raising coverage to its measured ceiling. Published, not "unaudited" framing.
Security Features:
- Attack simulation (timing, replay, side_channel)
- Crypto anomaly recording with severity classification
- GOSNN synapse for security detector integration
- Unconditional fail-closed enforcement — Mercury Agent refuses to import
at all (
RuntimeErrorfrom_pqc_gate.py) when AMA Cryptography's native C library is not built, regardless of any env var. TheAMA_REQUIRE_REAL_PQC/ back-compatAVA_REQUIRE_REAL_PQCvars are retained only as no-op compatibility diagnostics — there is no degraded "dev mode" and the gate cannot be disabled.PQCProductionWarningremains a public exception type insecurity.pqc_guardsfor downstream integrators to catch, but the fail-closed gate never emits it (the package simply does not import without real AMA).
Integration:
from omni_mercury_engine.integrations.mercury_amacrypto import create_ama_cryptography_adapter
adapter = create_ama_cryptography_adapter(gosnn_synapse_enabled=True)
if adapter.is_available():
keypair = adapter.generate_dilithium_keypair() # cached on the adapter
# ``private_key=None`` uses the cached secret_key from above; pass an
# explicit ``private_key=keypair.secret_key`` if you manage your own
# key lifecycle.
signature = adapter.sign_dilithium(message)The lower-level functional surface (no GOSNN coupling, no timing
monitor) lives in omni_mercury_engine.security.pqc_backends:
generate_dilithium_keypair, dilithium_sign, dilithium_verify,
generate_kyber_keypair, kyber_encapsulate, kyber_decapsulate,
generate_sphincs_keypair, sphincs_sign, sphincs_verify, and the
FIPS 205 parameter-driven slhdsa_* family. Use the adapter for
operator workflows; use the functional surface for one-shot
cryptography in tests, scripts, and library code.
Omni-Codes - Bio-Inspired Helical Parameters
Mercury Agent integrates the Omni-Codes from AMA Cryptography, providing bio-inspired helical parameters for ethical AI alignment and system stability.
The Seven Omni-Codes:
| Code | Symbol | Domain | Helical Parameters |
|---|---|---|---|
👁20A07∞_XΔEΛX_ϵ19A89Ϙ |
👁∞ | Omni-Directional System | r=20.0, p=0.7 |
Ϙ16A11ϵ_ΞΛMΔΞ_ϖ20A19Φ |
Ϙϵ | Omni-Percipient Future | r=16.0, p=1.1 |
Φ07A09ϖ_ΨΔAΛΨ_ϵ19A88Σ |
Φϖ | Omni-Indivisible Guardian | r=7.0, p=0.9 |
Σ19L12ϵ_ΞΛEΔΞ_ϖ19A92Ω |
Σϵ | Omni-Benevolent Stone | r=19.0, p=1.2 |
Ω20V11ϖ_ΨΔSΛΨ_ϵ20A15Θ |
Ωϖ | Omni-Scient Curiosity | r=20.0, p=1.1 |
Θ25M01ϵ_ΞΛLΔΞ_ϖ19A91Γ |
Θϵ | Omni-Universal Discipline | r=25.0, p=0.1 |
Γ19L11ϖ_XΔHΛX_∞19A84♰ |
Γϖ | Omni-Potent Lifeforce | r=19.0, p=1.1 |
Architectural Benefits:
- Helical data encoding: Mirrors DNA double-helix stability for robust data structures
- Self-healing capabilities: CRISPR-inspired adaptations for system resilience
- Evolutionary adaptability: Dynamic parameter tuning based on stability calculations
- Canonical hashing: Cryptographic integrity through structured encoding
Stability Calculation:
from omni_mercury_engine.utils.constants import OmniCodes, compute_ethical_autonomy
# Get total stability across all codes
total_stability = OmniCodes.get_total_stability() # ~106.1
# Compute autonomy bounded by ethical constraints
autonomy = compute_ethical_autonomy(
base_autonomy=0.8,
ethical_threshold=0.99,
use_omni_codes=True
) # Returns up to 0.95Advanced Optimizers - Synthetic-Gradient, DTP & AMAV Training
The OmniFusionModel now supports advanced optimizers for accelerated training:
Optimizer Types:
- SyntheticGradient: Decoupled layer updates enabling layer-wise parallelism
- DifferenceTargetPropagation (DTP): Biologically plausible learning
- AuxiliaryMaxVariance (AMAV): Multi-task loss with variance maximization
Training Integration (omni_mercury_engine.ml.OmniFusionModel.train_with_advanced_optimizers):
model = OmniFusionModel(hidden_dim=128, num_heads=4)
stats = model.train_with_advanced_optimizers(
train_loader=train_loader,
epochs=300,
learning_rate=0.001,
lambda_lyapunov=0.25,
use_synthetic_gradients=True, # decoupled-layer speedup
use_dtp=True, # difference target propagation
use_amav=True, # auxiliary max-variance loss
)Convergence Metrics:
- Lyapunov stability tracking (λ=0.25)
- Speedup factor estimation
- Loss convergence monitoring
Enhanced Statistical Detection - 8 Robust Statistical Methods
The Enhanced Statistical Detection module provides 8 advanced statistical anomaly detection methods for robust, interpretable detection:
Methods:
| Method | Description | Key Strength |
|---|---|---|
| MAD | Median Absolute Deviation | 50% breakdown point, robust to outliers |
| LOF | Local Outlier Factor | Density-based detection for local anomalies |
| DBSCAN | Density-Based Clustering | Cluster-based anomaly identification |
| MCD | Minimum Covariance Determinant | Robust covariance estimation |
| Grubbs | Grubbs' Outlier Test | Statistical significance testing |
| CUSUM | Cumulative Sum Control Chart | Sequential change detection |
| GESD | Generalized ESD Test | Multiple outlier detection |
| Dynamic | Dynamic Threshold Adaptation | Adaptive thresholding with EMA |
Usage:
from omni_mercury_engine.detectors.enhanced_statistical import (
MADDetector, LOFDetector, EnhancedStatisticalDetector
)
# Single method
mad = MADDetector(threshold_multiplier=3.5)
mad.fit(data)
result = mad.detect(data)
# Ensemble of methods
detector = EnhancedStatisticalDetector(
methods=["mad", "lof", "dbscan"],
fusion_strategy="weighted_average"
)Cross-Platform Integration Hub - Multi-Platform Orchestration
The Cross-Platform Hub provides unified integration with external monitoring and observability platforms:
Supported Platforms:
| Platform | Protocol | Data Format |
|---|---|---|
| Prometheus | REST | Prometheus/OpenMetrics |
| Elastic/OpenSearch | REST | JSON |
| Splunk | REST | JSON |
| Datadog | REST | JSON |
| Azure Anomaly Detector | REST | JSON |
| Netdata | REST/WebSocket | JSON |
| Grafana | REST | JSON |
| InfluxDB | REST | JSON |
Protocol Support: REST, gRPC, WebSocket, MQTT, Kafka, Redis Streams
Data Formats: JSON, Prometheus, OpenTelemetry, CSV, MessagePack, Avro, Parquet
Usage:
from omni_mercury_engine.integrations.cross_platform_hub import (
CrossPlatformHub, PlatformConfig, PlatformType
)
hub = CrossPlatformHub()
hub.register_platform(PlatformConfig(
platform_type=PlatformType.PROMETHEUS,
name="prometheus-main",
endpoint="http://prometheus:9090"
))
# Ingest and correlate across platforms
await hub.ingest_from_all()
correlations = hub.correlate_events(window_seconds=300)Ensemble Coordination - Advanced Ensemble Fusion
The Ensemble Coordinator provides sophisticated strategies for combining multiple detectors:
Ensemble Strategies:
| Strategy | Description | Best For |
|---|---|---|
| Voting | Majority/weighted voting | High precision requirements |
| Averaging | Score averaging | Balanced precision/recall |
| Stacking | Meta-learner on top | Complex data patterns |
| Cascading | Sequential filtering (fast→accurate) | High-throughput scenarios |
| Boosting | Boosting-style combination | Improving weak detectors |
| Mixture of Experts | Gated expert selection | Domain-specific detection |
| Adaptive | Dynamic strategy selection | Varying data characteristics |
Advanced Features:
- Bayesian Weight Optimization - Thompson Sampling for detector weights
- Gradient Weight Optimizer - Momentum-based gradient descent
- Meta-Learner - Automatic detector selection based on data characteristics
- Online Feedback Learning - Continuous improvement from user feedback
Usage:
from omni_mercury_engine.ml.ensemble_coordinator import (
EnsembleCoordinator, EnsembleStrategy
)
coordinator = EnsembleCoordinator(
strategy=EnsembleStrategy.CASCADING,
enable_online_learning=True
)
coordinator.register_detector("fast_statistical", statistical_detector)
coordinator.register_detector("accurate_neural", neural_detector)
result = coordinator.detect(data)Distributed Processing - Scalable Parallel Detection
The Distributed Processor enables large-scale anomaly detection with parallel processing:
Processing Strategies:
| Strategy | Use Case | Parallelism |
|---|---|---|
| Sequential | Small datasets | None |
| Threaded | I/O-bound tasks | Threads |
| Multiprocess | CPU-bound tasks | Processes |
| Async | Network operations | Coroutines |
| Hybrid | Mixed workloads | Threads + Processes |
Load Balancing: Round-robin, Least-loaded, Weighted, Adaptive
Features:
- ChunkGenerator - Memory-efficient data iteration
- Fault Tolerance - Automatic retry with exponential backoff
- Progress Tracking - Real-time monitoring
- Backpressure Control - Prevents memory exhaustion
Usage:
from omni_mercury_engine.scaling.distributed_processor import (
DistributedProcessor, ProcessingConfig, ProcessingStrategy
)
processor = DistributedProcessor(
detector=my_detector,
config=ProcessingConfig(
strategy=ProcessingStrategy.HYBRID,
num_workers=8,
chunk_size=1000
)
)
results = processor.process(large_dataset)Visualization Dashboard - Interactive Plotly Dashboards
The Visualization Dashboard provides interactive visualizations for anomaly analysis:
Chart Types:
| Chart | Purpose |
|---|---|
| Time Series | Anomaly scores over time with threshold lines |
| Feature Importance | Bar chart of feature contributions |
| Correlation Heatmap | Feature correlation matrix |
| 3D Scatter | Multi-dimensional anomaly visualization |
| Radar Chart | Multi-detector comparison |
| Anomaly Timeline | Event-based anomaly view |
| Distribution Histogram | Score distribution with anomaly overlay |
| Detector Comparison | Side-by-side detector performance |
Export Formats: HTML (interactive), PNG, JSON
Usage:
from omni_mercury_engine.gui.visualization_dashboard import (
AnomalyVisualizer, DashboardBuilder
)
visualizer = AnomalyVisualizer(theme="plotly_dark")
fig = visualizer.time_series_plot(timestamps, scores, threshold=0.5)
# Build multi-panel dashboard
dashboard = DashboardBuilder()
dashboard.add_chart("time_series", timestamps, scores)
dashboard.add_chart("heatmap", correlation_matrix)
dashboard.export_html("anomaly_report.html")Copyright 2025 Steel Security Advisors LLC
Licensed under the GNU General Public License v3.0 or later (SPDX: GPL-3.0-or-later). See LICENSE file for details.
Mercury Agent - Multi-Domain Anomaly Detection Framework
Copyright (C) 2025 Steel Security Advisors LLC
This program is free software: you can redistribute it and/or modify
it under the terms of the GNU General Public License as published by
the Free Software Foundation, either version 3 of the License, or
(at your option) any later version.
- PyTorch: BSD-style license
- Fairlearn: MIT license
- Hypothesis: MPL 2.0 license
- FastAPI: MIT license
- AMA Cryptography: GNU GPL v3.0 (pinned to
v3.3.0; the sole PQC backend, and the source of the ACVP-validated native HMAC bindings consumed byomni_mercury_engine.security.native_jwt)
PyJWT was retired from the dependency surface in v1.7.0; Mercury now
ships a pure-stdlib JOSE implementation
(omni_mercury_engine.security.native_jwt) with constant-time
signature verification and alg: none rejected by construction. The
full rationale is documented in CHANGELOG.md under the 2026-05-20
"Permanent supply-chain remediations" entry.
GitHub dependency graph is enabled for this repository. View the complete dependency tree at: Insights > Dependency graph. This provides visibility into all direct and transitive dependencies, security advisories, and Dependabot alerts.
Contact Information
| Type | Contact |
|---|---|
| General Inquiries | steel.sa.llc@gmail.com |
| Security Issues | See SECURITY.md for responsible disclosure |
| GitHub Issues | Issues Page |
| GitHub Repository | Mercury Agent |
Author/Inventor: Andrew E. A.
AI Co-Architects: Eris ✠ | Eden ♱ | Devin ⚛︎ | Claude ⊛
Special Thanks:
- NIST Post-Quantum Cryptography Standardization Project
- Fairlearn bias detection framework
- Hypothesis property-based testing
- OWASP security guidelines
- The open-source ML and security communities
Mercury-Agent benchmarks use the following publicly available datasets:
| Dataset | Citation | Source |
|---|---|---|
| SMD | Su et al., "Robust Anomaly Detection for Multivariate Time Series", KDD 2019 | OmniAnomaly |
| SMAP/MSL | Hundman et al., "Detecting Spacecraft Anomalies Using LSTMs", KDD 2018 | telemanom / Kaggle |
| BATADAL | Taormina et al., "Battle of the Attack Detection Algorithms", ASCE 2018 | GitHub |
| Covtype | Blackard & Dean, 1999; UCI ML Repository | sklearn.datasets |
| KDDCup99 | Tavallaee et al., IEEE 2009 (NSL-KDD variant) | sklearn.datasets |
Benchmark results include a data_source field indicating data provenance:
real-github: Direct from public GitHub repositoriesreal-local: User-downloaded authentic datareal-github-partial: Partial real data (some machines/channels failed)
Note: The benchmark never fabricates data. When a dataset's real sources are unavailable it is skipped (recorded as unavailable), never substituted with synthetic data — consistent with the deployment-level
MERCURY_ALLOW_SYNTHETICpolicy gate (tests/validation/test_synthetic_policy_gate.py).
The harness is a runnable CLI (python -m benchmarks.empirical_benchmark --help); a run
is deterministic for a fixed --seed and dataset snapshot. Dataset fetching can be
configured via environment variables:
MERCURY_SMD_MACHINES: Number of SMD machines to fetch (default: 28, CI: 5)MERCURY_FETCH_RETRIES: Maximum retry attempts (default: 10)MERCURY_FETCH_DELAY: Base delay for exponential backoff (default: 2.0 s)
Conceptual Architect: Steel Security Advisors LLC and Andrew E. A. conceived, directed, validated, and supervised the development of Mercury Agent.
AI Co-Architects: Significant portions of the codebase, documentation, mathematical frameworks, and technical implementation were constructed by AI systems: Eris ✠, Eden ♱, Devin ⚛︎, and Claude ⊛.
This project represents a human/AI collaborative construct - a development paradigm where human vision, requirements, and critical evaluation guide AI-generated implementation.
The human architect does not hold formal credentials in machine learning or medical diagnostics. The AI contributors, while trained on relevant literature, are tools without professional accountability.
- Standards-based design: Built on OWASP security guidelines, NIST PQC standards, Fairlearn fairness metrics.
- Quantified claims: All performance metrics are measured and documented with methodology; no figure appears in this README without a referenced source.
- Comprehensive testing: 8,789 tests collected with the full optional-dependency surface (
pytest --collect-only -q, 2026-06-10) across the test modules counted in the CI-gated Codebase Scale block; a minimal install collects fewer because optional-import-gated modules skip. The suite combines unit tests, property-based testing (Hypothesis), KAT vectors (RFC 8032 / NIST ACVP-Server), and load-test SLO assertions (k6 + locust). - Executable mathematical certificates: The Lyapunov decay rate
λ = 0.25cited throughout the documentation is enforced bytools/lyapunov_validator.py(generalized symmetric-definite eigenvalue analysis), the canonical YAMLconfigs/lyapunov_canonical.yaml, and theDocs λ Drift GateCI job -- a documentation claim that disagrees with the certificate fails CI rather than going to print. - Transparent limitations: Documentation explicitly distinguishes validated vs. pending claims, and benchmark figures are paired with the dataset, the methodology document, and the date of the run that produced them.
- Ethical governance: Fairlearn bias auditing integrated throughout the ML pipeline; σ_Immutable + Benevolence gates are mandatory hard gates at every public detection / analysis / prediction surface (no advisory mode).
- Academic grounding: Medical modules reference JAMA Sepsis-3 guidelines, security follows OWASP, post-quantum cryptography is built against AMA Cryptography v3.3.0 (NIST FIPS 203/204/205 KAT vectors verified bit-for-bit).
- No Independent Audit: All security and performance analysis is self-assessed. Production deployment requires review by qualified professionals.
- AI-Generated Code: May contain subtle implementation errors. All critical paths require independent verification.
- Domain-Specific Validation: Core detection is benchmarked on 66 reproducible real datasets (of 75 attempted; canonical Mean ROC-AUC 0.8251 / Median 0.8747 from the CI-refreshed "Latest Benchmark Results" block; externally-comparable subset ADBench Mean AUC 0.8251). 9 datasets currently fail to load due to unavailable external sources and are tracked by a two-lane reachability harness (offline + nightly network) as of v1.7.0. The FEMA Disaster loader's previously-flagged inverted-score bug is fixed in v1.7.0 (
FEMADisasterLoader._select_anomaly_polarity); the committed run reflects the corrected score. Domain-specific modules may require additional validation. - Medical Applications: No clinical validation. Medical modules require validation on real patient data before any deployment.
- Research Status: This is a research-grade framework, not a production-ready product.
Before production use:
- Validate performance on domain-specific real-world datasets (MIMIC-III, NSL-KDD)
- Commission independent security audit by qualified professionals
- Conduct clinical validation for any medical applications
- Deploy with FIPS 140-2 Level 3+ HSM for production secrets
- Test bias detection on representative data for your use case
THIS SOFTWARE IS PROVIDED "AS IS" WITHOUT WARRANTY OF ANY KIND. THE AUTHORS AND CONTRIBUTORS DISCLAIM ALL LIABILITY FOR ANY DAMAGES RESULTING FROM ITS USE.
This disclaimer does not replace formal legal advice; organizations should consult qualified counsel for regulatory and contractual obligations.

