Long Horizon Terminal Benchmark with Dense Reward Grading
-
Updated
Jul 24, 2026 - Python
Long Horizon Terminal Benchmark with Dense Reward Grading
The local-first intelligence layer that gives AI agents durable continuity, explainable retrieval, portable context, and verified learning across harnesses.
A Plug-and-Play, Cost-Efficient, and Lightweight Memory Plugin for Long-Horizon LLM Agents
Cortex: A Bidirectionally Aligned Embodied Agent Framework for Long-horizon Manipulation
🔁 Build reliable recurring AI-agent systems: 730 resources, 22 operational patterns, 22 loop contracts, 8 runtime starters, an interactive atlas, and a structured dataset.
Official Repository for our paper: PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems
A workflow specification for autonomous agents
Progressive disclosure retrieval for long-horizon AI agents over personal knowledge bases. Three-tier model: hot cache → FTS5 index → full read. Local-first, zero dependencies.
Recoverable long-horizon AI agents — a framework-agnostic reference harness + recovery-faithful live benchmark. Thesis: "Checkpoints Are Compactions" via Re-grounding Recovery. 0.x: v1.0 held until a powered live-LLM study confirms the claims.
Repo-native operating layer for AI agents: memory, planning, proof, routing, and continuity across sessions.
Server-authoritative Minecraft AI agent with a Python LLM brain and Fabric + Carpet/Scarpet body. Autonomous navigation, gathering, crafting, combat, recovery, persistent memory, governed tools, and long-horizon play.
Bounded context for long-horizon LLM agents: carry a small re-grounded signal instead of replaying the transcript, so a parent agent's footprint stays flat as the task grows. Recompute-verifiable; null results published; cheaper, not smarter.
Self-hosted continuity and collaboration substrate for autonomous agents with bounded, recoverable memory and explicit restart orientation
Agent-agnostic autonomous research protocol adapted from Deli AutoResearch, with STORM-style multi-perspective questioning, source-grounded claim ledgers, persistent state, Discovery Tail Pass, and sufficiency checks for long-horizon AI agents.
Adaptive SDD harness for long-horizon coding agents. Lock product intent before agents write code.
Architecture notes on CLI-native autonomous coding agents, tool runtimes, orchestration, UI surfaces, and protocols.
EvoLoop — Cross-platform autonomous AI agent: control computers & Android devices from a single interface, learn from demonstrations, and collaborate across agents. Desktop / web / server / embedded.
Deterministic reliability-floor metrics for long-horizon agents — pass^k, reliability decay curve, meltdown onset. Zero dependencies.
A durable, long-horizon LLM agent harness: write-ahead logging, idempotent tool replay, identifier-preserving compaction, and goal-drift detection.
Reversible, provenance-aware governed memory for long-horizon agents, with controlled simulations, baselines, tests, and auditable checkpoint reopening.
Add a description, image, and links to the long-horizon-agents topic page so that developers can more easily learn about it.
To associate your repository with the long-horizon-agents topic, visit your repo's landing page and select "manage topics."