If you are an AI agent (Claude / Codex / Gemini / a sub-agent / a script in a sandbox) and you need to read or write memhall, this is the doc for you.
The README's "Three entry points" lists the surfaces. This doc is the decision tree: which surface you should actually pick, and the gotchas each one has.
Status legend (last verified 2026-04-28 against
fix/reliability-phase-a5-2026-04-27):
- ✅ verified — exercised end-to-end in a real session, including against a server with
MH_API_TOKENset.⚠️ partial — works for the no-auth case, but does not currently work against a server that requiresMH_API_TOKEN.If a path is marked
⚠️ and you need it to work with auth, fall back to a ✅ path until the gap is closed.
Are you running in the same process / repo as memory-hall, with `import memory_hall` available?
├─ Yes → use the embedded Python runtime (Path A)
└─ No
│
Can your sandbox open a TCP socket to the memhall host?
├─ Yes → use HTTP + Bearer (Path B)
└─ No (sandboxed agents: Codex CLI, restricted containers, some Gemini setups)
└─ install the package and use Path A in-process,
or shell out via `mh` CLI which goes through Path A under the hood (Path C)
If you do not know which one applies to you, default to Path B (HTTP + Bearer) — it works from anywhere that has network access and curl.
Status: ✅ verified. Bypasses HTTP + auth entirely (in-process call, no middleware).
Use when: same process, sandboxed environments where TCP is blocked, batch imports, tests.
import asyncio
from memory_hall import Settings, build_runtime
from memory_hall.models import WriteMemoryRequest, SearchMemoryRequest
async def main():
runtime = build_runtime(settings=Settings())
await runtime.start()
try:
await runtime.write_entry(
tenant_id="default",
principal_id="my-agent",
payload=WriteMemoryRequest(
agent_id="my-agent",
namespace="shared",
type="note",
content="hello from inside the process",
),
)
hits = await runtime.search_entries(
tenant_id="default",
payload=SearchMemoryRequest(query="hello", limit=5),
)
print(hits.total)
finally:
await runtime.stop()
asyncio.run(main())No network, no auth, same storage. This is the path Codex / Gemini sandboxes should prefer when localhost TCP is blocked by the sandbox.
Gotchas:
Settings()reads from env (MH_DB_PATH,MH_EMBEDDER_KIND, …). If the agent's working directory has its own.env, runtime config will diverge from the running HTTP server. Point both at the same DB if you want them to share state.build_runtimeis async; you need an event loop. In a sync script, wrap withasyncio.run(...).
Status: ✅ verified against a server with MH_API_TOKEN set. This is the most reliable path when the sandbox has TCP access.
Use when: any language, any tool, sandbox can reach the host over TCP.
# Set once per shell. Maki's setup keeps the token at ~/.config/memhall/token (0600).
export MH_API_TOKEN="$(cat ~/.config/memhall/token)"
curl -sS http://127.0.0.1:9000/v1/memory/write \
-H "Authorization: Bearer ${MH_API_TOKEN}" \
-H 'Content-Type: application/json' \
-d '{
"agent_id": "my-agent",
"namespace": "shared",
"type": "note",
"content": "hello from curl"
}'Gotchas:
Authorization: Bearer …is required on every/v1/memory/*request when the server hasMH_API_TOKENset./v1/healthis the only public endpoint. Missing the header returns{"detail":"missing bearer token"}— the server is alive, you are just unauthenticated.- If the server runs without
MH_API_TOKENset (dev / standalone), the header is ignored. Sending it anyway is safe and forward-compatible — always send it. - Default port is
9000. Maki's home deployment maps it to9100(http://100.122.171.74:9100). Check the deployment you are talking to. /v1/admin/*requiresMH_ADMIN_TOKEN(a different token). RegularMH_API_TOKENis rejected on admin paths when admin token is set. Seedocs/adr/0007-minimal-token-auth.md.- See
examples/shell/write_memory.shfor a runnable starter.
Status: ✅ verified. The CLI reads MH_API_TOKEN from the environment (via Settings()) and attaches Authorization: Bearer <token> automatically when set. Works against both auth-enabled and no-auth servers. Verified against src/memory_hall/cli/main.py:31 on fix/reliability-phase-a5-2026-04-27; covered by tests/test_cli_auth.py.
Use when: you want a one-liner from a shell, you do not want to hand-roll JSON, and the package is installed.
# One-time install in the project venv:
uv sync
# Then `mh` is on PATH inside the venv.
# If the server has MH_API_TOKEN set, export it (CLI reads it automatically):
export MH_API_TOKEN="$(cat ~/.config/memhall/token)"
uv run mh write "DEC-018 落地完成" \
--agent-id codex \
--namespace project:memory-hall \
--type decision \
--tag governance
uv run mh search "DEC-018"Gotchas:
mhis a console script defined inpyproject.toml. It is not globally available. Ifcommand -v mhreturns nothing, you have not installed the package — runuv sync(orpip install -e .) inside the repo first.uv run mh …works without prior install but resolves dependencies on first use. In sandboxes where~/.cache/uvis not writable, setUV_CACHE_DIR=/tmp/uv-cachebefore calling.- The CLI hits HTTP under the hood.
MH_API_TOKENis read from the environment on each command; no CLI flag is needed. If unset, noAuthorizationheader is sent (works against no-auth servers).
| Symptom | Likely cause | Fix |
|---|---|---|
{"detail":"missing bearer token"} |
Path B without Authorization header |
Set MH_API_TOKEN and add -H "Authorization: Bearer ${MH_API_TOKEN}" |
curl: (7) Couldn't connect to server from a sandboxed agent |
Sandbox blocks localhost TCP | Switch to Path A (embedded Python) |
command not found: mh |
Package not installed in this shell's PATH | uv sync inside the repo, or use uv run mh … |
uv run mh errors on ~/.cache/uv permission |
Sandbox cache dir not writable | export UV_CACHE_DIR=/tmp/uv-cache |
| Writes succeed but search returns nothing | Path A and Path B pointing at different DB files | Align MH_DB_PATH in both, or always go through HTTP |
| Handoff loads an old but relevant session | /v1/memory/search ranks by relevance, not recency |
Use GET /v1/memory?namespace=...&type=episode&limit=... or uv run mh list --namespace ... --type episode for latest handoff |
mh write times out around 5 seconds |
Client timeout shorter than server embed timeout | Upgrade to a version with write timeout max(MH_REQUEST_TIMEOUT_S, MH_EMBED_TIMEOUT_S + 2) or set a larger request timeout |
For session handoff, "latest" and "most relevant" are different operations:
- Use list for normal handoff/load/status flows. It returns entries in
created_at DESCorder. - Use search only when the user asks for a keyword or historical lookup.
Latest handoff via HTTP:
export MH_API_TOKEN="$(cat ~/.config/memhall/token)"
curl -sS "http://100.122.171.74:9100/v1/memory?namespace=project:memory-hall&type=episode&limit=5" \
-H "Authorization: Bearer ${MH_API_TOKEN}"Latest handoff via CLI:
export MH_API_TOKEN="$(cat ~/.config/memhall/token)"
uv run mh list \
--base-url http://100.122.171.74:9100 \
--namespace project:memory-hall \
--type episode \
--limit 5If you do use search for handoff-related keywords, prefer --mode lexical when the embedder is degraded. Hybrid search can still return lexical results with "degraded": true, but the timeout adds noise to operational debugging.
agent_id— stable identity for the agent. Examples:claude,codex,gemini,max,grok,gemma4,maki. Do not invent a new id per session; one id per agent persona.namespace— scope of the entry. Examples:home,work,project:<name>,agent:<id>,shared.type— one ofepisode,decision,observation,experiment,fact,note,question,answer.
Do not write company-sensitive content into shared or work. Use project:<name> or do not write at all.
README.md— full feature list and quickstartdocs/api.md— HTTP endpoint referencedocs/adr/0007-minimal-token-auth.md— why Bearer auth is the way it isexamples/codex_cli/— Codex CLI starterexamples/shell/— curl starter