mrkeyoor.com_
Wed 30 Sept 06:23 UTC
LLM Toolsevaluationupdated 30 Sept 2026

agent-memory review

Agent Memory gives Claude Code, Codex CLI, and other shell-capable agents a shared long-term memory store. Memories stay as readable Markdown files, while a disposable SQLite index handles local search and optional vector retrieval.

Verdict

Our run installed Agent Memory in 27 seconds and built it in 10 seconds, then pytest exited 4 because agent_memory could not be imported. Trial it when readable files, shared Claude Code and Codex recall, and reversible history are worth working from a checkout. Wait for a tagged release and green reproduction if memory correctness is part of a production release gate.

We ran it

Lab card: what happened when we ran agent-memoryScreenshot of agent-memory (github.com/tigerless-labs/agent-memory)
Install✓ · 27s35 packages · 37 MB
Build✓ · 10s
Tests✗ · 12sran, no count parsed
Known vulns0(pip-audit)
Repo149 files~18,969 lines of source · 1 MB · 1 CI workflows · tests dir

Answers from our run

Does agent-memory build from source?

Dependencies installed in 27 seconds (35 packages), and the build succeeded in 10 seconds. We cloned commit 243868d into a clean Debian container with 3 CPUs and no project-specific setup.

Do agent-memory's tests pass?

The test command failed in our container, and its output did not report a pass or fail count.

Does agent-memory have known vulnerabilities in its dependencies?

pip-audit found none in the dependency tree at the time of our run.

Who should not use agent-memory?

Release pipelines that require a green clean-container suite: our pytest step exited 4 because agent_memory could not be imported from tests/conftest.py.

What are the alternatives to agent-memory?

Mem0, Graphiti, MemOS. Our run installed Agent Memory in 27 seconds and built it in 10 seconds, then pytest exited 4 because agent_memory could not be imported.

Setup2/5Install and build passed, but pytest stopped on an import error
Docs5/5Architecture, limits, experiments, hooks, and recovery are explicit
Community3/51,780 stars and active PRs, with less than a month of history
Maturity2/5Version 0.1.0, no release, failed test run, open data defects

Who it’s for

Developers who want Claude Code and Codex to remember the same decisions across sessions.
Teams that prefer inspectable Markdown and Git history over an opaque hosted memory database.
Agent builders who need provenance, corrections, superseded facts, and point-in-time recall.
Python 3.12 users willing to install from a checkout and inspect the hook configuration.

Who it’s NOT for

Release pipelines that require a green clean-container suite: our pytest step exited 4 because agent_memory could not be imported from tests/conftest.py.
Teams that need a stable packaged release: the README says there is no PyPI release, and GitHub has no tagged release.
Workloads with frequent concurrent corrections to the same memory: open issue 16 describes a lost-update race, and its proposed fix remains an open pull request.
Stores built from non-English-only names without explicit slugs: open issue 20 reports that non-ASCII-only input can produce an empty generated name.
MCP clients that need every operation through MCP: the README says context, sleep, and the proposal ledger are available only through the wider CLI.
Users unwilling to review hook changes or let a host CLI process archived conversations in the background.

Setup reality

Our run installed commit 243868d in 27 seconds, adding 35 packages and using 37 MB. The build passed in 10 seconds. Tests failed with exit 4 after 12 seconds because tests/conftest.py could not import agent_memory; pip-audit reported 0 known vulnerabilities.

You need Python 3.12 or newer, uv, and a checkout because there is no PyPI release. The CLI also needs its virtual environment on PATH, a shared AGENT_MEMORY_STORE, and a Claude Code or Codex login if background distillation will borrow that host's model.

mem setup changes the selected host's hooks and Codex asks you to trust new hooks. Optional vector recall downloads a model on first use. The repository has 1 CI workflow and a tests directory, but no Dockerfile, so container packaging is yours to define.

Markdown remains the truth after the SQLite index is deleted

Across 149 files, Agent Memory stores each durable fact or decision as Markdown with frontmatter. MEMORY.md acts as a compact starting index, while SQLite FTS5 and an optional vector layer rank recalls. Delete .index, run mem rebuild, and the knowledge should return from those files. That gives you ordinary file inspection, Git history, and a migration route that does not depend on a private database format.

The file boundary also defines what can expire. A memory has a validity interval, and correction, superseding, merging, or deletion changes which file is current without removing its history. Raw session evidence stays in an append-only archive. mem trace can then follow a memory back to cited message ranges, which is useful when a compact summary has lost a qualification.

Recall starts with 8 paths instead of pasting whole memories

A normal mem recall returns 8 memory records by default, each with an abstract, path, anchor, and score. The agent chooses whether to open an outline, a full memory, or the cited raw material. That staged read is a practical answer to context waste: irrelevant files can stop at one line, while a disputed detail can be checked against its source.

Search has 3 routes. The root MEMORY.md is injected at session start, BM25 ranks the local FTS5 index, and optional vectors join through reciprocal-rank fusion. Plain ls and grep remain available when ranking misses. The README reports that vector fusion raised Recall@5 from 79.0% to 86.6%, while median retrieval latency rose from 5.1 ms to 139.2 ms. Those are project measurements, not ours.

Boundary hooks share one store across Claude Code and Codex

On Python 3.12, mem setup wires 4 host events: session start, stop, session end, and pre-compaction. Full traces are copied before background distillation, so a poor summary should not erase its raw input. The core library has no built-in LLM client; distillation and sleep-time maintenance borrow the host CLI login or use a configured model endpoint.

Unattended management has 2 authority tiers. Rule-based T0 work can normalize metadata, while T1 reasoning may add or update files and submit deletion proposals. A person accepts or rejects those proposals. This separation, plus one report per sleep pass, is better suited to auditing than a memory system that silently rewrites its only copy.

What happened when we ran it

Our sandbox cloned commit 243868d with 3 CPUs, 8 GB of RAM, Python 3.12 on Debian, no secrets, and no elevated privileges. We measured 149 files, about 18,969 lines of source, and a 1 MB checkout. Installation succeeded in 27 seconds, adding 35 packages and using 37 MB on disk. The build then passed in 10 seconds, and pip-audit found 0 known vulnerabilities.

Under our test method, the test step failed with exit 4 after 12 seconds. Pytest stopped while loading tests/conftest.py; line 4 tried to import FrozenClock from agent_memory.core.clock, and Python reported ModuleNotFoundError: No module named 'agent_memory'. The log does not establish why that package was unavailable. The repository has 1 CI workflow and a tests directory, but no Dockerfile. Our result is a failed fresh-container test run, not a claim that the code has no tests.

Nine MCP tools leave context and sleep on the CLI

The MCP server exposes 9 named memory operations: recall, read, trace, record, correct, supersede, merge, delete, and feedback. That covers direct memory work from compatible agents. The README explicitly keeps context, sleep, and the proposal ledger on the CLI, so an MCP-only integration cannot operate the complete lifecycle without shell access.

Installation also asks more than uv sync. Python 3.12 or newer is required, the checkout's .venv/bin must reach your shell, and every host must agree on AGENT_MEMORY_STORE. There is no PyPI release. Optional vector search loads FastEmbed and ONNX, and its first use may download BAAI/bge-small-en-v1.5. Pinning the checkout and preparing that model cache are part of a repeatable deployment.

Three open correctness reports deserve a pre-production check

Open issue 16 describes concurrent corrections overwriting one another because the read and write do not share one lock. Issue 20 reports that a generated name from non-ASCII-only text can become empty. Issue 49 shows quoted frontmatter gaining backslashes across rewrites and scalar-looking strings returning as another type. None of the proposed fixes was merged into commit 243868d.

These defects touch the files that the architecture calls truth, so they matter more than a rough dashboard. Before adoption, test concurrent writers, your real languages, quotes, commas, and names such as true or null. The project has good rules around provenance and reversibility; those rules still depend on parsing and writes preserving the exact record.

The September push history is active, with no tagged release

GitHub showed 1,780 stars and 10 open issues and pull requests on September 30, 2026. The split was 4 issues and 6 pull requests. The repository was created September 1 and pushed September 28, when pull request 46 merged a fix for preserving group boundaries during maintenance. Activity is current, though the public history covers less than one month.

There is no GitHub release, and the README labels the workspace version 0.1.0. Mem0 is the easier choice when you want a packaged library, server, or cloud account. Graphiti fits a temporal knowledge graph built around entities and relationships. MemOS spans local plugins, self-hosting, and cloud. Agent Memory is the pick when plain files and host-shared recall are the product requirement, provided your own test run turns green first.

Alternatives

ProjectWhat it isPick it when
Mem0 gh↗A memory layer with library, self-hosted server, and managed cloud paths.pick this instead when you want a packaged SDK or hosted service rather than Markdown files as the primary store.
Graphiti gh↗A framework for temporal knowledge graphs with hybrid and graph retrieval.pick this instead when entity relationships and time-aware graph queries matter more than a browsable file tree.
MemOS gh↗A broader memory system with cloud, self-hosted, and local agent integrations.pick this instead when you need multimodal memory, a managed option, or a server backed by Neo4j and Qdrant.

What people are saying

  1. [github-trending] TencentCloud/TencentDB-Agent-Memory
  2. [velocity-scout] tigerless-labs/agent-memory
  3. [velocity-scout] okf-memory/okf-agent-memory

Sources

  1. Agent Memory README
  2. Measured commit 243868d
  3. Concurrent correction issue 16
  4. Non-ASCII name issue 20
  5. Frontmatter round-trip issue 49
  6. Mem0 README
  7. Graphiti README
  8. MemOS README

More llm tools reviews

codex-astra-luna-orchestrator · okf-agent-memory · mlc-llm · awesome-openclaw-skills · TensorFold · ai-evaluation-framework · the whole board →