mrkeyoor.com_
Wed 30 Sept 06:09 UTC
LLM Toolsevaluationupdated 30 Sept 2026

okf-agent-memory review

OKF Agent Memory stores project decisions, constraints, and runbooks as Markdown and YAML inside Git, then gives AI agents a CLI and MCP server to search and update them. It is curated project memory rather than automatic conversation recall, which makes every remembered fact reviewable in a pull request.

Verdict

Our OKF Agent Memory run built in 25 seconds and passed all 12 Go tests in 110 seconds, so the core is unusually easy to verify for a young agent-memory project. Use it when memory means curated project decisions that belong in Git and your agents can follow search-before-write rules. Choose a database-backed system for automatic personal memory, and inspect open pull request 43 before relying on its human-verification metadata for security decisions.

We ran it

Lab card: what happened when we ran okf-agent-memoryScreenshot of okf-agent-memory (okf-memory.dev)
Install✓ · 21s6 packages
Build✓ · 25s
Tests✓ · 110s12 passed · 0 failed of 12 (go test)
Repo213 files~17,675 lines of source · 1.1 MB · 2 CI workflows

Answers from our run

Does okf-agent-memory build from source?

Dependencies installed in 21 seconds (6 packages), and the build succeeded in 25 seconds. We cloned commit 007ab43 into a clean Debian container with 3 CPUs and no project-specific setup.

Do okf-agent-memory's tests pass?

Yes: 12 of 12 passed when we ran the project's own test command (go test). Some failures need services or credentials a bare container does not have.

Who should not use okf-agent-memory?

Products that need automatic extraction from large volumes of chat or user activity: OKF expects agents or people to curate concepts.

What are the alternatives to okf-agent-memory?

Mem0, Letta, Graphiti. Our OKF Agent Memory run built in 25 seconds and passed all 12 Go tests in 110 seconds, so the core is unusually easy to verify for a young agent-memory project.

Setup4/521-second install, successful build, and prebuilt binaries
Docs5/5Detailed CLI, MCP, security, format, and workflow guides
Community3/5739 stars with active issues and outside contributions
Maturity3/5v0.4.4 is tested and active, but the project is weeks old

Who it’s for

Engineering teams that want agent memory reviewed through ordinary Git diffs and pull requests.
Claude Code, Codex, Cursor, and other agent users willing to adopt a search-before-write habit.
Projects that need local retrieval without a vector database or embedding service.
Teams prepared to curate durable decisions instead of saving raw chat transcripts.

Who it’s NOT for

Products that need automatic extraction from large volumes of chat or user activity: OKF expects agents or people to curate concepts.
Teams requiring semantic vector search across vague or multilingual phrasing: retrieval uses lexical BM25, and issue 42 documents acronym and prefix false positives.
Users needing one search across personal and project memory: issue 11 says the MCP server currently handles one bundle per call.
Security-sensitive teams treating human: verification as an authorization boundary: open pull request 43 describes frontmatter smuggling and false human-provenance paths that it aims to harden.
Organizations unwilling to place project knowledge in repository files, including the review and secret-scanning work that creates.

Setup reality

Our sandbox installed commit 007ab43 in 21 seconds with 6 packages. The build succeeded in 25 seconds, and go test finished in 110 seconds with 12 passed and 0 failed of 12.

Local use needs Go to build from source, or a prebuilt v0.4.4 binary for macOS, Linux, or Windows. A project bootstrap writes a knowledge/ bundle, agent skill, AGENTS.md, and Makefile; MCP clients then point a stdio server at that bundle. Optional Hub sync needs its own token, password, and secret key.

The repository has 2 CI workflows but no Dockerfile and no tests directory at the root, despite the passing Go test command. The operational burden moves into keeping Markdown concepts current, resolving Git conflicts, and preventing secrets or unverified claims from entering memory.

Project memory becomes reviewable repository content

OKF Agent Memory gives an agent somewhere deliberate to put facts that should survive a chat window. Decisions, constraints, runbooks, and domain notes live in a knowledge/ directory as Markdown with structured YAML metadata. Git records who changed them and lets a teammate reject a bad memory in review. The model sees a short instruction layer at session start, then retrieves specific concepts when the task calls for them.

That is a different product from an automatic memory service. The getting-started guide tells agents to store durable facts rather than chat logs, search before creating a concept, update an existing concept instead of duplicating it, and validate the bundle before finishing. A bootstrap command writes 4 pieces into a project: the knowledge bundle, an agent skill, AGENTS.md, and Makefile shortcuts. The value comes from keeping that loop alive after installation.

The CLI and MCP server share the same plain-text bundle

The Go binary can initialize, validate, search, show, create, update, and relate concepts. Path-aware search can surface rules tied to a source file before an agent edits it. Lifecycle metadata includes status, provenance, review timing, and stale dates. Because the store is ordinary text, a broken index or disputed statement can be fixed with the same editor and Git workflow used for code.

MCP support exposes 6 documented tools for search, retrieval, creation, updates, relationships, and validation. Claude Code and other MCP clients run the server over stdio against one bundle. Codex and Gemini can call the CLI directly after reading AGENTS.md. There is no database process in the local path, and the latest v0.4.4 release provides binaries for macOS, Linux, and Windows on both common 64-bit architectures.

The single-bundle boundary matters. Issue 11 asks for combined user-level and project-level search because one MCP instance currently points at one root. Registering 2 servers is a workaround, but it leaves the agent to query both and produces separate rankings. If your memory model spans personal conventions, company policy, and repository facts, decide how those scopes will meet before bootstrapping hundreds of concepts.

What happened when we ran it

Our sandbox installed commit 007ab43 in 21 seconds, adding 6 packages. The repository contained 213 files, about 17,675 lines of source, and used 1.1 MB before installation. The build completed successfully in 25 seconds inside an unprivileged Debian container with 3 CPUs, 8 GB of RAM, Go 1.24, and no secrets.

The Go test step finished in 110 seconds. All 12 tests passed and 0 failed. That is a clean result for the exact commit we ran, though it does not verify the README's retrieval-speed, token-reduction, or memory-footprint tables. Those are project benchmarks with their own runner; our supplied lab block measured installation, build, and tests only.

We found 2 CI workflow files, no Dockerfile, and no root tests directory. Go tests can live beside packages, so the missing directory does not conflict with the 12 passing cases. A static binary also makes a Dockerfile optional for many users. The more useful operational test is whether your agents reliably search before editing, record only durable facts, and leave the bundle conformant after concurrent work.

Lexical search is cheap but has known ranking edges

OKF uses in-memory BM25 rather than embeddings. That keeps retrieval local and removes model credentials from ordinary search. It also means wording matters. Exact project names, API terms, filenames, and architecture labels are a good fit. A vague recollection expressed with different vocabulary may be easier for a vector or graph system to recover.

Open issue 42 identifies 3 concrete scoring problems: document frequency and term frequency use different matching behavior, short stems can match unrelated prefixes, and some metadata fields lack term-frequency caps. Its examples include ci inside longer words and log matching login. These are maintainer-documented search defects, not failures from our 12-test run. Teams with acronym-heavy knowledge should add their own retrieval checks before treating the top result as authoritative.

Provenance metadata still needs a security review

The format distinguishes generated material from verified material, and the documentation tells agents never to forge human verification. That is useful governance, but metadata is not self-enforcing merely because the schema names it. If a tool can write the file, its mutation path must reject a forged verifier and frontmatter that hides reserved fields.

Open pull request 43, updated September 27, describes 2 related weaknesses and proposed fixes: YAML frontmatter smuggling around reserved keys and agent-supplied human: verification entries. The pull request says it adds parser hardening and negative tests. Until that change is reviewed and lands in the version you deploy, do not use a verification label as permission to run commands, release code, or bypass a human approval step.

Active releases help, but v0.4.4 is still young

The repository was pushed on September 27, 2026, the same day v0.4.4 shipped. GitHub showed 739 stars and 3 combined open issues and pull requests on September 30. Recent release notes name outside contributors for mutation parity and a link-resolution bug, which is better evidence of community involvement than the star count alone. Two CI workflows provide visible automation.

The project itself was created on September 5, so a version number and clean tests should not be mistaken for years of compatibility history. Adopt it behind Git review, back up the knowledge/ directory with the rest of the repository, and watch how agents behave when two branches edit the same index. For curated engineering memory, the design is sensible. For automatic life history or cross-user personalization, Mem0, Letta, or a graph store addresses a different problem.

Alternatives

ProjectWhat it isPick it when
Mem0 gh↗A memory layer that extracts and retrieves facts for agents and applications.pick this instead when automatic user and conversation memory matters more than Git-native review.
LettaA platform for stateful agents with managed memory and long-running behavior.pick this instead when you need an agent runtime, not only a project knowledge format and search tool.
Graphiti gh↗A temporal knowledge-graph engine for changing agent facts and relationships.pick this instead when time-aware entity relationships and automatic graph ingestion outweigh plain-text Git review.

What people are saying

  1. [hackernews] OKF Agent Memory – Git-native persistent memory for AI coding agents
  2. [velocity-scout] okf-memory/okf-agent-memory

Sources

  1. OKF Agent Memory README
  2. Getting started guide
  3. v0.4.4 release
  4. BM25 scoring issue
  5. Provenance hardening pull request

More llm tools reviews

codex-astra-luna-orchestrator · mlc-llm · awesome-openclaw-skills · TensorFold · ai-evaluation-framework · llm-wiki-compiler · the whole board →