Project memory becomes reviewable repository content
OKF Agent Memory gives an agent somewhere deliberate to put facts that should survive a chat window. Decisions, constraints, runbooks, and domain notes live in a knowledge/ directory as Markdown with structured YAML metadata. Git records who changed them and lets a teammate reject a bad memory in review. The model sees a short instruction layer at session start, then retrieves specific concepts when the task calls for them.
That is a different product from an automatic memory service. The getting-started guide tells agents to store durable facts rather than chat logs, search before creating a concept, update an existing concept instead of duplicating it, and validate the bundle before finishing. A bootstrap command writes 4 pieces into a project: the knowledge bundle, an agent skill, AGENTS.md, and Makefile shortcuts. The value comes from keeping that loop alive after installation.
The CLI and MCP server share the same plain-text bundle
The Go binary can initialize, validate, search, show, create, update, and relate concepts. Path-aware search can surface rules tied to a source file before an agent edits it. Lifecycle metadata includes status, provenance, review timing, and stale dates. Because the store is ordinary text, a broken index or disputed statement can be fixed with the same editor and Git workflow used for code.
MCP support exposes 6 documented tools for search, retrieval, creation, updates, relationships, and validation. Claude Code and other MCP clients run the server over stdio against one bundle. Codex and Gemini can call the CLI directly after reading AGENTS.md. There is no database process in the local path, and the latest v0.4.4 release provides binaries for macOS, Linux, and Windows on both common 64-bit architectures.
The single-bundle boundary matters. Issue 11 asks for combined user-level and project-level search because one MCP instance currently points at one root. Registering 2 servers is a workaround, but it leaves the agent to query both and produces separate rankings. If your memory model spans personal conventions, company policy, and repository facts, decide how those scopes will meet before bootstrapping hundreds of concepts.
What happened when we ran it
Our sandbox installed commit 007ab43 in 21 seconds, adding 6 packages. The repository contained 213 files, about 17,675 lines of source, and used 1.1 MB before installation. The build completed successfully in 25 seconds inside an unprivileged Debian container with 3 CPUs, 8 GB of RAM, Go 1.24, and no secrets.
The Go test step finished in 110 seconds. All 12 tests passed and 0 failed. That is a clean result for the exact commit we ran, though it does not verify the README's retrieval-speed, token-reduction, or memory-footprint tables. Those are project benchmarks with their own runner; our supplied lab block measured installation, build, and tests only.
We found 2 CI workflow files, no Dockerfile, and no root tests directory. Go tests can live beside packages, so the missing directory does not conflict with the 12 passing cases. A static binary also makes a Dockerfile optional for many users. The more useful operational test is whether your agents reliably search before editing, record only durable facts, and leave the bundle conformant after concurrent work.
Lexical search is cheap but has known ranking edges
OKF uses in-memory BM25 rather than embeddings. That keeps retrieval local and removes model credentials from ordinary search. It also means wording matters. Exact project names, API terms, filenames, and architecture labels are a good fit. A vague recollection expressed with different vocabulary may be easier for a vector or graph system to recover.
Open issue 42 identifies 3 concrete scoring problems: document frequency and term frequency use different matching behavior, short stems can match unrelated prefixes, and some metadata fields lack term-frequency caps. Its examples include ci inside longer words and log matching login. These are maintainer-documented search defects, not failures from our 12-test run. Teams with acronym-heavy knowledge should add their own retrieval checks before treating the top result as authoritative.
Provenance metadata still needs a security review
The format distinguishes generated material from verified material, and the documentation tells agents never to forge human verification. That is useful governance, but metadata is not self-enforcing merely because the schema names it. If a tool can write the file, its mutation path must reject a forged verifier and frontmatter that hides reserved fields.
Open pull request 43, updated September 27, describes 2 related weaknesses and proposed fixes: YAML frontmatter smuggling around reserved keys and agent-supplied human: verification entries. The pull request says it adds parser hardening and negative tests. Until that change is reviewed and lands in the version you deploy, do not use a verification label as permission to run commands, release code, or bypass a human approval step.
Active releases help, but v0.4.4 is still young
The repository was pushed on September 27, 2026, the same day v0.4.4 shipped. GitHub showed 739 stars and 3 combined open issues and pull requests on September 30. Recent release notes name outside contributors for mutation parity and a link-resolution bug, which is better evidence of community involvement than the star count alone. Two CI workflows provide visible automation.
The project itself was created on September 5, so a version number and clean tests should not be mistaken for years of compatibility history. Adopt it behind Git review, back up the knowledge/ directory with the rest of the repository, and watch how agents behave when two branches edit the same index. For curated engineering memory, the design is sensible. For automatic life history or cross-user personalization, Mem0, Letta, or a graph store addresses a different problem.

