AI & AUTOMATION

AgentMemory: Persistent Memory for AI Coding Agents That Finally Remembers You

Key Takeaways:

AgentMemory is a one-command, self-hosted memory server that captures, compresses, and re-injects context across every AI coding agent you use, so you never have to re-explain your stack, preferences, or past decisions again.

What Is AgentMemory?

If you have ever spent the first ten minutes of every Claude Code or Cursor session re-explaining your project architecture, your preferred libraries, or last week’s bug fix, AgentMemory was built for you. It is an open-source persistent memory engine that silently captures what your coding agent does, compresses it into searchable long-term memory, and injects the right context when the next session starts. The same memory server is shared across every agent that speaks MCP or HTTP, so the knowledge you build with Claude Code instantly becomes available to Cursor, Codex CLI, Gemini CLI, OpenCode, Cline, Goose, Aider, Windsurf, and many more.

The project is published on GitHub at rohitg00/agentmemory and is built on top of the iii engine. It implements and extends Andrej Karpathy’s LLM Wiki pattern with confidence scoring, lifecycle management, knowledge graphs, and hybrid search – and it has quickly become one of the most starred memory projects in the agent ecosystem.

Why AgentMemory Matters

Built-in agent memory like CLAUDE.md or .cursorrules caps out around 200 lines and goes stale quickly. Paste-the-full-context approaches blow past LLM windows and burn millions of tokens per year. AgentMemory takes a smarter path with several key advantages.

Best-in-class retrieval. On the LongMemEval-S benchmark (ICLR 2025, 500 questions), AgentMemory achieves 95.2% Recall@5, 98.6% Recall@10, and an MRR of 88.2 – significantly outperforming a BM25-only baseline. On the in-house coding-agent benchmark it hits a 100% top-5 hit rate with a median latency of just 14 milliseconds.

Massive token savings. A typical full-context paste workflow burns more than 19 million tokens per year. LLM-summarized approaches still cost roughly $500 annually. AgentMemory drops the same usage to about 170K tokens per year – and to literally $0 if you run local embeddings with all-MiniLM-L6-v2. That is up to 92% fewer tokens for equivalent or better recall.

Works with every agent. Claude Code, Cursor, Codex CLI, Gemini CLI, Hermes, OpenClaw, pi, OpenHuman, OpenCode, Cline, Goose, Kilo Code, Aider, Claude Desktop, Windsurf, and Roo Code are all supported through 53 MCP tools, 12+ native lifecycle hooks, or a plain REST API on port 3111. One server, memories shared across all of them.

Zero external dependencies. No Postgres, no Redis, no Pinecone, no Weaviate. AgentMemory stores everything locally, ships with 950+ passing tests, and runs from a single npm package.

Privacy-first by design. Captures are filtered for secrets, deduplicated via SHA-256 hashes, and stored under an explicit audit policy that governs every delete path. Your code stays on your machine.

Installing AgentMemory in Under Five Minutes

Illustration of installing a Node.js CLI tool from the terminal

AgentMemory ships as a single npm package. You have three clean install paths.

Option 1: Global npm install (recommended)

Run once, then call agentmemory from anywhere on your PATH:

npm install -g @agentmemory/agentmemory

# If you hit EACCES on macOS or Linux system Node installs:
# sudo npm install -g @agentmemory/agentmemory

agentmemory                # start the memory server on :3111
agentmemory demo           # seed sample sessions and prove recall
agentmemory connect claude-code   # wire your agent (also: codex, cursor, gemini-cli, ...)

Option 2: One-shot run with npx

If you want to kick the tires without installing anything globally:

npx -y @agentmemory/agentmemory@latest

The first npx run from v0.9.16 onward will prompt you to install globally inline so the bare agentmemory command works everywhere afterwards. If a stale npx cache serves you an older release, force the latest with the @latest tag or clear ~/.npm/_npx once.

Option 3: From source

For contributors and self-hosters:

git clone https://github.com/rohitg00/agentmemory.git
cd agentmemory
npm install
npm run build
npm start

The REST API listens on :3111, the MCP server can be proxied through the running daemon, and a real-time viewer lets you inspect everything your agents are remembering.

Configuring AgentMemory

Configuration lives in ~/.agentmemory/.env. The defaults work out of the box with local embeddings and no API keys, but you can tune almost everything.

LLM providers. Set ANTHROPIC_API_KEY, OPENAI_API_KEY, or GEMINI_API_KEY if you want LLM-powered compression of captured observations.

Embeddings. Set EMBEDDING_PROVIDER to local (default, free, no key needed), openai, voyage, or another supported provider. Local mode uses all-MiniLM-L6-v2, which is small, fast, and runs entirely offline.

Behavior toggles. AGENTMEMORY_AUTO_COMPRESS controls whether observations are automatically summarized by an LLM. AGENTMEMORY_TOOLS switches between the core MCP toolkit and the full 53-tool surface.

How AgentMemory Actually Works

Under the hood, AgentMemory runs a 4-tier consolidation pipeline inspired by human memory – Working, Episodic, Semantic, and Procedural – with three core phases.

Capture

When your agent finishes a tool call, a PostToolUse hook (or an MCP memory_save call) fires. The raw observation is hashed with SHA-256 for deduplication, run through a privacy filter that strips secrets, and stored as a raw observation in the local database.

Compress

A background worker takes raw observations and compresses them into structured facts using either an LLM (if you provided keys) or a deterministic local pipeline. The result is embedded with your configured embedding model and indexed for hybrid BM25 + vector + graph retrieval.

Inject

On the next SessionStart, AgentMemory loads your project profile, runs a hybrid search over your accumulated memory, and injects the most relevant facts into the agent’s prompt within a configurable token budget. The agent starts the session already knowing your auth middleware lives in src/middleware/auth.ts, that you prefer jose over jsonwebtoken for Edge compatibility, and that the failing test from last Tuesday was about token expiry.

Using AgentMemory Day to Day

The typical workflow looks like this.

Step 1 – Start the server. Run agentmemory in a terminal (or set it up as a system service). It listens on :3111 and exposes both REST and MCP endpoints.

Step 2 – Connect your agents. Run agentmemory connect claude-code, agentmemory connect cursor, agentmemory connect codex, or agentmemory connect gemini-cli. Each connector wires native hooks or registers an MCP server entry so your agent finds AgentMemory automatically. For Cline, Goose, Kilo Code, Roo Code, Claude Desktop, and Windsurf, the MCP server is the standard hook-in point. Aider speaks to it through the REST API.

Step 3 – Work normally. Just write code as you always do. AgentMemory observes tool calls, file edits, and conversations in the background. Run agentmemory demo once at the start to seed sample sessions and watch how recall works.

Step 4 – Inspect via the real-time viewer. A local dashboard shows what is being captured, compressed, and recalled. The MCP server exposes 53 tools including memory_recall, memory_save, memory_smart_search, memory_sessions, memory_graph_query, and memory_snapshot_create for advanced workflows like time-travel debugging and multi-agent coordination via memory_lease and memory_signal_send.

Step 5 – Share memory across agents. This is where AgentMemory shines. The auth decision you made with Claude Code in the morning is automatically available to Cursor in the afternoon, to Codex CLI in CI, and to Gemini CLI on your laptop. One server, one memory, every agent.

Is AgentMemory Right for You?

If you live in a multi-agent world – flipping between Claude Code, Cursor, Codex, and Gemini CLI – AgentMemory is one of the most pragmatic upgrades you can make to your developer workflow. It is fast, local, free at the embedding layer, requires no external databases, and integrates with virtually every agent that matters today.

If you are deep in the OpenHuman ecosystem, AgentMemory is also the project that powers OpenHuman’s optional Memory trait backend – so investing in it pays dividends across your entire stack.

Star the GitHub repository, run a single npx command, connect your favorite agent, and stop re-explaining your codebase every morning. Your future sessions will thank you.

You may also like

Subscribe
Notify of
guest

0 Comments
Newest
Oldest Most Voted