Field report · · AgentConn Team
Agents Don't Need Memory, They Need Docs
claude-mem hits 96K stars the same day HN debates why agents need docs, not memory. The real architecture lesson for builders.
Agents Don’t Need Memory, They Need Docs
On the same day this week, two signals collided on every developer’s feed. On GitHub Trending, claude-mem hit 96,001 stars with 627 new ones that day alone — a persistent memory plugin that captures everything your coding agent does and replays it next session. On Hacker News, a post titled “Agents don’t need memory, they need documentation” climbed to 297 points with 173 comments, arguing the exact opposite: that most agent memory systems are solving the wrong problem entirely.
The tension is everywhere. Developer Shirish captured the frustration driving this debate:
This is not a coincidence. It is a real-time architectural fork, and which side you pick determines how your agent stack scales — or doesn’t.
We have argued memory is the new moat in coding agents. Today we are going to argue against ourselves, because the counter-thesis deserves a serious hearing. And the truth, as usual, is more interesting than either camp admits.
The Case Against Memory
The original blog post that lit up HN makes a sharp argument: most agent failures blamed on “poor memory” are actually documentation failures. The agent did not forget your architecture — you never wrote it down in a form the agent could use. You are not solving a memory problem. You are compensating for a documentation deficit with an expensive, lossy, hallucination-prone retrieval layer.
The numbers support the critique. Research published earlier this year found that providing full conversation history to agents results in 91% higher p95 latency and over 90% higher token costs. Models start degrading around 130,000 tokens, putting effective capacity at roughly 60-70% of what the context window advertises. You are paying a premium in latency, cost, and reliability to store information that could live in a version-controlled markdown file.
The HN discussion surfaced the sharpest version of this argument. One commenter put it bluntly: “You don’t need documentation or the 3rd party memory systems. The code IS the documentation. All this stuff is LLM rube goldberg machines. It just pollutes context.” Another offered a practical counter-architecture: “What’s wrong with README.md? A general one at the top level, a specific one for each subfolder that contains a specific subsystem. Next to it, a CLAUDE.md that only has ‘See @README.md.’ Claude Code knows how to use this.”
James Long, creator of Prettier and Actual Budget, crystallized the documentation-first intuition in a thread that caught significant attention:
His distinction matters: AGENTS.md is not memory. It is documentation. The difference is that documentation is loaded deliberately, structured intentionally, and never surprises you with stale context from a session you forgot about.
The Documentation-First Architecture
The documentation camp is not just complaining about memory. They are building an alternative infrastructure.
AGENTS.md — a simple, open format for guiding coding agents — is now used by over 60,000 open-source projects. It has quietly become the closest thing the agent ecosystem has to a shared standard. Codex CLI, Copilot CLI, Gemini CLI, Cursor, and Claude Code itself all recognize it. It solves the same problem claude-mem solves — agents forgetting your project conventions between sessions — without a database, without embeddings, without token overhead on every retrieval.
The research backs this up. A study on machine-readable project context found that repositories with structured agent instruction files achieve a 29% reduction in median agent runtime and a 17% reduction in output token consumption. The key variable is not whether you use CLAUDE.md or AGENTS.md — it is whether you use anything at all. The comparison between formats confirms this: “A good AGENTS.md is a model upgrade, but a bad one is worse than no docs at all.”
When we covered context files and what actually works earlier this year, we reported ETH Zurich’s finding that LLM-generated context files decrease agent performance by 3% and increase cost by 20-23%. The reconciling variable is quality: human-written, minimal context files improve performance by about 4%. This means the documentation-first approach works — but only when a human writes and maintains the docs with intention.
Here is what a documentation-first agent setup looks like in practice:
repo/
CLAUDE.md # Project conventions, arch decisions, test commands
AGENTS.md # Vendor-neutral copy (symlinked or identical)
docs/
architecture.md # System design, service boundaries, data flow
conventions.md # Naming patterns, error handling, API contracts
runbooks/ # Operational procedures agents can follow
src/
module/
README.md # Module-specific context, dependencies, gotchas
This entire tree is deterministic, versionable, reviewable, and costs exactly zero tokens until the agent needs it. No database. No vector index. No embedding model. No monthly bill.
But claude-mem Has 96,000 Stars for a Reason
Here is where the documentation purists lose the plot: dismissing the memory layer entirely ignores real, measured demand.
Alex Newman’s claude-mem — now at 96K stars and 7,500 forks — exists because documentation alone cannot solve every persistence problem. It captures what Claude does during coding sessions, compresses the observations using AI, and injects relevant context back into the next session. That last part is key: it does not just store information — it selects and compresses what to replay based on relevance to the current task.
This is a genuinely different capability from static documentation. Your CLAUDE.md tells the agent how your project works. Your memory layer tells the agent what you did last session, what failed, what you decided, and why. These are not the same question.
The State of AI Agent Memory 2026 report from Mem0 formalized this distinction: RAG answers “what does the documentation say about this?” while memory answers “what do I already know about this specific user or task?” Agent memory in 2026 is a first-class architectural component with its own benchmark suite, its own research literature, and a measurable performance gap between approaches.
Andrew Ng’s partnership with Oracle to build an entire course on agent memory signals how seriously the industry takes this layer:
The Agentic Context Management research from July 2026 proposes what may be the right framing: a three-layer architecture. A “hot memory” constitution encoding conventions and orchestration protocols (this is your CLAUDE.md). Specialized domain-expert agents (these are your documentation modules). And a “cold memory” knowledge base of on-demand specification documents (this is where real memory earns its keep — dynamic, session-specific, evolving).
When we wrote about memory as a cost center, we documented the economics: DRAM up 500%, Mem0’s graph tier at $249/month. The documentation camp’s implicit argument is that 80% of what memory systems store should be free — because it is project knowledge that belongs in version control, not a vector database.
The Real Architecture Lesson
The debate crystallizes around one question: what belongs in documentation, and what belongs in memory?
Here is our framework for where each type of knowledge belongs:
| Knowledge Type | Where It Belongs | Why |
|---|---|---|
| Project conventions (naming, patterns, error handling) | CLAUDE.md / AGENTS.md | Static, shared, reviewable |
| Architecture decisions | docs/architecture.md + ADRs | Versioned, auditable |
| API contracts and schemas | OpenAPI specs, type definitions | Machine-readable, tested |
| ”Last session we decided to…” | Memory layer | Session-specific, evolving |
| ”This user prefers…” | Memory layer | User-specific, not universal |
| ”This approach failed because…” | Memory layer and documentation | Start in memory, promote to docs when validated |
| Build/test commands | CLAUDE.md | Deterministic, everyone needs them |
The last row is the interesting one. Some knowledge starts as memory (an observation from a failed attempt) and graduates to documentation once validated. The best agent architectures treat memory as a staging area for documentation, not a replacement for it.
What Builders Should Do Now
The same-day collision between claude-mem’s star count and the HN documentation post is not a contradiction — it is a roadmap. Both camps are right about their own scope and wrong about the other’s.
Here is the concrete playbook:
Step 1: Documentation first. Before installing any memory plugin, write your CLAUDE.md. Include project structure, naming conventions, test commands, architecture decisions, and common pitfalls. This alone — per the research — cuts agent runtime by 29% and token use by 17%. You get that improvement on day one, for free. The Ramble Session approach can jumpstart this if you are staring at a blank file.
Step 2: Add memory for genuinely dynamic state. Session history, user preferences, failed-approach logs, and evolving project state are legitimate memory use cases. Tools like claude-mem, Mem0, and Letta exist because these problems are real. Use them for what they are good at — not as a substitute for writing things down.
Step 3: Build the graduation pipeline. When a memory observation proves useful across multiple sessions, promote it to documentation. “This test suite fails silently when run outside Docker” is a memory observation the first time. The fifth time, it belongs in CLAUDE.md. The teams that automate this graduation — memory to docs — will have the most reliable agent stacks.
Step 4: Measure both layers. Track token costs, latency, and task completion rates separately for documentation-loaded context and memory-retrieved context. If documentation-loaded context has a higher hit rate (it usually does), you know where to invest your engineering time.
The Real Moat Is Knowledge Infrastructure
We titled our earlier piece “Memory Is the New Moat.” Today’s counter-thesis demands an update to that framing. The moat is not memory alone. It is not documentation alone. It is the knowledge infrastructure that connects both: deterministic documentation as the foundation, dynamic memory as the optimization layer, and a graduation pipeline that continuously hardens observations into documentation.
The teams winning the agent game in late 2026 are not choosing between claude-mem and CLAUDE.md. They are using both — with documentation carrying the load and memory handling the edge cases. The HN post is right that most teams should start with documentation. The 96,000 developers starring claude-mem are right that documentation alone is not enough.
The architecture that survives is the one that treats documentation as the source of truth and memory as the cache. Get the layers right, and your agent stack becomes both reliable and adaptive. Get them backwards — all memory, no documentation — and you are building an expensive system to compensate for knowledge you never wrote down.







