Field report · · AgentConn Team
The Real Moat Is Being the Tool Your Agent Reaches For
Agent-infra repos top GitHub with 360K+ combined stars. Postiz hit $221K MRR by becoming agent-callable. The moat moved from models to infrastructure.
The Real Moat Is Being the Tool Your Agent Reaches For
Four repos on today’s GitHub Trending tell you everything you need to know about where the money is moving. agency-agents — a bundle of agent persona templates — sits at 157,000 stars. caveman — a proxy that makes coding agents talk in compressed fragments to cut token costs 65% — has 110,000. claude-mem — persistent memory that follows agents between sessions — just crossed 96,500. Agent-Reach — one CLI that gives agents eyes on Twitter, Reddit, YouTube, and GitHub with zero API fees — is pulling 1,156 new stars per day.
None of these repos ship a model. None of them ship an agent. They ship infrastructure that agents reach for — the memory, the persona definitions, the token optimization, the search layer. Combined: over 455,000 stars for tools that wrap around models, not tools that are models. That is not a trend. That is a thesis the market is pricing in real-time.
The competitive moat in AI has moved. It is no longer about having the best model — Sonnet 5.5, Opus 5, GPT-5, Gemini 3.8 are all “good enough” for most production workloads. The moat is about being the tool that the agent calls when it needs to do something in the real world: search the web, remember a conversation, post to social media, optimize its own token spend.
The Model Moat Is Dead — Long Live the Infrastructure Moat
Two Minute Papers put it bluntly in their recent video: “The Billion Dollar AI Advantage Is Disappearing.” The performance gap between frontier models has collapsed. When Sonnet 5.5 matches Opus 5 on most benchmarks at a fraction of the cost, and open-weight alternatives offer comparable reasoning, the model itself stops being a competitive advantage.
Harrison Chase made the same argument at LangChain’s Interrupt NYC keynote: the future belongs to teams that “own their intelligence” — not by training their own model, but by building domain-specific infrastructure around whatever model they use. The three pillars he identified: proprietary data connectors, custom evaluation loops, and an open model-neutral harness. Notice what is absent from that list: the model.
Karpathy himself demonstrated how commoditized models have become. His recent eval — simply asking an LLM “Land or Water?” with latitude/longitude coordinates 16,200 times and plotting the results as an image — showed that the models all know the same thing now. The differentiator is not what the model knows. It is what the infrastructure around the model enables it to do.
This is not a new thesis on AgentConn. We covered the harness dimension in Memory Is the New Moat in Coding Agents and the skills layer in Agent Skills Are the New Dotfiles. But what is new is the scale of the market signal. In July, agent-infra repos were pulling thousands of stars. Today they are pulling hundreds of thousands. The picks-and-shovels thesis is not speculative anymore — it is the dominant strategy on GitHub.
The 360K-star signal: agency-agents (157K) + caveman (110K) + claude-mem (96.5K) = 363,500 combined GitHub stars for three repos that ship zero models and zero agents. They ship infrastructure that makes agents better. That is the market telling you where the moat lives.
The GitHub Agent-Infra Gold Rush
Let us examine what these four repos actually do, because their success patterns reveal the infrastructure thesis at the nuts-and-bolts level.
View agency-agents on GitHub →
agency-agents: 157K Stars for Agent Personas
agency-agents does not ship code in the traditional sense. It ships 147+ specialized AI agent persona templates — markdown files with calibrated system prompts, tool permissions, and decision guardrails for engineering, design, marketing, sales, PM, and twelve other functions. It works with Claude Code, Cursor, Codex, Gemini CLI, and every major harness.
The repo has accumulated stars across nine months of steady growth rather than a single viral spike. That durability matters: it means the demand is structural, not a meme. Developers are not starring it because it went viral — they are starring it because they use the persona files daily in their agent workflows.
The lesson: the agent’s identity layer is infrastructure. If you standardize how agents introduce themselves, scope their tools, and make decisions, you have created a dependency that every downstream workflow relies on.
caveman: 110K Stars for Token Optimization
caveman started as a joke — make your coding agent talk like a caveman to save tokens. It cuts output tokens by 65%, which translates to roughly 33% lower provider-billed input costs per session. The proxy sits between the user and the agent, compressing responses without losing technical substance.
At 110,000 stars, it hit #1 on Hacker News and GitHub Trending. The caveats are real — the 65% cut only applies to output tokens, and the skill itself consumes input tokens. But the signal is clearer: developers will install infrastructure that makes agents cheaper to run, even when the savings are nuanced.
The lesson: cost optimization is an infrastructure layer. As agent usage scales from minutes per day to hours per day, the economics of token spend become a first-order engineering concern. The tool that sits in the middle and makes every agent call cheaper has a structural advantage.
claude-mem: 96.5K Stars for Persistent Memory
claude-mem captures everything an agent does during a session, compresses it with AI, and injects the relevant pieces into future sessions automatically. With v13.8.0, it added a server-beta backed by Postgres and BullMQ, which means persistent memory now scales to team deployments — not just individual developers.
At 96,500 stars, 288 releases, and 124 contributors, this is not a weekend project — it is a production memory system that works across Claude Code, Gemini CLI, Codex, and ten other harnesses. The fact that the most recent commit was co-authored by @claude tells you something about the feedback loop: the agents are helping build their own memory infrastructure.
The lesson: memory is the stickiest infrastructure layer. Once an agent has weeks of compressed session history, switching to a different memory provider means losing that context. This is the definition of a moat — a service that gets more valuable the longer you use it.
Agent-Reach: 91.7K Stars for Agent Perception
Agent-Reach gives agents the ability to read and search Twitter, Reddit, YouTube, GitHub, and more through a single CLI with zero API fees. At 1,156 new stars per day — the highest daily growth rate of any repo in this cluster — the demand for agent perception infrastructure is explosive.
The lesson: the sensory layer is infrastructure. Agents that cannot see the web are blind agents. The tools that give agents eyes and ears become dependencies that every agentic workflow needs.
Postiz: The SaaS That Became Agent-Callable and 7x’d
The GitHub repos tell the supply-side story. Postiz tells the demand-side story — and the revenue numbers make it impossible to ignore.
Postiz is an open-source social media scheduler that supports 28+ channels. It was doing fine as a standard SaaS: roughly $21,000 MRR, growing steadily, one-person operation. Then its creator made a pivotal decision: rebuild the product as agent-callable infrastructure by adding a CLI, an MCP server, and a well-documented API that agents could invoke programmatically.
Read the full Postiz case study on SuperFrameworks →
The result: revenue climbed from $21K to over $145K MRR in four months, eventually reaching a Stripe-verified $221,581 MRR. That is a 10x jump, driven primarily by one architectural decision — making the product callable by agents, not just usable by humans.
The Postiz Signal: $21K MRR (human-only) to $221K MRR (agent-callable) in under six months. The product did not change. The interface did — from browser-only to CLI/MCP/API. The customers who showed up were not humans scheduling tweets. They were agents scheduling tweets on behalf of humans.
This is the clearest demand-side evidence that being agent-callable is a growth strategy, not a feature. When agents can call your tool programmatically, your addressable market expands from “humans who visit your dashboard” to “every agent workflow that touches your domain.”
The implications for SaaS founders are concrete: if your product does not have an API that agents can discover and call, you are invisible to the fastest-growing customer segment — automated workflows. Adding an MCP server is not a nice-to-have. It is a distribution channel.
Cloudflare’s Search API: The Agent Tollbooth
View the discussion on Hacker News →
On October 2, 2026, Cloudflare launched its Web Search API — and the HN response tells you how the market read it. 306 points and 154 comments, with the top comment cutting straight to the infrastructure question: “My number one question about search APIs is always if they allow you to store and resyndicate results you get from them.”
Read the official announcement on Cloudflare →
The API is positioned inside Cloudflare’s AI Gateway, which means every agent search query flows through unified logging, billing, analytics, and access controls. The search providers — Ceramic.ai, Exa, and Linkup — committed to following Cloudflare’s crawling standards. The pricing model: pay-per-query through Cloudflare’s billing system.
This is the tollbooth strategy. Cloudflare is not building agents. Cloudflare is not building models. Cloudflare is building the infrastructure that sits between agents and the web — and taking a cut of every query.
The strategic logic is sound. Cloudflare’s own data shows that bot-generated traffic surpassed human traffic in July 2026, hitting 57.5% of all global web requests. Over half of that is AI agents. If your infrastructure already handles the majority of web traffic, positioning yourself as the search layer for agents is not a pivot — it is a natural expansion.
Another HN commenter noted: “For those developers out there, the best is still Gemini Flash Lite 2.5 — it gives you 1,000 Google searches per day for free.” This comparison tells you the pricing dynamics are already competitive, which reinforces the infrastructure thesis: when search-for-agents becomes commodity infrastructure, the moat belongs to whoever controls the billing and routing layer, not the search index itself.
Contrarian Corner: Infrastructure Commoditizes Too
Here is the counter-argument to the agent-infra-is-the-moat thesis: infrastructure commoditizes even faster than models. If Postiz can add an MCP server in a few weeks, so can Buffer, Hootsuite, and every social media incumbent. If caveman can cut tokens with a proxy, so can the next proxy. Today’s 157K-star persona bundle is tomorrow’s built-in feature of Claude Code.
The four-layer infrastructure stack — Memory, Execution, Tooling, Governance — is real, but only the stateful layers have lasting defensibility. A memory system with six months of compressed session history has a switching cost. A persona template has none. A billing-integrated search API with enterprise contracts has lock-in. A token optimizer is a commodity the moment the model providers ship their own compression.
The real moat might be neither model nor infrastructure, but the proprietary data and workflows built on top of both. The company that uses agent-callable infra to automate its unique business processes — and generates months of operational memory in the process — is harder to displace than the infra provider itself.
What the Community Is Saying
The convergence signal is not just GitHub stars. It is showing up across every platform where practitioners discuss tooling decisions.
Andrej Karpathy — one of the most credible voices in AI engineering — described the shift directly: the new paradigm for interacting with models is “significantly more inline with all the other human activity org-wide.” The key phrase is “under the hood engineering work to make this just work” — that engineering work is the infrastructure layer. It is not the model. It is everything around the model.
At the AI Engineer Summit NYC, six talks dropped in 24 hours, all circling the same theme: agents are leaving the demo stage and hitting real systems. The common thread across all six: the infrastructure challenges dominate. Sandhya Subramani from AWS presented agents that write their own tools at runtime. Karthik Ranganathan from Yugabyte presented on the gap between per-agent memory (solved) and cross-agent learning (unsolved). Dan Adler from Sourcegraph asked the question every enterprise team is grappling with: “Claude Code can make the change — but what if you have 90,000 repos?”
Meanwhile, Boris Cherny’s post about Tag — an agent that writes over 50% of his daily PRs, handles all data analysis, and fixes most product feedback and bugs — illustrates why agent-callable infrastructure matters at the individual team level:
When agents are writing half a team’s PRs, the infrastructure those agents call becomes the critical path for the entire engineering operation. The tools in the agent’s toolkit are not optional dependencies — they are production infrastructure.
Sam Altman added a meta-layer signal with his post about Sign In With ChatGPT/Plugin Extensions: “I think there is much more potential energy in Sign In With ChatGPT/Plugin Extensions than we realize.” Translation: the identity and access layer for agents is still wide open — and whoever captures it controls a chokepoint.
The Four-Layer Agent Infrastructure Stack
Where does defensibility actually live? The emerging consensus among infrastructure investors points to four layers, each with different moat characteristics:
Layer 1: Memory (Strong Moat) Persistent agent memory that accumulates over time. claude-mem, Mem0, and Cloudflare’s Agent Memory all target this layer. The switching cost is real — months of compressed context history cannot be trivially migrated. Mem0’s $24M Series A and Braintrust’s $80M Series B validate the investor thesis.
Layer 2: Execution (Moderate Moat) Durable agent execution with state management. Temporal’s “Human Is an Async API” talk at the AI Engineer Summit NYC signals that this layer is maturing. The moat comes from workflow history and reliability guarantees.
Layer 3: Tooling (Weak-to-Moderate Moat) The MCP servers, CLI wrappers, and API adapters that let agents interact with external services. This is where Postiz sits, where Cloudflare’s search API sits, and where most of the GitHub star growth is happening. The moat is weak at the individual tool level — any service can add an MCP endpoint — but strong at the aggregate level. An ecosystem of 1,000 tool integrations is hard to replicate.
Layer 4: Governance (Emerging Moat) Agent observability, evaluation, and guardrails. Arize’s $70M Series C targets this layer. The moat grows with the volume of eval data — a company with six months of production agent traces has institutional knowledge that a new entrant lacks.
The key insight: the moat is proportional to the state the layer accumulates. Stateless infrastructure — persona templates, token proxies, single-query search APIs — is valuable but commoditizable. Stateful infrastructure — memory graphs, execution histories, eval datasets — is defensible.
What This Means for Builders
Tech With Tim demonstrated this thesis concretely by building a coding agent harness from scratch. Start with a raw LLM call — useless. Add an agentic loop — it can chain tool calls. Add an MCP client — it instantly picks up tools it never knew existed. The infrastructure around the model is what makes the agent smart.
If you are building SaaS, the action items are clear:
- Add an MCP server. It is the minimum viable interface for agent-callable products. Postiz’s revenue data proves the ROI.
- Invest in stateful layers. Memory, execution history, and eval data are the layers that accumulate defensibility. Persona templates and token optimization are valuable but not moats.
- Think in terms of agent workflows, not human workflows. The agent’s decision about which tool to call is the new distribution channel. Being discoverable by agents — through tool registries, MCP hubs, and well-structured API schemas — is the new SEO.
- Watch the tollbooth plays. Cloudflare’s search API is the template: sit between agents and the resource they need, provide unified billing and access control, take a cut. Every cloud provider will attempt this.
The end of the app era that Nate B Jones describes is not about apps disappearing. It is about the interface layer moving from browser tabs to agent tool calls. The apps that survive are the ones agents can reach.
The Road Ahead
The agent infrastructure gold rush will not last forever — the same consolidation dynamics that hit every infrastructure wave will apply. Some of today’s 100K-star repos will be absorbed into platform features. Some will become the next Datadog or Stripe — foundational infrastructure that earns its place in every stack.
But the thesis itself is durable: the moat is not the model, and it is not the agent either. It is the tool the agent reaches for — the memory it loads, the API it calls, the search layer it queries, the execution environment it trusts.
If you are building in the agent space, ask yourself one question: when an agent needs to do something in your domain, does it reach for your tool? If the answer is no, you do not have a moat. You have a demo.
The Agent Skills Marketplace Land Grab is the distribution side of this same thesis. And The Harness Is the Moat covers the orchestration layer. Read both for the full picture.








