This week, two very different people went viral for the same reason. Andrej Karpathy, arguably the most credible AI engineer alive right now, shared how he built a self-maintaining personal wiki using LLMs. 41,000 people bookmarked it. A few days later, Milla Jovovich, of Resident Evil fame, co-launched MemPalace, an open-source AI memory tool inspired by the ancient "memory palace" technique used by Greek scholars. It hit 23,000 GitHub stars in two days before the developer community started tearing apart the benchmark claims.
Two very different people. Same underlying frustration: AI keeps forgetting things, and our own knowledge stays scattered.

Here's what's funny to me: I first heard this exact problem framed as a strategic challenge about 20 years ago, in a computer science lecture. My professors were emphatic about it. Knowledge management is one of the most important competitive advantages a company can build, and also one of the hardest, because it's not a tooling problem, it's a cultural one. You can buy the best software in the world and it won't help if people don't actually use it.
And they were right. Over the years, working with companies of all sizes, I kept running into the same pattern. Where do we document this? Debates about whether something belongs in the code, in the wiki, in the shared drive. SharePoint deployments that nobody touched after the first month. Notion workspaces that started beautifully organized and, six months later, looked like a junk drawer. The search tools, the tagging systems, the folder structures. Always a pain, never really solved.
Now we have AI systems trained on the sum of human knowledge, and somehow the problem is still here. Just wearing a different hat.
Modern LLMs have consumed effectively everything humanity has written. And yet, they cannot reliably remember what you told them last Tuesday.
The Paradox at the Heart of This
Here's the thing that genuinely puzzles me, and I think it should puzzle you too.
Modern LLMs have consumed effectively everything humanity has written. Every book, paper, forum post, code repository, and article that was publicly available at the time of training. They can explain quantum physics, write legal documents, and debug code in languages that barely existed five years ago.
And yet, they cannot reliably remember what you told them last Tuesday.
This sounds like a joke, but it's one of the core practical limitations that everyone building with AI hits sooner or later. And it's not because the models are dumb. It's because the problem of storing and retrieving personal, contextual, evolving knowledge is genuinely hard, in ways that are different from storing static training data.
Training data is curated, structured, and mostly frozen in time. Your knowledge isn't. It grows, contradicts itself, gets outdated, references things that only make sense in context. Figuring out what to keep, what to update, when to update it, and how to surface the right thing at the right moment, that's the hard part. And it turns out that's as hard for AI systems as it always was for humans.
What I've Actually Tried
I've been running my own experiments on this for a while, and I want to share what I've found, not because I've solved anything, but because I think the honest account of what works and what doesn't is more useful than another theoretical framework.
Claude Projects was my first serious attempt. When it launched, I moved everything there: context about the newsletter, preferences, how I like things written, ongoing projects. And it worked really well, at first. The first couple of months felt like a genuine step change. The responses were clearly informed by everything I'd put in.
Then things started degrading. By month three, I noticed the memory was getting stale. Facts that had changed were still being referenced as current. Details I'd explicitly shared weren't showing up in relevant conversations. I went into the memory settings and found a lot of entries that simply hadn't been updated, not because the system was broken, but because nobody, human or AI, had gone in to maintain them. The same cultural problem my professors warned me about, now playing out inside my AI setup.
OpenClaw was my next experiment. If you haven't come across it yet, it's the open-source personal AI agent that went viral earlier this year (formerly Clawdbot, then Moltbot, long story). It's essentially a persistent AI assistant that lives on your machine and talks to you through Telegram, WhatsApp, or whatever messaging app you use. I set one up, gave her a name (Rhea) and a personality, and she became genuinely useful for a lot of things.
But the memory problem followed me there too. I'd share something important in a conversation, come back a week later, and it simply wasn't there. The default memory structure OpenClaw ships with puts everything into one file, which starts feeling unwieldy fast. So I built a custom repo to manage Rhea's memory instead, breaking it into separate files by topic so I could actually audit what was in there and so the agent could write to specific places more accurately.
That helped. But the real improvement came when I added a scheduled job that ran every day. Every 24 hours, the agent would go back through our recent conversations, pull out anything that seemed worth remembering, and add it to the right memory files. It wasn't perfect, some things still slipped through, some entries ended up slightly wrong. But it moved the needle more than anything else I'd tried.
The lesson I took from that: passive memory doesn't work. You need an active process that goes back and harvests knowledge, not just a container that things get dropped into when someone remembers to do it.
What we call 'memory' in AI systems is really just structured context that gets injected at the start of a conversation.
This Is Bigger Than Memory
The more I sat with this, the more I realized that what we're actually talking about isn't just a memory problem. It's a learning problem.
Humans learn continuously. Every conversation, every experience, every mistake updates something in how we understand the world. We don't need to manually schedule a job to extract lessons from our day, it happens automatically, imperfectly, but constantly.
Agents don't work like that. They have their capabilities up to a training cutoff, and then everything after that is context you give them. The model itself doesn't change. What we call "memory" in AI systems is really just structured context that gets injected at the start of a conversation, a way of making the agent seem like it knows more than it did last time you spoke.
That's a meaningful distinction. Better knowledge management doesn't make agents actually learn, it makes the context they receive better. Which is still genuinely valuable, especially as the systems we're building become more complex. But it's worth being honest that it's a workaround, not a solution to the underlying gap.
The gap is this: we need agents that can build up a model of the world they operate in, update it as things change, and reason about what they know and don't know. That's much closer to how human expertise actually works than anything we have today.
What's Actually Happening in This Space
The approaches people are converging on right now fall into three fairly distinct camps, and they represent genuinely different philosophies about the problem.
The file-based camp (Karpathy, OpenClaw, plain markdown). Instead of using a RAG pipeline to search chunks of raw documents, the idea is to have an LLM actively maintain a structured wiki: human-readable articles, each focused on a single topic, with an index file the agent reads to decide what to pull for any given query. The wiki gets updated and "linted" regularly, contradictions get flagged, gaps get filled. It's active maintenance by design, not passive storage. Simple, auditable, and surprisingly effective, but it requires discipline to maintain and hits limits as knowledge grows.
The memory-as-a-layer camp (Mem0 and similar tools). You keep your existing agent framework and bolt a dedicated memory service on top of it. The interesting development here is the shift from pure vector memory to graph memory. Vector memory retrieves things that are semantically similar to your query. Graph memory retrieves things connected through relationships, which is actually much closer to how human recall works. If you know that someone used React until last year and switched to Vue, a vector search might return both facts with equal weight. A graph query can encode the temporal relationship and return the current state correctly. Mem0 added graph memory in early 2026, and it's a meaningful step up for anything involving preferences, people, or evolving context.
The memory-as-a-runtime camp (Letta, formerly MemGPT, from a UC Berkeley research project). This is the most architecturally radical approach. The idea is to treat the LLM like an operating system: core memory is like RAM, always in context, small and fast; archival memory is like disk, queried on demand when needed. The agent actively manages its own memory through tool calls, deciding what to promote into core memory and what to archive. It's more complex to set up, but it solves the degradation problem by design rather than hoping someone remembers to maintain the files.
One pattern that cuts across all three camps is what Letta calls "sleep-time compute": a background process that goes through recent conversations asynchronously and consolidates them into structured memory, without blocking the main agent loop. I find this validating because, as I mentioned earlier, independently building a daily scheduled job for Rhea that did exactly this was the single thing that moved the needle most in my own experiments. It turns out there's a name for it, and it's becoming standard practice.
MemPalace sits somewhat apart from these camps, drawing on the ancient "method of loci" technique where information is organized into a spatial architecture of wings, halls, and rooms. The project launched claiming a perfect score on the LongMemoryEval benchmark, which the developer community promptly tore apart. Within 24 hours the score had been revised down and accusations were flying about who actually built what. But honestly, the controversy is almost beside the point. The fact that 23,000 developers starred a memory tool co-built by a Hollywood actress in two days tells you something about how hungry people are for a real solution here. The problem is real enough that even a flawed answer gets that kind of response.
What all of these approaches share, and what my own experiments kept confirming, is the same basic insight: you need structure, you need active maintenance, and you need to design explicitly for retrieval, not just storage.
Where to Start if You're Building with Agents Today
If you're running agents and knowledge is already a problem for you, here's what I'd actually recommend based on what's worked:
Start by auditing what you already have. Before adding more tooling, look at what your agent currently knows. If you're using Claude Projects, go into the memory settings. If you're using OpenClaw or a similar setup, open the memory files. You'll probably find things that are outdated, contradictory, or just wrong. Clean that up first.
Separate your memory by topic. A single memory file or a flat context block is hard to maintain and hard for the agent to use well. Break it into logical areas, each one short enough that a human could read and verify it in a few minutes.
Build an active retrieval loop. Don't rely on the agent passively accumulating knowledge during conversations. Set up something, even a simple daily prompt, that goes back through recent exchanges and explicitly asks: what here is worth remembering? What contradicts something we already have? This is the highest-leverage thing I've done.
Write for machines, not just humans. This sounds obvious but it's easy to get wrong. Your knowledge entries should be unambiguous, specific, and single-topic. "Pedro prefers a conversational tone" is harder to use correctly than "Newsletter articles should not use em dashes, bullet points should only appear when the content is genuinely list-like, and the opening paragraph should not reference AI tools as amazing or transformative." The more specific you are, the more reliably the agent applies it.
A messy knowledge base doesn't just slow down a human, it actively misleads an agent, producing confident wrong answers that are worse than no answer at all.
The Bigger Picture
I think we're in an interesting transitional moment. The tools for building agents have matured a lot. The models are genuinely capable. But the knowledge infrastructure underneath them is still being figured out, and it shows.
The people going viral right now for building personal knowledge bases, Karpathy, Jovovich, the whole second brain ecosystem, are pointing at something real. The problem my professors identified 20 years ago hasn't gone away. It's just become more urgent because now the stakes are higher. A messy knowledge base doesn't just slow down a human, it actively misleads an agent, producing confident wrong answers that are worse than no answer at all.
Getting this right is, I'd argue, one of the more important unsexy problems in AI right now. It's not as exciting as a new model release, but it's what determines whether the agents you build are actually useful six months after you launch them.
We don't have a clean solution yet. But we're getting closer to understanding the shape of the problem. And that's usually where the good work starts.
References & Further Reading
Knowledge Management & Memory Systems
- Andrej Karpathy. LLM Wiki (April 2026)
- Cybernews. Milla Jovovich creates MemPalace AI memory tool (April 2026)
- Corey Ganim. Karpathy's Second Brain clearly explained (April 2026)
Agent Memory & Context Engineering
- RAGFlow. From RAG to Context: A 2025 Year-End Review
- Mem0. Context Engineering in 2025: The Complete Guide
- Regal AI. The RAG Playbook: Structuring Knowledge Bases for AI Agents
OpenClaw & Personal Agents
- OpenClaw. Official Documentation
- freeCodeCamp. How to Build and Secure a Personal AI Agent with OpenClaw
Broader Context
- Google Cloud. Lessons from 2025 on Agents and Trust
- Composio. Why AI Agent Pilots Fail and the 2026 Integration Roadmap