In April 2026, Andrej Karpathy posted a gist on GitHub with just three words in the title: "LLM Wiki." No code, no package, just an idea. But this idea may be rewriting our fundamental assumptions about personal knowledge management.
His core argument is simple: Stop making AI start from scratch every time it looks for an answer. Let it build a continuously growing knowledge base for you.
RAG's Fundamental Problem: Reinventing the Wheel Every Time
The mainstream AI knowledge tools today — NotebookLM, ChatGPT file uploads, various RAG systems — all work roughly the same way: you dump a bunch of documents in, and when you ask a question, the AI retrieves relevant fragments and stitches together an answer.
This is called Retrieval-Augmented Generation (RAG), and it has a fundamental problem: knowledge never accumulates.
Karpathy described this dilemma perfectly:
"The LLM re-discovers knowledge from scratch on every question. There is no accumulation. Ask a nuanced question requiring synthesis across five documents, and the LLM has to re-find and re-assemble the relevant pieces every time. Nothing is built up."
This isn't just a theoretical concern. Academic research shows RAG systems achieve only 9-17% accuracy in real enterprise scenarios (arXiv:2604.02640). Even the best-performing NotebookLM still has a 13% hallucination rate; ChatGPT's document Q&A hallucination rate is as high as 40% (arXiv:2509.25498). Worse still, RAG suffers from "attribution drift" — a senator's personal opinion can, through RAG processing, become an unquestioned universal fact.
Barnett et al. (2024) identified seven failure points in RAG: missing content, ranking errors, context truncation, extraction failures, format errors, insufficient precision, and incomplete answers. Every one of them stems from the same structural flaw: chopping documents into fragments and reassembling them via vector similarity is inherently discarding context.
LLM Wiki: From "Retrieval" to "Compilation"
Karpathy's proposed alternative is conceptually closer to "compilation" than "interpretation" in programming language terms.
Traditional RAG is like an interpreter: it re-parses the source code every time it executes. LLM Wiki is like a compiler: it processes knowledge into structured intermediate artifacts at write time, and subsequent queries read directly from the compiled output.
Three-Layer Architecture
The entire system has three layers:
Raw Sources — the articles, papers, images, and data files you collect. This layer is immutable; the LLM only reads, never writes. This is your source of truth.
Wiki Layer — a directory of Markdown files generated and maintained by the LLM. Summary pages, entity pages, concept pages, comparative analyses, all with cross-references and backlinks. The LLM owns this layer entirely. You read it; the LLM writes it.
Schema Layer — a configuration file (like Claude Code's CLAUDE.md) that defines the Wiki's structure, conventions, and workflows. This is the key to transforming a general-purpose AI into a disciplined knowledge curator.
Karpathy used an elegant analogy: "Obsidian is the IDE; the LLM is the programmer; the Wiki is the codebase."
Three Core Operations
Ingest — feed in a new source, and the LLM doesn't just index it — it integrates key information into the existing Wiki: updating entity pages, revising topic summaries, flagging contradictions between new and old data. A single source typically touches 10-15 Wiki pages.
Query — when you ask a question, the LLM first reads index.md (a structured page directory), finds relevant pages, then synthesizes a cited answer. Critical insight: good answers can be saved back to the Wiki as new pages. Your exploration itself thickens the knowledge base.
Lint — periodically have the LLM run health checks: find contradictions, flag outdated claims, discover orphaned pages, fill in missing cross-references. This is the knowledge base equivalent of a code review.
From Memex to LLM Wiki: 80 Years of Knowledge Management Evolution
Karpathy explicitly referenced Vannevar Bush's 1945 Memex concept in his gist. This wasn't a casual citation — LLM Wiki truly is the technical realization of the Memex vision, 80 years later.
In 1945, Bush envisioned in "As We May Think" a desk-sized machine that could store personal information and link documents through "associative trails." His insight was that human thinking is associative, not hierarchically classified. But the problem he couldn't solve was: who does the maintenance?
Over the following 80 years, every generation of knowledge management tools tried to answer this question:
| Era | Tool/Method | Key Innovation | Unsolved Problem |
|---|---|---|---|
| 1945 | Memex (Bush) | Associative linking instead of hierarchical classification | Never built |
| 1950s | Zettelkasten (Luhmann) | Atomic notes + emergent structure | Required decades of manual maintenance (Luhmann accumulated 90,000 cards) |
| 1965 | Hypertext (Nelson) | Non-linear linked text | Project Xanadu never completed |
| 1995 | WikiWikiWeb (Cunningham) | Anyone can edit and link | Maintenance cost explodes with scale |
| 2017 | Second Brain (Forte) | Methodological standardization (CODE/PARA) | Still requires extensive manual organization; most of 500K readers abandon it midway |
| 2020 | Obsidian / Roam | Bidirectional links + local Markdown | Creating links is still manual |
| 2026 | LLM Wiki (Karpathy) | LLM handles all maintenance | Scale limits, hallucination accumulation (see below) |
Each generation lowered the "friction" of knowledge management, but none eliminated the fundamental bottleneck of "maintenance cost." LLM Wiki's breakthrough: it's the first to reduce maintenance cost to near zero.
As Karpathy put it: "Humans give up maintaining wikis because the growth rate of maintenance burden exceeds the growth rate of value. LLMs don't get bored, don't forget to update cross-references, and can modify 15 files at once."
Community Implementations: From Idea to Tool Ecosystem
Karpathy's gist had a deliberate design: he shared an idea, not code. Because in the Agent era, describing a pattern is more valuable than publishing a fixed GitHub repo — each person's LLM agent will customize the implementation to their own needs.
But the community quickly built out a tool ecosystem:
qmd (developed by Shopify CEO Tobi Lutke) — a local Markdown search engine combining BM25 keyword search, vector semantic search, and LLM re-ranking, all running on-device. 19,700 GitHub stars. It's not a Wiki builder but the search layer for LLM Wiki — when the Wiki scales beyond what index.md can handle, qmd takes over.
sage-wiki — an LLM Wiki compiler written in Go, appearing within a day of Karpathy's post. Five-stage progressive processing pipeline (diff, summarize, extract, write, images), with a strict typed entity system to prevent concept duplication. Benchmarks show keyword search latency of just 411 microseconds, graph traversal at 1 microsecond.
openaugi — converts an Obsidian vault into a SQLite knowledge graph, exposed to Claude via MCP. Supports bidirectional operation: the Agent can not only read but also write back to the vault. The developer's philosophy: "write-back is the key to knowledge compounding."
LLM Wiki v2 (rohitg00) — the most substantive extension of Karpathy's original design. Adds confidence scores (every fact gets a credibility label), knowledge decay (retention mechanism based on the Ebbinghaus forgetting curve), four-layer memory integration (working memory -> episodic memory -> semantic memory -> procedural memory), and multi-Agent synchronization. His core criticism of v1: "treats all Wiki content as eternally equally valid."
Comparison with Existing Tools: Nobody Is Doing "Knowledge Compilation"
Placing LLM Wiki on the spectrum of existing AI knowledge tools, it occupies a unique position.
NotebookLM — Google's flagship product, free with an excellent experience. But it's essentially a polished version of RAG: every conversation re-retrieves from your uploaded documents with no persistent knowledge synthesis. Audio Overviews (podcast-style summaries) are a killer feature, but knowledge doesn't accumulate across sessions.
Notion AI — workspace-level AI search, integrating Slack and Google Drive, multi-model support (GPT-5, Claude Opus 4.1, o3). Strongest collaboration capabilities, but AI is an add-on feature, not a core architecture. It answers questions but won't autonomously construct a knowledge graph.
Mem.ai — the closest tool to automated organization (automatic linking, temporal context), but still retrieval-oriented. It links notes but doesn't compile them into synthesized knowledge.
Khoj — open-source, self-hostable, supports multiple LLMs, 17,000 GitHub stars. But the underlying mechanism is still RAG retrieval.
Key observation: No tool on the market currently does what Karpathy calls "knowledge compilation." Every commercial tool stops at "let AI help you search" rather than "let AI help you build understanding."
A Hacker News user captured this distinction perfectly: "A vector database is only useful to a machine. You can't open a .faiss file and browse it. But a Wiki is useful to both humans and machines."
Honestly Facing the Limitations
LLM Wiki is not a silver bullet. In Hacker News discussions and actual user reports, several real issues have surfaced:
Scale limits. index.md starts breaking down at 100-200 pages — the file itself exceeds the model's context window. While models nominally support 200K tokens, research shows quality begins degrading around 130K tokens (the "lost-in-the-middle" phenomenon). Karpathy's own Wiki (approximately 400K words / 530K tokens) already exceeds the reliable range for a single model read.
Hallucinations get permanently embedded. When LLM-generated content builds on LLM-generated content, errors get "compiled" into the knowledge base. An HN commenter cited a Nature paper on "model collapse," warning this approach invites degeneration — "what's called the compounding effect may just be continuously rewriting valid information into lower-quality information."
The cost isn't trivial. One user reported that compiling three books (155K words) consumed approximately 12 million tokens. At Sonnet 4.6 rates, that's roughly 200. And Anthropic has step-function pricing at the 200K token threshold — a 199K token request costs 1.21.
The loss of learning itself. This is perhaps the deepest criticism. One HN user wrote: "The process of writing documentation is itself updating your own mental model." Another quoted Ezra Klein: "If you don't have something human to say, then don't say it." Outsourcing all knowledge organization to AI may mean sacrificing the very cognitive process you need most.
Not enterprise-ready. No access controls, no compliance-grade audit trails, the entire knowledge base is just a bunch of plain text files. Epsilla's enterprise analysis bluntly called this filesystem implementation "dangerously naive" for enterprises.
The Bigger Picture: From Bookkeeper to Curator
LLM Wiki's significance transcends a single tool or methodology. It represents a fundamental shift in knowledge work.
Tiago Forte — author of "Building a Second Brain" and one of the most influential thinkers in personal knowledge management — stopped teaching the BASB course in early 2023 because he felt his methods had "suddenly become obsolete." After three years of silence, he launched the "AI Second Brain" course in April 2026, with a core thesis:
"Personal Context Management is replacing Personal Knowledge Management — the new bottleneck isn't AI's capability, but your ability to give AI the right information."
This perfectly echoes Karpathy's view: the human role shifts from "organizing knowledge" to "curating context." You no longer need to manually build indexes, write summaries, or maintain cross-references — that's the LLM's job. What you need to do is: choose what's worth reading, ask the right questions, and judge what matters.
Microsoft Research's "Tools for Thought" project at CHI 2025 found a similar dynamic: "The higher the confidence in AI, the less critical thinking; the higher the confidence in one's own abilities, the more critical thinking." They recommend AI should play the role of "thinking partner" rather than "answer machine" — LLM Wiki fits this positioning exactly. It's not a tool that gives you answers; it's a system that helps you build understanding.
From a broader perspective, McKinsey estimates generative AI can unlock $4.4 trillion in annual productivity; 60-70% of knowledge work activities are technically automatable. But Harvard/Metaintro research shows that automation-intensive positions decreased by 17%, while human-AI collaborative positions increased by 22%. The future isn't AI replacing knowledge workers — it's knowledge workers' job description shifting from "bookkeeping" to "curation."
Conclusion: Your Memex Can Finally Be Built
In 1945, Vannevar Bush dreamed of a machine that could store, link, and organize personal knowledge. He described "associative trails" — human-curated links between documents — and foresaw that these links would be as important as the documents themselves. The one problem he couldn't solve: who does the maintenance?
80 years later, LLMs have answered that question.
Karpathy's LLM Wiki isn't a product, nor even a complete solution. It's a pattern — a design intuition about how AI should interact with human knowledge. It has real limitations: scale bottlenecks, cost thresholds, hallucination risks, cognitive trade-offs. But it points in the right direction: AI shouldn't just be an upgraded search engine — it should be your knowledge curator.
Under this paradigm, your job changes. You're no longer knowledge's bookkeeper, but its curator. You choose what to read, what to ask, what to track — and the LLM handles everything else.
As Karpathy said: "The most boring part of maintaining a knowledge base isn't the reading or thinking — it's the bookkeeping."
Now, the bookkeeper has clocked out. The age of the curator has begun.
References
- Andrej Karpathy — LLM Wiki (GitHub Gist) — Karpathy's original LLM Wiki pattern document
- Vannevar Bush — As We May Think (The Atlantic, 1945) — The classic article that first proposed the Memex personal knowledge machine concept
- Overcoming the "Impracticality" of RAG (arXiv:2604.02640) — Research showing RAG systems achieve only 9-17% accuracy in real enterprise scenarios
- Not Wrong, But Untrue: LLM Overconfidence in Document-Based Queries (arXiv:2509.25498) — Research measuring hallucination rates of NotebookLM (13%) and ChatGPT (40%)
- Barnett et al. — Seven Failure Points When Engineering a RAG System (arXiv:2401.05856) — Experience report identifying seven structural failure points in RAG systems
- Shumailov et al. — AI Models Collapse When Trained on Recursively Generated Data (Nature, 2024) — Nature study on model collapse
- qmd — Tobi Lutke (GitHub) — Local Markdown search engine by Shopify CEO, combining BM25, vector search, and LLM re-ranking
- sage-wiki — xoai (GitHub) — LLM Wiki compiler written in Go with a five-stage progressive processing pipeline
- openaugi — bitsofchris (GitHub) — Converts Obsidian vault into SQLite knowledge graph exposed to Claude via MCP
- rohitg00 — LLM Wiki v2 (GitHub Gist) — Extended LLM Wiki design with confidence scores, knowledge decay, and multi-layer memory integration
- Tiago Forte — The AI Second Brain — Forte's new course proposing Personal Context Management as the successor to PKM
- Microsoft Research — Tools for Thought (CHI 2025) — Research on AI's impact on critical thinking
- McKinsey — The Economic Potential of Generative AI — Report estimating $4.4 trillion annual productivity potential from generative AI
- Harvard Business School — Research: How AI Is Changing the Labor Market (HBR, 2026) — Research showing 17% decrease in automation-intensive roles and 22% increase in human-AI collaborative roles
- Epsilla — Did Karpathy's LLM Wiki Just Kill RAG? The Enterprise Verdict — Enterprise perspective analysis of LLM Wiki limitations


