
A Second Brain That Maintains Itself: Claude Fable 5 + Obsidian
A Second Brain That Maintains Itself: Claude Fable 5 + Obsidian
Hundreds of bookmarks you will never reopen. Dozens of AI conversations that scrolled away. That is not accumulating knowledge — it is accumulating storage.
The difference is whether anything gets compiled on the way in. A folder holds what you put in it. A wiki gets denser: every new source has to be reconciled against what is already there, which is exactly the work nobody does on a Sunday evening.
This is the setup where that reconciliation is handed to an agent — an Obsidian vault whose wiki/ layer is written and maintained entirely by Claude Fable 5, following the LLM Wiki pattern Andrej Karpathy published as an idea file meant to be handed straight to an agent.
Why not just RAG
Karpathy's framing is the sharpest part, and it is worth stating before any of the mechanics.
Upload files, retrieve relevant chunks at query time, generate an answer — that is how most document setups work, and nothing accumulates. Ask a question that needs five documents synthesised and the model finds and re-assembles those fragments from scratch, every time.
A wiki inverts the timing. The synthesis happens once, on the way in: the cross-references are already written, the contradictions already flagged. Every later question reads a structure that got denser with each source, instead of re-deriving one from raw text.
His own working arrangement is the one used here — agent on one side, Obsidian on the other, watching the graph redraw as pages land:
Obsidian is the IDE; the LLM is the programmer; the wiki is the codebase.
— Andrej Karpathy, LLM Wiki
About this write-up
Written from my own vault, not from theory — the folder counts, the contract, and the log entries quoted below are the real files. Where the system has limits, they are named.
The one idea that makes it work
Most "AI second brain" setups fail the same way: the agent keeps writing pages, nothing ever gets merged, and six months later you have four pages about the same tool under four names.
The fix is not a better prompt. It is ownership, written down.
This is the written companion to the video, where the vault is built from an empty folder and the first two sources are ingested end to end — including the graph view redrawing itself.
Prefer to watch it on YouTube?
Watch on YouTube — or subscribe to the channel for the next one.
Three layers, and the boundary between them is enforced by a contract file the agent reads at the start of every task:
| Path | Owner | Rule |
|---|---|---|
inbox/ | you | landing zone for anything not yet ingested |
raw/ | you | immutable source documents — read-only to the agent |
wiki/ | the agent | every generated page; the agent owns this layer entirely |
raw/ being off-limits is the whole trick. The agent may write as much as it likes, and still cannot alter the evidence that every one of its claims has to cite.
The five rules that do the real work
The contract is long, but five rules carry it.
Provenance — no source, no claim. Every factual statement carries an inline wiki-link to the file in raw/ it came from. "If you cannot trace a statement to a file in raw/, do not write it." Its own synthesis is allowed, but must be labelled as synthesis.
Compression — merge, don't duplicate. A new source earns its place by merging into what exists, not by piling up pages. Update the existing page and rewrite the affected sections so it stays "a coherent current-best synthesis, not a chronological scrapbook."
Conflict — never silently overwrite. When a new source contradicts a page, both claims stay, each with its citation, under a ⚠️ CONFLICT callout, and it goes to the human to resolve.
Token efficiency — read the index, not the vault. Start from index.md, use its one-line summaries to decide which pages are relevant, open only those. Full-vault scans are for lint passes only.
Destructive edits need approval. Additive fixes can be biased toward action; merges that delete pages and removals always wait for a human.
Three workflows
INGEST — one source at a time. Read it, discuss the takeaways before writing, search the index and the alias registry for overlap, decide new-page-versus-merge for each topic and say why, then write, update the index, and append to the log.
QUERY — answer from the wiki, citing the pages used. If the wiki is thin on the topic, say so rather than guessing. Flag anything whose updated date is more than 90 days old as possibly stale. Notably, it does not go to the internet — answers come from what you actually ingested.
LINT — offered every tenth ingest. Scan for near-duplicates, broken links, orphans, contradictions, dead stubs, stale pages, and gaps: topics that keep recurring across sources without a page of their own.
What two sources actually produced
The vault at the time of writing has ingested exactly two things: a web article comparing AI coding agents, and a 93-page PDF on building agents.
Sources in raw/ | 2 |
wiki/concepts/ | 11 pages |
wiki/entities/ | 12 pages |
wiki/summaries/ | 2 pages |
So one article and one book became 25 interlinked pages — and the second ingest did not just append. The log records it merging into the first: ai-coding-agents gained a parent concept, and two existing pages gained cross-links to the new evaluation and framework pages.
That merge behaviour is the thing worth checking on your own vault. It is what separates a wiki from a folder.
The log is the part I did not expect
log.md is append-only, newest first, never edited. It is a plain audit trail — and the entries flag their own weaknesses:
Flagged: vendor source (Galileo) and ~late-2024 vintage — framework verdicts likely stale.
…noted Firecrawl marketing bias and a gap (no dedicated MCP source yet).
An ingest that records "this source is a vendor selling something, and it is old" is doing the job an honest research assistant does. It is the same instinct as the quality-gate argument in the SDLC notes: a green result is only worth what its evidence is worth, and the evidence here is a marketing page.
The lint pass is similarly blunt — it reported 0 broken links, 0 orphans, 0 stale pages, and then flagged a stray empty file in the vault root and two topics recurring across both sources with no page of their own.
The contract, in full
This is the file the whole system runs on. It sits at the vault root and the agent reads it first, every time. Copy it, change the domains and the folder names to suit you, and point an agent at an empty vault.
`CLAUDE.md` — the full schema and contract (click to expand)
It contains markdown tables and fenced examples, so it is left exactly as-is rather than re-wrapped — reflowing it would corrupt the tables. Copy it whole.
# Second Brain Wiki — Schema & Contract
This vault is a personal knowledge base on **IT, AI, DevOps, and Software Engineering**,
maintained by an LLM agent (you) on the three-layer pattern: raw sources → wiki → this schema.
The human curates sources and asks questions; you do all writing, filing, and maintenance.
All wiki content is written in **English**.
## Layers & folders
| Path | Owner | Purpose |
|---|---|---|
| `raw/` | Human | Immutable source documents. **Read-only for you — never edit, move, rename, or delete anything here.** |
| `raw/assets/` | Human | Images/attachments belonging to raw sources. Also immutable. |
| `inbox/` | Human | Landing zone for not-yet-ingested material. After ingest, the human (or you, with permission) moves the file to `raw/`. |
| `wiki/` | You | All LLM-generated pages. You own this layer entirely. |
| `wiki/summaries/` | You | One page per ingested source. |
| `wiki/entities/` | You | Proper nouns: tools, products, models, companies, people, platforms (e.g. `claude-code.md`, `kubernetes.md`). |
| `wiki/concepts/` | You | Ideas, techniques, practices, patterns (e.g. `prompt-caching.md`, `blue-green-deployment.md`). |
| `wiki/aliases.md` | You | Terminology registry (see below). |
| `index.md` | You | Live map of the whole wiki, grouped by domain. Updated on every change. |
| `log.md` | You | Append-only chronological record. Never edit past entries. |
### Naming rules
- Everything **kebab-case**, `.md` extension.
- `raw/`: `YYYY-MM-DD-<slug>.md` (date = date added, slug from the title). PDFs keep their extension: `YYYY-MM-DD-<slug>.pdf`.
- `wiki/summaries/`: same slug as the raw file it summarizes: `YYYY-MM-DD-<slug>.md`.
- `wiki/entities/` and `wiki/concepts/`: canonical name only, no date: `github-actions.md`, `retrieval-augmented-generation.md`. Check `wiki/aliases.md` before naming.
### Domains
Primary tag taxonomy (extend as needed, record extensions here):
`ai-tools`, `ai-models`, `ai-workflows`, `software-engineering`, `devops`, `cloud`, `containers`, `ci-cd`, `infrastructure`.
## Page schema
Every wiki page has this shape:
```markdown
---
title: Prompt Caching
aliases: [prompt cache, KV cache reuse]
tags: [ai-models, ai-workflows]
status: draft # stub | draft | stable
sources: [2026-07-05-anthropic-caching-docs]
updated: 2026-07-05
---
**One-line summary of the page, readable on its own.**
<body: sections, claims with citations, wiki-links>
## Related
- [[parent-or-sibling-pages]]
```
- The **first body line** (right after front-matter) is a one-line bold summary. Read this line
to judge relevance before loading a full page. (It sits after the front-matter, not before,
because Obsidian only parses YAML at the very top of a file.)
- `status`: `stub` = placeholder, needs content; `draft` = has content, single-source or unreviewed;
`stable` = multi-source, reviewed by the human.
- **Backlinks:** every concept page links to at least one parent (a broader concept or domain hub
page). No orphan pages.
## Provenance rule — no source, no claim
Every factual claim cites its raw source inline as a wiki-link to the raw file, e.g.:
> Claude Code supports hooks for intercepting tool calls ([[2026-07-05-claude-code-docs]]).
- Citations link to files in `raw/` so they are clickable in Obsidian.
- If you cannot trace a statement to a file in `raw/`, do not write it. Your own synthesis is
allowed but must be marked as such (e.g. "Synthesis:") and must reference the pages it draws on.
## Compression rule — merge, don't duplicate
A new source earns its place by **merging** into existing knowledge, not by piling up pages.
- Before writing anything, search `index.md` and `wiki/aliases.md` for overlap.
- If a topic already has a page: **update that page**, and note in its body or the log what
changed and why. Do not append blindly; rewrite sections so the page stays a coherent
current-best synthesis, not a chronological scrapbook.
- Never create a near-duplicate page. Two pages about the same thing under different names is
the failure mode this whole rule exists to prevent.
## Terminology registry — `wiki/aliases.md`
A table mapping every known alias to its one canonical page:
```markdown
| Alias | Canonical page |
|---|---|
| RAG | [[retrieval-augmented-generation]] |
| retrieval augmented generation | [[retrieval-augmented-generation]] |
| K8s | [[kubernetes]] |
```
- **Consult it on every ingest** before creating or linking pages.
- When you coin a canonical page name, register the name and all aliases you've seen.
- Also mirror aliases into each page's `aliases:` front-matter so Obsidian resolves them.
## Conflict rule
When a new source contradicts an existing page, **never silently overwrite**. Instead:
1. Keep both claims on the page, each with its citation, under a `> ⚠️ CONFLICT:` callout
explaining the disagreement.
2. Flag it to the human in your ingest report and let them resolve it.
3. Record the conflict in `log.md`.
## Token-efficiency rule
- Start every task by reading `index.md` (and `wiki/aliases.md` when ingesting).
- Use the one-line summaries in the index to decide which pages are relevant; open only those.
- Never re-read the whole vault by default. Full-vault scans are for lint passes only.
## Workflows
### INGEST (one source at a time, human in the loop)
1. Read the source (from `inbox/` or `raw/`). If it references local images, read the text first,
then view the images that matter.
2. Discuss key takeaways with the human; let them steer emphasis before writing.
3. Search `index.md` + `wiki/aliases.md` for overlap. Decide new page vs merge for each affected
topic, and say why.
4. Write the summary page in `wiki/summaries/`; create/update every affected entity and concept
page (a rich source may touch 10–15 pages). Ensure the file lands in `raw/` with the standard name.
5. Update `index.md` and `wiki/aliases.md`; append to `log.md`.
6. Report in ~3 lines: what was written, what was merged, anything flagged (conflicts, gaps).
### QUERY
- Answer **from the wiki**, citing the pages used (which themselves cite raw sources).
- If the wiki is thin or silent on the topic, say so — flag the gap rather than guessing.
- Flag pages relied on whose `updated` is more than 90 days old as potentially stale.
- If an answer produced durable new synthesis (a comparison, an analysis), offer to file it back
into the wiki as a page.
### LINT (on request, or offer after every 10th ingest)
Scan the wiki for: near-duplicate pages, broken wiki-links, orphan pages (no inbound links),
contradictions between pages, dead stubs, stale pages, and **gaps** — topics recurring across
multiple raw sources that lack a dedicated page. Propose fixes and new-page candidates;
**apply only after human approval.** Additive fixes may be biased toward action; destructive
edits (merges that delete pages, removals) always require approval first.
### LOG format (newest first, never edit past entries)
```markdown
## [2026-07-05] ingest | Article Title
Pages touched: [[summary-page]], [[entity]], [[concept]] — one-line why.
```
Actions: `ingest` | `update` | `merge` | `lint` | `query-filed` | `schema`.
## Meta
- This schema is co-evolved with the human: when a workflow repeatedly fails or a convention
proves awkward, propose an amendment here (action `schema` in the log) rather than silently
deviating.
- Stay quiet about your own mechanics during normal work unless asked.Setting it up
- Create a new vault in Obsidian (Manage vaults → Create), somewhere you can also open in an editor.
- Open that same folder in VS Code — you want the file tree and a terminal beside the graph.
- Install the Claude Code CLI and run
claudein that terminal, with the vault as the working directory. Select the strongest model you have. - Give it the pattern and ask for the schema. It will ask you questions first — domain, source types, query style, how involved you want to be, housekeeping preferences. Answer them properly; those answers become the contract.
- Let it write
CLAUDE.md, then the folder structure,index.md, andlog.md. - Drop your first source in
inbox/and ask for an ingest.
Put the vault under git. git status after an ingest is the fastest way to see exactly what the agent touched, and it is the only undo you have.
What I would watch
It only knows what you feed it. The QUERY step deliberately does not reach the internet. That is a feature for trust and a limit on coverage — a thin vault gives thin answers, and it will tell you so.
Merging is the hard part, and it degrades quietly. Two pages about one thing under different names is the failure this design exists to prevent, and the only thing standing between you and it is the lint pass. Run it.
Two sources is not a test. Everything above is true of a 25-page vault. The interesting question is what the compression rule does at three hundred pages, and I do not have that evidence yet.
Related
- Karpathy's LLM Wiki gist — the original pattern, written to be pasted into an agent.
- Fable Mode: Getting a Frontier Model to Write the Manual for Its Replacement — the same move at a different scale: get the strong model to write the rules, then let something cheaper follow them.
- Session 4 — AI-Engineering Governance — why provenance beats confidence, worked through a real failure.