---
title: "The Knowledge Base My Agents Actually Keep Up"
summary: "A Markdown-and-Git knowledge base needs upkeep. Here is how I divide the work between tools and agents, from filing and retrieval to corrections and retirement, and where that division still breaks."
author: "Tim Sherstyuk"
author_title: "Founder"
publisher: "General Frontier"
published: 2026-09-28
canonical: https://generalfrontier.com/writing/the-knowledge-base-my-agents-actually-keep-up
---

# The Knowledge Base My Agents Actually Keep Up

Keeping written knowledge in step with the work it describes is the job that kills most setups. Dex Horthy made the sharpest version of that case in [an August post about spec-driven development](https://x.com/dexhorthy/status/2090091396917248209). He was talking about spec-driven development, where you keep written specs in sync with the code, and his verdict was that "you basically have just manufactured a new problem that doesn't actually yield that much leverage."

He wasn't talking about knowledge bases for agents. But the cost he names is the one anyone running one will recognize: whatever you write down starts drifting the moment the work moves, and someone has to keep pulling it back.

I've been building one of these anyway. In my own words from a few weeks ago:

> Everyone has seen Karpathy's LLM-wiki idea: keep your notes in files, let an agent maintain them into a handbook, and sessions stop starting from zero. I built a version of that; now at 12,000+ .md file in one Markdown + Git workspace, maintained with the coding agents I use daily.
> The files alone changed nothing. My agents still started every session lost, because nothing decided what they read, or when.

(It now has more than 18,000 tracked Markdown files.) Horthy's objection is fair. Tools can file records, track source changes, and flag what needs attention. An agent still has to decide what changed and revise the knowledge. My view: tools do the mechanics, agents do the judgment, and both work inside a structure and process. The tools are designed to make the right way to do the work the easy way. Where a check is built into a tool, it can refuse a wrong move without relying on the agent to remember the rule.

The general split isn't new. Garry Tan describes [latent versus deterministic work](https://x.com/garrytan/status/2042925773300908103): the model interprets and decides; code handles repeatable operations. Here I'm applying that distinction to the upkeep of a knowledge base. [Karpathy's post](https://x.com/karpathy/status/2039805659525644595) says: "the LLM writes and maintains all of the data of the wiki, I rarely touch it directly." My agents write too. Tools take over specific parts of the bookkeeping.

Everyone's building a knowledge base for their agents. Here's how to build one they actually keep up: let the tools do the bookkeeping, and leave the judgment to the agent.

One limit up front. This is one person's system: plain Markdown files in Git, maintained by the coding agents I use every day, often several at once. I'm not claiming it's faster or better than a simpler setup. I'm showing where the split sits at each stage, and where it still breaks, including one rule I still haven't gotten agents to follow.

Figure: The life of a piece of knowledge: gets in, gets found, gets used, gets corrected, gets retired. Tools file and validate records, search and return pointers, check supported work paths, flag changed sources, and filter retirement markers they recognize. Agents decide what is worth keeping, inspect sources, choose and do the work, interpret corrections, and decide what is replaced.

*Tools handle specific parts of the bookkeeping; agents supply and revise the meaning. Concurrent work adds coordination at the point where changes are made.*

## "Everything I save goes into one pile"

This stage is short because nothing dramatic happened here. When an agent saves a record through the relevant tool, it goes into an area with its own rules instead of one shared pile. Research findings, notes on outside projects, project decisions, and short lessons each live in their own area with their own rules for what gets in. The agent supplies the content and labels: what the record is about and what it claims. The tool writes the file and checks the required metadata. Source relationships are recorded alongside it.

The area for short lessons is the narrowest. It takes only things like dead ends (what we tried and why it failed), patterns that held up, and corrections to documentation. Each entry has to say what it couldn't find, and an agent can never mark its own claim as high confidence. What the agent decides is whether something is worth keeping at all, and what it actually means.

## "It re-decided something we already settled"

If your agent keeps reopening settled questions, it probably isn't reading the answer at the moment it needs it. The usual fix is a bigger rules file. Mine went the other way.

The file every agent loads at the start is a map, about 200 lines. Most of it points to where the depth lives: this kind of work, start with that guide. Not all of it is routing. About a quarter is instructions for starting and finishing work, which is the part I trust least in that file, for reasons the next section covers. None of this is new. Karpathy's wiki idea, Garry Tan's ["routing table for context,"](https://x.com/garrytan/status/2044479509874020852) and a lot of public "map, not manual" advice all say the same thing.

The depth arrives mostly through tools. In sessions with the prompt hook enabled, a small script searches short summaries of lessons on each message and injects the matching claims, each labelled with how much it has been verified. When an agent resumes a project, the initial project read gives a compact summary and pointers to the next files to open.

The agent's job is to read a hit as "go inspect this," not "this is true." That rule isn't mine either. [@himanshustwts put it plainly](https://x.com/himanshustwts/status/2038924027411222533) when writing about Claude Code's memory: "memory is a hint, not truth."

## "I wrote a rule and my agent ignored it"

This is the stage where the split matters most, and the one with the most failures behind it.

### The rule that only asked

I built a docs tool, fugg, that lets an agent look up current library documentation instead of relying on what it remembers from training. By April it was listed in the always-loaded file. During one long session that month (13 commits, including work on libraries the tool covered) I added an explicit line to that file: when you write code against these libraries, check the docs instead of relying on training data.

The agent never used the tool once. At the end, it said so itself: "I didn't use fugg once during the session."

That wasn't a clean test. The file loads at the start, and the explicit line came in partway through. But the tool had been in that file from the first message, and the line changed nothing.

Looking back on it later, my read on why:

> yeah, it was annoying that fugg was not used because I was expecting it to get used. I think the problem that happened was that it seems like, from when we investigated this, agents were just overly confident in the things that they knew or thought that they knew.

That's my impression, not something I measured. In June I tried again with a skill, a packaged instruction the agent loads when a task matches it. Of that attempt:

> That seemed to help a little bit, but still, I think it didn't address the main foundational thing there

"A little bit" is also an impression, not a measurement. The ticket for this is still open. I haven't solved it.

### The rule that refused

Compare a rule enforced by a tool. Several agents share one copy of my repository, the main checkout. Each one is supposed to do its work in its own separate copy (a git worktree: a second checkout of the same repo), so their changes can't mix. Agents kept writing into the shared copy anyway. Back in August I was still correcting this by hand:

> No, you should 100% be using the project tool to write in the worktree.

The first idea was a warning that noticed when an agent drifted into the wrong copy. Before building it, we checked 30 days of past sessions. It would have fired about 30 times a day, and still about 15 after narrowing. A warning that fires that often turns into noise, so it was never built.

What shipped on 20 August was a refusal at write time: the tools that create tickets and project records refuse to write into the shared copy. A check at commit time already existed; the new refusal catches the mistake earlier. The refusals also point the agent toward the separate workspace it should use. The agent doesn't have to remember the rule. The wrong move fails, and the error hands it the right one.

The April rule asked the agent to remember. The refusal sits where the work already passes through. These controls do different jobs, though: a write guard can block a forbidden action on the paths it covers. It doesn't force an agent to take an omitted action, such as looking up the docs.

### Where the refusal doesn't reach

It isn't complete. Direct file edits and several tools were outside the write-time refusal. The commit-time check is a later guard, not protection against every earlier write. On 22 September, a tool that saves video transcripts ran before the agent had started its own copy and wrote six record files into the shared one. Nothing stopped it. That tool was fixed the next day. The others are still on an open list.

### Tools that won't guess

The other half of the split is what tools refuse to decide. Say a project has three unfinished threads, each with a note one session left for the next. An agent asked to "pick up where we left off" will happily resume one of them, usually the newest, which may not be the one you meant. The project entry lists the open threads and makes the ambiguity visible. Its operating rule requires a choice before resuming; a suggested starting point is not a decision about which work matters. Which thread matters is a judgment call, so it stays with the agent, or with me.

## "Two agents stepped on each other"

At the point where knowledge gets changed, concurrent work adds another problem. Separate copies stop agents from overwriting each other's files. They don't stop agents from slowing each other down when their work lands. This was me in August, after a long check run was thrown away because another agent had landed something unrelated ("mastery" is dictation for master, the main branch):

> Specifically the absolute most annoying thing happens when a repo runs the closing test or whatever and then mastery gets updated and has to rerun all that shit. This is incredibly annoying and slows down our workflow like crazy. We need to have a more consistent coherent way to be able to do this properly. When multiple agents are working in the same repo, all that shit should not be getting in the way dude.

The design that reset the landing step was accepted the same day, and its list of incidents names this exact case: agents rerunning their full task checks every time the main branch moved. The fix split the checking into two jobs. Each agent checks its own changes. When someone else lands first, a short repository safety check runs on the combined result; unrelated movement on the main branch doesn't automatically repeat the full task test run. That check is a safety floor, not proof of every possible interaction. Changes that affect the task's behavior can still need more testing. A September change added a queue: agents finishing their work now take turns in the order they arrived, and the tool combines their changes itself when they don't clash.

What the tool doesn't do is resolve a conflict. Those come back to the agent as ordinary git work, along with the call on whether a change deserves more testing.

## "I corrected something and it still gets quoted"

Notes go stale. The obvious fix is to have the agent regenerate or re-lint them on a schedule, and most of the public knowledge-base pieces I've saved do some version of that. Karpathy's setup has the model maintain the whole wiki. Others run a weekly lint or a nightly consolidation pass. Some push parts of it into plain code: [GBrain's auto-linking](https://github.com/garrytan/gbrain) uses pattern matching rather than model calls.

I split it differently. A few maintained explanation pages track the exact files they were built from. When one of those files changes, the tool marks the page stale. It never rewrites the page. An agent rereads the sources and decides what the change means.

The failure here is instructive. A research finding was formally refuted on 4 September, with the reason and a successor recorded on the finding itself. On 21 September, a fresh audience scan in the same repository still repeated its numbers. The correction reached the record. It didn't reach everything that had copied from it.

## "I retired a note and it keeps coming back"

Short, because the honest answer is a gap. The system has a marker for "this has been replaced." The search that feeds lessons into every prompt honours it and skips those files. Another part of the system reads a different field. So one retirement is respected by some readers and not others. One retirement marker does not yet guarantee that every consumer stops using a note. The tool hides what's marked. The agent decides what's retired and why, and for now it has to know the marker doesn't reach everywhere.

## The split, stage by stage

| Stage                      | The problem you'll notice                             | What the tool does                                                                                    | What the agent decides                                        |
| -------------------------- | ----------------------------------------------------- | ----------------------------------------------------------------------------------------------------- | ------------------------------------------------------------- |
| Gets in                    | Everything goes into one pile                         | Files supported records by kind, validates metadata and confidence, records source links              | Whether it's worth keeping, and what it means                 |
| Gets found                 | It re-decides something already settled               | Searches in sessions with the prompt hook; returns a compact project entry with source pointers       | Whether a hit is actually true, by opening the source         |
| Used at the moment of work | You wrote a rule and it got ignored                   | Refuses covered wrong moves and prints the fix; exposes ambiguous options                             | Which option; the work itself                                 |
| Two agents at once         | One agent's landing forces another to redo its checks | Separate copies; takes turns; merges clean changes; checks the combined result against a safety floor | How to resolve conflicts; whether a change needs more testing |
| Gets corrected             | A fixed mistake still gets quoted                     | Marks explanations stale when sources change; never rewrites them                                     | What the change means; the rewrite                            |
| Gets retired               | A retired note keeps showing up                       | Hides what's marked retired, where it honours the marker                                              | What to retire and why                                        |

## Where to start

A few things you can set up this week:

1. **Cut the always-loaded file to a map.** Keep the routes. Move the depth to files the agent opens when the task calls for it.
2. **Pick one rule your agent keeps missing, and find the moment it should apply.** If a command runs at that moment, make it refuse the wrong move and print the right one in the error. That can block a covered wrong move. It cannot, by itself, force a step the agent never takes.
3. **Before you build a warning, count how often it would fire.** Run it over a few weeks of your own session logs first. Mine would have gone off about thirty times a day, which is noise.
4. **Split flagging from rewriting.** Have a script mark a note stale when its source changes, and have the agent do the rewrite.
5. **When you retire something, check every reader.** Search, links, and summaries may each read a different field.

If you want the prior art, start with [Karpathy's LLM-wiki idea](https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f), Garry Tan's ["Resolvers: The Routing Table for Intelligence"](https://x.com/garrytan/status/2044479509874020852) and ["Thin Harness, Fat Skills"](https://x.com/garrytan/status/2042925773300908103), the ["memory is a hint, not truth" thread](https://x.com/himanshustwts/status/2038924027411222533), and [Horthy's objection about keeping specs in sync](https://x.com/dexhorthy/status/2090091396917248209). This piece is one account of where the bookkeeping went.
