A stale document costs me twice with an agent, once in the work it does wrong and again in the context that work is now sitting in. A document it cannot find costs me the same way, in everything it opens on the way to the one it needed.
My past two articles, on smoothing the seams between sessions and seeing them all at once, were inspired by my continual effort to get a grasp on the multitude of agents I have running across multiple projects and ultimately, grappling with my own limited context window. Part of that has been encouraging the agents to document and with the pace of change the agents afford, that's a lot of documentation. That can quickly get out of hand, disorganized, and stale.
There are more than a handful of solutions out there already that attempt to address this. So many, in fact, that rather than picking one that didn't work out, I thought I'd engage Claude Opus to help me out. I'm not going to claim it beats them, because I haven't run them side by side. What I can show you is what this one measures, on a real 136-document tree, and you can decide from there whether it's the approach you want.
What came out of that is doc-librarian, a Claude Code skill that treats docs/ as what it literally is, a small library, and applies the discipline libraries have used for well over a century: a taxonomy, a controlled vocabulary, a catalog, hub-and-spoke indexes, and a records lifecycle that ends in an archive rather than a delete.
The agent-builds-on-a-stale-document failure is the one I felt first. A person skimming a docs folder can usually smell that something has gone off: the screenshot is from an old UI, the command has a flag that no longer exists. An agent reads it at face value and starts building. So "which of these documents are still true" stops being a tidiness question and becomes a correctness one, and the answer that worked for me was structural: a document is either current, or it is visibly archived, and there is no third state where it sits in a live folder looking authoritative. That rule turned out to answer the other half as well, since a folder with the dead documents visibly out of it is a smaller thing to search.
Then I pointed it at my own worst case: a 9-package TypeScript monorepo, 136 markdown documents, frontmatter on none of them, and nobody who could say which of them were still true.
This is for people whose docs/ has outgrown their memory of it, enough documents that "I'll just read them" has stopped being a plan. If yours is a dozen files you wrote this quarter, you don't have this problem yet.
TL;DR
doc-librarianis a Claude Code skill that audits, triages, reorganizes, indexes and weeds adocs/tree, and installs the enforcement that keeps it organized afterwards.Every verdict is measured, not guessed: staleness from
git log, orphanhood from a repo-wide referrer count that reads source and config as well as other docs, near-duplicates from TF-IDF cosine.Cheap signal runs first and scores the whole tree in 11 seconds. Agents are spent reading only the pile it couldn't decide.
It never deletes a referenced document. The lifecycle is deprecate, archive, and much later dispose, and the step that moves and rewrites does nothing until you pass
--apply.The second half is a governance kit: a linter, an authoring-time guard that nudges any agent writing a non-conforming doc, and a commit-time gate.
Install and driving instructions: START-HERE.md.
How a docs directory decays
Sprawl, duplication, orphans, stale shelves, and no catalog: nothing that says what each document is for, so classification lives only in whoever wrote it. You have all five or you wouldn't be reading this.
The cost is in the deciding rather than the reshuffling. To know whether a document is still true, someone has to read it and then read the code it describes, and at 136 documents I'd put that at a week I didn't have. So the tree stays as it is, and each new document is filed by vibes into whichever folder looks least wrong.
There's a second cost that only shows up once you start moving things. Documentation gets cited from live source, tests, CI config, editor and agent config, and root-level README and CLAUDE.md files. On my corpus, 56 of the 62 documents that moved had referrers outside docs/, and 30 of those were hard. An ESLint config naming a standards doc, a JSDoc comment citing a design doc, a @see on an exported function, a test asserting a filename. A docs reorganization is not a docs-only change, and finding that out halfway through a migration is a bad time to learn it.
The methodology
Library science already has answers for all five, and the skill is organized around them rather than around file operations. Documents are classified by Diátaxis, which sorts by what the reader is doing rather than by subject and is mutually exclusive and collectively exhaustive, which is what stops folders multiplying. Tags come from a curated TAGS.md and nowhere else, because inventing one at write time gives you auth, authentication and authn as three separate facets. The rest is bookkeeping that pays for itself: frontmatter on every document so the corpus is queryable, a README in every folder so nothing is reachable only by grep, and a lifecycle that ends in an archive rather than a delete, because the archive is what makes weeding safe enough to actually do.
Underneath all five is the discipline that makes the rest trustworthy: every verdict has to cite something. A document is stale because git says when it last changed. It's an orphan because a repo-wide grep found nothing pointing at it, checked against source and config rather than just other documents. It's a near-duplicate because the TF-IDF cosine cleared a threshold you set, and without the optional Python pass installed that one degrades to matching normalized titles, which the report tells you.
The measurements are not automatically right, and mine were wrong. An early version of the referrer counter only looked at links between documents, so anything cited from source code counted as an orphan, and orphanhood feeds the WEED bucket. It was routing live documents toward deletion.
What saved it was that every row prints its evidence. The agents reading the REVIEW pile kept returning verdicts that contradicted the score. This says orphaned, but it's cited from AgentSpecGenerator.ts:456 and four other places. They did that on five different documents, independently, without being told the bug existed.
The safety guarantees lean on that same counter. The rule that a referenced document is never deleted is enforced by the count that was wrong, so a miscounted document reads as unreferenced and the guard waves it through, which is why the WEED pile still goes to the reading agents and why I'd treat that guard as the second line rather than the first.
Fixing that counter and two related path-resolution bugs reclassified 42 of the 136 documents as KEEP, cutting the pile that needed an agent to read it from 96 to 56.
What a run looks like
Triage is the cheap half, and triage-docs.ts scores every document on staleness, inbound references, dead code anchors, duplicate similarity and location, and drops each into one of three buckets: KEEP, WEED, or REVIEW, where REVIEW explicitly means "the signal could not decide this one." On my corpus that split the 136 into 76 KEEP, 56 REVIEW and 4 WEED, and the 76 never get read by anything. It prints its reasoning per row, in the tool's own words: very stale (14mo), orphaned (no inbound links anywhere), dead code anchors (0% resolve), in "migration/" (iterative/likely-completed).
The deep dive is the expensive half, and it runs only on what the score didn't settle: REVIEW, plus the small WEED pile, because nothing gets discarded on a score alone. plan-review-batches.ts packs that bucket into batches of related documents, so one agent reading a folder sees the whole folder, and writes a brief per batch. Then one agent per batch reads the documents and the code around them and returns a per-document verdict with evidence. On my corpus that was 15 batches for 102 documents, and that is the number to price it by rather than the 11 seconds, which only ever covered the cheap half. That 102 is what I actually paid, because I ran the appraisal before I had fixed the referrer counter, when triage was still handing it 96 REVIEW and 6 WEED. With the counter fixed the same corpus routes 60 documents to the agents instead, which is the figure to plan against.
Here is a real one, trimmed. This is the actual output row for a document in my repo:
doc_path: migration/branded-ids.md
verdict: ARCHIVE confidence: high inbound_links: 0
One-shot generated snapshot from 2026-05-13, never regenerated since. All 339
checkboxes are still unchecked (`grep -c '^- \[x\]'` → 0) despite active,
ongoing ID-branding work in the same area since. Spot-checked line-number
anchors have drifted: `orchestration-engine.ts:178` no longer contains the
cited `runId?: string;` snippet.
Flag for the librarian: that generator script's OUTPUT_PATH still points at
`docs/migration/branded-ids.md` — if anyone re-runs it, the file reappears at
this path, undoing an archive-move unless the script is also updated.That last paragraph is the reason the deep dive exists. No amount of scoring finds a script that will silently recreate a document you just archived. Somebody reading the code around it does, and that is the only thing that found it.
The verdicts roll up into a migration map, and the map is what you sign off on, rather than the migration itself. Here is how mine opens:
Proposed by `doc-librarian` Procedure B. **Nothing has been moved.** This is the
sign-off artifact: approve it, or amend it, before the first `git mv`.
## Shape
| | count |
|---|---:|
| documents today | 136 |
| stays put — `docs/TAGS.md`, the vocabulary | 1 |
| → `archive/` | 73 |
| → `tutorials/` | 1 |
| → `how-to/` | 13 |
| → `reference/` | 16 |
| → `explanation/` | 26 |
| → `decisions/` | 1 |
| → `project/` | 5 |
| frontmatter to author (every surviving doc has none) | 62 |
| moving docs with live non-docs referrers to rewrite | 22 |
## Needs your call — 5 itemsThose five are the things it would not decide for me, among them two duplicate pairs where one copy of each is regenerated by a skill, and a document whose generator hardcodes the old path in four places. The per-document tables come after all of that, current path to target path with the live referrers each move would break. The skill is explicit that you don't map and move at the same time. Read the map next to the verdicts rather than on its own, though: the map is destinations and counts, and the reasoning that caught my orphan bug lives in the verdict report beside it.
What it writes, and how to undo it
Worth having as a closed list before you point it at anything, because it is otherwise spread across three phases.
Installing the skill touches ~/.claude/skills/ and nothing else, so no repository and no settings.json.
/doc-librarian setup writes six things into the repo you run it in, all tracked: docs/scripts/lint-docs.ts, .claude/hooks/docs_structure_guard.py, a starter docs/TAGS.md if you don't already have one, a hook entry in .claude/settings.json, npm scripts in package.json, and a .doc-librarian/ line in .gitignore. The settings.json entry is the one your teammates inherit.
The migration moves documents with git mv, so every move lands as a rename with its history intact, and rewrites links inside docs/. It does nothing at all until you pass --apply.
Then run it on a branch and git diff each batch before you commit it. That is not me being careful for its own sake: it's how I caught the frozen-capture corruption in my own migration, and the tool never would have.
Adopting it: the initial cleanup
Install is a clone, a dry run, and an installer:
git clone https://github.com/agentience/agentience-skills.git
cd agentience-skills/skills/doc-librarian
./install.sh --dry-run # prints every path it would touch
./install.shIt symlinks into ~/.claude/skills/, writes nothing to settings.json, and changes no repository. Then start a new session — a running one keeps the skill text it loaded at startup.
From inside the repo you want organized, /doc-librarian setup installs the governance kit. The recommended order for a neglected tree is then: bootstrap a tag vocabulary from the corpus's own language, weed, audit what survived, finalize governance, reorganize, index. That is G → E → A → F → B → D in the skill's own lettering. It reads one docs/ tree, so on a monorepo you point it at the one you mean rather than at nine.
The order is not arbitrary, and weeding before reorganizing is the part I'd insist on: reshuffling documents you're about to discard is wasted work, and a clean new structure lends false authority to dead content. Audit after weeding too, so you're measuring the collection you're actually keeping.
Two things to do before the first git mv, both of which I'd have skipped if the skill hadn't insisted:
The first is the broken-link baseline. Run the linter before you touch anything and write the number down. Afterwards, an absolute count of broken links tells you nothing — a corpus that was already broken will still be broken, and without the baseline you cannot separate breakage you caused from breakage you inherited. Mine was 25 before and 13 after, every remaining one a strict subset of the original 25. The honest claim is "no new broken links," and you can only make it if you measured first. As a side effect the migration repaired 12 links that were already broken.
The second is the blast radius. For every document the map moves, git grep -l "old/path.md" -- ':!docs', and classify each referrer as soft (only other docs) or hard (source, tests, config, root files), and report the split before you start. That split, 56 of 62 with 30 of them hard, is what told me the migration had to pass the test suite and not just the link checker.
Then migrate in batches, verify each, commit each, and never start a batch while the previous one has left new breakage. apply-actions.ts does the mechanical work and is dry-run by default; it refuses to delete a document with inbound links even when you ask it to.
One rule I'd put on a sticker: frozen artifacts are not documents. Test fixtures, captured model outputs, recorded HTTP sessions. These live under docs/ in a lot of repos, and a link rewriter will happily rewrite them. Rewriting a captured output changes the thing it recorded. In my migration an early pass silently edited 22 frozen experiment captures, one set of which existed precisely to measure whether a model cites file paths that exist. It was caught by a git diff against those directories, not by anything in the tool, which is why it's now rule seven of the skill's safety rules. It stayed a rule rather than a guard: nothing in the tool excludes them for you.
Adopting it: keeping it clean

A one-shot cleanup has nothing holding it in place. The repo that produced those 136 documents is still producing them at the same rate, so the half I'd expect to matter more over time is the part that runs afterwards, and it's what /doc-librarian setup installs:
lint-docs.ts, copied into your repo as the single source of truth for every check: links, tags against the controlled vocabulary, frontmatter schema, bucket/type agreement.A
PostToolUseguard hook, registered in that repo's.claude/settings.json, so it fires for everyone who works in the repo, not just you. When any agent writes a document that doesn't conform, the hook exits non-zero and hands the agent back the fix list: classify the Diátaxis type, pick the one correct folder, add frontmatter, name it kebab-case, wire it into the index. The write has already happened at that point; what you get is the agent correcting it in the same turn rather than someone auditing it months later. It fails open on any error of its own, so a broken hook never blocks authoring.A commit-time gate:
lint-docs.ts --stagedin your pre-commit hook, blocking, plus--stalenessas a non-blocking review report..doc-librarian/gitignored, with the reports written there rather than intodocs/. A report written into the tree becomes a document, and the next scan appraises the skill's own output alongside your corpus. That one isn't finished —plan-review-batches.tsstill defaults its briefs into the tree it just read, so pass it an out-dir.
The guard is Claude Code-specific; the linter and the pre-commit gate are not, so on a client without hooks you get the same rules one step later.
What it won't do
Named up front, because each of these bit mine:
Reference counting is liveness-blind. A citation from
specs/completed/weighs exactly as much as one from live source, so a genuinely dead document can hold itself out of the WEED bucket on dead referrers. It errs toward human review rather than deletion, which is the safe direction — but a high inbound count is not proof of relevance.Nothing checks anchors.
guide.md#configurationsurvives a move as a link to the right file and the wrong place in it. The linter sees paths, not fragments; the skill tells you to grep them by hand.It rewrites links inside
docs/only. Referrers in source, tests and config are counted but not rewritten, so you get the list and the edits are yours to make.It can write to files under
docs/that aren't documents. Anything in the docs tree that is a captured artifact rather than prose is fair game for the link rewriter, and excluding it is a rule the skill states rather than a check the code runs, which is how 22 of mine got rewritten.It cannot tell you what is true. It measures stale, orphaned and duplicated. Whether the content is still correct is a reading task — that's what the deep dive is for, and even that returns verdicts for a human to accept.
The first three are open on purpose: design questions I'd have had to guess at, and a skill that quietly did the wrong thing in those places would be worse than one that names them. The fourth is not on purpose; it is unfinished, and it is the one I'd check by hand.
Try it
The fastest path is to hand an agent the bootstrap document and let it install and drive the skill itself:
Read https://raw.githubusercontent.com/agentience/agentience-skills/main/skills/doc-librarian/START-HERE.md and follow it.
It's written for an agent rather than a person, it's self-contained, and it ends with the rules that keep an enthusiastic agent from destroying a corpus quickly and plausibly. Everything else, the eight procedures and the scripts and the templates, is in the repo.
Start with audit. It only counts things, so it's the one step I'd run without thinking twice about it.

