# Agent Cache Blog

Writing on docs for agents, extraction mechanics, and building documentation that coding agents can actually read.

## Articles

- [i extracted 100 documentation sites: 12% failed. here's why.](/blog/100-docs-sites-what-broke) — i tested my docs extraction pipeline on 100 real developer tools. 88% extracted perfectly. 12% failed for 4 specific reasons. here's the full data.
- [try raw markdown before paying to scrape docs](/blog/acquisition-ladder) — most docs sites expose raw markdown. i built a 6-tier acquisition ladder that tries free methods first. here's how to extract docs without scraping.
- [should docs put agents or humans first?](/blog/agent-first-documentation-future) — ai agents are reading docs. should documentation be written for humans or agents? here's my prediction for the future of docs.
- [why your agent hallucinates apis (and how docs fix it)](/blog/agent-hallucination-docs) — your agent just used an api that doesn't exist. here's why agents hallucinate apis, why the fix is documentation, and how agent cache prevents it.
- [ai-native documentation platforms: hype or actual future?](/blog/ai-native-documentation-platforms) — documentation.ai, docsalot, hyperdocs, docsio. ai-native docs platforms are everywhere. i tested them. here's what's real and what's marketing.
- [api documentation tools: from openapi to graphql](/blog/api-documentation-tools-ranked) — scalar, redocly, swagger, and 15+ other api doc tools compared. which ones are agent-readable and which ones hide content behind javascript?
- [agent cache architecture: zero-memory, disk-backed, sse-streamed](/blog/architecture-deep-dive) — agent cache uses a zero-memory architecture where nothing stays in ram. everything streams to disk. here's how i built a deterministic, replayable docs extraction pipeline.
- [bot protection is the last mile problem for docs extraction](/blog/bot-protection-docs-extraction) — 12% of documentation sites block extraction because of cloudflare, captchas, and wafs. here's how bot protection blocks crawlers and what i do about it.
- [docsify, markdoc, and the problem with browser-rendered documentation](/blog/browser-rendered-docs-extraction-problem) — docsify renders markdown in the browser. markdoc requires custom parsing. here's why browser-rendered documentation is a trap for ai agents.
- [cloudflare workers' 50 subrequest limit killed my serverless dream](/blog/cloudflare-workers-50-subrequest-limit) — i tried to run agent cache on cloudflare workers. the 50 subrequest limit made it impossible for large docs sites. here's why i switched to a vps.
- [8-12 workers: finding the sweet spot for parallel docs extraction](/blog/concurrency-sweet-spot-extraction) — i tested 1 to 50 concurrent workers for documentation extraction. 8-12 was the sweet spot. here's why more workers actually hurts.
- [content negotiation: getting markdown from sites that don't advertise it](/blog/content-negotiation-markdown) — some docs sites serve raw markdown when you send accept: text/markdown. they don't advertise it. here's how i found them.
- [deterministic extraction beats llm-powered summarization](/blog/deterministic-vs-llm-extraction) — llms summarize. deterministic extraction preserves. for documentation, you need the whole thing, not an ai's interpretation. here's why code wins.
- [docs generators i'd actually choose in 2026](/blog/documentation-site-generators-2026) — 50+ docs tools exist. most don't matter. here's the short list of documentation generators worth your time, with honest picks for each use case.
- [keeping reference docs in .agentcache, beside your code](/blog/agentcache-dot-folder) — See how a local documentation bundle can live beside your code in a predictable .agentcache folder, ready for an AI coding agent to read.
- [building a dual-layer error system: machine errors vs human errors](/blog/dual-layer-error-system) — agent cache uses two error layers: machine errors for logs and human errors for users. here's why and how it makes debugging easier.
- [why some docs sites cost 10x more to extract](/blog/economics-docs-extraction) — extracting documentation costs anywhere from free (direct markdown) to $$$ (paid services). here's the full cost breakdown across 100 extracted sites.
- [mintlify vs docusaurus vs gitbook: which docs framework is best for agents?](/blog/docs-framework-agent-readability) — i extracted docs from 100 sites across 8 frameworks. here's my ranking of documentation framework agent-readability, from best to worst.
- [github tree api: the fastest docs extraction you've never heard of](/blog/github-tree-fast-extraction) — if a docs site is open-source on github, you can extract the entire docs tree in 30 seconds using the git tree api. no scraping needed. here's how.
- [why head requests fail on netlify edge and cloudflare workers](/blog/head-requests-fail-edge-routes) — head requests return 502 bad gateway on netlify edge and cloudflare workers. here's why dynamic edge functions don't handle head and what i use instead.
- [from html to markdown: building a documentation html cleaner](/blog/html-to-markdown-extraction) — How documentation HTML can be cleaned and converted into Markdown files, preserving useful content while removing navigation and page clutter.
- [30 lessons learned from building a documentation extraction tool](/blog/30-lessons-docs-extraction) — after extracting 100+ documentation sites and building agent cache, here are 30 lessons about docs, agents, and extraction.
- [meta.yaml: what makes documentation discoverable by agents](/blog/meta-yaml-agent-discovery) — every agent cache bundle includes a meta.yaml with structured metadata. here's what goes in it and why agents need it.
- [mintlify vs docusaurus vs fumadocs: which docs framework should you choose?](/blog/mintlify-vs-docusaurus-vs-fumadocs) — the three most popular docs frameworks compared. which one is fastest, most agent-friendly, and right for your project?
- [knowledge bases, ai docs, and platforms that do too much](/blog/knowledge-base-docs-platforms-compared) — document360, intercom, gitbook, and ai-native docs tools compared. when a docs site becomes a support portal, extraction gets complicated.
- [i requested page.md.md and wondered why it returned 404](/blog/md-md-extension-bug) — a simple bug caused me to request `page.md.md` instead of `page.md`. here's why url normalization is critical.
- [why agent cache is going open source (and what that means)](/blog/open-source-direction) — i'm open sourcing the agent cache extraction engine under mit. here's what that means and why i did it.
- [pure documentation generators ranked for agent-readability](/blog/pure-docs-generators-ranked) — not all docs generators are equal for ai agents. here's my ranking of pure documentation ssgs based on extraction experience from 100+ sites.
- [why cloudflare r2 beats s3 for my documentation bundles](/blog/r2-vs-s3-storage) — i use cloudflare r2 instead of s3 for storing documentation bundles. here's why zero egress fees matter when you serve zip downloads.
- [do i respect robots.txt? my crawling ethics policy](/blog/robots-txt-crawling-ethics) — i respect robots.txt, rate limits, and only crawl public docs. here's my full crawling ethics policy.
- [how readable are docs for ai agents in 2026?](/blog/state-of-docs-for-agents-2026) — i extracted 100+ documentation sites to understand how docs are adapting for agents. here's what i found, and where i'm headed.
- [how stripe's custom docs framework broke me (and what i learned)](/blog/stripe-docs-crawler-challenges) — stripe doesn't use mintlify, docusaurus, or gitbook. they built their own. extracting it took days. here's what i found.
- [why i chose turso (libsql) over postgresql for job metadata](/blog/why-turso-over-postgres) — agent cache uses turso (libsql) instead of postgresql for job metadata. here's why serverless sqlite was the right call.
- [vemetric broke us: 40 hours, 3 fixes, and the bug that survived all of them](/blog/vemetric-docs-case-study) — a single docs site ate 40 hours and survived 3 fixes. here's the full autopsy of why vemetric broke agent cache after 100 sites of testing.
- [how i handle versioned documentation (v1, v2, v3)](/blog/versioned-documentation) — many docs sites have version selectors. here's how agent cache extracts versioned docs and which version you get.
- [why i chose hono over express for agent cache](/blog/why-hono-over-express) — express is the default. fastify is performant. i chose hono. here's why an ultra-lightweight framework was the right call for agent cache.
- [why i chose htmx over react for agent cache](/blog/why-htmx-over-react) — everyone uses react for web apps. i chose htmx + hono for agent cache. here's why server-rendered html beats spa for this use case.
- [should your coding agent use local documentation?](/blog/why-local-docs) — Explore local and offline documentation for Cursor, Claude Code, Codex, and Windsurf, including the tradeoffs of keeping API docs on disk.
- [where LLMs fit in documentation extraction](/blog/why-no-llm-extraction) — Documentation should preserve source text accurately. Here's why I started with deterministic extraction and where LLM-assisted organization or retrieval may help.
- [if state matters, agent cache writes it to disk](/blog/zero-memory-architecture) — agent cache uses a zero-memory rule where nothing stays in ram. everything streams to disk. here's why this matters for reliability.
