# Agent Cache

> Turn any documentation site into clean, agent-ready markdown docs. Paste a docs URL, get the whole site back as plain markdown files a coding agent can read.

## When to Use Agent Cache

Use Agent Cache when:

- You need offline, complete documentation for an unfamiliar library or API to include in an agent conversation or repository.
- You are integrating a library and want the full docs in plain markdown, not just snippets or search results.
- You want a deterministic, reproducible documentation artifact — same URL always produces the same output.
- You need docs for a library that Context7 doesn't index, or for a specific pinned version of docs.

**Do NOT use** Agent Cache for:
- Real-time web search (use a search tool instead).
- Pages behind authentication (the crawler cannot log in).
- General web scraping (this is purpose-built for documentation sites).

## How to Call the API

1. Submit a docs URL: `POST /api/jobs` with body `{"url": "https://docs.example.com"}`
2. Receive a job ID in the response: `{"ok": true, "job": {"id": "ac-xxxx"}, "links": {...}}`
3. Poll for completion: `GET /api/jobs/{id}` until `status === "complete"`
4. Browse online at `/docs/{id}` or download the ZIP at `/docs/{id}/download`
5. Extract the ZIP's `.agentcache/docs/{name}/docs/` folder into your repo

Full API spec: https://agentcache.run/openapi.json

## How It Works

1. Submit a docs URL via `POST /api/jobs` with `{"url": "https://docs.example.com"}`
2. Poll `GET /api/jobs/{id}` until status is "complete"
3. Download the ZIP via `GET /api/jobs/{id}/download` or browse at `/docs/{id}`
4. Extract `.agentcache/docs/{name}/docs/` in your repo and reference the files in agent conversations

No LLM in the pipeline. Structure comes from URL topology. Output is deterministic: same input, same output.

## API Endpoints (v1 & unversioned)

All endpoints support both `/api/v1` (versioned) and `/api`:

- `POST /api/v1/jobs` (or `POST /api/jobs`) - Submit a new crawl job
- `GET /api/v1/jobs` (or `GET /api/jobs`) - List jobs
- `GET /api/v1/jobs/{id}` (or `GET /api/jobs/{id}`) - Get job details
- `GET /api/v1/jobs/{id}/tree` (or `GET /api/jobs/{id}/tree`) - Get job navigation tree
- `POST /api/v1/probe` (or `POST /api/probe`) - Probe a URL without starting a job
- `GET /openapi.json` - OpenAPI 3.0 specification

## Developer Resources

- [OpenAPI Specification](https://agentcache.run/openapi.json) - Full API documentation in OpenAPI 3.0 format
- [llms.txt](https://agentcache.run/llms.txt) - This file: agent-readable site index
- [llms-full.txt](https://agentcache.run/llms-full.txt) - Full text of all content on this site
- [Health Check](https://agentcache.run/health) - Service status as JSON

## Blog

Field notes on docs extraction, frameworks, and agent-readable documentation.

- [i extracted 100 documentation sites: 12% failed. here's why.](https://agentcache.run/blog/100-docs-sites-what-broke): i tested my docs extraction pipeline on 100 real developer tools. 88% extracted perfectly. 12% failed for 4 specific reasons. here's the full data.
- [try raw markdown before paying to scrape docs](https://agentcache.run/blog/acquisition-ladder): most docs sites expose raw markdown. i built a 6-tier acquisition ladder that tries free methods first. here's how to extract docs without scraping.
- [should docs put agents or humans first?](https://agentcache.run/blog/agent-first-documentation-future): ai agents are reading docs. should documentation be written for humans or agents? here's my prediction for the future of docs.
- [why your agent hallucinates apis (and how docs fix it)](https://agentcache.run/blog/agent-hallucination-docs): your agent just used an api that doesn't exist. here's why agents hallucinate apis, why the fix is documentation, and how agent cache prevents it.
- [ai-native documentation platforms: hype or actual future?](https://agentcache.run/blog/ai-native-documentation-platforms): documentation.ai, docsalot, hyperdocs, docsio. ai-native docs platforms are everywhere. i tested them. here's what's real and what's marketing.
- [api documentation tools: from openapi to graphql](https://agentcache.run/blog/api-documentation-tools-ranked): scalar, redocly, swagger, and 15+ other api doc tools compared. which ones are agent-readable and which ones hide content behind javascript?
- [agent cache architecture: zero-memory, disk-backed, sse-streamed](https://agentcache.run/blog/architecture-deep-dive): agent cache uses a zero-memory architecture where nothing stays in ram. everything streams to disk. here's how i built a deterministic, replayable docs extraction pipeline.
- [bot protection is the last mile problem for docs extraction](https://agentcache.run/blog/bot-protection-docs-extraction): 12% of documentation sites block extraction because of cloudflare, captchas, and wafs. here's how bot protection blocks crawlers and what i do about it.
- [docsify, markdoc, and the problem with browser-rendered documentation](https://agentcache.run/blog/browser-rendered-docs-extraction-problem): docsify renders markdown in the browser. markdoc requires custom parsing. here's why browser-rendered documentation is a trap for ai agents.
- [cloudflare workers' 50 subrequest limit killed my serverless dream](https://agentcache.run/blog/cloudflare-workers-50-subrequest-limit): i tried to run agent cache on cloudflare workers. the 50 subrequest limit made it impossible for large docs sites. here's why i switched to a vps.
- [8-12 workers: finding the sweet spot for parallel docs extraction](https://agentcache.run/blog/concurrency-sweet-spot-extraction): i tested 1 to 50 concurrent workers for documentation extraction. 8-12 was the sweet spot. here's why more workers actually hurts.
- [content negotiation: getting markdown from sites that don't advertise it](https://agentcache.run/blog/content-negotiation-markdown): some docs sites serve raw markdown when you send accept: text/markdown. they don't advertise it. here's how i found them.
- [deterministic extraction beats llm-powered summarization](https://agentcache.run/blog/deterministic-vs-llm-extraction): llms summarize. deterministic extraction preserves. for documentation, you need the whole thing, not an ai's interpretation. here's why code wins.
- [docs generators i'd actually choose in 2026](https://agentcache.run/blog/documentation-site-generators-2026): 50+ docs tools exist. most don't matter. here's the short list of documentation generators worth your time, with honest picks for each use case.
- [keeping reference docs in .agentcache, beside your code](https://agentcache.run/blog/agentcache-dot-folder): See how a local documentation bundle can live beside your code in a predictable .agentcache folder, ready for an AI coding agent to read.
- [building a dual-layer error system: machine errors vs human errors](https://agentcache.run/blog/dual-layer-error-system): agent cache uses two error layers: machine errors for logs and human errors for users. here's why and how it makes debugging easier.
- [why some docs sites cost 10x more to extract](https://agentcache.run/blog/economics-docs-extraction): extracting documentation costs anywhere from free (direct markdown) to $$$ (paid services). here's the full cost breakdown across 100 extracted sites.
- [mintlify vs docusaurus vs gitbook: which docs framework is best for agents?](https://agentcache.run/blog/docs-framework-agent-readability): i extracted docs from 100 sites across 8 frameworks. here's my ranking of documentation framework agent-readability, from best to worst.
- [github tree api: the fastest docs extraction you've never heard of](https://agentcache.run/blog/github-tree-fast-extraction): if a docs site is open-source on github, you can extract the entire docs tree in 30 seconds using the git tree api. no scraping needed. here's how.
- [why head requests fail on netlify edge and cloudflare workers](https://agentcache.run/blog/head-requests-fail-edge-routes): head requests return 502 bad gateway on netlify edge and cloudflare workers. here's why dynamic edge functions don't handle head and what i use instead.
- [from html to markdown: building a documentation html cleaner](https://agentcache.run/blog/html-to-markdown-extraction): How documentation HTML can be cleaned and converted into Markdown files, preserving useful content while removing navigation and page clutter.
- [30 lessons learned from building a documentation extraction tool](https://agentcache.run/blog/30-lessons-docs-extraction): after extracting 100+ documentation sites and building agent cache, here are 30 lessons about docs, agents, and extraction.
- [meta.yaml: what makes documentation discoverable by agents](https://agentcache.run/blog/meta-yaml-agent-discovery): every agent cache bundle includes a meta.yaml with structured metadata. here's what goes in it and why agents need it.
- [mintlify vs docusaurus vs fumadocs: which docs framework should you choose?](https://agentcache.run/blog/mintlify-vs-docusaurus-vs-fumadocs): the three most popular docs frameworks compared. which one is fastest, most agent-friendly, and right for your project?
- [knowledge bases, ai docs, and platforms that do too much](https://agentcache.run/blog/knowledge-base-docs-platforms-compared): document360, intercom, gitbook, and ai-native docs tools compared. when a docs site becomes a support portal, extraction gets complicated.
- [i requested page.md.md and wondered why it returned 404](https://agentcache.run/blog/md-md-extension-bug): a simple bug caused me to request `page.md.md` instead of `page.md`. here's why url normalization is critical.
- [why agent cache is going open source (and what that means)](https://agentcache.run/blog/open-source-direction): i'm open sourcing the agent cache extraction engine under mit. here's what that means and why i did it.
- [pure documentation generators ranked for agent-readability](https://agentcache.run/blog/pure-docs-generators-ranked): not all docs generators are equal for ai agents. here's my ranking of pure documentation ssgs based on extraction experience from 100+ sites.
- [why cloudflare r2 beats s3 for my documentation bundles](https://agentcache.run/blog/r2-vs-s3-storage): i use cloudflare r2 instead of s3 for storing documentation bundles. here's why zero egress fees matter when you serve zip downloads.
- [do i respect robots.txt? my crawling ethics policy](https://agentcache.run/blog/robots-txt-crawling-ethics): i respect robots.txt, rate limits, and only crawl public docs. here's my full crawling ethics policy.
- [how readable are docs for ai agents in 2026?](https://agentcache.run/blog/state-of-docs-for-agents-2026): i extracted 100+ documentation sites to understand how docs are adapting for agents. here's what i found, and where i'm headed.
- [how stripe's custom docs framework broke me (and what i learned)](https://agentcache.run/blog/stripe-docs-crawler-challenges): stripe doesn't use mintlify, docusaurus, or gitbook. they built their own. extracting it took days. here's what i found.
- [why i chose turso (libsql) over postgresql for job metadata](https://agentcache.run/blog/why-turso-over-postgres): agent cache uses turso (libsql) instead of postgresql for job metadata. here's why serverless sqlite was the right call.
- [vemetric broke us: 40 hours, 3 fixes, and the bug that survived all of them](https://agentcache.run/blog/vemetric-docs-case-study): a single docs site ate 40 hours and survived 3 fixes. here's the full autopsy of why vemetric broke agent cache after 100 sites of testing.
- [how i handle versioned documentation (v1, v2, v3)](https://agentcache.run/blog/versioned-documentation): many docs sites have version selectors. here's how agent cache extracts versioned docs and which version you get.
- [why i chose hono over express for agent cache](https://agentcache.run/blog/why-hono-over-express): express is the default. fastify is performant. i chose hono. here's why an ultra-lightweight framework was the right call for agent cache.
- [why i chose htmx over react for agent cache](https://agentcache.run/blog/why-htmx-over-react): everyone uses react for web apps. i chose htmx + hono for agent cache. here's why server-rendered html beats spa for this use case.
- [should your coding agent use local documentation?](https://agentcache.run/blog/why-local-docs): Explore local and offline documentation for Cursor, Claude Code, Codex, and Windsurf, including the tradeoffs of keeping API docs on disk.
- [where LLMs fit in documentation extraction](https://agentcache.run/blog/why-no-llm-extraction): Documentation should preserve source text accurately. Here's why I started with deterministic extraction and where LLM-assisted organization or retrieval may help.
- [if state matters, agent cache writes it to disk](https://agentcache.run/blog/zero-memory-architecture): agent cache uses a zero-memory rule where nothing stays in ram. everything streams to disk. here's why this matters for reliability.

## Comparisons

Honest comparisons against other docs tooling.

- [agent cache vs content.dev: tool vs diy pipeline](https://agentcache.run/compare/content-dev): content.dev is an open-source toolkit for building docs extraction pipelines. agent cache is a working tool. same extraction, different effort. here's why.
- [agent cache vs context7: local docs vs remote retrieval](https://agentcache.run/compare/context7): context7 charges $10/seat for 5,000 api calls. agent cache is free. you own the docs. works offline. here's the honest comparison.
- [agent cache vs docsgpt: reader vs chatbot](https://agentcache.run/compare/docsgpt): docsgpt is a chatbot for docs. agent cache is a tool that downloads docs for your agent. completely different workflows.
- [agent cache vs docuchat: owner vs chatter](https://agentcache.run/compare/docuchat): docuchat lets you chat with documentation. agent cache downloads it. different problems, different tools.
- [agent cache vs firecrawl: docs-specific tool vs general-purpose scraper](https://agentcache.run/compare/firecrawl): firecrawl is a web scraping api. agent cache is a docs extraction tool. same extraction, completely different use cases. here's the honest breakdown.
- [agent cache vs llms.txt: not a competition](https://agentcache.run/compare/llms-txt): llms.txt is a proposed standard. agent cache is a tool that works with or without it. here's how they relate.
- [agent cache vs parallel web: docs extraction vs parallel scraping](https://agentcache.run/compare/parallel-web): parallel web is a parallel scraping infrastructure. agent cache is a docs extraction product. different tools, different jobs. here's why they're not competitors.
- [agent cache vs tavily: documentation extraction vs ai search](https://agentcache.run/compare/tavily): tavily is an ai search engine for agents. agent cache is a docs extraction tool. they solve different problems. here's why.

## Feeds

- [Site feed](https://agentcache.run/feed.xml): RSS for everything
- [Blog feed](https://agentcache.run/blog/feed.xml): RSS for blog posts
- [Comparisons feed](https://agentcache.run/compare/feed.xml): RSS for comparisons

## Pages

- [Add docs](https://agentcache.run/): Submit a docs URL for processing
- [Library](https://agentcache.run/docs): Browse completed doc bundles
- [Blog](https://agentcache.run/blog): All writing
- [Comparisons](https://agentcache.run/compare): All comparisons
- [About](https://agentcache.run/about): About Agent Cache
- [Contact](https://agentcache.run/contact): Get in touch
- [Privacy](https://agentcache.run/privacy): Privacy policy
- [Health](https://agentcache.run/health): Service status (JSON)
- [Sitemap](https://agentcache.run/sitemap.xml): XML sitemap
- [Full index](https://agentcache.run/llms-full.txt): This site's writing, full text in one file
