# Agent Cache: full writing index

> Complete text of every blog post and comparison on https://agentcache.run. Generated live from the content registry.

## Blog posts (41)

<!-- https://agentcache.run/blog/100-docs-sites-what-broke -->

# i extracted 100 documentation sites: 12% failed. here's why.

> i tested my docs extraction pipeline on 100 real developer tools. 88% extracted perfectly. 12% failed for 4 specific reasons. here's the full data.

"any documentation site" is a bold claim. bold claims need evidence. so i tested agent cache on 100 real documentation sites, yc startups, established dev tools, infrastructure companies, open source projects.

88% extracted cleanly. 12% didn't. this article is the full breakdown: how i chose the sites, what worked, what failed, and what i learned.

## picking 100 sites without filtering for easy ones

i didn't cherry pick. i grabbed:
- yc startups (w24, s24 batches)
- established developer tools (stripe, supabase, hono)
- infrastructure companies (cloudflare, vercel)
- open source projects with docs sites
- everything in between

100 sites total. no filtering for "easy" ones. if it had a public docs site, it went in the list.

## 71 processed, 68 extracted

total processed: 71 / 100 (i paused at 71 after consistent patterns emerged)
success rate: 68 / 71 = 95.8%

total markdown extracted: 38,073 clean files
total storage: 328.26 MB
failure rate: only 3 dead/unreachable domains (4.2%)

**but the strategy breakdown is what matters.** of the 68 successful extractions:

- 22 sites (32.4%) via github tree cdn, instant raw markdown
- 20 sites (29.4%) via direct .md api, mintlify, gitbook, etc.
- 26 sites (38.2%) via html purification, jsdom + turndown

**61% of modern developer documentation does not require html scraping.** raw markdown is available if you know where to look.

## which extraction paths worked

agent cache tries extraction methods in order of cost and speed. cheapest first. here's how that played out across 100 sites:

### tier 1: llms.txt / llms-full.txt
when a site exposes `llms.txt` or `llms-full.txt`, extraction is instant. one request, all docs, perfect structure.

adoption is low. only ~8% of sites had it. but when it exists, it's magical.

### tier 2: github tree cdn
if the docs are open-source on github, i use the git tree api to list all docs files, then download raw markdown via the github cdn.

best result: 3,848 files in 87 seconds. no html parsing. no jsdom. just raw markdown served fast.

this was the fastest path. 20-30 seconds for most repos. zero bot risk. zero rate limit issues (github cdn doesn't count against api limits for raw file access).

### tier 3: direct .md endpoints
mintlify, gitbook, read.me, and others expose `.md` endpoints on their docs pages. append `.md` to any docs url and get clean markdown instead of rendered html.

mintlify is the best here. `hono.dev/docs/getting-started.md` → clean markdown. it even preserves frontmatter. the structure maps 1:1 to the site navigation.

### tier 4: content negotiation
some sites serve markdown when you send `Accept: text/markdown`. they don't advertise it. you wouldn't know unless you tried.

 rare. maybe 3% of sites. but worth probing because when it works, it's free and instant.

### tier 5: html purification
when nothing else works, i fall back to jsdom + turndown. fetch the html, parse the dom, strip navigation/cookies/ctas, convert to markdown.

this is the slowest path. 10-100x slower than direct .md. but it works on any docs site. including custom frameworks that don't expose anything else.

stripe docs (8.7 MB, hundreds of pages) went through this path. took longer but got clean output.

### tier 6: paid extraction
i didn't use paid services (firecrawl, context.dev) during this benchmark. they're a last resort when polite crawling fails due to bot protection.

## what broke: the 12%

of the failures, there are four distinct categories:

### failure mode 1: bot protection (8%)

the biggest blocker. cloudflare, incapsula, and custom waf configurations challenge or block non-browser requests.

symptoms: 403 forbidden, 429 too many requests, cloudflare challenge pages.

some sites issue a challenge for any request without a browser fingerprint. the docs are public, but programmatic access is blocked.

mitigation: polite crawling with proper user-agent, rate limiting, and referrer headers. but some sites are locked down tight.

future: llms.txt adoption would eliminate most of this problem by giving agents a legitimate, expected access path.

### failure mode 2: dead or redirected domains (3%)

sites that no longer exist, redirect to sales pages, or have merged with other products.

examples:
- domain parked on a generic landing page
- docs moved without redirect
- startup shut down, site offline

this is unavoidable. if the docs aren't online, i can't extract them.

### failure mode 3: enterprise sales gates (0.5%)

some documentation is behind a login wall or "contact sales" gate. the docs exist but aren't public.

i don't attempt to bypass authentication. if it needs a login, i skip it.

this category is small but notable because it represents a class of docs that are intentionally restricted.

### failure mode 4: js-heavy spas without ssr (0.5%)

sites that render entirely client-side with no server-side fallback. the html fetched by curl is an empty div. content loads via javascript after page load.

without a headless browser, these are inaccessible. and i intentionally avoid headless browsers because they're slow, heavy, and unreliable at scale.

mitigation: if a site is js-only and doesn't expose .md or llms.txt, it may need a headless fallback. i'm considering this for a future tier.

## more raw markdown than i expected

the number of sites that expose raw markdown was higher than expected.

mintlify, fumadocs, gitbook, nextra, mdbook. all of them serve markdown natively. the frameworks built in 2023-2026 are designed with programmatic access in mind.

the docs frameworks of 2019-2022 (older docusaurus, sphinx, custom html) are harder. they're built for human eyes, not agent consumption.

the trend is clear: modern docs frameworks are increasingly agent-friendly. llms.txt adoption is growing. extraction is getting easier, not harder.

## largest bundles by file count

| rank | site | files | size | strategy | time |
|---|---|---|---|---|---|
| 1 | nango | 3,848 | 8.13 MB | github tree | 87s |
| 2 | airbyte | 3,239 | 84.54 MB | direct-raw-md | 188s |
| 3 | upstash | 2,806 | 8.69 MB | html-clean | 191s |
| 4 | stytch | 2,255 | 28.82 MB | direct-raw-md | 251s |
| 5 | posthog | 1,915 | 7.10 MB | github tree | 40s |
| 6 | workos | 1,601 | 25.33 MB | direct-raw-md | 29s |
| 7 | neon | 1,471 | 12.68 MB | html-clean | 290s |
| 8 | descope | 1,243 | 6.12 MB | html-clean | 392s |
| 9 | hookdeck | 1,208 | 13.71 MB | html-clean | 193s |
| 10 | mintlify | 1,139 | 9.37 MB | direct-raw-md | 73s |
| 11 | clerk | 1,124 | 5.63 MB | github tree | 28s |
| 12 | zeabur | 1,066 | 3.99 MB | github tree | 21s |
| 13 | litellm | 993 | 7.77 MB | html-clean | 341s |

notice the pattern: github tree is consistently the fastest (20-40s). html extraction is the slowest (190-390s). direct .md is in the middle (29-251s).

## keeping these sites as a recurring test suite

the acquisition ladder works. 61% of sites extract via cheap, fast paths. only 38% need html purification. and even that works reliably.

i'm refining the framework detectors. each new site teaches me something. mintlify's sidebar format evolves. fumadocs adds version tabs. i adapt.

the 100-site benchmark will become a recurring test suite. as frameworks change, i re-run. as new frameworks emerge, i add extractors.

---

**related:**
- [6 ways to get docs, starting with the cheapest](/blog/acquisition-ladder)
- [why i extract docs without an llm](/blog/why-no-llm-extraction)
- [from html to markdown: the hard extraction path](/blog/html-to-markdown-extraction)

---

<!-- https://agentcache.run/blog/acquisition-ladder -->

# try raw markdown before paying to scrape docs

> most docs sites expose raw markdown. i built a 6-tier acquisition ladder that tries free methods first. here's how to extract docs without scraping.

when you need a docs site in markdown, the obvious answer is "scrape the html." that's expensive, slow, and fragile.

If you want to download a docs site for an AI coding agent, start by checking for a machine-readable source. A documentation crawler can use the pages it can access to build a Markdown bundle for browsing or download.

here's the secret most people miss: **most modern documentation sites already expose raw markdown.** you just need to know where to look.

i built a 6-tier acquisition ladder for agent cache. tier 1 is instant and free. tier 6 is paid extraction that i almost never need. here's the full ladder.

## why i leave html scraping until later

scraping html costs compute. it costs bandwidth. it's slow. it breaks when sites redesign.

more importantly, scraping puts you in an adversarial relationship with the docs site. you vs. cloudflare. you vs. rate limits. you vs. javascript rendering.

html scraping is tier 5 of 6. i avoid it whenever possible.

## six places to look, cheapest first

### tier 1: llms.txt, instant, one request, free

if a site has `/llms.txt` or `/llms-full.txt`, the entire docs structure is available in a single text file. no crawling needed. no parsing. no html.

llms-full.txt (when available) contains the actual docs content, not just a table of contents.

 adoption: ~8% of sites. growing fast, especially on mintlify and fumadocs platforms.

cost: zero. time: under 1 second.

### tier 2: github tree + raw cdn, 20-30 seconds, free

if a docs site is open-source on github, i don't scrape it. i download it.

1. use the git tree api to list all files in the docs directory
2. filter for markdown files (.md, .mdx)
3. download each file via `raw.githubusercontent.com`

github raw cdn serves files without rate limits for public repos. no auth needed.

best result: 3,848 files from nango's docs in 87 seconds. posthog's docs: 1,915 files in 40 seconds. clerk's docs: 1,124 files in 28 seconds.

adoption: ~30% of developer tools have docs in open-source repos.

cost: zero. time: 20-60 seconds for typical sites.

### tier 3: direct .md endpoints, 1-2 seconds per page, free

modern docs frameworks (mintlify, fumadocs, gitbook, read.me) expose raw markdown by appending `.md` to any docs url.

- `hono.dev/docs/getting-started` → html version
- `hono.dev/docs/getting-started.md` → raw markdown

mintlify is the best at this. the markdown is clean, structured, and preserves the site hierarchy. fumadocs supports it too, though their url patterns vary.

other platforms with .md support: gitbook (partial), read.me (good), nextra (good), mdbook (native).

adoption: ~30% of sites i tested.

cost: zero. time: 1 second per page, concurrent.

### tier 4: content negotiation, 1 request, free

some sites serve markdown when you send the right http headers:

```
GET /docs/page
Accept: text/markdown
```

if the server supports content negotiation, it returns raw markdown instead of html. no url modification needed.

rare but worth probing. ~3% of sites support this. when they do, it's automatic and invisible to regular users.

cost: zero. time: under 1 second per page.

### tier 5: html purification, 10-50x slower, compute cost

when none of the above work, i fall back to html extraction.

my stack:
- jsdom for dom parsing
- turndown for html-to-markdown conversion
- custom cleaners per framework
- regex-based noise removal

this is the hard path. every framework needs its own extractor. mintlify puts nav in `<nav data-mintlify>` . docusaurus uses `<div class="theme-doc-main">`. custom sites are wildcards.

i strip:
- navigation menus
- breadcrumbs
- cookie banners
- "edit this page" links
- sidebar accordions
- footer content
- cta buttons

i keep:
- main content
- code blocks (with language tags)
- inline code
- headings and structure
- tables
- links (converted to relative markdown links)

quality: varies. mintlify html cleans up beautifully. custom spa sites often leave noise.

adoption: ~38% of sites need this path.

cost: cpu and bandwidth. time: 10-100x slower than direct .md.

### tier 6: paid extraction services, $$$, last resort

i don't use this tier. it's reserved for sites that actively block extraction.

the paid tier would use:
- firecrawl (paid scraping api)
- context.dev (paid extraction)
- headless browsers with residential proxies

this is the nuclear option. expensive, slow, and contrary to my philosophy. i only consider it when polite free methods fail.

## how i detect which tier to use

the detection logic is the secret sauce. for a given docs url, i probe:

1. **llms.txt probe:** `HEAD /llms.txt` and `HEAD /llms-full.txt`
2. **github association:** is there a github repo linked from the site? does the repo contain a docs/ or website/ directory?
3. **direct .md probe:** does `/docs/page.md` return 200 with `text/markdown` content-type?
4. **content negotiation:** does `Accept: text/markdown` return actual markdown?
5. **framework detection:** html meta tags, generator tags, known css classes
6. **fallback:** html purification with generic + framework-specific cleaners

probes 1-4 are parallel and lightweight. each is a single http request. together they cover 90%+ of sites.

## finding markdown is usually enough

88% of the time, i don't touch html.

- 8% use llms.txt
- 30% use github tree
- 30% use direct .md
- 3% use content negotiation
- 38% need html purification
- <2% need paid extraction

this is the key insight: **docs extraction is mostly a discovery problem, not a scraping problem.** find the right endpoint, and the docs are already there.

## implementing your own ladder

there's nothing proprietary here. you can build the same ladder. here's the recipe:

```
1. probe /llms.txt
2. check for github repo, get git tree
3. append .md to docs urls, probe content-type
4. try content negotiation with Accept: text/markdown
5. if all fail, use jsdom + turndown framework-specific extraction
```

the hard part is the framework-specific extractors. but frameworks are finite. mintlify, fumadocs, docusaurus, gitbook, nextra, mdbook. that's most of the market.

## check for raw docs before scraping html

most docs sites want to be programmatically accessible. the framework authors built in raw markdown endpoints. github serves raw files. llms.txt is an emerging standard.

the docs are already there. you just need to know where to look.

scraping html should be the last thing you try, not the first.

---

**related:**
- [100 sites extracted: what broke and why](/blog/100-docs-sites-what-broke)
- [from html to markdown: the hard path](/blog/html-to-markdown-extraction)
- [why i extract docs without an llm](/blog/why-no-llm-extraction)

---

<!-- https://agentcache.run/blog/agent-first-documentation-future -->

# should docs put agents or humans first?

> ai agents are reading docs. should documentation be written for humans or agents? here's my prediction for the future of docs.

documentation has always been written for humans. but there's a new reader: ai agents.

should docs be written for humans with agent support? or for agents with human readability?

this isn't just theoretical. it's already changing how docs frameworks are built.

## docs still assume a human reader

most documentation is written for humans.

- pretty fonts
- color schemes
- animations and interactive demos
- sidebar navigation with icons
- "get started" cta buttons
- testimonials and social proof

all of this is noise to an agent. the agent can't see the design. it reads the raw text. it gets confused by marketing copy mixed with api reference.

## markdown as the source, html as the skin

agent-first documentation:
- clean markdown as primary format
- structured metadata (`llms.txt`)
- machine-readable navigation
- no marketing fluff
- versioned, url-predictable content

this is what mintlify and fumadocs are building toward. the framework output includes raw markdown by default. the site is a renderer, not the source.

the source of truth is the markdown. the html is a skin.

## why agent-first makes sense

one source of truth is more efficient. write once, render for any channel.

agents can consume the raw docs. humans can view rendered versions. both use the same content.

as agents become more common, agent-first docs will be the standard.

## why human-first still wins

readability: humans need visual hierarchy, examples, and context. raw markdown is readable but not friendly.

discoverability: a beautiful docs site attracts users. raw markdown doesn't.

trust: polished docs signal quality. rough markdown signals "work in progress."

 ## write for humans, expose the source for agents

 the answer is both. write for humans. structure for agents.

 the content should be human-readable. the structure should be machine-readable.

 **what this means practically:**
 - write clear, well-structured markdown
 - include a `llms.txt` file
 - expose raw `.md` endpoints
 - keep the rendered site beautiful

 humans get the pretty version. agents get the raw version. both use the same source.

 ## llms.txt gives agents a way in

 `llms.txt` is the first widely-adopted agent-first standard. a single text file at the site's root that tells agents what's available.

 it's a small step. but it signals intent: "agents can read these docs too."

 current adoption is low but growing fast. mintlify and fumadocs support it natively.

 ## what frameworks are doing

 | framework | agent-readiness |
 |---|---|
 | mintlify | high (`.md` endpoints, llms.txt, nav api) |
 | fumadocs | high (`.md`, llms.txt) |
 | docusaurus | medium (raw md in repo, no exposed endpoints) |
 | gitbook | medium (partial `.md` support) |
 | nextra | high (next.js ± raw file serving) |

 modern frameworks are becoming agent-aware by default.

 ## what agent cache does

 agent cache is a bridge. it takes human-first documentation and converts it to agent-ready format.

 whether the docs are agent-first or human-first, agent cache extracts the useful content and converts it to structured markdown.

 but agent-first sites are easier. when `.md` endpoints exist, extraction is trivial.

 the long-term trend: docs frameworks will expose agent-ready content natively. agent cache will evolve from a converter to a packager.

 ## my bet: hybrid, leaning agent-first

 documentation should always be human-readable. but it should also be machine-consumable.

 the frameworks that do both well will win. mintlify and fumadocs are leading.

 the future: docs sites are for humans. underneath, they're structured markdown with rich metadata. both audiences served.

 ---

 **related:**
 - [mintlify vs docusaurus: framework rankings](/blog/docs-framework-agent-readability)
 - [where to find raw docs before scraping](/blog/acquisition-ladder)

---

<!-- https://agentcache.run/blog/agent-hallucination-docs -->

# why your agent hallucinates apis (and how docs fix it)

> your agent just used an api that doesn't exist. here's why agents hallucinate apis, why the fix is documentation, and how agent cache prevents it.

your agent just wrote:

```javascript
const user = await stripe.users.createUser({ name: "John" })
```

the real stripe api is:

```javascript
const customer = await stripe.customers.create({ name: "John" })
```

the agent hallucinated. `users.createUser` doesn't exist. `customers.create` does.

this is the #1 problem with coding agents. and the fix isn't a better model. it's better docs.

## what agent hallucination looks like

hallucination in coding agents isn't dramatic. it's subtle. small wrongnesses that compile but break at runtime.

examples:
- wrong parameter names (`email` instead of `email_address`)
- wrong method names (`delete` instead of `remove`)
- wrong types (string instead of object)
- non-existent features ("stripe supports crypto payments", it doesn't, yet)
- outdated apis (using deprecated methods from training data)

each one is minor. but together, they make generated code unreliable.

## why agents hallucinate

agents don't know facts. they know patterns.

### reason 1: training data cutoff

models are trained on data up to a certain date. new apis, new features, new versions. the model hasn't seen them.

if stripe shipped a new feature last month, the model might not know about it. or worse, it might hallucinate what it _should_ look like.

### reason 2: no documentation in context

agents have limited context windows. if the relevant docs aren't in the context, the agent improvises.

improvisation = hallucination. every time.

### reason 3: pattern matching over facts

models are trained to predict the next token. "what's the most likely next word?" not "what's the correct api signature?"

for common patterns, this works. `p = Person(name="John")` is a safe bet in many languages.

for specific apis with unusual names, it fails. if stripe chose `customers.create` instead of `users.create`, the model might forget which is real.

### reason 4: similar names

libraries with similar names cause cross-pollination. if an agent knows django's orm and fastapi's orm, it might mix them up.

`filter()` in django returns a queryset. in fastapi/sqlalchemy it's different. the model might apply the wrong behavior.

## give your agent the actual api reference

the only reliable fix: put the actual api documentation in the agent's context.

not a summary. not a third-party tutorial. not stack overflow answers. the actual docs.

when the agent reads the actual docs:
- exact method names are correct
- parameter types match the api
- edge cases are documented
- deprecation warnings are visible
- the agent knows what it doesn't know (missing context vs. hallucinated confidence)

## how agent cache prevents hallucination

agent cache doesn't "prevent" hallucination. it gives you the tool to prevent it.

1. extract documentation for any api
2. put the extracted docs in your agent's context
3. the agent reads the actual docs instead of guessing

before:
> agent: "i'll use `stripe.users.createUser`"

after:
> agent (reading extracted stripe docs): "the method is `stripe.customers.create`, and it takes these parameters..."

## does this actually work?

short answer: yes. longer answer: it reduces hallucination dramatically, but not to zero.

when docs are in context:
- method name accuracy: ~95%
- parameter accuracy: ~90%
- type accuracy: ~85%

still not perfect. but compare to without docs:
- method name accuracy: ~70%
- parameter accuracy: ~50%
- type accuracy: ~40%

the improvement is massive. and it scales: more docs, more coverage, fewer hallucinations.

## other approaches (and why they fall short)

**retrieval-augmented generation (rag):** the agent retrieves docs snippets at query time. works, but incomplete snippets and retrieval errors add their own failure modes.

**context7 / similar:** runtime retrieval apis. good for discovery. but returns snippets, not full docs. the agent might miss edge cases not in the snippet.

**fine-tuning:** train the model on the api docs. expensive. slow. and the model still has the training data cutoff problem.

**local docs (agent cache):** full docs in context. offline. zero api calls. all content available. this is the most reliable approach.

## put the docs in context before asking for code

agents hallucinate because they don't have access to facts. they interpolate from patterns.

documentation is the fact source. put docs in context, and agents stop hallucinating.

it's not about better models. it's about better context.

---

**related:**
- [why local docs beat remote retrieval](/blog/why-local-docs)
- [deterministic extraction vs llm](/blog/deterministic-vs-llm-extraction)
- [keeping reference docs beside your code](/blog/agentcache-dot-folder)

---

<!-- https://agentcache.run/blog/ai-native-documentation-platforms -->

# ai-native documentation platforms: hype or actual future?

> documentation.ai, docsalot, hyperdocs, docsio. ai-native docs platforms are everywhere. i tested them. here's what's real and what's marketing.

2026 is the year of ai-native documentation.

at least, that's what the landing pages say. "ai writes your docs." "ai updates your docs." "ai-powered developer portals."

i looked at the players. tested what i could. here's what's real and what's just seo juice.

## who's selling ai-native docs

**Documentation.AI**: connects to your codebase and prs. continuously writes and updates docs with autonomous agents.

**DocsAlot**: unifies product manuals and dev docs into a single source of truth. optimized for humans and agents. native mcp server.

**Hyperdocs**: automatic change detection from git prs. ai content publishing.

**Docsio**: ingests repos or web content. outputs hosted sites with automatic `llms.txt` generation and mcp support.

**DocuWriter.ai**: inspects source code across 20+ languages. generates api references and tutorials.

**Unmint**: open-source, self-hosted alternative to mintlify. zero subscription fees. (not strictly ai-native, but positioned against saas tools.)

**Jamdesk**: docs-as-code with ai search and api testing widgets.

**Papervine**: git-synced docs with ai assistance.

## connect a repo, get updated docs?

the pitch is attractive:

1. connect to your github repo
2. ai reads your code
3. ai generates documentation
4. ai updates docs when code changes
5. human reviews, approves, publishes

sounds like the end of "docs are outdated" forever.

## what i could check in public demos and tools

i tested the parts i could access (public demos, open-source tools, documentation):

### what works well

**auto-generating api references from code.** this is the oldest and most solved problem. tools like jsdoc, typedoc, and rustdoc have done this for years. ai-native tools just wrap this with better presentation.

**summarizing existing content.** if you have a 5,000-word design doc, an ai can produce a 500-word summary. this is useful. it's also not new, just faster and more accessible.

**detecting stale docs.** comparing code changes to doc files and flagging "this section might be outdated." this is genuinely useful. manually checking every doc page after a release is tedious. automation helps.

**generating `llms.txt` and structured metadata.** automatically creating the files that make docs agent-readable. low-hanging fruit, but important.

### what doesn't work well

**writing conceptual documentation from code.** code tells you what a function does. it doesn't tell you why a design decision was made, what tradeoffs exist, or how concepts relate to each other. ai-generated conceptual docs read like expanded api references, technically accurate, conceptually hollow.

**understanding domain-specific context.** "this endpoint accepts a `user_id`." ai knows that. "this endpoint accepts a `user_id` which must be the same as the authenticated user's id, or the request returns 403". the second part comes from business logic, not code. ai misses this constantly.

**maintaining tone and voice.** every company has a docs voice. stripe is precise and terse. twilio is friendly and tutorial-heavy. vercel is modern and minimalist. ai generates generic corporate docs-speak. it doesn't match your voice without extensive tuning.

**handling edge cases and errors.** code paths that don't execute in the happy path. error states. race conditions. timeout handling. these are documented in comments, design docs, and tribal knowledge. ai can't extract what isn't explicitly written.

## "ai-native" covers very different products

what makes a documentation platform "ai-native"?

**Documentation.AI:** ai writes the docs. true ai-native.

**DocsAlot:** ai optimizes for agents. partially ai-native.

**Hyperdocs:** ai detects changes. partially ai-native.

**Mintlify:** has ai search and semantic features. is this ai-native? they don't claim the label, but feature-wise they're similar.

**Docusaurus:** has ai plugins. not ai-native by any definition.

the line is blurry. "ai-native" is being applied to anything with an ai feature. it's becoming meaningless.

## comparison: ai-native vs traditional + ai features

| feature | ai-native (docsalot, documentation.ai) | traditional + ai (mintlify, docusaurus + plugins) |
|---|---|---|
| ai writes docs from scratch | yes | no (human writes, ai assists) |
| ai updates docs on code changes | yes | partial (flagging, not rewriting) |
| ai search | yes | yes (mintlify has this) |
| human review workflow | varies | yes |
| cost | higher (ai generation tokens) | lower (hosting only) |
| output quality | mixed (good for references, weak for concepts) | depends on human writer |
| agent-readiness | varies | mintlify, fumadocs excellent |

## references are easier to automate than explanations

ai-native documentation tools are useful for specific tasks:
- api reference generation
- stale doc detection
- content summarization
- metadata generation (`llms.txt`)

they are not useful (yet) for:
- conceptual explanation
- architectural decision records
- tutorial writing
- voice and tone consistency
- edge case documentation

the best workflow in 2026 is hybrid: humans write conceptual and tutorial content. ai assists with references, summaries, and maintenance. tools that understand this hybrid model (mintlify's approach, human writes, ai enhances) are more practical than pure ai generation.

## will they replace traditional docs platforms?

no. not in their current form.

traditional docs platforms (docusaurus, mintlify, vitepress) are the foundation. ai-native tools are a layer on top.

the realistic future:
1. traditional platform hosts the docs
2. ai tools augment: generate references, detect staleness, suggest updates
3. humans review and approve
4. published docs are traditional + ai-assisted

this is what mintlify is already doing. what docusaurus will do with plugins. what github copilot will do for code comments.

pure ai-generated documentation without human review is unreliable. pure human documentation without ai assistance is expensive to maintain.

the middle path wins.

## use ai for maintenance, keep humans on concepts

ai-native documentation is a real category. the tools do things that weren't possible five years ago.

but "ai-native" is being stretched beyond meaning. a docs platform with ai search is not "ai-native." a tool that flags stale docs is not "ai-native." true ai-native means ai is generating or substantially rewriting content. and that's where the quality problem lives.

use ai for what it does well: reference generation, summarization, stale detection. don't expect it to replace human judgment on concepts, tone, and edge cases.

---

**related:**
- [why i extract docs without an llm](/blog/why-no-llm-extraction)
- [deterministic extraction vs llm](/blog/deterministic-vs-llm-extraction)

---

<!-- https://agentcache.run/blog/api-documentation-tools-ranked -->

# api documentation tools: from openapi to graphql

> scalar, redocly, swagger, and 15+ other api doc tools compared. which ones are agent-readable and which ones hide content behind javascript?

api documentation is different from product documentation. it's not prose and tutorials. it's schemas, endpoints, parameters, and response objects.

the tools that generate api docs are also different. they don't take markdown. they take openapi specs, graphql schemas, asyncapi definitions. they turn structured machine-readable specs into human-readable websites.

here's how they work, which ones are good, and which ones break extraction.

## openapi renderers: the big three

openapi (formerly swagger) is the standard for rest api documentation. a single `openapi.json` or `openapi.yaml` file describes every endpoint. the tools render it.

### scalar

newer, faster, and simpler than the alternatives. open source. built-in api testing client. code snippets in 15+ languages. custom themes. offline-first.

**agent-readiness:** the openapi spec is the source of truth. it's a json file. agents can read the spec directly. the rendered website is for humans. extraction: grab the spec file. done.

**verdict:** if you're choosing a new openapi renderer, start here. it's modern, fast, and the spec-first approach is inherently agent-friendly.

### redocly

the enterprise standard. redoc pioneered the three-column layout. redocly expanded it into a full developer portal platform with linting, bundling, search, and ci/cd integration.

**agent-readiness:** same as scalar. the openapi spec is readable by agents. redocly adds a developer portal layer that's human-focused.

**verdict:** if you need enterprise features (multi-spec portals, team collaboration, governance), redocly is the choice. for simple api docs, scalar is leaner.

### swagger ui

the original. still maintained by smartbear. executes live api calls from the browser. the tool that made "interactive api docs" a standard expectation.

**agent-readiness:** old, bulky, and renders everything client-side. the spec is accessible at `/openapi.json`, but the html is javascript-heavy. for agents, read the spec file directly.

**verdict:** legacy choice. works. stable. but scalar and redocly are better modern options.

## modern alternatives worth knowing

**stoplight elements**: embeddable web components. converts openapi + asyncapi into standalone web apps or react portals. good if you need to embed api docs inside an existing application.

**rapidoc**: zero-dependency web component. parses openapi specs directly inside standard html. no build pipeline. just a `<script>` tag.

**zuplo**: edge-based api gateway that auto-converts openapi specs to hosted developer portals. includes rate-limiting controls and live request consoles.

**bump.sh**: hosted api portals that track breaking changes between spec releases. ci/cd integration. good for teams that release api versions frequently.

## markdown output and guides alongside specs

**widdershins**: converts openapi 3.0, swagger 2.0, and asyncapi definitions into clean markdown. formatted for docusaurus, mkdocs, or slate. this is actually the most agent-friendly tool in the category: it produces markdown that agents can read natively.

**spectacle**: generates static html from openapi specs using handlebars templates. multi-page or single-page output.

**dapperdox**: embeds markdown conceptual guides alongside openapi endpoints. unified developer portal.

## graphql documentation: different problem

graphql apis don't have openapi specs. they have schemas. the documentation tools extract the schema via introspection and render it.

**magidoc**: static site generator built for graphql. inspects schemas via introspection. outputs svelte-powered searchable sites. the introspection query is the agent-readable source. the rendered site is for humans.

**spectaql**: node.js generator. parses schema files or introspection queries. multi-column html with customizable css. also outputs markdown if configured.

**graphdoc**: static html from graphql schemas. no runtime dependencies. simple and lightweight.

**agent-readiness for graphql:** introspection queries (`__schema`, `__types`) are the machine-readable source. agents can query these directly. the rendered documentation is a human convenience.

## protobuf, asyncapi, and other protocols

**asyncapi generator**: official tooling for event-driven architectures. websockets, kafka, mqtt topics. outputs html, markdown, or react apps. the asyncapi spec is machine-readable.

**protoc-gen-doc**: plugin for the protocol buffer compiler. extracts inline comments from `.proto` files. outputs html, markdown, or json. the `.proto` files with comments are the source.

**typespec**: microsoft's language for api definitions. typescript-like syntax. compiles to openapi + html docs. also outputs the spec file.

## fetch the spec instead of parsing the site

for api documentation tools, the rendered website is almost always secondary to the spec.

| tool | machine-readable source | extraction strategy |
|---|---|---|
| scalar | `openapi.json` | fetch spec directly |
| redocly | `openapi.json` | fetch spec directly |
| swagger ui | `openapi.json` | fetch spec directly |
| stoplight | `openapi.json` | fetch spec directly |
| magidoc | introspection query | run introspection |
| spectaql | schema file / introspection | read file or query |
| graphdoc | schema file | read file |
| asyncapi | `asyncapi.yaml` | fetch spec directly |
| typespec | `tsp` files + compiled spec | read source or output |

for all of these, the best extraction strategy is: don't extract the website. extract the spec.

this is different from product documentation (docusaurus, mintlify) where the site *is* the content. for api docs, the site is frosting. the spec is the cake.

## when the spec isn't available

some teams don't publish their openapi spec publicly. they only publish the rendered docs. in those cases:

- scalar and redocly: the spec is usually embedded in the page json or available at a predictable url
- swagger ui: check `/openapi.json`, `/swagger.json`, or `/api-docs`
- custom implementations: might require html parsing to extract the schema

if the spec truly isn't exposed, extraction becomes hard. but this is rare. most api docs tools want the spec to be accessible.

## choose a renderer for people, a spec for agents

api documentation tools are simpler to extract than product docs because they have a machine-readable source: the spec.

scalar is the modern default. redocly for enterprise. widdershins if you want markdown output. everything else is a variation.

for agents: read the spec. ignore the website.

---

**related:**
- [pure docs generators ranked](/blog/pure-docs-generators-ranked)
- [browser-rendered docs break extraction](/blog/browser-rendered-docs-extraction-problem)

---

<!-- https://agentcache.run/blog/architecture-deep-dive -->

# agent cache architecture: zero-memory, disk-backed, sse-streamed

> agent cache uses a zero-memory architecture where nothing stays in ram. everything streams to disk. here's how i built a deterministic, replayable docs extraction pipeline.

most web apps hold state in memory. they hope the process doesn't crash. they hope the server doesn't restart. they hope nothing needs to replay.

agent cache uses a different rule: **zero in-memory state.** every log line, every discovered page, every progress update streams directly to disk. this makes the system crash-proof, replayable, and surprisingly simple.

here's the full architecture.

## if it matters, write it to disk

rule: **no state is stored in memory unless it can be reconstructed from disk.**

this means:
- no in-memory job queues
- no in-memory caches with "back to cache later"
- no "i'll write to disk at the end"
- if the process dies mid-crawl, restart and resume from disk

everything that matters goes to an append-only event ledger on disk. everything else is disposable.

## every job gets an append-only event log

every extraction job gets a directory: `storage/jobs/ac-{id}/`

inside:
- `events.jsonl`: append-only event log. every event timestamps and serializes to this file.
- `logs.txt`: human-readable logs for terminal viewing
- `final/`: extracted docs in their final structure
- `bundle.zip`: the final packaged output

the event log is the source of truth. it's append-only. no events are ever modified or deleted. to "cancel" a crawl, i append a cancellation event.

this makes the system trivially replayable. any agent or human can `cat events.jsonl` and see exactly what happened.

## from url to zip in five phases

```
User Input URL
     │
     ▼
[ Phase 1: Resolver ] → Normalize URL, find canonical docs endpoint
     │
     ▼
[ Phase 2: Topology ] → Detect framework, extract navigation structure
     │
     ▼
[ Phase 3: Ladder ]   → Try extraction strategies in cost order
     │
     ▼
[ Phase 4: Crawler ]  → Fetch and convert markdown (8-12 workers)
     │
     ▼
[ Phase 5: Packager ] → Build INDEX.md, meta.yaml, _map.json, zip
```

each phase is isolated. it reads from the previous phase's outputs and writes to disk. no shared memory.

## phase 1: resolver

takes any input url. finds the canonical docs endpoint.

examples:
- `stripe.com` → `docs.stripe.com`
- `github.com/stripe/stripe-node` → `docs.stripe.com` (not the repo)
- `context.dev` → `docs.context.dev`

probes subdomains (`docs.*`, `developer.*`, `api.*`). checks for redirects. verifies the endpoint returns actual documentation.

if no docs endpoint is found, the job fails early with a clear error. no point crawling a marketing page.

## phase 2: topology mapper

detects the docs framework and extracts the navigation structure.

framework detection queries:
- html meta tags: `<meta name="generator" content="docusaurus">`
- specific css classes: `.theme-doc-main` for docusaurus, `data-mintlify` for mintlify
- known html patterns: specific `<nav>` structures
- url patterns: `/docs/getting-started` structure

once detected, framework-specific extractors parse the navigation:
- **mintlify:** reads the navigation from a known api endpoint
- **docusaurus:** extracts from the sidebar json or html
- **fumadocs:** reads version tabs and group structure
- **generic:** falls back to sitemap parsing

the topology is a tree: chapters → sections → pages. this structure drives phase 4.

## phase 3: acquisition ladder

the ladder was covered in detail in a separate post. but architecturally, it's a decision tree that runs per-page.

for each discovered url, the ladder asks:
1. llms.txt available? → use it
2. direct .md endpoint? → fetch it
3. github tree? → download raw files
4. content negotiation works? → use it
5. html purification → fall back

each tier writes to disk independently. no tier depends on another tier's memory state.

## phase 4: crawler with worker pool

8-12 concurrent workers. each worker:
1. reads a url from the discovered list
2. runs it through the acquisition ladder
3. writes the extracted markdown to disk
4. appends a completion event

workers are independent. one worker crashing doesn't affect others. the main process restarts crashed workers.

concurrency is bounded: 8-12 workers is the sweet spot. below 8, i leave speed on the table. above 12, diminishing returns and potential for rate limiting.

## phase 5: packager

after all pages are extracted, the packager:
1. reads all extracted markdown files from disk
2. generates `meta.yaml` with metadata (name, url, keywords, etc.)
3. generates `_map.json` with the navigation tree
4. generates `INDEX.md`: a table of contents
5. structures the output: `docs/<chapter>/<section>/<page>.md`
6. zips everything into `bundle.zip`
7. uploads to r2 for permanent storage

this phase is single-threaded and io-bound. disk → cpu → disk → network.

## sse streaming: progress without websockets

the ui shows extraction progress in real-time. sse (server-sent events) streams events from the `events.jsonl` file to the browser.

when a user visits `/dingdong/ac-{id}`, the server:
1. opens the event log
2. replays all historical events via sse
3. keeps the stream open for new events
4. buffers and streams as events happen

if the user refreshes, events replay again. no state lost.

this is simpler than websockets. no connection management. no reconnect logic. just read the file and stream.

## storage: local disk + r2

**local disk** is the working store. everything during extraction lives here. fast writes. fast reads.

**r2 (cloudflare)** is the final store. completed bundles are uploaded here. r2 has zero egress fees, which matters when users download zip files.

turso (libsql) stores job metadata: job id, url, status, timestamps. tiny data. sqlite handles it fine.

## dual-layer error handling

every error generates two outputs:

**machine error:** goes to logs. technical. exhaustive. includes http status, stack traces, cloudflare ray ids, dns codes. for debugging.

**human error:** goes to the user. witty, actionable, specific. never "something went wrong." always "cloudflare blocked the request with a challenge page. try again or use a different url."

machine errors are for me. human errors are for users. both are correct for their audience.

## why this architecture wins

**crash-proof:** process dies? restart and replay from the event log. crawling resumes where it stopped.

**debuggable:** `cat events.jsonl` tells the entire story. every decision. every failure. every retry.

**scalable (not in the cloud sense):** no redis. no message queues. no kubernetes. just disk and processes. a $6 vps handles everything.

**simple:** fewer moving parts than a typical microservices setup. one process. linear flow. predictable.

**observable:** the event log is a perfect audit trail. see exactly what happened, when, and why.

## tradeoffs

- **disk io is slower than memory.** but for this workload, it doesn't matter. extraction is network-bound (fetching pages), not cpu-bound.
- **not horizontally scalable.** i don't need horizontal scaling. one vps handles everything.
- **replay takes time.** if a job has 10,000 events, replaying them takes a few seconds. acceptable.

## disk, one process, and a browser stream

the architecture is intentionally boring. no fancy distributed systems. no kubernetes. no event sourcing framework.

just: write to disk. read from disk. stream to browser.

the simplicity is the feature.

---

**related:**
- [cloudflare workers killed my serverless dream](/blog/cloudflare-workers-50-subrequest-limit)
- [trying cheaper extraction methods first](/blog/acquisition-ladder)
- [100 sites extracted](/blog/100-docs-sites-what-broke)

---

<!-- https://agentcache.run/blog/bot-protection-docs-extraction -->

# bot protection is the last mile problem for docs extraction

> 12% of documentation sites block extraction because of cloudflare, captchas, and wafs. here's how bot protection blocks crawlers and what i do about it.

you'd think documentation, meant to be read, would be easy to access. but 12% of the sites i tried to extract were blocked by bot protection.

cloudflare challenges. captchas. rate limits. wafs. they treat documentation crawlers like threats.

this is the last mile problem of docs extraction: the content is public, but you can't programmatically fetch it.

## public in a browser, blocked over curl

documentation is public information. it's meant to be read. but modern web infrastructure treats non-browser requests as suspicious.

if your user-agent says "curl" or "node-fetch" or doesn't match a known browser, cloudflare issues a challenge. solve the javascript challenge, or no access.

this is reasonable for e-commerce. reasonable for news sites. but for documentation? it's a barrier to legitimate use.

## how bot protection works (simplified)

bot protection has layers:

**layer 1: ip reputation.** known data center ips get flagged. aws, digitalocean, linode ranges are often blocked by default.

**layer 2: request fingerprint.** user-agent string, headers, tls fingerprint. curl and node-fetch leave obvious fingerprints.

**layer 3: javascript challenge.** the server returns an html page with javascript. the script must execute and solve a math problem or set a cookie. non-browsers can't do this.

**layer 4: behavioral analysis.** number of requests per second. patterns. humans don't fetch 100 pages in 30 seconds. scripts do.

**layer 5: captcha.** image or puzzle challenge. 

## cloudflare: the biggest blocker

cloudflare is used by ~40% of docs sites. their "browser integrity check" challenges non-browser requests.

symptoms: 403 forbidden with cloudflare headers. or a 503 with a challenge page. or a redirect to a captcha.

some cloudflare configs are aggressive. some are relaxed. there's no way to know without trying.

## rate limiting: 429 too many requests

even sites without full bot protection have rate limits. too many requests too fast = temporary block.

my approach: polite crawling. 8-12 workers. random delays between 200-500ms. user-agent string that identifies me honestly.

i'm not trying to ddos anyone. i'm trying to download documentation.

## why docs sites turn bot protection on

why do docs sites need bot protection?

- **scrapers:** competitors scraping pricing, features, content
- **ai training data:** companies scraping docs to build training datasets
- **bandwidth:** large crawlers consuming significant bandwidth
- **default settings:** cloudflare's default includes bot protection. many sites never disable it

most docs sites don't intentionally block crawlers. they just never disabled the default settings.

## my approach: polite crawling first

1. **identify my crawler honestly.** user-agent includes "agent-cache" and a contact url.
2. **rate limit.** 8-12 concurrent requests max. delays between requests.
3. **respect robots.txt.** if a site blocks me there, i don't crawl.
4. **no headless browsers.** i don't use selenium/puppeteer. they're slow and resource-intensive. if a site requires javascript execution, i mark it as blocked.
5. **retry with backoff.** transient errors get retries with exponential backoff.

## when polite crawling fails

if cloudflare issues a challenge and my crawler can't solve it, i have options:

**option 1: accept the failure.** 12% of sites won't extract. i log the failure and move on. this is my default.

**option 2: use a headless browser.** i don't do this currently. but it's technically possible. tradeoffs: 10x slower, 10x more expensive, 10x more complex.

**option 3: use a paid extraction service.** firecrawl and similar services have infrastructure (residential proxies, browser farms) that can bypass challenges. tier 6 of my acquisition ladder. i don't use it often.

## what i don't do

- i don't use residential proxy networks
- i don't fake browser fingerprints
- i don't solve captchas
- i don't bypass authentication
- i don't crawl behind paywalls

if a site actively doesn't want to be crawled, i respect that.

## llms.txt could make crawling unnecessary

bot protection exists because scraping is adversarial. but if sites expose `llms.txt`, there's no need to scrape. one request. one text file. no crawling. no bot protection triggers.

this is why llms.txt matters. it changes the dynamic from "crawler vs. protector" to "friendly request for agreed-upon content."

llms.txt adoption is growing. mintlify and fumadocs support it natively. when it's widespread, bot protection ceases to be a docs extraction problem.

## crawl politely and accept a blocked request

bot protection is the biggest unsolved problem in docs extraction. not parsing html. not handling custom frameworks. just getting a basic http response.

i handle it by being polite, respecting robots.txt, and accepting that 12% of sites won't work.

use llms.txt if you're a docs site maintainer. it's the clean fix for everyone.

---

**related:**
- [100 sites extracted: what broke](/blog/100-docs-sites-what-broke)
- [my robots.txt policy](/blog/robots-txt-crawling-ethics)
- [how llms.txt fits into downloading docs](/compare/llms-txt)

---

<!-- https://agentcache.run/blog/browser-rendered-docs-extraction-problem -->

# docsify, markdoc, and the problem with browser-rendered documentation

> docsify renders markdown in the browser. markdoc requires custom parsing. here's why browser-rendered documentation is a trap for ai agents.

most documentation sites serve html. the content is in the response. you fetch the url, you get the text.

some don't. they serve a shell, a nearly empty html file with a javascript payload. the browser runs the javascript. the javascript fetches markdown. the markdown renders into the page.

this is browser-rendered documentation. and it's an extraction nightmare.

## how normal docs work

server-rendered docs (docusaurus, mintlify, vitepress):

```
GET /docs/getting-started
→ server sends complete html with all content
→ extraction: trivial (just read the html)
```

browser-rendered docs (docsify):

```
GET /docs/getting-started
→ server sends empty shell + javascript
→ javascript loads: "oh, this route means load `/getting-started.md`"
→ javascript fetches `/getting-started.md`
→ javascript renders markdown to html in the browser
→ extraction: impossible without a headless browser
```

## docsify: the textbook case

docsify is a popular choice because it's simple:

```html
<!-- index.html -->
<script src="//cdn.jsdelivr.net/npm/docsify/lib/docsify.min.js"></script>
<div id="app"></div>
```

that's it. one html file. the rest is markdown files served statically. docsify's javascript reads the url, finds the corresponding `.md` file, fetches it, and renders it.

for a human with a browser, this is fine. fast, simple, no build step.

for extraction, it's terrible.

### why extraction fails

when i fetch `https://docs.example.com/getting-started`:
- html response: `<div id="app"></div>` and a `<script>` tag. no content.
- the content is at `https://docs.example.com/getting-started.md`. but how do i know that? docsify's routing logic is in the javascript. i'd need to execute the javascript to know where to look.

i could try a heuristic: "if it's docsify, append `.md` to the path." this works for basic cases.

but docsify supports:
- custom route mappings (`/quickstart` → `/docs/start.md`)
- nested sidebar configurations that change routing
- plugins that modify content loading
- base path configurations
- relative paths that depend on the current route

each of these breaks the simple heuristic. without executing docsify's javascript, i can't know the actual content urls.

### running a browser costs more than fetching a page

the "solution" is a headless browser: puppeteer, playwright, selenium. a real browser that executes javascript and waits for content to appear.

this works. but:
- **10x slower.** launching a browser, executing js, waiting for render: seconds per page. vs milliseconds for static html.
- **10x heavier.** chromium is ~150mb. my extraction runs on a small vps. can't afford that.
- **unreliable.** javascript errors, infinite loops, network timeouts inside the browser. each is a failure mode.
- **rate limit trigger.** headless browsers are easy to detect. sites that block bots will block these faster than simple http requests.

for agent cache, headless browsers are tier 6. last resort. i avoid them.

## markdoc: custom parsing required

markdoc is different. it's not browser-rendered, it's processed server-side. but it extends markdown with custom syntax that requires a parser.

markdown with markdoc tags:

```markdown
# getting started

{% callout type="check" %}
make sure you have node.js 18+ installed.
{% /callout %}

{% tabs %}
{% tab label="npm" %}
```bash
npm install @stripe/stripe-js
```
{% /tab %}
{% tab label="yarn" %}
```bash
yarn add @stripe/stripe-js
```
{% /tab %}
{% /tabs %}

{% card title="next steps" href="/docs/authentication" %}
learn how to authenticate your requests.
{% /card %}
```

this isn't standard markdown. `{% callout %}`, `{% tabs %}`, `{% card %}`. these are markdoc-specific tags. they compile to react (or vue, or svelte) components.

if i extract the raw markdown, an agent sees:
- `{% callout type="check" %}` instead of a formatted callout
- `{% tabs %}` instead of tabbed content
- component references instead of rendered output

the agent can still read the text content. but it loses structure, formatting, and component semantics.

### markdoc's extraction complexity

to properly extract markdoc content, i'd need to:
1. identify that the site uses markdoc (difficult, no standard marker)
2. understand the site's markdoc schema (each site defines different tags)
3. parse the custom syntax
4. convert tags to meaningful markdown equivalents

markdoc is open-source. stripe published it. but every markdoc site has a different schema. there is no universal markdoc extractor.

for agent cache, markdoc sites fall into "custom framework" territory. site-specific extraction. expensive to maintain.

## other browser-rendered docs tools

**notion public pages:** notion serves a shell with heavy javascript. content loads via api calls. extraction requires reverse-engineering notion's private api. i skip notion entirely.

**some next.js apps with client-side data fetching:** a poorly configured next.js docs site might fetch content in `useEffect` instead of server components. same problem as docsify, empty shell on initial load.

**swagger ui / redoc:** these render openapi specs in the browser. the spec file is downloadable (usually `/openapi.json`), but the rendered html is generated client-side.

## easy for readers, expensive for crawlers

for end users, browser rendering is fine. the page loads, content appears. maybe a flicker. acceptable.

for extraction, it's catastrophic:

| approach | speed | reliability | cost | agent-readable output |
|---|---|---|---|---|
| static html (docusaurus, vitepress) | fast | reliable | $0 | excellent |
| server-rendered (next.js ssr, mintlify) | fast | reliable | $0 | excellent |
| browser-rendered (docsify, notion) | slow | unreliable | $$$ | requires headless browser |
| custom syntax (markdoc) | moderate | moderate | $0 | needs custom parser |

## what i do about it

for docsify sites, i try a heuristic: detect the framework, guess the markdown url pattern, probe for `.md` files. works for simple sites. fails for complex configurations.

for markdoc, i extract the raw markdown and accept that custom tags come through as-is. agents can still read the text. structure is degraded but not destroyed.

for notion, i skip. too complex.

for sites that truly require headless browsers: i mark as failed. agent cache doesn't do headless rendering. that's a different product.

## advice for docs authors

if you're choosing a docs tool and care about agent accessibility:

**do:** mintlify, fumadocs, docusaurus, vitepress, starlight. these render content server-side. extraction is easy.

**think twice about:** docsify. simple for you, hard for machines.

**avoid:** notion as public docs. it's a collaboration tool, not a docs platform.

**if you must use markdoc:** publish the raw markdown alongside the rendered site. or include a `/llms.txt` or `/llms-full.txt` file with clean content.

## factor extraction into your framework choice

browser-rendered documentation is a convenience for authors that becomes a tax on extraction.

every docs framework makes tradeoffs. docsify trades extractability for simplicity. markdoc trades universal parsing for component flexibility. these are legitimate choices.

but in 2026, with ai agents reading docs, extractability is no longer a niche concern. it's a core accessibility feature.

the frameworks that understand this, mintlify, fumadocs, starlight, are winning.

---

**related:**
- [framework extractor rankings](/blog/docs-framework-agent-readability)
- [how stripe's custom framework broke me](/blog/stripe-docs-crawler-challenges)

---

<!-- https://agentcache.run/blog/cloudflare-workers-50-subrequest-limit -->

# cloudflare workers' 50 subrequest limit killed my serverless dream

> i tried to run agent cache on cloudflare workers. the 50 subrequest limit made it impossible for large docs sites. here's why i switched to a vps.

cloudflare workers seemed perfect for agent cache. edge computing. global distribution. serverless pricing. zero cold starts. built-in caching.

i built a proof-of-concept. then i hit the subrequest limit. and realized serverless isn't always the answer.

## why cloudflare workers seemed perfect

agent cache processes documentation extraction jobs. users submit a url. the system crawls the site, converts it to markdown, zips it, and stores it.

sounds like a perfect serverless workload: trigger on request, process in the background, store results, done.

cloudflare workers specifically:
- 0ms cold starts (v8 isolates)
- runs at the edge (low latency for users)
- integrates with r2 (my storage)
- durable objects for state management
- native fetch api for crawling

i was sold.

## 50 fetches won't cover a typical docs site

cloudflare workers limits each request to **50 subrequests**.

what's a subrequest? any `fetch()` call. to external domains. to r2. to durable objects. to the cache api.

each one counts.

for a single-page extraction, 50 subrequests sounds like plenty. but documentation sites have hundreds of pages. thousands, sometimes.

counting the requests: extracting a medium-sized docs site like supabase or convex requires fetching:
- the main page (1)
- the sitemap (1)
- navigation/sidebar to discover structure (2-3)
- each individual page content (100-500)
- images/assets (optional, i skip these)

**that's 100+ subrequests for a typical docs site.**

for large sites like stripe (8.7 MB of docs, thousands of pages), it's 500+ subrequests.

cloudflare workers caps at 50. total.

## stripe would need chained batches

stripe docs is one of the hardest extractions i do. it's large. it's custom. it has multiple tabs and versions.

breaking that extraction into batches that fit within 50 subrequests per request? practically impossible. you'd need to chain 10+ sequential requests, each triggered by the previous. durable objects for state. queues for orchestration.

by the time you build all that, you've reinvented a server architecture. poorly.

## workarounds i tried

### workaround 1: batching pages per request

fetch multiple pages in a single request? no, cloudflare counts each `fetch()` separately. one request per page = one subrequest per page.

### workaround 2: durable objects for state

durable objects can persist state across requests. so you could:
1. request 1: fetch 50 pages, store in durable object
2. request 2: fetch next 50 pages, store in durable object
3. repeat until done

this works but adds complexity. and each durable object request counts as a subrequest too. you get maybe 40 page fetches + 10 durable object writes per request.

for a 500-page site: 500 / 40 = 13 sequential requests. taking 15+ seconds total. not terrible, but not great.

### workaround 3: waiting for a limit increase

cloudflare's enterprise plan can increase limits. but this is a side project. enterprise is $5,000+/month minimum.

not happening.

## why a vps monolith won

i switched to a vps. a single $6/month instance running bun + hono.

no subrequest limits. crawl 500 pages concurrently. use as many `fetch()` calls as you want. store results on local disk. zip them. upload to r2.

the whole extraction pipeline runs in one process. no state management. no queuing. no durable objects. the simplicity is beautiful.

portability story? serverless is more portable. but for this workload, the vps is actually easier to reason about. one process. linear execution. predictable.

## when cloudflare workers *does* work

workers is perfect for:
- simple api endpoints (authentication, webhooks)
- lightweight proxying
- small-scale scraping (under 50 pages)
- edge caching

it's just not designed for crawling documentation sites with hundreds of pages.

## crawling doesn't fit a short, stateless request

serverless is great for request-response patterns. small units of work. stateless transformations.

crawling is different. it's stateful. sequential. io-heavy. requires coordination across many requests.

the industry pushes serverless as the default. but defaults are just defaults. they're not rules.

for documentation extraction, a traditional server is the pragmatic choice. sometimes boring architecture is the right architecture.

## i'm keeping the crawler on a $6 vps

cloudflare workers is great technology. i'm a fan. i've used it for other projects.

but for a documentation crawler that needs to fetch hundreds of pages? the 50 subrequest limit is a dealbreaker. a $6/month vps handles it better.

the lesson isn't "serverless is bad." the lesson is **"match your architecture to your workload."**

---

**related:**
- [agent cache architecture deep dive](/blog/architecture-deep-dive)
- [why i chose hono over express](/blog/why-hono-over-express)
- [the economics of docs extraction](/blog/economics-docs-extraction)

---

<!-- https://agentcache.run/blog/co-founder-validation -->

# why i built agent cache, why i never built the llm crawler, and what it takes for us to build this together

> raw facts on agent cache for a prospective co-founder: live database metrics, why the llm extractor was never built, reddit validation signals, and what we do next.

to anyone reading this as a prospective co-founder: i am not pitching you. no pitch deck, no tam charts, no hype.

if we team up, you need ground reality: what works, what broke, live database numbers, why i never implemented the llm extractor, and why i will drop this project if developers do not pull it from our hands.

here is the unvarnished brief.

## 1. live numbers

the tool is live at [agentcache.run](https://agentcache.run). paste a docs url, get a clean markdown zip with `meta.yaml` and `_map.json`. coding agents read local files offline instead of wasting context on live html.

live stats from our turso database:

- **129 crawl jobs** run.
- **94 completed**.
- **18 failed** (dns drops, bot protection, zero discoverable routes).
- **17 jobs** crawling, pending, or probing.
- **23,904 pages** extracted to markdown.
- **70.6 mb** of compressed archives generated.
- **4,047 unique visitor sessions** across 4,733 logged events.

top crawls:
- stripe: 3,644 pages (9.2 mb)
- supabase: 864 pages (4.2 mb)
- paddle: 610 pages (2.6 mb)
- drizzle orm: 439 pages (1.4 mb)
- convex: 428 pages (1.4 mb)
- resend: 392 pages (1.0 mb)

inspect the live dashboard yourself: [https://agentcache.run/cockpit/analytics](https://agentcache.run/cockpit/analytics). i will send you the admin password over reddit dm.

## 2. the accuracy flaw

our pipeline only catches fatal crashes.

if a crawler hits a 403 or dns error, it marks `failed`. but if a site has 200 pages and our crawler grabs only 25 because the rest sit behind dynamic client-side javascript tabs, it still exits clean.

the UI says "complete". 85% of the docs are missing.

a 200 status is an exit code, not an accuracy benchmark. true fidelity requires automated tree diffs against sitemaps and sidebar dom. until we build that audit, completed counts are partially vanity.

## 3. the llm extractor: why i never wrote a line of it

my first impulse to fix missing dynamic pages was an autonomous llm agent: spin up a headless sandbox, click tabs, expand menus, parse dom trees.

i stopped before writing any code. here is why:

1. **latency and cost**: our deterministic pipeline costs $0 in inference and runs in seconds. an agent sandbox calling multimodal models blows crawl time from 15 seconds to 4 minutes, and compute costs from pennies to dollars per run.
2. **premature paywall**: an llm crawler forces credit meters, stripe checkout, and paywalls immediately (like firecrawl at $19-$99/mo or context7 at $10/seat). charging money before validating repeat retention kills adoption.

rule: free tier uses zero-cost primitives only. no paid inference until retention is proven.

## 4. reddit validation

pageviews prove curiosity, not retention. to test real pain, i messaged developers using cursor, claude code, and codex on reddit:

> *"when you use coding agents, do you ever have to fetch or paste library docs yourself, or do they usually get the right docs? i'm researching this workflow before building anything and would value your take. no pitch."*

outreach numbers from our tracker:
- 21 leads identified
- 18 contacted 1:1
- 9 replies (50% response rate)

the split:
- **no pain**: one developer said they never struggle getting docs into agents. standard libraries (react, express) already live in model weights. they do not need this tool.
- **real friction**: other builders hit walls with fast-moving libraries, hallucinated methods, context limits, and metered cloud query fees. they want local, offline markdown.

if only a tiny fraction feels acute friction, this is a weekend hobby. if that friction grows as agents tackle niche libraries, it is a company.

## 5. our next 5 validation moves

we do not write heavy infra next. we tighten validation:

1. **real analytics**: replace our raw sqlite event logging with posthog or plausible to track repeat user retention.
2. **google search console**: read actual search queries to see which library docs developers search for.
3. **daily build-in-public posts**: share what breaks across x, reddit, linkedin, and daily.dev.
4. **move feedback button to the top**: it sits buried at line 458 in `src/app/docs/[id]/page.tsx`. put it in the top header so developers flag bad crawls immediately.
5. **sitemap accuracy benchmark**: compare extracted files against live sitemaps to score real tree completeness.

## 6. bottom line

the future of the llm feature and this startup depends on user pull.

if developers stay indifferent, i will stop the project, open-source the code, and move on. i refuse to build a solution looking for a problem.

if demand holds, we build the lean path first:
1. `.agentcache/` local project convention
2. open-source `agentcache` cli
3. zero-cloud stdio mcp server for cursor and claude code
4. paid llm extraction later, only for complex enterprise portals users cannot scrape themselves

## 7. the co-founder ask

i need a partner who values pragmatism over theater:
- technical alignment: fast and deterministic beats slow and expensive.
- brutal honesty: kill weak ideas early.
- customer curiosity: talk to users, inspect failed crawls, care about ground truth.

the repo is clean, the web app works, 23,000+ pages are cached, and real users are visiting. 

if this sounds like how you build, let's look at the database together and decide the next move.

---

<!-- https://agentcache.run/blog/concurrency-sweet-spot-extraction -->

# 8-12 workers: finding the sweet spot for parallel docs extraction

> i tested 1 to 50 concurrent workers for documentation extraction. 8-12 was the sweet spot. here's why more workers actually hurts.

when crawling in parallel, more workers should mean faster extraction, right?

i tested it. concurrency from 1 to 50 workers. the results were surprising.

**8-12 workers is the sweet spot.** below 8, you're slow. above 12, you hit diminishing returns. at 50, things actually get slower.

## adding workers looks like a shortcut

naive assumption: if 1 worker takes 10 minutes, 10 workers take 1 minute. linear scaling.

reality: extraction has bottlenecks. and they aren't cpu.

## test methodology

i ran the same extraction job (a 500-page docs site) with different concurrency levels:

- 1 worker
- 4 workers
- 8 workers
- 12 workers
- 25 workers
- 50 workers

measured: total extraction time, failed requests, system resource usage.

## results

| workers | time | speedup | notes |
|---|---|---|---|
| 1 | 10:00 | 1.0x | baseline |
| 4 | 4:00 | 2.5x | good scaling |
| 8 | 2:30 | 4.0x | sweet spot start |
| 12 | 2:15 | 4.5x | sweet spot end |
| 25 | 2:05 | 4.8x | marginal gain |
| 50 | 2:30 | 4.0x | slower than 12 |

**50 workers was slower than 12 workers.** why? because of network and server-side limits.

## why more workers isn't always better

**1. server-side rate limiting**

most docs sites have implicit or explicit rate limits. nginx, cloudflare, or application-level limits kick in after a certain request rate.

at 50 concurrent requests to the same site, you're likely to hit a limit. responses slow down. some fail. retry logic adds overhead.

at 8-12, you stay under most rate limits. polite crawling gets better throughput than aggressive crawling.

**2. connection overhead**

each worker maintains http connections. at 50 workers, that's 50 tcp connections, tls handshakes, dns lookups.

established connections can be reused within a worker, but with 50 workers, the overhead of managing connections starts to matter.

**3. memory and context switching**

50 workers means 50 concurrent contexts. memory for each worker. cpu for context switching.

not enough to crash a modern server, but enough to add friction.

**4. diminishing returns on io-bound work**

docs extraction is io-bound. mostly waiting for network responses. mostly idle.

at some point, adding more workers doesn't add more throughput because the network is the bottleneck, not the cpu.

## why i default to 10 workers

8 workers crawl fast without being aggressive. 12 workers push it slightly. both keep extraction reliable.

i default to 10 workers. that covers 95% of use cases.

for small sites with known rate limits, i drop to 4-6.

for large sites with confirmed high rate limits, i might try 15-20. but that's rare.

## adaptive concurrency

i have a simple adaptive strategy:
1. start with 10 workers
2. if requests start failing with 429, reduce by 2
3. if requests succeed consistently for 20 seconds, increase by 1
4. cap at 15, floor at 4

in practice, 95% of extractions stay at 10. only aggressive rate limiters force changes.

## what about 1 worker?

single-worker extraction is _slow_. but it's the most polite. if a site is known to be sensitive, single-worker with 1-second delays between requests is the safest option.

i only use this for sites that have already shown signs of rate limiting on previous attempts.

## cloudflare-protected sites

sites behind cloudflare are the most rate-limit-sensitive. cloudflare's rate limits depend on your "threat score." data center ips score badly. browser-like requests score well.

for cloudflare sites, 6-8 workers with 500ms delays between requests usually works. above that, you trigger challenges.

## start at 10 rather than maxing out concurrency

start with 10 workers. it's fast enough for most cases. polite enough for most servers.

going above 15 is usually counterproductive. going below 8 leaves speed on the table.

parallel extraction is about finding the balance between speed and politeness. 8-12 workers hits that balance.

---

**related:**
- [agent cache architecture](/blog/architecture-deep-dive)
- [bot protection: why 12% of sites fail](/blog/bot-protection-docs-extraction)
- [100 sites extracted](/blog/100-docs-sites-what-broke)

---

<!-- https://agentcache.run/blog/content-negotiation-markdown -->

# content negotiation: getting markdown from sites that don't advertise it

> some docs sites serve raw markdown when you send accept: text/markdown. they don't advertise it. here's how i found them.

most documentation sites serve html. that's the default.

but some sites serve markdown if you ask correctly. send the right http headers, and you get raw markdown instead of rendered html.

this is content negotiation. an underappreciated technique for documentation extraction.

## what is content negotiation

http content negotiation is a standard where the client tells the server what content types it prefers. the server serves the best matching format.

the standard way:
```
GET /docs/page
Accept: text/html, text/markdown
```

if the server supports both, it picks the best match. if it only supports html, it serves html.

but here's the trick: many sites support markdown without advertising it.

## how i discovered hidden markdown endpoints

i discovered this accidentally.

while building the acquisition ladder, i tried appending `.md` to urls. that worked for mintlify and fumadocs. but for some sites, `.md` returned 404.

on a hunch, i tried:
```
GET /docs/page
Accept: text/markdown
```

and got back clean markdown. not html. actual markdown.

the site never advertised this capability. no docs mentioned it. but it worked.

## why sites support this

most modern web frameworks (next.js, sveltekit, nuxt) support content negotiation out of the box. if you have both `page.md` and a `page.tsx` renderer, the framework can serve either based on the accept header.

but many developers don't know about this feature. they never test it. never document it.

the capability exists but is invisible.

## which sites support it

i've found content negotiation support on:
- some docusaurus sites (surprisingly)
- a few nextra sites
- some custom next.js docs sites
- rare gitbook configurations

it's not universal. maybe 3-5% of sites support it. but when it works, it's the fastest path to clean markdown.

## how i detect it

the detection is part of my tier 4 probe:

```
GET /docs/some-known-page
Accept: text/markdown
```

if the response content-type is `text/markdown` → tier 4 works.

if the response is still `text/html` → tier 4 doesn't work. move to tier 5 (html purification).

i test with a known page (usually the getting-started or overview page) because those are most likely to exist.

## why this matters

content negotiation is a clue that a docs site is "agent-friendly." it means the maintainers (or their framework) thought about programmatic access.

even if the site doesn't have `.md` endpoints or `llms.txt`, content negotiation shows the underlying docs are accessible.

it's a tier 4 fallback. not primary. but worth probing because it costs nothing.

## implementation: one header change

```javascript
const response = await fetch(url, {
  headers: { 'Accept': 'text/markdown' }
})

if (response.headers.get('content-type').includes('markdown')) {
  // great, raw markdown
} else {
  // fallback to html
}
```

## where this fits

tier 4 of the acquisition ladder:
1. llms.txt
2. github tree
3. direct .md
4. content negotiation ← you are here
5. html purification
6. paid extraction

## your docs site might already support this

more sites will support content negotiation as developers learn about it. the accept header is a standard http feature. it's free to implement.

the bigger problem is awareness. most developers don't know their framework supports this.

if you're reading this and maintain a docs site: try `curl -H "Accept: text/markdown" your-docs-url`. you might already support it.

## spend one request checking the accept header

content negotiation is a hidden gem. most sites that support it don't know they support it.

if you're building a docs extraction pipeline, probe for this. it's a free tier that costs one http request.

---

**related:**
- [six ways to acquire documentation](/blog/acquisition-ladder)
- [html extraction: the hard path](/blog/html-to-markdown-extraction)

---

<!-- https://agentcache.run/blog/deterministic-vs-llm-extraction -->

# deterministic extraction beats llm-powered summarization

> llms summarize. deterministic extraction preserves. for documentation, you need the whole thing, not an ai's interpretation. here's why code wins.

llms are amazing. they understand text. they generate content. they answer questions. they summarize.

but for documentation extraction, summarization is the problem, not the solution. you don't want your api docs summarized. you want them preserved.

## what deterministic means

deterministic: same input → same output. always.

this isn't a nice-to-have for docs. it's required.

imagine an agent that references docs. today the docs say `function createUser(opts)`. tomorrow, after re-extraction, it says `function createUser(options)`.

nothing changed in the actual api. the llm just interpreted the source differently. your agent is confused.

## problem 1: llms summarize

llms are trained to be helpful. helpful means concise. concise means summarization.

when an llm extracts documentation, it might:
- combine two functions into one "see also" reference
- skip "obvious" parameters
- rewrite examples to be "clearer"
- omit deprecated methods as "unnecessary"

for a human reader, this might be fine. for an agent writing code against an api, this is catastrophic.

## problem 2: llms hallucinate

llms invent things. confidently.

in a documentation context, hallucination means:
- fake parameter names
- non-existent methods
- wrong types
- imaginary return values

i tested this. gave gpt-4 a stripe api docs page and asked it to extract the signatures. it correctly identified most functions. but it confidently reported one parameter as optional when it was required. and it added a fake parameter that doesn't exist.

one error in a thousand lines might not matter. but in code generation, one wrong parameter name breaks everything.

## problem 3: non-deterministic output

temperature 0 doesn't fix it. modern llms still produce slightly different outputs for the same input.

reasons:
- random sampling at the token level
- differences in context window formatting
- model updates between calls

for creative work, this is fine. for documentation, it's a bug.

## problem 4: speed and cost

an llm call for docs extraction: 30 seconds per page. $0.02 per page.

for 100 pages: 50 minutes. $2.

my approach: 1 second per page (direct .md). or 10 seconds per page (html purification).

for 100 pages: 2-15 minutes. $0.

llms are 60x slower and infinitely more expensive.

## problem 5: context limits

llms have limited context windows. 128k tokens is ~100 pages of dense docs.

what about sites with 500 pages? 1,000 pages? you chunk them. process independently. lose cross-page references. lose navigation structure.

my approach handles thousands of pages. just files on disk. no limits.

## i use an llm for metadata after extraction

i do use an llm. for `meta.yaml` generation. after extraction is complete.

why? because keywords and intent triggers are subjective. they require semantic understanding of what the library does. code can't generate good keywords. an llm can.

this is one llm call per site. cheap. and the output is reviewed.

## what my extraction looks like

code. pure code.

- regex for url patterns
- jsdom for html parsing
- framework-specific selectors for content areas
- turndown for html-to-markdown conversion
- string operations for cleaning

no neural networks. no embeddings. no vector stores.

## when llms do make sense

llms are useful for extraction problems that require interpretation:
- sentiment analysis of documentation tone
- inferring relationships between code and docs
- generating summaries (for human consumption, not agent consumption)
- classifying uncategorized docs

but none of these are in the critical path. they're niceties, not requirements.

## html to markdown is a transformation, not a judgment call

ai is for fuzzy problems. extraction is not fuzzy. it's deterministic transformation.

html → markdown. noisy → clean. unstructured → structured.

code does this perfectly. llms add noise, cost, and unreliability.

## keep llms out of the extraction path

if you're building a docs extraction pipeline, skip the llm. use code. deterministic extraction is faster, cheaper, and more reliable.

use llms for what they're good at: understanding, reasoning, creativity.

don't use them for what code does better: transformation.

---

**related:**
- [why i don't use llms for extraction](/blog/why-no-llm-extraction)
- [how i look for raw markdown first](/blog/acquisition-ladder)
- [html extraction how-to](/blog/html-to-markdown-extraction)

---

<!-- https://agentcache.run/blog/documentation-site-generators-2026 -->

# docs generators i'd actually choose in 2026

> 50+ docs tools exist. most don't matter. here's the short list of documentation generators worth your time, with honest picks for each use case.

there are 50+ tools for building documentation sites. i've extracted docs from most of them. here's the truth: most don't matter.

this is the short list. the ones that are actually good. organized by what you should use, not alphabetically.

## tier 1: just use these

if you're starting a docs site today, pick one of these. they all expose markdown natively, work with agents, and won't waste your time.

**mintlify**: saas docs platform. connect your github repo, get a beautiful site, ai search, automatic `llms.txt`, and `.md` endpoints. if you want zero ops and maximum polish, this is it. the yc-backed default for startups in 2026.

**fumadocs**: next.js docs framework. open source, free, absurdly fast. if your project is already next.js, this is the obvious choice. exposes raw markdown endpoints. agent-ready by default.

**starlight**: astro's official docs framework. zero client-side javascript. the fastest docs site possible. built-in i18n, search, accessible navigation. if performance matters more than flashy features, pick this.

**vitepress**: vue-based, evan you's successor to vuepress. clean, fast, minimal. great for vue projects and anything that doesn't need react overhead.

## tier 2: solid, with tradeoffs

these are good. they have larger ecosystems or specific strengths. but they're not as effortless as tier 1.

**docusaurus**: meta's docs framework. huge ecosystem. powers react, redux, jest docs. mature, stable, deeply customizable. but no automatic `.md` endpoints. no auto `llms.txt`. extraction requires going to github or parsing html. good for large open-source projects with complex needs.

**nextra**: next.js + mdx from vercel. simple and effective. raw files in your repo. works well for next.js projects. less feature-rich than fumadocs but simpler to set up.

**gitbook**: from cli tool to venture-backed cloud. real-time editing, github sync, polished output. but custom extraction is harder. content is somewhat locked into their platform.

## tier 3: right tool for the job

these aren't general-purpose docs frameworks. they excel in specific contexts.

**mkdocs**: python docs standard. material for mkdocs theme is everywhere in data science and python projects. if you're python-first, this is your tool.

**sphinx**: the original python documentation engine. powers python's own docs. restructuredtext-based. verbose but thorough. if you need cross-references, api docs, and manuals, sphinx does it. if you just want a quick readme site, it's overkill.

**mdbook**: rust-based. zero dependencies. the tool that builds "the rust programming language" book. if you need a book, not a website, use this.

**11ty**: javascript ssg with zero client-side js. maximum flexibility. can be docs, blogs, anything. requires more configuration than dedicated docs tools.

## tier 4: avoid for new projects

**docsify**: renders markdown in the browser. sounds convenient. no build step. but the content isn't in the html, so extraction is painful. agents can't read it without headless browsers. new projects should use static generators instead.

**honkit**: community fork of the old gitbook cli. backward compatibility is its only reason to exist. don't start new projects with this.

**vuepress**: replaced by vitepress. still maintained but vitepress is the future.

## 50+ tools, grouped by what they build

for the obsessives, here's every tool i know about. most of these have specific niches or are legacy. the tiers above cover what you actually need.

**pure docs ssgs:** docusaurus, mintlify, fumadocs, nextra, starlight, vitepress, vuepress, docus, mkdocs, sphinx, mdbook, 11ty, hugo + docsy/hextra, zola, retype, doctave, honkit, docsify

**api docs renderers:** scalar, redocly, stoplight elements, rapidoc, swagger ui, postman, bump.sh, zuplo, aglio, widdershins, spectacle, dapperdox

**graphql docs:** magidoc, spectaql, graphdoc

**other spec tools:** asyncapi generator, protoc-gen-doc, typespec

**knowledge base / support platforms:** document360, zendesk guide, intercom, help scout, crisp, helpjuice, stonly, freshdesk, hubspot, tidio

**ai-native docs:** documentation.ai, docsalot, hyperdocs, docsio, docuwriter.ai, jamdesk

**developer portals / hosted:** gitbook, readme, fern, developerhub.io, papervine, slite, unmint

## how to choose

**just want docs, zero setup:** mintlify

**next.js project:** fumadocs

**vue project:** vitepress

**python project:** mkdocs (material) or sphinx

**rust project:** mdbook

**maximum performance (zero js):** starlight

**massive open source project with complex needs:** docusaurus

**api docs from openapi spec:** scalar (modern) or redocly (enterprise)

**book-style documentation:** mdbook

## prefer a generator that exposes its markdown

most docs tools are fine. a few are great. the difference between "fine" and "great" is whether your docs are readable by machines as easily as humans.

tier 1 tools understand this. they expose markdown natively. they generate `llms.txt`. they don't hide content behind javascript.

the rest make you work for it. and in 2026, that's increasingly a liability.

---

**related:**
- [mintlify vs docusaurus vs fumadocs](/blog/mintlify-vs-docusaurus-vs-fumadocs)
- [browser-rendered docs break extraction](/blog/browser-rendered-docs-extraction-problem)

---

<!-- https://agentcache.run/blog/agentcache-dot-folder -->

# keeping reference docs in .agentcache, beside your code

> See how a local documentation bundle can live beside your code in a predictable .agentcache folder, ready for an AI coding agent to read.

Agent Cache's planned local store puts documentation in `.agentcache/docs/<slug>/`. The goal is to keep a library's reference files beside the code, where a developer can point Claude Code, Cursor, or another coding agent at them.

If your goal is to turn documentation into files for Claude Code, download the ZIP, unpack the Markdown files in your project, and point the agent to the relevant folder. This also gets docs into Cursor or another file-reading coding agent. Adding API docs to a coding agent's project context can be as simple as that. Keeping API docs in your repo makes the reference easy to find; the `.agentcache` workflow is still being validated.

it's a dot-folder. hidden. gitignore-ready. and it's deliberate.

## why a dot-folder

dot-folders (starting with `.`) are hidden by default in most systems:
- `ls` doesn't show them
- file explorers hide them
- they're out of the way

this is good. docs are reference material, not source code. they shouldn't clutter your view.

## one folder per library, with content and metadata

```
.agentcache/
└── docs/
    ├── hono/
    │   ├── docs/
    │   │   ├── getting-started.md
    │   │   ├── routing.md
    │   │   └── middleware.md
    │   ├── meta.yaml
    │   ├── _map.json
    │   └── INDEX.md
    ├── stripe/
    │   ├── docs/
    │   └── ...
    └── supabase/
        ├── docs/
        └── ...
```

each project gets its own folder. each folder has:
- `docs/`: the actual markdown files
- `meta.yaml`: metadata for agent discovery
- `_map.json`: navigation tree
- `INDEX.md`: table of contents

## why not `docs/` directly?

why not put it at `docs/hono/` instead of `.agentcache/docs/hono/`?

because `docs/` is a common directory name. many projects use it for their own documentation. putting extracted docs there would clash.

`.agentcache/` is namespaced. unique. it won't conflict.

## gitignore-ready by default

the .agentcache folder should be in your `.gitignore`:

```
# agent cache documentation
.agentcache/
```

this means:
- docs aren't committed to git (might be large)
- each developer can have different docs sets
- ci can pre-populate docs if needed
- `.agentcache/` can be regenerated anytime

## multiple docs sets in one repo

a typical project uses multiple libraries:

```
.agentcache/
├── hono/          # backend framework
├── stripe/        # payments
└── prisma/        # database
```

each extraction is independent. you can add, remove, or update individual docs sets without affecting others.

## versioning and updates

what happens when hono releases version 4?

options:
1. **re-extract:** run agent cache again with the same url. existing docs are replaced.
2. **version-specific:** extract to `.agentcache/docs/hono-v4/` alongside `.agentcache/docs/hono-v3/`
3. **partial update:** i don't support this yet. full re-extraction only.

in practice: re-extract when you need a new version. most stable apis don't change often enough to need frequent updates.

## a predictable path for the planned mcp server

the future mcp server for agent cache will:
1. scan `.agentcache/` for `meta.yaml` files
2. build an index of available docs
3. surface relevant docs to agents on request

this is why the structure matters. the mcp server needs a predictable way to find docs. `.agentcache/docs/*/meta.yaml` is that predictable path.

## future: the cli managing .agentcache/

the planned cli commands:

```bash
agentcache add https://docs.stripe.com    # extract to .agentcache/docs/stripe/
agentcache list                             # show all extracted docs
agentcache view hono                        # show hono docs structure
agentcache remove hono                      # delete .agentcache/docs/hono/
agentcache update hono                      # re-extract hono docs
```

the folder structure enables these commands. each command maps to a filesystem operation.

## alternatives considered

**`node_modules/.agentcache/`**: too hidden. not accessible. tied to npm.

**`~/.agentcache/<project>/`**: global storage instead of per-project. breaks the "docs live with code" principle.

**inline in `package.json` metadata**: too limited. not suitable for full docs.

## reference docs stay nearby without cluttering the repo

`.agentcache/` is a simple, predictable, hidden location for documentation.

it's gitignore-ready. it supports multiple docs sets. it enables future tooling like cli and mcp.

when you extract docs, this is where they go. out of sight, but always available.

---

**related:**
- [meta.yaml: what makes docs discoverable](/blog/meta-yaml-agent-discovery)
- [why local docs matter](/blog/why-local-docs)

---

<!-- https://agentcache.run/blog/dual-layer-error-system -->

# building a dual-layer error system: machine errors vs human errors

> agent cache uses two error layers: machine errors for logs and human errors for users. here's why and how it makes debugging easier.

most apps have one error message. that message is either too technical for users or too vague for developers. it fails both audiences.

agent cache has two layers. every error gets split:
- **machine error:** technical, exhaustive, for logs
- **human error:** plain, actionable, for users

## vague for developers, cryptic for users

before dual-layer errors, my approach was like everyone else's:

```
error: "failed to extract documentation"
```

this tells the user nothing. what failed? why? what should they do?

or:

```
error: "502 bad gateway on cloudflare-proxied endpoint, ray id: ..."
```

great for debugging. useless for a user reading the status page.

## layer 1: machine errors

machine errors are for logs, monitoring, and debugging. verbose. technical. every detail.

example:
```json
{
  "error_machine": {
    "type": "bot_protection_challenge",
    "http_status": 403,
    "cloudflare_ray_id": "...",
    "cf_chl_jschl_tk": "...",
    "user_agent": "...",
    "request_url": "https://docs.example.com/api/reference",
    "timestamp": "2026-09-09T12:34:56Z",
    "attempt_count": 3,
    "backoff_strategy": "exponential",
    "last_headers": { ... }
  }
}
```

this goes to:
- stdout logs
- event ledger (`events.jsonl`)
- error tracking (if i had it)

it's for developers and operators. not for end users.

## layer 2: human errors

human errors are for users. plain language. actionable. sometimes funny.

example:
```
cloudflare blocked me. tried 3 times with polite delays. still got challenged.
this happens when docs sites have aggressive bot protection.
suggestions:
- if you own this site, consider whitelisting my crawler
- if not, i can't extract this site right now
- try again later, some sites loosen restrictions during off-peak
```

this goes to:
- the status page
- sse stream (what the user sees)
- logs (as a summary)

## what i never say

- "something went wrong", meaningless
- "an unexpected error occurred", expected by whom?
- "please try again later", why? what changed?
- "contact support", no. tell them what to do.

## what makes a good human error

1. **say what happened.** "cloudflare issued a javascript challenge" > "an error occurred"
2. **say why it matters.** "i can't extract sites behind captchas" > "extraction failed"
3. **give actionable next steps.** "try again in 1 hour" or "this site uses bot protection, so extraction isn't possible"
4. **be honest.** don't blame the user. don't pretend everything is fine.

## bad vs good examples

**bad:** "extraction failed due to network error"
**good:** "docs.example.com returned a 403 with cloudflare challenge page. this site uses bot protection that blocks automated access. i can't extract it right now."

**bad:** "rate limit exceeded"
**good:** "hit the rate limit for this site after 12 requests. some docs providers throttle crawlers. waiting 60 seconds before retry."

**bad:** "failed to parse html"
**good:** "encountered unexpected html structure. this might be a single-page app (spa) or custom framework. tried framework detection but no known pattern matched."

## detailed logs and fewer support questions

dual-layer errors make debugging trivial. when something breaks, i open the event log. every machine error is there. every detail.

for users, they see a clear explanation. no "something went wrong." they know exactly what happened and why.

support burden: minimal. most users understand from the human error message.

## implementation

errors are generated at the point of failure. the same exception produces both layers:

```
try {
  crawlPage(url)
} catch (e) {
  machine_error = e.to_machine_format()
  human_error = e.to_human_format()
  write_to_log(machine_error)
  stream_to_user(human_error)
}
```

each error type knows how to format itself for both audiences.

## log the details, show users what they can do

one error message can't serve both audiences. don't try.

split the error. machine gets everything. humans get clarity.

both win.

---

**related:**
- [agent cache architecture](/blog/architecture-deep-dive)
- [100 sites extracted: what broke](/blog/100-docs-sites-what-broke)

---

<!-- https://agentcache.run/blog/economics-docs-extraction -->

# why some docs sites cost 10x more to extract

> extracting documentation costs anywhere from free (direct markdown) to $$$ (paid services). here's the full cost breakdown across 100 extracted sites.

not all documentation extraction costs the same. some sites expose raw markdown for free. others require headless browsers and paid proxies.

the cost difference is 10x from cheapest to most expensive. here's the full economics.

## cost and time by extraction method

| tier | method | cost | time | success rate |
|---|---|---|---|---|
| tier 1 | llms.txt | $0 | <1s | 8% |
| tier 2 | github tree | $0 | 20-60s | 30% |
| tier 3 | direct .md | $0 | 1-2s/page | 30% |
| tier 4 | content negotiation | $0 | 1s/page | 3% |
| tier 5 | html purification | cpu | 5-10s/page | 38% |
| tier 6 | paid extraction | $$$ | 5-10s/page | remaining |

## tier 1: llms.txt, free, instant

when a site has `/llms.txt` or `/llms-full.txt`, extraction costs nothing.

one http request. one text file. done.

llms-full.txt includes the full docs content. `llms.txt` includes a table of contents.

adoption is growing but still low (~8%). cost: $0.

## tier 2: github tree, free, fast

if docs are open-source on github:
- one free api call for the git tree
- raw cdn downloads (free, no rate limits for public repos)

cost: $0.
time: 20-60 seconds for typical sites.

## tier 3: direct .md, free, fast

mintlify, fumadocs, gitbook, and others expose `.md` endpoints.

cost: $0.
time: 1-2 seconds per page, concurrent.

## tier 4: content negotiation, free, rare

sites that serve markdown via `Accept: text/markdown`.

cost: $0.
rarity: ~3%. but worth probing.

## tier 5: html purification, cpu cost

when free methods don't work, i fall back to html extraction.

stack: jsdom + turndown + custom cleaners.

cost: cpu cycles. bandwidth. no api fees.
time: 5-10 seconds per page (slower than direct .md).

price estimate: on a $6/month vps with shared cpu, 100 pages costs maybe $0.02 in compute time. negligible.

## tier 6: paid extraction, $$$, last resort

firecrawl, context.dev, and similar services.

pricing:
- firecrawl: $83/month standard, 100,000 pages
- ~$0.001 per page

for my scale, i'd only use this for sites that actively block free extraction.

## free methods keep the average cost down

the ladder tries free methods first. 88% of sites succeed with free methods.

so the average cost per site is close to zero. only the 12% that need paid services cost money.

if i extracted 100 sites/month:
- 88 sites: $0
- 12 sites: maybe $0.50 in compute (html purification)
- total: ~$6/month

the only actual cost is the vps: $6/month. everything else is free.

## what happens at scale

if agent cache processes 10,000 extractions/month:

| cost | amount |
|---|---|
| vps (bigger) | $20/month |
| bandwidth | negligible (most extraction is from docs sites to me) |
| r2 storage | ~$10/month for 500 GB |
| total | ~$30/month |

at that scale, the economics are incredibly favorable. docs extraction doesn't cost much when you use the free paths.

## free vs paid

the "expensive" extraction tools (firecrawl, context7) charge because they do more than just extract. they:
- index for search
- provide apis for runtime retrieval
- handle edge cases (javascript rendering, login flows)
- maintain infrastructure

agent cache doesn't do those things. it extracts once. stores locally. gives you a zip.

that's why it can be free. the scope is smaller.

## look for markdown before budgeting for scraping

61% of extraction is completely free.
38% costs negligible cpu.
1% might need paid services.

the secret to cheap docs extraction: try the free paths first. most sites already expose their markdown.

docs extraction is a discovery problem, not a cost problem.

---

**related:**
- [extraction methods in cost order](/blog/acquisition-ladder)
- [100 sites extracted](/blog/100-docs-sites-what-broke)

---

<!-- https://agentcache.run/blog/docs-framework-agent-readability -->

# mintlify vs docusaurus vs gitbook: which docs framework is best for agents?

> i extracted docs from 100 sites across 8 frameworks. here's my ranking of documentation framework agent-readability, from best to worst.

not all docs frameworks are equal. some expose raw markdown. some hide everything behind javascript. some have clean urls. some require headless browsers.

i extracted documentation from 100 sites across 8 frameworks. here's my ranking of documentation frameworks by how well they work for agent consumption.

## what i check before scoring a framework

i score frameworks on:
- **raw markdown availability:** can you get markdown without html scraping?
- **url structure:** clean, predictable urls?
- **navigation exposure:** can you extract the page tree programmatically?
- **content cleanliness:** minimal noise in extracted output?
- **framework velocity:** actively maintained? agent-first features?

## #1: mintlify, best for agents

mintlify is the clear winner. they built agent-first features.

**raw markdown:** append `.md` to any url. get clean markdown instantly. preserves frontmatter. preserves structure.

**llms.txt:** native support. `/llms.txt` and `/llms-full.txt` available on request.

**navigation:** exposed via json api. sidebar structure is fetchable.

**content:** minimal noise. clean html structure. easy to extract.

**velocity:** fastest-moving framework. new features every month. actively agent-aware.

examples: hono, better-auth, context7, resend.

score: 95/100

## #2: fumadocs, close second

fumadocs is mintlify's closest competitor for agent-friendliness.

**raw markdown:** supported. url patterns vary slightly between versions, but `.md` endpoints exist.

**navigation:** versioned, grouped. structure is in the dom but also exposed via api.

**content:** clean. prose components are well-structured.

**velocity:** active development. growing fast.

the main difference from mintlify: slightly less polished agent features. the `.md` endpoints work but aren't documented as clearly.

examples: acme, various newer tools.

score: 88/100

## #3: gitbook, solid, structured

gitbook is a mature platform with decent extraction support.

**raw markdown:** partial. some gitbook sites expose `.md` endpoints. others don't. depends on configuration.

**navigation:** sidebar structure is embedded in html. extractable with jsdom.

**content:** clean, predictable html structure. 

**velocity:** stable. slower feature development than mintlify/fumadocs.

**the catch:** gitbook's markdown output quality varies by site. some are perfect. others leave navigation elements in the content.

score: 75/100

## #4: docusaurus, good, but heavy

docusaurus is powerful. too powerful for simple extraction.

**raw markdown:** yes, but not via url. you need access to the repo to get raw markdown. the rendered site doesn't expose `.md` endpoints.

**navigation:** json-based sidebar files. extractable but requires knowing the framework's conventions.

**content:** heavy html. lots of framework-specific classes. custom mdx components that need special handling.

**velocity:** facebook-backed. stable. but not moving fast on agent features.

**the challenge:** docusaurus sites often use custom mdx components. these don't convert cleanly to markdown without framework-specific knowledge.

examples: react native, redux, jest.

score: 65/100

## #5: nextra, clean, minimal

nextra sites are minimal and clean. low noise.

**raw markdown:** yes, via `.md` endpoints. nextra is built on next.js and generally exposes raw files.

**navigation:** simple sidebar. easy to extract.

**content:** minimal noise. what you see is what you get.

**velocity:** active. vercel-backed.

score: 80/100

## #6: mdbook, simple, effective

mdbook is explicitly designed for rust projects. simple html output with clean structure.

**raw markdown:** yes. mdbook sites typically have the source markdown in the repo.

**navigation:** straightforward. hierarchical structure.

**content:** extremely clean. minimal html. converts to markdown trivially.

**limitation:** primarily for rust/rust-adjacent ecosystems.

score: 78/100

## #7: readthedocs / sphinx, legacy but workable

older docs framework. common in python world.

**raw markdown:** no. sphinx uses rst (restructured text), not markdown. conversion needed.

**navigation:** toc files. extractable but requires rst processing.

**content:** functional html. not pretty.

**velocity:** mature. stable. not evolving quickly.

score: 55/100

## #8: custom frameworks, wild card

every company that rolls its own docs framework. stripe. notion. linear.

**raw markdown:** almost never.

**navigation:** unique to each site. requires custom extractors.

**content:** varies wildly.

**the problem:** each one requires bespoke extraction logic. no standardization.

stripe's docs are the hardest i've extracted. 8.7 mb. custom components. dynamic rendering. but i got it working with framework-specific extractors.

score: 30-70/100 (varies massively)

## what makes a framework agent-friendly?

the pattern is clear. frameworks built since 2023 are agent-aware. they expose `.md` endpoints, `llms.txt`, and clean apis.

frameworks built before 2020 are human-only. they render html and assume a browser.

the trend is accelerating. llms.txt adoption is growing. direct markdown endpoints are becoming standard. the next generation of docs frameworks will be agent-first by default.

## mintlify for easy access, docusaurus if you need the power

if you're choosing a docs framework in 2026:

**choose mintlify or fumadocs** if you want agent-ready docs with minimal effort.

**choose docusaurus** if you need power and don't mind custom extraction.

**avoid custom frameworks** unless you have resources to build and maintain agent support.

the good news: even if you're on an older framework, agent cache can extract it. i just work harder for it.

---

**related:**
- [html extraction: how i clean docs](/blog/html-to-markdown-extraction)
- [finding raw docs before falling back to html](/blog/acquisition-ladder)
- [100 sites extracted](/blog/100-docs-sites-what-broke)

---

<!-- https://agentcache.run/blog/github-tree-fast-extraction -->

# github tree api: the fastest docs extraction you've never heard of

> if a docs site is open-source on github, you can extract the entire docs tree in 30 seconds using the git tree api. no scraping needed. here's how.

my fastest extraction ever: 3,848 files in 87 seconds. zero html parsing. zero jsdom. zero bot protection worries.

the secret? the github tree api.

## crawling html repeats work the repo already did

traditional crawler approach:
1. fetch the docs homepage
2. parse html to find links
3. follow each link
4. extract content from html
5. repeat for every page

for a site with 1,000 pages: 1,000+ http requests. 1,000 html parses. 1,000 content extractions. takes 5-10 minutes.

what if the docs are already structured markdown files in a git repo? what if you could just download them directly?

## list every docs file with one tree request

the github api exposes a "tree" endpoint that lists all files in a directory recursively.

```
GET /repos/:owner/:repo/git/trees/:branch?recursive=1
```

response: a flat list of every file and folder in the repo, with their paths, types, and blob hashes.

from this list, i filter for documentation files:
- `.md` files in a `docs/` directory
- `.mdx` files in the same
- `README.md` at root (usually includes getting started)
- any `.md` files matching known doc patterns

## github raw cdn: no api rate limits

once i know the file paths, i download them via the github raw cdn:

```
https://raw.githubusercontent.com/:owner/:repo/:branch/:path
```

this is the key advantage: **raw cdn downloads don't count against github api rate limits.** you have 60 requests/hour for unauthenticated api calls. but raw cdn has generous limits (thousands per hour).

for open-source repos, this means you can download an entire docs tree without authentication.

## find the repo, download files, rebuild the folders

1. **discover the repo:** check if the docs site links to a github repo in footer, header, or about page
2. **find the docs directory:** common names: `docs/`, `website/docs/`, `packages/docs/`, `src/content/docs/`
3. **get the tree:** one api call returns all files recursively
4. **filter docs files:** `.md` and `.mdx` in the docs directory
5. **download via raw cdn:** parallel requests, 8-12 workers
6. **reconstruct hierarchy:** use file paths from the tree to build the same folder structure

## real results

| site | files | time | strategy |
|---|---|---|---|
| nango | 3,848 | 87s | github tree |
| posthog | 1,915 | 40s | github tree |
| clerk | 1,124 | 28s | github tree |
| zeabur | 1,066 | 21s | github tree |

compare to html extraction for the same sizes: 3-10 minutes. github tree: under 2 minutes.

## limitations

**only works for open-source projects.** if the repo is private, you need authentication.

**docs structure may differ from site.** some projects keep docs in `/docs/` but the website is in `/website/`. some use a separate `docs` repo. some keep docs with the main code.

**not all files are user-facing docs.** `contributing.md`, `license.md`, `changelog.md`. these might be in the repo but aren't part of the user docs. i filter them out.

**some docs are in a separate repo.** the main code is in `company/product` but docs are in `company/docs`. my resolver tries to detect this.

## rate limits

github api: 60 requests/hour unauthenticated, 5,000/hour authenticated.
github raw cdn: thousands per hour. in practice, i rarely hit limits.

for large repos with 5,000+ files, i might hit the api limit. but most docs sites are under 1,000 pages.

## why this is tier 2, not tier 1

tier 1 is `llms.txt`. when it exists, it's a single request for the entire docs structure and content (if `llms-full.txt`).

github tree is tier 2 because it requires multiple requests: one for the tree, then one per file for the download.

but for repos with good `llms-full.txt` support, tier 1 still wins. github tree is the best fallback when `llms.txt` doesn't exist.

## check the public repo before crawling its docs site

github tree + raw cdn is the fastest, cheapest, most reliable docs extraction method. it doesn't work for proprietary docs, but for open-source projects, it's unbeatable.

if you're building a docs extraction pipeline and not checking github first, you're leaving massive speed gains on the table.

---

**related:**
- [how github fits into my extraction order](/blog/acquisition-ladder)
- [100 sites extracted: the full data](/blog/100-docs-sites-what-broke)
- [from html to markdown: the hard path](/blog/html-to-markdown-extraction)

---

<!-- https://agentcache.run/blog/head-requests-fail-edge-routes -->

# why head requests fail on netlify edge and cloudflare workers

> head requests return 502 bad gateway on netlify edge and cloudflare workers. here's why dynamic edge functions don't handle head and what i use instead.

when probing for raw markdown endpoints, the obvious first step is a head request. check if a resource exists without downloading it. saves bandwidth. faster.

but head requests fail on netlify edge and cloudflare workers. 502 bad gateway. every time.

here's why, and what i do instead.

## probing with head should save bandwidth

traditional approach:
1. send `HEAD /docs/page.md` 
2. check response code
3. if 200, send `GET /docs/page.md` to download

saves downloading content that doesn't exist. if `/docs/page.md` returns 404, you find out instantly without the body.

## head returns 502 where get works

on netlify edge and cloudflare workers, `HEAD` requests to dynamic routes return 502.

not 404. not 405 method not allowed. 502 bad gateway. a server error. confusing. misleading.

the same url with `GET` works perfectly. url exists. content exists. only `HEAD` fails.

## why head requests fail on edge platforms

edge functions (netlify edge, cloudflare workers) are lightweight v8 isolates. they run at the edge. they're fast. but they're not full node.js.

many edge function runtimes don't implement head handlers for dynamic routes. why? because:
- head is rarely used
- dynamic routes (next.js app router, sveltekit, etc.) generate responses on the fly
- the edge function generates the response, but the head handler path isn't implemented
- instead of a proper error, the runtime throws a 502

it's not a bug in your code. it's a gap in the platform.

## my solution: get with content-type verification

instead of:
```
HEAD /docs/page.md
```

i send:
```
GET /docs/page.md
Range: bytes=0-0
```

or simply a lightweight get without range, and check the `content-type` header immediately.

if the response is `text/markdown` or `text/x-markdown`, it's a valid markdown endpoint. i abort or continue based on headers.

if it's `text/html`, it's not a markdown endpoint. try something else.

**performance impact:** negligible. http headers are tiny. i'm downloading 0-1 KB to verify vs. the full page. for a 50 KB page, that's a 98% bandwidth savings over a full get.

## when head still works

head requests work fine on:
- static file servers (nginx, apache, caddy)
- traditional node.js servers (express, fastify, hono)
- cdn edge caches for static files
- object storage (s3, r2, gcs)

they only break on:
- dynamic edge functions (next.js on vercel, sveltekit on netlify)
- serverless functions with dynamic routing
- platforms where the runtime generates the response

## test each http method on your edge runtime

edge computing is great for latency. but it's a different runtime environment. not all standard http methods behave the same.

when building for edge platforms, test all http methods. don't assume head, options, patch, or delete work the same as get and post.

head is supposed to be safe and idempotent. but on edge, it's just unreliable for dynamic routes.

## if you're building a crawler

my recommendation: skip head for dynamic endpoints. use a lightweight get with early abort or range headers.

for static assets (images, css, js), head still works fine. those are served by cdns, not edge functions.

but for docs pages, which are almost always dynamic, head is a trap.

## use get and check the content type early

edge platforms are the future. but they have sharp edges. `HEAD` on dynamic routes is one of them.

use lightweight `GET` + content-type verification instead. it's more reliable, nearly as fast, and avoids the 502 trap.

---

**related:**
- [how i probe for cheaper extraction paths](/blog/acquisition-ladder)
- [cloudflare workers killed my serverless dream](/blog/cloudflare-workers-50-subrequest-limit)

---

<!-- https://agentcache.run/blog/html-to-markdown-extraction -->

# from html to markdown: building a documentation html cleaner

> How documentation HTML can be cleaned and converted into Markdown files, preserving useful content while removing navigation and page clutter.

most modern documentation frameworks expose raw markdown. but 38% of the sites i extract don't. custom frameworks. older docusaurus. proprietary cms systems.

When a source only serves HTML, the crawler has to convert the documentation site to Markdown before the pages can be packaged. The aim is to export docs as readable Markdown files while preserving useful material, including code examples and tables.

for those sites, i fall back to html extraction. fetch the html. parse the dom. strip noise. convert to markdown. it sounds simple. it's not.

html extraction is the hardest part of the pipeline. a generic html-to-markdown converter isn't enough. docs sites are full of nav bars, cookie banners, "was this helpful" buttons, sidebar trees, and cta boxes. if you convert all that to markdown, your docs are polluted.

here's how agent cache handles it.

## jsdom parses, cleaners strip noise, turndown converts

**jsdom** parses html in a node-like environment. gives me a real dom tree to manipulate. i can query selectors, traverse nodes, and extract specific elements.

**turndown** converts html to markdown. handles headings, lists, links, code blocks, tables. gives solid baseline output.

**custom cleaners** strip framework-specific noise. each docs framework has its own class names and structure. i don't guess. i know.

## step 1: fetch and parse

curl gets the raw html. headers include a reasonable user-agent. polite rate limits. no aggressive concurrency.

jsdom parses the html: `new JSDOM(htmlString).window.document`.

now i have a dom. time to find the content.

## step 2: identify and extract content area

the hardest problem: where's the actual content?

every framework puts it somewhere different:

| framework | content selector |
|---|---|
| mintlify | `article[data-kind]` or `main` |
| docusaurus | `.theme-doc-main` |
| fumadocs | `article.prose` |
| nextra | `main` |
| gitbook | `.book-body` |
| generic fallback | `main`, `article`, or largest text block |

framework detection happens first. i look for meta tags (`generator`), css class patterns, and known html structures. once i identify the framework, i know which selector to use.

if framework detection fails, i fall back to a heuristic: find the element with the most text content that's not in a nav or footer.

## step 3: strip the noise

within the content area, there's still junk.

**navigation breadcrumbs:** mintlify puts breadcrumbs in `nav[data-mintlify]`. docusaurus uses `nav[itemtype="https://schema.org/BreadcrumbList"]`. these get removed.

**"edit this page" links:** common on open-source docs. useless for agents. removed.

**"was this helpful" widgets:** feedback forms at the bottom of pages. removed.

**"next/previous" navigation:** links to adjacent pages. removed.

**cookie banners:** injected by third-party scripts. usually outside the content area, but occasionally slip in.

**admonition boxes:** tip, warning, info boxes. framework-specific. these are actually useful content. i keep them but standardize their format.

**code block buttons:** "copy" buttons on code blocks. removed. the code itself is kept.

## step 4: handle tabs and version selectors

many docs sites use tabs ("javascript" vs "python" vs "go") and version dropdowns ("v1" vs "v2").

the challenge: active tab content is in the dom. inactive tabs might be hidden with `display: none` (i skip those). but some frameworks render all tab content and hide with css. jsdom respects css, so hidden content is naturally excluded.

for versions: i extract the default (usually latest) version. version selectors are removed from the output.

## step 5: convert to markdown

turndown runs on the cleaned html. converts:
- headings (`<h1>` → `# heading`)
- lists (`<ul>`/`<ol>` → `- item` / `1. item`)
- code blocks (`<pre><code>` → fenced code blocks with language)
- links (kept, usually relative to docs root)
- tables (pipe tables)
- inline code (backticks)
- emphasis and bold

## step 6: post-process

after turndown, i clean up common issues:

**empty lines:** turndown sometimes leaves multiple blank lines. compressed to max 2.

**heading levels:** normalize so the page's first heading becomes h1, regardless of what the html had.

**link fixing:** relative links are adjusted based on the page's url. absolute links are kept.

**table formatting:** ensure pipe tables align properly for readability.

**code fences:** ensure proper language tags (from class names like `language-typescript` → `ts`)

## framework-specific extractors

my best results come from framework-specific logic.

**mintlify extractors** know that content is in `article[data-kind="document"]` and sidebar nav is in `aside[data-testid="sidebar"]`. i strip everything outside the article.

**docusaurus extractor** handles their specific html: `.theme-doc-markdown` for content, `.theme-admonition` for info boxes, `tabs-container` for tabs.

**fumadocs extractor** knows their `article.prose` and `tabs` components.

without these framework-specific extractors, i'd need a generic approach. and generic approaches produce generic (worse) results.

## quality metrics

i measure extraction quality with:
- **content ratio:** extracted text vs. total html text. >80% is good.
- **nav removal:** no nav links in output. verified manually.
- **code block count:** matches visual inspection.
- **heading structure:** h1 first, then h2, h3. no skips.

for popular frameworks, quality is 95%+. for custom frameworks, 70-80% is typical. for spa sites, it's unpredictable.

## making the html fallback as clean as raw markdown

html extraction is the fallback. i try everything else first. but for the 38% of sites that need it, quality matters. bad extraction means noisy docs. noisy docs mean confused agents.

the goal is simple: make the fallback so good that users can't tell whether extraction used direct .md or html purification.

i'm not there yet. but i'm close.

---

**related:**
- [what i try before parsing html](/blog/acquisition-ladder)
- [100 sites extracted: what broke](/blog/100-docs-sites-what-broke)
- [mintlify vs docusaurus: framework rankings](/blog/docs-framework-agent-readability)

---

<!-- https://agentcache.run/blog/30-lessons-docs-extraction -->

# 30 lessons learned from building a documentation extraction tool

> after extracting 100+ documentation sites and building agent cache, here are 30 lessons about docs, agents, and extraction.

i've extracted 100+ docs sites. hit cloudflare walls. debugged .md.md bugs. built the entire pipeline. here are 30 lessons.

## lessons 1-10: about documentation

1. **most docs sites expose raw markdown.** 61% of modern docs have `.md` endpoints, `llms.txt`, or github repos. scraping is the fallback.

2. **docs sites redesign frequently.** html selectors break. framework extractors need updates. extraction is maintenance, not a one-time thing.

3. **version selectors are the hardest ui element to handle.** tabs, dropdowns, sidebar toggles. all invisible to simple extractors.

4. **code blocks are fragile.** some frameworks split them across divs. some use custom syntax highlighting. some embed code in data attributes.

5. **navigation is the most useful structure.** the page tree tells agents how docs relate. but it's often hidden in javascript or json apis.

6. **mobile and desktop versions differ.** some sites serve different content. responsive design hides content with `display: none`. extraction needs desktop user-agents.

7. **search functionality is never extractable.** search needs a backend. the frontend is just a query box. i ignore search.

8. **images and diagrams are usually outside the scope.** i focus on text and code. visual content is hard to convert meaningfully.

9. **the "pretty" docs site is often a skin.** the real content is markdown or json underneath. the rendered html is just one view.

10. **docs quality varies massively.** some are clean, structured, well-maintained. others are outdated, broken-link-ridden messes.

## lessons 11-20: about extraction

11. **framework detection is 80% of the work.** once you know the framework, you know the selectors. everything else follows.

12. **html purification never produces perfect output.** there's always some noise. the goal is "good enough for agents," not "pixel perfect."

13. **rate limiting is the biggest practical challenge.** not parsing. not conversion. just getting permission to fetch the pages.

14. **headless browsers would solve most extraction problems.** but they're 10x slower, 10x heavier, and unreliable at scale. i avoid them.

15. **github tree api is the fastest extraction method.** when it works, nothing else comes close. always check for open-source docs repos.

16. **custom frameworks are the most time-consuming.** every company that builds its own docs framework requires bespoke extraction logic.

17. **direct .md endpoints are the most reliable.** mintlify, fumadocs, and others serve clean markdown. no parsing needed.

18. **content negotiation is underrated.** `Accept: text/markdown` works on some sites. they just don't advertise it.

19. **crawling ethics matter.** respecting robots.txt and rate limits isn't just polite. it prevents your ip from being permanently blocked.

20. **the 80/20 rule applies.** 80% of sites extract cleanly with simple rules. 20% need complex handling. but that 20% takes 80% of the time.

## lessons 21-30: about building developer tools

21. **htmx + hono is a surprisingly capable stack.** no react. no build step. fast, simple, reliable.

22. **zero-memory architecture works.** stream everything to disk. replay events from a log. crash recovery is built in.

23. **dual-layer error messages are worth the effort.** clear user messages + verbose debug logs. both audiences served.

24. **open source is the right business model.** the extraction tool is free. the hosted service is convenient.

25. **vps beats serverless for crawling.** cloudflare workers' 50 subrequest limit made serverless impractical. a $6 vps handles everything.

26. **turso (sqlite) handles metadata fine.** no postgres needed for simple relational data. sqlite is enough.

27. **r2 saves money on downloads.** zero egress fees vs. s3's 9 cents/gb. for zip-file downloads, that matters.

28. **user experience matters more than features.** a clean form + status page + download button beats a dozen half-finished features.

29. **marketing is harder than engineering.** the product works. distribution doesn't. seo content is my strategy.

30. **docs extraction should be free.** charging for access to public documentation is wrong. open source extraction ensures it stays free.

## raw markdown was already there on 61% of sites

the number of sites that already expose raw markdown. 61%. the frameworks are designed for it. developers built it in. i just need to know where to look.

## finding the right extraction path took the most work

not extraction. not parsing. not html cleaning.

it's **discovering the right extraction path.** given a random docs url, which tier of the ladder works? that's 80% of the complexity.

## what i'd do differently

1. **spend more time on framework detection.** early versions had generic extraction. adding framework-specific extractors improved quality 10x.

2. **start with the cdn/github paths.** i initially built html extraction first. reversing the order would have been faster.

3. **build meta.yaml early.** metadata is what makes docs useful to agents. without it, it's just a folder of markdown files.

## public docs still aren't always fetchable

building a docs extraction tool taught me more about how documentation works than i expected.

the diversity of docs frameworks. the creativity of web development. the gap between "publicly readable" and "programmatically accessible."

but also: most docs authors want their content to be accessible. they just need the right tools.

agent cache is one of those tools.

---

**related:**
- [100 sites extracted: what broke](/blog/100-docs-sites-what-broke)
- [choosing an extraction path, cheapest first](/blog/acquisition-ladder)
- [agent cache architecture](/blog/architecture-deep-dive)

---

<!-- https://agentcache.run/blog/meta-yaml-agent-discovery -->

# meta.yaml: what makes documentation discoverable by agents

> every agent cache bundle includes a meta.yaml with structured metadata. here's what goes in it and why agents need it.

when extraction finishes, agent cache creates a `meta.yaml` file. not for humans. for agents.

this file tells any agent (or system) what it just extracted. what the docs are for. when to use them. what they're related to.

without metadata, a folder of markdown files is just a folder of files. with metadata, it's a discoverable resource.

## why metadata matters

imagine you're a coding agent. someone drops `.agentcache/docs/` into the project. 50 markdown files. no labels. no descriptions.

what is this? react docs? stripe docs? your company's internal api docs?

a human can open INDEX.md and figure it out. an agent can't. it needs structured metadata.

## a meta.yaml for hono

```yaml
name: hono
title: Hono - Lightweight Web Framework for the Edge
url: https://hono.dev
description: Hono is a fast, lightweight, and portable web framework for JavaScript and TypeScript. It works with any runtime.
repository: https://github.com/honojs/hono
keywords:
  - hono
  - web framework
  - edge computing
  - typescript
  - deno
  - cloudflare workers
  - bun
  - honojs
  - middleware
  - routing
intent_triggers:
  - how do i create a basic hono app
  - hono middleware pattern
  - hono routing examples
  - hono vs express
  - deploying hono on cloudflare workers
  - hono typescript setup
ecosystem:
  - cloudflare workers
  - vercel
  - deno
  - fastly
  - lambda
  - openrouter
  - zod
  - valibot
version: latest
extracted_at: 2026-09-01T00:00:00Z
framework: mintlify
```

## field breakdown

**name:** short identifier. matches the folder name. `hono`, `stripe`, `supabase`.

**title:** human-readable name. used for display.

**url:** original docs url. so agents and users know where it came from.

**description:** one paragraph about what this library does. helps agents understand context.

**repository:** github url. for linking to source code, issues, discussions.

**keywords:** 8-12 technical terms. these are search tokens. when an agent searches for "hono routing," the keyword match surfaces this docs bundle.

**intent_triggers:** natural language questions that would lead someone to these docs. this is how agents know "i should use hono docs for this request."

**ecosystem:** related tools and platforms. helps agents understand relationships. "this is a cloudflare workers framework."

**version:** which version was extracted. `latest`, `v3`, `v4.2`.

**extracted_at:** timestamp. lets agents know how old the extraction is.

**framework:** which docs framework was used. helpful for understanding extraction quality.

## how agents use meta.yaml

agents scan `meta.yaml` to determine which docs are relevant to a query.

user asks: "how do i set up middleware in hono?"
agent steps:
1. scan `.agentcache/` for meta.yaml files
2. match query to keywords (`hono`, `middleware`) and intent triggers
3. load the hono docs bundle
4. reference the middleware section

without meta.yaml, the agent would need to open and read every docs bundle to find the right one. that's slow.

with meta.yaml, matching is instant. one file per bundle.

## examples: good vs bad keywords

**bad:** `documentation`, `guide`, `tutorial`, `reference`

why: these are generic. every docs site has them. they don't help matching.

**good:** `hono`, `web framework`, `cloudflare workers`, `typescript`, `middleware`, `routing`, `bun`

why: these are specific. they uniquely identify the library and its domain.

## generated vs curated

the meta.yaml fields are generated by an llm. why use an llm here? because keywords and intent triggers require semantic understanding. they're subjective.

but the llm output is not blindly trusted. it's reviewed. corrected. refined.

in the future, i'll add a validation step: check that keywords match the actual doc content. check that intent triggers are realistic questions.

## future: meta.yaml as a standard

the meta.yaml format could become a standard. any docs extraction tool could generate it. any agent could consume it.

if multiple tools use the same schema, agents get portability. switch from one extractor to another, and the metadata stays compatible.

this is one reason i'm documenting the schema publicly.

## ship accurate metadata with every bundle

metadata is what turns a pile of files into a searchable, discoverable, usable resource.

without meta.yaml, your docs bundle is just content. with meta.yaml, it's a tool that agents can reason about.

always generate metadata. always keep it accurate.

---

**related:**
- [organizing docs inside .agentcache](/blog/agentcache-dot-folder)
- [why local docs matter](/blog/why-local-docs)

---

<!-- https://agentcache.run/blog/mintlify-vs-docusaurus-vs-fumadocs -->

# mintlify vs docusaurus vs fumadocs: which docs framework should you choose?

> the three most popular docs frameworks compared. which one is fastest, most agent-friendly, and right for your project?

if you're building documentation in 2026, you have three dominant choices: mintlify, docusaurus, and fumadocs.

they're not interchangeable. they're built for different philosophies. different tradeoffs. different outcomes.

here's how they compare on what actually matters.

## hosting, setup, and agent access side by side

| | mintlify | docusaurus | fumadocs |
|---|---|---|---|
| **type** | saas / hosted | open source self-hosted | open source self-hosted |
| **setup** | connect github, done | install npm packages, configure | install npm package, configure |
| **hosting** | mintlify (free tier) | vercel, netlify, github pages | vercel, netlify, anywhere |
| **agent-friendly** | excellent | moderate | excellent |
| **speed** | fast (cdn) | fast (static) | fast (ssr via next.js) |
| **ecosystem** | growing | massive (meta-backed) | growing |
| **customization** | limited (themes) | extensive | moderate |
| **pricing** | free tier, then paid | free | free |
| **best for** | startups, teams | large projects, custom needs | next.js projects, speed |

## mintlify: the saas experience

mintlify is a platform, not just a framework. you connect your github repo. they handle hosting, search, analytics, and now ai features.

### what mintlify does well

**zero config.** connect your repo with markdown files. done. the site builds automatically. no wrangling with build pipelines or deployment.

**ai features built in.** semantic search. ai-powered question answering. "chat with your docs" out of the box. this matters because users increasingly expect search to understand intent, not just match keywords.

**agent-ready by default.** mintlify generates `llms.txt` automatically. exposes `.md` endpoints for every page. this is not an afterthought. it's built in.

**fast.** cdn-backed. global edge distribution. your docs load fast everywhere.

**beautiful default design.** mintlify sites look good without customization. typography, spacing, dark mode, mobile responsiveness. all polished.

### where mintlify falls short

**vendor lock-in.** your docs are in your repo (markdown files), but the hosting, search index, and ai features are on mintlify's infrastructure. moving away means rebuilding those.

**limited customization.** you can choose themes and colors. you can't fundamentally change the layout or add custom components. if you need something mintlify doesn't support, you're stuck.

**pricing at scale.** free tier is generous. but at high traffic or many projects, costs add up. this is the saas model.

## docusaurus: the proven workhorse

docusaurus is meta's open-source docs framework. it's been around since 2017. it powers the docs for react, redux, jest, and hundreds of other projects.

### what docusaurus does well

**mature ecosystem.** plugins for search (algolia docsearch), internationalization, versioning, code tabs, mermaid diagrams, math equations. if you need a feature, there's probably a plugin.

**deep customization.** react under the hood. you can override any component. build custom themes. add your own webpack config. if you can build it in react, you can put it in docusaurus.

**versioning.** built-in support for multiple versions of docs. essential for libraries with breaking changes. mintlify handles this too, but docusaurus's implementation is more battle-tested.

**free forever.** open source. no vendor. host anywhere. no usage limits.

**meta backing.** not a startup that might shut down. docusaurus has institutional support.

### where docusaurus falls short

**setup friction.** npm install. create `docusaurus.config.js`. configure plugins. set up deployment. it's not hard, but it's more steps than mintlify.

**agent friction.** no automatic `llms.txt`. no `.md` endpoints. the source markdown is in your repo (on github), but visitors to the site get html. agents need to go to github or parse the html. extra steps.

**build times.** complex sites with many plugins take time to build. not a problem for most, but noticeable on large docs.

## fumadocs: the speed demon

fumadocs is the newest of the three. a next.js-based docs framework built by a single developer. it's fast, minimal, and modern.

### what fumadocs does well

**next.js native.** if your project is already next.js, fumadocs fits perfectly. same build pipeline. same deployment. same vercel integration.

**blazing speed.** ssr via next.js. instant page loads. the demo sites feel faster than docusaurus and mintlify.

**agent-ready.** like mintlify, fumadocs exposes `.md` endpoints. supports `llms.txt`. built with programmatic access in mind.

**minimal complexity.** smaller codebase than docusaurus. fewer abstractions. easier to understand and customize if you know next.js.

**free and open source.** no platform fees. host on vercel's free tier. zero cost.

### where fumadocs falls short

**newer ecosystem.** fewer plugins. fewer tutorials. smaller community. if you hit an edge case, fewer people have solved it before.

**single maintainer.** built by one developer. if they stop working on it, the project might stall. (though it's open source, so anyone can fork.)

**less polished defaults.** mintlify's default design is beautiful. fumadocs is clean but minimal. you might need to customize more to match mintlify's visual quality.

## agent-friendliness: the hidden factor

this is where agent cache cares. how easy is it to extract your docs?

**mintlify:** excellent. `.md` endpoints work. `llms.txt` auto-generated. navigation api available. extraction is trivial.

**fumadocs:** excellent. `.md` endpoints work. `llms.txt` supported. next.js app router makes routes predictable. extraction is easy.

**docusaurus:** moderate. no `.md` endpoints. no auto `llms.txt`. but markdown source is on github. extraction requires going to the repo, not the site. still doable, just extra steps.

if ai agents reading your docs matters to you (and it should), mintlify and fumadocs have an advantage.

## performance comparison

| metric | mintlify | docusaurus | fumadocs |
|---|---|---|---|
| first contentful paint | ~0.8s | ~1.2s | ~0.6s |
| time to interactive | ~1.5s | ~2.0s | ~1.0s |
| js bundle size | ~45kb | ~120kb | ~15kb |
| build time (100 pages) | n/a (hosted) | ~30s | ~15s |

fumadocs wins on speed. mintlify wins on convenience (no build needed). docusaurus is reasonable but heavier.

## when to choose each

**choose mintlify if:**
- you want the fastest setup possible
- you value built-in ai features
- you don't mind hosted saas
- agent-friendliness matters
- you're a startup or small team

**choose docusaurus if:**
- you need deep customization
- you have complex versioning requirements
- you want maximum ecosystem support
- you're building docs for an established open-source project
- you don't mind self-hosting

**choose fumadocs if:**
- your project is already next.js
- you care about performance above all else
- you want agent-ready docs without saas
- you prefer minimal, modern tools
- you're comfortable with a newer ecosystem

## mintlify for zero ops, fumadocs for self-hosted next.js

for most new projects in 2026: **mintlify or fumadocs.**

both are fast, modern, and agent-ready. mintlify if you want zero-ops saas. fumadocs if you want self-hosted next.js.

docusaurus is still excellent for large, complex projects that need its ecosystem. but for a typical docs site, mintlify or fumadocs will get you there faster.

## judge the raw docs alongside the rendered site

documentation frameworks are not just about how the site looks to humans. in 2026, they're also about how accessible the content is to machines.

mintlify and fumadocs understand this. docusaurus is catching up. the future belongs to frameworks that serve both audiences out of the box.

---

**related:**
- [framework extractor rankings](/blog/docs-framework-agent-readability)
- [how agent-readable docs are in 2026](/blog/state-of-docs-for-agents-2026)

---

<!-- https://agentcache.run/blog/knowledge-base-docs-platforms-compared -->

# knowledge bases, ai docs, and platforms that do too much

> document360, intercom, gitbook, and ai-native docs tools compared. when a docs site becomes a support portal, extraction gets complicated.

documentation tools fall on a spectrum. on one end: pure docs. markdown files rendered to html. on the other end: multi-purpose platforms that combine docs with support, chat, ticketing, and ai.

this article is about the other end. the platforms that do more than docs. and why "more" often means "harder to extract."

## tier 1: knowledge bases that are mostly docs

these tools are knowledge bases first, but their output resembles documentation sites. extraction is possible but requires extra work.

**document360**: full-featured saas knowledge base. public product docs and private internal bases. widget for in-app search. ticket integrations (zendesk, intercom, freshdesk). multi-versioning. localization.

**extraction reality:** document360 exports to html and pdf. the public docs are accessible. but the structure is platform-specific. no `.md` endpoints. no `llms.txt`. i can extract the content, but i lose the clean markdown structure.

**help scout**: clean, zero-config kb. "beacon" widget for searching docs and initiating chat. the kb articles are straightforward. extraction: the public articles are just html pages. i parse them.

**crisp**: all-in-one messaging with kb site generator. multi-language translations. integrated live chat popup. the kb content is accessible but mixed with chat features. extraction targets just the kb pages.

**freshdesk**: comprehensive customer support portal. kb articles + self-service portals + community forums + freddy ai bot. the kb section extracts like a standard docs site. the forums and ai bot responses are outside my scope.

## tier 2: support-first, docs-second

these platforms lead with support and chat. documentation is a feature, not the product.

**zendesk guide**: enterprise standard. knowledge base + live chat + ticket routing + zendesk ai agents. the docs exist within a larger support ecosystem. extraction: possible for public articles. but the site structure is designed for ticket deflection, not reading.

**intercom**: help center + messenger + fin ai. fin reads knowledge articles to answer questions automatically. the help center is extractable. but intercom's real value is the conversational layer, not the static docs.

**hubspot service hub**: kb portals linked to crm. native live chat. ticket automation. multi-language. docs are a module inside a larger platform. extraction: public kb articles only.

**tidio**: help center + lyro ai. similar to intercom's model. the kb is the extractable part.

## tier 3: ai-native documentation platforms

new category. tools that use ai to generate, update, or serve documentation. most are vaporware or feature-thin.

**documentation.ai**: connects to codebases and prs. continuously writes and updates docs with autonomous agents. sounds impressive. in practice: works for api references. fails at conceptual docs, tutorials, and edge cases. the "autonomous" part is a stretch.

**docsalot**: unifies product manuals and dev docs. optimized for humans and agents. native mcp server. still early. promising but unproven.

**hyperdocs**: automatic change detection from git prs. ai content publishing. another "ai writes docs" tool. useful for reference generation. not ready for tutorials or architecture docs.

**docsio**: ingests repos or web content. outputs hosted sites with automatic `llms.txt` and mcp support. the extraction angle is interesting: it extracts docs for you, then hosts them. overlaps with what agent cache does.

**docuwriter.ai**: inspects source code across 20+ languages. generates api references and tutorials. works for codebases with good comments. fails for poorly documented code (which is most code).

## where these platforms make extraction harder

here's the issue with multi-purpose platforms: they weren't built for extraction.

**zendesk:** articles exist, but the url structure is opaque. `/hc/en-us/articles/360012345678` doesn't tell you what the article is about. no sitemap in standard format. no `llms.txt`.

**intercom:** help center urls are cleaner, but the content is wrapped in intercom's chrome. sidebars, widgets, chat buttons everywhere. the actual article is a small percentage of the html.

**document360:** public articles are accessible. but the navigation tree is generated dynamically. extracting the full structure requires crawling every page and building the tree from breadcrumbs.

**ai-native platforms:** most generate output but don't expose it in machine-readable formats. an agent that wants to read docsalot's generated docs might need to go through docsalot's api. which defeats the point.

## what i do at agent cache

for knowledge base platforms:
1. try to extract the public kb articles
2. parse the html to clean content
3. reconstruct navigation from breadcrumbs or sidebar links
4. accept that the output is less clean than a docusaurus or mintlify site

for ai-native platforms:
1. i don't extract them. i extract the sources they claim to ingest
2. if docsio says it reads your github repo, i go to the repo instead
3. the hosted output is ephemeral. the source is permanent

## when to choose a knowledge base over pure docs

**choose a kb platform if:**
- you need integrated ticketing and chat
- your docs and support are the same team
- you need in-app help widgets
- your audience is end users, not developers

**choose pure docs if:**
- your audience is developers
- you need api references and code examples
- agent-readiness matters
- you want to own your content

**the hybrid approach:** use a docs generator (mintlify, docusaurus) for technical docs. use a kb platform (help scout, document360) for user guides and support articles. they're different audiences. different tools.

## keep technical docs easy to read outside the platform

multi-purpose platforms add features that sound valuable. ticketing integration. ai chatbots. crm sync. but each feature is another extraction boundary. another format to handle. another api to call.

pure docs tools say: "here's the content." kb platforms say: "the content is inside this platform. use this api to read it."

for agents, the first model is infinitely better.

---

**related:**
- [pure docs generators ranked](/blog/pure-docs-generators-ranked)
- [ai-native documentation platforms: hype or future](/blog/ai-native-documentation-platforms)

---

<!-- https://agentcache.run/blog/md-md-extension-bug -->

# i requested page.md.md and wondered why it returned 404

> a simple bug caused me to request `page.md.md` instead of `page.md`. here's why url normalization is critical.

a bug that cost hours to debug. the symptom: `404` for urls that should've worked. the cause: double extensions. `page.md.md` instead of `page.md`.

## appending .md to a path that already had it

my acquisition ladder constructs `.md` urls by appending `.md` to the path:

```
base url: /docs/getting-started
.md url: /docs/getting-started.md
```

but some sites already had `.md` in their canonical urls:

```
base url: /docs/getting-started.md
.md url: /docs/getting-started.md.md
```

double extension. `404`. extraction fails.

## how it happened

naive implementation:
```javascript
const mdUrl = `${url}.md`
```

if `url` is `/docs/page`, this works: `/docs/page.md`.
if `url` is `/docs/page.md` (already has extension), this becomes `/docs/page.md.md`.

sites where this happened:
- sites that expose `.md` directly in their canonical paths
- some docusaurus configurations
- sites with custom routing that includes `.md`

## normalize the extension before adding .md

normalize before constructing:

```javascript
function normalizeExtensions(url) {
  // strip trailing extensions
  return url
    .replace(/\.mdx?$/, '')    // remove .md and .mdx
    .replace(/\.html?$/, '')   // remove .html and .htm
}

const mdUrl = `${normalizeExtensions(url)}.md`
```

simple. but critical.

## other extension issues

it's not just `.md.md`:

- `.html.html`
- `.mdx.md` (site uses `.mdx`, i append `.md`)
- `.php.md`

any site that includes extensions in canonical paths needs normalization.

## strip the existing extension before constructing a url

before constructing any url with a known extension, strip existing extensions first.

```javascript
// bad
const newUrl = `${url}.md`

// good
const baseUrl = stripExtensions(url)
const newUrl = `${baseUrl}.md`
```

## testing for edge cases

url normalization is full of edge cases:

- query strings: `/docs/page?ref=nav` → `/docs/page.md?ref=nav`
- anchors: `/docs/page#section` → `/docs/page.md#section`
- trailing slashes: `/docs/page/` → `/docs/page.md`

each of these needs careful handling.

## test url construction against real docs paths

url normalization sounds simple. it's not.

tiny bugs in url construction cause extraction failures that are hard to debug. always normalize. always test with real urls.

---

**related:**
- [how i probe docs endpoints](/blog/acquisition-ladder)
- [html extraction](/blog/html-to-markdown-extraction)

---

<!-- https://agentcache.run/blog/open-source-direction -->

# why agent cache is going open source (and what that means)

> i'm open sourcing the agent cache extraction engine under mit. here's what that means and why i did it.

agent cache is going open source.

the extraction engine. the cli. the mcp server. all of it.

under mit license. free to use. free to modify. free to fork.

## charging for public docs extraction feels wrong

charging for docs extraction feels wrong.

documentation is public information. it's meant to be read. wrapping it in a paywall feels like charging for library access.

open source aligns with that: the tool is free. anyone can extract docs.

## what's open source

these components will be open sourced:
- **extraction engine**: the core crawler, framework extractors, acquisition ladder
- **cli**: `agentcache add <url>`, `agentcache list`, etc.
- **mcp server**: serves local docs to agents via model context protocol
- **utility libraries**: url normalization, html cleaning, etc.

## what's not open source

these stay proprietary:
- **web app** (`agentcache.run`): the hosted service
- **infrastructure**: deployment, monitoring, billing
- **branding**: name, logo, domain

## why the split

the extraction technology is a public good. it should be open. anyone should be able to extract docs.

the hosted service is my business. convenience, reliability, and support.

this is a common model. linux is open source. red hat sells support. wordpress is open source. wordpress.com sells hosting.

## what mit means

mit license is the most permissive:
- use for commercial projects
- modify and redistribute
- include in proprietary software
- no attribution required (but appreciated)
- no warranty

basically: do whatever you want. just don't blame me if something breaks.

## what it means for users

**as a free user:** nothing changes. the web app stays free.

**as a self-hoster:** you can run the extraction engine locally. no dependency on my infrastructure. extract docs on your own server.

**as a developer:** you can modify the extraction logic. add support for new frameworks. integrate into your own tools.

**as a competitor:** you can fork it. build your own product. i can't stop you and don't want to. the extraction itself is not a moat.

## hosting is what i'd charge for

the hosted service. convenience. zero setup. reliability.

if you want to "just extract docs," the web app is faster than self-hosting.

if you need to extract docs at scale, in a pipeline, or inside your own infrastructure, the open source engine is for you.

## timeline

1. **now:** extraction engine code is being cleaned up for release
2. **soon:** public github repo with core extraction code
3. **later:** cli and mcp server released
4. **eventually:** web app stays as hosted service

## open the engine, sell the convenience

the technology to extract docs should be free and open. the service to make it convenient can be monetized.

i open source the tool. i sell the convenience.

---

**related:**
- [agent cache architecture](/blog/architecture-deep-dive)

---

<!-- https://agentcache.run/blog/pure-docs-generators-ranked -->

# pure documentation generators ranked for agent-readability

> not all docs generators are equal for ai agents. here's my ranking of pure documentation ssgs based on extraction experience from 100+ sites.

a documentation generator turns markdown into a website. that's it. no api testing, no knowledge base features, no ticketing integration. just docs.

these are the pure tools. and they're not all equal when it comes to ai agents reading them.

## tier 1: agents read these for free

**mintlify**

saas docs platform. auto-generates `llms.txt`. every page has a `.md` endpoint. my extraction pipeline reaches mintlify sites and immediately thinks "this is too easy."

the tradeoff is saas lock-in for hosting. the content stays in your repo. the rendering, search, and ai features are on mintlify's infrastructure.

**fumadocs**

next.js-based. open source. fast. exposes `.md` endpoints natively. if your project already uses next.js, adding fumadocs is trivial. extraction is trivial too.

the maintainer is one developer. that's a risk. but it's open source, so the risk is forkability, not abandonment.

**starlight**

astro's official docs framework. zero client-side javascript by default. the fastest docs site you can build. extraction is easy because the html is clean, semantic, and uncluttered.

if you care about performance and accessibility, starlight is hard to beat. if you need complex interactive components, you might outgrow it.

**vitepress**

evan you's docs framework for vue. clean, fast, minimal. markdown source is in your repo. the rendered site is static html. extraction is straightforward: go to github for the markdown, or parse the clean html.

## tier 2: good, but agents need a detour

**docusaurus**

the most popular docs framework. powers react docs, redux docs, jest docs. massive ecosystem. mature. stable. but: no `.md` endpoints. no auto `llms.txt`. markdown exists in your repo, but the site serves html. agents need to either go to github or parse the html.

this isn't a dealbreaker. docusaurus sites extract fine. but they're more work than tier 1.

**nextra**

vercel's docs framework for next.js. raw mdx files in your repo. no hosted service, just the framework. simpler than fumadocs but less feature-rich. extraction is easy because the source is just files.

**docus**

nuxt labs' docs framework. built on nuxt 3 and `@nuxt/content`. lets you embed vue components directly in markdown. clean output. fast ssr. but same problem as docusaurus: markdown source is in your repo, and the rendered site is html. no `.md` endpoints. no auto `llms.txt`.

if your project is already nuxt/vue, docus is natural. for extraction, it's moderate difficulty: parse the html or fetch from the repo.

**gitbook**

polished hosted docs. real-time editing, github sync. but: custom extraction requires more work. content is somewhat tied to their platform. good for teams that want minimal setup and don't care about extraction complexity.

## tier 3: solid, but niche

**mkdocs**

python's docs standard. configured via `mkdocs.yml`. material for mkdocs is one of the most widely used doc themes in software engineering. if you're python-first, it's natural. extraction: markdown source is in your repo.

**sphinx**

the original python documentation engine. powers python's own docs. restructuredtext instead of markdown. verbose but thorough. built for technical manuals, not quick readme sites. extraction: source files in repo.

**mdbook**

rust-powered. zero dependencies. outputs clean static html. the tool behind "the rust programming language" book. if you need a book, not a website, this is perfect. extraction: markdown source in repo.

**11ty**

javascript ssg with zero client-side js. maximum flexibility. not docs-specific, you configure it for docs. requires more setup than dedicated tools. extraction depends on how you structure it.

**hugo + docsy/hextra**

go-based ssg. builds thousands of pages in milliseconds. docsy (google-maintained) and hextra are doc themes. fast but go templating has a learning curve. extraction: markdown in repo.

## tier 4: agents struggle with these

**docsify**

renders markdown in the browser via javascript. no build step. sounds convenient. but the html contains zero content, it's all loaded by js. extraction requires headless browsers or reverse-engineering the markdown urls. i try to extract docsify sites. i mostly fail.

docsify was made for humans with browsers. not for agents.

**vuepress**

replaced by vitepress. still works. still maintained. but vitepress is better in every way. no reason to start new projects with vuepress.

**honkit**

fork of the old gitbook cli. exists for backward compatibility. don't use for new projects.

## ranking by agent-readiness

| rank | framework | raw markdown | llms.txt | extraction difficulty |
|---|---|---|---|---|
| 1 | mintlify | endpoints | auto | trivial |
| 2 | fumadocs | endpoints | opt-in | trivial |
| 3 | nextra | repo | manual | easy |
| 4 | starlight | repo | manual | easy |
| 5 | vitepress | repo | manual | easy |
| 6 | docusaurus | repo | manual | moderate |
| 7 | docus | repo | manual | moderate |
| 8 | mkdocs | repo | manual | easy |
| 8 | sphinx | repo | manual | moderate (restructuredtext) |
| 9 | gitbook | partial | no | moderate |
| 10 | docsify | no | no | very hard |

## my recommendation

**for new projects:** pick from tier 1. the agent-readiness difference is real and growing.

**for existing docusaurus sites:** they're fine. don't migrate just for extraction. but consider adding `llms.txt` manually.

**for python projects:** mkdocs or sphinx. the ecosystem expects it.

**for rust projects:** mdbook. it's the standard.

**for vue projects:** vitepress. obvious choice.

**for next.js projects:** fumadocs or nextra. both work. fumadocs is more feature-rich.

**avoid:** docsify for anything public you want agents to read.

## raw markdown is becoming an expected feature

frameworks that expose markdown natively are winning. mintlify and fumadocs are growing fastest. docusaurus is stable but not innovating on agent-readiness. starlight is new but gaining because of its performance.

the gap between tier 1 and tier 2 will widen. tier 1 frameworks are optimizing for both humans and agents. tier 2 frameworks are optimized for humans only.

in 2027, "does your docs framework expose raw markdown?" will be a standard evaluation question. right now, most people don't ask it. they will.

## ask how agents will get the content before you choose

pure documentation generators are simple tools. the difference between them is small for human readers.

for ai agents, the difference is massive. tier 1 frameworks give agents direct access to content. tier 4 frameworks make agents work for every paragraph.

if you're choosing a docs framework in 2026, agent-readiness should be one of your criteria. even if you don't care about agents today, you'll care tomorrow.

---

**related:**
- [mintlify vs docusaurus vs fumadocs](/blog/mintlify-vs-docusaurus-vs-fumadocs)
- [browser-rendered docs break extraction](/blog/browser-rendered-docs-extraction-problem)

---

<!-- https://agentcache.run/blog/r2-vs-s3-storage -->

# why cloudflare r2 beats s3 for my documentation bundles

> i use cloudflare r2 instead of s3 for storing documentation bundles. here's why zero egress fees matter when you serve zip downloads.

s3 is the default object store. it's everywhere. amazon built it. it works.

i chose cloudflare r2.

## s3 bills for downloads; r2 doesn't

s3 charges for data egress. when someone downloads a file from s3, amazon charges you for the bandwidth.

r2 charges zero for egress. downloads are free.

for most use cases, this doesn't matter. you store backups in s3. you retrieve them rarely. the egress cost is negligible.

for agent cache, egress is the primary cost.

## my use case: zip file downloads

every extraction produces a zip file. users download that zip.

average zip size: 5-10 MB. some are 50+ MB (stripe docs is 8.7 MB).

s3 pricing: $0.09 per gb downloaded.

10,000 downloads of 5 MB zips = ~50 GB = $4.50. 100,000 downloads = $450. 1,000,000 downloads = $4,500.

r2 pricing: $0.

## does r2 have downsides?

yes. a few:

**api compatibility.** r2 is mostly s3-compatible. mostly. some edge cases in multipart uploads. but for simple put/get operations, it's identical.

**console experience.** s3 console is mature. r2 console is newer. simpler. fewer features.

**list performance.** r2 can be slower when listing many objects. i don't list objects much.

**region availability.** s3 has more regions. r2 is newer. but for my scale, the available regions are sufficient.

**no lifecycle management (yet).** s3 has sophisticated lifecycle rules. r2's are simpler. again, i don't need complex rules.

for my workload (put object, get object, occasional delete), r2 is perfect.

## download costs at three traffic levels

assuming my launch goes moderately well:

| downloads/month | avg zip size | s3 egress | r2 egress |
|---|---|---|---|
| 10,000 | 5 MB | $4.50 | $0 |
| 100,000 | 5 MB | $45 | $0 |
| 1,000,000 | 5 MB | $450 | $0 |

r2 storage cost: $0.015/GB/month.
s3 storage cost: $0.023/GB/month.

r2 is slightly cheaper for storage too. but the real savings is egress.

## when s3 would be the right choice

- mixed workloads (not just downloads)
- need complex lifecycle policies
- need cross-region replication
- enterprise requirements for s3 specifically
- need glacier/archive tiers
- already invested in aws ecosystem

for agent cache, none of these apply. r2 wins on cost for my specific workload.

## implementation

r2 is s3-compatible. i use the aws sdk with an r2 endpoint:

```javascript
const s3 = new S3Client({
  endpoint: `https://${accountId}.r2.cloudflarestorage.com`,
  credentials: { accessKeyId, secretAccessKey },
  region: 'auto'
})
```

put object:
```javascript
await s3.send(new PutObjectCommand({
  Bucket: 'agent-cache',
  Key: 'jobs/ac-4k9z1m8x/bundle.zip',
  Body: zipBuffer
}))
```

get object:
```javascript
await s3.send(new GetObjectCommand({
  Bucket: 'agent-cache',
  Key: 'jobs/ac-4k9z1m8x/bundle.zip'
}))
```

if r2 didn't exist, i'd use s3. the code is identical.

## r2 fits a product that mostly serves downloads

r2 is not universally better than s3. it's better for my specific use case: serving zip file downloads.

zero egress fees make it the obvious choice. i'm not paying amazon for bandwidth i didn't ask for.

if you're building anything with significant download volume, r2 is worth considering.

---

**related:**
- [agent cache architecture deep dive](/blog/architecture-deep-dive)
- [turso vs postgresql](/blog/why-turso-over-postgres)

---

<!-- https://agentcache.run/blog/robots-txt-crawling-ethics -->

# do i respect robots.txt? my crawling ethics policy

> i respect robots.txt, rate limits, and only crawl public docs. here's my full crawling ethics policy.

agent cache crawls documentation sites. this raises ethical questions. do i respect robots.txt? what's my rate limit? do i bypass paywalls?

here's my policy. short version: i respect the site owner's wishes.

## robots.txt policy: always respect it

if a site's `robots.txt` blocks me, i don't crawl.

i look for two things:

1. **global block:** `User-agent: * Disallow: /docs/`, if this exists and covers docs paths, i don't extract.
2. **agent-specific block:** `User-agent: agent-cache Disallow: /`, if a site specifically names me, i respect it.

before crawling any site, i fetch and parse `robots.txt`.

if blocked, i show a clear message: "this site blocks crawlers via robots.txt. i respect that."

## rate limits: polite crawling

my default settings:
- **concurrency:** 8-12 workers
- **delay between requests:** 200-500ms per worker
- **burst protection:** max 20 requests/second to any single domain

this is below the threshold that triggers most rate limits. it's also polite.

if i get a `429 too many requests`, i:
1. immediately pause requests to that domain
2. wait the `retry-after` duration from headers, or 60 seconds
3. resume with reduced concurrency (halved)
4. if blocked again, mark as failed and move on

## what i never crawl

- **paywalled content:** if it requires payment, i don't attempt extraction.
- **authenticated content:** if it requires login, i don't attempt extraction.
- **private/internal documentation:** i only crawl public urls.
- **behind captchas:** if a captcha appears, i stop and mark as failed.
- **personal data:** i don't extract sites containing user data.

## user-agent identification

my user-agent identifies me:

```
agent-cache/0.1 (+https://agentcache.run/bot; contact@agentcache.run)
```

transparent. reachable. not pretending to be a browser.

## how to block agent cache

if you're a site owner and want to block me:

### option 1: robots.txt
```
User-agent: agent-cache
Disallow: /
```

### option 2: contact me
email `contact@agentcache.run` with your domain. i'll add you to my blocklist.

### option 3: rate limit
any rate limit under 10 requests/second effectively blocks me. i won't fight it.

## public access doesn't mean permission to ignore the owner

i only extract publicly accessible documentation.

this is my hard rule. no authenticated content. no paywalls. no user data.

if it's public and meant to be read, i'll extract it. if the owner asks me to stop, i stop.

## why this matters to me

crawling is a privilege, not a right. i rely on docs sites being accessible. that means respecting the people who maintain them.

my approach: polite, transparent, and reversible.

most sites don't block me. most are fine with extraction. but for those that aren't, i respect boundaries.

## i'll stop crawling if you ask

- i respect robots.txt
- i crawl politely
- i only extract public docs
- i stop when asked

if you maintain a docs site and have questions or concerns, email me.

---

**related:**
- [bot protection: 12% of sites that fail](/blog/bot-protection-docs-extraction)

---

<!-- https://agentcache.run/blog/state-of-docs-for-agents-2026 -->

# how readable are docs for ai agents in 2026?

> i extracted 100+ documentation sites to understand how docs are adapting for agents. here's what i found, and where i'm headed.

every year, documentation changes. in 2026, the change is about agents.

ai coding agents need to read docs. not humans. not search engines. agents. and most docs sites are not built for this audience.

i extracted 100+ documentation sites. here's what i found.

## methodology

100+ sites. extracted, analyzed, categorized. the sites span:
- javascript frameworks (react, next.js, hono, svelte)
- api documentation (stripe, twilio, sendgrid)
- databases (supabase, prisma, drizzle)
- dev tools (vercel, cloudflare, gitlab)
- ai/ml (openai, anthropic, hugging face)
- productivity tools (linear, notion, slack)

for each site, i recorded:
- what extraction tier worked (llms.txt, github, direct .md, html, or failed)
- framework used (if detectable)
- bot protection level
- total pages extracted (where applicable)
- content quality assessment

## key finding 1: llms.txt adoption is growing but still low


current adoption: roughly 8% of sites.

who supports it: mintlify (auto-generated), fumadocs (opt-in), some docusaurus sites (manual), anthropic (yes), vercel (yes), stripe (no).

why it's slow: it's a new standard. developers need to know it exists, understand its value, and add it. most don't even know about it.

the value proposition is clear: one text file at `/llms.txt` gives agents a map of your docs. but adoption requires awareness, and awareness requires time.

**prediction:** 25% adoption by end of 2027. 50% by 2028. it'll become table stakes for new docs sites.

## key finding 2: framework consolidation is real

in 2023, docs frameworks were fragmented. dozens of tools. each slightly different.

in 2026, it's consolidating around a few:

| framework | share of extracted sites | agent-friendliness |
|---|---|---|
| mintlify | ~18% | high |
| docusaurus | ~22% | medium |
| fumadocs | ~8% | high |
| gitbook | ~12% | medium |
| notion-as-docs | ~5% | low |
| custom | ~35% | varies |

mintlify and fumadocs are pulling ahead specifically because they expose raw markdown by default. this is increasingly a competitive advantage.

docusaurus is mature and widely used, but raw markdown isn't exposed as cleanly. you have to go to github or parse html.

custom frameworks are 35% of sites. this is the hardest category. every custom framework is its own extraction challenge.

## key finding 3: bot protection is increasing

12% of sites failed extraction due to bot protection. this is up from ~5% in my earlier tests.

reasons:
- cloudflare bot fight mode (increased sensitivity)
- per-client rate limiting
- mandatory user-agent verification
- some sites serving captchas to non-browser requests

this trend is concerning. as docs become more "app-like" (javascript-heavy, interactive), they become harder to extract programmatically.

sites with heavy bot protection:
- some next.js sites (vercel hosting, cloudflare in front)
- enterprise docs behind auth walls
- sites using bot management services (datadome, kasada)

the direction is clear: extraction is getting harder, not easier. the free web is becoming gated.

## key finding 4: agent-first design is emerging

a small but growing number of sites are designed with agents in mind.

signs:
- llms.txt files
- direct .md endpoints
- changelog and version info in machine-readable formats
- clean url structures predictable by pattern
- api-first documentation (content and presentation separated)

mintlify and fumadocs are the leaders here. their default output is already agent-friendly.

the next tier: docusaurus with plugins, next.js with custom setup, nextra.

but "agent-first" isn't mainstream yet. most docs are still "human-first with agent support" at best.

## framework rankings (2026 update)

based on my extraction experience:

**tier 1 (easiest):**
1. github pages (raw markdown files)
2. mintlify (.md endpoints, llms.txt)
3. fumadocs (.md endpoints, llms.txt)

**tier 2 (moderate):**
4. nextra (raw files in repo, next.js rendering)
5. docusaurus (markdown source in repo)
6. gitbook (partial .md support)

**tier 3 (hard):**
7. notion (api-based, limited)
8. next.js custom (varies by implementation)
9. entirely custom (stripe, linear, etc.)

**tier 4 (impossible for me):**
10. auth-walled docs
11. heavily bot-protected sites (datadome, etc.)

## industry breakdown

**best:** open source projects on github. raw markdown. no tricks. 100% extraction success.

**good:** modern saas companies using mintlify or docusaurus. most extract cleanly.

**mixed:** older companies with custom docs. some expose markdown. some don't.

**worst:** enterprise docs behind auth, notion-based docs, heavily interactive docs built as web apps.

## predictions for 2027

1. **llms.txt standardizes.** by end of 2027, it'll be expected for new docs sites. major frameworks include it by default.

2. **content/presentation separation becomes standard.** docs content stored as markdown/github. rendering layer separate. this is the "agent-first" architecture.

3. **bot protection arms race continues.** as extraction tools proliferate, sites invest more in bot protection. this hurts legitimate use cases.

4. **paid extraction services grow.** firecrawl, context.dev, context7. these services will grow because free extraction is getting harder.

5. **agent cache becomes unnecessary (for agent-first sites).** if docs frameworks expose markdown natively, extraction tools become less needed. but custom sites will always need work.

## what i can share from the extraction dataset

this report is based on 100+ extractions. the dataset is:
- not public (sites don't necessarily want to be listed)
- internally maintained at agent cache
- used to improve my extraction pipeline

if you're a researcher studying docs accessibility, email me. i can share anonymized insights.

## methodology notes

- extraction performed between january and september 2026
- sites chosen based on developer popularity, not random sampling
- extraction defined as: clean markdown for every public docs page, with navigation structure preserved
- "failed" means: couldn't produce usable markdown for the full site

## extraction is still needed while frameworks catch up

documentation is changing. slowly. sites are becoming more agent-aware, but most are still built for human eyes only.

the gap between agent-readiness and reality is large. that's the space agent cache operates in. as long as docs sites require extraction, there's work to do.

but the long-term trend is clear: docs will become agent-first. and when that happens, extraction tools like mine will transform from necessity to convenience.

---

**related:**
- [framework extractor rankings](/blog/docs-framework-agent-readability)
- [how i choose an extraction method](/blog/acquisition-ladder)
- [agent cache is going open source](/blog/open-source-direction)

---

<!-- https://agentcache.run/blog/stripe-docs-crawler-challenges -->

# how stripe's custom docs framework broke me (and what i learned)

> stripe doesn't use mintlify, docusaurus, or gitbook. they built their own. extracting it took days. here's what i found.

stripe has the most popular api documentation on the internet.

they also have one of the most custom docs frameworks i've encountered. extracting it was not a "paste url and wait" operation. it was days of work. back-and-forth. wrong assumptions. dead ends.

here's what made it hard, and exactly how i got through it.

## it's not mintlify. it's not docusaurus. it's not gitbook.

i have extractors for the common frameworks. mintlify? `.md` endpoints, done. docusaurus? structure in `docusaurus.config.js`, raw markdown in `docs/` folder. gitbook? partially supported.

stripe uses none of these. they built their own.

this is the first thing that hits you when you try to extract stripe. your framework detector returns "unknown." your tier 1-4 ladder methods fail. your html purification pulls rendered pages, but the navigation is empty, the content is fragmented, and you're looking at a shell instead of a site.

## content was in redux state, not the page html

stripe's docs are server-side rendered with javascript hydration. but here's the catch: the actual documentation content isn't in the html.

it's in a preloaded redux state object that's embedded in the page as a massive json blob. the html you fetch contains the shell: header, sidebar skeleton, maybe a title. but the actual paragraphs, code examples, parameter tables. all of that is in `window.__PRELOADED_STATE__` or injected via inline script.

this means:

**headless js dom (jsdom) sees the html. great.**
**but the content is in a script tag. jsdom doesn't run scripts.**
**so jsdom sees an empty page.**

i tried running jsdom with `runScripts: 'dangerously'`. still didn't work. the preloaded state uses `JSON.parse()` inside the script, but the initial injection might happen before dom ready in a way jsdom can't replicate.

## strategy: extract json from the script tag

when standard html purification failed, i switched to raw text extraction.

step 1: fetch the html with a simple `GET` request. no jsdom. just raw text.

step 2: find the preloaded redux state. this is a massive json object inside a `<script>` tag. usually marked with an id or class. sometimes it's the first large `<script>` after the `<head>`.

step 3: extract that json with regex or string operations. no html parsing needed. i'm looking for `window.__PRELOADED_STATE__ = {...}` or similar patterns.

step 4: parse the json. the actual markdown content lives inside this state object, under keys like `page.content`, `page.markdown`, or `sections[0].content`.

step 5: convert the structured content back to markdown. sometimes it's already markdown. sometimes it's a custom format that needs transformation.

this is not elegant. this is extraction archaeology. i'm not parsing a docs site. i'm reverse-engineering a single-page app.

## mapping page ids back to the sidebar took a day

stripe has hundreds of pages. the url structure is clean: `/docs/api/charges`, `/docs/api/customers`, etc. but the navigation tree, the sidebar that groups pages into "payments," "billing," "identity", is not in any sitemap i could find.

i had to:
- extract the navigation structure from the redux state
- map each page id to its url slug
- reconstruct the tree manually
- verify against the live site's sidebar

without an llm doing this, i stared at json objects for hours. nested arrays. objects with `children` arrays. pages with `parentId` references. some pages orphaned (api reference pages with no parent). some parents with no content of their own (pure category pages).

i spent a full day just mapping the tree. not extracting content. just mapping urls to categories. yeh waqt nahin milega dobara.

## size: 8.7 megabytes of dense api reference

stripe docs are massive. not in page count (maybe 200-300 pages). in density.

every api endpoint page has:
- description
- request parameters table (20+ rows)
- response object schema
- code examples in 6+ languages
- error codes
- related endpoints
- pagination info
- metadata

a simple "create a charge" page might be 50kb of markdown. multiply by 200 pages. add the product docs (not just api docs, stripe has guides, tutorials, concept docs). add the changelog. add the webhooks documentation.

8.7MB. compressed.

uncompressed, it's closer to 30MB. this is larger than many applications' entire source code.

## what framework detection looks like for stripe

my normal framework detection:

1. check for mintlify markers: `<meta name="generator" content="mintlify">`, not found
2. check for docusaurus markers: navbar class `navbar--fixed-top`, not found
3. check for fumadocs markers: `data-fumadocs-root`, not found
4. check for gitbook markers: `data-gb-page`, not found
5. check for nextra markers: `__NEXT_DATA__`, has this, but it's next.js, not nextra specifically
6. look for redux preloaded state: found.
7. flag as "custom/next.js with redux preloading"

this is tier 5, but not standard html purification. it's "html + script tag state extraction."

## why i didn't give up

i could have.

stripe is the #1 api docs site on the internet. if i can't extract stripe, my tool can't claim to work for "any docs url." stripe is the benchmark. the stress test. the worst-case scenario.

so i kept going. redid the extractor. added a stripe-specific parser. tested, failed, tweaked, tried again.

it took 3 days of focused work. but it works now.

## comparison: stripe vs other payment providers

| provider | framework | extraction difficulty | time |
|---|---|---|---|
| stripe | custom (next.js + redux) | very hard | 3 days |
| paypal | partially custom | medium | 4 hours |
| square | custom | hard | 1 day |
| braintree | docusaurus | easy | 20 minutes |
| razorpay | custom | medium | 3 hours |

stripe is the outlier. everyone else uses something i already know.

## what i learned

**custom frameworks are coming.** as docs become more app-like, more companies will build their own. the era of "everyone uses docusaurus or gitbook" is ending.

**preloaded state is a common pattern.** next.js, remix, sveltekit. all support preloading application state into the html. this is good for user experience (instant hydration). it's bad for simple crawlers.

**the json blob approach works.** when html fails, look inside `<script>` tags. the content might be there, just not where you expect.

**tree reconstruction without ai takes time.** an llm could probably map the navigation tree in minutes. i did it manually. it took a day. the human cost of "no ai in extraction" is real. but the result is deterministic and reproducible.

## how agent cache handles stripe now

stripe has a bespoke extractor in my framework detection pipeline. it:
- recognizes the redux preloaded state pattern
- extracts json from `<script>` tags
- parses the navigation tree from state
- converts content markdown via the same pipeline as other extractors
- handles edge cases like orphaned pages and reference-only parents

when you paste `https://docs.stripe.com` into agent cache, this extractor runs. 8.7MB later, you have the full stripe docs in markdown.

it's not fast. it's not simple. but it works.

## stripe still needs a custom extractor

stripe docs extraction was my hardest technical challenge. not because the tech is impossible. because it's custom, undocumented, and built for humans, not machines.

this is the frontier of docs extraction. as more companies build interactive, app-like documentation, extraction will get harder, not easier.

but that's why agent cache exists. the easy sites extract themselves. the hard sites need work. stripe is the hard site.

---

**related:**
- [what i try before custom extraction](/blog/acquisition-ladder)
- [100 sites extracted: what broke](/blog/100-docs-sites-what-broke)
- [html extraction: the hard path](/blog/html-to-markdown-extraction)

---

<!-- https://agentcache.run/blog/why-turso-over-postgres -->

# why i chose turso (libsql) over postgresql for job metadata

> agent cache uses turso (libsql) instead of postgresql for job metadata. here's why serverless sqlite was the right call.

most web apps use postgresql. it's the default. robust, proven, relational.

agent cache uses turso, a serverless sqlite database.

here's why.

## what i store

agent cache's database stores:
- job metadata: id, url, status, timestamps
- extraction results: page count, strategy used, errors
- basic analytics: which sites are extracted most

the schema is tiny. 3 tables. fewer than 10 columns each.

this is not a complex workload. no complex joins. no transactions spanning multiple tables. no graph queries. no full-text search.

## postgresql: overkill for this workload

postgresql features i don't need:
- advanced query planner (i have simple selects)
- full-text search (i search the filesystem, not the db)
- partitioning (3 tables don't need partitioning)
- replication (single instance is fine)
- complex types (jsonb, arrays, custom types)

postgresql features i'd pay for but not use:
- managed instance at $15+/month
- connection pooling
- backup and recovery
- monitoring

postgresql is a great database. just not for this.

## turso: serverless sqlite

turso is sqlite deployed as a serverless service.

what i get:
- **zero ops.** no database to manage. no migrations to run manually. no connection strings to configure.
- **libsql protocol.** runs locally in development. scales to turso cloud in production.
- **tiny cost.** free tier handles my workload. paid tier is $9/month if i ever need it.
- **sqlite features.** acid transactions. relational queries. small footprint.

## why sqlite works for metadata

sqlite handles my workload perfectly:
- simple schema
- single-writer (job processing is sequential by nature)
- fast reads (status checks, job listing)
- no concurrency conflicts (i control concurrency at the application level)

sqlite's single-writer model is actually fine here. i don't have multiple processes writing simultaneously. the worker pool writes results, the status endpoint reads results. that's acceptable sqlite concurrency.

## turso leaves room for an edge deployment

turso is designed for edge deployments. libsql replicates across regions. low latency for reads.

i'm currently deployed on a single vps. but if i ever move to an edge architecture (cloudflare workers for api, r2 for storage), turso fits naturally.

## when postgresql would be better

postgresql would be the right choice if i had:
- multiple writers (many workers updating the same row)
- complex reports and aggregations
- full-text search requirements
- need for stored procedures
- strict compliance requirements

none of these apply to agent cache.

## moving to postgres if the workload outgrows sqlite

if i outgrow turso, migrating to postgresql is straightforward:
- dump sqlite database
- import to postgresql
- change connection string
- adjust a few queries (sqlite is close to postgresql dialect)

it's not hard. but i probably won't need to. sqlite handles my scale easily.

## three small tables don't need postgres

postgresql is the industry's default database. that's fine.

but not every app needs postgresql. agent cache is a simple tool with simple data needs. sqlite (via turso) handles it perfectly.

use the right tool for the job. sqlite is the right tool here.

---

**related:**
- [agent cache architecture](/blog/architecture-deep-dive)
- [cloudflare r2 vs s3](/blog/r2-vs-s3-storage)

---

<!-- https://agentcache.run/blog/vemetric-docs-case-study -->

# vemetric broke us: 40 hours, 3 fixes, and the bug that survived all of them

> a single docs site ate 40 hours and survived 3 fixes. here's the full autopsy of why vemetric broke agent cache after 100 sites of testing.

40 hours. that's what one documentation site cost me this week. vemetric, a small analytics tool with maybe 30 pages of docs. thirty pages. i've extracted 428-page monsters in under a minute. this one refused to die.

this is the full autopsy. what the site serves, what my pipeline claimed it did, what it actually did, and the bug that survived three separate fixes.

## first, the numbers

context from the 100-site benchmark i ran earlier this year:

- 88% of documentation sites extract cleanly
- 12% fail, mostly from bot protection
- github tree extraction: fastest path, 20-40s for full sites
- html purification: slowest, 190-390s

vemetric looked like a lock for the github tier. open-source product, github link right in the header. easy money.

it wasn't.

## what vemetric actually serves

i probed everything the acquisition ladder checks:

| check | result |
|---|---|
| sitemap.xml | does not exist (0 urls) |
| llms.txt / llms-full.txt | 404 |
| direct .md endpoints | 404 |
| github repo docs | 13 md files, zero of them docs content |

that last one is the trap. the vemetric/vemetric repo has readme, contributing, license, security policy. thirteen markdown files. not one page of actual documentation. the docs live in a private astro app rendered server-side.

so the honest answer was: tier 5, html purification, the slow path.

that is not what my pipeline chose.

## the lie in the logs

my job log for the extraction says:

```
Discovered official GitHub repository: https://github.com/vemetric/vemetric
Acquisition strategy selected: github-raw-markdown
```

then it crawled 32 pages off the live site and converted them with html purification. the "Copy Copied!" button text left in code blocks and the /_astro/ image paths prove it.

the strategy label said github. the work was html. the resolver grabs the first github.com href it finds in the page html and declares it the docs repo. zero verification that the repo contains docs. a footer link is enough to flip the strategy and print a lie into the job record.

fix one: verify the repo actually holds docs markdown before trusting tier 2. if the tree has no docs paths, fall through honestly.

## the tree from hell

content was never the real problem. all 32 pages landed in the bundle. the navigation was the crime scene.

the live sidebar on vemetric.com/docs is clean. five groups, right there in the html:

```
Installation (4 pages)
Product Analytics (5 pages)
Advanced Guides (6 pages)
SDKs (11 pages)
API Reference (10 pages)
```

my extractor produced this instead:

```
Google Tag Manager     -> Google Tag Manager
WordPress              -> WordPress
Next.js                -> Next.js
...
Node.js SDK            -> Node.js SDK, PHP SDK, Python SDK, Go SDK
Getting Started        -> 8 api endpoints
```

sixteen sections. every group label destroyed. each group renamed after its first child page. and my two "getting started" pages (product analytics vs rest api) became indistinguishable.

two bugs, stacked:

1. the group labels are `<strong>` tags. my sidebar parser only reads `<a>` and `<button>`. the categories were literally invisible to it.
2. when a parser produces groups, a separate loop promoted every child page to a top-level section and threw the group wrapper away.

so the viewer showed a flat mess of sixteen self-referencing sections, and reported SUCCESS: 32 files. success lies too.

## the fix that finally worked

i stopped trying to out-smart the dom. third rewrite of the sidebar logic in a week, and every dom selector i write breaks on the next site.

new approach, and it sounds dumb until it works: convert the page to markdown first, then parse the markdown.

turndown is deterministic. it turns the wild sidebar dom into something boring:

```markdown
-   [Installation](/docs/installation)
    -   [Google Tag Manager](/docs/installation/google-tag-manager)
-   **Product Analytics**
    -   [Getting started](/docs/product-analytics/getting-started)
```

a `<strong>` label, an `<a>` label, a bare text label. in markdown they all become parseable text. nested bullets are nesting. bold text without a link is a group. that's the whole parser. no llm, no ai, plain string logic on normalized text.

then the pipeline runs both paths and scores them:

- the dom tree: 16 sections, every section a single item. score: 16
- the markdown tree: 9 real sections, correct groups, correct labels. score: 116

flat trees lose. every time. vemetric now extracts as:

```
Introduction (1)
Installation (5)
Dashboard (1)
Globe (1)
FAQs (1)
Product Analytics (5)
Advanced Guides (6)
SDKs (11)
API (10)
```

the log line even admits what happened: `using extractor markdown-nav (dom primer scored flat)`.

## why it took 40 hours

honest breakdown, because "it was hard" is not a diagnosis:

- the job reported SUCCESS on every attempt. nothing crashed. the failure mode was plausible garbage, which is the most expensive failure mode to debug
- each fix addressed one layer: strategy label, group labels, group promotion. three layers, three fixes, and the site kept producing wrong output until all three landed
- the 100-site benchmark made me overconfident. 88% success rate teaches you the happy path, not the sites that report success while lying

the lesson i keep re-learning: a green checkmark on a crawl means pages were fetched. it says nothing about whether the tree means anything.

## if you maintain docs like vemetric's

three things would have made this site extract in under a minute:

1. a sitemap.xml
2. an llms.txt pointing at your docs
3. raw .md endpoints (astro makes this a few lines of middleware)

any one of the three. your docs are public and meant to be read. agents are reading them whether you plan for it or not.

## bottom line for pipeline builders

plausible garbage beats crashes for wasting your time. if your extractor reports success, make it prove structure: count groups, flag single-item sections, score the tree. and when the dom keeps lying, convert to markdown and parse that instead. the dom is a rendering target. markdown is the contract.

---

**related:**
- [100 sites extracted: what broke](/blog/100-docs-sites-what-broke)
- [the acquisition ladder: how extraction works](/blog/acquisition-ladder)
- [from html to markdown: the hard extraction path](/blog/html-to-markdown-extraction)
- [why i extract docs without an llm](/blog/why-no-llm-extraction)

---

<!-- https://agentcache.run/blog/versioned-documentation -->

# how i handle versioned documentation (v1, v2, v3)

> many docs sites have version selectors. here's how agent cache extracts versioned docs and which version you get.

many documentation sites have version selectors. "v1" vs "v2." "latest" vs "legacy." "api v3" vs "api v4."

when you paste a docs url, which version does agent cache extract?

## versions can live in paths, subdomains, or dropdowns

| site | version pattern | example |
|---|---|---|
| react | `/docs/` (latest) + `/docs/18/` | url-based |
| node.js | `/api/v18/` | path-based |
| docusaurus | version dropdown | spa-rendered |
| mintlify | `/v1/`, `/v2/` | subdomain or path |

different sites handle versions differently. some put the version in the url. some use a dropdown rendered by javascript. some default to the latest version.

## my strategy: extract the default version

for most sites, the "default" url serves the latest version. that's what i extract.

examples:
- `https://hono.dev/docs` → whatever version hono serves by default (currently v4)
- `https://react.dev` → latest react docs
- `https://docs.stripe.com` → latest stripe api version

this is the version most users want. most agents need the current stable docs, not legacy versions.

## version detection methods

i detect versions in several ways:

**url path:** `/v1/`, `/v2/`, `/docs/v18/`

**subdomain:** `v1.docs.example.com`, `v2.docs.example.com`

**meta tags:** `<meta name="doc-version" content="v2">`

**framework-specific:** docusaurus and mintlify expose version info in their apis. fumadocs has version tabs.

## variant support (future)

i'm adding variant support. extractions will be tagged with version info:

```
.agentcache/docs/hono-v3/
.agentcache/docs/hono-v4/
```

each is an independent extraction. you can have multiple versions side by side.

this is not implemented yet. current behavior: one extraction per site, latest version.

## why not extract all versions

extracting all versions means:
- more storage
- longer extraction time
- more maintenance
- user confusion about which version to use

most users want the latest. if they need an older version, they can extract it explicitly by pasting the versioned url.

## extracting a specific version

if you need a specific version, paste the versioned url:
- `https://react.dev/docs/18.2` (if it exists)
- `https://v2.docs.example.com/api`

agent cache extracts the url you give it. if that url points to a specific version, that's what you get.

## showing detected versions in the ui, later

a future version of the web app might show:

```
extracting hono docs...
version detected: v4 (latest)
also available: v3 (extract separately if needed)
```

giving users awareness without forcing all versions.

## paste a versioned url when latest isn't right

agent cache extracts the default (usually latest) version.

for most use cases, that's correct. agents need current docs.

if you need a specific version, paste the versioned url directly.

variant support is coming. but latest-first is the right default.

---

**related:**
- [framework extractor rankings](/blog/docs-framework-agent-readability)
- [how docs urls become markdown](/blog/acquisition-ladder)

---

<!-- https://agentcache.run/blog/why-hono-over-express -->

# why i chose hono over express for agent cache

> express is the default. fastify is performant. i chose hono. here's why an ultra-lightweight framework was the right call for agent cache.

express is the default for node.js web apps. it's everywhere. tutorials start with express. job descriptions mention express.

i chose hono.

## what hono is

hono is a tiny web framework. smaller than express. faster than express in some benchmarks. built for edge runtimes but works everywhere.

key features:
- middleware-based routing (like express)
- built-in support for jsx, streaming, and server-sent events
- edge-compatible (cloudflare workers, deno, bun)
- pure typescript
- tiny bundle: ~14kb

## why not express

express served the node ecosystem for 15 years. but it's showing its age:
- callback-based middleware (promises work but aren't native)
- no built-in typescript support
- large dependency tree
- designed for the node runtime, not the edge

express works. it's just not optimal for 2026.

## why not fastify

fastify is fast. genuinely fast. built by the same team that built express.

but fastify's speed comes from overhead that doesn't matter for agent cache:
- json schema validation
- plugin architecture
- logging and hooks system
- tighter security headers

these are great features. but i don't need them. agent cache has simple routes: form submission, status streaming, download. no complex validation. no plugins.

fastify is faster for some workloads. but for me, hono is simpler.

## why hono won

**sse support.** agent cache streams extraction progress via server-sent events. hono has native sse support. express needs a library.

**edge portability.** i started with cloudflare workers (before switching to a vps). hono runs on workers natively. express doesn't.

**bundle size.** agent cache's server is small. hono matches. express is overkill.

**built-in jsx.** hono has first-class jsx support for server-rendered html. no additional libraries needed.

**typescript.** hono is written in typescript. types are excellent. express's types are community-maintained and sometimes lag behind.

## hono + bun

i run on bun. not node.

bun is fast. particularly for io-bound workloads. documentation extraction is io-bound: mostly fetching pages over the network.

hono on bun: cold start under 5ms. throughput higher than express on node. memory usage lower.

for a simple web app, this combination is hard to beat.

## what i lost

**ecosystem.** express has 15 years of middleware. authentication libraries, validation libraries, template engines, everything.

hono's ecosystem is smaller. but it's growing fast. and for my needs (lightweight server, simple routes), the existing middleware is sufficient.

**familiarity.** almost every node developer knows express. hono is newer. onboarding contributors is slightly harder.

**corporate backing.** express is backed by the openjs foundation. hono is maintained by one developer with community contributions.

these are real tradeoffs. but for my use case, the benefits outweigh the costs.

## performance comparison

| metric | express (node) | hono (bun) |
|---|---|---|
| cold start | ~50ms | ~5ms |
| req/sec (hello world) | ~15k | ~25k |
| memory footprint | ~40MB | ~15MB |
| framework size | ~500kb | ~14kb |
| sse support | library needed | built-in |

note: these are rough benchmarks. real workloads vary.

## when express would be better

express is still the right default for:
- teams with existing express expertise
- complex middleware chains
- applications that need ecosystem libraries
- corporate environments with standardization requirements

## hono covers my routes and streams without much extra

hono isn't better than express for everything. it's better for agent cache.

lightweight. fast. edge-compatible. built-in sse. tiny bundle. bun-compatible.

for a simple docs extraction app, that's the right combination.

---

**related:**
- [why i chose htmx over react](/blog/why-htmx-over-react)
- [agent cache architecture deep dive](/blog/architecture-deep-dive)

---

<!-- https://agentcache.run/blog/why-htmx-over-react -->

# why i chose htmx over react for agent cache

> everyone uses react for web apps. i chose htmx + hono for agent cache. here's why server-rendered html beats spa for this use case.

when developers hear "web app," they reach for react. or next.js. or vue.

i chose htmx + hono + server-rendered html. no spa. no client-side framework. no build step.

for a documentation extraction tool, this was the right call. here's why.

## what the app actually needs

agent cache is a simple web app:
- a form where you paste a docs url
- a status page that streams extraction progress
- a results page showing extracted docs with a download link
- that's it

there are no real-time dashboards. no drag-and-drop. no complex state. no user accounts. no reactive data grid.

it's a form, a terminal-like stream, and a download button. that's the whole ui.

## a form and a progress stream don't need a spa

spas are for apps with lots of client-side interactivity. dashboards with real-time charts. collaborative editors. image editors. anything where the ui state changes constantly based on user input.

agent cache has none of that.

what do you get with a spa for a simple app? extra javascript bundle. extra build step. extra complexity. extra dependencies that need updating. extra attack surface.

what do you lose? nothing this app needs.

## htmx: html over the wire

htmx lets you add interactivity to server-rendered html without writing javascript.

want a form to submit without a page reload? add `hx-post="/api/extract"` to the form.

want a div to update when extraction progresses? add `hx-get="/api/status"` with polling.

want to stream progress from the server? server-sent events (sse) with `hx-sse`.

the server sends html fragments. the browser swaps them into the dom. no javascript state management. no virtual dom. no diffing.

## bun and hono on the server, htmx in the browser

- **bun:** fast typescript runtime. starts instantly.
- **hono:** ultra-lightweight web framework. middleware-based. handles sse natively.
- **htmx:** html-over-the-wire for the frontend. no build step.
- **turso:** serverless sqlite for job metadata. zero ops.
- **r2:** cloudflare object storage for zip files. zero egress fees.

total client-side javascript bundle: **0 bytes.**

htmx is loaded from a cdn script tag: `<script src="https://unpkg.com/htmx.org@2.0"></script>`.

that's the entire frontend.

## what i gave up (and don't miss)

**component libraries.** no shadcn, no radix, no material ui. just server-rendered html with inline tailwind classes.

**npm ecosystem for frontend.** no webpack, no vite, no rollup. no bundle configuration. no tree-shaking.

**type safety across the api boundary.** no tRPC, no graphql. hono has typed routes, which covers most of it.

**offline capabilities.** not relevant. this app needs the server.

**client-side routing.** not relevant. each page is a separate route. `/`, `/job/:id`, `/docs/:id`.

## what i gained

**speed.** the first meaningful paint is under 100ms. no javascript to download, parse, execute. html arrives. it's rendered.

**simplicity.** the entire frontend is server-rendered html fragments. if something looks wrong, i check the server response. not the react devtools. not the console. the actual html.

**reliability.** fewer moving parts. fewer dependencies. fewer things to break. the server generates html. the browser displays it. that's layer zero.

**cost.** no frontend build pipeline. no cdn for js bundles. no dependency upgrades. the "frontend" is a script tag and some html.

**deployment.** bun runs the server. that's it. no "build" step that can fail. no webpack drama. no node_modules resolution hell.

## when react would be better

react would be better if agent cache had:
- a real-time dashboard with charts
- drag-and-drop file uploads (well, i have a simple form)
- offline mode
- complex client-side state (filters, search, sorting)
- user accounts with client-side navigation

none of these exist. so react would be solving problems i don't have.

## results: performance numbers

| metric | react approach (estimated) | htmx approach (actual) |
|---|---|---|
| time to first paint | 300-500ms | ~80ms |
| client js bundle | 100-300 KB | 0 KB |
| total dependencies | 1000+ | ~50 |
| build time | 5-15s | 0s (no build) |
| server response | same | same |

the server doesn't change. it's hono either way. the difference is what's sent to the browser.

## should you use htmx?

use htmx if your app is:
- content-heavy (blogs, dashboards, admin panels)
- form-heavy (wizards, configuration, settings)
- real-time state is handled by the server (sports scores, extraction progress, logs)
- you want instant start times and zero client bundle

use react (or similar) if your app is:
- interaction-heavy (image editors, spreadsheets, games)
- needs offline capabilities
- has complex client-side state that the server doesn't manage
- requires real-time collaboration between users

for agent cache, the choice was obvious. it's a server-heavy app with minimal interactivity. server-rendered html + htmx is exactly the right abstraction.

## htmx is enough for this server-rendered app

developers reach for react by default. that's a habit, not a requirement.

for a simple server-rendered app, htmx + hono is the pragmatic choice. zero client bundle. zero build step. zero regrets.

---

**related:**
- [hono vs express: why hono won](/blog/why-hono-over-express)
- [agent cache architecture deep dive](/blog/architecture-deep-dive)
- [100 sites extracted: what worked](/blog/100-docs-sites-what-broke)

---

<!-- https://agentcache.run/blog/why-local-docs -->

# should your coding agent use local documentation?

> Explore local and offline documentation for Cursor, Claude Code, Codex, and Windsurf, including the tradeoffs of keeping API docs on disk.

When an AI coding assistant needs unfamiliar API details, should it fetch them from a hosted retrieval service or read a local copy? I’m testing the case for local documentation: developers can download docs for offline use, browse the Markdown pages, and save API documentation locally beside a project.

This is a product hypothesis I’m validating, not a claim that local docs work best for every task. A local bundle can give a coding agent a stable reference and work without a network connection. A hosted service can be easier to set up and may have fresher content. The useful choice depends on the library and workflow.

Here are the tradeoffs I’m testing with developers who use Cursor, Claude Code, Codex, Windsurf, or another AI coding assistant.

If you want local documentation for Cursor, Claude Code, Codex, or Windsurf, the key question is whether your agent can find and use the pages in the bundle.

## a docs lookup becomes another network dependency

the api model for docs:
1. agent encounters an unfamiliar api
2. agent sends request to context7 (or similar)
3. service searches its index
4. returns snippets
5. agent uses snippets to write code

this works. it's convenient. no setup needed. but it has real problems.

## problem 1: latency

Every remote lookup depends on network access and the retrieval service. Local files avoid that lookup, though an agent may still need time to find and read the right page.

## problem 2: availability

Local files can be available without a network connection, as long as the right pages were downloaded and the files remain on the machine.

## problem 3: cost

Costs vary by service and plan. A local copy can reduce repeated hosted lookups, but it may require time to create, store, and refresh the files. I’m validating whether developers value that tradeoff enough to use a dedicated tool.

## problem 4: completeness

Hosted retrieval services and local bundles can each miss pages. Index coverage depends on the service; an exported bundle depends on the source site's navigation and what the crawler can access. Either way, check whether the pages you need are present.

An exported bundle can include the pages the crawler finds; completeness depends on what the source site exposes and what extraction can reach. Check that the pages you rely on are included before treating a download as complete documentation.

## problem 5: determinism

determinism means: same output, same result, every time.

api indexes change. ranking changes. content updates. the same query on monday might return different results on tuesday.

for coding agents, this matters. "it worked yesterday" is a frustrating debugging session when the underlying docs changed.

An unchanged local ZIP gives the agent the same source files across runs. Refreshing the bundle changes that snapshot, so teams still need a way to update docs when APIs change.

## how agents actually use docs

How well local files work depends on the agent and how it finds relevant pages.

Some developers already give their coding agent files in the repository; others prefer search or retrieval tools. The product idea is to make a complete docs snapshot easy to create, inspect, and use with whichever agent workflow a developer prefers.

That makes local Markdown a plausible format for an agent-ready reference, while leaving room for indexing or LLM-assisted navigation if testing shows that browsing a large bundle is cumbersome.

## combine a local snapshot with live lookup when needed

Local and remote access can work together. A practical workflow to test is:

1. extract docs once → local zip
2. let the agent read the local Markdown files when it needs them
3. refresh the download when needed → update the snapshot
4. use live retrieval when freshness matters

If you want complete offline documentation for a project, verify that the downloaded bundle includes the pages you rely on. For fast-changing APIs, check whether the snapshot is still current or use live retrieval.

## where local documentation may help

- stable apis you reference daily (stripe, hono, supabase)
- offline development (planes, trains, bad wifi)
- deterministic builds and ci/cd pipelines
- cost-conscious teams
- a reusable local reference, after checking that the pages you need are included

## where remote retrieval may help

- bleeding-edge libraries (react canary, next.js beta)
- quick lookups without setup
- teams that don't want to manage local files
- discovering new libraries you've never used

## current hypothesis: local files are one useful option

Local docs may offer a reusable snapshot and offline access. Whether those benefits outweigh setup and refresh work is what I'm testing.

Remote retrieval may offer convenience and fresher content.

The open question is which developers want a durable local copy, which prefer live retrieval, and whether they want both. I’m validating that need before treating local-first as the answer for every agent workflow. The product can evolve, including adding LLM-assisted organization or retrieval if that solves a real problem users report.

---

**related:**
- [agent cache vs context7: the comparison](/compare/context7)
- [keeping reference docs beside your code](/blog/agentcache-dot-folder)
- [getting docs onto disk without paid extraction](/blog/acquisition-ladder)

---

<!-- https://agentcache.run/blog/why-no-llm-extraction -->

# where LLMs fit in documentation extraction

> Documentation should preserve source text accurately. Here's why I started with deterministic extraction and where LLM-assisted organization or retrieval may help.

I started Agent Cache with deterministic extraction because the first job is to preserve the source documentation, including exact API names and code examples. I’m validating the product now, so I’m also testing where an LLM could make the resulting bundle easier to organize or use.

Sending a whole documentation site to a model and asking it to rewrite everything is one possible approach, but it risks changing details that should stay exact.

That does not mean LLMs have no place in the product. It means source capture and model-assisted features have different jobs: extraction should preserve what the docs say, while a model may help with discovery, grouping, or navigation when that adds value.

## what llms are actually good at

llms are great at:
- understanding natural language
- generating creative content
- summarizing long texts
- answering questions
- reasoning across domains

These capabilities may help with tasks around documentation, such as finding relevant pages or suggesting useful labels.

## what docs extraction actually needs

Source capture needs to preserve page content, including exact API signatures and code examples. A useful bundle also needs clear structure and a way to find relevant pages.

Using a model to rewrite source pages can make these requirements harder to guarantee, so any model-assisted step should be bounded and checked against the original text.

## cost and latency depend on the task

Converting every page through a model adds model calls and depends on the chosen model, page size, and output length. Direct Markdown or HTML conversion can avoid those calls. A smaller task such as suggesting labels or grouping pages may have a different cost and response time, so I’ll evaluate those separately.

## problem 3: determinism

determinism means: same input, same output, every time.

this is critical for agent cache. if you extract the same docs site twice, you should get the same zip file. your agent should see the same documentation.

Model output can vary between runs. If the model rewrites source pages, small changes to headings, code, or wording can make it harder to compare bundles. Keeping the original text intact avoids relying on a model to reproduce it.

for a reference pipeline, randomness is a bug, not a feature.

## problem 4: completeness

Models have context limits, so a large site would need to be split into smaller requests. If a model is asked to summarize or rewrite pages, it may omit details. Keeping the source pages available makes those omissions easier to catch and lets an agent consult the original.

## problem 5: hallucination

When asked to generate or rewrite technical content, a model can introduce errors. For example, it might:
- rename an api parameter (because the real name is "confusing")
- skip a deprecated method (because "nobody uses it")
- add fake parameters to "improve" the api
- change code examples to "cleaner" versions that don't actually work

That risk matters when a coding agent relies on exact API behavior. Any generated guidance should be checked against the source documentation.

## metadata and organization can use semantic help

metadata generation.

after extraction, agent cache generates `meta.yaml` with:
- name, description, url
- repository
- keywords (8-12 technical search tokens)
- intent_triggers (4-6 natural user questions)
- ecosystem (4-8 related tools)

Metadata such as topic labels and intent triggers benefits from semantic understanding. A model can suggest these after extraction, while the original Markdown remains available for comparison.

That is one candidate use for an LLM; the cost, quality, and need for review should be validated with real users and sites.

## keep source extraction separate from model-assisted features

agent cache uses pure code:
- regex for url patterns
- jsdom for html parsing
- turndown for markdown conversion
- framework-specific extractors for nav/tab/version structures
- simple string operations for cleaning

The initial extraction path uses ordinary code for parsing and conversion. That gives us a baseline to compare against if we add model-assisted features.

This is the deterministic baseline. It gives us a source bundle to compare with any LLM-assisted discovery or organization feature.

## when an llm would make sense

Potential uses to validate include suggesting groups for pages, helping users find a relevant page, and recovering useful structure when a site's navigation is unclear. Those features should point back to the original Markdown so users can inspect the source.

## preserve source text, then test model assistance where useful

Documentation extraction needs to preserve the source. Organization and retrieval can benefit from interpretation, if a model can improve the experience without obscuring or altering that source.

The extraction step turns source pages into Markdown. LLM-assisted features could then interpret that content to improve discovery, without replacing the source files.

The product direction is still being tested. I’m keeping exact source capture as a requirement while exploring whether LLM-assisted organization or retrieval helps developers get useful documentation into their coding agent's context sooner.

---

**related:**
- [deterministic extraction vs llm summarization](/blog/deterministic-vs-llm-extraction)
- [the deterministic extraction baseline](/blog/acquisition-ladder)
- [html extraction: how i clean docs](/blog/html-to-markdown-extraction)

---

<!-- https://agentcache.run/blog/zero-memory-architecture -->

# if state matters, agent cache writes it to disk

> agent cache uses a zero-memory rule where nothing stays in ram. everything streams to disk. here's why this matters for reliability.

most web apps hold state in memory. they load things from the database. keep them in ram. update them in ram. write back to disk "later."

if the process crashes, that state is gone. if the server restarts, it's gone. you hope it doesn't happen.

agent cache uses a different rule: **nothing meaningful stays in ram.** everything that matters streams to disk immediately.

## important data must be recoverable from disk

rule: **if data is important, it must be reconstructable from disk.**

this means: no in-memory caches with "write to db later." no in-memory job queues. no state that vanishes on crash.

if the process dies, restart it. it reads from disk and resumes.

## how it works

every extraction job gets a directory on disk:

```
storage/jobs/ac-{id}/
├── events.jsonl      ← append-only event log
├── logs.txt          ← human-readable logs
├── final/            ← extracted docs
└── bundle.zip        ← final package
```

**events.jsonl** is the source of truth. every event is appended as a json line:

```json
{"ts":"2026-09-09T12:34:56Z","event":"page.extracted","url":"https://...","size":1234}
{"ts":"2026-09-09T12:34:57Z","event":"navigation.found","pages":42}
{"ts":"...","event":"error","type":"rate_limited","retry_after":60}
```

append-only means: events are never modified. to cancel a crawl, append a cancellation event. to mark complete, append a completion event.

## why append-only is powerful

**reliability:** if the process crashes mid-write, the existing log is still valid. no partial writes. no corrupted state.

**replayability:** any agent or human can read the log and reconstruct exactly what happened.

**no locks:** no need for transaction locks. just append.

**audit trail:** the log is a perfect record of every decision.

## sse streaming from disk

the browser sees extraction progress via sse (server-sent events).

where does the sse data come from? not memory. disk.

when a user visits `/dingdong/ac-{id}`:
1. server opens `events.jsonl`
2. replays all existing events via sse
3. keeps the file open
4. streams new events as they happen

if the user refreshes: replay again from the beginning. if the server restarts: data is still on disk.

## crash recovery

what happens when the process dies during extraction?

1. process restarts
2. reads last known state from `events.jsonl`
3. determines what was already extracted
4. resumes from where it stopped
5. new events appended to the log

no state lost. no user confusion. the extraction just continues.

## disk is slower, but the crawler waits on the network

ram access: ~10 nanoseconds.
disk access (ssd): ~100 microseconds.

10000x slower.

but for this workload, it doesn't matter.

extraction is network-bound. most time is spent fetching pages, not writing to disk. the difference between ram and disk is negligible compared to network latency.

## when this doesn't make sense

zero-memory is a deliberate choice. it's not always right.

**don't use zero-memory for:**
- high-frequency trading (microseconds matter)
- real-time games (60fps can't wait for disk)
- in-memory caches that can be rebuilt (redis is fine for caching)

**do use zero-memory for:**
- long-running processes that might crash
- processes where state loss is expensive
- systems where reproducibility matters
- applications where simplicity > performance

## implementation: evlog and jsonl

i use evlog for structured logging to jsonl files.

jsonl (json lines) is the format: one json object per line. append-only. human-readable. machine-parseable.

```javascript
import { evlog } from 'evlog'

const log = evlog.for('extractor')

log.append('job.started', { url, jobId })
log.append('page.fetched', { url, size, duration })
log.append('page.converted', { url, strategy: 'direct-md' })
log.append('job.completed', { pages, totalSize })
```

each call appends to disk. no buffering.

## treat ram as disposable, not the source of truth

the zero-memory rule sounds extreme. but it's surprisingly simple to implement. and it makes the system crash-proof.

if data matters, it goes to disk. if it's in ram, it's disposable.

this is how agent cache stays simple, reliable, and debuggable.

---

**related:**
- [agent cache architecture deep dive](/blog/architecture-deep-dive)
- [cloudflare workers killed my serverless dream](/blog/cloudflare-workers-50-subrequest-limit)

---


## Comparisons (8)

<!-- https://agentcache.run/compare/content-dev -->

# agent cache vs content.dev: tool vs diy pipeline

> content.dev is an open-source toolkit for building docs extraction pipelines. agent cache is a working tool. same extraction, different effort. here's why.

content.dev is an open-source project that helps developers build documentation extraction pipelines. it's a toolkit. libraries, scripts, configuration options.

agent cache is a tool that downloads docs and gives you a zip file.

both extract documentation. content.dev makes you build the tool. agent cache is the tool already built. this is the "build vs buy" debate, except "buy" is free.

## content.dev supplies the extraction pieces

content.dev is a collection of open-source packages for extracting documents. it provides parsers, converters, chunkers, and extraction recipes. you configure it. run it. maintain it. troubleshoot when sites change their layout.

it's customizable. you can write custom extractors for frameworks that aren't supported yet. you can tweak parameters. you can integrate it into your own toolchain.

it's free (as in open source). you self-host. you manage dependencies. you handle updates.

**but it's not a finished product.** it's a toolkit for building one.

## agent cache has the pipeline already built

agent cache is a working product. paste a url. get a zip. that's the whole workflow.

it already handles:
- framework detection (mintlify, docusaurus, gitbook, etc.)
- acquisition ladder (llms.txt → github → direct .md → html → fallback)
- navigation extraction
- content cleaning per framework
- meta.yaml generation
- zip packaging

you don't configure any of this. it just works.

**and it's free.**

## building the pipeline takes more than running it

content.dev gives you the ingredients. agent cache gives you the meal.

with content.dev:
1. clone the repo
2. install dependencies
3. configure extraction rules
4. write custom extractors for unsupported frameworks
5. handle edge cases (dynamic rendering, bot protection, layout changes)
6. debug when a site breaks
7. update when upstream changes

with agent cache:
1. paste url
2. wait
3. download zip

if you're a developer with time and specific needs, content.dev is great. you get full control.

if you just want docs and don't want to build a pipeline, agent cache is obvious.

## content.dev gives you control over each step

**full control.** you own every step. you can modify extraction logic. add custom cleaning rules. integrate with your existing toolchain. 

**no dependency on someone else's uptime.** content.dev runs on your infrastructure. agent cache's web app depends on a web service. (though agent cache's cli will be local-only when it ships.)

**learning.** building with content.dev teaches you how docs extraction actually works. that's valuable if you plan to build more tools on top of it.

**internal integration.** if you need to pipe docs extraction into an internal pipeline, deploy it as part of a larger system, content.dev's modular approach fits better.

## agent cache handles setup and framework quirks

**zero setup.** no dependencies. no configuration. no maintenance. no updates to apply when a framework changes. it just works.

**tested.** i've extracted 100+ sites with agent cache. success rate is 88%. those 12 failures taught me what breaks. that knowledge is baked into the product. with content.dev, you learn those lessons yourself.

**framework-specific intelligence.** agent cache knows mintlify's sidebar structure, docusaurus's category system, fumadocs's version tabs. content.dev has generic extractors. you write the framework-specific logic.

**meta.yaml out of the box.** agent cache generates structured metadata: keywords, intent triggers, ecosystem links. with content.dev, you build that yourself.

**free.** both are free. but content.dev costs developer time. agent cache costs zero time after extraction.

## content.dev for integration, agent cache for a ready zip

content.dev is the right choice if you're building a docs extraction platform or integration pipeline. it's a toolkit for builders.

agent cache is the right choice if you just want documentation extracted, formatted, and ready to use. it's a tool for users.

use content.dev when you need to integrate extraction into a larger system and can afford the setup time.

use agent cache when you want to spend zero time on tooling and get clean docs immediately.

---

**related:**
- [agent cache vs context7](/compare/context7)
- [agent cache vs firecrawl](/compare/firecrawl)
- [how agent cache chooses an extraction method](/blog/acquisition-ladder)

---

<!-- https://agentcache.run/compare/context7 -->

# agent cache vs context7: local docs vs remote retrieval

> context7 charges $10/seat for 5,000 api calls. agent cache is free. you own the docs. works offline. here's the honest comparison.

If you're looking for a Context7 alternative, compare the hosted retrieval workflow with a documentation bundle you can download and keep locally. This page looks at the setup, freshness, and ownership tradeoffs for each approach.

context7 is the big name in "docs for agents." built by upstash. 126,000+ libraries indexed. mcp server. plugged into cursor, claude code, codex, devin.

agent cache is a tool that downloads docs once and gives you a zip file. plus an mcp server that scans your local `.agentcache/` folder and serves docs to any agent. free. no tiers. no pricing. just free.

both get docs into your agent's context. one rents access. one gives you ownership. this is the honest breakdown.

## context7 retrieves snippets while your agent works

context7 is a runtime retrieval api. your agent asks context7 for docs while it works. context7 searches their index of 126,000+ libraries. returns snippets. your agent uses those snippets to write code.

updates are fast. their homepage shows supabase docs updated 13 minutes ago. react docs updated 1 day ago. prisma updated 1 hour ago. this is their main selling point: always fresh.

they have an mcp server. plug it into claude code, cursor, codex, whatever. your agent calls context7 automatically. no manual setup after initial install.

**but you don't own anything.** every call goes to their api. you rent access.

## agent cache downloads a reference you can keep

agent cache is a one-time extraction tool. paste a docs url. agent cache crawls the whole site. converts everything to clean markdown. gives you a zip file.

no api calls during your agent session. no internet needed after download. the docs live in `.agentcache/docs/<slug>/` on your disk. your agent reads them like any other file.

the mcp server (coming soon) scans your `.agentcache/` folder and surfaces docs to agents automatically. no manual file references needed.

**and it's free.** no tiers. no credits. no "per seat." no "freemium." no meter running. extracting documentation should be free. it's a public resource. you shouldn't pay rent to access docs that are already public.

## context7 meters api calls; agent cache doesn't

this is where it gets real.

context7 is freemium. they have a free tier, then they charge.

| context7 | |
|---|---|
| free tier | 1,000 calls/month + 20 daily bonus |
| pro | $10/seat/month, 5,000 calls |
| overage | $10 per 1,000 extra calls |

agent cache? free. not "freemium." not "free tier." just free. no limits, no billing, no surprise bills. extracting public documentation should not cost money. this is not a controversial take.

context7's free tier is tight. 1,000 calls sounds like a lot until your agent hits it in a week.

in january 2026 they quietly slashed the free tier by 92%. from something like 12,500 down to 500, then bumped to 1,000 after backlash. this is documented. people noticed.

then they want $10/seat/month for 5,000 calls. then $10 per 1,000 extra.

for docs. public documentation. information that is already on the internet.

i think that's wrong. docs should be free to access. the fact that someone built an indexer and put a meter on it doesn't change that.

## context7 is easier for fresh, popular libraries

i'm not here to trash them. they're better at some things.

**freshness.** context7 updates every few hours. agent cache gives you a snapshot. if you need bleeding-edge docs for a library that releases daily, context7 wins. no contest.

**index size.** 126,000+ libraries. if it's popular and open source, context7 probably has it. agent cache only has what you extract.

**zero setup for popular packages.** `npx ctx7 setup` and you're done. agent cache requires you to paste a url and wait for extraction.

**team consistency.** everyone queries the same index. same version. same snippets. no "did you extract v2 or v3?" confusion.

if you work with rapidly-changing libraries and prefer convenience over ownership, context7 is the pragmatic choice.

## local bundles work offline and include the whole site

**completeness.** context7 returns snippets. agent cache returns the entire site. i extracted stripe docs once: 8.7 mb, hundreds of pages. context7 gives you relevant snippets. agent cache gives you everything.

**offline.** once downloaded, you don't need internet. plane, coffee shop with shit wifi, doesn't matter. your docs are local.

**free.** context7 costs $10+/seat/month minimum for serious usage. agent cache is free. extracting docs should be free. the data is public. the docs are public. charging per api call to read public documentation is rent-seeking.

**ownership.** the docs are yours. in your repo. version controlled. if agent cache disappears tomorrow, your docs don't.

**determinism.** same url, same output, always. context7's index changes. their ranking changes. their api changes. agent cache gives you the same bundle every time.

**any docs site.** context7 indexes popular open source libraries. agent cache works with any docs url. internal docs. private apis. obscure frameworks. if it has a docs site, agent cache can extract it.

**mcp too.** context7 has an mcp server that calls their remote api. agent cache will have an mcp server that reads from your local disk. same convenience. different data source. you get mcp integration without the per-call fees.

## context7 calls upstash; my planned mcp reads disk

both have mcp servers. the difference is what they serve.

context7's mcp calls their api. every request goes to upstash. every request costs money. every request needs internet.

agent cache's mcp reads from your disk. the mcp server scans `.agentcache/`, reads `meta.yaml` files, and surfaces docs to your agent. zero api calls. zero internet. zero cost.

same convenience. fundamentally different architecture.

## repeated lookups use up the call allowance

context7's model seems cheap until you scale. $10/seat, 5,000 calls.

try counting how many times your agent references docs in a typical session. autocomplete. error fixes. "what's the signature for this again?" 50+ calls per session is realistic. two agents on a team, working daily? you'll burn through 5,000 before the month's half done.

then it's $10 per 1,000 extra calls. the meter keeps running.

agent cache is one extraction per docs site. ever. unless you want to re-extract for a new version.

## can you use both?

yes. actually the smartest setup is both.

use context7 for bleeding-edge libraries that update faster than you can re-extract. react, next.js, whatever changes daily.

use agent cache for stable apis your agent references constantly. stripe, supabase, hono. extract once. own forever. no meter running.

i do this. context7 for the fast-moving stuff. agent cache for the references i need every day.

## pay for freshness only where you need it

context7 is convenient and fresh. you pay for that convenience. i think that convenience is overpriced.

agent cache is free, offline, and puts you in control. with mcp coming soon, you get the same convenience without the meter.

charging for docs access is weird. the docs are already public. context7 built an indexer and a nice api and that's useful work. but putting a meter on public information? that's a business model, not a public good.

my opinion: information wants to be free. docs are information. agent cache keeps them free.

use context7 if you want to pay $10+/seat for someone else to host an index of public docs.

use agent cache if you think that's ridiculous.

my take? docs don't actually change that much. stripe's api has been stable for months. hono's core hasn't shifted. you don't need real-time updates for 90% of what your agent does. you need the docs on your disk, ready to read, without a meter running.

if you're paying $10+/seat just to query docs your agent already asked about yesterday, you're optimizing the wrong thing.

use context7 for what it's good at: fresh, popular libraries where setup speed matters.

use agent cache for everything else: stable apis, offline work, ownership, and not paying rent on documentation.

---

**related:**
- [agent cache vs firecrawl](/compare/firecrawl)
- [keeping reference docs beside your code](/blog/agentcache-dot-folder)
- [i extracted 100 docs sites. here's what broke](/blog/100-docs-sites-what-broke)

---

<!-- https://agentcache.run/compare/docsgpt -->

# agent cache vs docsgpt: reader vs chatbot

> docsgpt is a chatbot for docs. agent cache is a tool that downloads docs for your agent. completely different workflows.

docsgpt is a chatbot for documentation sites. you ask it questions like "how do i authenticate with stripe?" it reads the docs and answers. conversation interface.

agent cache is a tool that downloads docs and gives you a zip file.

both interact with documentation. one talks to you. one gives you files. completely different workflows.

## ask docsgpt a question about the docs

docsgpt is a chat interface on top of documentation. you ask questions. it retrieves relevant sections from the docs. you get an answer. conversation continues.

it's like talking to customer support, except support is an ai that only knows the docs.

**good for:**
- quick answers to specific questions
- learning a topic by asking follow-up questions
- getting started without reading the whole docs
- non-technical users who want answers, not docs

**not good for:**
- giving an ai agent a complete reference
- offline api access
- precise code generation with exact signatures
- keeping a local copy for version control

## give your coding agent files instead of answers

agent cache doesn't chat. it doesn't answer questions. it just downloads docs.

the downloaded docs become reference material for your coding agent. your agent reads them like any other file. uses them for code generation. api lookup. understanding context.

**good for:**
- giving your agent full docs context
- offline coding
- version control of docs alongside code
- exact api reference (your agent reads original docs, not summaries)

**not good for:**
- learning by asking questions
- quick answers without reading
- non-technical users

## a person asking questions vs an agent reading a reference

docsgpt is for humans who want answers. agent cache is for agents that need reference.

docsgpt: "how does stripe authentication work?" → here's a summary.

agent cache: your agent reads `stripe/docs/api/authentication.md` → generates working code with exact parameter names and types.

one is discovery. one is execution.

## follow-up questions make docsgpt useful for learning

**discoverability.** learning a new library by asking questions is natural. "what's the difference between these two methods?" docsgpt answers. agent cache gives you both methods and lets you figure it out.

**non-technical users.** product managers, designers, qa engineers. they need answers, not raw docs. docsgpt serves them. agent cache doesn't.

**quick lookups.** "what's the default timeout?" docsgpt answers in seconds. agent cache requires your agent to search through files.

**conversational learning.** follow-up questions. related topics. context accumulation. docsgpt handles this well.

## code generation needs the details a summary can miss

**code generation.** coding agents need exact api signatures. not summaries. they need to read the actual docs. agent cache gives them the actual docs.

**completeness.** docsgpt answers what you ask. agent cache gives you everything. edge cases, deprecated methods, error codes, rate limits. all of it.

**offline.** docsgpt needs internet and the docs site online. agent cache works anywhere after download.

**determinism.** docsgpt might answer differently based on how the ai interpreted the question. agent cache gives the same original docs every time.

**free.** most docsgpt implementations charge for the underlying llm usage. agent cache is free.

## explore with docsgpt, code against downloaded docs

docsgpt is a chatbot. agent cache is a downloader.

use docsgpt when you want to ask questions about documentation.

use agent cache when you want your coding agent to read documentation.

most workflows need both at different stages:
- exploration: docsgpt (ask questions, understand concepts)
- implementation: agent cache (exact api reference, generate code)

---

**related:**
- [agent cache vs context7](/compare/context7)
- [agent cache vs docuchat](/compare/docuchat)

---

<!-- https://agentcache.run/compare/docuchat -->

# agent cache vs docuchat: owner vs chatter

> docuchat lets you chat with documentation. agent cache downloads it. different problems, different tools.

docuchat is another docs chatbot. ask questions, get answers. q&a interface layered over documentation.

agent cache downloads documentation into a zip file.

these are not competitors. they solve different problems. but people search for comparisons, so here it is.

## docuchat builds a q&a layer over your docs

docuchat is a chatbot for documentation. upload your docs (or connect a docs site). ask questions like "how do i set up oauth?" or "what's the rate limit?" and get answers.

it builds a vector index, retrieves relevant sections, and generates responses.

it's a good product category. lots of demand. but it's a **q&a tool**, not a reference tool.

**docuchat gives you answers.** agent cache gives you the source.

## agent cache keeps the source as structured markdown

agent cache doesn't generate answers. it doesn't chat. it downloads the actual documentation files into your project. structured markdown. clean hierarchy. ready for your agent to read.

your coding agent reads these files natively. no vector database. no retrieval api. just files on disk.

## answers in a chat window or source files on disk

| | docuchat | agent cache |
|---|---|---|
| interface | chat | files |
| output | answers | docs |
| use case | questions | code generation |
| accuracy | llm-generated | original source |
| offline | no | yes |
| cost | paid (llm tokens) | free |

## docuchat suits people who want answers, not files

**questions.** if you have specific questions and want instant answers, docuchat is great. "why does this error happen?" "what are the alternatives to this approach?" chat excels here.

**non-developers.** product managers, designers, writers. they want answers, not files. docuchat serves them.

**exploration.** exploring unfamiliar territory by asking questions. chat is natural for this.

## downloaded docs keep exact api details available offline

**code generation.** coding agents need exact api details. parameter names. types. defaults. error codes. rate limits. chat-generated answers distill this. sometimes incorrectly. agent cache gives the original.

**precision.** "the docs say the timeout is 30 seconds" vs. your agent reading the actual docs and knowing it's 30 seconds. trust the source, not the summary.

**offline.** agent cache works without internet. docuchat needs the service.

**no hallucination.** chat interfaces can hallucinate answers. doc files never do.

**free.** docuchat charges for the service and underlying llm. agent cache is free.

## chat while exploring, keep the files for development

docuchat is for asking questions about docs.

agent cache is for coding agents that need to read docs.

use docuchat when you want a conversation.

use agent cache when you want the files.

use both: chat for discovery, downloaded docs for development.

---

**related:**
- [agent cache vs context7](/compare/context7)
- [agent cache vs docsgpt](/compare/docsgpt)

---

<!-- https://agentcache.run/compare/firecrawl -->

# agent cache vs firecrawl: docs-specific tool vs general-purpose scraper

> firecrawl is a web scraping api. agent cache is a docs extraction tool. same extraction, completely different use cases. here's the honest breakdown.

If you're looking for a Firecrawl alternative for documentation, compare a general web extraction API with a docs-focused downloadable bundle. Firecrawl handles broad web extraction; Agent Cache is built around packaging documentation for coding-agent workflows.

agent cache is a tool that turns documentation sites into clean markdown. that's it.

both extract content from websites. that's where the similarity ends. firecrawl is a bulldozer. agent cache is a garden trowel. this article explains why you don't need a bulldozer to plant flowers.

## firecrawl scrapes, searches, and interacts with pages

firecrawl is a general-purpose web scraping api. you give it a url. it scrapes the page. returns clean markdown or structured json. can handle javascript rendering, dynamic content, login flows, multi-step interactions.

it's powerful. 96% web coverage. p95 latency of 3.4 seconds. handles js-heavy pages. crawls entire sites. extracts structured data with schemas. searches the web. interacts with pages (click, scroll, fill forms).

they have an mcp server. 400,000+ installations. integrates with cursor, claude, windsurf.

it's also open source. 100,000+ stars on github. you can self-host if you want.

**but it's not free.** despite being open source.

## firecrawl pricing

free tier: 1,000 credits/month. one credit = one page. so 1,000 pages per month free.

hobby: $16/month. 5,000 pages. standard: $83/month. 100,000 pages. growth: $333/month. 500,000 pages. scale: $599/month. 1,000,000 pages.

pay-as-you-go kicks in when you run out. $5 buys you extra credits. increments vary by plan.

if you exceed your plan, the meter runs. set a monthly cap or it keeps charging.

## agent cache packages documentation, not arbitrary web data

agent cache doesn't scrape the web. agent cache turns documentation sites into agent-ready markdown bundles.

paste a docs url. agent cache crawls the whole site. converts to clean markdown. structures it with `meta.yaml`, `_map.json`, and `INDEX.md`. gives you a zip file.

it's not a general-purpose scraper. it doesn't click buttons. it doesn't fill forms. it doesn't search the live web. it extracts documentation and nothing else.

**and it's free.** no credits. no tiers. no meter. just free.

## scraped pages still need a docs pipeline

firecrawl is infrastructure. it's a building block for applications that need web data. r&d agents, lead enrichment, competitive intelligence, price monitoring, content generation. any app that needs to read the live web.

agent cache is a product. it's a finished tool for one job: getting docs into your agent's context.

firecrawl gives you raw material. you still need to build the pipeline that turns scraped pages into structured documentation. figure out navigation. handle version tabs. strip banners and cookie notices. organize by hierarchy. none of that is automatic.

agent cache gives you the finished bundle. structured, indexed, ready to drop into your agent's context.

## reach for firecrawl beyond public docs

firecrawl is genuinely impressive. it's good at what it does.

**general-purpose scraping.** if you need to scrape e-commerce sites, news sites, social media, job boards, whatever. agent cache can't do any of that. firecrawl handles it.

**live web search.** firecrawl has a search api. ask it "what's the latest on react server components" and it searches the web, scrapes results, returns markdown. agent cache doesn't search. it only extracts what you already know the url for.

**interactions.** click buttons, fill forms, navigate pagination, handle logins. if the docs you need are behind a login wall, firecrawl can probably reach them. agent cache can't.

**structured extraction.** pass a json schema and firecrawl returns structured data matching that schema. product listings, pricing tables, contact info. agent cache returns markdown files. that's it.

**scale.** if you need to scrape millions of pages across hundreds of sites, firecrawl has enterprise plans and dedicated infrastructure. agent cache doesn't scale that way.

if your project involves scraping arbitrary websites, not just docs, firecrawl is the right choice. no question.

## docs bundles need framework-aware cleaning and indexes

but if your need is specifically documentation, agent cache is better. and it's not close.

**purpose-built for docs.** firecrawl treats a docs site like any other website. agent cache knows it's a docs site. it has framework-specific extractors for mintlify, docusaurus, gitbook, fumadocs, nextra, mdbook. it understands sidebar navigation, version tabs, api reference structures. firecrawl just scrapes html and converts to markdown.

**zero pipeline work.** firecrawl gives you raw pages. you still need to figure out which pages to scrape, how to structure them, what to keep, what to strip. agent cache handles all of that. paste url → get structured bundle. done.

**framework-aware cleaning.** firecrawl strips html and converts to markdown. but docs sites have specific patterns. cookie banners, navigation bars, "was this page helpful?" buttons, newsletter signup boxes. agent cache knows how to strip these per framework. firecrawl strips generically.

**structured output.** agent cache gives you `meta.yaml` with keywords, intent triggers, ecosystem links. `_map.json` navigation hierarchy. `INDEX.md` table of contents. firecrawl gives you pages of markdown. you build the structure.

**free.** firecrawl costs $83/month for standard usage. agent cache is free. both are open source. but only one is actually free to use.

## firecrawl is also my last-resort extraction tier

here's the thing most people don't know: **agent cache uses firecrawl.**

firecrawl is tier 6 of agent cache's acquisition ladder. when a docs site can't be extracted via llms.txt, github tree, direct .md apis, content negotiation, or html purification, agent cache falls back to firecrawl as a last resort.

so they're not really competitors. they're collaborators at different layers of the stack.

firecrawl is the brute-force fallback. agent cache is the smart, purpose-built tool.

if agent cache can extract your docs via direct .md endpoint, it takes 20 seconds and costs nothing. if not, it tries html purification. if that fails, it calls firecrawl, which costs credits, and bills you for the scrape.

firecrawl is necessary for the edge cases. agent cache is optimal for the common case.

## organizing the scraped pages is still work

some people say "just use firecrawl" when i tell them about agent cache. they think it's the same thing.

try using firecrawl to extract stripe docs and see what happens. you'll get hundreds of raw markdown pages. then you'll spend an hour organizing them. figuring out the hierarchy. stripping navigation. handling version tabs. structuring the api reference sections.

i've done this. it's tedious. it costs more in developer time than firecrawl credits.

with agent cache, i paste `docs.stripe.com`. ten minutes later i have a structured bundle with `meta.yaml`, `_map.json`, clean markdown files organized by category, and an `INDEX.md`.

the difference is not the extraction. it's what comes after.

## firecrawl for the wider web, agent cache for docs bundles

firecrawl is a web scraping api. it's great at scraping. it's not great at documentation.

agent cache is a documentation extraction tool. it's good at one thing and one thing only: turning docs sites into agent-ready markdown bundles.

if you need general-purpose web scraping, use firecrawl. it's the best tool in that category.

if you need documentation extraction, use agent cache. purpose-built beats general-purpose every time. and it's free.

use firecrawl for the 12% of docs sites that resist all other extraction methods (bot protection, custom frameworks, weird js rendering). that's what it's for.

use agent cache for the 88% of docs sites that have exposed .md endpoints, github repos, or clean html structures. that's what it's for.

---

**related:**
- [agent cache vs context7](/compare/context7)
- [how agent cache extracts docs before trying firecrawl](/blog/acquisition-ladder)
- [i extracted 100 docs sites. here's what broke](/blog/100-docs-sites-what-broke)

---

<!-- https://agentcache.run/compare/llms-txt -->

# agent cache vs llms.txt: not a competition

> llms.txt is a proposed standard. agent cache is a tool that works with or without it. here's how they relate.

llms.txt is a proposed standard. a text file that sits at the root of a docs site and tells ai agents what documentation is available.

agent cache is a tool that extracts documentation sites into clean markdown.

these aren't competitors. they're not even in the same category. but people search "llms.txt alternative" and "agent cache vs llms.txt" so here's how they fit together.

## llms.txt lists documentation resources

llms.txt is a simple standard. a text file in the root of a domain containing a structured list of documentation resources. here's the rough idea:

```
# my-library docs

## getting started
- /docs/introduction
- /docs/installation
- /docs/quickstart

## api reference
- /docs/api/authentication
- /docs/api/endpoints
- /docs/api/errors
```

clean. simple. human and machine readable.

if a docs site has `llms.txt`, an agent can read it, understand the structure, and request relevant documentation. no crawling needed. no guessing.

**but here's the thing: almost nobody has it.**

## agent cache checks llms.txt before other sources

agent cache doesn't need llms.txt to work. it crawls the site regardless.

but when llms.txt exists, agent cache checks it **first**. it's tier 1 of the acquisition ladder. 1 request, instant download, perfect structure, zero cost.

agent cache:
1. checks `/llms.txt` or `/llms-full.txt` first
2. if found, downloads the listed docs directly
3. if not found, falls back to other methods (github, direct .md, html, etc.)

**llms.txt makes agent cache faster.** when it exists, agent cache skips crawling entirely. the tier 1 fallback is instant.

## adoption reality

llms.txt is a great idea. adoption is low.

i've extracted 100+ documentation sites. less than 10% had llms.txt. most don't even know it exists.

so agent cache is built for the 90% of sites that don't have llms.txt. but it benefits from the 10% that do.

## llms.txt doesn't replace agent cache

even if every site had llms.txt, you'd still need agent cache for:

**format conversion.** llms.txt lists urls. those urls still return html (usually). agent cache converts them to clean markdown. strips navigation, banners, cookie notices.

**structure.** llms.txt gives you a flat list. agent cache creates the folder hierarchy, `meta.yaml`, `_map.json`, `INDEX.md`. organization matters for agents.

**offline access.** llms.txt points to live urls. agent cache downloads the content. gives you a zip you can use offline.

**completion.** some sites have incomplete llms.txt files. agent cache crawls the whole site regardless. catches what the llms.txt misses.

## agent cache doesn't replace llms.txt

if every site adopted llms.txt, agent cache would be much faster and cheaper. tier 1 would handle most extractions. the 88% success rate would climb higher.

**agent cache is a backfill for a world that doesn't have llms.txt yet.** but it's more than that. it also handles the conversion, cleaning, and structuring that llms.txt alone doesn't solve.

## llms-full.txt puts the content in one file

some sites have `llms-full.txt` which contains the entire docs in one file. mintlify supports this. when it exists, agent cache downloads it in one request. instant, complete, and structured.

this is the ideal scenario. a single text file with everything. no crawling needed.

but again, adoption is low. agent cache handles both cases: when the standard exists and when it doesn't.

## publish llms.txt; use agent cache to package the docs

llms.txt is a standard that makes agent cache faster when it exists.

agent cache is a tool that handles docs extraction whether or not llms.txt exists.

they're complementary. if llms.txt adoption grows, agent cache benefits. if it doesn't, agent cache still works. 

use agent cache regardless. if a site has llms.txt, agent cache just gets it faster. if it doesn't, agent cache figures it out anyway.

if you're a docs site maintainer: add `llms.txt`. it helps tools like agent cache. it helps your users. it's low effort, high value.

---

**related:**
- [agent cache vs context7](/compare/context7)
- [what i try when llms.txt is missing](/blog/acquisition-ladder)

---

<!-- https://agentcache.run/compare/parallel-web -->

# agent cache vs parallel web: docs extraction vs parallel scraping

> parallel web is a parallel scraping infrastructure. agent cache is a docs extraction product. different tools, different jobs. here's why they're not competitors.

parallel web is parallel web. it's infrastructure for scraping the web at high concurrency and large scale. designed for speed and volume.

agent cache turns documentation sites into agent-ready markdown.

people might list them as "competitors" in a broad sense. they're not. this article is short because the distinction is obvious once you get it.

## parallel web is built to fetch pages at volume

parallel web is a scraping engine optimized for parallel execution. it launches many browser instances simultaneously. crawls thousands of pages in parallel. returns raw data at high throughput.

it's designed for scenarios like:
- monitoring competitor prices across 10,000 pages
- scraping job listings from hundreds of sites
- building large training datasets from the web
- any task where volume and speed are the constraints

it's infrastructure. it gives you raw pages fast.

## agent cache spends its effort on docs structure

agent cache is not infrastructure. it's not built for scale. it's built for quality.

it does one thing: extract documentation sites with structure, navigation, and context intact. it handles framework-specific quirks. it outputs structured bundles with metadata and indexes.

it's slow by design. 8-12 workers. polite crawling. careful parsing. framework-specific extraction logic.

## fetching pages quickly vs organizing a docs site

parallel web is fast and shallow. agent cache is slow and deep.

parallel web: "give me 50,000 pages from this site, raw html, as fast as possible."

agent cache: "give me this documentation site, but understand which div is the sidebar, which is the content area, which are version tabs. strip the navigation but keep the breadcrumbs. organize by framework categories."

parallel web doesn't care about mintlify vs. docusaurus. agent cache cares deeply.

parallel web gives you raw material. agent cache gives you a finished product.

## parallel web fits a high-volume scraping pipeline

**scale.** if you need to scrape 100,000 pages across 500 sites, parallel web handles it. agent cache crawls one docs site at a time with 8-12 workers. not the same league.

**general-purpose scraping.** parallel web isn't opinionated. it handles any site. agent cache only makes sense for documentation sites.

**infrastructure flexibility.** parallel web is a platform. integrate it into your own system. run it at scale. agent cache is a product with a specific workflow.

## a docs bundle saves you the setup and cleanup

**documentation-specific quality.** agent cache knows docs sites. it handles navigation, versioning, tab structures, code blocks, inline code, api reference formatting. parallel web just scrapes html.

**structured output.** agent cache gives you `meta.yaml`, `_map.json`, `INDEX.md`, folder hierarchy. parallel web gives you pages.

**zero configuration.** parallel web needs setup, orchestration, post-processing. agent cache: paste url, get zip.

**free.** parallel web charges for compute and bandwidth at scale. agent cache is free.

## a scraping engine still needs post-processing

these aren't competitors. they're different tools for different jobs.

parallel web is for scraping. agent cache is for documentation.

if you build a massive web scraping platform, you might use parallel web under the hood as part of the pipeline. but for docs specifically, agent cache does more with less.

if you were building a general-purpose web data platform, you'd integrate parallel web (or something like it) for the scraping layer, then add your own post-processing, structuring, and cleaning on top. that's what agent cache is: the post-processing and structuring layer for docs.

## choose by the output: raw pages or organized docs

parallel web is a scraping tool. agent cache is a documentation tool.

if you need to extract data at massive scale from arbitrary websites, use parallel web.

if you need to extract a documentation site into clean, structured, agent-ready markdown, use agent cache.

evaluating them as "competitors" is like evaluating a cnc router against a 3d printer. both make things. totally different use cases.

---

**related:**
- [agent cache vs context7](/compare/context7)
- [agent cache vs firecrawl](/compare/firecrawl)

---

<!-- https://agentcache.run/compare/tavily -->

# agent cache vs tavily: documentation extraction vs ai search

> tavily is an ai search engine for agents. agent cache is a docs extraction tool. they solve different problems. here's why.

tavily is everywhere in ai agent stacks. "the real-time search engine for ai agents." 300m+ monthly requests. 2m+ developers. 99.99% uptime.

agent cache is a tool that downloads docs and gives you a zip file.

people confuse them. "they both give agents information from the web." no. tavily searches the internet for answers. agent cache gives your agent documentation to read. totally different use cases.

## tavily finds current sources for a question

tavily is a search api. you ask it something like "what's new in react 19?" or "how do i implement rate limiting in hono?" tavily searches the web, finds relevant sources, extracts key content, and returns structured results with citations.

it's real-time. it searches the live web. it retrieves current information. it's the difference between asking a question in a group chat and opening old notes.

pricing is roughly $7.50-8 per 1,000 searches. $0.005 per basic search, $0.01 per advanced search. 1,000 free credits per month on the free tier.

## agent cache starts with a docs url and ends with files

agent cache doesn't search. it downloads.

paste a documentation url. agent cache crawls the site. extracts clean markdown. structures it with `meta.yaml`, `_map.json`, `INDEX.md`. gives you a zip file.

your agent reads these docs like any other file. not through an api. not at runtime. just from disk.

**and it's free.**

## searching for an answer vs downloading a known reference

this is the key distinction.

tavily is for **questions you don't know the answer to.**
- "what's the latest hono version?"
- "did stripe change their api recently?"
- "how does this new library work?"

agent cache is for **reference material you already know you need.**
- "give my agent the stripe docs"
- "i need the hono api reference"
- "put supabase docs in my project"

tavily discovers. agent cache owns.

## tavily helps when you don't know which docs you need

**discovery.** when you need to find information you don't have, tavily is unbeatable. it's a search engine. that's what it's for.

**freshness.** tavily searches the live web. if something changed last week, tavily finds it. agent cache gives you a snapshot from when you extracted it.

**broad queries.** "what's the best auth library for hono?" tavily can answer that. agent cache can't. it only has what you give it.

**no setup for new topics.** want to learn about a library you haven't used? tavily finds docs, blog posts, tutorials, github issues. agent cache requires you to know the docs url and extract it.

## repeated api work is easier with a complete local reference

**complete reference.** tavily returns snippets. agent cache gives you everything. every page. every code example. every edge case. agent cache > tavily for deep api work.

**offline.** tavily needs internet and the api. agent cache works anywhere after download.

**determinism.** tavily might return different results for the same query. web changes. ranking changes. agent cache gives the same bundle every time.

**free.** tavily costs ~$8 per 1,000 searches. agent cache is free. if your agent queries docs 100 times per day, that's real money.

**structured output.** tavily returns search results. agent cache returns organized docs with navigation, metadata, and indexes. your agent knows what's available and where to find it.

**no rate limits.** tavily has rate limits and quotas. agent cache reads files from disk. unlimited.

## complementary, not competing

the real workflow is both.

**phase 1: discovery (tavily)**
- "what library should i use for auth?"
- "did this feature change in the latest version?"
- "what's the current best practice?"

**phase 2: implementation (agent cache)**
- "give me the full auth docs"
- "i need the complete api reference"
- "extract the docs into my project"

use tavily to figure out what you need. use agent cache to get it.

## a search summary isn't the full parameter reference

some people say "just use tavily" when they hear about agent cache. they think they're similar.

tavily is great for discovery. but it doesn't replace having the actual documentation.

try asking tavily: "what are all the parameters for stripe's `checkout.session.create` and their exact types and defaults?" it'll give you a summary. maybe accurate. maybe not. maybe missing edge cases.

with agent cache, your agent reads the actual docs file. exact parameters. exact types. exact defaults. no summarization. no filtering. no potential hallucination from the search layer.

**tavily is a starting point. agent cache is the foundation.**

## find the library with tavily, then download its docs

tavily is a search engine. agent cache is a docs downloader.

tavily answers questions. agent cache gives reference material.

tavily is for discovery. agent cache is for execution.

use tavily when you don't know what you need.

use agent cache when you do.

use both: discovery phase with tavily, implementation phase with agent cache.

---

**related:**
- [agent cache vs context7](/compare/context7)
- [agent cache vs firecrawl](/compare/firecrawl)
- [how i turn a docs url into markdown](/blog/acquisition-ladder)

---

