# meta.yaml: what makes documentation discoverable by agents

> every agent cache bundle includes a meta.yaml with structured metadata. here's what goes in it and why agents need it.

when extraction finishes, agent cache creates a `meta.yaml` file. not for humans. for agents.

this file tells any agent (or system) what it just extracted. what the docs are for. when to use them. what they're related to.

without metadata, a folder of markdown files is just a folder of files. with metadata, it's a discoverable resource.

## why metadata matters

imagine you're a coding agent. someone drops `.agentcache/docs/` into the project. 50 markdown files. no labels. no descriptions.

what is this? react docs? stripe docs? your company's internal api docs?

a human can open INDEX.md and figure it out. an agent can't. it needs structured metadata.

## a meta.yaml for hono

```yaml
name: hono
title: Hono - Lightweight Web Framework for the Edge
url: https://hono.dev
description: Hono is a fast, lightweight, and portable web framework for JavaScript and TypeScript. It works with any runtime.
repository: https://github.com/honojs/hono
keywords:
  - hono
  - web framework
  - edge computing
  - typescript
  - deno
  - cloudflare workers
  - bun
  - honojs
  - middleware
  - routing
intent_triggers:
  - how do i create a basic hono app
  - hono middleware pattern
  - hono routing examples
  - hono vs express
  - deploying hono on cloudflare workers
  - hono typescript setup
ecosystem:
  - cloudflare workers
  - vercel
  - deno
  - fastly
  - lambda
  - openrouter
  - zod
  - valibot
version: latest
extracted_at: 2026-09-01T00:00:00Z
framework: mintlify
```

## field breakdown

**name:** short identifier. matches the folder name. `hono`, `stripe`, `supabase`.

**title:** human-readable name. used for display.

**url:** original docs url. so agents and users know where it came from.

**description:** one paragraph about what this library does. helps agents understand context.

**repository:** github url. for linking to source code, issues, discussions.

**keywords:** 8-12 technical terms. these are search tokens. when an agent searches for "hono routing," the keyword match surfaces this docs bundle.

**intent_triggers:** natural language questions that would lead someone to these docs. this is how agents know "i should use hono docs for this request."

**ecosystem:** related tools and platforms. helps agents understand relationships. "this is a cloudflare workers framework."

**version:** which version was extracted. `latest`, `v3`, `v4.2`.

**extracted_at:** timestamp. lets agents know how old the extraction is.

**framework:** which docs framework was used. helpful for understanding extraction quality.

## how agents use meta.yaml

agents scan `meta.yaml` to determine which docs are relevant to a query.

user asks: "how do i set up middleware in hono?"
agent steps:
1. scan `.agentcache/` for meta.yaml files
2. match query to keywords (`hono`, `middleware`) and intent triggers
3. load the hono docs bundle
4. reference the middleware section

without meta.yaml, the agent would need to open and read every docs bundle to find the right one. that's slow.

with meta.yaml, matching is instant. one file per bundle.

## examples: good vs bad keywords

**bad:** `documentation`, `guide`, `tutorial`, `reference`

why: these are generic. every docs site has them. they don't help matching.

**good:** `hono`, `web framework`, `cloudflare workers`, `typescript`, `middleware`, `routing`, `bun`

why: these are specific. they uniquely identify the library and its domain.

## generated vs curated

the meta.yaml fields are generated by an llm. why use an llm here? because keywords and intent triggers require semantic understanding. they're subjective.

but the llm output is not blindly trusted. it's reviewed. corrected. refined.

in the future, i'll add a validation step: check that keywords match the actual doc content. check that intent triggers are realistic questions.

## future: meta.yaml as a standard

the meta.yaml format could become a standard. any docs extraction tool could generate it. any agent could consume it.

if multiple tools use the same schema, agents get portability. switch from one extractor to another, and the metadata stays compatible.

this is one reason i'm documenting the schema publicly.

## ship accurate metadata with every bundle

metadata is what turns a pile of files into a searchable, discoverable, usable resource.

without meta.yaml, your docs bundle is just content. with meta.yaml, it's a tool that agents can reason about.

always generate metadata. always keep it accurate.

---

**related:**
- [organizing docs inside .agentcache](/blog/agentcache-dot-folder)
- [why local docs matter](/blog/why-local-docs)
