# LLMs.txt Generator — Full Site Content
> Complete text content of llmstxtgenerator.dev in a single Markdown file: free llms.txt generator and validator tools, plus every guide, tutorial, comparison and troubleshooting article. Generated automatically from the deployed pages.
> Source: https://llmstxtgenerator.dev/llms-full.txt · Companion index: https://llmstxtgenerator.dev/llms.txt
## Contents
- [LLMs.txt Generator — Create AI-Ready Website Files for Free](https://llmstxtgenerator.dev/)
- [LLMs.txt Generator — Create Spec-Compliant llms.txt Files](https://llmstxtgenerator.dev/generator/)
- [LLMs.txt Validator — Check & Score Your llms.txt File](https://llmstxtgenerator.dev/checker/)
- [Blog — llms.txt Guides, Tutorials & Examples](https://llmstxtgenerator.dev/blog/)
- [What is llms.txt? — The Complete Guide for Developers & Site Owners](https://llmstxtgenerator.dev/guide/)
- [LLMs-full.txt Generator — Full Site Content From Your Sitemap](https://llmstxtgenerator.dev/llms-full-txt-generator/)
- [AGENTS.md vs llms.txt: Two Files, Two Readers (2026)](https://llmstxtgenerator.dev/blog/agents-md-vs-llms-txt/)
- [ai.txt vs llms.txt: Which AI File Should You Ship in 2026?](https://llmstxtgenerator.dev/blog/ai-txt-vs-llms-txt/)
- [Does llms.txt Affect SEO? What Google Actually Says (2026)](https://llmstxtgenerator.dev/blog/does-llms-txt-affect-seo/)
- [How to Create an llms.txt File: Step-by-Step Guide](https://llmstxtgenerator.dev/blog/how-to-create-llms-txt/)
- [How to Validate an llms.txt File: Spec, Links and Delivery (2026)](https://llmstxtgenerator.dev/blog/how-to-validate-llms-txt/)
- [How to Add llms.txt to an Astro Site (Static or Generated, 2026)](https://llmstxtgenerator.dev/blog/llms-txt-astro/)
- [llms.txt Best Practices 2026: The 9 Rules for a File AI Engines Trust](https://llmstxtgenerator.dev/blog/llms-txt-best-practices/)
- [llms.txt Case Studies 2026: Inside the Files of Vue.js, Nuxt, Cloudflare, OpenAI and Anthropic](https://llmstxtgenerator.dev/blog/llms-txt-case-studies/)
- [How to Add llms.txt to Cloudflare Pages (Astro, React, Plain HTML)](https://llmstxtgenerator.dev/blog/llms-txt-cloudflare-pages/)
- [How to Add llms.txt to Docusaurus (Docs Sites, 2026)](https://llmstxtgenerator.dev/blog/llms-txt-docusaurus/)
- [llms.txt Examples: Real Files for Every Site Type](https://llmstxtgenerator.dev/blog/llms-txt-examples/)
- [llms.txt Format Guide — Complete Specification & Structure](https://llmstxtgenerator.dev/blog/llms-txt-format/)
- [How to Add llms.txt to Hugo (Static or Auto-Generated, 2026)](https://llmstxtgenerator.dev/blog/llms-txt-hugo/)
- [How to Add llms.txt to Jekyll (GitHub Pages, 2026)](https://llmstxtgenerator.dev/blog/llms-txt-jekyll/)
- [llms.txt Maintenance: How to Keep Your File From Going Stale (2026)](https://llmstxtgenerator.dev/blog/llms-txt-maintenance/)
- [6 llms.txt Mistakes That Make AI Engines Ignore Your File](https://llmstxtgenerator.dev/blog/llms-txt-mistakes/)
- [How to Add llms.txt to MkDocs (Material for MkDocs, 2026)](https://llmstxtgenerator.dev/blog/llms-txt-mkdocs/)
- [llms.txt for Multilingual Websites: One File or One Per Language? (2026)](https://llmstxtgenerator.dev/blog/llms-txt-multilingual-sites/)
- [How to Add llms.txt to Next.js (App Router, 2026)](https://llmstxtgenerator.dev/blog/llms-txt-nextjs/)
- [llms.txt Not Working? 7 Delivery-Level Fixes (Content-Type, Redirects, WAF, 2026)](https://llmstxtgenerator.dev/blog/llms-txt-not-working/)
- [How to Add llms.txt to Nuxt (public/, Server Routes, nuxt-llms, 2026)](https://llmstxtgenerator.dev/blog/llms-txt-nuxt/)
- [How to Add llms.txt to Shopify: Replace the Default File (2026)](https://llmstxtgenerator.dev/blog/llms-txt-shopify/)
- [llms.txt v2: What Changed in the 2026 Spec Update](https://llmstxtgenerator.dev/blog/llms-txt-v2/)
- [How to Add llms.txt to VitePress (Static or Auto-Generated, 2026)](https://llmstxtgenerator.dev/blog/llms-txt-vitepress/)
- [llms.txt vs llms-full.txt: What's the Difference and Which Do You Need?](https://llmstxtgenerator.dev/blog/llms-txt-vs-llms-full-txt/)
- [llms.txt vs MCP: The Read Layer vs the Action Layer for AI Agents (2026)](https://llmstxtgenerator.dev/blog/llms-txt-vs-mcp/)
- [llms.txt vs robots.txt vs sitemap.xml: What's the Difference?](https://llmstxtgenerator.dev/blog/llms-txt-vs-robots-txt/)
- [llms.txt vs sitemap.xml: Why AI Engines Need a Different File](https://llmstxtgenerator.dev/blog/llms-txt-vs-sitemap-xml/)
- [How to Add llms.txt to WordPress (2026 Guide)](https://llmstxtgenerator.dev/blog/llms-txt-wordpress/)
- [Markdown for Agents vs llms.txt: Content Negotiation or Curated Index? (2026)](https://llmstxtgenerator.dev/blog/markdown-for-agents-vs-llms-txt/)
- [How to Measure Whether AI Engines Read Your llms.txt (Logs, GA4 and Cloudflare, 2026)](https://llmstxtgenerator.dev/blog/measure-llms-txt-traffic/)
- [robots.txt for AI Crawlers in 2026: Allow, Block or Feed GPTBot, ClaudeBot and PerplexityBot](https://llmstxtgenerator.dev/blog/robots-txt-ai-crawlers/)
- [Which AI Engines Actually Read llms.txt? (2026 Crawler Data)](https://llmstxtgenerator.dev/blog/which-ai-engines-read-llms-txt/)
---
# LLMs.txt Generator — Create AI-Ready Website Files for Free
> Source: https://llmstxtgenerator.dev/
Free · No sign-up · Spec-compliant (llmstxt.org)
# Make your website
AI-ready in 60 seconds
**llms.txt** is the file that tells AI engines — ChatGPT, Claude, Perplexity, Gemini — what your site is about, the same way robots.txt guides search crawlers. Generate yours free, validate it, and deploy in minutes.
[Generate your llms.txt](/generator) [Validate an existing file](/checker)
No account needed · Your content never leaves your browser
[
### Generator
Enter your site info, import URLs from a sitemap, and generate a clean llms.txt — titles auto-derived from URLs.
Try it now →](/generator)[
### Validator
Paste or upload your file. Get a 0–100 AI-readiness score with a detailed, line-by-line issue report.
Check a file →](/checker)[
### Learn the spec
What llms.txt is, why it matters for Generative Engine Optimization (GEO), and how to deploy it.
Read the guide →](/guide)
## Learn llms.txt inside out
From the exact format to real-world examples — everything you need to ship a spec-perfect file.
[
### 📐 The Complete Format Guide
Every element of the llmstxt.org v2 spec, element-by-element, with a valid end-to-end example and common mistakes.
Read the spec →](/blog/llms-txt-format/)[
### 📚 Real-World Examples
Complete llms.txt files for SaaS docs, e-commerce, blogs and corporate sites — copy the structure that fits you.
See examples →](/blog/llms-txt-examples/)[
### ⚖️ vs robots.txt & sitemap.xml
Three root files, three different jobs. A 9-dimension comparison and how they work together.
Compare the files →](/blog/llms-txt-vs-robots-txt/)[
### 🛠️ Step-by-Step Tutorial
Create and deploy your llms.txt in 10 minutes — with per-platform instructions for WordPress, Next.js, Hugo and more.
Start the tutorial →](/blog/how-to-create-llms-txt/)[
### 🔌 WordPress Guide
The largest CMS on the web — here's how to add llms.txt via upload, plugins (Yoast, Rank Math) or a child-theme snippet.
WordPress how-to →](/blog/llms-txt-wordpress/)[
### 📖 What is llms.txt?
Start here if you're new: what llms.txt is, why AI engines need it, and the basics of GEO.
Beginner's guide →](/guide)
## What llms.txt looks like
A single Markdown file at your site root. That's it — no schema, no XML, no server config.
```
# Acme Docs
> Developer documentation for the Acme API,
> covering REST, GraphQL, SDKs and best practices.
## Important links
- [Home](https://acme.com/): Documentation home
- [Getting Started](https://acme.com/getting-started): 5-minute guide
- [API Reference](https://acme.com/api): Full REST reference
- [GraphQL](https://acme.com/graphql): GraphQL endpoint & schema
```
## Why it matters
### AI is the new front door
More and more users get answers from ChatGPT, Claude and Perplexity instead of Google. Without llms.txt, AI agents crawl your whole site and guess what matters — often missing your best content.
### A content map for LLMs
llms.txt gives AI engines a curated, structured summary of your pages so they can understand and cite you accurately. That's the foundation of Generative Engine Optimization (GEO).
### Deploys in minutes
Generate → upload to your site root → done. Same effort as robots.txt, with zero dependencies or configuration.
## llms.txt vs robots.txt vs sitemap.xml
Three tiny files, three very different jobs. You probably want all three.
File
Audience
What it does
Format
`robots.txt`
Search engine crawlers
Allows or blocks crawling of specific paths
Plain text rules
`sitemap.xml`
Search engine crawlers
Lists every URL for indexing
XML
llms.txt
AI engines (ChatGPT, Claude, Perplexity…)
Structured content summary for LLM understanding & citation
Markdown
## Give AI engines a map of your site
Generate a spec-compliant llms.txt for your website in under a minute — free, forever.
[Start generating →](/generator) [Validate first](/checker)
---
# LLMs.txt Generator — Create Spec-Compliant llms.txt Files
> Source: https://llmstxtgenerator.dev/generator/
# LLMs.txt Generator
Generate a file that follows the [llmstxt.org](https://llmstxt.org) specification, so ChatGPT, Claude, Perplexity and other AI engines can understand your website.
## Site details
Site name \* (H1, the only required element)
Tagline — optional, recommended. One sentence that tells AI what your site is.
llms-full.txt URL — optional. Reference a full-documentation version (great for docs sites).
## Page links
Add URLs one by one, auto-crawl your whole site, or import from your sitemap.
Auto-crawl & auto-pick — enter your homepage URL; we discover your pages, pull real titles & descriptions, and pick the most important ones, grouped by section
Import from sitemap.xml — paste your sitemap URL to extract every page automatically
Add a link
No links yet — add a URL above or import a sitemap.
## Your llms.txt
✓ Valid Markdown · spec-compliant
**Deployment:** save the file as `llms.txt` and upload it to your site root (next to `robots.txt`). Verify it's live at `https://your-domain.com/llms.txt`, then run it through the [validator](/checker) for a full score.
---
# LLMs.txt Validator — Check & Score Your llms.txt File
> Source: https://llmstxtgenerator.dev/checker/
# LLMs.txt Validator
Paste or upload your `llms.txt` and check it against the [llmstxt.org](https://llmstxt.org) specification — with a **0–100 AI-readiness score** and a detailed issue report.
Paste your llms.txt content
or Upload file
—
AI-readiness score
## Issue report
Errors break spec compliance; warnings are recommendations for better AI readability.
---
# Blog — llms.txt Guides, Tutorials & Examples
> Source: https://llmstxtgenerator.dev/blog/
Blog & Guides
# llms.txt Guides & Examples
Everything you need to ship a spec-compliant llms.txt — from the exact format to real-world examples.
[
Validation 2026-09-20 · 7 min
## How to Validate an llms.txt File: Spec, Links and Delivery (2026)
Validation reads the text of your file, not the response that delivers it — which is why a clean validator run can still leave you with a file nobody reads. The three-layer check: structure and the only element the spec requires, a curl loop that catches 404s, redirects and duplicates, and the path, content-type, status and cache tests that catch the rest.
Read article →](/blog/how-to-validate-llms-txt/)[
AI Policy 2026-09-19 · 7 min
## ai.txt vs llms.txt: Which AI File Should You Ship in 2026?
ai.txt declares permissions — may this media be mined or used for training? llms.txt declares content — here is what the site contains. How the two files differ, where the W3C TDM Reservation Protocol, the TDM-Reservation header and noai fit, the December 2025 German ruling that made machine-readable opt-outs the standard, and the four-signal stack worth shipping.
Read article →](/blog/ai-txt-vs-llms-txt/)[
AI Agents 2026-09-18 · 8 min
## Markdown for Agents vs llms.txt: Content Negotiation or Curated Index? (2026)
llms.txt is a curated index that tells agents what exists; Markdown for Agents is HTTP content negotiation that changes the format one page arrives in. The Feb 2026 data on which agents send Accept: text/markdown, how Cloudflare and Vercel implement the delivery layer, and how Markdown twins + rel="alternate" connect the two.
Read article →](/blog/markdown-for-agents-vs-llms-txt/)[
AI Agents 2026-09-17 · 8 min
## AGENTS.md vs llms.txt: Two Files, Two Readers (2026)
Both are plain Markdown conventions and both get called "the file that tells AI about your project" — but one is read inside your repository and the other is fetched from your website. Side-by-side comparison, what belongs in each, the new link relations in llms.txt v2, the Claude Code holdout and the documented imports that fix it, plus what the ETH Zurich study of AGENTS.md files actually found.
Read article →](/blog/agents-md-vs-llms-txt/)[
Maintenance 2026-09-16 · 7 min
## llms.txt Maintenance: How to Keep Your File From Going Stale (2026)
A published llms.txt decays quietly: renamed URLs, deleted pages, a summary describing a product you repositioned months ago. A three-layer maintenance system — a link check that fails your build, a split between what you generate and what you curate, and a quarterly review list you can finish in ten minutes.
Read article →](/blog/llms-txt-maintenance/)[
Troubleshooting 2026-09-15 · 7 min
## llms.txt Not Working? 7 Delivery-Level Fixes (Content-Type, Redirects, WAF, 2026)
Seven delivery-level reasons a published llms.txt never gets read — a catch-all rewrite answering text/html, rewrite rules that map the path onto robots.txt, redirect chains, case-sensitive paths, WAF and robots.txt blocks, files that never shipped in the build, and a stale edge cache — each with the curl command that proves which one you have.
Read article →](/blog/llms-txt-not-working/)[
Global SEO 2026-09-14 · 7 min
## llms.txt for Multilingual Websites: One File or One Per Language? (2026)
A file covers the URLs under its path and the most specific one wins — that single spec rule is what makes per-language llms.txt files work. Three structures (single root file, root index plus /en/ and /de/ files, root plus Optional section), how to keep it aligned with hreflang and canonicals, a curl-verified table of what Stripe, Cloudflare, Shopify, GitLab and Wikipedia actually ship on September 14, 2026, and a five-minute verification pass.
Read article →](/blog/llms-txt-multilingual-sites/)[
AI Agents 2026-09-13 · 7 min
## llms.txt vs MCP: The Read Layer vs the Action Layer for AI Agents (2026)
llms.txt answers "what is on this site?", MCP answers "what can I do here?" — and the 2026-07-28 MCP revision made the action layer stateless. How the two protocols differ, why MCP has no discovery of its own, where WebMCP and Lighthouse's Agentic Browsing audits fit, and the five-step adoption order for the next quarter.
Read article →](/blog/llms-txt-vs-mcp/)[
Measurement 2026-09-12 · 8 min
## How to Measure Whether AI Engines Read Your llms.txt (Logs, GA4 and Cloudflare, 2026)
A three-layer measurement framework for llms.txt: count direct /llms.txt fetches in your access logs, track AI crawler hits on the pages your file lists, and read AI referral sessions in the GA4 AI Assistant channel (added May 13, 2026) plus Cloudflare AI Crawl Control — with the exact log commands, the two GA4 caveats, and a 30-day benchmark table.
Read article →](/blog/measure-llms-txt-traffic/)[
AI Crawlers 2026-09-11 · 7 min
## robots.txt for AI Crawlers in 2026: Allow, Block or Feed GPTBot, ClaudeBot and PerplexityBot
Training, search-index and user-triggered: the three classes of AI crawler decide what blocking actually costs you. The 2026 token reference for GPTBot, OAI-SearchBot, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended and CCBot, two ready-to-paste robots.txt policies, Cloudflare Content Signals, and where your llms.txt file takes over.
Read article →](/blog/robots-txt-ai-crawlers/)[
Case Studies 2026-09-10 · 7 min
## llms.txt Case Studies 2026: Inside the Files of Vue.js, Nuxt, Cloudflare, OpenAI and Anthropic
Six production llms.txt files, curl-verified on September 10, 2026: how Vue.js, Nuxt, Cloudflare Developers, OpenAI and Anthropic structure their files — hierarchical tables of contents, markdown twins, subpath hubs and llms-full.txt — plus the five patterns worth copying.
Read article →](/blog/llms-txt-case-studies/)[
Tutorial 2026-09-09 · 7 min
## How to Add llms.txt to Nuxt (public/, Server Routes, nuxt-llms, 2026)
Nuxt ships no llms.txt output, but its public/ directory serves a static file at the site root in 60 seconds, a server/routes handler generates one per request, and NuxtLabs' nuxt-llms module prerenders llms.txt and llms-full.txt at every build — with Nuxt Content 3.2+ integration. Includes nuxt.com's own file, curl-verified.
Read article →](/blog/llms-txt-nuxt/)[
E-commerce 2026-09-08 · 6 min
## How to Add llms.txt to Shopify: Replace the Default File (2026)
Shopify now auto-generates an llms.txt for every store at yourstore.com/llms.txt — but the default reads like Shopify's pitch, not yours. Override it with a templates/llms.txt.liquid theme file (the robots.txt.liquid mechanism, confirmed on the Shopify developer community): what to include, exact admin clicks, a ready-to-copy store sample and a curl check.
Read article →](/blog/llms-txt-shopify/)[
Tutorial 2026-09-07 · 7 min
## How to Add llms.txt to Jekyll (GitHub Pages, 2026)
Jekyll has no built-in llms.txt support, but its root-file rules make it easy: a static file in 60 seconds, a Liquid template that re-lists every post at every build (no plugin, so GitHub Pages accepts it), or the jekyll-llms-output plugin for llms.txt + llms-full.txt. Includes real config, the GitHub Pages plugin whitelist caveat and a curl check of a live Jekyll site.
Read article →](/blog/llms-txt-jekyll/)[
Tutorial 2026-09-06 · 7 min
## How to Add llms.txt to VitePress (Static or Auto-Generated, 2026)
No native llms.txt support in VitePress, but public/ ships a static file in 60 seconds and vitepress-plugin-llms — used in production by Vue.js, Vite and Vitest — generates llms.txt, llms-full.txt and per-page Markdown at every build. Real config plus a curl check.
Read article →](/blog/llms-txt-vitepress/)[
Tutorial 2026-09-05 · 6 min
## How to Add llms.txt to MkDocs (Material for MkDocs, 2026)
MkDocs has no built-in llms.txt support, but docs/ files are copied verbatim to the site root and the mkdocs-llmstxt plugin generates a spec-compliant file at every build — with .md page versions, llms-full.txt and Read the Docs subpath support. Includes real mkdocs.yml config and a curl check.
Read article →](/blog/llms-txt-mkdocs/)[
Tutorial 2026-09-04 · 7 min
## How to Add llms.txt to Hugo (Static or Auto-Generated, 2026)
Hugo has no built-in llms.txt support as of 2026, but its static/ folder and custom output formats make it one of the easiest SSGs to add one: a hand-written file in 60 seconds, or a file auto-generated from your content at build time. Includes real hugo.toml config and a curl check.
Read article →](/blog/llms-txt-hugo/)[
Tutorial 2026-09-03 · 6 min
## How to Add llms.txt to Docusaurus (Docs Sites, 2026)
No official plugin yet, but docs sites are llms.txt's best use case: ship a static file in static/ in 60 seconds, or generate llms.txt and llms-full.txt from your docs with the docusaurus-plugin-llms community plugin — plus a checklist and curl verification.
Read article →](/blog/llms-txt-docusaurus/)[
Tutorial 2026-09-02 · 6 min
## How to Add llms.txt to an Astro Site (Static or Generated, 2026)
Astro has no llms.txt file convention, but public/ files and src/pages endpoints make it trivial: a static public/llms.txt, a build-time .ts endpoint, or the astro-slop integration — with examples and a curl check.
Read article →](/blog/llms-txt-astro/)[
Tutorial 2026-09-01 · 7 min
## How to Add llms.txt to Next.js (App Router, 2026)
No native llms.txt convention in Next.js yet — here is how to ship one today: a static public/llms.txt file, a dynamic app/llms.txt route handler, or ISR, plus verification.
Read article →](/blog/llms-txt-nextjs/)[
Research 2026-08-31 · 7 min
## Which AI Engines Actually Read llms.txt? (2026 Crawler Data)
Two years of crawler audits show 97% of llms.txt files get zero requests — here is the evidence from 500M+ AI bot requests, Google's on-the-record stance, and the realistic playbook.
Read article →](/blog/which-ai-engines-read-llms-txt/)[
Guide 2026-08-30 · 6 min
## 6 llms.txt Mistakes That Make AI Engines Ignore Your File
Placement, format, links, curation, tone and metrics: the six errors that get your llms.txt skipped — plus a five-minute verification checklist.
Read article →](/blog/llms-txt-mistakes/)[
Format & Spec 2026-08-29 · 7 min
## llms.txt v2: What Changed in the 2026 Spec Update
The first major revision of the llmstxt.org spec: two markdown URL forms, officially defined subpath files, rel=alternate and rel=describedby link relations, and what happened to llms\_txt2ctx.
Read article →](/blog/llms-txt-v2/)[
Guide 2026-08-27 · 6 min
## llms.txt Best Practices 2026: The 9 Rules for a File AI Engines Trust
Format, curation, and maintenance rules that separate an llms.txt agents actually use from one they ignore — with a complete reference example.
Read article →](/blog/llms-txt-best-practices/)[
SEO 2026-08-27 · 5 min
## Does llms.txt Affect SEO? What Google Actually Says (2026)
Google says llms.txt has no effect on rankings. Here is the evidence, the nuance, and where the file genuinely pays off for AI visibility.
Read article →](/blog/does-llms-txt-affect-seo/)[
Explainer 2026-08-27 · 5 min
## llms.txt vs sitemap.xml: Why AI Engines Need a Different File
sitemap.xml lists every URL for search engines; llms.txt curates the pages that matter for AI agents. When you need both and how they work together.
Read article →](/blog/llms-txt-vs-sitemap-xml/)[
Tutorial 2026-08-27 · 5 min
## How to Add llms.txt to Cloudflare Pages (Astro, React, Plain HTML)
Publish llms.txt on Cloudflare Pages in 5 minutes — static file, dashboard upload, or a Pages Function for dynamic generation — plus verification.
Read article →](/blog/llms-txt-cloudflare-pages/)[
Explainer 2026-08-27 · 5 min
## llms.txt vs llms-full.txt: What's the Difference and Which Do You Need?
The curated index vs the full Markdown dump: what each is for, when you need llms-full.txt, and how to link them correctly.
Read article →](/blog/llms-txt-vs-llms-full-txt/)[
Format & Spec 2026-08-15 · 9 min
## The Complete llms.txt Format Guide
Every element of the llmstxt.org v2 spec, element-by-element, with a valid end-to-end example and the mistakes people make.
Read article →](/blog/llms-txt-format/)[
Tutorial 2026-08-15 · 8 min
## How to Create an llms.txt File: Step-by-Step Guide
Create your llms.txt in 10 minutes: pick your 5–15 best pages, format them per the spec, publish at your site root, and validate.
Read article →](/blog/how-to-create-llms-txt/)[
Examples 2026-08-15 · 7 min
## llms.txt Examples: Real Files for Every Site Type
Complete llms.txt files for SaaS docs, e-commerce, blogs and corporate sites — copy the structure that fits your site.
Read article →](/blog/llms-txt-examples/)[
Explainer 2026-08-15 · 6 min
## llms.txt vs robots.txt vs sitemap.xml: What's the Difference?
What each root-level file does, who reads it, and how the three work together for AI engines and search crawlers.
Read article →](/blog/llms-txt-vs-robots-txt/)[
WordPress 2026-08-15 · 6 min
## How to Add llms.txt to WordPress (2026 Guide)
Three ways to ship llms.txt on WordPress — upload the file, use a plugin, or add it with code — plus the pitfalls that break it.
Read article →](/blog/llms-txt-wordpress/)[
Beginner Guide 2026-08-15 · 4 min
## What is llms.txt? The Complete Guide
What llms.txt is, why it matters for AI engines like ChatGPT and Claude, and how it fits with robots.txt and sitemap.xml.
Read article →](/guide/)
## Need it done for you?
Paste your sitemap.xml into the generator and get a spec-compliant llms.txt in seconds.
[Open the Generator →](/generator)
---
# What is llms.txt? — The Complete Guide for Developers & Site Owners
> Source: https://llmstxtgenerator.dev/guide/
Complete guide · Updated for llmstxt.org v2
# What is llms.txt?
The complete guide
llms.txt is the "content map" that AI engines read to understand your website. This guide explains what it is, why it matters, and how to create and deploy it in minutes.
## What is llms.txt?
**llms.txt** is a Markdown file placed in the root of your website, proposed in 2024 by AI researcher [Jeremy Howard](https://en.wikipedia.org/wiki/Jeremy_Howard_\(entrepreneur\)) as an open convention for large language models (LLMs). Just as robots.txt gives search-engine crawlers directions, llms.txt gives AI engines a **structured summary and index** of your site's content — so they can understand and cite it accurately instead of guessing.
The canonical specification lives at [llmstxt.org](https://llmstxt.org).
## Why llms.txt matters now
AI is becoming a primary entry point to the web. Users increasingly ask ChatGPT, Claude, Perplexity and Gemini questions instead of typing keywords into Google. These AI engines need an efficient, reliable way to understand your site:
- **Without llms.txt:** AI agents crawl your entire site, guess at what's important, and may miss or misread your core content.
- **With llms.txt:** AI reads your curated content map directly and cites your pages accurately — the foundation of **Generative Engine Optimization (GEO)**.
## The llms.txt format
An llms.txt file is deliberately simple — plain Markdown. A minimal valid file looks like this:
```
# Site Name
> One-sentence description of the site (optional).
## Important links
- [Page title](https://your-site.com/page): Short description
- [Another page](https://your-site.com/other): Another description
```
The rules per the llmstxt.org spec:
- **The H1 heading (`# Site Name`) is the only mandatory element.** Everything else is optional.
- Links must be **absolute URLs** (`https://…`) — no relative paths.
- Group links under **H2 headings** (`## Section`); use Markdown list items for links.
- Optionally add a **blockquote tagline** (`> …`) describing the site.
- For documentation-heavy sites, you can reference a full version at `/llms-full.txt`.
## llms.txt vs robots.txt vs sitemap.xml
They complement each other — most sites should have all three.
File
Audience
Purpose
`robots.txt`
Search engine crawlers
Allows / blocks crawling
`sitemap.xml`
Search engine crawlers
Lists all URLs for indexing
llms.txt
AI engines (ChatGPT, Claude, Perplexity…)
Structured content summary for LLM understanding & citation
## How to create an llms.txt
1. Use [our free generator](/generator) — enter your site name, tagline and URLs (or import a sitemap) and it produces a spec-compliant file instantly.
2. Download the file and upload it to your website root, next to `robots.txt`.
3. Verify it's publicly accessible at `https://your-domain.com/llms.txt`.
4. Run it through the [validator](/checker) to confirm the structure and get an AI-readiness score.
5. Update it whenever your site structure changes significantly.
## Frequently asked questions
### Does llms.txt affect my Google ranking?
No, not directly. Traditional search engines like Google do not use llms.txt. Its value is in how AI engines (ChatGPT, Claude, Perplexity, etc.) understand and cite your content — the GEO channel.
### How often should I update it?
Whenever your site structure changes in a meaningful way — new sections, renamed pages, removed content. For most sites, reviewing it monthly is plenty.
### Is llms.txt an official standard?
Not yet formally. It's a community-driven open convention that isn't standardized by the W3C or IETF, but adoption is growing quickly — tools, frameworks and AI products increasingly recognize it.
### Where should the file live?
In the root of your website, exactly like `robots.txt`. The canonical URL is `https://your-domain.com/llms.txt`. AI engines look for it there first.
## Keep reading
[
### 📐 The Complete Format Guide
Every element of the v2 spec explained, with a valid end-to-end example.
](/blog/llms-txt-format/)[
### 📚 Real-World Examples
Copy-paste llms.txt files for SaaS, e-commerce, blogs and corporate sites.
](/blog/llms-txt-examples/)
## Ready to get started?
Generate your first llms.txt in under a minute — free, no sign-up.
[⚡ Generate your llms.txt](/generator)
---
# LLMs-full.txt Generator — Full Site Content From Your Sitemap
> Source: https://llmstxtgenerator.dev/llms-full-txt-generator/
Paid tool · Full-content knowledge package
# LLMs-full.txt Generator
Give us your `sitemap.xml` and we crawl your pages, extract the real text and build a single **llms-full.txt** knowledge file that ChatGPT, Claude, Perplexity and Gemini can read directly — not just a list of links.
## Step 1 — Scan your sitemap
We read your sitemap, list every page with its real title, and show how much content we can extract.
## Step 2 — Choose pages
Content mode
Summary — ~300 words per page (lean file) Full content — up to 2,000 words per page
## Step 3 — Preview & download
### Unlock the full file — one-time $1
The preview shows the first 2 pages. Unlock to download the complete llms-full.txt with every page you selected.
License key — sent to you after purchase
## What is llms-full.txt?
An **llms-full.txt** file is the deep-content companion to a standard [llms.txt](/blog/llms-txt-vs-llms-full-txt/). Where llms.txt gives AI engines a curated list of links, llms-full.txt contains the actual text of your pages — articles, docs, product copy — in one Markdown file. Models can then answer questions about your site accurately without crawling every URL.
llms.txt (free)
llms-full.txt (this tool)
**Content**
Titles, links, short descriptions
Full page text as Markdown
**Typical size**
2–10 KB
100 KB – 2 MB
**Best for**
Discoverability
Answer accuracy and depth
**Recommended**
Every site
Docs, blogs, content-heavy sites
## How the generator works
1. **Scan** — we fetch your sitemap.xml (including sitemap index files) and list every page with its real title and content size.
2. **Select** — untick anything that shouldn't be in the AI knowledge file: tag pages, legal boilerplate, thin pages.
3. **Generate** — we crawl each selected page, strip navigation, footers and scripts, and convert the main content to clean Markdown.
4. **Publish** — download the file, upload it to your site root as `/llms-full.txt`, and link it from your llms.txt.
## Already have an llms.txt?
Validate it first with our free [llms.txt validator](/checker/) (AI Readiness Score, line-level fixes), or build the index file with the free [llms.txt generator](/generator/) — it auto-crawls your site or imports your sitemap. Our own site publishes both files: [/llms.txt](/llms.txt) and [/llms-full.txt](/llms-full.txt).
Privacy note: we fetch the pages you submit server-side to build your file. Submitted URLs are processed for that request only and are not stored. See our [Privacy Policy](/privacy-policy/).
---
# AGENTS.md vs llms.txt: Two Files, Two Readers (2026)
> Source: https://llmstxtgenerator.dev/blog/agents-md-vs-llms-txt/
Comparison · AI agents · Updated 2026
# AGENTS.md vs llms.txt: Two Files, Two Readers (2026)
Both are plain Markdown, both are conventions rather than enforced standards, and both are frequently described as "the file that tells AI about your project." They have almost nothing else in common. One is read inside your repository, the other is fetched from your website.
## The Short Answer
AGENTS.md
llms.txt
Reader
Coding agents — OpenAI Codex, Jules, Cursor, Aider, Zed, Warp and others
External AI crawlers, chat assistants with search, IDE agents browsing docs
Lives in
Repository root, plus nested copies per package
Site root, or any path such as `/docs/llms.txt`
Transport
A git checkout
A plain HTTP GET, cacheable at the edge
Contains
Build and test commands, code style, boundaries, PR rules
Site name, one-line summary, curated links to the pages that matter
Governance
Open format, stewarded by the Agentic AI Foundation under the Linux Foundation
Community proposal by Jeremy Howard, documented at llmstxt.org
## AGENTS.md Is a Repository Contract
The [agents.md](https://agents.md/) project defines AGENTS.md as a "README for agents" — the extra, sometimes detailed context that would clutter a README written for humans: exact build steps, test invocations, conventions that differ from tool defaults, and the files an agent should never touch. The site reports adoption across more than 60,000 open-source projects.
Real files look like real engineering notes. The root AGENTS.md in [openai/codex](https://github.com/openai/codex/blob/main/AGENTS.md) runs about 320 lines and is mostly Rust-specific constraints and sandbox warnings. Apache Airflow's copy opens with an SPDX license header before its first instruction. Shape, not length, is what carries over:
```
# AGENTS.md
## Setup commands
- Install deps: `pnpm install`
- Run tests: `pnpm test`
## Code style
- TypeScript strict mode
- Use functional patterns where possible
## Boundaries
- Never edit files under `src/generated/`
- Do not modify the sandbox environment-variable checks
```
In a monorepo, place a second AGENTS.md inside each package. Agents read the nearest file in the directory tree, so the closest one takes precedence and every subproject can ship tailored instructions.
## llms.txt Is a Site Contract
The [llms.txt](https://llmstxt.org/) proposal was updated to v2 in August 2026, and the v2 text is explicit about placement: the file can sit at the site root or at any path, covering the pages under that path. It also recommends publishing a clean Markdown twin for each page — either with `.md` appended (`page.html.md`) or with the extension replaced (`page.md`) — so the links inside your file point at content an agent can read without stripping HTML.
```
# Acme Docs
> API documentation for Acme's payments platform, covering charges,
> refunds and webhooks.
## Guides
- [Quickstart](https://docs.acme.com/quickstart.md): first charge in 5 minutes
- [Refunds](https://docs.acme.com/refunds.md): partial and full refund flows
## Optional
- [Changelog](https://docs.acme.com/changelog.md)
```
The same v2 document leans on HTTP to help clients find both artefacts: a `rel="alternate" type="text/markdown"` link points at the Markdown version of a page, and `rel="describedby"` points at the llms.txt file that covers it — expressed either as HTML elements or as a server-level `Link:` response header, which requires no page changes. Chrome's Lighthouse agentic browsing checks now audit sites for the file, and the labs themselves ship one: OpenAI, Anthropic and Gemini all publish llms.txt for their developer documentation.
## The Discovery Difference
llms.txt has two discovery channels: the `/llms.txt` convention and the `describedby` link relation. AGENTS.md has exactly one — the folder it sits in. There is no registry, no HTTP header, no sitemap entry. That single fact explains the nesting rule, the per-package files, and why the format's own documentation stresses that the closest file wins.
## The Claude Code Exception
One major agent still does not read AGENTS.md natively. Anthropic's memory documentation states it plainly: "Claude Code reads CLAUDE.md, not AGENTS.md." The documented fixes are cheap, so a repository does not have to choose between tools:
- Create a `CLAUDE.md` that imports it — a line containing `@AGENTS.md` expands the file into Claude's context at session start, and Claude-specific rules can follow below it.
- Or symlink instead of duplicating: `ln -s AGENTS.md CLAUDE.md` (on Windows a symlink needs Administrator rights or Developer Mode, so use the import).
- Or run `/init` once, which reads AGENTS.md when `CLAUDE_CODE_NEW_INIT=1` is set, and `/import` (Claude Code v2.1.213+) which copies a supported agent's configuration into Claude Code.
Two size limits are worth knowing because they are enforcement, not advice: the docs target under 200 lines per memory file and say Claude Code skips a file over 4 MiB. Imported files still load at launch, so splitting into imports organises a file without shrinking the context it consumes.
## What the Evidence Says About Context Files
An ETH Zurich and LogicStar.ai preprint, ["On the Impact of AGENTS.md Files on the Efficiency of AI Coding Agents"](https://arxiv.org/abs/2601.20404) (revised June 2026), measured generated and developer-written context files against real GitHub pull requests. The result is unromantic: context files consistently increased the number of steps agents needed, developer-written files produced a marginal gain in task success, and LLM-generated files were marginally negative.
The practical rule that follows is the same one that keeps llms.txt useful: write down only what the reader cannot infer. For a repository that means exact commands, non-obvious constraints and past failures. For a site it means a short summary and a curated list of the handful of pages that answer questions — not a mirror of your sitemap. Curation is the whole product in both cases.
## Which One Do You Actually Need?
- **Public docs site plus a repo:** both files, and keep them separate — never paste an AGENTS.md into llms.txt or the reverse.
- **Marketing site, no code:** llms.txt only. Draft it from your sitemap with the [generator](/generator) and validate before publishing with the [checker](/checker).
- **Private repository, no public site:** AGENTS.md alone. There is nothing for external crawlers to read.
- **Open-source project with hosted docs:** both — AGENTS.md for contributors' agents, llms.txt and its Markdown twins for everyone else's.
Related reading: how llms.txt differs from the action layer in [llms.txt vs MCP](/blog/llms-txt-vs-mcp/), the [nine rules for a file AI engines trust](/blog/llms-txt-best-practices/), and which assistants actually fetch it, per [which AI engines read llms.txt](/blog/which-ai-engines-read-llms-txt/).
---
# ai.txt vs llms.txt: Which AI File Should You Ship in 2026?
> Source: https://llmstxtgenerator.dev/blog/ai-txt-vs-llms-txt/
Comparison · AI permissions · Updated 2026
# ai.txt vs llms.txt: Which AI File Should You Ship in 2026?
Both files sit at your site root, both are plain text, and both get described as "the file that tells AI what to do with my site." They are not interchangeable. ai.txt is a permission file that restricts use; llms.txt is an index that invites reading. Mixing them up is the most common mistake in the AI-files conversation.
## The Short Answer
ai.txt
llms.txt
Declares
Permissions: may this content be mined or used for commercial AI training?
Content: here is what the site contains and which pages matter
Direction
Restrictive — an opt-out at the media level
Invitational — an opt-in to a better answer
Origin
Spawning AI, distributed through its API to partners including Hugging Face and Stability AI
The Answer.AI proposal published at llmstxt.org
Read when
At the moment media is downloaded, including downloads triggered by old dataset links
When an AI reader chooses to fetch your index before browsing
Legal weight
Pitched as a machine-readable opt-out aligned with the EU DSM Article 4 TDM exception
None. It is a convention, not a standard, and never a rights reservation
## What ai.txt Actually Does
ai.txt lives in your root directory and sets machine-readable permissions for **commercial text and data mining**. Spawning's generator exposes five media types — text, images, audio, video and code — each set to block or allow, and the default posture is opt-out.
The interesting design decision is the timing. A robots.txt file is read when a crawler walks your site; ai.txt is read when your media is downloaded. That difference matters if your images already sit inside a public dataset: adding a robots.txt today cannot retroactively remove the links, but a permission file checked at download time can affect what happens the next time someone trains from those links.
What ai.txt cannot do is reach copies of your work hosted on domains you do not control, and it only binds tools that actually check it. Treat it as a consent layer, not a lock.
## What llms.txt Actually Does
llms.txt is a Markdown file at your site root that opens with an H1 site name, an optional blockquote summary, and a list of curated links with one-line descriptions. The 2026 revision of the spec added two Markdown link forms, defined subpath files, and two link relations — `rel="alternate"` and `rel="describedby"`.
There is nothing in that spec about permission. llms.txt does not block a crawler, does not reserve rights, and does not change rankings. Its entire job is to shorten the path between an AI reader and the pages that answer questions. If you want the restriction side of the story, that lives in [the robots.txt token reference for AI crawlers](/blog/robots-txt-ai-crawlers/).
## The Legal Layer: TDMRep and Machine-Readable Opt-Outs
The reason "just write it in your terms of service" stopped being advice: a German appeals court decision handed down in December 2025 (OLG Hamburg, 5 U 104/24) held that natural-language opt-outs buried in terms of use are insufficient, because a machine cannot read them. What counts is a machine-readable reservation.
The W3C Community Group specification for that is the TDM Reservation Protocol. It publishes a JSON file at `/.well-known/tdmrep.json` where a reservation flag of 1 reserves rights under CDSM Article 4(3), and it also defines a `TDM-Reservation` HTTP response header as a header-only fallback:
```
{
"tdm-reservation": 1,
"tdm-policy": "https://example.com/ai-policy.html"
}
```
One property in that file has no robots.txt equivalent: `tdm-policy` points at a URL where the actual policy lives, so the machine-readable signal and the human-readable explanation stay in sync. The community group's own documentation is explicit that robots.txt remains the right place to signal an _opt-in_ for search, and TDMRep is the opt-out that sits beside it — which is exactly what you would expect from a site that wants AI citations but not AI training.
## Where noai and X-Robots-Tag Fit
The `noai` and `noimageai` directives, usually delivered as an `X-Robots-Tag` response header, are the third signal and the least formalised. They are not part of an IETF or W3C standard; the TDMRep group notes that `noai` was proposed but does not cover the full range of mining practices publishers want to address. Use them as an extra header on responses, not as your primary reservation.
- **robots.txt** (RFC 9309) — crawl and search permissions, named AI tokens such as `GPTBot` or `ClaudeBot`.
- **tdmrep.json** — purpose-level rights reservation under CDSM Article 4(3).
- **ai.txt** — media-type permissions for commercial mining, read at download time.
- **X-Robots-Tag: noai** — supplementary per-response signal, no standard behind it.
## What to Ship: A Decision List
- **You want to be cited in AI answers:** llms.txt plus a sitemap. Draft it from your sitemap with the [generator](/generator) and run the result through the [checker](/checker) before publishing. Add nothing about permissions to this file.
- **You publish original media or proprietary text:** add ai.txt or tdmrep.json. This is the decision that actually changes what happens to your content, so make it deliberately rather than by default.
- **You are EU-facing and commercially minded:** lead with `/.well-known/tdmrep.json` and keep robots.txt as the search opt-in, per the community group's own guidance.
- **You run a documentation site:** llms.txt is doing real work here, and permission files are largely irrelevant. Keep the index curated and the [expectations about which engines actually fetch it](/blog/which-ai-engines-read-llms-txt/) realistic.
## Verifying All Four Signals in Five Minutes
Every file in this stack is a plain HTTP GET, so verification is one command per file. Replace the host and check the status line and the body — a 200 with an HTML error page is the classic failure mode:
```
curl -sI https://example.com/llms.txt | head -1
curl -sI https://example.com/ai.txt | head -1
curl -s https://example.com/.well-known/tdmrep.json
curl -sI https://example.com/ | grep -i 'x-robots-tag\|tdm-reservation'
```
If the first line comes back as `text/html`, your rewrite rules are serving your catch-all page instead of the file — a delivery problem, not a content problem, and the same one that breaks most llms.txt launches. The [robots.txt vs llms.txt breakdown](/blog/llms-txt-vs-robots-txt/) covers which file belongs in which job if you are still deciding what to publish at all.
---
# Does llms.txt Affect SEO? What Google Actually Says (2026)
> Source: https://llmstxtgenerator.dev/blog/does-llms-txt-affect-seo/
SEO myth-busting · AI visibility · Updated 2026
# Does llms.txt Affect SEO? What Google Actually Says
Short answer: **no — not for Google rankings.** But that's not the whole story. Here's the evidence, the nuance, and where llms.txt genuinely pays off in 2026.
## Google's Official Position
Google has been consistent: it does not use llms.txt for ranking, indexing, or crawling decisions. In 2026, Google's search advocate John Mueller addressed the file directly, likening it to the `meta keywords` tag — a file that sounds authoritative but is outside Google's systems entirely. The official line is that Google will not treat llms.txt as a ranking signal.
So if your goal is purely "higher Google rankings," llms.txt will not move the needle. No amount of H1 optimization or absolute URLs in the file changes Googlebot's behavior.
## Why the Confusion Exists
llms.txt sounds like robots.txt, which _does_ affect Google. It sits in the same root directory, it's the same kind of plain-text file, and it's marketed as "the file for AI." The naming and placement create an intuitive leap — "if robots.txt controls crawlers and this is for AI, Google must care" — that the technical reality doesn't support.
## Where llms.txt Actually Matters
Google is one consumer of your content. AI engines are increasingly the others:
- **ChatGPT / Claude / Perplexity / Gemini browsing** — these products' agents check for llms.txt and use it to decide what to read about your site. When they answer "is X good for Y?", your curated index improves the odds the answer reflects your best content.
- **AI-native discovery** — the major AI companies publish their own llms.txt files and have publicly supported the spec, making it an emerging standard for how sites describe themselves to agents.
- **Developer tooling** — Chrome's Lighthouse (13.3+) includes an Agentic Browsing audit that checks for llms.txt, and documentation platforms like Mintlify and GitBook generate one automatically. It's becoming a baseline expectation for technical sites.
The measurable reality is still modest: third-party analyses (e.g. Limy's study of billions of AI crawler requests) found most AI crawlers fetch HTML pages directly rather than llms.txt, and only about 10% of sites have adopted the file. It's an emerging signal with real but limited current penetration.
## The Honest 2026 Playbook
1. **Don't expect rankings.** Build llms.txt for AI assistants, not for Google.
2. **Do the cheap, real SEO first** — crawlable pages, fast Core Web Vitals, one clear topic per URL, and an accurate sitemap. That's what feeds Google's AI Overviews anyway.
3. **Ship llms.txt anyway** — it takes 10 minutes, costs nothing, and positions your site for the channel that's growing: AI-referenced answers. Use our [generator](/generator) to create it and the [validator](/checker) to keep it correct.
4. **Keep it current** — a stale llms.txt pointing at old pages is worse than none. Regenerate after major content changes.
## The Bottom Line
llms.txt will not improve your Google rankings, and anyone selling it as a "Google ranking hack" is wrong. But it's a legitimate, low-cost piece of AI visibility infrastructure — the same way you'd claim your brand on emerging platforms early. Combined with solid [sitemap hygiene](/blog/llms-txt-vs-sitemap-xml/) and genuinely useful content, it makes your site easier for AI engines to represent accurately. That's a bet worth taking even when the payoff is a year out.
---
# How to Create an llms.txt File: Step-by-Step Guide
> Source: https://llmstxtgenerator.dev/blog/how-to-create-llms-txt/
llms.txt tutorial · 10-minute setup
# How to Create an llms.txt File: Step-by-Step Guide
An **llms.txt** file tells AI engines — ChatGPT, Claude, Gemini, Perplexity, and coding agents — exactly which pages of your site are worth reading. This llms.txt tutorial walks you from zero to a live, validated file in about **10 minutes**, no matter what your site is built with.
## Why Create an llms.txt File in the First Place?
Every day, AI assistants answer questions with content from your site. Without an llms.txt file, those models have to guess what to crawl, what to skip, and which of your pages actually matter. The result: they often land on your marketing pages instead of your documentation, or miss your best content entirely. An llms.txt file fixes that by giving agents a small, curated Markdown index of your most important pages — the LLM equivalent of `robots.txt`, but for what _should_ be read instead of what shouldn't.
The payoff is concrete: Chrome's Lighthouse now audits sites for an llms.txt file, documentation platforms like Mintlify and GitBook generate one automatically, and major AI labs publish their own. Sites with a valid file get referenced more accurately — and more often — in AI-generated answers. Creating one is one of the cheapest SEO moves available: a single text file, no plugins, no redesign. Here's the exact process.
## Step 1: Pick Your Site Name and Write a One-Line Summary
The first two lines of your file are the most important. Line one is the **H1**: your site or project name, prefixed with a single `#`. Line two is a **blockquote** (a line starting with `>`) summarizing what the site is and who it's for, in one or two sentences. This is the only context an agent gets before deciding whether to read further, so be concrete: what you do, for whom, and what makes you different.
```
# Acme Corp
> Acme Corp builds time-tracking software for small teams. Free plan available for up to 10 users.
```
Keep the summary short and factual — a one-sentence pitch works better than marketing fluff. The H1 here is the only element the llmstxt.org v2 specification marks as required, so never skip it.
## Step 2: List Your 5–15 Most Important Pages
An llms.txt file is a curated index, not a sitemap. Resist the urge to list everything — a well-chosen 5–15 links beats 200 links that drown agents in noise. Include pages an AI should actually quote: product overviews, documentation, API references, pricing, setup guides. Skip boilerplate, thin blog posts, login pages, and anything behind an auth wall. Write down the full absolute URLs for your candidates:
```
# Candidate pages
- https://acme.dev/product.md → Product overview
- https://acme.dev/pricing.md → Plans and pricing
- https://acme.dev/docs/quickstart.md → Quick start guide
- https://acme.dev/docs/api.md → REST API reference
- https://acme.dev/blog/ → Skip: low value for AI answers
```
Prefer clean Markdown versions of your pages when they exist (most documentation sites serve `.md` variants of every page — same URL with `.md` appended). Plain HTML pages work too, but Markdown is cheaper for LLMs to parse and less likely to be truncated.
## Step 3: Organize Everything in the llms.txt Format
Now turn your list into the spec structure: H1 at the top, blockquote summary, then `##` sections (H2 headings) each containing a Markdown link list. Every link is written as `[text](url): short description` — the description after the colon is optional but recommended, because it tells agents what's behind each link without fetching it. Use a `## Optional` section at the end for secondary links agents can skip when context is short. Here's the complete skeleton:
```
# Acme Corp
> Acme Corp builds time-tracking software for small teams.
## Important links
- [Product overview](https://acme.dev/product.md): What Acme does and why teams choose it
- [Pricing](https://acme.dev/pricing.md): Plans, per-seat pricing, and the free tier
## Documentation
- [Quick start](https://acme.dev/docs/quickstart.md): Set up Acme in 10 minutes
- [API reference](https://acme.dev/docs/api.md): Complete REST API reference
## Optional
- [Changelog](https://acme.dev/changelog.md): Release notes for all versions
```
Three rules keep the file machine-readable: all URLs must be absolute (`https://…`, never `/relative/paths`), there is no comment syntax (a line starting with `#` is a heading, so don't use hash comments), and H2 sections must contain lists of links — not prose. Prose belongs in the free-form notes area between the blockquote and the first `##`.
## Step 4: Generate the File Automatically (or Write It by Hand)
If you've already followed Steps 1–3, writing the file by hand takes two minutes — it's plain Markdown, nothing more. That said, hand-writing is exactly where mistakes creep in: missing H1s, relative URLs, sections without descriptions. If you want a spec-perfect file in seconds, paste your site URL into our [free llms.txt generator](/generator) — it crawls your most important pages, builds the formatted file for you, and lets you tweak the section names and descriptions before downloading:
```
# Generated by llmstxtgenerator.dev
# Acme Corp
> Acme Corp builds time-tracking software for small teams.
## Important links
- [Product overview](https://acme.dev/product.md): ...
- [Pricing](https://acme.dev/pricing.md): ...
```
Either path produces the same plain-text file. The generator is fastest for large or frequently changing sites; writing by hand gives you full control over tone and link order. For a first file, we recommend the generator, then manual edits — the combination of speed and control is hard to beat. Save the result as `llms.txt` (lowercase, no extension, UTF-8 encoding).
## Step 5: Put the File at Your Website's Root
The file must be publicly reachable at `https://your-domain.com/llms.txt`. Exactly where that file lives in your project depends on your platform — here's where to place it for the most common setups:
Platform
Where to place llms.txt
**WordPress**
Upload to `public_html/llms.txt` via the file manager or FTP — the web root, not your theme folder.
**Next.js**
Add `llms.txt` to the `public/` folder; Next.js copies it to the site root on every build.
**Hugo**
Put the file in `static/llms.txt` — Hugo copies everything in `static/` verbatim to the site root.
**Shopify**
Upload the file via Files, then create a URL redirect from `/llms.txt` to the uploaded file URL so the standard path works.
**Cloudflare Pages**
Add `llms.txt` to your build output folder (e.g. `public/` for Astro) — it ships to the site root automatically.
On a static host, the deployment step is usually as simple as dropping the file next to your `index.html` and redeploying — for example:
```
cp llms.txt public/llms.txt
# then rebuild / redeploy as usual
```
## Step 6: Verify the File Is Actually Reachable
A file that isn't publicly accessible is the same as no file at all — and this is the step most people skip. Open a browser (or an incognito window) and visit `https://your-domain.com/llms.txt`. You should see your plain-text file, not a 404, a redirect, or a login page. From the terminal, the equivalent check is:
```
curl -I https://your-domain.com/llms.txt
# Expect: HTTP/2 200 and content-type: text/plain
```
A `200` status with a `text/plain` content type is exactly what you want. If you get a `404`, double-check the file name (it's `llms.txt`, not `llms.txt.txt` or `Llms.txt`) and the deploy path from Step 5. A redirect is acceptable on platforms like Shopify, but a 404 isn't.
## Step 7: Validate the File With the llms.txt Checker
The last step is a quality gate. Paste the URL or file contents into the [free llms.txt checker](/checker), which validates the structure against the llmstxt.org v2 specification — required H1, absolute URLs, correctly formed sections — and scores AI-readiness:
```
Checking https://your-domain.com/llms.txt ...
✔ H1 heading present
✔ Blockquote summary found
✔ 3 sections, all with link lists
✔ All 9 links use absolute URLs
✔ 0 errors — structure score: 10/10
```
Fix anything the checker flags, redeploy, and re-check. Once you see zero errors, you're done: agents will start finding the file the next time they crawl your domain. That's the whole process — seven steps, ten minutes, one small text file.
## Hand-Written or Generated: Which llms.txt Should You Use?
Both approaches are valid, and the file format is identical either way — so this is a workflow choice, not a quality one. Hand-writing gives you total control over wording and ordering, which matters if your descriptions need a human editorial touch. The generator wins on speed, completeness, and consistency: it never forgets the H1, never emits a relative URL, and regenerates in seconds when your site changes. The pragmatic pattern is a hybrid: generate an initial file, review and polish the descriptions, then keep the generator around for monthly refreshes. If you're creating files for multiple sites or products, the generator's consistency alone is worth it.
## Common llms.txt Creation Mistakes to Avoid
Most problems in hand-written files come down to a handful of recurring mistakes:
- **Missing H1.** Starting the file with prose or an H2 means the required title element is gone. Always begin with `# Site Name`.
- **Relative URLs.** `- [Docs](/docs)` breaks parsers and agents. Every link must be an absolute `https://…` URL.
- **Treating it like a sitemap.** Listing hundreds of pages dilutes the file's value. Curate: 5–15 links that actually help agents answer questions about you.
- **Hash comments.** A line like `# TODO: add pricing` is parsed as a heading. There is no comment syntax — use HTML comments (``) in the notes area if you must annotate.
- **Links without descriptions.** Bare links force agents to fetch each page to learn what it is. Add a `: short description` to every item.
- **Never re-checking.** Links go stale, pages get renamed, sections disappear. Set a monthly reminder to re-run the checker, or regenerate from scratch after any redesign.
## llms.txt Creation FAQ
### How long does it take to create an llms.txt file?
Around 10 minutes for a typical site: 2 minutes to pick your site name, summary, and top pages, 2 minutes to format them, 3 minutes to deploy, and 3 minutes to verify and validate. Using the generator cuts the formatting step to seconds.
### Do I still need an llms.txt file if I have a sitemap.xml?
Yes — they serve different purposes. A sitemap lists every URL for search-engine crawlers; llms.txt is a curated, prioritized index written in Markdown for LLMs. Agents can consume both, but llms.txt tells them what matters, which is what produces accurate AI answers.
### Can I create an llms.txt file for a WordPress or Shopify site?
Absolutely. On WordPress, upload the file to `public_html/`. On Shopify, upload it via Files and add a URL redirect from `/llms.txt`. Both make the standard `https://your-domain.com/llms.txt` URL work — see the placement table in Step 5 for details.
## Keep reading
[
### 📐 The Complete Format Guide
Every element of the v2 spec explained, with a valid end-to-end example.
](/blog/llms-txt-format/)[
### 🔌 llms.txt on WordPress
Three ways to ship llms.txt on WordPress, plus the pitfalls that break it.
](/blog/llms-txt-wordpress/)
## Ready to create your llms.txt file?
Generate a spec-perfect file in under a minute, then validate it with the checker — free, no sign-up.
[⚡ Generate your llms.txt](/generator) [Check my file](/checker)
---
# How to Validate an llms.txt File: Spec, Links and Delivery (2026)
> Source: https://llmstxtgenerator.dev/blog/how-to-validate-llms-txt/
Validation · September 2026
# How to Validate an llms.txt File: Spec, Links and Delivery
Validation is where most llms.txt projects quietly go wrong. A file can be pasted into a validator, come back clean, and still never be read — because the text of the file and the way the file is served are two different problems. Here is the three-layer check worth running on every llms.txt you publish.
## What "Valid" Actually Means
The [specification](https://llmstxt.org/) is deliberately short: one H1 with the site name, a blockquote summary, optional free-form details, then H2-delimited lists of links in `name: description` form. Of those, only the H1 is genuinely required. That minimalism is why three different layers of checking exist, and why a file that passes one can fail another.
Layer
Question it answers
How to test it
**1\. Structure**
Does the file follow the format?
Parser / validator, or a markdown lint pass
**2\. Link integrity**
Does every listed URL resolve to the page you meant?
HTTP status check on each link
**3\. Delivery**
Can a crawler actually fetch it as text?
Headers, content type, cache and WAF behaviour
## Layer 1: Structure — the Checks a Parser Can Make
These are mechanical and objective. A file passes or it does not:
- **Exactly one H1, on the first line.** Multiple H1s, or a heading that arrives after the summary, confuse the parse.
- **A blockquote summary.** Optional in the spec, but the single highest-value line in the file — it is the context an assistant reads before deciding whether to keep going.
- **Links as list items, with display text and an absolute URL.** `- [Docs](https://example.com/docs): API reference` is valid; a bare path or a bullet with no link is a formatting error.
- **No prose in link sections.** Paragraphs between the H2 sections are where parsers report unrecognised lines — put narrative text above the first H2.
- **`Optional` used for the right thing.** That H2 is reserved for resources that can be skipped when a consumer is short on context. Filling it with your best pages defeats the point.
```
# Acme Docs
> Acme is the developer documentation site for the Acme API,
> with REST and GraphQL references, SDKs and migration guides.
## Important links
- [Home](https://acme.com/): Documentation home
- [Quick start](https://acme.com/getting-started): 5-minute setup
- [API reference](https://acme.com/api): Complete REST reference
## Optional
- [Changelog](https://acme.com/changelog): Release history
```
If you would rather not eyeball it, the [llms.txt validator](/checker) runs this layer and more: it flags a missing H1, an H1 that is not on the first line, multiple H1s, links with no display text, relative URLs, unrecognised lines and a file with no links at all, then returns a 0–100 AI-readiness score so you can see progress rather than a binary verdict.
## Layer 2: Link Integrity — the Check Nobody Runs
Structure checks never touch your URLs. A file can be spec-perfect and full of 404s, redirect chains and pages that moved six months ago — see [llms.txt maintenance](/blog/llms-txt-maintenance/) for why that decay is the normal state of a published file. The audit is a loop:
```
# 1. pull the file
curl -s https://example.com/llms.txt -o /tmp/llms.txt -w '%{http_code} %{content_type}\n'
# 2. status of every absolute URL in it
grep -oE 'https?://[^)]+' /tmp/llms.txt | sort -u | while read -r u; do
printf '%s %s\n' "$(curl -s -o /dev/null -w '%{http_code}' -L "$u")" "$u"
done
# 3. anything still redirecting or gone
grep -oE 'https?://[^)]+' /tmp/llms.txt | sort -u | while read -r u; do
curl -sI -o /dev/null -w '%{http_code} %{url_effective}\n' "$u"
done
```
- A **404** means the entry is worse than nothing: the consumer spent a fetch and learned nothing.
- A **301 or 302** where the effective URL differs from the one you listed means you are sending agents through a detour. List the final URL.
- **Duplicates** — the same page under `www` and non-`www`, or with tracking parameters — split your authority across entries that look like separate resources.
- A **gated or login-walled** link is a dead end. If the page needs a session, describe it in the file but link to the public overview instead.
## Layer 3: Delivery — Where Most Failures Hide
The last layer is the one a text validator cannot see, because it never leaves the file. Four checks catch almost everything:
- **Path.** The file must be reachable at `/llms.txt` on the host that serves your pages. A copy on a `cdn.` subdomain or a docs subdomain only covers the URLs under that subdomain.
- **Content type.** `text/plain` or `text/markdown`. A catch-all rewrite that answers `text/html` is the classic failure, and it usually means an SPA fallback swallowed the path.
- **Status.** 200, directly — not a 302 to the homepage. Redirects to marketing pages are counted as missing files by anything doing a strict fetch.
- **Content.** Confirm the bytes you published are the bytes being served. Edge caches happily serve the previous version of a file that changed an hour ago.
Each of those has a specific curl test and a specific fix — the [delivery-level troubleshooting guide](/blog/llms-txt-not-working/) walks through all seven variants, from WAF rules to case-sensitive paths.
## Make Validation Part of the Build, Not a Launch Task
A file validated once at launch is a file that is correct for one day. Put the cheapest layer in CI so a rename or a deleted page fails the build instead of silently rotting in production:
```
- name: Validate llms.txt
run: |
test -f public/llms.txt || exit 1
head -1 public/llms.txt | grep -qE '^# ' || exit 1
node scripts/check-llms-links.mjs public/llms.txt
```
Keep the link check non-blocking at first while you clean up the existing entries, then make it blocking once the file is green. It is the same rhythm described in the [six llms.txt mistakes](/blog/llms-txt-mistakes/) that get a file ignored — only this time it is enforced by tooling rather than memory.
## What a Validator Cannot Tell You
Two things, and both matter more than the mechanical checks:
1. **Whether your summary is any good.** A grammatically valid blockquote can still be a slogan. Rewrite it as a factual description of what the site covers and who it is for.
2. **Whether you listed the right pages.** A perfectly formatted file of your ten least important URLs passes every check and helps nobody. That judgement is editorial, and it is what separates a [trusted file](/blog/llms-txt-best-practices/) from a compliant one.
Use the [generator](/generator) to get a structurally valid first draft from your sitemap, then cut it down by hand before publishing. Validation is the last step of that workflow, not the first.
---
# How to Add llms.txt to an Astro Site (Static or Generated, 2026)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-astro/
Astro tutorial · Static + generated · 6 minutes
# How to Add llms.txt to an Astro Site
Astro has no llms.txt file convention — and it is one of the few frameworks that genuinely does not need one. Between the `public/` directory and `src/pages` endpoints, there are three clean ways to ship a spec-compliant `llms.txt`, from a 30-second static file to a fully generated index.
## Why Astro and llms.txt Are a Natural Fit
The llms.txt convention exists so AI engines and coding agents can find the handful of pages that matter on your site. Astro sites are usually built from Markdown and frontmatter — exactly the structured content llms.txt wants to point at. On top of that, Astro's default output is static: whatever the build produces is what your CDN serves, which makes a root-level text file trivial.
Two Astro behaviors do all the work. First, every file in `public/` is copied verbatim into the output root, so `public/llms.txt` becomes `/llms.txt` automatically. Second, any `.ts` file in `src/pages/` becomes a route, so `src/pages/llms.txt.ts` serves plain text at the same path. This very site — llmstxtgenerator.dev — is an Astro 5 site and serves its own [llms.txt](/llms.txt) from the `public/` directory.
## Method 1: Static File in public/ (30 Seconds)
Create `public/llms.txt` with your H1, a blockquote summary, and curated sections of markdown links. This follows the [llmstxt.org v2 spec](/blog/llms-txt-v2/):
```
# Example Docs
> Example Docs is the official documentation for the Example platform.
> It covers installation, configuration, the API and troubleshooting.
## Getting Started
- [Quickstart](https://example.com/docs/quickstart): first project in 5 minutes
- [Installation](https://example.com/docs/install): system requirements and setup
- [Authentication](https://example.com/docs/auth): API keys and access tokens
## API Reference
- [REST API](https://example.com/docs/api): every endpoint, parameter and error
- [Webhooks](https://example.com/docs/webhooks): events your server can subscribe to
```
Run `npm run build` and check the output — the file appears at `dist/llms.txt`, ready for any static host. Deploying to Cloudflare Pages, Netlify or GitHub Pages requires no server config at all. Keep this method when your important links rarely change; move to Method 2 when they do.
## Method 2: Generate It at Build Time from a .ts Endpoint
For a file that should stay in sync with your site, create `src/pages/llms.txt.ts`. Astro maps the filename to the `/llms.txt` route, and on a static site the endpoint runs during the build — the result is a real file on your CDN, with nothing executed at request time:
```
// src/pages/llms.txt.ts
import type { APIRoute } from 'astro';
const pages = [
['Quickstart', 'https://example.com/docs/quickstart', 'first project in 5 minutes'],
['Authentication', 'https://example.com/docs/auth', 'API keys and access tokens'],
['REST API', 'https://example.com/docs/api', 'every endpoint, parameter and error'],
];
export const GET: APIRoute = () {
const lines = [
'# Example Docs',
'',
'> Example Docs is the official documentation for the Example platform.',
'',
'## Documentation',
'',
];
for (const [name, url, summary] of pages) {
lines.push('- ' + name + ': ' + summary + ' (' + url + ')');
}
return new Response(lines.join('\n') + '\n', {
headers: { 'Content-Type': 'text/plain; charset=utf-8' },
});
};
```
Two details matter. The `Content-Type` must be `text/plain` — agents and validators expect plain text, not HTML. And if you later add an adapter for server-side rendering, add `export const prerender = true` to keep generating the file at build time instead of on every request.
To go further, pull from your content collections instead of a hard-coded array. Import `getCollection` from `astro:content`, map each post's frontmatter title, description and URL into a markdown link line, and join them. Your `llms.txt` then updates itself every time you publish content.
## Method 3: The astro-slop Integration (Maintenance-Free)
If you run a large docs or blog site, the community [astro-slop](https://github.com/yaroslav/astro-slop) integration automates the whole pipeline. It generates an `llms.txt` index from your pages, per-page Markdown versions of every route, a companion `llms-full.txt`, and adds content negotiation plus `rel="alternate"` link headers. It is the closest thing to a "just works" llms.txt setup for Astro today — worth a look before you hand-roll Method 2 on a big site.
## Astro-Specific Checklist
- **One source of truth** — a `public/llms.txt` file and a `src/pages/llms.txt.ts` route that both target /llms.txt will drift apart. Pick one.
- **Absolute URLs only** — `https://example.com/docs/start`, never `/docs/start`. Agents resolve links from the file's location.
- **Serve text/plain** — endpoints must set `Content-Type: text/plain`; a validators may reject HTML responses.
- **Curate, don't dump** — 10–50 links to canonical docs pages beat 500 links to marketing pages. See the [format guide](/blog/llms-txt-format/) for structure rules.
- **Add llms-full.txt for deep docs** — if your reference section is large, link a generated [llms-full.txt](/blog/llms-txt-vs-llms-full-txt/) companion file.
- **Keep it fresh** — rebuild and redeploy whenever a listed page changes; stale files erode agent trust. The [best practices guide](/blog/llms-txt-best-practices/) covers maintenance cadence.
## Verify and Deploy
After deploying, confirm the route serves plain text with a 200 status:
```
curl -s -o /dev/null -w "%{http_code} %{content_type}\n" https://your-domain.com/llms.txt
# Expect: 200 text/plain
```
If you host on Cloudflare Pages, our [Cloudflare Pages deployment guide](/blog/llms-txt-cloudflare-pages/) covers the build configuration. Then paste your URL into the free [llms.txt checker](/checker) to validate the H1, the blockquote summary, absolute URLs and section structure before AI engines start reading it.
Build the initial curated file in seconds by running your sitemap through the [free llms.txt generator](/generator), then validate the result with the [checker](/checker) before deploying. A correct, current `llms.txt` is one of the highest-leverage AI visibility changes most Astro sites can make this week.
---
# llms.txt Best Practices 2026: The 9 Rules for a File AI Engines Trust
> Source: https://llmstxtgenerator.dev/blog/llms-txt-best-practices/
Spec guide · AI readiness · Updated 2026
# llms.txt Best Practices 2026: The 9 Rules for a File AI Engines Trust
The llms.txt spec is short — one required H1, a few conventions — but the difference between a file agents use and a file they ignore is in the details. These nine rules cover format, curation, and maintenance.
## 1\. Start With the H1 — It's the Only Required Element
The first line of llms.txt must be a single Markdown H1 with your site name. Everything else is optional by the spec — but a file with no H1 fails the very first check an agent (or a validator) runs.
```
# Example Site
```
## 2\. Write a One-Line Summary as a Blockquote
Right under the H1, a `>` blockquote telling agents what the site is and who it's for. This is the context an AI engine uses before reading anything else — make it specific, not marketing-speak.
```
# Example Site
> A SaaS API documentation site for developers building with Example SDK.
```
## 3\. Curate, Don't Dump
5 to 15 links. The whole point of llms.txt is that an agent can read it in milliseconds and know what matters. If your site needs more material, publish [llms-full.txt](/blog/llms-txt-vs-llms-full-txt/) and link to it — don't turn the index into a sitemap.
## 4\. Use Absolute URLs — Always
Every link must be a full `https://your-domain.com/…` URL. Relative paths like `/docs/start` or protocol-relative `//example.com/x` break agents that fetch the file directly. The one allowed exception is the relative link to `llms-full.txt` described in the spec.
## 5\. Group Links With H2 Sections
Use `##` headings to group related links — Documentation, Blog, Pricing, Support. Agents use these sections to match the user's question to the right part of your site.
```
## Documentation
- [Getting Started](https://example.com/docs/getting-started)
- [API Reference](https://example.com/docs/api)
## Blog
- [How We Built X](https://example.com/blog/how-we-built-x)
```
## 6\. Order Links by Importance
The spec has no ranking mechanism beyond order — so your most valuable pages go first within each section. First impressions matter for agents the same way they do for humans.
## 7\. Keep Link Text Descriptive
Use link text that says what the page is ("Getting Started Guide"), not the URL or generic words ("click here"). Agents quote link text when they cite your site.
## 8\. Maintain It Like Code
- Regenerate after major content changes (new flagship pages, removed sections).
- Validate after every change — our [free validator](/checker) scores spec compliance and flags broken structure.
- Generate from your sitemap with the [generator](/generator) when you add many pages at once.
- Serve it as `text/plain` at the site root with no auth, no redirects.
## 9\. Know What It Is — and Isn't
llms.txt is an AI-visibility signal, not a [Google ranking factor](/blog/does-llms-txt-affect-seo/). Don't neglect robots.txt, sitemap, and on-page SEO because you have it — run all three together ([see how they compare](/blog/llms-txt-vs-robots-txt/)), and you're covered for search crawlers _and_ AI agents.
## A Complete Reference Example
```
# Example Site
> A SaaS API documentation site for developers building with Example SDK.
## Documentation
- [Getting Started](https://example.com/docs/getting-started)
- [Authentication](https://example.com/docs/auth)
- [API Reference](https://example.com/docs/api)
## Blog
- [How We Built X](https://example.com/blog/how-we-built-x)
- [Pricing Philosophy](https://example.com/blog/pricing)
[llms-full.txt](llms-full.txt)
```
That's a file an agent can act on in under a second. Build yours with the [generator](/generator), check it with the [validator](/checker), and keep it honest.
---
# llms.txt Case Studies 2026: Inside the Files of Vue.js, Nuxt, Cloudflare, OpenAI and Anthropic
> Source: https://llmstxtgenerator.dev/blog/llms-txt-case-studies/
Case studies · Real production files · September 2026
# llms.txt Case Studies: Inside Real Production Files
Specs and templates tell you how llms.txt _should_ look. Real files show you how teams actually ship it. We fetched six production llms.txt files with curl on September 10, 2026 and broke down exactly what Vue.js, Nuxt, Cloudflare, OpenAI and Anthropic put in theirs — and what you can copy.
## Six Production Files, Verified Live
Every file below was requested over HTTPS during research for this article. All of them responded with HTTP 200 and spec-shaped content:
Site
File
Distinctive pattern
Vue.js
`/llms.txt`
Hierarchical table of contents, every link a .md page
Nuxt
`/llms.txt` + 4.7 MB llms-full.txt
Agent-first: MCP server, sitemap.md, openapi.json pointers
Cloudflare Developers
`/llms.txt` + per-product files
Hub-and-spoke index of subpath llms.txt files
OpenAI API docs
`/docs/llms.txt` + llms-full.txt
Markdown twin for every page, single combined export
Anthropic docs
`/llms.txt`
Multilingual hub with per-language page counts
llmstxt.org
`/llms.txt`
The spec site's own minimalist example
## Vue.js: A Clean Hierarchical Table of Contents
The first three lines of `vuejs.org/llms.txt` state the site name and tagline, then the file becomes a pure navigation document — every framework topic grouped under headings like `Getting Started` and `Essentials`, each link pointing at the Markdown version of a page:
```
# Vue.js
Vue.js - The Progressive JavaScript Framework
### Getting Started
- [Introduction {#introduction}](/guide/introduction.md)
```
Note the details: URLs end in `.md`, and the anchor metadata in brackets matches the HTML page's heading IDs so agents can link both formats. For full depth, Vue.js also serves `llms-full.txt` (about 0.9 MB at the time of testing) — exactly the two-file setup described in our [llms.txt vs llms-full.txt](/blog/llms-txt-vs-llms-full-txt/) guide.
## Nuxt: The Agent-First File
Nuxt's file shows where the format is heading. After a one-sentence summary it tells agents that every page has a Markdown twin, then points to an unusually rich set of machine-readable resources:
```
# Nuxt Docs
> Nuxt is an open source framework that makes web development intuitive and powerful.
Every documentation, blog and deploy page is also Markdown. Append `.md` to a URL...
```
The rest of the file lists `llms-full.txt` — 4,756,574 bytes of single-file documentation, which we downloaded to confirm — plus a `sitemap.md`, a public OpenAPI spec and an MCP server endpoint. For an LLM agent, one llms.txt now opens the door to the whole ecosystem. The file even includes a `When to use this` section that tells the agent which tasks the docs are for — curation as copywriting.
## Cloudflare Developers: A Hub of Subpath Files
Cloudflare's documentation is huge, so its root file does not try to list everything. It is a hub: one link per product area, each pointing at that product's _own_ llms.txt one level down:
```
## Application performance
- [DNS](https://developers.cloudflare.com/dns/llms.txt): Deliver excellent
performance and reliability to your domain
```
This is the subpath pattern the [llms.txt v2 spec](/blog/llms-txt-v2/) formally defines: a file covers the URLs under its own path, and agents use the most specific one. Cloudflare pairs the hub with a root `llms-full.txt`, giving agents both a browseable directory and a complete dump.
## OpenAI and Anthropic: AI Labs Publish Their Own Files
The biggest irony in the llms.txt ecosystem is that the AI labs themselves are early adopters. OpenAI's API docs (`platform.openai.com/docs/llms.txt`, HTTP 200) opens with `# OpenAI API docs`, states that every entry has a Markdown twin, and links a combined single-file export. Anthropic's file goes multilingual, listing page counts per language — English at 700 pages, others at 249 — so an agent knows which translations are complete before fetching anything.
Neither file is a ranking hack: crawler studies we covered in [Which AI Engines Actually Read llms.txt?](/blog/which-ai-engines-read-llms-txt/) show most AI bots never request the file. These companies ship it because their docs are Markdown anyway and the file costs nothing to generate. Google's `ai.google.dev/llms.txt` returns 404 — a reminder that adoption is real but not universal.
## What the Best Production Files Have in Common
Strip away the branding and five patterns recur in every case study above:
- **One-line identity.** H1 with the product name plus a sentence saying what the docs cover.
- **Markdown-first links.** Point agents at `.md` versions, not HTML pages.
- **A description on every link**, so agents can decide before fetching — Cloudflare's one-liners and Nuxt's task guidance are both this idea.
- **Curated H2 groups** of five to forty links, never a URL dump.
- **An escape hatch**: llms-full.txt (or subpath files) for sites too big for one page.
None of these files is longer than a couple of hundred lines, and all of them are plain Markdown. Structure your file the same way — then check it the way we checked these, with curl, a validator, or both. Generate a spec-compliant draft from your sitemap with the [llms.txt generator](/generator), run it through the [checker](/checker), and publish it at your site root. If you need a refresher on any element, the [complete llms.txt format guide](/blog/llms-txt-format/) covers every field.
---
# How to Add llms.txt to Cloudflare Pages (Astro, React, Plain HTML)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-cloudflare-pages/
Deployment tutorial · Cloudflare Pages · 5 minutes
# How to Add llms.txt to Cloudflare Pages
Cloudflare Pages is one of the easiest hosts to serve llms.txt from — you get it at the site root with zero server config. Here are the three ways to do it, from simplest to most flexible.
## Method 1: Static File (Simplest — Works for Any Framework)
Every Cloudflare Pages project serves files from a build output directory. Put an `llms.txt` file there and it's automatically live at your domain root.
**Astro / SvelteKit / 11ty (static dir):** drop the file in the framework's static folder, which is copied to the output as-is:
```
# Astro: place this at public/llms.txt
# (public/ contents are copied to dist/ at build time)
```
**Plain HTML / Vite / React (manual):** add `llms.txt` next to your `index.html` in the output folder (e.g. `dist/llms.txt`), or configure your build tool to copy it there.
Build and deploy as usual — either with `wrangler pages deploy` or via your connected Git repository. The file is served at `https://your-domain.com/llms.txt`.
## Method 2: Dashboard Direct Upload
If you manage the project through the Cloudflare dashboard (Workers & Pages → your project → **Upload assets**), you can upload an `llms.txt` file directly without a build pipeline at all. This is handy for a quick test or for sites not built from a repo.
## Method 3: Pages Function (Dynamic Generation)
When your llms.txt should reflect live data — a fresh sitemap, a changing content list, or a dynamic `llms-full.txt` — generate it with a [Pages Function](https://developers.cloudflare.com/pages/functions/):
```
// functions/llms.txt.js — serves https://your-domain.com/llms.txt
export async function onRequest() {
const llms = `# Example Site
> A concise description of the site.
## Docs
- [Getting Started](https://example.com/docs/start)`;
return new Response(llms, {
headers: { "Content-Type": "text/plain; charset=utf-8" },
});
}
```
The function can fetch your `sitemap.xml`, filter sections, or pull from a CMS — anything you can do in JavaScript. Just keep the response under the platform's limits and serve `text/plain`.
## Verify It's Live
After deploying, check from the command line or browser:
```
curl -I https://your-domain.com/llms.txt
# Expect: HTTP/2 200, Content-Type: text/plain
```
Then validate the content — run the URL through our free [llms.txt checker](/checker) to confirm the H1, absolute URLs, and Markdown structure are spec-compliant before AI engines start reading it.
## Common Mistakes
- **Wrong folder** — the file must end up in the _build output_, not your source root. Check `dist/` (or your output dir) after building.
- **Relative URLs** — llms.txt links must be absolute (`https://…`), not `/docs/start`.
- **Missing H1** — the file must start with a single `#` heading with your site name; the blockquote summary below it is what agents read for context.
- **Caching surprises** — static files on Pages are cached at the edge; if you regenerate via a Function, add `Cache-Control: no-store` (as above) or purge when content changes.
Once it's live, add the file to your footer or homepage so human visitors (and the agents that read pages) discover it — and consider adding an [llms-full.txt](/blog/llms-txt-vs-llms-full-txt/) companion for content-heavy sections.
---
# How to Add llms.txt to Docusaurus (Docs Sites, 2026)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-docusaurus/
Docusaurus tutorial · Static or generated · 6 minutes
# How to Add llms.txt to Docusaurus
Docusaurus has no official `llms.txt` plugin yet — but it is one of the best-behaved frameworks for the format anyway. Everything in `static/` is copied to your site root at build time, so a spec-compliant file is a 60-second job. For larger docs, a community plugin generates `llms.txt` and `llms-full.txt` from your actual pages. Both routes are covered here.
## Why Docs Sites Are llms.txt's Best Use Case
The llms.txt convention exists so AI engines, chatbots and coding agents can find the pages that answer their questions without crawling an entire site. Documentation is the content type where that matters most: docs are factual, densely linked and constantly referenced in AI answers, yet their navigation is often sidebar-heavy and hostile to crawlers.
The docs-tooling world noticed early. The llmstxt.org integrations list includes Mintlify, which generates an `llms.txt` and Markdown page versions for every site it hosts, and GitBook, which serves the file for published docs. Real-world adopters include dub.co (via Mintlify) and Cloudflare's developer docs (via a build script). Docusaurus, used by React, Jest and thousands of open-source projects, has no native option yet — an official plugin proposal has sat as an open feature request in the facebook/docusaurus repo since February 2025. In practice you do not need to wait for it.
## Method 1: A Hand-Written File in static/ (60 Seconds)
Docusaurus treats `docs/` as content and `static/` as assets: every file in `static/` is copied verbatim to the output root on every build. That is the entire mechanism — create `static/llms.txt` and follow the [v2 spec](/blog/llms-txt-v2/) structure:
```
# Acme Docs
> Acme Docs is the official documentation for the Acme platform.
> It covers installation, configuration, the REST API and troubleshooting.
## Getting Started
- [Quickstart](https://acme.com/docs/quickstart): first project in five minutes
- [Authentication](https://acme.com/docs/auth): API keys and access tokens
- [Deployment](https://acme.com/docs/deploy): production setup for self-hosting
## API Reference
- [REST API](https://acme.com/docs/api): every endpoint, parameter and error
- [Webhooks](https://acme.com/docs/webhooks): events your server can subscribe to
```
Build with `npm run build` and the file lands at the root of `build/`, ready for Netlify, Vercel, Cloudflare Pages or any static host. Keep this method when your information architecture changes rarely. The file must point to canonical pages with absolute URLs — never relative paths, which agents resolve inconsistently.
## Method 2: Generate llms.txt and llms-full.txt From Your Docs
For a file that stays in sync with a growing sidebar, the community plugin [docusaurus-plugin-llms](https://github.com/rachfop/docusaurus-plugin-llms) (MIT-licensed and listed on llmstxt.org) walks your docs at build time. With no options it writes a curated table of contents to `llms.txt` and combines the full text of every page into `llms-full.txt`:
```
// docusaurus.config.js
const config = {
// ... your existing config ...
plugins: [
'docusaurus-plugin-llms',
],
};
export default config;
```
The plugin is designed for exactly the situations a static file cannot handle: versioned docs (each version can emit its own file, such as `/stable/llms.txt`), custom LLM files for specific audiences, and large references where a full-text dump saves an agent dozens of requests. If your reference section is deep, link the generated dump the way the [llms.txt vs llms-full.txt guide](/blog/llms-txt-vs-llms-full-txt/) describes.
## The Docusaurus Checklist
- **One file, one owner** — a hand-written `static/llms.txt` and a plugin that also writes `/llms.txt` will fight each other on rebuild. Pick one approach.
- **Curate docs, not marketing pages** — an agent landing on your docs wants the quickstart, the auth guide and the API reference, not your pricing page. See the [best practices guide](/blog/llms-txt-best-practices/) for curation rules.
- **Absolute URLs only** — `https://acme.com/docs/quickstart`, never `/docs/quickstart`.
- **Watch your baseUrl** — sites deployed under a subpath like `/docs/` should generate the file with the plugin so URLs and placement stay consistent with the spec's subpath-file rules.
- **Ship it with every release** — Docusaurus is a static generator: nothing exists on your host until you deploy. Rebuild after sidebar changes so the file never points at removed pages.
## Verify Before Agents Do
After deploying, confirm the file serves with a 200 status and a text content type:
```
curl -s -o /dev/null -w "%{http_code} %{content_type}\n" https://your-domain.com/llms.txt
# Expect: 200 text/plain
```
Docusaurus serves the file from the same CDN edge as your docs, so check the version you actually deployed — a stale build is the most common reason an llms.txt goes missing after a redesign. Then paste the URL into the free [llms.txt checker](/checker) to validate the H1, the blockquote summary and the section structure.
Draft the initial file in seconds by feeding your sitemap into the [free llms.txt generator](/generator), then verify the output with the [checker](/checker) before you wire up the plugin. Documentation sites are the highest-leverage place an `llms.txt` can live — a correct, current file is cheap insurance that your docs are the version AI engines read.
---
# llms.txt Examples: Real Files for Every Site Type
> Source: https://llmstxtgenerator.dev/blog/llms-txt-examples/
Copy-ready examples · Updated 2026
# llms.txt Examples: Real Files for Every Site Type
The fastest way to learn a format is to read a good example — so here are four complete, spec-valid **llms.txt example** files for the site types that publish them most: SaaS documentation, e-commerce, personal blogs, and corporate websites.
## Why Examples Are the Fastest Way to Learn llms.txt
The llms.txt specification is deliberately minimal — one required `# H1`, an optional blockquote summary, free-form notes, and `## H2` sections that each contain a list of links. Reading the spec tells you the rules; reading real **llms.txt examples** tells you what the rules look like in practice. The format leaves real decisions to you: which pages deserve a link, how to group them into sections, how many links is too many, and where the optional stuff goes.
Every example below follows the llmstxt.org v2 specification exactly: exactly one H1 at the top, a blockquote summary, short notes without headings, H2 file lists with absolute URLs and a `: description` on every link, and a trailing `## Optional` section where it helps. The site names are fictional — swap in your own domain, pages, and descriptions. If you want the rules themselves, the [complete llms.txt format guide](/blog/llms-txt-format/) walks through every element.
The second thing examples teach you is that the format scales without changing shape: the same skeleton that fits a ten-link blog file also carries a multilingual corporate index. Once you can read one of these files end to end, you can write one for any site — which is exactly what the generator behind this page automates.
## llms.txt Example 1: SaaS Documentation Site
Documentation is the canonical use case for llms.txt — AI coding agents answer questions from docs more than from any other page type. This example shows three organizational habits worth copying: linking straight to `.md` versions of pages, referencing an `llms-full.txt` for agents that want everything, and moving secondary links into `## Optional`.
```
# Pulseboard Analytics
> Pulseboard is a product analytics platform for SaaS teams — event tracking, funnels, retention, and dashboards from a single SDK.
Pulseboard has SDKs for JavaScript, Python, and mobile. All documentation is published as Markdown, so every link below points directly at a .md file.
## Important links
- [Quick start](https://pulseboard.dev/docs/quickstart.md): Install the SDK and send your first event in 10 minutes
- [Event tracking API](https://pulseboard.dev/docs/api/events.md): Track, batch, and alias events, with code examples
- [Dashboards & reports](https://pulseboard.dev/docs/dashboards.md): Build funnels, retention cohorts, and custom reports
- [llms-full.txt](https://pulseboard.dev/llms-full.txt): The complete documentation in a single file
## Guides
- [Tracking plan design](https://pulseboard.dev/docs/guides/tracking-plan.md): Structure your events before you ship
- [Data privacy & consent](https://pulseboard.dev/docs/guides/privacy.md): GDPR, CCPA, and cookieless tracking
- [Migrating from Amplitude](https://pulseboard.dev/docs/guides/migrate.md): Step-by-step migration guide
## Optional
- [Changelog](https://pulseboard.dev/changelog.md): Release notes for all past versions
- [Pricing](https://pulseboard.dev/pricing.md): Plans, limits, and billing FAQ
```
Notice what makes this file useful to an agent: the blockquote explains what the product is in one breath, the `llms-full.txt` link offers a deep-read option without bloating this file, and `## Optional` keeps pricing and changelog out of the default context. The `.md` URLs matter too — Markdown renders cleanly inside an LLM's context window, with none of the nav, cookie banners, or layout noise of the HTML versions.
## llms.txt Example 2: E-commerce & Product Site
For stores, the goal is different: agents should be able to answer product questions — what you sell, what it costs, how shipping works — and recommend pages to shoppers. This **llms.txt example** organizes by product line, adds an HTML comment in the notes area (the only safe place for comments), and keeps time-sensitive pages like sales in `## Optional`.
```
# Terra Outfitters
> Terra Outfitters is an online outdoor gear store — tents, backpacks, sleeping bags, and hiking apparel with free shipping over $75.
## Product lines
- [Tents](https://terra-outfitters.com/products/tents.md): 2–6 person tents, shelters, and footprints
- [Backpacks](https://terra-outfitters.com/products/backpacks.md): Daypacks and multi-day packs from 20L to 75L
- [Sleeping bags](https://terra-outfitters.com/products/sleeping-bags.md): Down and synthetic bags rated to -20°F
- [Hiking apparel](https://terra-outfitters.com/products/apparel.md): Jackets, base layers, and trail pants
## Help & policies
- [Shipping & returns](https://terra-outfitters.com/shipping.md): Delivery times, costs, and 60-day returns
- [Sizing guide](https://terra-outfitters.com/sizing.md): Fit guidance for packs, apparel, and footwear
- [Product care](https://terra-outfitters.com/care.md): Cleaning and repair instructions
## Optional
- [Sale & clearance](https://terra-outfitters.com/sale.md): Current discounts and limited offers
```
The shape is the same as the SaaS example, but the content priorities flipped: category pages and policies carry the weight instead of API references. The HTML comment survives parsing because it sits in the notes area, where the spec allows it. If your store runs seasonal promotions, keep them out of the main sections — a sale page that goes stale in `## Optional` costs far less trust than a dead link at the top of the file.
## llms.txt Example 3: Personal Blog
Blogs have the opposite problem from docs sites: too much content, most of it not worth an agent's context. A good blog **llms.txt example** is a curated list — recent or best posts only — with the archive tucked into `## Optional`. Keep the whole file under a couple of hundred words and agents will actually read it.
```
# Field Notes by Maya Chen
> Field Notes is a personal blog about distributed systems, Rust, and marathon running, written by Maya Chen, a backend engineer at a fintech startup.
Posts are written in Markdown and archived by category. The list below covers the most useful posts; the full archive is under Optional.
## Recent posts
- [Why we moved from Postgres to TiDB](https://mayachen.dev/posts/tidb-migration.md): A production postmortem with real query numbers
- [Rust async in 15 minutes](https://mayachen.dev/posts/rust-async.md): A practical walkthrough for beginners
- [My 2025 reading list](https://mayachen.dev/posts/reading-2025.md): Books on systems, history, and endurance training
- [Training for a sub-3 marathon](https://mayachen.dev/posts/marathon-training.md): The exact plan I used, week by week
## Optional
- [About](https://mayachen.dev/about.md): Bio, speaking, and contact details
- [All posts](https://mayachen.dev/archive.md): Complete chronological archive
- [Talks](https://mayachen.dev/talks.md): Slides and recordings from conference talks
```
Ten links, two sections, one blockquote — that's the whole file. If you publish a lot, rotate the featured posts as new ones land, or point `## Optional` at category pages instead of individual posts. "Most useful" beats "most recent": an evergreen tutorial outranks yesterday's link roundup, because agents are usually answering a question, not catching up on news.
## llms.txt Example 4: Corporate Website
Corporate sites serve a different audience again: investors, journalists, and job seekers asking factual questions. This **llms.txt example** demonstrates a multilingual layout — a section per language, each with its own absolute URLs — plus company-critical pages up front and legal pages at the end.
```
# Nordhelm Group
> Nordhelm Group is a Munich-based industrial automation company — robotics, control software, and factory integration services for manufacturers across Europe.
The site is published in English and German. Links are grouped by language, and every page is available as Markdown.
## English
- [About Nordhelm](https://nordhelm-group.com/en/about.md): Company history, leadership, and locations
- [Products & solutions](https://nordhelm-group.com/en/products.md): Robot arms, vision systems, and control software
- [Industries](https://nordhelm-group.com/en/industries.md): Automotive, electronics, and logistics use cases
- [Newsroom](https://nordhelm-group.com/en/news.md): Press releases and company news
## Deutsch
- [Über Nordhelm](https://nordhelm-group.com/de/ueber-uns.md): Unternehmensgeschichte, Führung und Standorte
- [Produkte & Lösungen](https://nordhelm-group.com/de/produkte.md): Roboterarme, Bildverarbeitung und Steuerungssoftware
- [Karriere](https://nordhelm-group.com/de/karriere.md): Offene Stellen und Bewerbungsprozess
## Optional
- [Investor relations](https://nordhelm-group.com/en/investors.md): Annual reports and financial results
- [Imprint & privacy](https://nordhelm-group.com/en/imprint.md): Legal notices (Impressum) and privacy policy
```
The language-split sections are the takeaway: agents can fetch the German pages when the answer should be in German and ignore them otherwise. Note the German descriptions still follow the `[text](url): note` convention — keep descriptions in the page's own language so the link list stays meaningful to every reader. The legal section at the end matters for European sites, where an Impressum is a legal requirement and a common question for agents to answer.
## What to Highlight in Your llms.txt, by Site Type
Each site type earns agent traffic for different reasons. Use this comparison to decide what belongs in your file:
Site type
What agents look for
What to prioritize
**SaaS documentation**
Setup steps, API reference, guides
Quick start + API links + an `llms-full.txt` reference
**E-commerce / product**
What you sell, prices, shipping and returns
Product categories, help & policy pages, FAQ
**Personal blog**
Recent and evergreen posts, author context
A curated list of recent/best posts — keep it short
**Corporate website**
Company facts, products, news, contact info
About, products, newsroom — one section per language if multilingual
Hybrid sites — a blog with a docs section, a store with a corporate About page — can mix and match: pick the section structure that matches your dominant content type, then add one section for the secondary type. The file stays valid either way, because validity never depends on the topic, only on the order and shape of the elements.
## The Universal llms.txt Template: A Copy-Paste Checklist
Strip away the site-specific choices and all four examples share the same skeleton. Work through this checklist in order and you'll have a valid file every time:
Step
Element
What to include
1
`# Site Name` (H1)
The only required element — exactly one, at the very top.
2
Blockquote
One or two lines of key context about the site.
3
Free-form notes
0–3 short paragraphs; no headings allowed here.
4
`## Section` (H2)
Two to five file lists, grouped by topic.
5
Link items
`[text](absolute URL): description` — never relative, always described.
6
`## Optional`
Secondary links agents can skip when context is short.
7
`llms-full.txt`
One link to the expanded file, if you publish one.
8
Validation
Run the finished file through the free [llms.txt checker](/checker).
When you're done, your file should look like a short README: roughly 40 to 150 lines, one H1, one blockquote, a few notes, and every link absolute and described. If a section is thinner than three links, fold it into another section — the checklist is a guide, not a contract.
## llms.txt Examples FAQ
### Can I copy these llms.txt examples as-is?
The structure is copy-safe — that's the point of a template. Replace the site name, links, and descriptions with your own, keep the section order, and run the result through a validator. The fictional domains in these examples will fail link checks if you leave them in.
### How many links should my llms.txt file have?
There is no hard limit, but most sites do best with 10–40 links across two to five sections. The file is meant to fit comfortably in a context window — if it reads like a sitemap, cut it down and move the long tail into `## Optional` or `llms-full.txt`.
### Should my llms.txt list every page on my site?
No. An llms.txt file is a curated index, not a sitemap — agents trust it because every link is worth reading. List the pages that deserve attention, group them sensibly, and publish the full expansion separately as `llms-full.txt`.
## Keep reading
[
### 📐 The Complete Format Guide
Every element of the v2 spec explained, with a valid end-to-end example.
](/blog/llms-txt-format/)[
### 🛠️ How to Create llms.txt
A step-by-step tutorial: build your file in 10 minutes and validate it.
](/blog/how-to-create-llms-txt/)
## Ready to write your own llms.txt?
Generate a spec-perfect file from your URLs in under a minute — free, no sign-up. Then read the full format guide for the details.
[⚡ Generate your llms.txt](/generator) [Read the full format spec](/blog/llms-txt-format/)
---
# llms.txt Format Guide — Complete Specification & Structure
> Source: https://llmstxtgenerator.dev/blog/llms-txt-format/
llmstxt.org v2 spec · Updated 2026
# The Complete llms.txt Format Guide
Everything you need to write a spec-perfect **llms.txt** file: the official format, every element explained, a valid end-to-end example, common mistakes, and answers to the questions developers ask most.
## What Is the llms.txt Format?
**llms.txt** is a plain Markdown file, conventionally placed at the root of a website (`https://your-domain.com/llms.txt`), that gives AI engines — ChatGPT, Claude, Gemini, Perplexity, coding agents — a curated, human- and LLM-readable index of your content. The proposal was published in 2024 by AI researcher [Jeremy Howard](https://en.wikipedia.org/wiki/Jeremy_Howard_\(entrepreneur\)) and the canonical specification lives at [llmstxt.org](https://llmstxt.org) (v2, August 2026). It's the LLM equivalent of `robots.txt`: instead of telling crawlers what _not_ to visit, it tells agents exactly what _is_ worth reading — in one small file that fits comfortably in a context window.
Why should developers and site owners care about the format itself? Because adoption is no longer hypothetical: Chrome's Lighthouse audits sites for an llms.txt file, documentation platforms such as Mintlify and GitBook generate one automatically, and OpenAI, Anthropic, and Google all publish llms.txt files for their own docs. An invalid or badly structured file fails the audit, confuses parsers, and costs you visibility in AI answers — so knowing the exact structure matters.
## The llms.txt Specification, Element by Element
The format is deliberately minimal: it uses Markdown as the structure language, and it can be parsed with plain regex and a list parser. Per the llmstxt.org v2 specification, a valid file contains the following elements **in this order**:
1. **H1 heading (`# Site Name`)** — the name of the project or site. This is the _only required element_ in the whole file. An optional byte-order mark (BOM) may precede it.
2. **Blockquote summary (`> ...`)** — an optional one- or two-line summary containing the key context needed to understand everything else in the file.
3. **Free-form notes** — zero or more Markdown sections (paragraphs, lists, and so on) _of any type except headings_, giving agents more detail about the project and how to interpret the links below.
4. **H2 sections (`## Section Name`)** — zero or more sections that each contain a "file list": a Markdown list of links where further detail lives. A section named `## Optional` is used by convention for secondary links that agents can skip when context is short.
5. **Link list items** — each entry is a Markdown hyperlink `[text](url)`, optionally followed by a colon and a short note about the file: `- [API reference](https://site.dev/api.md): Complete REST API docs`.
Two additional rules keep files machine-readable:
- **URLs must be absolute.** Use full `https://…` addresses — relative paths like `/docs/api` break both LLM consumption and regex-based parsers. Prefer linking to clean Markdown versions of pages (same URL with `.md` appended or substituted).
- **There is no comment syntax.** A line starting with `#` is parsed as a heading, not a comment — so never use hash-comments to annotate the file. If you want human-readable annotations, use HTML comments (``) in the free-form notes area above the first `##`; parsers ignore them.
One more scoping rule: an llms.txt file doesn't have to live at the site root. It can be placed at any path (e.g. `/docs/llms.txt`) and then covers all URLs under that path. Where multiple files apply, agents should use the most specific one.
## llms.txt Structure: Element Reference Table
Element
Required?
Description
Byte-order mark (BOM)
Optional
May precede the file content; parsed and skipped by tools.
H1 heading (`# Name`)
Yes — the only required element
Name of the site or project.
Blockquote (`> ...`)
Optional
Short summary with the key context for interpreting the file.
Free-form notes
Optional
Paragraphs and lists with more detail; any Markdown _except headings_.
H2 sections (`## Name`)
Optional
Delimit "file lists" of links; `## Optional` is reserved by convention for secondary links.
Link list items (`[text](url)`)
Inside H2 sections
Required Markdown hyperlink, optionally followed by `: description`.
Absolute URLs
Yes
Every link must be a full `https://…` address; no relative paths.
## A Complete, Valid llms.txt Example
Here is a full example that follows the v2 specification exactly — H1, blockquote, notes, H2 file lists, absolute URLs, and a trailing `## Optional` section:
```
# Acme Documentation
> Acme is a developer platform for building and deploying serverless applications in Python and TypeScript.
Important notes:
- All APIs below are stable; existing endpoints never break compatibility.
- Examples assume the latest CLI, installable with `pip install acme-cli`.
## Important links
- [Acme quick start](https://acme.dev/docs/quickstart.md): Set up your first app in 10 minutes
- [REST API reference](https://acme.dev/docs/api.md): Complete API reference with request/response examples
- [CLI reference](https://acme.dev/docs/cli.md): Every command and flag of the acme CLI
## Guides
- [Deploying to production](https://acme.dev/docs/guides/production.md): Scaling, monitoring, and rollbacks
- [Authentication](https://acme.dev/docs/guides/auth.md): API keys, OAuth, and JWT handling
## Optional
- [Changelog](https://acme.dev/changelog.md): Release notes for all past versions
```
Notice the structure: one H1 at the top, one blockquote, a short notes block (no headings), then only H2 sections containing link lists. The file stays small enough to fit in a context window — agents fetch the linked Markdown pages only when they need detail.
## llms.txt vs llms-full.txt: What's the Difference?
The official specification defines only `llms.txt`. `llms-full.txt` is a widely adopted _convention_ (pioneered by sites like OpenAI's and Anthropic's docs) for publishing the complete, expanded version of your documentation — every section in full, not just links to it.
`llms.txt`
`llms-full.txt`
**Status**
Defined by the llmstxt.org proposal
Unofficial, community convention
**Content**
Curated index: H1, summary, links to detail
The full content itself, expanded and complete
**Size**
Small — fits in an LLM context window
Large — fetched on demand for deep answers
**Usage**
Entry point: agents read it first
Referenced from llms.txt; fetched when needed
In practice you add one link to `llms-full.txt` inside your `llms.txt` (for example under an `## Important links` section), so agents that need the full picture know where to find it — without loading it into context by default.
## Common llms.txt Format Mistakes (and How to Fix Them)
Here is a file that fails the specification in several ways at once:
```
Acme Documentation
## What is Acme?
- [Homepage](/)
- [API reference](/docs/api)
# this line was meant to be a comment
## Pricing info
Acme costs $20 per month for the Pro plan.
```
What's wrong, and how to fix it:
- **Missing H1.** "Acme Documentation" has no leading `#`, so the file has no required title element. Fix: start with `# Acme Documentation`.
- **Relative URLs.** `(/)` and `(/docs/api)` are paths, not addresses. Fix: use absolute URLs like `(https://acme.dev/docs/api.md)`.
- **Hash-comment misuse.** The `# this line was meant to be a comment` line starts with `#`, so parsers treat it as a heading — llms.txt has no comment syntax. Fix: remove it, or use an HTML comment (``) in the notes area.
- **H2 used for prose.** `## Pricing info` contains a paragraph, but H2 sections must contain file lists of links. Fix: move the prose into the free-form notes above the first `##`, and keep pricing links (absolute, with descriptions) in the list.
- **Missing descriptions.** Links without a `: note` give agents no hint of what's behind them. Fix: add a short description to every list item.
Not sure if your file is valid? Run it through the [free llms.txt validator](/checker) — it checks the structure against the spec and scores AI-readiness in seconds.
## llms.txt Format FAQ
### Is llms.txt an official standard?
Not yet. It's an open, community-driven proposal maintained on GitHub by AnswerDotAI (the llmstxt.org authors), not a W3C or IETF standard. That said, adoption is broad and growing — Chrome Lighthouse audits for it, and platforms like Mintlify, GitBook, Wix, and Yoast generate the file automatically.
### Where should the llms.txt file be placed?
At the site root, next to `robots.txt`: `https://your-domain.com/llms.txt`. It can also live at any subpath (e.g. `/docs/llms.txt`) to cover only the pages under that path — agents use the most specific file that applies.
### How often should I update my llms.txt file?
Whenever your content structure changes meaningfully: new sections, renamed pages, removed or updated documentation. A monthly review is plenty for most sites, and a full regeneration after any redesign. Stale links are the fastest way to lose an agent's trust in your file.
## Keep reading
[
### 📚 Real-World Examples
Copy-paste llms.txt files for SaaS, e-commerce, blogs and corporate sites.
](/blog/llms-txt-examples/)[
### 🛠️ How to Create llms.txt
A step-by-step tutorial: build your file in 10 minutes and validate it.
](/blog/how-to-create-llms-txt/)
## Ready to write a spec-perfect llms.txt?
Generate your file in under a minute, then validate it against the specification — free, no sign-up.
[⚡ Generate your llms.txt](/generator) [Check my file](/checker)
---
# How to Add llms.txt to Hugo (Static or Auto-Generated, 2026)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-hugo/
Hugo tutorial · Static or generated · 7 minutes
# How to Add llms.txt to Hugo
Hugo is a static site generator written in Go that ships as a single binary and builds plain files into `public/`. It has no native `llms.txt` feature — and it barely needs one. A file in `static/` becomes your site root file automatically, and a custom output format can generate an `llms.txt` from your real content at build time. Both routes are covered here.
## Hugo and llms.txt: No Built-In Support (Yet)
As of 2026 there is no native llms.txt feature in Hugo and no Hugo entry on the integrations list at llmstxt.org. The official Hugo forum hosts a thread titled "Support for llms.txt standard for AI crawlers" that tracks community interest, but the practical answer is that you do not need a plugin. Two long-standing Hugo mechanisms do all the work: the `static/` directory, which is copied verbatim into the output root on every build, and custom output formats, which let Hugo emit any file type next to your HTML.
Hugo's model also keeps the format honest: what `public/` contains is exactly what your CDN serves, with no server-side rendering in between. Hugo powers a large share of documentation and blog sites on the web, and several Hugo developers already publish the file this way — for example [Duncan Mackenzie's Hugo blog](https://www.duncanmackenzie.net/blog/llms-txt-support/) serves `/llms.txt` plus a Markdown version of every post.
## Method 1: static/llms.txt, the 60-Second Route
Create `static/llms.txt` in your project root. Hugo copies every file in `static/` to the root of `public/` without touching it, so on the next `hugo` build the file is live at `https://your-domain.com/llms.txt`. Use the [llmstxt.org v2 spec](/blog/llms-txt-v2/) structure:
```
# Acme Docs
> Acme Docs is the official documentation for the Acme platform.
> It covers installation, configuration, the REST API and troubleshooting.
## Getting Started
- [Quickstart](https://acme.com/docs/quickstart): first project in five minutes
- [Authentication](https://acme.com/docs/auth): API keys and access tokens
- [Deployment](https://acme.com/docs/deploy): production setup for self-hosting
## API Reference
- [REST API](https://acme.com/docs/api): every endpoint, parameter and error
- [Webhooks](https://acme.com/docs/webhooks): events your server can subscribe to
```
Run `hugo` and check `public/llms.txt` appeared, then deploy `public/` to Netlify, Cloudflare Pages or any static host. Keep this method when your curated list changes rarely; a hand-written file is the easiest thing to audit and the easiest to forget. Remember absolute URLs — never relative paths, which agents resolve inconsistently.
## Method 2: Auto-Generate llms.txt at Build Time
When your content changes weekly, a generated file stays honest. Hugo's custom output formats emit arbitrary file types. First, register a plain-text format in `hugo.toml` whose `baseName` is `llms`:
```
[mediaTypes."text/plain"]
suffixes = ["txt"]
[outputFormats.llms]
mediaType = "text/plain"
baseName = "llms"
isPlainText = true
root = true
```
Next, create a content page that opts into the format. Hugo renders it through a template named for the format's suffix (`.txt`), so add `layouts/_default/single.txt` that outputs the page body with `.RawContent` and then walks your site to append link lines:
```
# content/llms.md
+++
title = "Acme Docs"
outputs = ["llms"]
+++
# Acme Docs
> Acme Docs is the official documentation for the Acme platform.
> It covers installation, configuration, the REST API and troubleshooting.
```
```
{{ .RawContent | safeHTML }}
## Pages
{{- range .Site.Pages -}}
- [{{ .Title }}]({{ .Permalink }})
{{- end -}}
```
Run `hugo` and the file is written to `public/llms.txt`. Filter `.Site.Pages` in the template to include only the sections you curate — dumping every tag and taxonomy page defeats the purpose of a curated index. A minimal, ready-to-read reference implementation is the [roverbird/llms-hugo](https://github.com/roverbird/llms-hugo) repository on GitHub, and the same technique can emit Markdown versions of individual pages for agents that want full text, the pairing the [llms.txt vs llms-full.txt guide](/blog/llms-txt-vs-llms-full-txt/) explains.
## The Hugo-Specific Checklist
- **One source of truth** — a `static/llms.txt` file and a generated `content/llms.md` page that both target /llms.txt will overwrite each other unpredictably. Pick one route.
- **Set baseURL before building** — generated links use `.Permalink`, which is built from your `baseURL`. A placeholder baseURL ships broken links into your llms.txt. Also avoid `relativeURLs = true`, which produces relative paths agents resolve inconsistently.
- **Absolute URLs only** — `https://acme.com/docs/quickstart`, never `/docs/quickstart`.
- **Mind multilingual sites** — a file in `static/` is served for every language version. If each language needs its own file, give it its own directory via the `staticDir` option in the language block.
- **Curate, don't dump** — list the 5–15 pages that answer real questions and skip marketing and tag pages. Curation rules live in the [best practices guide](/blog/llms-txt-best-practices/).
- **Rebuild to publish** — Hugo has no server runtime. Every change to the file or your content requires a build and deploy, so wire llms.txt into your existing release pipeline.
## Verify Before AI Engines Read It
After deploying, confirm the file serves plain text with a 200 status:
```
curl -s -o /dev/null -w "%{http_code} %{content_type}\n" https://your-domain.com/llms.txt
# Expect: 200 text/plain
```
Hugo sites are usually deployed through CI, so check the build that is actually live — a stale `public/` folder is the most common reason a fresh llms.txt goes missing. Then paste the URL into the free [llms.txt checker](/checker) to validate the H1, the blockquote summary, absolute URLs and the section structure.
Draft the initial file in seconds by feeding your Hugo sitemap into the [free llms.txt generator](/generator), then verify the output with the [checker](/checker) before wiring up a template. Hugo makes the file easy to ship and easy to keep honest — a correct, current `llms.txt` is one of the highest-leverage AI visibility changes a static site can make this week.
---
# How to Add llms.txt to Jekyll (GitHub Pages, 2026)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-jekyll/
Jekyll tutorial · GitHub Pages · 7 minutes
# How to Add llms.txt to Jekyll
Most GitHub Pages blogs run on Jekyll, which has no native `llms.txt` feature. Its file rules make it one of the easiest SSGs to fix, though: a static file in 60 seconds, auto-generation from your posts with a few lines of Liquid (no plugin, so GitHub Pages accepts it), or the `jekyll-llms-output` plugin for `llms.txt` plus `llms-full.txt`.
## The Jekyll File Rule That Makes This Easy
Every root-level file in Jekyll meets one of two fates: with YAML front matter it is rendered through Liquid and written to the same path in `_site`; without front matter it is copied verbatim. A root file named `llms.txt` with a `layout: null` front matter block gets its Liquid processed and published at `/llms.txt` as plain text — no custom output formats needed.
## Real-World Case: jtemporal.com
Developer advocate Jessica Temporal (Jekyll + GitHub Pages) [published her exact setup](https://jtemporal.com/how-to-create-llms-txt-in-jekyll). Her live file, curl-verified on 2026-09-07:
```
$ curl -sI https://jtemporal.com/llms.txt
HTTP/2 200
content-type: text/plain; charset=UTF-8
content-length: 32551
```
The `_Generated` timestamp proves Liquid wrote the file at build time. Her setup: `layout: null` front matter, a `{% for post in site.posts %}` loop enumerating every post, and `llms.txt` in the `_config.yml` `include` list (custom include lists omit unlisted files).
## Method 1: A Static File in 60 Seconds
If your site is small and stable, create `llms.txt` at the Jekyll root _without_ front matter — it is copied into `_site` unchanged:
```
# Acme Blog
> Acme Blog publishes practical tutorials on web performance and SEO.
## Essentials
- [About](https://acme.com/about/): who writes here and why
- [Archive](https://acme.com/archive/): every post, newest first
## Popular Posts
- [Core Web Vitals in 2026](https://acme.com/cwv-2026/): what changed and what to fix
- [Heading Hierarchy](https://acme.com/headings/): H1 to H6 done right
```
Caveat: if your `_config.yml` defines a custom `include` list, add `llms.txt` to it — otherwise the file never reaches the build output. The cost of this route is maintenance: every new post is a hand edit, and a file that drifts from your content is one agents follow into stale territory.
## Method 2: Auto-Generate from site.posts with Liquid
For a blog, the no-plugin template is the sweet spot: no gem, so it runs inside the stock GitHub Pages build, and it regenerates on every `bundle exec jekyll build`:
```
---
layout: null
---
# Acme Blog
> Acme Blog publishes practical tutorials on web performance and SEO.
> Generated: {{ site.time | date: "%Y-%m-%d" }}
## Posts
{% for post in site.posts %}
- [{{ post.title }}](https://acme.com{{ post.url }}): {{ post.excerpt | strip_html }}
{% endfor %}
```
`layout: null` emits plain text instead of your theme's HTML wrapper. URLs are hard-coded because `site.url` stays empty until you set `url:` in `_config.yml`. A one-line description per post becomes the link description agents read.
Variant: link to raw Markdown sources using `post.path`: `raw.githubusercontent.com/USER/REPO/main/{{ post.path }}`. Clean Markdown is the most LLM-friendly form — like the per-page `.md` files the [llms.txt v2 spec](/blog/llms-txt-v2/) defines.
## Method 3: jekyll-llms-output for llms.txt + llms-full.txt
For spec-grade output without writing Liquid, [jekyll-llms-output](https://github.com/abhinavs/jekyll-llms-output) generates `/llms.txt` and `/llms-full.txt`. In curated mode, drop a `_data/llms.yml` mapping sections to links; without it, auto mode emits one `## Section` per collection, one bullet per document. Add the gem to your Gemfile:
```
# Gemfile
gem "jekyll-llms-output", group: :jekyll_plugins
# terminal
bundle install
```
Paired with [jekyll-markdown-output](https://github.com/abhinavs/jekyll-markdown-output), it writes a clean `.md` sibling for every page, and `llms-full.txt` concatenates every document's full body under `# Title` headers. See the [llms.txt vs llms-full.txt guide](/blog/llms-txt-vs-llms-full-txt/).
**The GitHub Pages catch:** Pages only allows a whitelist of Jekyll plugins, and this gem is not on it — building from a `gh-pages` branch means it silently never runs. Build in CI and deploy the resulting `_site` directory instead.
## Static, Liquid or Plugin?
Approach
Setup
GitHub Pages
Auto-updates
llms-full.txt
Static root file
1 minute
Yes
No
No
Liquid template
~5 minutes
Yes, no plugin needed
Every build
No
jekyll-llms-output
~10 minutes
Via CI only
Every build
Yes
For a typical blog the Liquid route offers the best return: it stays inside GitHub Pages' constraints and your post list can never go stale. Choose the plugin when you also want `llms-full.txt` or per-page Markdown and already build in CI. Either way, follow [llms.txt best practices](/blog/llms-txt-best-practices/) — a summary per link beats a raw title.
## Verify and Maintain
- **Confirm the include list** — with a custom `include`, add `llms.txt` or it never publishes.
- **Set `url:` in `_config.yml`** — or hard-code absolute URLs.
- **Use `layout: null`** for the Liquid route — without it Jekyll serves HTML as text.
- **Verify the deployed file** — confirm it serves as plain text:
```
curl -s -o /dev/null -w "%{http_code} %{content_type}\n" https://yourblog.com/llms.txt
# Expect: 200 text/plain
```
Paste the URL into the free [llms.txt checker](/checker), or feed your blog's sitemap into the [free llms.txt generator](/generator), then let Liquid keep it current — a zero-maintenance byproduct of the build that already publishes your blog.
---
# llms.txt Maintenance: How to Keep Your File From Going Stale (2026)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-maintenance/
Maintenance · September 2026
# llms.txt Maintenance: Keeping the Map Honest
Publishing the file takes ten minutes. The interesting question is the one nobody writes about: what happens to it over the next two years, while your URLs, your product names and your documentation structure keep moving.
## Staleness Is a Content Bug, Not a Delivery Bug
Delivery problems — the wrong content type, a rewrite rule, a blocked crawler — are loud and fixable in an afternoon; they are covered in [the delivery-level troubleshooting guide](/blog/llms-txt-not-working/). Staleness is quiet. The file still returns `200`, the format is still valid, and every link is still syntactically perfect — but three of them now 404, and the H2 section called _Pricing_ points at a page you retired last spring. An agent that follows your map and hits 404s learns that your curated file is less reliable than your sitemap, which is the opposite of the file's purpose.
There is a second, slower cost: summaries drift. If the blockquote under your H1 still describes a product you repositioned six months ago, that sentence travels with every retrieval that touches your domain. Curation is a claim about your site, and claims expire.
## What Actually Changes (and What Each Change Forces)
Most site changes need no edit at all. These six do:
Change
Edit required
A new top-level section or product line
Add an H2 and its 3–6 links; update the summary if the positioning moved
URL restructuring or a rename
Replace the URL — do not keep the old one alongside the redirect
A page in the file was deleted
Remove the line; leave the section if it still has three or more links
A page is now your best entry point
Move it into the first section — order is editorial information
A new language or regional site
Decide root file versus per-language files before adding URLs
Content moved to a subdomain you do not advertise
Point at the canonical host, never the alias an agent cannot index
Nothing on that list requires a new format, a version bump or a migration. It is editing a markdown file you already own, which is why the expensive part is remembering to do it.
## Layer 1: Make Dead Links Fail Your Build
The single highest-value habit is a link check that runs on every deploy and turns a 404 into a failed pipeline. A file that lists twenty URLs is small enough for a plain shell script — no dependency, no framework:
```
#!/usr/bin/env bash
# scripts/check-llms-links.sh — exits 1 if any URL in llms.txt is not 200
fail=0
while read -r url; do
code=$(curl -s -o /dev/null -w '%{http_code}' -L -A "GPTBot/1.5" "$url")
if [ "$code" != "200" ]; then
echo "DEAD LINK $code $url"
fail=1
fi
done < <(grep -oE 'https?://[^ )]+' public/llms.txt | sort -u)
exit $fail
```
Run it locally before you push, and once after every deploy against the production host. The `-L` flag matters: it follows redirects, so a link that only works through a hop shows up as a pass. If you would rather not allow redirects at all — several engines will not follow them for this file — drop `-L` and treat a `301` as a failure too.
Add the same script to your CI workflow and to a weekly scheduled run. Weekly is enough: a link that breaks on Monday and is reported on Friday costs you four days of failed fetches, not four months.
## Layer 2: Generate What Can Be Generated, Curate the Rest
Split the file into the part a machine can keep current and the part only you can write. Sitemap derived lists of every URL belong to the automated side; the `/generator` builds exactly that from a `sitemap.xml` if you want a starting point. The header summary, the section names and the ordering are the editorial side, and they should be reviewed by a human — automation that refreshes the whole file nightly tends to produce an alphabetical dump that loses the curation that made the file useful in the first place. The [nine best-practice rules](/blog/llms-txt-best-practices/) are mostly about that editorial layer.
## Layer 3: A Quarterly Review You Can Actually Finish
Ten minutes, once a quarter, in this order:
- Run the link check and delete or fix every failure.
- Read the summary line aloud. Would a stranger understand what the site is today?
- Check the first section: is your single best page still in it?
- Compare against the sitemap — did a whole section appear that the file never learned about?
- Confirm the file still returns `200` with `text/plain` from the canonical host.
- Check the cache headers on that response, so you know how long a wrong version persists.
- If you measure fetches at all, compare the last quarter's numbers against the previous one — see [how to measure llms.txt traffic](/blog/measure-llms-txt-traffic/).
## When Not to Touch It
Not every week is a version. Adding a blog post that is not one of your best pages is not a reason to edit the file — that is what `sitemap.xml` is for, and the two files have different jobs; [the comparison](/blog/llms-txt-vs-sitemap-xml/) covers the split. Constant churn also fights caching: agents and proxies that hold a copy of your file have to reconcile changes, and a file that reorders itself weekly is noise. Keep the URL fixed, keep absolute links, and let the file improve by accretion.
Once the automation is in place, staleness stops being a memory problem. The build fails, you fix the link, and the map stays honest. If you have not published a file yet, start with the [generator](/generator), then push the production URL through the [checker](/checker) before you automate anything.
---
# 6 llms.txt Mistakes That Make AI Engines Ignore Your File
> Source: https://llmstxtgenerator.dev/blog/llms-txt-mistakes/
Common errors · August 2026 · AI readiness
# 6 llms.txt Mistakes That Make AI Engines Ignore Your File
You published `/llms.txt`, so why do AI engines seem not to care? In most cases the file is technically "there" but fails in one of six ways: wrong placement, format errors, dead links, a site-wide dump, marketing language, or the wrong success metric. Here is how to find and fix each one.
## Mistake 1: The file is where agents can't find it
AI agents look for llms.txt at the root of your domain over HTTPS — the same place as robots.txt. If your file only works on `www` while the crawler checks the apex domain, or a redirect chain swallows the request, the agent gets a 404 and moves on. The [llmstxt.org](https://llmstxt.org) spec also allows subpath files such as `/docs/llms.txt`, which cover the pages under their path — useful when you only control a directory, but a root file is what most agents check first.
- Serve the file at `https://yourdomain.com/llms.txt` — lowercase, exact name, root path.
- Make sure HTTPS redirects work before that URL, with no infinite loops.
- Confirm the status code, not just that the file exists on disk:
```
curl -I https://example.com/llms.txt
# expect: HTTP/2 200
curl -s https://example.com/llms.txt | head -5
```
## Mistake 2: The format fails validation
The spec is deliberately small, and the most common failures are the simplest. The H1 with your site name must be the _very first line_ — the only required element. A byte-order mark, an HTML comment, or a blank line pushed in front of it by your CMS is enough to break the file for strict validators.
- First line: `# Your Site Name`, nothing before it.
- Blockquote summary directly after the H1.
- Absolute URLs only: `https://...` — never `/about`.
- One H1; markdown link list items under H2 headings such as `## Important links`.
Run every revision through a validator and fix everything reported as an error, not just the warnings. The [free validator](/checker) checks exactly these rules — missing H1, H1 not on the first line, relative URLs, unrecognized lines — and scores the result 0–100.
## Mistake 3: The links lead nowhere useful
llms.txt is a list of links. If those links 404, redirect to a login wall, or point at HTML pages buried in navigation and cookie banners, an agent that follows them gets no clean context. Links rot fastest after a redesign — a fact every site with an old llms.txt learns the hard way.
- Link to markdown versions of pages when you can; both `.md` and `.html.md` URL forms are valid in the v2 spec.
- Verify every listed URL returns 200, and re-check after each deploy.
- Add a short description after each link so an agent knows what the page is about before fetching it.
## Mistake 4: You dumped the whole site in
sitemap.xml lists _every_ URL for search crawlers; llms.txt should curate the handful of pages that best represent your site for an AI agent that reads the file inside its context window. Hundreds of links dilute the signal — an agent skims, and your five most important pages disappear into the noise. See [llms.txt vs sitemap.xml](/blog/llms-txt-vs-sitemap-xml/) for the full comparison.
llms.txt
sitemap.xml
Purpose
Curated index for AI agents
Complete URL list for search crawlers
Typical size
5–15 links
Thousands of URLs
Readers
AI engines and agents
Google and other search engines
Detail
One-line description per link
Metadata and priorities
Keep the file under a few dozen links. Secondary material can live in an `Optional` section, which agents may skip — put your core value in the first sections.
## Mistake 5: You wrote marketing copy, not context
Agents summarize from what the file actually says. "Industry-leading," "next-generation," and "empowering businesses" carry zero information — an AI model has no way to verify them, so it ignores them. Write like you are explaining the site to a smart colleague:
```
# Acme Inc
> Acme is the industry-leading provider of next-generation
> solutions that empower businesses worldwide.
```
Compare that with a summary an agent can actually use:
```
# Acme Inc
> Acme builds payroll software for US small businesses,
> with 40,000 customers and a 12-person support team.
```
The same rule applies to link descriptions: "About us" is weaker than "About: founding story, team, and 12 years of payroll history." Specificity is what makes an agent confident enough to cite you.
## Mistake 6: You expect rankings — and never re-check
In June 2026, Google updated its AI optimization guide to state plainly that llms.txt has no effect — positive or negative — on Search rankings or AI Overviews. Independent data points the same way: ALLMO pulled 94,614 cited URLs from 11,867 AI answers across five AI platforms and found exactly one /llms.txt page among them (0.001%), and Limy's analysis of 500M+ LLM bot traffic events shows GPTBot, ClaudeBot, PerplexityBot and Google-Extended overwhelmingly crawl HTML directly and skip the file. We covered the evidence in detail in [does llms.txt affect SEO?](/blog/does-llms-txt-affect-seo/).
None of that makes llms.txt useless — it is a control layer. Adoption is real and rising (about 5.6% of the top 10,000 sites had a valid file in June 2026, and Shopify pushed one to every store by default in 2026), and Chrome Lighthouse audits for the file by default. The mistake is judging it by ranking metrics instead of by whether it stays correct. Every quarter, and after every site change, run this checklist:
- `curl -I` returns 200 at the exact root URL.
- First line is still the H1 with your site name.
- All links are absolute and return 200.
- The file is under a few dozen links and still curated.
- Summary and descriptions are factual, not hype.
- Re-validated after every deploy.
Fix the six mistakes above and your llms.txt does its real job: it gives agents that do check it a fast, clean, accurate path to your best content. Generate a compliant file from your sitemap with the [generator](/generator), then confirm every point above with the [validator](/checker).
---
# How to Add llms.txt to MkDocs (Material for MkDocs, 2026)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-mkdocs/
MkDocs tutorial · Plugin or static · 6 minutes
# How to Add llms.txt to MkDocs
MkDocs is the Python-based documentation generator behind a large share of open-source docs sites, usually dressed in the Material for MkDocs theme. It has no native `llms.txt` feature — but it barely needs one. A plain file in `docs/` is copied straight to your site root on every build, and the community `mkdocs-llmstxt` plugin generates a spec-compliant file from your own config. Both routes are covered here.
## MkDocs and llms.txt: No Native Support (Yet)
Software documentation is where llms.txt gets used most heavily — the llmstxt.org spec notes that coding agents follow the file to find API references and tutorials — and MkDocs is one of the most common engines behind that documentation. Yet the integrations list at llmstxt.org still has no MkDocs entry: it names hosted platforms like Mintlify and GitBook and CMS plugins like Yoast, but nothing for self-hosted MkDocs.
The gap is visible in the ecosystem. A community discussion on the Material for MkDocs repository ([squidfunk/mkdocs-material #8384](https://github.com/squidfunk/mkdocs-material/discussions/8384)) asks for native llms.txt generation, and the answer so far is a plugin. Timothée Mazzucotelli's `mkdocs-llmstxt` is the reference implementation, and real projects use it in production — the [Instructor documentation](https://python.useinstructor.com/blog/2025/08/29/mkdocs-llmstxt-plugin-integration) pipeline regenerates its llms.txt automatically on every deploy. Note that MkDocs pages can sit at any subpath, so the file works for versioned docs too — the [llms.txt v2 spec](/blog/llms-txt-v2/) defines how subpath files and Markdown page versions behave.
## Method 1: Ship a Static docs/llms.txt
MkDocs copies every non-Markdown file in `docs/` verbatim into the output directory, preserving its path — the same mechanism that makes `docs/robots.txt` work. Create `docs/llms.txt` and your next `mkdocs build` serves it at `https://docs.example.com/llms.txt`:
```
# Acme Docs
> Acme Docs is the official documentation for the Acme platform.
> It covers installation, configuration and the REST API.
## Getting Started
- [Quickstart](https://docs.example.com/getting-started/): first project in five minutes
- [Authentication](https://docs.example.com/authentication/): API keys and scopes
## API Reference
- [REST API](https://docs.example.com/api/): every endpoint, parameter and error
- [Webhooks](https://docs.example.com/webhooks/): events your server can subscribe to
```
This is the right route for small sites whose structure changes rarely. The cost is manual maintenance: every new page is a hand edit, and when your docs drift from the file, agents follow a stale index. Keep URLs absolute — never relative paths — and remember the file only describes pages under its own path, so docs served at `/en/latest/` want a file in that same tree.
## Method 2: Auto-Generate with the mkdocs-llmstxt Plugin
Install the plugin and add it to `mkdocs.yml`. Three keys matter: `site_url` is required (all links are built from it), `site_description` becomes the blockquote, and your `sections` define the curated heading groups, with optional per-file descriptions and glob support:
```
# terminal
pip install mkdocs-llmstxt
```
```
# mkdocs.yml
site_name: Acme Docs
site_url: https://docs.example.com/
plugins:
- search
- llmstxt:
markdown_description: Long-form context an agent needs before navigating.
sections:
User Guide:
- index.md: What Acme Docs covers
- getting-started.md
- deployment.md
API Reference:
- api/*.md
```
Every `mkdocs build` now writes `site/llms.txt`. The plugin parses the rendered HTML (BeautifulSoup + Markdownify) and converts it back to clean Markdown, so executed code blocks, Jinja-generated snippets and API documentation survive in the output — and each page you list is published as its own Markdown file, with llms.txt linking those `.md` URLs. Set `full_output: llms-full.txt` to also emit the full-text dump, the pattern the [llms.txt vs llms-full.txt guide](/blog/llms-txt-vs-llms-full-txt/) compares. For versioned builds on Read the Docs, the `base_url` option rewrites every generated link to a subdirectory such as `https://docs.example.com/en/0.1.34`.
One honesty note: the plugin's repository is in maintenance mode — the author moved on to another project and is looking for a maintainer. It is widely used and stable, but pin your version and keep the static-file fallback in mind. An alternative, `mkdocs-llmstxt-md`, takes a raw-Markdown approach (enabled as `llmstxt-md`) and generates both llms.txt and llms-full.txt from your source files by default, deriving sections from your `nav` when you configure nothing.
## Which Route Fits Your Docs?
Approach
Setup
Stays current
Extra outputs
Static `docs/llms.txt`
1 minute
Manual edits only
None
`mkdocs-llmstxt`
~5 minutes
Every build
.md page versions, optional llms-full.txt
`mkdocs-llmstxt-md`
~5 minutes
Every build
llms.txt + llms-full.txt by default
A static file is fine while your sitemap is small and stable. The moment docs change weekly — new guides, moved pages, added API endpoints — a generated file stays honest, because the llms.txt an agent reads is only as trustworthy as its last update.
## The MkDocs-Specific Checklist
- **Set a canonical site\_url** — plugin links are derived from it. A placeholder `site_url` ships broken or relative links into your llms.txt.
- **Curate, don't dump** — list the pages that answer real questions: getting started, auth, deployment, the API. Skip the 404 page, the changelog and search indexes. Curation rules live in the [best practices guide](/blog/llms-txt-best-practices/).
- **Write useful descriptions** — "what problem this page solves" beats "documentation for module X" when an agent decides which link to fetch.
- **Handle versioned docs per tree** — each language or version wants its own subpath file; use `base_url` in per-version builds rather than one file claiming to cover all of them.
- **Rebuild to publish** — MkDocs is static: llms.txt changes only when CI runs `mkdocs build`, so wire it into your existing deploy pipeline.
## Verify Before Agents Read It
After deploying, confirm the file serves as plain text with a 200 status:
```
curl -s -o /dev/null -w "%{http_code} %{content_type}\n" https://docs.example.com/llms.txt
# Expect: 200 text/plain
```
Then paste the URL into the free [llms.txt checker](/checker) to validate the H1, blockquote summary, absolute URLs and section structure against the v2 spec. Draft the initial file in seconds by feeding your MkDocs sitemap into the [free llms.txt generator](/generator), then switch to the plugin once your sections are settled. MkDocs makes the file trivial to ship and easy to keep current — for a documentation site, that is the highest-leverage AI visibility change you can make this week.
---
# llms.txt for Multilingual Websites: One File or One Per Language? (2026)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-multilingual-sites/
Global SEO · September 2026
# llms.txt for Multilingual Websites: One File or One Per Language?
Multilingual sites are the awkward case for `llms.txt`. English-only sites write one file and move on; global sites have to decide whether a single curated index can represent five languages, or whether each language version deserves its own. The spec does not answer the question — but it does give you the one rule that settles it.
## The Rule That Settles It: Scope
The [specification](https://llmstxt.org/) is explicit about what a file covers: an `llms.txt` file describes the URLs _under its own path_, and where more than one file applies to a given page, agents should use the **most specific** one. The file may sit at the root or at any subpath.
That is the whole ballgame. A root `/llms.txt` covers your entire domain, every language included. A file at `/de/llms.txt` covers everything under `/de/` and **wins** for those URLs. The v2 spec revision (modified August 10, 2026) kept this wording unchanged — see [what changed in v2](/blog/llms-txt-v2/) for the rest of the update, including the `rel="describedby"` link relation that lets a page point at the file covering it.
Note also that the format is language-neutral: nothing in it requires English. Descriptions should be written in the language of the content they describe, because the consumer is often a localised assistant answering a question in that language.
## Three Structures That Work
Structure
Best for
Maintenance cost
**A.** Single root file, one H2 per language
2–3 languages, shorter sites
Lowest — one file to edit
**B.** Root index + `/en/llms.txt`, `/de/llms.txt` per locale
4+ languages, separate locale paths
Higher — one file per locale, but each stays short
**C.** Root file + `Optional` section for secondary languages
One dominant market, one or two minor ones
Low, but secondary languages are second-class
## Structure B in Practice
The pattern that keeps showing up in practitioner write-ups is the sitemap-index shape: a small root file at the domain root whose job is to route agents to the right language file, plus one real file per language root. The root file stays stable for years because it changes only when you add or retire a language:
```
# Example Store
> Global retailer selling household goods in 6 languages.
## Language editions
- [English](https://example.com/en/llms.txt): full product and support content, primary market
- [German](https://example.com/de/llms.txt): full product and support content, DACH market
- [French](https://example.com/fr/llms.txt): product catalogue and shipping policy
- [Japanese](https://example.com/ja/llms.txt): product catalogue and support centre
```
Each locale file then does the real work under its own path — the top-level sections an assistant needs, with one-line descriptions, exactly as the [best-practice rules](/blog/llms-txt-best-practices/) describe. Keep the count of curated links per locale in the 20–40 range. If you are generating a first draft from a sitemap, [the generator](/generator) can start you off, but split the output by locale before you publish it — dumping six languages into one file is the failure mode this whole structure exists to avoid.
One caveat specific to subpaths: your locale directories must be real paths (`/de/`, `/ja/`), not cookie- or `Accept-Language`\-based switching. If the language is decided at request time, there is no stable path for a file to cover, and there is no URL for an agent to cite.
## Keep It Aligned With hreflang and Canonicals
- **List canonical URLs only.** If `/de/produkte/` is the German canonical, do not also list the parameterised or tracking variants.
- **One entry per language version, not per region.** Your `hreflang` annotations already tell engines that `de-DE` and `de-AT` share a page. Duplicating it in `llms.txt` just adds noise.
- **Never list a page that redirects.** A locale link that 301s to another language is worse than no entry at all — the agent burns a request and may cite the wrong version.
- **Self-reference with `rel="describedby"`.** Adding the link relation on localised pages (in the HTML `head` or as an HTTP `Link:` header) tells agents which file covers that page, which is how a deep `/ja/...` URL gets routed to `/ja/llms.txt`.
## What Big Multilingual Sites Actually Ship
We checked a sample of global sites with `curl` on September 14, 2026. The picture is consistent: the root file is becoming normal, per-language files are not there yet — and the most multilingual site of them all publishes nothing at all.
URL checked
Status
`docs.stripe.com/llms.txt`
200, `text/markdown`, ~90 KB
`docs.stripe.com/de/llms.txt`
404
`developers.cloudflare.com/llms.txt`
200, `text/plain`
`developers.cloudflare.com/zh-cn/llms.txt`
404
`www.shopify.com/llms.txt`
200, `text/plain`
`www.shopify.com/de/llms.txt`, `/fr/llms.txt`
404 (both)
`about.gitlab.com/llms.txt`
200, `text/plain`, ~10 KB
`en.wikipedia.org/llms.txt`, `/es/`, `/ja/`
404 (all three)
`www.ikea.com/us/en/llms.txt`
404
Two lessons. First, publishing any valid file puts you ahead of most of the web. Second, if you already run locale paths, per-language files are a differentiator rather than a catch-up move: Stripe's single 90 KB file is genuinely useful for English-speaking developers, and useless to a German-speaking one asking about payment methods in German.
## A Five-Minute Verification Pass
Check every path you claim to support, from one shell, and confirm the file is served as text and not as an HTML error page:
```
curl -sI https://example.com/llms.txt
curl -sI https://example.com/en/llms.txt
curl -sI https://example.com/de/llms.txt
# content type and first lines
curl -s https://example.com/de/llms.txt | head -20
```
- Every URL returns **200** with `text/plain` or `text/markdown` — a 302 to your homepage counts as missing.
- Every link inside each file resolves, and every link under `/de/` really is German content.
- The H1 and summary in each locale file are written in that locale, not machine-translated boilerplate.
Run the files through the [llms.txt checker](/checker) after each language launch, and re-run it whenever you add a locale. Structure follows content: add the language version first, then give it a file.
The short answer to the question in the title: one file if you have two or three languages, one file per language once locale paths are a real part of your site. Either way, make sure the file that applies to a page is the most specific one available — that is the only rule the spec asks you to respect.
---
# How to Add llms.txt to Next.js (App Router, 2026)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-nextjs/
Next.js tutorial · App Router · 6 minutes
# How to Add llms.txt to Next.js
Next.js still has no built-in llms.txt file convention — but shipping one takes minutes. Here are the three ways to do it, from a 30-second static file to a fully dynamic, ISR-cached route handler.
## Why Next.js Sites Are the #1 llms.txt Use Case
The llms.txt convention exists mainly for software documentation, where coding agents follow the file to find API references and tutorials. That is exactly the content Next.js is most often used to build. The early adopters read like a who's who of the Next.js ecosystem: Vercel publishes [vercel.com/docs/llms.txt](https://vercel.com/docs/llms.txt), the Next.js docs themselves serve [nextjs.org/docs/llms.txt](https://nextjs.org/docs/llms.txt), and Strapi, Vultr, Langbase, Checkly and CircleCI all publish one at their docs roots.
If your Next.js project is a SaaS site, a docs portal, or a dev-tool marketing page, your audience is literally the people asking ChatGPT and Claude how your product works. An `llms.txt` is the cheapest way to make sure the answer they get links back to your canonical pages — this is core [GEO (generative engine optimization)](/blog/llms-txt-v2/) for framework-built sites.
## Method 1: Static File in public/ (30 Seconds)
The fastest option works on both App Router and Pages Router projects. Create a file at `public/llms.txt` (or `src/../public/llms.txt` with the standard `src/` layout — Next.js copies the whole `public/` folder to the root of the build output):
```
# Example SaaS Docs
> Example SaaS is a hosted API platform. This file maps the pages
> coding agents should read first: quickstart, API reference, guides.
## Quickstart
- [Quickstart](https://example.com/docs/quickstart): 5-minute setup with your first API call
- [Authentication](https://example.com/docs/auth): API keys and token scopes
## API Reference
- [REST API Reference](https://example.com/api-reference): all endpoints, params and errors
- [Webhooks](https://example.com/docs/webhooks): events your server can subscribe to
```
Deploy and the file is live at `https://your-domain.com/llms.txt`. One caveat: a static `public/` file always wins over a route handler with the same path — delete the static file if you later switch to Method 2.
## Method 2: Dynamic Route Handler (App Router)
For a file that should reflect live data — your CMS, your database, your sitemap — use a route handler. A folder named `llms.txt` works fine in the App Router (no rewrites needed):
```
// app/llms.txt/route.ts
import { siteName, siteDescription, sections } from '@/lib/llms-config';
export const dynamic = 'force-static'; // optional: pre-render at build time
export async function GET() {
const lines = [
'# ' + siteName,
'',
'> ' + siteDescription,
'',
];
for (const section of sections) {
lines.push('## ' + section.title, '');
for (const link of section.links) {
lines.push('- ' + link.name + ': ' + link.description + ' (' + link.url + ')');
}
lines.push('');
}
return new Response(lines.join('\n'), {
headers: { 'Content-Type': 'text/plain; charset=utf-8' },
});
}
```
Two details matter. First, the `Content-Type` must be `text/plain` — agents and validators expect plain text, not HTML. Second, with `export const dynamic = 'force-static'` the file is generated once at build time and served from the edge, so it costs nothing at request time.
## Method 3: ISR — Keep It Fresh Without Rebuilding
Static is fast but goes stale. If your docs change hourly (release notes, changelogs, pricing), swap `force-static` for an ISR interval:
```
// app/llms.txt/route.ts
export const revalidate = 3600; // regenerate at most once per hour
```
The route handler then rebuilds lazily on the next request after the interval expires, so your `llms.txt` never describes pages that no longer exist. You can even fetch your `sitemap.xml` inside the handler and filter it down to the genuinely important pages — a good way to avoid hand-maintaining the file. If you prefer a ready-made solution, the community [next-llms-txt](https://github.com/bke-daniel/next-llms-txt) plugin adds this as a build-time step.
## What Next.js Doesn't Ship (Yet)
In July 2025 a feature request ([vercel/next.js discussion #81182](https://github.com/vercel/next.js/discussions/81182)) asked for an `llms.txt` file convention in the app directory, mirroring `sitemap.ts` and `robots.ts`. A community PR (#90580) is in progress, but native support has not shipped. The good news: the file convention only matters for ergonomics. The `public/` file and route-handler approaches above produce exactly the same result at `/llms.txt`, so there is no reason to wait.
## Verify Your File
After deploying, check that the route serves plain text with a 200 status:
```
curl -s -o /dev/null -w "%{http_code} %{content_type}\n" https://your-domain.com/llms.txt
# Expect: 200 text/plain
```
Then paste the URL into our free [llms.txt checker](/checker) to validate the H1, the blockquote summary, absolute URLs, and section structure before AI engines start reading it.
## Best Practices Checklist for Next.js
- **Absolute URLs only** — `https://example.com/docs/start`, never `/docs/start`. Agents resolve links from the file's location.
- **Curate, don't dump** — 10–50 links beats 500. Point agents at canonical docs pages, not marketing landing pages.
- **Watch for interceptors** — make sure no middleware, redirect, or `next.config` rewrite catches `/llms.txt` before your route does.
- **Treat it like a README** — update the file whenever the underlying page changes; stale files erode agent trust over time.
- **Add llms-full.txt for deep docs** — for large reference sections, link a generated [llms-full.txt](/blog/llms-txt-vs-llms-full-txt/) companion so agents can fetch the full content on demand.
Once it's live, run your sitemap through the [free llms.txt generator](/generator) to build the initial curated file in seconds, then validate it with the [checker](/checker) before you deploy. A correct, current `llms.txt` is the single highest-leverage AI visibility change most Next.js sites can make this week.
---
# llms.txt Not Working? 7 Delivery-Level Fixes (Content-Type, Redirects, WAF, 2026)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-not-working/
Troubleshooting · September 2026
# llms.txt Not Working? Seven Delivery-Level Fixes
The file is in the repository, it passed validation, and no AI engine ever asks for it. Usually the problem is not the format but delivery: what your server actually returns when something requests `/llms.txt`.
## The 30-Second Test
Ask the server, not your editor. Two commands answer the status code, the content type, whether a redirect sits in front of the file, and whether the body is markdown or an HTML shell:
```
# headers only: status, Content-Type, Location, cache status
curl -sI https://example.com/llms.txt
# fetch the way an AI crawler would, and look at the first lines
curl -sA "GPTBot/1.2 (+https://openai.com/gptbot)" https://example.com/llms.txt | head -20
```
A healthy response is `200`, `content-type: text/plain; charset=utf-8` (or `text/markdown`), no `Location` header, and a body starting with a markdown H1 and the summary line. Anything else maps onto the table below.
Symptom
What it usually means
200, but `text/html`
Catch-all rewrite served your app shell — fix 1
`text/x-robots` or the wrong body
A rule mapped `/llms.txt` onto `/robots.txt` — fix 1
301 or 302 chain, or a redirect to the homepage
Canonical host mismatch — fixes 2 and 3
404 for that path only
Path case, extension not whitelisted, or wrong subpath — fix 3
403, 429 or an empty reply
WAF, bot rule or a `robots.txt` block — fix 4
200 locally, 404 in production
The file never shipped in the build — fix 5
Old content, or a 404 that comes and goes
Edge cache holding a previous response — fixes 6 and 7
## Fix 1: The Response Is HTML, Not Markdown
This failure is invisible in a browser, because the browser renders the app shell anyway. A single-page app or a rewrite-everything rule answers _every_ unknown path with `index.html` — including `/llms.txt`. A Common Crawl survey of published llms.txt files found the two signatures: sites that serve their SPA for any unknown path answer `text/html`, and sites whose rewrite sends `/llms.txt` to `/robots.txt` answer `text/x-robots`.
A static file in `public/` (or `static/`, or your host's root directory) is served as `text/plain` automatically — add an exception so the SPA fallback never sees that path. If a route handler generates the file, set the header yourself. The v2 spec, modified August 10, 2026, names no required MIME type, but it does require a markdown file at that path, and a response labelled `text/html` is not one.
## Fixes 2 and 3: Redirect Chains and Path Case
Redirects turn a working file into an unreliable one. Some agents follow them, some do not, and a chain across hosts doubles the failure surface. Publish the file on the exact hostname you advertise. A redirect whose destination is your homepage is worse than a 404 — a missing file that looks like a success at the HTTP level.
Path case is the quieter version of the same problem: object storage and Linux hosts are case-sensitive, so `/LLMS.TXT` and `/Llms.txt` are 404s even though `/llms.txt` exists, and NGINX rules that whitelist extensions can 404 a newly added `.txt` file. Subpath files follow the spec's scope rule — a file describes the URLs _under its path_, most specific wins — so `/docs/llms.txt` must be reachable at that path, not only at the root.
## Fix 4: Something Is Blocking the Request
Check the two layers in front of your origin. First, `robots.txt`: the major AI crawlers honour it, so a broad `Disallow: /` or a pattern such as `Disallow: /*.txt` blocks the llms.txt fetch itself — the file never gets read. Second, the firewall: bot protection, WAF rules and aggressive rate limiting return 403 or 429 to clients that do not look like browsers. Test with a crawler user agent and compare:
```
curl -s -o /dev/null -w '%{http_code}\n' -A "GPTBot/1.2" https://example.com/llms.txt
curl -s -o /dev/null -w '%{http_code}\n' -A "ClaudeBot/1.0" https://example.com/llms.txt
```
A 403 for the crawler while your browser sees 200 puts the block in the CDN or WAF, not your site. Staging has the same trap in another form: basic authentication on the whole hostname fails every fetch, which is harmless on staging and fatal if that rule is ever applied to production.
## Fix 5: The File Never Reached Production
Local success is not deployment success. The file may sit in a source folder the build does not copy into the output directory, the deploy may have run from a commit that predates it, or a dynamic route may catch `/llms.txt` and return something else. The only reliable check is against the production URL after every deploy — and if you deploy manually, confirm the file made it into that build, not the previous one.
## Fixes 6 and 7: Caches, Copies and the Verification Pass
An edge cache can keep serving a 404 your origin stopped returning, or a copy from before your last edit. Purge the exact path, and read the cache headers while you are there: an `age` measured in days on a file you changed an hour ago tells you which copy agents get. Watch for a second hostname on the same site, too — a staging or preview alias that answers `/llms.txt` is a file agents can find and you will forget.
Then run the pass once, and again after every deploy that touches the file:
- Canonical host, `curl -sI`: **200**, `text/plain` or `text/markdown`, no `Location`.
- Crawler user agent: same status as your browser, not a 403.
- Body: a markdown H1 and summary line, not an HTML document.
- Every link inside the file returns 200 — dead links waste the curation even when delivery is perfect.
- No `robots.txt` or firewall rule matches the path.
If the content is the suspect rather than delivery, the [six mistakes that get a file skipped](/blog/llms-txt-mistakes/) cover that side, the [crawler reference](/blog/robots-txt-ai-crawlers/) shows which tokens matter per engine, and the [measurement walkthrough](/blog/measure-llms-txt-traffic/) answers whether anyone reads the file at all. Run the URL through the [llms.txt checker](/checker) the way an agent would, and build the file itself with the [generator](/generator) if it is still a draft.
---
# How to Add llms.txt to Nuxt (public/, Server Routes, nuxt-llms, 2026)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-nuxt/
Nuxt tutorial · Nitro · nuxt-llms · 7 minutes
# How to Add llms.txt to Nuxt
Nuxt ships no `llms.txt` output of its own — but its `public/` directory makes Nuxt apps one of the easiest to fix: a static file in 60 seconds, a `server/routes/` handler for a live-generated file, or the official `nuxt-llms` module for build-time `llms.txt` and `llms-full.txt`. Nuxt's own documentation site, curl-verified for this guide, shows what a good file looks like.
## The Nuxt Rule That Makes This Easy
Nuxt serves every file inside `public/` at the site root, untouched by the build — the same mechanism that publishes `robots.txt` and `favicon.ico` (the directory was called `static/` in Nuxt 2). That one rule means an `llms.txt` file on disk becomes your live `/llms.txt` with zero routes, zero config and zero plugins. There is no native Nuxt convention beyond it, so your choice is really about how much of the file you want to maintain by hand.
## Method 1: A Static public/llms.txt in 60 Seconds
Create the file and publish — the whole setup is one command and one edit. A spec-compliant file needs only an H1; add a summary and grouped link lists for the context AI agents need:
```
# Acme Docs
> Acme Docs covers the Acme API: REST reference, SDKs and step-by-step tutorials.
## Essentials
- [Quickstart](https://acme.com/docs/quickstart): first API call in five minutes
- [API Reference](https://acme.com/api): every endpoint, error and rate limit
- [Guides](https://acme.com/docs/guides): auth, webhooks and data models
## Optional
- [Changelog](https://acme.com/changelog): what changed in each release
```
Links must be absolute URLs, and every entry benefits from a short description — the same rules as the full [llms.txt format guide](/blog/llms-txt-format/). A second file named `public/llms-full.txt` gives agents the full-text sibling; see [llms.txt vs llms-full.txt](/blog/llms-txt-vs-llms-full-txt/) for when that pays off. The cost of this route is maintenance: every new page is a hand edit, and a file that drifts from your content is one agents follow into stale territory.
## Real-World Case: nuxt.com/llms.txt
The Nuxt project dogfoods the format. On 2026-09-09 the official docs site served a 58,277-byte file at `/llms.txt`:
```
$ curl -s -o /dev/null -w "HTTP %{http_code}, %{size_download} bytes\n" https://nuxt.com/llms.txt
HTTP 200, 58277 bytes
```
The file itself is a good model for a framework site. It opens with the `# Nuxt Docs` H1 and a blockquote that states exactly when the docs should be used. Then it links `nuxt.com/llms-full.txt` (the whole documentation in one file), advertises per-page Markdown — append `.md` to any URL or send `Accept: text/markdown` — and points agents at a public MCP server under `/mcp`. Clean Markdown and `llms-full.txt` matter enough to both [the v2 spec](/blog/llms-txt-v2/) and Nuxt's own file.
## Method 2: Generate It on the Fly with server/routes/
For content that changes faster than your deploy cadence, serve the file from a Nitro route. Nuxt maps `server/routes/llms.txt.get.ts` to `/llms.txt` — server routes outside `server/api/` get no `/api` prefix — and auto-imports `defineEventHandler`. Query Nuxt Content or your CMS and return plain text:
```
// server/routes/llms.txt.get.ts
const header = '# Acme Docs\n\nDocumentation for the Acme API and SDKs.\n\n## Documentation\n\n'
export default defineEventHandler(async () => header + (await getPages()).join('\n'))
```
Return a string and Nitro serves it as text; return a `Response` to set the `Content-Type` header explicitly. Every request rebuilds the file, so a new CMS entry is reflected instantly with no redeploy — at the cost of a query per request, which is why build-time generation is the better default for most sites.
## Method 3: nuxt-llms — Build-Time llms.txt and llms-full.txt
[nuxt-llms](https://github.com/nuxtlabs/nuxt-llms), maintained by NuxtLabs and published on npm since February 2025 (v0.2.0, January 2026), generates and prerenders `/llms.txt` at build time from a small `nuxt.config.ts` block:
```
# terminal
npm i nuxt-llms
// nuxt.config.ts
export default defineNuxtConfig({
modules: ['nuxt-llms'],
llms: {
domain: 'https://acme.com',
title: 'Acme Docs',
description: 'API reference and guides for the Acme platform.',
sections: [{ title: 'Guides', links: [{ title: 'Quickstart', href: '/docs/quickstart' }] }],
},
})
```
What you get for that config:
- **llms.txt prerendered at build** — set `prerender: false` to serve it dynamically instead (for example behind ISR).
- **llms-full.txt on demand** — add a `full.title` and `full.description` and the module emits the second file too.
- **Runtime hooks** — `llms:generate` lets modules and server plugins push extra sections, so CMS content can extend the file.
- **Nuxt Content integration** — with `@nuxt/content` 3.2.0 or newer, every file in your `content/` directory is added to `llms.txt` and `llms-full.txt` automatically.
## Which One Should You Pick?
Approach
Setup
Auto-updates
llms-full.txt
Best for
Static public/ file
1 minute
No
Second file
Small, stable sites
server/routes handler
~10 minutes
Every request
Your code
Content that changes between deploys
nuxt-llms module
~5 minutes
Every build
Yes, via full option
Docs and Nuxt Content sites
For a documentation site on Nuxt Content, the module is the clear winner: the file cannot drift, because the build writes it from the same content your pages render. Static suits a brochure site that rarely changes; a server route only earns its per-request cost when freshness beats the deploy interval.
## Verify and Maintain
- **Confirm the file is at the root** — `public/llms.txt`, never `public/llms/llms.txt`.
- **Use absolute URLs** and one descriptive line per link — see the [llms.txt best practices guide](/blog/llms-txt-best-practices/).
- **Check the deployed file** after every deploy:
```
curl -s -o /dev/null -w "HTTP %{http_code}, %{content_type}\n" https://yoursite.com/llms.txt
# Expect: HTTP 200, text/plain
```
Paste the URL into the free [llms.txt checker](/checker), or feed your sitemap into the [llms.txt generator](/generator) for a first draft, then hand the maintenance to nuxt-llms — a zero-maintenance byproduct of the build that already publishes your Nuxt site.
---
# How to Add llms.txt to Shopify: Replace the Default File (2026)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-shopify/
Shopify tutorial · E-commerce · 6 minutes
# How to Add llms.txt to Shopify: Replace the Default File
Good news: Shopify now publishes an `llms.txt` for every store automatically. Bad news: the default is Shopify's pitch, not yours. This guide shows how to swap in a curated file with a single theme template — no app, no Cloudflare Worker — and verify it is live.
## Shopify Now Serves llms.txt Natively
For most of 2025, adding `llms.txt` to a Shopify store meant a workaround: intercept `/llms.txt` in a Cloudflare Worker fronting the store, or upload the file to Shopify Files and create a URL redirect from `yourstore.com/llms.txt` to the CDN link. Both approaches are dead now. Shopify serves the route at the platform level, and every store gets a generated file — open `yourstore.com/llms.txt` and you will get a real response, no setup required.
That sounds like the problem is solved. Then you read the file and find the catch: it was generated by Shopify, _for_ Shopify — and it shows.
## Why the Default Shopify llms.txt Undersells Your Store
- **It leads with Shopify, not you.** The default recommends the Shop app, links Shopify's developer tools, and points to pages for starting a new store — in some cases before a single product appears. An AI engine that reads the file first learns about the platform, not about your brand.
- **It skips what matters.** Your top collections, best sellers, FAQ and about page are not prioritized — and may be missing entirely.
- **It has no descriptions.** Bare URLs force an AI to fetch each page and guess what it covers, which is exactly what `llms.txt` exists to prevent.
- **It wastes space on developer links.** Lines aimed at people building on Shopify carry zero value for a shopper asking ChatGPT or Perplexity which store sells what they want.
Aspect
Shopify default
Your custom file
Leads with
Shopify ecosystem
Your brand and catalog
Catalog links
Generic, unprioritized
Top collections and best sellers
Link descriptions
Rare
One line per link
Update control
Platform decides
You decide
The fix is a curated file — the same principle as the [e-commerce examples in our llms.txt sample library](/blog/llms-txt-examples/), but written around your store.
## Step 1: Write Content an AI Can Act On
Start from your sitemap: Shopify generates one at `yourstore.com/sitemap.xml`. Paste it into the free [llms.txt generator](/generator) to draft the file, then curate — delete out-of-stock lines, thank-you pages and thin pages until only the links you want AI engines to reach remain. A proven shape for a store:
```
# Meadow & Stone Tea Co.
> Meadow & Stone sells loose-leaf tea and brewing kits, ships worldwide from Portland, OR.
## Collections
- [Best Sellers](https://meadow.example.com/collections/best-sellers): the 20 teas customers reorder most
- [Gift Sets](https://meadow.example.com/collections/gift-sets): curated tea and kit bundles under $60
## Products
- [Jasmine Silver Needle](https://meadow.example.com/products/jasmine-silver-needle): single-origin jasmine green tea, 50 g pouch
- [Cold Brew Kit](https://meadow.example.com/products/cold-brew-kit): carafe, filter and two sampler tins
## Useful Pages
- [About](https://meadow.example.com/pages/about): sourcing, story and the team
- [Shipping & Returns](https://meadow.example.com/pages/shipping): delivery times and the return policy
```
Curated beats exhaustive: a dozen strong links with clear descriptions outperform a dump of hundreds of SKUs, and AI engines reward files that follow the [llms.txt best practices](/blog/llms-txt-best-practices/) — short summaries per link, grouped sections, no dead URLs.
## Step 2: Add a templates/llms.txt.liquid Theme File
Shopify treats `llms.txt` the way it treats `robots.txt`: a special template overrides the generated file. Confirmed on the Shopify developer community and in 2026 guides, the method is:
1. In your Shopify admin go to **Online Store → Themes**.
2. Open the three-dot menu next to your live theme and choose **Edit code**.
3. Inside the `templates` folder, click **Add a new template**.
4. Pick a plain **template** (not JSON), and name it exactly `llms.txt.liquid` — the `.liquid` extension is mandatory; a file named `llms.txt` will not be served.
5. Paste your curated Markdown content into the template and click **Save**.
Your content is now served at `yourstore.com/llms.txt`, replacing the Shopify default. Theme code edits do not affect your store front end: the template only answers that one route.
## Step 3: Verify, Then Keep It Fresh
Confirm the file is live and served correctly:
```
curl -s -o /dev/null -w "%{http_code}" https://your-store.com/llms.txt
# Expect: 200 — then view the URL and check it is your content
```
Paste the URL into the free [llms.txt checker](/checker) to validate structure and links against the spec. Then set a maintenance cadence: refresh the file after major catalog or pricing changes, seasonal collection swaps, and product retirements. Each subdomain or regional storefront needs its own file at its own root, since URLs, currency and inventory differ between them.
Once the default is replaced, Shopify stops speaking for you and your catalog speaks first. Draft the file from your sitemap with the [free llms.txt generator](/generator), publish it through the template, and validate with the [checker](/checker) — a five-minute job that puts your products ahead of the platform in every AI answer.
---
# llms.txt v2: What Changed in the 2026 Spec Update
> Source: https://llmstxtgenerator.dev/blog/llms-txt-v2/
Spec update · August 2026 · AI readiness
# llms.txt v2: What Changed in the 2026 Spec Update
Two years after the original proposal, the llmstxt.org specification has its first major revision. v2 (August 2026) adds discoverability, defines subpath files, and changes how agents are expected to use the file. Here is exactly what changed — and what you should update on your site.
## Why a v2 at All?
The original llms.txt proposal was published by Jeremy Howard on September 3, 2024, when the idea that language models would routinely read websites was still speculative. By 2026 that is routine: thousands of sites publish an llms.txt file, documentation platforms such as Mintlify generate one automatically, Chrome's Lighthouse audits sites for one as part of its agentic browsing checks, and OpenAI, Anthropic, and Gemini all publish llms.txt files for their own developer docs. The revision (updated August 10, 2026) reflects what two years of adoption taught the community — and the most common request was _discoverability_.
## Change 1: Two URL Forms for Markdown Versions
v1 specified one way to expose a machine-readable version of a page: append `.md` to the full URL, so `page.html` becomes `page.html.md`. In practice, many publishing tools replace the extension instead. v2 blesses both forms:
- Appended: `https://example.com/docs/page.html.md`
- Replaced: `https://example.com/docs/page.md`
For URLs without a filename, append `index.html.md` or `index.md`. Agents that follow a link from your llms.txt to a markdown page should now accept either form.
## Change 2: Subpath Files Are Officially Defined
v1 allowed llms.txt in subpaths without saying what that meant. v2 defines it: a file covers the pages under its path, and where more than one file applies, agents should use the most specific one. A root `/llms.txt` covers the whole site; `/docs/llms.txt` covers only the documentation.
This matters for sites that control just a directory — a GitHub Pages project site, for example, can publish `llms.txt` in its own path even though it can never touch the host root. The FastHTML project does exactly this, placing its file at `/docs/llms.txt` to cover only its documentation pages.
## Change 3: Standard Link Relations for Discoverability
Given a page, how does an agent find its markdown version — or the llms.txt file that covers it — without guessing? v2 answers with two standard link relations:
- `rel="alternate" type="text/markdown"` — points to the page's markdown version
- `rel="describedby"` — points to the llms.txt file that covers the page
You can provide them as HTML link elements in the page head:
```
```
Or as an HTTP `Link:` response header, which also works for non-HTML resources and can be added in web server or CDN configuration without touching any pages:
```
Link: ; rel="alternate"; type="text/markdown", ; rel="describedby"
```
## Change 4: Agents View, Then Follow Links
v1 described a tool, `llms_txt2ctx`, that expanded a file into an LLM context, and the `Optional` section carried special meaning for it. v2 drops the context-expansion tooling from the proposal and states the expectation directly: agents view or search the llms.txt to find what they need, then follow the relevant links — which should point to LLM-friendly content such as the markdown versions above.
The `Optional` section is still allowed and remains a useful convention for secondary links an agent can skip, but it no longer has mechanical semantics. The practical takeaway: keep the file small enough to fit in context, and put the detail behind the links.
## What Did Not Change
The core format is untouched: an optional byte-order mark, the H1 with your site name (the only required element), a blockquote summary, optional context paragraphs, and H2 sections with markdown file lists. Root placement still works. Everything you learned from the [format guide](/blog/llms-txt-format/) and [best practices](/blog/llms-txt-best-practices/) remains valid — existing files don't break.
## Your v2 Update Checklist
- Publish markdown versions of your key pages (either `.md` or `.html.md` form).
- Point the links in your llms.txt at the markdown versions, not the HTML pages.
- Add `rel="alternate"` and `rel="describedby"` via HTML link elements or an HTTP `Link:` header — see the [Cloudflare Pages guide](/blog/llms-txt-cloudflare-pages/) for CDN-level headers.
- Use a subpath file such as `/docs/llms.txt` when you control only a directory.
- Keep secondary links in an `Optional` section; keep the file under a few dozen links.
- Re-validate after every change with the [free validator](/checker).
The v2 update is small but meaningful: it makes llms.txt easier for agents to find, easier to consume, and usable by sites that only control a path. Generate a spec-compliant file with the [generator](/generator), check it with the [validator](/checker), and you're ready for the agentic web.
---
# How to Add llms.txt to VitePress (Static or Auto-Generated, 2026)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-vitepress/
VitePress tutorial · Static or plugin · 7 minutes
# How to Add llms.txt to VitePress
VitePress is the Vue-powered static site generator behind the documentation of Vite, Vue.js, Vitest and many other flagship open-source projects. It has no native `llms.txt` feature — yet its own documentation ships one, and so do the docs of every major project in its ecosystem. Two routes get you there: a static file in `public/`, or build-time generation with `vitepress-plugin-llms`. Both are covered here, with real config and a curl check.
## VitePress Docs Sites That Already Publish llms.txt
Documentation is where llms.txt earns its keep — the llmstxt.org spec notes that coding agents follow the file to reach API references and tutorials — and the VitePress ecosystem was among the first to dogfood it. These are not hypothetical examples. Verified with curl on 2026-09-06:
Site
Built with VitePress for
/llms.txt on 2026-09-06
vuejs.org
Vue.js
HTTP 200
vite.dev
Vite
HTTP 200
vitest.dev
Vitest
HTTP 200
vitepress.dev
VitePress itself
HTTP 200
The shape of these files is worth studying. The one on vitepress.dev starts with a single H1, a one-line blockquote summary, and a Table of Contents:
```
# VitePress
> Vite & Vue Powered Static Site Generator
Markdown to beautiful docs in minutes
## Table of Contents
- [Getting Started](/guide/getting-started.md): Get up and running with VitePress.
Learn how to install, scaffold, and start developing your documentation site.
- [Asset Handling](/guide/asset-handling.md): Learn how to reference and handle
static assets such as images, media, and fonts in VitePress.
```
Notice the links: every page is listed as a `.md` URL. That is the markdown page form the [llms.txt v2 spec](/blog/llms-txt-v2/) defines — a VitePress page rendered as clean Markdown is exactly the format an LLM wants to ingest, and it is what `vitepress-plugin-llms` outputs.
## Method 1: Drop a Static File in the public Directory
VitePress gives you a shortcut that most SSGs do: everything in the `public` directory — `docs/public` by default, directly under your source directory — is copied verbatim to the root of the build output. The official docs name `robots.txt` and favicons as the typical use case; `llms.txt` works the same way. Create `docs/public/llms.txt`:
```
# Acme Docs
> Acme Docs is the official documentation for the Acme platform.
> It covers installation, configuration and the REST API.
## Getting Started
- [Quickstart](https://docs.example.com/getting-started/): first project in five minutes
- [Authentication](https://docs.example.com/authentication/): API keys and scopes
## API Reference
- [REST API](https://docs.example.com/api/): every endpoint, parameter and error
```
The next `npm run docs:build` serves it at `https://docs.example.com/llms.txt`. Two notes: if you set a custom `srcDir`, the public directory follows it, and if the site deploys under a subpath via the `base` option, the file still lands at the output root while your listed URLs stay absolute. This route is fine while your docs are small and stable — the cost is that every new page means a hand edit, and a file that drifts from your content is a file agents follow into stale territory.
## Method 2: Auto-Generate with vitepress-plugin-llms
The ecosystem workhorse is [okineadev/vitepress-plugin-llms](https://github.com/okineadev/vitepress-plugin-llms) — roughly 400 stars as of September 2026, and adopted across the VoidZero ecosystem with the blessing of Vue.js. The flagship sites above use it: the file served on vitepress.dev is its output. Install and register it:
```
# terminal
npm install vitepress-plugin-llms --save-dev
```
```
// .vitepress/config.ts
import { defineConfig } from 'vitepress'
import llmstxt from 'vitepress-plugin-llms'
export default defineConfig({
vite: {
plugins: [llmstxt()],
},
})
```
Zero configuration required. Every build now writes three things into `.vitepress/dist`: `llms.txt` (the curated index with section links), `llms-full.txt` (all documentation merged into one file — see how the two relate in the [llms.txt vs llms-full.txt guide](/blog/llms-txt-vs-llms-full-txt/)), and a clean Markdown version of every page, which llms.txt links to. The plugin authors call out one practice that visibly improves the output: add a `description` to each page's frontmatter and it becomes the link description in llms.txt — which is exactly what an agent reads when deciding which link to fetch.
Two finer points from the project's documentation. First, you can mark sections of a Markdown source file as `llm-only` or `llm-exclude` to control what reaches the LLM versions — useful for instructions aimed at agents, or human-only notes. Second, for repositories with documentation in multiple languages the author recommends enabling the plugin for the English docs only; that is all an LLM needs.
## Alternative: vitepress-plugin-llmstxt
If you need finer control than zero-config gives you, [vitepress-plugin-llmstxt](https://www.npmjs.com/package/vitepress-plugin-llmstxt) (npm 0.5.x, zero dependencies) also generates `llms.txt`, `llms-full.txt` and per-page Markdown files, then adds: glob-based `ignore` patterns to exclude routes, a `transform` callback that rewrites generated content, dynamic-route and i18n support, and experimental compatibility with VitePress 2.0 alpha. A configured example:
```
// .vitepress/config.ts
import { defineConfig } from 'vitepress'
import llmstxt from 'vitepress-plugin-llmstxt'
export default defineConfig({
vite: {
plugins: [
llmstxt({
hostname: 'https://docs.example.com',
ignore: ['**/api/**/*'],
}),
],
},
})
```
The `hostname` option pins the absolute URLs in your file (useful when VitePress infers the wrong origin), and `ignore` keeps internal or generated pages out of the index.
## Static File or Plugin?
Approach
Setup
Per-page .md
llms-full.txt
Stays current
Static `docs/public/llms.txt`
1 minute
No
No
Manual edits only
`vitepress-plugin-llms`
~5 minutes
Yes
Yes
Every build
`vitepress-plugin-llmstxt`
~5 minutes
Yes
Yes
Every build
A hand-written file is fine while your sitemap is small and your structure rarely moves. The moment docs change weekly — new guides, migrated pages, added API endpoints — generation wins, because the llms.txt an agent reads is only as trustworthy as its last update, and a plugin cannot forget to update itself.
## The VitePress Checklist
- **Confirm the public directory path** — it is `docs/public` by default but moves with `srcDir`; a static file in the wrong folder quietly never gets published.
- **Register the plugin under `vite.plugins`**, not at the top level of the VitePress config — these are Vite plugins and VitePress only loads them there.
- **Write a `description` in every page's frontmatter** — it becomes the link description agents use to choose pages. "What problem this page solves" beats a module name.
- **Keep it English-only for LLM purposes** — for multi-language docs, generate the file from the English tree only, per the plugin's own guidance.
- **Verify the deployed file, not just the build** — redirect rules, domain moves and docs restructures break llms.txt in production after the build goes green. Confirm it serves as plain text:
```
curl -s -o /dev/null -w "%{http_code} %{content_type}\n" https://docs.example.com/llms.txt
# Expect: 200 text/plain
```
Then paste the URL into the free [llms.txt checker](/checker) to validate the H1, blockquote summary, absolute URLs and section structure against the spec. To draft the initial file in seconds, feed your docs sitemap into the [free llms.txt generator](/generator), then let the plugin keep it current from there. VitePress made llms.txt a zero-maintenance default for the biggest projects in the Vue ecosystem — your docs can be next.
---
# llms.txt vs llms-full.txt: What's the Difference and Which Do You Need?
> Source: https://llmstxtgenerator.dev/blog/llms-txt-vs-llms-full-txt/
llms.txt spec · AI content strategy · Updated 2026
# llms.txt vs llms-full.txt: What's the Difference and Which Do You Need?
**llms.txt** is a curated index of your best pages. **llms-full.txt** is the entire site content in one Markdown file. Different jobs, different sizes — here's how to choose.
## The Short Answer
`llms.txt` gives AI agents a _map_ of your site: a small, curated Markdown list of the pages that matter most. `llms-full.txt` gives them the _territory_: the full text of your content, formatted as Markdown, in a single file. The spec at [llmstxt.org](https://llmstxt.org) treats llms.txt as the primary file — llms-full.txt is an optional companion for sites where the map alone isn't enough.
## What llms.txt Is (and Isn't)
An llms.txt file is a short Markdown document at your site root. It opens with an H1 (your site name) and a one-line `>` blockquote summary, followed by a curated list of links — usually your 5 to 15 most important pages, grouped under `##` headings.
```
# Example Site
> A concise description of what the site is and who it's for.
## Documentation
- [Getting Started](https://example.com/docs/getting-started)
- [API Reference](https://example.com/docs/api)
## Blog
- [How We Built X](https://example.com/blog/how-we-built-x)
```
The job of llms.txt is _prioritization_. An agent that lands on your site can read this file in milliseconds and know exactly which pages are worth fetching — instead of guessing from a crawl or landing on your marketing pages.
## What llms-full.txt Is
llms-full.txt is the same idea scaled up: your site's full content, converted to Markdown and merged into one file, served from `https://your-domain.com/llms-full.txt`. It's meant for situations where the curated index isn't enough material for the agent:
- **Documentation sites** — an agent needs the actual API docs, not just a link to them.
- **Long-form content** — tutorials and guides that span thousands of words.
- **JS-rendered pages** — content that doesn't exist in the raw HTML an agent fetches.
- **Research-heavy sites** — papers, data pages, reference material agents should read in full.
You link to it from your llms.txt with a relative Markdown link:
```
# Example Site
> A concise description of what the site is and who it's for.
## Documentation
- [Getting Started](https://example.com/docs/getting-started)
- [API Reference](https://example.com/docs/api)
[llms-full.txt](llms-full.txt)
```
## Key Differences at a Glance
llms.txt
llms-full.txt
**Content**
Curated link list
Full site content as Markdown
**Typical size**
1–5 KB
10 KB – several MB
**Purpose**
Prioritize what agents read
Give agents everything at once
**Required?**
Yes (the primary file)
Optional
**Best for**
Every site
Docs, long-form, JS-heavy sites
## Which One Do You Need?
Start with **llms.txt alone**. It's the file every AI engine checks first, it's trivial to maintain, and for most sites the curated list is exactly what an agent needs. Add **llms-full.txt** when you observe that agents need more material — typically once your site has substantial documentation or long-form content, or when you want AI engines to be able to answer from your content without fetching every page.
There's no penalty for having both, and no penalty for having just llms.txt. The only real mistake is treating llms-full.txt as a replacement for the curated index — an agent expects the index first, and llms-full.txt is reached _through_ it.
## Validate Both Files
After publishing either file, run it through our free [llms.txt validator](/checker) — it checks the H1, absolute URLs, Markdown formatting, and structure, and gives you an AI Readiness Score so you know the file will parse cleanly for agents. Then use the [generator](/generator) to produce a fresh llms.txt from your sitemap in seconds.
---
# llms.txt vs MCP: The Read Layer vs the Action Layer for AI Agents (2026)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-vs-mcp/
AI Agents & Discovery · September 2026
# llms.txt vs MCP: The Read Layer vs the Action Layer
Two protocols get mentioned in the same breath and solve opposite problems. `llms.txt` answers _"what is on this site?"_ — a static map an agent reads once. MCP answers _"what can I do here?"_ — a live interface an agent calls. Confusing them leads to the wrong investment: a file you never needed to make dynamic, or a server nobody can find.
## Two Questions, Two Protocols
Ask a chat assistant to "summarise how this company handles refunds" and it needs **content**: which page, in what form. Ask an agent to "find the cheapest plan for eight seats" and it needs **capabilities**: a query it can run against your data or your UI. `llms.txt` is the accepted answer to the first, MCP to the second, and a third, browser-level answer — WebMCP — is emerging underneath both.
## llms.txt Is the Read Layer
`/llms.txt` is a Markdown file at your site root, proposed by Jeremy Howard in September 2024 and maintained by the community as a convention rather than a formal standard — the current shape of it is documented at [llmstxt.org](https://llmstxt.org/). It carries an H1 name, a summary blockquote and curated sections of links with one-line descriptions. Nothing executes; nothing authenticates.
Its value is proportional to how bad your website is to parse. An agent landing on an HTML marketing page burns context on navigation, cookie banners and markup before it reaches the answer; a curated file replaces that with a handful of links. That is why documentation-heavy sites adopted first — and why independent crawl studies put overall adoption at roughly **10% of domains**, far below robots.txt or XML sitemaps.
Treat it as the cheap, static, cacheable layer. [The format guide](/blog/llms-txt-format/) covers every element and the [generator](/generator) turns a sitemap into a first draft. Google has said repeatedly that the file has no effect on Search rankings — see [what Google actually says](/blog/does-llms-txt-affect-seo/). That makes it an agent-facing asset, not a ranking asset.
## MCP Is the Action Layer
The Model Context Protocol, launched by Anthropic in late 2024 and now maintained as an open specification at [modelcontextprotocol.io](https://modelcontextprotocol.io/), defines how a client (an assistant or IDE) connects to a server that exposes tools, resources and prompts. Instead of reading your documentation, the agent calls it: a search tool, a pricing lookup, a status endpoint.
The **2026-07-28 revision** is the biggest change since remote MCP shipped. The `initialize` handshake and the `Mcp-Session-Id` header are retired: requests are stateless, each carrying its protocol version, client identity and capabilities in request metadata. A capability inventory is now an optional `server/discover` call, and the new `Mcp-Method` and `Mcp-Name` headers let a gateway route on the operation without parsing the body:
```
POST /mcp HTTP/1.1
MCP-Protocol-Version: 2026-07-28
Mcp-Method: tools/call
Mcp-Name: search
Authorization: Bearer
Content-Type: application/json
```
The practical consequence: with no session store to share, a server can run behind a plain round-robin load balancer — or, as Cloudflare describes in [its MCP v2 write-up](https://blog.cloudflare.com/mcp-v2/), entirely inside a single Worker. Remote servers are expected to use OAuth 2.0 with PKCE, and updated TypeScript, Python, Go and C# SDKs shipped with the spec.
MCP has one weakness llms.txt does not: **the protocol has no discovery**. Agents connect only to servers they were configured with, so yours can be excellent and unknown. This is where the read layer earns its keep: a curated file naming your MCP endpoint, what it exposes and who it is for is the most reliable "here is my server" signal you control.
## WebMCP: The Third, In-Browser Layer
WebMCP moves the same idea into the page. Rather than a server the agent connects to, the browser exposes tools the site declares: _declaratively_ by annotating existing HTML forms, or _imperatively_ with `navigator.modelContext.registerTool`. The agent then operates the interface the user is already looking at, inside the visitor's own session — no separate endpoint, no separate auth.
WebMCP is still experimental. Chrome's documentation for [Lighthouse's Agentic Browsing category](https://developer.chrome.com/docs/lighthouse/agentic-browsing/scoring) — which verifies tool registration, checks the accessibility tree, measures layout stability and inspects `llms.txt` discoverability — requires Chrome 150 or later, and its WebMCP audits require the origin trial. Not a ranking factor, but a clear signal of where browser vendors are heading.
## Side by Side
Layer
What it is
Who consumes it
Cost to ship
`llms.txt`
Static Markdown index at the site root
Reading crawlers and assistants
An afternoon, then occasional edits
MCP server
Live tools and resources over a stateless protocol
Configured clients: IDEs, assistants, agents
Engineering work, auth, ongoing maintenance
WebMCP
Browser-exposed tools from an existing page
Agents acting in a live session
Form annotations today, experimental API
## The Adoption Order That Makes Sense
Layer these in sequence, and only move down when the layer above is measurably working:
- **1\. Publish and validate the file.** Run it through the [llms.txt checker](/checker): every link must resolve, every entry must carry a description, and the header block must follow the spec.
- **2\. Measure whether anything reads it.** Start with access logs. If fetches are still zero after two months, the [crawler data](/blog/which-ai-engines-read-llms-txt/) and the [measurement framework](/blog/measure-llms-txt-traffic/) tell you whether that is normal for your niche.
- **3\. Make the file point at real capabilities.** Documentation subsections, a machine-readable pricing page, a status endpoint, your MCP server if you have one. This is where the two layers compound.
- **4\. Build a server only for repeated, high-value questions.** Support deflections, catalogue lookups, quote generation. If nobody asks twice a week, a smaller file beats a bigger server.
- **5\. Watch the browser layer.** If your value sits in a form, configurator or checkout flow, WebMCP is the place to prototype — but not on the roadmap until the audits leave origin-trial status.
The short version: `llms.txt` is how an agent learns what you have, MCP is how it acts on it, and neither replaces the other. Ship the file, instrument it, and let real request data — not vendor enthusiasm — decide whether a server comes next.
---
# llms.txt vs robots.txt vs sitemap.xml: What's the Difference?
> Source: https://llmstxtgenerator.dev/blog/llms-txt-vs-robots-txt/
Site root files · Search & AI discovery · Updated 2026
# llms.txt vs robots.txt vs sitemap.xml: What's the Difference?
Three tiny files, three very different jobs. **robots.txt** controls which crawlers may access your site, **sitemap.xml** tells search engines which URLs exist, and **llms.txt** tells AI engines what's worth reading.
Open the root directory of almost any serious website and you will find a small set of plain files: `robots.txt`, `sitemap.xml`, and — increasingly — `llms.txt`. They look similar — tiny, text-based files in the same folder, served to automated visitors — so it's tempting to assume they do the same job. They don't: each is read by a different consumer and answers a different question.
Below we compare all three and show how they work together.
## What Is robots.txt? — Crawler Permission Control
**robots.txt** is the oldest of the three, dating back to the Robots Exclusion Protocol of 1994. It is a plain-text file placed at the root of a domain (`https://example.com/robots.txt`) that tells compliant crawlers which parts of the site they may or may not crawl.
**Who reads it:** search-engine crawlers — Googlebot, Bingbot, YandexBot, DuckDuckBot — and increasingly AI crawlers such as GPTBot, ClaudeBot, and PerplexityBot. Any compliant crawler checks `robots.txt` before it fetches anything else on your domain.
**What it does:** `robots.txt` does not rank, index, or describe anything — it is a permission system. Rules like `Disallow: /admin/` stop compliant crawlers from requesting those paths, keeping private sections out of the crawl. It can also point crawlers at your sitemap via a `Sitemap:` directive.
```
User-agent: *
Disallow: /admin/
Disallow: /private/
User-agent: GPTBot
Disallow: /
Sitemap: https://example.com/sitemap.xml
```
One crucial caveat: `robots.txt` only governs well-behaved crawlers, and blocking a URL does _not_ remove it from Google — it can even leave a bare URL-only result with no snippet. For true removal, use a `noindex` meta tag or HTTP header instead.
## What Is sitemap.xml? — The Complete URL Inventory
**sitemap.xml** is an XML file that follows the Sitemap Protocol from sitemaps.org. It is an inventory of your site: a machine-readable list of every URL you want search engines to know about, optionally annotated with ``, ``, and `` hints.
**Who reads it:** search-engine indexers. Google, Bing, and others check the sitemap (typically referenced from `robots.txt` or submitted via Search Console) and use it to discover and prioritize new or rarely linked URLs.
**What it does:** a sitemap has no effect on rankings — it improves _discovery_. Pages that would otherwise wait weeks to be found through internal links get crawled quickly, and a fresh `` helps crawlers notice updates. Large sites split sitemaps into multiple files via a sitemap index.
```
https://example.com/2026-08-01https://example.com/docs/getting-started2026-08-10
```
Keep in mind a sitemap is a machine inventory, not a recommendation: it lists every URL with no notion of what is most important for a human or an AI to read first.
## What Is llms.txt? — A Content Index Built for AI Engines
**llms.txt** is the newest of the three. Proposed by AI researcher [Jeremy Howard](https://en.wikipedia.org/wiki/Jeremy_Howard_\(entrepreneur\)) in 2024 and specified at [llmstxt.org](https://llmstxt.org), it is a small Markdown file at the root (`https://example.com/llms.txt`) that gives AI engines a curated, human- and LLM-readable index of your content. The file opens with an H1 title and an optional blockquote summary, followed by sections of Markdown links, each with a short description. It is designed to fit in a model's context window, so an agent can understand your whole site from one small file.
**Who reads it:** AI engines and agents — ChatGPT, Claude, Gemini, Perplexity, coding assistants, and retrieval tools. Adoption is accelerating: Chrome Lighthouse audits for it, platforms like Mintlify and GitBook generate it automatically, and OpenAI, Anthropic, and Google publish llms.txt files for their own docs.
**What it does:** where `robots.txt` says what not to crawl and `sitemap.xml` lists every URL, `llms.txt` tells an AI engine what the site is about and which pages to read first. It is curation, not inventory: you order links by importance, add descriptions, and point agents at clean Markdown versions of your pages. There is no ranking or crawling behavior attached — it is simply a content index.
```
# Example Docs
> Example is a developer platform for building serverless applications in Python and TypeScript.
## Important links
- [Quick start](https://example.com/docs/quickstart.md): Set up your first app in 10 minutes
- [API reference](https://example.com/docs/api.md): Complete REST API reference
## Optional
- [Changelog](https://example.com/changelog.md): Release notes for all past versions
```
## llms.txt vs robots.txt vs sitemap.xml: Side-by-Side Comparison
Here is the full comparison. The short version: robots.txt manages crawler access, sitemap.xml feeds indexing, and llms.txt serves AI engines — three different consumers, three different jobs.
Dimension
`robots.txt`
`sitemap.xml`
`llms.txt`
**Intended reader**
Search-engine and AI crawlers (Googlebot, Bingbot, GPTBot)
Search-engine indexers and crawlers
AI engines and agents (ChatGPT, Claude, Gemini, Perplexity)
**File format**
Plain text (Robots Exclusion Protocol)
XML (Sitemap Protocol)
Markdown (llmstxt.org spec)
**Location**
Root only: `/robots.txt`
Any URL; typically referenced from robots.txt
Root by convention: `/llms.txt` (subpaths allowed)
**Primary purpose**
Control which crawlers may access which paths
List every URL you want discovered and indexed
Curate an index of what AI engines should read
**Impact on Google rankings**
None
None
None — built for AI engines, not ranking
**How it is discovered**
Checked automatically by compliant crawlers
Via `Sitemap:` directive or Search Console submission
Fetched by convention at `/llms.txt` or found via links
**Typical size**
A few lines to a few kilobytes
Up to 50,000 URLs per file (usually split)
Small enough to fit a context window
**Typical example**
`Disallow: /admin/`
`` with `` entries
`# Site name` + sections of Markdown links
**If it is wrong or missing**
Crawlers may access things you wanted private
New pages are discovered slowly
Agents misread your site or miss your best content
## Can You Use robots.txt, sitemap.xml, and llms.txt Together?
Absolutely. None of these files substitutes for another — they work best as a pipeline answering three different questions:
robots.txt
Read by crawlers (Googlebot, GPTBot)
Answers: “May I crawl this?”
sitemap.xml
Read by search-engine indexers
Answers: “Which URLs exist?”
llms.txt
Read by AI engines (ChatGPT, Claude, Perplexity)
Answers: “What should AI read?”
Crawler flow: robots.txt → sitemap.xml → pages. AI flow: llms.txt → linked Markdown pages. `robots.txt` points crawlers at the sitemap via a `Sitemap:` directive; `llms.txt` is served directly to AI engines.
A practical setup for a modern site:
- Point crawlers at your sitemap with a `Sitemap:` directive in `robots.txt`, so discovery is automatic.
- Allow GPTBot, ClaudeBot, and PerplexityBot instead of blanket-blocking them, then control what they consume with a well-curated `llms.txt`.
- Link to `llms.txt` from your homepage or footer, and keep it in sync with your sitemap — add new guides to both files.
- Use `robots.txt` to protect genuinely private paths, and `noindex` — not robots.txt — to remove indexed pages from Google.
## Common Misconceptions About These Three Files
- **“llms.txt will boost my Google rankings.”** — No. Google does not read `llms.txt` for ranking; the file targets AI engines such as ChatGPT, Claude, and Perplexity. For better Google rankings, invest in content quality, internal links, and Core Web Vitals — and treat `llms.txt` as a separate channel for AI visibility.
- **“robots.txt removes pages from Google.”** — No. `robots.txt` blocks crawling, not indexing. A blocked URL can still appear in results. Use `noindex` for removal and robots.txt for crawl control.
- **“llms.txt is just a sitemap in Markdown.”** — No. A sitemap is an exhaustive machine inventory; `llms.txt` is a curated, ordered recommendation with a summary and descriptions. Sitemaps are read by indexers; llms.txt by agents that reason about your content.
- **“A complete sitemap.xml is enough for AI engines.”** — Increasingly not. AI engines can crawl sitemaps, but increasingly look for `llms.txt` by convention because it tells them what matters in one small file — different consumers, complementary files.
## FAQ: llms.txt vs robots.txt vs sitemap.xml
### What's the difference between robots.txt and sitemap.xml?
`robots.txt` is a permission file for crawlers (“may I crawl, and where?”); `sitemap.xml` is a URL inventory for indexing (“these URLs exist”). They interact via the `Sitemap:` directive in robots.txt, but permission control and discovery are two different jobs.
### Does llms.txt replace my sitemap.xml?
No. `sitemap.xml` feeds search-engine indexing with every URL you own; `llms.txt` feeds AI engines with the curated subset that matters most — different consumers, formats, and goals. Most sites should publish both, plus `robots.txt`.
### Where should llms.txt live, and how do AI engines find it?
At the site root by convention — `https://your-domain.com/llms.txt`, next to `robots.txt`. Agents that support the convention check that URL automatically; linking from your homepage or footer helps discovery. Subpaths (e.g. `/docs/llms.txt`) work too, covering just that section.
## Keep reading
[
### 🧭 What is llms.txt?
The beginner-friendly overview: what it is and why it matters now.
](/guide/)[
### 📐 The Complete Format Guide
Every element of the v2 spec explained, with a valid end-to-end example.
](/blog/llms-txt-format/)
## Ready to add llms.txt to your site?
Generate a spec-compliant llms.txt in under a minute, then validate it for AI-readiness — no sign-up needed.
[⚡ Generate your llms.txt](/generator) [Check my file](/checker)
---
# llms.txt vs sitemap.xml: Why AI Engines Need a Different File
> Source: https://llmstxtgenerator.dev/blog/llms-txt-vs-sitemap-xml/
Site root files · AI discovery · Updated 2026
# llms.txt vs sitemap.xml: Why AI Engines Need a Different File
Your `sitemap.xml` is a complete inventory for search engines. Your `llms.txt` is a curated recommendation for AI agents. Here's why one can't do the other's job.
## What sitemap.xml Does
A sitemap is an XML file that lists _every URL_ you want search engines to index — product pages, articles, category pages, sometimes thousands of entries — with metadata like last-modified dates and change frequency. Googlebot and Bingbot read it to discover pages they might otherwise miss.
```
https://example.com/docs/getting-started2026-08-01
```
Its strengths: complete, machine-readable, and the industry standard for indexing. Its weakness for AI: it's a _flat list_. Nothing says which pages matter, what each one is about, or what order an agent should read them in.
## What llms.txt Does
An llms.txt file is a short Markdown document — typically 5 to 15 curated links — that tells an AI agent exactly what your site is and which pages are worth reading:
```
# Example Site
> A concise description of what the site is and who it's for.
## Documentation
- [Getting Started](https://example.com/docs/getting-started)
- [API Reference](https://example.com/docs/api)
## Blog
- [How We Built X](https://example.com/blog/how-we-built-x)
```
It's a _curation_ file. The agent gets your H1, a one-line summary, and a prioritized list in seconds — so it fetches your best pages instead of crawling blindly or landing on a tag archive.
## The Comparison
sitemap.xml
llms.txt
**Format**
XML
Markdown
**Audience**
Search crawlers (Googlebot, Bingbot)
AI engines (ChatGPT, Claude, Perplexity, Gemini)
**Scope**
Every URL
5–15 curated pages
**Priorities**
No
Yes (order + H2 groups)
**Human-readable?**
No
Yes
**Effect on Google rankings**
Yes (indexing)
None (Google doesn't read it)
## Why You Need Both
Different consumers, different jobs. Your sitemap keeps Google's index complete; your llms.txt makes AI answers about your site accurate and grounded in your best content. The two also support each other operationally: many teams [generate llms.txt from the sitemap](/blog/llms-txt-cloudflare-pages/) automatically, and our [generator](/generator) does exactly that — paste your sitemap URL and get a curated, spec-compliant llms.txt in seconds.
The realistic caveat: AI crawler adoption of llms.txt is still ramping up, and Google has explicitly said the file does not affect its rankings. Treat llms.txt as an [AI-visibility signal](/blog/does-llms-txt-affect-seo/), not a ranking factor — and keep your sitemap healthy for the search side of the equation.
## Quick Checklist
- Submit your sitemap in Google Search Console and Bing Webmaster Tools.
- Serve llms.txt at the site root, spec-compliant (H1 + blockquote + absolute URLs).
- Validate it with our [free checker](/checker) after every content change.
- Regenerate when you add or remove important pages — a stale index is worse than none.
---
# How to Add llms.txt to WordPress (2026 Guide)
> Source: https://llmstxtgenerator.dev/blog/llms-txt-wordpress/
WordPress tutorial · Updated 2026
# How to Add llms.txt to WordPress (2026 Guide)
WordPress powers more than 40% of the web — but adding **llms.txt** to WordPress is different from a static site. This guide covers three proven ways to add llms.txt to WordPress in 2026: uploading a generated file, using a WordPress llms.txt plugin such as Yoast SEO, or adding a few lines of code. No matter which method you pick, you'll be done in under ten minutes.
## Why llms.txt Is Different on WordPress
On a static site, adding llms.txt is trivial: drop a file into the build folder and redeploy. On WordPress it's trickier, for three reasons. First, WordPress sites are **dynamic** — posts, pages, categories, tags, archives and pagination are generated from the database, so a site can easily expose hundreds or thousands of URLs. Second, the **plugin ecosystem** is where WordPress SEO actually happens: the tools that already manage your `robots.txt` and sitemap are the natural home for llms.txt too. And third, WordPress adds infrastructure concerns that static sites don't have — caching plugins, multisite networks and subdirectory installs can all silently break an llms.txt file.
The good news: because llms.txt is just a plain Markdown file served at `https://your-domain.com/llms.txt`, every method below works with any WordPress host — shared hosting, managed WordPress, or a VPS.
## Method 1 (Recommended): Upload a Generated llms.txt to Your Site Root
The fastest, most reliable way to add llms.txt to WordPress requires no plugins and no code: generate the file, then upload it to the root of your WordPress installation. You keep full control over what goes into the file, and there is nothing to break later.
1. **Generate your file.** Use the free [llms.txt generator](/generator) — enter your URL, pick the pages you want AI engines to read, and download the resulting `llms.txt` file.
2. **Find your WordPress root.** The root is the folder that contains `wp-config.php` — usually `/public_html` (cPanel and most shared hosts), `/var/www/html`, or whichever folder your hosting file manager opens first. If WordPress is installed in a subdirectory (for example `example.com/blog/`), the root is _that_ subdirectory, not the document root above it.
3. **Upload the file there.** Three ways, pick one:
- **Hosting file manager:** open `public_html` → _Upload_ → select `llms.txt`.
- **cPanel File Manager:** navigate to `public_html` → _Upload_, then move the file if needed.
- **FTP/SFTP:** connect with FileZilla or any FTP client, open `/public_html/`, and drag the file in.
4. **Verify.** Open `https://your-domain.com/llms.txt` in a browser — you should see your file rendered as plain text, not a 404 or the homepage.
Two notes. Don't upload the file inside `wp-content/themes/…`: theme updates can delete it. And if your site serves through Cloudflare or another CDN, purge the cache after uploading so visitors and AI crawlers get the fresh copy.
## Method 2: Use a WordPress llms.txt Plugin
If you prefer automation — especially on a busy site where content changes daily — a WordPress llms.txt plugin is the way to go. Plugin-generated files stay in sync with your content, respect your `noindex` rules, and need zero maintenance:
- **Yoast SEO (free and premium).** Yoast shipped built-in llms.txt support in 2025 and it's fully part of the plugin in 2026. To enable it: go to **Yoast SEO → Settings → Site Features**, scroll to the **AI tools** section, switch the **LLMS.txt** toggle on, and save. Yoast automatically picks your most important and recently updated pages, respects `noindex` settings, and regenerates the file weekly — no code, no manual uploads.
- **Website LLMs.txt (WordPress.org).** A dedicated llms.txt plugin that generates the file from your published content, pulls titles and meta descriptions from Yoast, Rank Math, SEOPress or AIOSEO, excludes anything marked `noindex`, and serves the file from your site root. Good choice if you don't use Yoast.
- **Rank Math.** Also offers llms.txt support in its SEO suite, useful if Rank Math is already your SEO plugin of choice.
Method
Difficulty
Best for
Upload a generated file (Method 1)
Easy — no plugins
Full control, one-time setup, any host
WordPress llms.txt plugin (Method 2)
Easiest — automatic
Frequently updated sites, non-technical owners
Code in a child theme (Method 3)
Intermediate — PHP
Developers who want zero extra plugins
## Method 3: Add llms.txt to WordPress with Code
If you'd rather not install another plugin, a few lines of PHP serve your llms.txt file directly from your theme. Two rules first: **never edit WordPress core files** (`wp-includes`, `wp-admin`, or the bundled default themes), and **always use a child theme** so a parent theme update can't wipe your changes.
Step 1 — create the file `llms.txt` inside your child theme folder (for example `wp-content/themes/my-theme-child/llms.txt`) and paste in your generated content.
Step 2 — add this snippet to your child theme's `functions.php`:
`` Step 3 — flush the rewrite rules once: go to **Settings → Permalinks** and click _Save Changes_ (no other change needed). The snippet registers `/llms.txt` as a route, then serves your theme's file with the correct `text/plain` content type. ## What to Put in a WordPress llms.txt File Here is the temptation to resist: your WordPress site can produce hundreds of URLs — posts, pages, categories, tags, author archives, pagination, media attachments, WooCommerce products. **Do not list all of them.** An llms.txt file is a _curated index_: AI agents read it first and treat it as "these pages are worth reading." A giant auto-generated list dilutes your most important content and wastes the agent's context window. Include your homepage, about page, key landing pages, best articles, documentation and top products. Exclude tag and category archives, paginated listings, admin pages, and anything marked `noindex`. A healthy WordPress llms.txt is usually 10–30 links: ``` # Example Coffee Co. > A specialty coffee roaster in Portland, Oregon — beans, brewing guides, and our story. ## Important links - [Home](https://example.com/): Storefront, featured products, and seasonal offers - [Shop](https://example.com/shop/): All coffees and brewing equipment - [Brewing guides](https://example.com/guides/): Step-by-step recipes for every method - [About us](https://example.com/about/): Our story and sourcing philosophy ``` ## Common llms.txt WordPress Pitfalls - **Caching plugins serve stale files.** WP Rocket, W3 Total Cache, LiteSpeed Cache and Cloudflare can all cache `/llms.txt` and keep serving an old version after you update it — which is exactly how an AI engine ends up citing a deleted page. After every change, purge the page cache and the CDN cache, and consider excluding `/llms.txt` from caching entirely. - **WordPress in a subdirectory.** If your site lives at `example.com/blog/`, the file goes in the `blog/` folder, not `public_html` — otherwise agents fetch `example.com/llms.txt` and get a 404. - **Multisite networks.** Each site in a WordPress multisite should have its own llms.txt at its own root: a subdomain site's file goes in that subdomain's root, a subdirectory site's file in that subdirectory. - **Redirects and plugins hijacking the path.** Some hosts or security plugins redirect unknown paths to the homepage. After deploying, confirm `/llms.txt` returns a `200` with a `text/plain` content type — not a redirect to `/`. ## WordPress llms.txt FAQ ### Does Yoast SEO support llms.txt in 2026? Yes. llms.txt generation is built into both the free and premium versions of Yoast SEO. Enable it under Yoast SEO → Settings → Site Features → AI tools → LLMS.txt. Yoast picks your most important and recently updated pages, respects noindex rules, and regenerates the file weekly. ### Will a caching plugin break my llms.txt file? It can — a page cache may keep serving an old copy after you update the file. Purge your page cache and CDN cache after every change, and if you can, exclude /llms.txt from caching so crawlers always get the current version. ### Do I need llms.txt if I already have a sitemap? Keep both — they serve different purposes. An XML sitemap lists every URL for search-engine crawlers; llms.txt is a short, curated, Markdown index that AI agents read directly. A WordPress site should ship both: Yoast handles the sitemap, and any of the three methods above handles llms.txt. ## Keep reading [ ### 🛠️ How to Create llms.txt A step-by-step tutorial: build your file in 10 minutes and validate it. ](/blog/how-to-create-llms-txt/)[ ### 📚 Real-World Examples Copy-paste llms.txt files for SaaS, e-commerce, blogs and corporate sites. ](/blog/llms-txt-examples/) ## Ready to add llms.txt to WordPress? Generate your file in under a minute, then validate it against the spec — free, no sign-up, works with any of the three methods above. [⚡ Generate your llms.txt](/generator) [Check my file](/checker) ``
---
# Markdown for Agents vs llms.txt: Content Negotiation or Curated Index? (2026)
> Source: https://llmstxtgenerator.dev/blog/markdown-for-agents-vs-llms-txt/
Comparison · Content negotiation · Updated 2026
# Markdown for Agents vs llms.txt: Content Negotiation or Curated Index?
Two 2026 conventions are fighting for the same phrase: "serve Markdown to AI." One is a file that describes your site. The other is a request header that changes how a single page arrives. They are not competitors, and treating them as such is why a lot of agent-readiness work misses half the surface.
## The Short Answer
llms.txt
Markdown for Agents
Question it answers
What exists on this site, and what is worth an agent's context window?
Can I have this one page as clean Markdown, please?
Unit of work
The whole site, or every page under a path such as `/docs/llms.txt`
A single URL per request
Mechanism
A static Markdown file you curate
HTTP content negotiation on the `Accept` header
Discovery
The `/llms.txt` convention and the `rel="describedby"` link relation
None — the client has to know to ask
Where you implement it
Your repo, CMS or static assets
Your CDN, edge worker, or framework routes
The one-line version: llms.txt handles **discovery**, Markdown for Agents handles **delivery**. An agent that never finds your page does not benefit from your perfect Markdown endpoint, and an agent that finds your page but receives 180,000 tokens of div soup pays for the privilege.
## llms.txt Is a Map, Not a Pipe
The proposal by Jeremy Howard was updated to v2 in August 2026 and still describes a plain Markdown file: one H1, a blockquote summary, then H2 sections of linked resources with one-line descriptions. The v2 text formalises what practitioners had already started doing — the file can sit at the site root or at any path, covering the pages beneath it. That is why Cloudflare's developer documentation ships a root `developers.cloudflare.com/llms.txt` that is mostly a list of per-product files (Workers, R2, D1, AI Gateway and so on), each with its own full index:
```
> Each product below links to its own llms.txt, which contains a full
> index of that product's documentation pages and is the recommended
> way to explore a specific product's content.
## Developer platform
- [Workers](https://developers.cloudflare.com/workers/llms.txt): Build, deploy, and scale serverless applications
- [D1](https://developers.cloudflare.com/d1/llms.txt): Create managed, serverless databases with SQL semantics
```
What the file cannot do is hand over content. It points. Whether the agent then fetches HTML or Markdown is a separate decision — which is exactly the gap content negotiation fills. For the spec details, see the [llms.txt v2 breakdown](/blog/llms-txt-v2/).
## Markdown for Agents Is a Delivery Format
Cloudflare shipped "Markdown for Agents" on 12 February 2026: for zones with the feature enabled, a request carrying `Accept: text/markdown` makes the network convert the origin's HTML to Markdown in real time and return it from the same URL. Vercel published a similar pattern on 3 February 2026, implemented as a rewrite in `next.config.ts` that routes Markdown-preferring requests to a route handler, plus Markdown sitemaps at URLs such as `vercel.com/docs/sitemap.md`.
```
curl https://developers.cloudflare.com/workers/ \
-H "Accept: text/markdown"
HTTP/2 200
content-type: text/markdown; charset=utf-8
vary: accept
x-markdown-tokens: 725
content-signal: ai-train=yes, search=yes, ai-input=yes
```
Three headers on that response are worth more than the Markdown itself. `Vary: accept` keeps caches from serving the Markdown version to browsers. The `content-type` tells the client it got what it asked for, so it skips its own conversion step. And `x-markdown-tokens` is a hint about context cost that an agent can use for chunking decisions.
## Which Agents Actually Ask for Markdown
Content negotiation only pays off when the client sends the header. Checkly tested seven widely used coding agents in February 2026 by pointing their built-in web-fetch tools at a header echo endpoint. The result was a near-even split:
Agent
Sends `Accept: text/markdown`?
Claude Code (Fetch, v2.1.38)
Yes — `text/markdown, text/html, */*`
Cursor (WebFetch, 2.4.28)
Yes — `text/markdown` with q-values
OpenCode (webfetch, 1.2.5)
Yes — `text/markdown;q=1.0`
OpenAI Codex (Feb 2026 build)
No — browser-style HTML header
Gemini CLI (0.28.2)
No — `*/*`
GitHub Copilot (fetch\_webpage)
No
Windsurf (read\_url\_content)
No
Treat the table as a snapshot, not a law: tool versions move monthly, and crawlers such as GPTBot or OAI-SearchBot are a different population from IDE agents. The underlying mechanics are also unchanged — a server that ignores the header loses nothing, because the client will fall back to HTML and convert it itself.
## Where the Two Meet: Markdown Twins
The interesting design work happens when the map points at the pipe. llms.txt v2 recommends publishing a Markdown twin for each listed page — either by appending the extension (`page.html.md`) or replacing it (`page.md`) — and it defines two link relations to make the pairing machine-readable: `rel="alternate" type="text/markdown"` for the Markdown version of a page, and `rel="describedby"` for the llms.txt file that covers it.
Checkly's live file, fetched in September 2026, is a clean example of both layers in one place: it documents `Accept: text/markdown` support for every page URL, lists a set of hand-authored twins (`/pricing.md`, `/cli.md`), and ships scoped indexes at `/docs/llms.txt` and `/product/llms.txt` — so an agent can arrive through the index and then read each page in its cheapest form. Vercel's post describes the fallback for agents that never send the header: a `link rel="alternate" type="text/markdown"` tag in the HTML head.
## What Neither One Buys You
- **Rankings.** Google has consistently said llms.txt is not a Search signal, and changing the representation served to a requesting agent is not a ranking tactic either. Both belong to agent readiness, not to ranking.
- **Citation share.** Getting cited in an AI answer still depends on the content, not the container. Markdown reduces cost and latency; it does not make a thin page authoritative.
- **Coverage.** Markdown for Agents helps the agents that already browse your site. It does nothing for the ones that never discovered it — that is the index's job, and most published files still see very few fetches.
## How to Ship Both in an Afternoon
1. **Publish the index first.** Draft `/llms.txt` from your sitemap with the [generator](/generator), keep it to the 10–50 pages that answer real questions, and validate it with the [checker](/checker) before you touch the CDN.
2. **Add the delivery layer.** On Cloudflare, enable Markdown for Agents in the zone's Quick Actions. Off Cloudflare, read the `Accept` header and return Markdown with `Content-Type: text/markdown`, or pre-generate twins at build time if your content is already authoring in Markdown.
3. **Connect them.** Point the links inside llms.txt at the Markdown twin where one exists, and add `rel="alternate" type="text/markdown"` to the pages you care about most.
4. **Verify both paths.** One request per page: a plain fetch that must still return HTML, and an `Accept: text/markdown` request that must return Markdown plus `Vary: accept`. If the second returns HTML, the conversion never triggered — that is the failure to watch for.
Related reading: the same "two layers" split in the action domain, covered in [llms.txt vs MCP](/blog/llms-txt-vs-mcp/), and the measured reality of who fetches these files in [which AI engines read llms.txt](/blog/which-ai-engines-read-llms-txt/).
---
# How to Measure Whether AI Engines Read Your llms.txt (Logs, GA4 and Cloudflare, 2026)
> Source: https://llmstxtgenerator.dev/blog/measure-llms-txt-traffic/
Measurement & analytics · September 2026
# How to Measure Whether AI Engines Read Your llms.txt
Publishing `/llms.txt` is the easy half. Proving anyone read it is the half that decides whether you keep maintaining the file. Here is a three-layer measurement stack — raw logs, crawler behaviour, and AI referral traffic — that separates what actually happened from what your analytics dashboard implies happened.
## Why llms.txt Measurement Is Different
Most web measurement assumes a browser: JavaScript loads, a session starts, a referrer gets recorded. AI crawlers do none of that. When `OAI-SearchBot` fetches your `/llms.txt`, no GA4 session exists, no cookie is set, and no channel group fires. The only evidence is a line in your access log.
That single fact reorganises the whole reporting problem into three independent questions:
- **Was the file read?** Direct fetches of `/llms.txt` — the only signal that proves the file works.
- **Is your content being read?** AI crawler hits on your pages, which happen with or without an llms.txt file.
- **Is it producing anything?** Humans arriving from AI assistants in GA4.
These signals move on different timelines and disagree for months — heavy crawler traffic with zero AI referrals is normal. Report them separately and you will never confuse a crawl with a citation.
## Layer 1 — Count Direct llms.txt Fetches
On your own server, the file is one path. Filter for it and group by user agent:
```
# Did any AI agent request the file?
grep -i "llms.txt" /var/log/nginx/access.log
# Group the hits by crawler identity
grep -i "llms.txt" access.log | awk '{print $1, $12}' | sort | uniq -c | sort -rn
# Include rotated logs from the last month
zgrep -i "llms.txt" /var/log/nginx/access.log.*.gz | grep -iE "GPTBot|OAI-SearchBot|ClaudeBot|PerplexityBot"
```
What you are looking for is a small number of non-zero hits, not volume. The baseline across the wider web is grim: our breakdown in [Which AI Engines Actually Read llms.txt?](/blog/which-ai-engines-read-llms-txt/) covers the crawler audits showing that most published files are never requested even once. Two hits a month is a working file; zero after eight weeks means the file is [being skipped for a reason](/blog/llms-txt-mistakes/) — usually placement, an unparsable format, or links the agent cannot resolve.
User-agent strings are trivially spoofed, so confirm identity before reporting: resolve the requesting IP to a hostname, resolve that hostname back to an IP, and compare against the operator's published ranges. If they do not match, it was not that bot.
## Layer 2 — Track AI Crawler Hits on Content Pages
llms.txt is a map. Whether agents follow it shows up as requests to the pages it lists. Pull the top paths per crawler token:
```
# Top pages fetched by AI agents this month
grep -iE "GPTBot|OAI-SearchBot|ClaudeBot|Claude-SearchBot|PerplexityBot" access.log \
| awk '{print $12}' | sort | uniq -c | sort -rn | head -20
```
This is the most useful diagnostic in the stack, because it tests curation directly. If the paths agents fetch match the pages you listed in `/llms.txt`, the file is working as a prioritisation layer. If they keep fetching pages you deliberately left out — category archives, paginated tag pages, thin landing pages — the file is not steering anything yet, and the fix is structural, not a matter of publishing it again.
## Layer 3 — Read AI Referral Traffic in GA4
On May 13, 2026, Google added an **AI Assistant** entry to GA4's Default Channel Group. Sessions whose referrer matches a recognised AI assistant are tagged automatically:
Field
Value assigned
Default channel group
`AI Assistant`
Session medium
`ai-assistant`
Campaign
`(ai-assistant)`
Find it in _Reports → Acquisition → Traffic acquisition_ with Session default channel group as the primary dimension. Two caveats matter before you put the number in a report. First, Google's live Default Channel Group documentation lists ChatGPT, Gemini, Deepseek, Copilot and Grok, and states the channel excludes AI Overviews and AI Mode — so if you want Claude, Perplexity or AI Mode measurement, you still need a custom channel group with a regex on session source above the generic Referral channel. Second, this channel counts humans who clicked through. A busy AI crawler fleet produces exactly zero of these sessions.
## Layer 4 — Use Platform-Side Tooling Where You Have It
If your site sits behind Cloudflare, AI Crawl Control exposes a crawler-level view you cannot get from GA4: which AI services are accessing the domain, request patterns over time, whether a crawler respects your robots.txt, and per-crawler enforcement rules. Its Metrics tab also reports a Content Format breakdown — what content types AI systems request versus what your origin serves — and the changelog documents an option that redirects verified AI training crawlers to a page's canonical URL while humans and search crawlers see the original.
Cloudflare's older "Block AI bots" toggle is documented as deprecating on September 15, 2026, so plan a migration if your enforcement depends on it. And these dashboards report what reached the edge — not what an assistant chose to cite. Useful governance, not a truth machine.
## A 30-Day Benchmark Table
Put the three signals in one table and review monthly. This is the whole reporting system:
Signal
Where to read it
What good looks like
llms.txt fetches
Access logs, filtered on `llms.txt`
Any non-zero count from a verified AI agent
Crawler hits on listed pages
Access logs, grouped by user agent and path
Top paths overlap the sections in your file
AI referral sessions
GA4 AI Assistant channel, plus a custom channel group
Any stable, non-seasonal trickle worth segmenting
Before you spend a quarter tracking it, spend five minutes validating the file itself. Run it through the [llms.txt checker](/checker) to confirm each link resolves, every entry carries a description, and the header block follows the spec. If you are rebuilding from scratch, the [generator](/generator) turns a sitemap into a first draft. A file that fails validation will never produce a log line worth reporting — and no dashboard will tell you why.
---
# robots.txt for AI Crawlers in 2026: Allow, Block or Feed GPTBot, ClaudeBot and PerplexityBot
> Source: https://llmstxtgenerator.dev/blog/robots-txt-ai-crawlers/
AI crawlers & robots.txt · September 2026
# robots.txt for AI Crawlers: Who to Allow and Who to Block
robots.txt was designed to referee search engines. In 2026 it is refereeing a dozen AI agents from six different companies — and getting the policy right depends on one thing most guides skip: which class of crawler you are talking to. Here is the token list, two pastable policies, and how your llms.txt file picks up where robots.txt stops.
## The Three Classes of AI Crawler
Every AI user agent worth configuring falls into one of three job descriptions. The class decides what blocking actually costs you:
Class
What it does
Cost of blocking
**Training**
Collects pages for model pre-training corpora
None for today's citations — it affects future model weights
**Search index**
Indexes pages so the assistant can retrieve and cite them
Your pages stop being discoverable in AI answers
**User-triggered**
Fetches one page live when a user pastes a URL or asks about it
Often unavoidable — these behave like a browser on a user's behalf
That split is the whole game. Publishing teams that block "AI" wholesale by blocking one or two famous tokens usually block the harmless training crawler while leaving the retrieval agents untouched — or the reverse, and quietly remove themselves from AI answers.
## The 2026 User-Agent Token Reference
These are the tokens site owners ask about most. "Honours robots.txt" reflects each operator's published documentation, not a guarantee — verify against your own logs.
Token
Operator
Class
`GPTBot`
OpenAI
Training
`OAI-SearchBot`
OpenAI
Search index
`ChatGPT-User`
OpenAI
User-triggered
`ClaudeBot`
Anthropic
Training
`Claude-SearchBot`
Anthropic
Search index
`Claude-User`
Anthropic
User-triggered
`PerplexityBot`
Perplexity
Search index
`Perplexity-User`
Perplexity
User-triggered
`Google-Extended`
Google
Training opt-out token
`CCBot`
Common Crawl
Training corpus
`meta-externalagent`
Meta
Training
Two traps hide in that table. First, `Google-Extended` is not a crawler at all — it is a robot token that opts you out of Gemini and Vertex training while leaving Googlebot, Search and AI Overviews completely unaffected. You cannot opt out of AI Overviews without blocking Google Search itself. Second, Anthropic formalised the split in February 2026, documenting ClaudeBot, Claude-SearchBot and Claude-User separately and warning that blocking Claude-SearchBot reduces your site's visibility in its search results. Meanwhile OpenAI's user-triggered `ChatGPT-User` is documented as fetching on a user's behalf, which makes it effectively unmanageable from robots.txt.
## Two robots.txt Policies You Can Paste Today
### Policy A — Get cited, do not get trained
The split most publishers want: retrieval allowed, training denied.
```
# Retrieval agents: allow
User-agent: OAI-SearchBot
Allow: /
User-agent: Claude-SearchBot
Allow: /
User-agent: PerplexityBot
Allow: /
# Training crawlers: deny
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
User-agent: meta-externalagent
Disallow: /
# Google/Gemini training opt-out token
User-agent: Google-Extended
Disallow: /
```
### Policy B — Maximum visibility
If your goal is AI citations above all — documentation, tools, service pages — allow the agents that can recommend you, and keep them out of private paths only:
```
User-agent: GPTBot
User-agent: OAI-SearchBot
User-agent: ClaudeBot
User-agent: Claude-SearchBot
User-agent: PerplexityBot
Allow: /
Disallow: /admin/
Disallow: /checkout/
```
Rules for a specific token always override the wildcard group for that bot, so add a bot's group deliberately rather than relying on `User-agent: *`. And always test the published file, not the draft — `curl https://yourdomain.com/robots.txt` is the only version agents see.
## Cloudflare Content Signals: Usage Terms Inside robots.txt
Since Cloudflare introduced its Content Signals Policy, robots.txt has grown a second layer that describes _how_ content may be used, not just whether it may be fetched. Three signals are defined — `search` (building a search index), `ai-input` (feeding content into models for live answers) and `ai-train` (training or fine-tuning) — and sites can declare preferences with comma-delimited yes/no values:
```
# Content Signals Policy
User-agent: *
Content-Signal: search=yes, ai-train=no
Allow: /
```
Cloudflare's documentation states that domains without a robots.txt of their own may be served a default signal set of search permitted and AI training declined, and that content signals are a stated reservation of rights rather than a technical enforcement layer — enforcement still happens at the edge, through bot management rules. Treat signals as documentation of intent that strengthens your position; treat the allow/disallow rules above as the part that actually changes behaviour.
## Where llms.txt Fits: Permission vs Curation
robots.txt answers "may you read this?" llms.txt answers "what is worth reading?" Once you have allowed the retrieval agents, an `/llms.txt` file gives them a curated map: your best pages, one-line descriptions, and Markdown links that skip the HTML scaffolding. For large sites, llms.txt v2 formalises subpath files such as `/docs/llms.txt` and an `llms-full.txt` dump for full-context reads — Cloudflare's own developer docs run exactly this hub-and-spoke pattern.
The practical sequence is: audit your robots.txt groups, allow the retrieval tokens you want, then publish the map. Generate a first draft from your sitemap.xml with the [llms.txt generator](/generator), validate it with the [checker](/checker), and compare the two files properly in [llms.txt vs robots.txt](/blog/llms-txt-vs-robots-txt/). Remember that permission plus a map still leaves you at square one if agents never fetch: our crawl-data breakdown in [Which AI Engines Actually Read llms.txt?](/blog/which-ai-engines-read-llms-txt/) shows how uneven that coverage still is in 2026.
## Verify the Bots With Your Own Logs
Every claim above is testable on your own server. The monthly log check:
```
# Which AI agents actually reached you?
grep -iE "GPTBot|OAI-SearchBot|ClaudeBot|Claude-SearchBot|PerplexityBot" access.log | awk {print $1, $12} | sort | uniq -c | sort -rn | head -20
```
Then confirm identity — user-agent strings are trivially spoofed. The standard check is a double reverse-DNS lookup: resolve the requesting IP to a hostname, resolve the hostname back to an IP, and compare against the operator's published ranges. If the values do not match, it was not that bot. Run this audit once a month, and you will know within a single cycle which of your robots.txt rules are reshaping real traffic — and whether the llms.txt file you published is being read at all.
---
# Which AI Engines Actually Read llms.txt? (2026 Crawler Data)
> Source: https://llmstxtgenerator.dev/blog/which-ai-engines-read-llms-txt/
AI visibility · Crawler research · August 2026
# Which AI Engines Actually Read llms.txt?
llms.txt is pitched as the file that helps ChatGPT, Claude, Perplexity and Gemini understand your site. Two years of independent crawler audits tell a more honest story — here is the evidence, straight from server logs, and what it means for your site.
## The Short Answer: Almost No One Fetches It
If you publish an llms.txt today, the realistic outcome is that nothing automated will ever request it. In May 2026, Ahrefs analyzed 137,000 domains that publish the file: **97% of those files received zero requests** during the entire month — no crawler, no agent, no human. The 3% that were fetched were hit almost entirely (96%) by bots, and only about a fifth of those fetches came from named AI tools.
This is not one outlier study. Independent audits keep arriving at the same number: zero.
## What 500 Million Bot Requests Reveal
Two large log-based studies are worth knowing about because they measure what bots _actually do_, not what press releases say:
- **Weekerp (2025):** across two production sites, 68,759 AI bot requests in a month — and `/llms.txt` received **0 requests**. In the same period, `robots.txt` was fetched 1,938 times and sitemaps 1,224 times, proving AI bots do fetch supporting files when they care about them.
- **Limy (2026):** analysis of 515 million LLM bot traffic events found that GPTBot, ClaudeBot, PerplexityBot, OAI-SearchBot and Google-Extended "overwhelmingly skip the file and crawl HTML directly."
Crawler
Owner
Fetches llms.txt?
Evidence
`GPTBot`
OpenAI
Rarely — the most common named fetcher
Ahrefs 2026: top named AI fetcher
`ClaudeBot` / `Claude-Code`
Anthropic
Training crawler: not observed; Claude-Code: second most common
Ahrefs 2026; Weekerp: 0 requests
`PerplexityBot`
Perplexity
Not observed
Weekerp, Limy audits
`OAI-SearchBot`
OpenAI
Not observed
Weekerp, Limy audits
`Google-Extended` / Googlebot
Google
No — officially unsupported
Illyes, July 2025
## Google's Position Is On the Record
Google is the only major player that has answered the question publicly and unambiguously. In July 2025, Search Relations' Gary Illyes said Google has **"no plans to support llms.txt"**, and John Mueller compared the file to the keywords meta tag — a self-declared signal that was abandoned because it was trivially gameable. Google's 2026 guidance for generative AI visibility lists llms.txt by name as a tactic that does not help.
Yes, some Google properties (like `ai.google.dev`) still serve the file — Mueller confirmed that is an internal content system artifact, not an endorsement. We covered the full nuance in [Does llms.txt Affect SEO?](/blog/does-llms-txt-affect-seo/)
## Adoption Is Growing Anyway
Despite the crawl data, adoption keeps climbing. Casey Burridge's crawl of the top 10,000 sites found **5.61% with a valid llms.txt in June 2026**, up from 1.04% in July 2025 — roughly 5.4x growth in a year, or about 39,000 sites across the top million. SE Ranking measured 10.13% across 300,000 domains, and Rankability found 8.7% of the top 1,000 (15.8% of the 549 it could actually reach). Documentation platforms like Mintlify generate the file automatically, which is why the [v2 spec](/blog/llms-txt-v2/) and so many dev-tool sites adopted it early.
So who does read your file, realistically? Three groups:
- **Occasional fetchers.** GPTBot and Claude-Code are the only named AI tools with meaningful request records (Ahrefs).
- **Humans and agents on demand.** Users and AI assistants that are asked to summarize your site will often read llms.txt directly — this is the file's most reliable audience.
- **Future adopters.** The ecosystem is young; the v2 spec, subpath files and markdown link relations were added specifically to make adoption easier.
## Why Bots Skip Your llms.txt
- **Training runs from datasets, not live fetches.** Most model training uses pre-built corpora like Common Crawl, so there is no crawl step where llms.txt would matter.
- **It costs crawl budget.** Probing /llms.txt on every domain adds a request per host for little confirmed benefit, so most crawlers don't bother.
- **It is still unofficial.** No LLM lab has committed to honoring the file, which makes it a speculative investment for crawler teams.
- **robots.txt already works.** Access control is handled by a 30-year-old standard every crawler respects — llms.txt controls nothing and blocks nothing.
## What This Means for Your Site
The honest playbook in 2026: publish a spec-compliant llms.txt because it is cheap and because the file's on-demand audience is real — then spend your remaining SEO time on things with proven returns, like crawlable HTML content, a clean robots.txt, and site speed. There is no evidence that llms.txt moves rankings, traffic or citations today.
- Generate a valid file with the [llms.txt generator](/generator) and validate it with the [checker](/checker).
- Keep it small (5–15 links) and point links at markdown versions per the v2 spec.
- Check your own logs once a month to see if anything requests it:
```
grep -c "llms.txt" /var/log/nginx/access.log
grep -i "GPTBot" /var/log/nginx/access.log | head -20
```
If the second command shows hits, your file is being discovered — but don't expect it. Ship the file, keep it valid, and let the ecosystem catch up. You're ready for the agentic web either way.
---