llms.txt Generator

Enter your domain and this reads your sitemap, groups the pages into sections, and links the large ones to their index rather than listing every post. What comes back is a curated map of your site rather than a copy of your sitemap.

Enter the domain and the sitemap is found at its usual address, or paste the sitemap URL if yours lives elsewhere. A sitemap index is followed into its child sitemaps forty at a time, taking one of each kind first so that no part of the site is missed, and you may read further batches once the first is in.

Your file will appear here

Reading a sitemap takes a moment, and reading the pages that will appear in the file takes longer, since each one is fetched for its real title and description. When it lands you get a control for every section, so you decide which ones list their pages and which collapse to a single link, and a preview you may copy or download.

A sitemap is an inventory, an llms.txt is a map

The two files answer different questions. A sitemap tells a crawler every URL it may fetch, in whatever order the generator wrote them, with no notion of which pages matter. An llms.txt tells a reader which pages are worth reading. Converting one into the other is therefore a curation problem, not a formatting one, and that is where most generators stop short.

The distance between them is usually the long tail. A site with four hundred blog posts and a dozen product pages has a sitemap that is ninety-seven per cent blog, so a tool that takes pages in the order it finds them produces a file of arbitrary posts and quietly leaves out the pricing page. Linking to the blog once fixes that in a single line, and the space it frees goes to the pages a reader came for.

The same logic runs the other way. On a documentation site the docs are the substance rather than the tail, and collapsing them would empty the file. So the rule is relative: a section collapses when listing it would crowd out the rest, and a site that is one big section keeps its pages listed, because there is nothing for them to crowd out.

What to do with the file

Upload it to the root of your site so it is served at https://yourdomain.com/llms.txt, the same place robots.txt lives. It is a plain text file with no headers to configure and nothing to register.

Then edit it. The descriptions come from your meta descriptions, which were written to earn a click in a search result rather than to explain a page to a reader who has already arrived. Rewriting the dozen entries that matter most will make the file considerably more useful than the generated draft.

The file is a recommendation, not a rule. If you want to control which AI crawlers may read your site at all, that is robots.txt, and the robots.txt generator has presets for the AI user agents. The robots.txt tester will confirm the rules landed, and the sitemap validator will check the sitemap this was built from.

Frequently Asked Questions

What is llms.txt?

A Markdown file at the root of a site, at /llms.txt, that lists the pages worth reading and says in one line what each one covers. The specification asks for an H1 naming the site, an optional blockquote summary, and H2 sections holding lists of links in the form - [Name](url): notes. It was proposed in 2024 as a way to give AI assistants a curated map of a site rather than leaving them to infer one.

Why does it link to my blog instead of listing the posts?

Because a file of blog posts is not a map of your site. Most sitemaps are dominated by their long tail: this site's own sitemap holds 354 blog posts against seventeen root-level pages, so any tool that simply takes the first pages it finds produces a file of arbitrary articles and omits the pricing page. One link to /blog says the same thing in one line and leaves room for the pages a reader actually asks about. You may switch any section back to listing its pages.

How does it decide what to collapse?

Pages are grouped by the first segment of their path, and a section larger than about a dozen pages collapses to its index. Blogs, changelogs and news sections collapse sooner, because their value to a reader is the section rather than any one entry. Root-level pages are never collapsed, since those are the pages that say what the site is. Every decision is shown above the preview with a control to change it.

What if my site has no index page for a section?

The tool derives the likely URL, checks it, and only collapses when it answers. That check matters more than it sounds: plenty of sites list thousands of pages under /docs while never listing /docs itself in the sitemap. A section whose index does not answer falls back to listing its pages rather than pointing at a URL that does not exist.

Do AI assistants actually read it?

Support is not universal and no major AI company has committed to reading it as a standard. Treat it as a low-cost bet: the file takes minutes to produce, costs nothing to serve, and does no harm if it is ignored. What it does not do is control training or crawling, which is what robots.txt is for.

What is the Optional section for?

The specification gives that heading a defined meaning: a reader working with a shorter context may skip everything under it. Careers, legal and press pages are real pages that almost no one asks an assistant about, so they are filed there rather than left to compete with the pages that matter.

Where do the titles and descriptions come from?

Each page that will appear in the file is fetched for its real title and meta description. Only those pages are read, so switching a section from listing to a single link makes the build faster rather than slower. Edit the text before publishing: a meta description is written to earn a click in a search result, and a line here is written to tell a reader what the page contains.

How much of a large sitemap index is read?

Forty child sitemaps per pass, and the panel above the preview always says how many the index named against how many were read. Reading an index whole is often not possible: Search Engine Journal's index names 231 children and TechCrunch's names 2,058, and each one is a separate request against a shared per-minute limit. So the pass reads a batch and offers a button to read the next one.

Which forty matters more than the number. A sitemap index is written in runs of one kind, and Search Engine Journal names 189 post sitemaps before the three that hold its actual pages, so anything reading the list in order reads nothing but blog posts. One child of each kind is read before any kind is read a second time, which puts a page sitemap in the first handful whatever order the index was written in. Author, tag and category sitemaps are skipped outright, since none of those pages earn a line in the file.

Does the file follow the llms.txt specification?

Yes. The specification at llmstxt.org asks for an H1 naming the site, which is the only required part, then an optional blockquote summary, then H2 sections each holding a markdown list of - [Name](url): notes entries. That is what the preview writes, with the Optional heading last so that a reader working with a shorter context may stop there. Brackets in a page title and parentheses in a URL are escaped, since either would cut a markdown link short.

Get Google and ChatGPT traffic on autopilot.

Start today and generate your first article within 15 minutes.

Content Plan