Key takeaways

  • robots.txt controls which crawlers may fetch which parts of a site.
  • llms.txt is a proposed convention, not a standard, and AI services differ in whether they read it.
  • Neither file makes thin content citable.

robots.txt controls access

robots.txt tells crawlers which parts of a site they may fetch. AI companies often run separate crawlers for different purposes, for example one for training and another for search. You can allow one and block the other.

Check that you are not blocking search crawlers by accident

A rule written to keep training crawlers out can also shut out the crawler that feeds AI search. Read your file rule by rule and confirm which crawler each one applies to.

llms.txt is a proposal

llms.txt is a suggested convention: a plain file that summarizes a site for language models. It is not a standard, and AI services differ in whether they read it. Treat it as a low-cost addition, not a requirement.

robots.txtControls access. A crawler checks it before fetching pages, and each rule applies to the crawler it names.
llms.txtSummarizes a site for language models. It is a proposal, and a service may or may not read it.

What these files cannot do

Neither file makes thin content citable. If a page does not answer the question, access rules will not help.

A sensible order

First make sure search crawlers can reach the pages you want cited. Then improve the pages. Add llms.txt last.

Hyungmook Kimflowgen