Glossary

robots.txt

The plain-text file at a site's root that tells crawlers which paths they may fetch, including AI crawlers.

In one sentence

robots.txt is a plain-text file at the root of a site that tells crawlers which paths they may fetch. Compliant AI crawlers read it too.

What it means

Rules are grouped by user agent, so you can allow one bot and block another. This is how sites separate GPTBot from search agents, or set a rule for Google-Extended.

robots.txt is a request, not enforcement. Well-behaved crawlers follow it. Others may not. Separately, firewalls and bot-protection services can block crawlers that robots.txt permits, and that is a common reason a site is invisible to AI search despite a permissive file.

What to do about it

Review the file for broad rules such as a blanket Disallow for unknown bots. List the AI user agents you want to allow and test them from the outside with the AI crawler access checker. Pair it with an LLMs.txt if you want to point crawlers at key pages.

Keep the file short and commented. Each rule should have an owner and a reason, so that nobody deletes a needed rule or leaves an old block in place when the bot policy changes.

How Pineprompt measures it

The AI crawler access checker reads your robots.txt and tests reachability for the major AI user agents.

Frequently asked

What is robots.txt?
robots.txt is a plain-text file at the root of a site that tells crawlers which paths they may fetch. Compliant AI crawlers read it too.
How does Pineprompt measure robots.txt?
The AI crawler access checker reads your robots.txt and tests reachability for the major AI user agents.

Related terms

Test your robots.txt for AI bots.

Run the free AI crawler access checker.