Research any topic before you write.

Find related topics. | Discover entities. | See connections. | Build a topical map.

Robots.txt

The Robots Exclusion Protocol (often referred to by the filename used to implement it, robots.txt) is a standard used by websites to indicate to visiting web crawlers and other web robots which portions of the website they are allowed to visit.

Standards & History

Use the mouse wheel or two fingers (on touchscreens) to zoom in and out of the map.

Research this topic

Explore the main themes, entities and connections around Robots.txt. Start with the topic map, then use the sections below for research and deeper semantic analysis.

Explore this topic

Start with a few of the strongest sections from the source topic. These are research directions, not a list of keywords you must use.

Key facts & relationships

High-confidence facts extracted from structured source data. Use them as anchors for further research.

Authors
Martijn Koster (original author) · Gary Illyes, Henner Zeller, Lizzi Sassman (IETF contributors)
First published
1994 published, formally standardized in 2022
Status
Proposed Standard

Topics to explore

Browse the full topic structure. Each item opens a new analysis centered on that subject.

Overview

History

Standard

Compliance

Security

Alternatives

Nonstandard extensions

Meta tags and headers

Advanced semantic analysis

Deeper signals for content research, entity SEO and topical coverage. The plain-language headings explain what each technical view is useful for.

Map overview Semantic statistics

Robots.txt

Nodes84
Edges83
Triples103
Avg. degree1.98
Density0.02381
Components1

How this topic connects Entity context

See the strongest relationship patterns around the current topic before diving into the raw triples.

Robots.txt

Top relations

related to Artificial intelligence · 20
Robots.txt → AI, Also, Anthropic, BBC, Cloudflare, Denying, Google's Google-Extended, GPTBot, In, Internet, Many, Media, Medium, OpenAI's GPTBot, Originality, Perplexity, Starting, The New York Times, The Verge's David Pierce, To
related to history · 13
Robots.txt → AltaVista, By June, Charles Stross, February, Koster, Koster's, Lycos, Martijn Koster, Nexor, RobotsNotWanted, The, WebCrawler, WWW-related
see also · 11
Robots.txt → ArchiveMeta, Bidder's EdgehiQ Labs, Digital Library Program, Internet, LinkedInAutomated Content Access Protocol, National Digital Information Infrastructure, NDIIPP, NDLP, Now, Preservation Program, Simple LicensingSitemapsSpider
related to Security · 10
Robots.txt → Despite, In, Malicious, NIST, Standards, System, Technology, The National Institute, United States, While
related to Meta tags and headers · 8
Robots.txt → HTML, In, On, PDF, Robots, The, X-Robots-Tag, X-Robots-Tag HTTP
related to Archival sites · 7
Robots.txt → According, Archive Team, Co-founder Jason Scott, Digital Trends, In, Internet Archive, Some
related to Compliance · 4
Robots.txt → Bidder's Edge, March, May, The
related to Standard · 4
Robots.txt → Google, If, Robots, This
related to A "noindex" HTTP response header · 3
Robots.txt → The X-Robots-Tag, Thus, X-Robots-Tag
Authors · 2
Robots.txt → Gary Illyes, Henner Zeller, Lizzi Sassman (IETF contributors), Martijn Koster (original author)

Important terminology Word statistics

Use these terms to understand the vocabulary surrounding the topic, not as a checklist for keyword stuffing.

Important terminology

robots txt file standard web files website search pages crawlers bots use example protocol google access used websites server exclusion

Entity relationships Subject–Predicate–Object triples

SubjectPredicateObjectConfidenceSrc
Robots.txtAuthorsMartijn Koster (original author)1.00infobox
Robots.txtAuthorsGary Illyes, Henner Zeller, Lizzi Sassman (IETF contributors)1.00infobox
Robots.txtFirst published1994 published, formally standardized in 20221.00infobox
Robots.txtStatusProposed Standard1.00infobox
Robots.txtWebsiterobotstxt.org, RFC 93091.00infobox
WebCrawlerinstance ofincluding those operated by search engines0.80text
Lycosinstance ofincluding those operated by search engines0.80text
and AltaVista.On July 1instance ofincluding those operated by search engines0.80text
2019instance ofincluding those operated by search engines0.80text
Google announced the proposal of the Robots Exclusion Protocol as an official standard under Internet Engineering Task Forceinstance ofincluding those operated by search engines0.80text
Google.A robots.txt file on a website will function as a request that specified robots ignore specified files or directories when crawling a siteinstance ofRobots.txt files are particularly important for web crawlers from search engines0.80text
the BBCinstance ofDenying access to GPTBot was common among news websites0.80text

Related concept clusters Concept neighborhoods

These clusters group vocabulary that occurs around closely connected concepts in the source material.

    Connections between topic areas Semantic bridges

    Bridges can reveal useful research angles that are easy to miss in a flat list of related terms.

    Min side: 3
    For writers, content strategists, SEOs, marketers and creators — from quick topic research to advanced semantic analysis.