Research any topic before you write.

Find related topics. | Discover entities. | See connections. | Build a topical map.

Web crawler

A web crawler, sometimes called a spider or spiderbot and often shortened to crawler, is an Internet bot that systematically browses the World Wide Web and that is typically operated by search engines for the purpose of Web indexing (web spidering).

History & Applications

Use the mouse wheel or two fingers (on touchscreens) to zoom in and out of the map.

Research this topic

Explore the main themes, entities and connections around Web crawler. Start with the topic map, then use the sections below for research and deeper semantic analysis.

Explore this topic

Start with a few of the strongest sections from the source topic. These are research directions, not a list of keywords you must use.

Topics to explore

Browse the full topic structure. Each item opens a new analysis centered on that subject.

Overview

Nomenclature

  • FOAF FOAF (software)

Crawling policy

Architectures

Security

Crawler identification

Crawling the deep web

Historical web crawlers

In-house web crawlers

Commercial web crawlers

Open-source crawlers

Advanced semantic analysis

Deeper signals for content research, entity SEO and topical coverage. The plain-language headings explain what each technical view is useful for.

Map overview Semantic statistics

Web crawler

Nodes149
Edges148
Triples143
Avg. degree1.99
Density0.013423
Components1

How this topic connects Entity context

See the strongest relationship patterns around the current topic before diving into the raw triples.

Web crawler

Top relations

related to Open-source crawlers · 29
Web crawler → AGPL, Amazon CloudSearch, Apache Hadoop, Apache License, Apache Nutch, Apache Solr, Apache Storm, BSD, Dig, Elasticsearch, FTP, GNU Wget, GPL, Grub, Heritrix, HTTrack, Internet Archive's, It, Java, Microsoft Azure Cognitive Search
related to In-house web crawlers · 25
Web crawler → Apple's, Applebot, Baidu's, Baiduspider, Bingbot, DuckDuckBot, DuckDuckGo's, During, Googlebot, If, It, Mercator, Microsoft's Bing, Msnbot, Python, Siri, The, There, URL, URLs
related to Further reading · 18
Web crawler → Archived, Cho, Current Challenges, Denis, History, ICWE'13, Intelligent Web Crawling, July, Junghoo, OWASP, Search Engines, Shestakov, UCLA Computer Science Department, Wayback Machine, Web Crawling, Web Crawling Project, WI-IAT'13, WileyWIVET
related to Historical web crawlers · 14
Web crawler → Bingbot, California, Civil Engineering, Davis, Mani Singh, Microsoft, Slurp, The, University, Unix, URLs, WolfBot, World Wide Web Worm, Yahoo
related to Crawler identification · 9
Web crawler → Examining Web, HTTP, Identification, In, Spambots, The, URL, User-agent, Web
related to Commercial web crawlers · 7
Web crawler → APISortSite, Diffbot, Mac OSSwiftbot, Search, Swiftype's, The, Windows
related to overview · 6
Web crawler → As, If, The, Those, URLs, Web
related to Politeness policy · 5
Web crawler → As, Crawlers, If, Koster, The
related to Re-visit policy · 5
Web crawler → By, From, The, The Web, Web
is a · 3
Web crawler → highly extensible Web Crawler written in Java and released under an Apache License, outcome of a combination of policies, server and the Web sites are the queues

Important terminology Word statistics

Use these terms to understand the vocabulary surrounding the topic, not as a checklist for keyword stuffing.

Important terminology

web crawler pages crawlers crawling search url crawl also page engines urls may use server given engine resources used policy

Entity relationships Subject–Predicate–Object triples

SubjectPredicateObjectConfidenceSrc
Web crawleris aoutcome of a combination of policies0.90text
Web crawleris aserver and the Web sites are the queues0.90text
Web crawleris ahighly extensible Web Crawler written in Java and released under an Apache License0.90text
.htmlinstance ofa crawler may examine the URL and only request a resource if the URL ends with certain characters0.80text
.htminstance ofa crawler may examine the URL and only request a resource if the URL ends with certain characters0.80text
.aspinstance ofa crawler may examine the URL and only request a resource if the URL ends with certain characters0.80text
.aspxinstance ofa crawler may examine the URL and only request a resource if the URL ends with certain characters0.80text
.phpinstance ofa crawler may examine the URL and only request a resource if the URL ends with certain characters0.80text
.jspinstance ofa crawler may examine the URL and only request a resource if the URL ends with certain characters0.80text
.jspx or a slashinstance ofa crawler may examine the URL and only request a resource if the URL ends with certain characters0.80text
Apache Solrinstance ofIt can be used with many repositories0.80text
Elasticsearchinstance ofIt can be used with many repositories0.80text

Related concept clusters Concept neighborhoods

These clusters group vocabulary that occurs around closely connected concepts in the source material.

    Connections between topic areas Semantic bridges

    Bridges can reveal useful research angles that are easy to miss in a flat list of related terms.

    Min side: 3
    For writers, content strategists, SEOs, marketers and creators — from quick topic research to advanced semantic analysis.