Website logs include visits made by automated programmes as well as by people. These programmes are often called crawlers, robots or spiders. They request web pages to discover or collect information.
Useful and abusive crawling
The original article groups crawlers into “good” and “bad” categories. Search-engine crawlers are an example of useful crawling: they discover content that may later be indexed. Crawling alone does not guarantee indexing. See Google's crawling and indexing overview.
Automated collection can also be abusive, for example harvesting published email addresses for spam. Purpose and behaviour matter; the list below is not a declaration that every named bot is harmful or trustworthy.
Crawler profiles published on this site
The names and version strings below are retained from the original article and may be historical. Their links open the site's individual profiles.
- 2ip bot/1.1 (+http://2ip.io)
- AdsBot-Google (+http://www.google.com/adsbot.html)
- Buck/2.3.2; (+https://app.hypefactors.com/media-monitoring/about.html)
- CCBot/2.0 (https://commoncrawl.org/faq/)
- Sogou web spider/4.0(+http://www.sogou.com/docs/help/webmasters.htm#07)
- webprosbot/2.0 (+mailto:abuse-6337@webpros.com)
If you have more information about these crawlers, share it in the comments and distinguish what you observed from what you assume.
Comments (0)
Comments are shown in their original language.
No comments have been published yet. Be the first to join the conversation.