Overview
Scrapy is a web scraper used for public web data collection, page extraction, content monitoring, and third-party crawler activity.
Its primary user-agent pattern is Scrapy; related patterns include Scrapy/VERSION (+https://scrapy.org); a representative HTTP user-agent is Scrapy/VERSION (+https://scrapy.org).
Scrapy is verified with High confidence. The identity type is Official Documented, and the evidence basis is official operator documentation.
Robots.txt behavior is not currently confirmed.
Scrapy should be reviewed against site policy, source evidence, crawl rate, and requested paths before a permanent allow or block rule is created.
Identity
- User-Agent
Scrapy- Aliases
- Scrapy/VERSION (+https://scrapy.org)
- HTTP Agent Examples
Scrapy/VERSION (+https://scrapy.org)- Robots.txt Token
Scrapy- Identity Type
- Official Documented
- Evidence Method
- Verify Scrapy by matching `Scrapy` to official operator documentation, then checking reverse DNS, IP ownership, request behavior, and crawl consistency.
Classification
- Type
- Scraper
- Kind
- Crawler
- Family
- Scrapy
- Purpose
- web scraping
Behavior and handling
- Common Use
- Scrapy is used for public web data collection, page extraction, content monitoring, and third-party crawler activity.
- Detection Notes
- Scrapy traffic is primarily detected by the `Scrapy` user-agent pattern; related patterns include `Scrapy/VERSION (+https://scrapy.org)`; a representative HTTP user-agent is `Scrapy/VERSION (+https://scrapy.org)`. Compare source IPs, reverse DNS, request paths, and crawl cadence before trusting the traffic.
- Respects robots.txt
- Unknown
- Spoofing Risk
- Scrapy has medium spoofing risk because the user-agent can be copied, even when the bot has strong source or documentation support.
- Risk
- Neutral
- Recommended Handling
- Depends
Rules and controls
- Robots.txt Snippet
# robots.txt behavior is unconfirmed. Do not rely on this rule without verification.