Overview
Arquivo Web Crawler is a feed fetcher from Arquivo used for RSS or Atom feed polling, syndication refresh, subscription delivery, and content update checks.
Its primary user-agent pattern is Arquivo-web-crawler; a representative HTTP user-agent is Arquivo-web-crawler (compatible; heritrix/3.4.0-20200304 +https://arquivo.pt/faq-crawling).
Arquivo Web Crawler is verified with Medium confidence. The identity type is Verified Bot, and the evidence basis is a public crawler reference or source-linked documentation.
Arquivo Web Crawler is marked as respecting robots.txt directives for crawler access control.
Arquivo Web Crawler should be reviewed against site policy, source evidence, crawl rate, and requested paths before a permanent allow or block rule is created.
Identity
- User-Agent Pattern
-
Arquivo-web-crawler - HTTP Agent Examples
-
Arquivo-web-crawler (compatible; heritrix/3.4.0-20200304 +https://arquivo.pt/faq-crawling) - Robots Token
- Arquivo-web-crawler
- Identity Type
- Verified bot
- Evidence Method
- Verify Arquivo Web Crawler by matching `Arquivo-web-crawler` to Arquivo evidence, then checking reverse DNS, source-network ownership, signed request data, or published crawler documentation when available.
Classification
- Type
- Feed
- Kind
- Fetcher
- Family
- Arquivo
- Purpose
- Feed fetch
Behavior and handling
- Common Use
- Arquivo Web Crawler is used for RSS or Atom feed polling, syndication refresh, subscription delivery, and content update checks.
- Detection Notes
- Arquivo Web Crawler traffic is primarily detected by the `Arquivo-web-crawler` user-agent pattern; a representative HTTP user-agent is `Arquivo-web-crawler (compatible; heritrix/3.4.0-20200304 +https://arquivo.pt/faq-crawling)`. Compare source IPs, reverse DNS, request paths, and crawl cadence with Arquivo infrastructure before trusting the traffic.
- Respects robots.txt
- Yes
- Spoofing Risk
- Arquivo Web Crawler has medium spoofing risk because user-agent strings can be copied; pair the match with DNS, IP, behavior, or operator evidence.
- Risk
- Neutral
- Recommended Handling
- Depends
Rules and controls
- Robots.txt Snippet
-
User-agent: Arquivo-web-crawler Disallow: /
Relationships
- Operator
- Arquivo Checked 2026-06-23
Relationships without an Evidence link are normalized from the canonical directory record. They should not be interpreted as independent proof of physical presence or request origin.