Arquivo Web Crawler
Feed Directory evidence: Verified

Arquivo Web Crawler

Arquivo Web Crawler is a feed fetcher from Arquivo used for RSS or Atom feed polling, syndication refresh, subscription delivery; it appears in server logs as `Arquivo-web-crawler`.

Arquivo-web-crawler
Operator Arquivo
Risk Neutral

Overview

Arquivo Web Crawler is a feed fetcher from Arquivo used for RSS or Atom feed polling, syndication refresh, subscription delivery, and content update checks.

Its primary user-agent pattern is Arquivo-web-crawler; a representative HTTP user-agent is Arquivo-web-crawler (compatible; heritrix/3.4.0-20200304 +https://arquivo.pt/faq-crawling).

Arquivo Web Crawler is verified with Medium confidence. The identity type is Verified Bot, and the evidence basis is a public crawler reference or source-linked documentation.

Arquivo Web Crawler is marked as respecting robots.txt directives for crawler access control.

Arquivo Web Crawler should be reviewed against site policy, source evidence, crawl rate, and requested paths before a permanent allow or block rule is created.

Identity

User-Agent Pattern
Arquivo-web-crawler
HTTP Agent Examples
Arquivo-web-crawler (compatible; heritrix/3.4.0-20200304 +https://arquivo.pt/faq-crawling)
Robots Token
Arquivo-web-crawler
Identity Type
Verified bot
Evidence Method
Verify Arquivo Web Crawler by matching `Arquivo-web-crawler` to Arquivo evidence, then checking reverse DNS, source-network ownership, signed request data, or published crawler documentation when available.

Classification

Type
Feed
Kind
Fetcher
Family
Arquivo
Purpose
Feed fetch

Behavior and handling

Common Use
Arquivo Web Crawler is used for RSS or Atom feed polling, syndication refresh, subscription delivery, and content update checks.
Detection Notes
Arquivo Web Crawler traffic is primarily detected by the `Arquivo-web-crawler` user-agent pattern; a representative HTTP user-agent is `Arquivo-web-crawler (compatible; heritrix/3.4.0-20200304 +https://arquivo.pt/faq-crawling)`. Compare source IPs, reverse DNS, request paths, and crawl cadence with Arquivo infrastructure before trusting the traffic.
Respects robots.txt
Yes
Spoofing Risk
Arquivo Web Crawler has medium spoofing risk because user-agent strings can be copied; pair the match with DNS, IP, behavior, or operator evidence.
Risk
Neutral
Recommended Handling
Depends

Rules and controls

Robots.txt Snippet
User-agent: Arquivo-web-crawler Disallow: /

Relationships

Operator
Arquivo Checked 2026-06-23

Relationships without an Evidence link are normalized from the canonical directory record. They should not be interpreted as independent proof of physical presence or request origin.

Similar Bots

Feed Verified

FeedOtter

FeedOtter

FeedOtter
Feed Unverified

WMF Citoid

Wikimedia Foundation

Citoid/WMF