Overview
Internet Archive – Archive-It is a feed fetcher from Archive-It used for RSS or Atom feed polling, syndication refresh, subscription delivery, and content update checks.
Its primary user-agent pattern is Archive-It; related patterns include Mozilla/5.0 (X11; Linux x86_64; archive.org_bot; Archive-It; +http://archive-it.org/files/site-owners-special.html) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36.
Internet Archive – Archive-It is verified with Medium confidence. The identity type is Verified Bot, and the evidence basis is a public crawler reference or source-linked documentation.
Internet Archive – Archive-It is marked as not reliably governed by robots.txt directives; use server-side rules if the traffic should be restricted.
Internet Archive – Archive-It should be reviewed against site policy, source evidence, crawl rate, and requested paths before a permanent allow or block rule is created.
Identity
- User-Agent Pattern
-
Archive-It - Aliases
- Mozilla/5.0 (X11; Linux x86_64; archive.org_bot; Archive-It; +http://archive-it.org/files/site-owners.html) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36
- HTTP Agent Examples
-
Mozilla/5.0 (X11; Linux x86_64; special_archiver; Archive-It; +http://archive-it.org/files/site-owners-special.html) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36 Mozilla/5.0 (X11; Linux x86_64; archive.org_bot; Archive-It; +http://archive-it.org/files/site-owners.html) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36 Mozilla/5.0 (compatible; special_archiver; Archive-It; +@http://archive-it.org/files/site-owners-special.html) Mozilla/5.0 (compatible; archive.org_bot; Archive-It; +@http://archive-it.org/files/site-owners.html) - Robots Token
- Archive-It
- Identity Type
- Verified bot
- Evidence Method
- Verify Internet Archive - Archive-It by matching `Archive-It` to Archive-It evidence, then checking reverse DNS, source-network ownership, signed request data, or published crawler documentation when available.
Classification
- Type
- Feed
- Kind
- Fetcher
- Family
- Archive-It
- Purpose
- Feed fetch
Behavior and handling
- Common Use
- Internet Archive - Archive-It is used for RSS or Atom feed polling, syndication refresh, subscription delivery, and content update checks.
- Detection Notes
- Internet Archive - Archive-It traffic is primarily detected by the `Archive-It` user-agent pattern; related patterns include `Mozilla/5.0 (X11; Linux x86_64; archive.org_bot; Archive-It; +http://archive-it.org/files/site-owners-special.html) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/130.0.0.0 Safari/537.36`. Compare source IPs, reverse DNS, request paths, and crawl cadence with Archive-It infrastructure before trusting the traffic.
- Respects robots.txt
- No
- Spoofing Risk
- Internet Archive - Archive-It has medium spoofing risk because user-agent strings can be copied; pair the match with DNS, IP, behavior, or operator evidence.
- Risk
- Neutral
- Recommended Handling
- Depends
Rules and controls
- Robots.txt Snippet
-
# This agent may ignore robots.txt. Use authenticated access controls or network policy when blocking is required.
Relationships
- Operator
- Archive-It Checked 2026-06-23
Relationships without an Evidence link are normalized from the canonical directory record. They should not be interpreted as independent proof of physical presence or request origin.