Overview
CriteoBot is a search crawler from Criteo used for search indexing, URL discovery, page rendering, result freshness, and search-quality checks.
Its primary user-agent pattern is CriteoBot; related patterns include CriteoBot/; a representative HTTP user-agent is CriteoBot/0.1 (+https://www.criteo.com/criteo-crawler/).
CriteoBot is verified with Medium confidence. The identity type is Verified Bot, and the evidence basis is a public crawler reference or source-linked documentation.
CriteoBot is marked as not reliably governed by robots.txt directives; use server-side rules if the traffic should be restricted.
CriteoBot should be reviewed against site policy, source evidence, crawl rate, and requested paths before a permanent allow or block rule is created.
Identity
- User-Agent
CriteoBot- Aliases
- CriteoBot/
- HTTP Agent Examples
CriteoBot/0.1 (+https://www.criteo.com/criteo-crawler/)- Robots.txt Token
CriteoBot- Identity Type
- Verified Bot
- Evidence Method
- Verify CriteoBot by matching `CriteoBot` to Criteo evidence, then checking reverse DNS, source-network ownership, signed request data, or published crawler documentation when available.
Classification
- Type
- Search
- Kind
- Crawler
- Family
- Criteo
- Purpose
- indexing
Behavior and handling
- Common Use
- CriteoBot is used for search indexing, URL discovery, page rendering, result freshness, and search-quality checks.
- Detection Notes
- CriteoBot traffic is primarily detected by the `CriteoBot` user-agent pattern; related patterns include `CriteoBot/`; a representative HTTP user-agent is `CriteoBot/0.1 (+https://www.criteo.com/criteo-crawler/)`. Compare source IPs, reverse DNS, request paths, and crawl cadence with Criteo infrastructure before trusting the traffic.
- Respects robots.txt
- No
- Spoofing Risk
- CriteoBot has medium spoofing risk because user-agent strings can be copied; pair the match with DNS, IP, behavior, or operator evidence.
- Risk
- Neutral
- Recommended Handling
- Depends
Rules and controls
- Robots.txt Snippet
# This agent may ignore robots.txt. Use authenticated access controls or network policy when blocking is required.