Overview
Terracotta is a web scraper from Ceramic used for public web data collection, page extraction, content monitoring, and third-party crawler activity.
Its primary user-agent pattern is Terracotta; related patterns include Terracotta-News; a representative HTTP user-agent is Terracotta Terracotta-News.
Terracotta is Unverified at the identity-evidence level. The listed identity remains useful for detection, but this record does not currently contain authoritative evidence sufficient to authenticate the identity claim.
Terracotta is marked as not reliably governed by robots.txt directives; use server-side rules if the traffic should be restricted.
Terracotta can usually be allowed after confirming the source and monitoring request volume.
Identity
- User-Agent
Terracotta- Aliases
- Terracotta-News
- HTTP Agent Examples
Terracotta Terracotta-News- Robots.txt Token
Terracotta- Identity Type
- Observed
- Evidence Method
- Treat `Terracotta` as an identity signal only. Confirm it with current operator documentation, cryptographic verification, forward-confirmed reverse DNS, source-network ownership, or other authoritative evidence before trusting the claimed identity.
Classification
- Type
- Scraper
- Kind
- Crawler
- Family
- Terracotta
- Purpose
- scraping
Behavior and handling
- Common Use
- Terracotta is used for public web data collection, page extraction, content monitoring, and third-party crawler activity.
- Detection Notes
- Terracotta traffic is primarily detected by the `Terracotta` user-agent pattern; related patterns include `Terracotta-News`; a representative HTTP user-agent is `Terracotta Terracotta-News`. Compare source IPs, reverse DNS, request paths, and crawl cadence with Ceramic infrastructure before trusting the traffic.
- Respects robots.txt
- No
- Spoofing Risk
- Terracotta has medium spoofing risk because user-agent strings can be copied; pair the match with DNS, IP, behavior, or operator evidence.
- Risk
- Safe
- Recommended Handling
- Depends
Rules and controls
- Robots.txt Snippet
# This agent may ignore robots.txt. Use authenticated access controls or network policy when blocking is required.