Crawler Analytics is available on Business, Enterprise, and Agency Business plans, or through an explicit feature override.
Choose one delivery source
The setup page exposes these collectors:
AWS CloudFront is an active connector in the current setup flow. It is not a coming-soon integration.
A brand uses one activated delivery source. After activation, changing provider requires an explicit migration. A pending, unactivated connector can be replaced from the setup flow.
Create and protect a key
Create a dedicated key for the selected source. The plaintext secret is shown once; later the interface shows only its prefix. You can keep up to 5 active keys for an integration. To rotate, create the replacement, update the source, confirm delivery, and then delete the old key. Choose an IANA time zone when creating the first key. It defines daily snapshots and becomes immutable after activation.Provider-specific boundaries
Managed streams
- Vercel sends Static and Runtime production logs as JSON. Vercel cannot filter the drain by user agent, so forwarded human rows are discarded by Qwairy but can still count toward Vercel delivery volume. Sampling makes totals partial.
- Cloudflare Logpush should send only the fields listed by the setup page and apply its crawler pre-filter. The job must use
max_upload_bytes: 5000000andmax_upload_records: 10000. - CloudFront uses the exact Standard Logging v2 field list shown in setup, JSON output, a Firehose HTTP buffer of 1 MiB and 60 seconds, and S3 failure backup. Shorter delivery intervals are unsupported.
- Akamai uses the listed DataStream 2 fields, JSON output, Log User-Agent Header, and a push frequency of 60 seconds or longer.
- Fastly uses the provided minimal JSON format. Its setup allows up to 10,000 entries and 2,000,000 bytes per request; JSON arrays, NDJSON, and gzip are supported.
- Netlify sends its full Traffic Logs stream and has no crawler-only drain filter. User-Agent must remain available for classification. Qwairy ignores the client IP field and does not retain queries, human rows, or raw payloads.
Edge and application collectors
The Generic HTTP API acceptsPOST requests with Content-Type: application/json and X-API-Key. Each useful record contains method, path, hostname, user agent, timestamp, and optionally status code.
Filter for known AI crawler user agents in your server or edge layer before sending. A generic batch can contain up to 1,000 records and the normalized batch ceiling is 2,000,000 bytes.
The Cloudflare Worker and WordPress recipes apply this collector pattern at the edge or application layer. Review generated code before deployment and keep the key in a server-side secret.
Processing rules
The setup page lets you exclude paths and user agents before aggregation. Patterns are case-insensitive and support wildcards, one per line. You can configure up to 32 path rules and 32 user-agent rules. Each pattern is limited to 256 bytes and each rule set to 4,096 bytes.Read the dashboard
The available ranges are Last 24 hours, Last 7 days, and Last 30 days. Crawler Analytics uses analytic days in the integration’s configured time zone. Last 24 hours means the current analytic day, from its local midnight through the current snapshot; it is not a rolling 24-hour window. The other presets include the current analytic day and the preceding 6 or 29 analytic days. Dashboard totals, crawler breakdowns, charts, and page tables use the public set of supported observable AI-specific User-Agent identities. The ingestion registry can recognize additional identities for classification and compatibility, but those identities are excluded from the current public totals and tables. Classification matches a supported token in the HTTP User-Agent. Qwairy does not verify source IP or reverse DNS, so an occurrence is evidence of the received identity string, not proof of the requester’s origin. Traditional search crawlers and robots-policy tokens such asGoogle-Extended are excluded from public occurrence totals.
The current public identity set is:
Claude-SearchBot,MistralAI-Index,OAI-SearchBot, andPerplexityBot;ChatGPT-User,Google-GeminiNotebook,MistralAI-User,Perplexity-User,Claude-User, andGoogle-CloudVertexBot;ClaudeBot,GPTBot, andGrokBot.
Google-NotebookLM is accepted as an HTTP alias of the canonical Google-GeminiNotebook identity.
The page shows:
- observed occurrences;
- pages that returned at least one HTTP 2xx to a classified crawler;
- crawler identities observed;
- average frequency across observed daily buckets;
- trends, crawler distribution, page state, status-code groups, and page details.
- New: the first retained page observation falls within the last 7 days, including the boundary (
firstSeenAt >= now - 7 days). - Hot: the page is in the top decile by supported-crawler occurrences over the current 7 analytic days, has activity on at least 2 of those days, and is not neglected.
- Neglected: the most recent retained page observation is strictly older than the 30-day boundary.
New and Hot can appear together. Neglected is exclusive and cannot appear with either state. The hot calculation always uses the current 7-analytic-day window, independently of the dashboard range selected for displayed occurrences.
Missing delivery is not inferred as zero. The interface distinguishes observed, not measured, and incomplete-ingestion states. Quota or provider gaps produce incomplete_ingestion. The ≥ marker is reserved for a count whose numeric precision is capped; at-least-once delivery is reported separately because repackaged retries can duplicate observations.
Page-by-crawler daily rollups are retained for 30 analytic days. Delivery continuity evidence is retained for 13 months. Durable page identities keep firstSeenAt, lastSeenAt, and lifetimeCount after the daily rollups expire. A page identity is created only after at least one HTTP 2xx response.
Path privacy and matching
Before aggregation, Qwairy removes query strings and fragments. Sensitive-looking segments, including long numeric identifiers, email addresses, UUIDs, long hexadecimal values, JWTs, and high-entropy tokens, are replaced with an integration-scoped one-way pseudonym. Paths beyond the supported length are also bounded with a pseudonym. Occurrence counts and lifecycle timestamps remain real, but a pseudonymized path is not reconstructible or linkable across integrations. Page Performance does not attach a title to it and excludes it from GA4, Search Console, and Bing matching. Technical ingestion ceilings are 10,000,000 events per integration per day and 200,000,000 events per billed team per day. They protect shared infrastructure and do not cap page cardinality.Validate the connection
- Use the endpoint and authentication scheme displayed for the chosen source.
- Send the provider’s test event or a permitted crawler request.
- Wait for the connector to move from pending to connected after accepted rollup data.
- Check delivery-continuity warnings before interpreting a zero.
- Compare equivalent, unsampled source logs when reconciling totals.

