While building Agent Analytics we processed raw server logs from sites across the Flowgen Network. The goal was simple: understand what AI crawlers actually do when they land on a website, rather than what documentation says they do. Several of the findings contradict assumptions that most marketing and SEO teams still operate on.

Key findings

  • AI crawlers now account for a meaningful and fast-growing share of bot traffic, with GPTBot the most active AI crawler on the sites we analysed.
  • None of the major AI crawlers rendered JavaScript. Content that only exists after client-side rendering was effectively invisible to them.
  • Crawlers overwhelmingly favoured text-rich pages: documentation, guides, product detail and comparison pages, rather than homepages or navigation hubs.
  • On-demand retrieval bots (such as ChatGPT-User and PerplexityBot) re-fetched pages within minutes of a user question, which makes freshness a real-time concern.
  • A material share of requests claiming to be AI bots did not originate from the provider's published IP ranges.

Why server logs, not analytics

Google Analytics and similar tools instrument the browser. They fire a JavaScript tag, set a cookie and report what a human did. AI crawlers do not run that tag, so from an analytics dashboard it looks as if they never visited. Server logs, by contrast, record every request the origin or CDN handles, including the user agent, IP address, path and response code. They are the only honest record of machine traffic.

This matters because the question "is AI reading my site?" cannot be answered with the tools most teams already have. It can only be answered from the log.

Who is crawling the web

Among verified AI user agents, OpenAI's GPTBot generated the most requests on the sites in our sample, followed by crawlers from Anthropic, Perplexity, Meta and Amazon. Google-Extended, which governs whether Google can use content for Gemini, appeared alongside ordinary Googlebot traffic. Microsoft's Bingbot continues to serve as the retrieval layer for Copilot.

Share of verified AI crawler requests, by crawler
Flowgen Network sample, Q4 2024
  • GPTBot (OpenAI)~38%
  • ClaudeBot (Anthropic)~23%
  • PerplexityBot~14%
  • Meta-ExternalAgent~10%
  • Amazonbot~8%
  • Other AI crawlers~7%
Approximate shares of verified requests across the sampled sites. Googlebot and Bingbot excluded.

Just as important as the volume is the mix. Indexing crawlers (building training corpora) behave very differently from retrieval bots (fetching a page to answer a specific question). The first group crawls broadly and periodically; the second group arrives in bursts tied directly to user demand.

AI crawlers do not render JavaScript

This was the most consequential finding. Across every AI crawler we examined, the request pattern was the same: fetch the HTML document, and occasionally linked assets, but never execute scripts or request the API calls a single-page application makes after load. Googlebot has rendered JavaScript for years; AI crawlers, so far, do not.

For sites that render product information, pricing or article bodies on the client, the AI crawler sees a shell. Whatever the page looks like to a human, the model is trained on, or answers from, the empty version.

Client-rendered pageHTML arrives with an empty root element. Content is fetched by JavaScript the crawler never runs. AI sees nothing useful.
Server-rendered pageFull text, headings and structured data are in the initial response. AI crawlers index and cite the real content.

Which pages they fetch

Conventional SEO wisdom concentrates effort on homepages and high-authority landing pages. AI crawlers showed a different appetite. The pages fetched most often were dense with text and answered specific questions: documentation, help-centre articles, how-to guides, product specification and comparison pages. Thin landing pages and navigational hubs received a fraction of the attention.

Relative crawl intensity by page type
Requests per page, indexed to homepage = 1.0
4.2xDocs / help
3.1xGuides
2.6xProduct detail
2.0xBlog
1.0xHomepage
0.6xCategory hubs
Approximate values from the sampled sites.

Crawl frequency and freshness

Indexing crawlers returned on schedules that ranged from daily for large, frequently updated sites to several weeks for small ones. Retrieval bots were a different story: when a user asked an answer engine about a product, the corresponding page was often fetched within seconds, and fetched again on the next question. That means a stale or broken page is not a problem that waits for the next crawl; it is exposed to users immediately.

Retrieval bots<1 minFrom user prompt to page fetch
Large sitesDailyTypical indexing revisit
Mid-size sites~1 wkTypical indexing revisit
Small sites2–4 wkTypical indexing revisit

Spoofed crawlers are common

A surprising share of requests carrying an AI user agent string did not come from the IP ranges the provider publishes. Some were scrapers borrowing a trusted name; some were monitoring tools; some were simply misconfigured. Counting user agents without verification would have overstated AI traffic substantially on several sites. Every provider now publishes its ranges, and verification against them should be the default.

If you are not verifying the IP, you are not measuring AI crawlers. You are measuring whoever decided to call themselves one.

What this means for your site

  • Render on the server. If content is not in the initial HTML response, assume AI systems cannot see it.
  • Invest in the pages AI reads. Documentation, guides and detailed product pages are the surfaces answer engines draw from.
  • Treat freshness as real time. Retrieval bots fetch pages at the moment of the question, so errors and outdated content are shown to users immediately.
  • Verify before you count. Cross-check user agents against published IP ranges, or your numbers will be wrong.
  • Watch the log, not the tag. Server-side measurement is the only way to see this traffic at all.

These findings shaped the design of Agent Analytics, which ingests server logs from AWS, Cloudflare, Vercel and other platforms, verifies every AI request, and reports crawler behaviour in real time.

Methodology

We analysed raw origin and CDN request logs from a sample of sites in the Flowgen Network during Q4 2024, spanning software, media, retail and financial services. Requests were attributed to an AI crawler only when both the user agent matched a known pattern and the source IP fell inside the provider's published range. Page types were classified by URL pattern and template. JavaScript rendering was inferred from the absence of script, XHR and fetch requests following the document request. Figures in this post are approximate and vary by site.

Charles ZhouMember of Technical Staff, Flowgen