Most AI Search measurement stops at visibility: did the engine mention us? But answer engines write long, opinionated responses, and much of what they say about a brand was never requested. We call this the Parrot Problem: engines repeat and embellish whatever the web says about you, accurate or not.
Key findings
- Nearly half (47%) of AI Search response content is unsolicited editorial: comparisons, rankings and recommendations the user did not ask for.
- Answer engines favour recent content: the top 50% of cited content is less than 13 weeks old.
- Specific, measurable claims are far more extractable than generic marketing language.
- Roughly 11% of claims about one brand in answer engine responses were inaccurate or false.
- Agents make monitoring and fixing accuracy possible at the scale engines operate.
Here’s what we found about the Parrot Problem
We analyzed 50,000 prompts across seven industries and five prompt types. Responses are long: ChatGPT averaged 3,043 characters per response, Gemini 2,921 and Claude 1,590. Nearly 47% of that content was unsolicited editorial material such as comparisons, referrals and rankings.
The editorial content comes in seven recognisable types:
A real world example
Southwest Airlines ended its famous no-assigned-seating policy in January 2026. Months later, answer engines were still describing open seating as a reason to choose the airline. Nobody asked about seating; the engines volunteered it, and they were wrong. Editorial additions spread inaccurate brand information at scale.
Three ways to show up correctly in AI Search
1/ Find gaps in accuracy and poor sentiment
Audit what engines actually say, not just whether they mention you. Analysing millions of responses, one wearable manufacturer discovered that approximately 11% of claims made about it in answer engines were inaccurate or false. Those claims are a fixable backlog once they are visible.
2/ Focus on original, fresh, specific marketing
Answer engines reward recency: the top 50% of content they cite is less than 13 weeks old. They also reward specificity. "Enterprise grade" contains no extractable claim; "P95 latency at 38 milliseconds" does, and models retrieve and repeat it.
3/ Deploy agents to run this process at scale
Checking millions of claims and refreshing hundreds of pages is not a job for a spreadsheet. Flowgen built a publication outreach agent that pitches corrections and new data to the sources engines cite, and a content refresh agent that keeps owned pages inside the freshness window. Humans stay in the loop to approve every change.
The parrots aren’t going anywhere
The Parrot Problem is not a bug engines will patch. Synthesising, comparing and recommending is what they are for. The answer for brands is operational: audit claims regularly, invest in specific and fresh content, and use agents so the work scales with the engines.