Dashboard tracking brand mentions across ChatGPT, Perplexity, Google AI Overviews, and Gemini

GEO Case Study: Tracking Brand Mentions Across AI Search Tools

A mid-sized software company noticed something odd a few months back. Their organic traffic from Google had held roughly steady, nothing dramatic either way, yet they were suddenly getting more direct sign-ups from people who mentioned “I asked ChatGPT and it recommended you” in onboarding surveys. Nobody on the marketing team had run a campaign around that. Nobody had even been tracking it. It was just happening, quietly, in the background, completely invisible to their existing analytics.

That gap is the whole reason to track brand mentions AI search tools generate, separately from traditional traffic and ranking reports. This case study walks through how that team actually built a process to catch what their dashboards were missing, and what they found once they started looking properly.

The Problem: Invisible Wins and Losses

Before building any tracking process, the team had no idea whether AI tools were citing them accurately, citing a competitor instead for the same questions, or not mentioning either brand at all. That uncertainty made it impossible to know whether their content strategy was actually working in this new channel or just coasting on assumptions.

Traditional analytics couldn’t help here. Google Analytics shows traffic that arrives from a click, and a citation inside an AI answer often generates no click at all, or a delayed one days later that’s nearly impossible to attribute back to that specific mention. The team needed a different kind of visibility entirely, one built around direct observation rather than passive measurement.

Step One: Building a Question Set

The first real step was compiling a list of maybe thirty specific questions their target customers would plausibly ask an AI assistant, questions genuinely related to their product category, not just their brand name directly. This mattered because most people don’t type a company’s name into ChatGPT looking for a recommendation. They ask a broader question and see who gets suggested.

Questions ranged from direct comparisons, like “what’s the best tool for X,” to more specific troubleshooting queries where their own documentation was strong. This mix mattered because different question types revealed different things: direct comparison queries showed competitive positioning, while specific technical queries showed whether their own content was being trusted as a reference.

Step Two: Manual Testing Across Multiple Tools

With the question set built, the team ran each question through ChatGPT, Perplexity, and Gemini, logging exactly what came back each time, including whether their brand appeared, whether a competitor appeared instead, and whether any source citation was included at all. This part was genuinely manual and a little tedious, since polished tracking tools for this space are still maturing, but it produced real, concrete data instead of a guess.

They repeated this process every two weeks rather than once, since answers weren’t perfectly consistent between runs. That repetition mattered more than expected. A single test run made results feel noisier and less reliable than they actually were once patterns emerged over several cycles.

What They Found

The results were more specific than anyone expected going in. For direct comparison questions, they were mentioned in roughly half of responses across the three tools, usually alongside two or three competitors, rarely alone. For troubleshooting and how-to questions tied to their documentation, they showed up far more consistently, cited by name in the large majority of relevant queries.

That split told them something concrete. Their comparison-focused marketing content wasn’t distinct or specific enough to consistently win a citation among several similar competitors. But their technical documentation, written plainly, updated regularly, with real specificity, was already performing well in this channel without anyone having deliberately optimized it for AI citation at all.

Perplexity cited them most consistently and transparently of the three tools, likely tied to its heavier reliance on live retrieval. ChatGPT’s behavior was the least predictable, sometimes citing them clearly, sometimes referencing information without any visible source at all. Gemini’s results tracked reasonably closely with their existing Google search rankings, suggesting real overlap between traditional ranking signals and its citation behavior.

What They Changed as a Result

Rather than overhauling everything at once, the team focused first on their weakest area: comparison content. They rewrote their top comparison pages to include more specific, original detail, real benchmark numbers instead of vague claims, rather than the generic feature-list format most competitors were also using. That specificity was the piece most clearly missing compared to their stronger-performing technical content.

They also began repeating this tracking process monthly going forward, treating it as an ongoing metric rather than a one-time investigation, specifically to track brand mentions AI search tools generate over time and catch shifts before they became a larger, unnoticed problem.

What This Case Study Suggests More Broadly

The gap between what traditional analytics shows and what’s actually happening across AI tools can be significant, and it often hides in exactly the areas a team assumes are already fine. This company had no idea their technical documentation was quietly outperforming their marketing content in this specific channel until they built a process to actually look.

Manual, repeated testing across multiple tools, even without sophisticated software, is enough to surface genuinely useful patterns. The effort required is real, but it’s a fraction of what most teams already spend on traditional SEO tracking, and it currently captures visibility that would otherwise stay completely hidden.

Reference

The Bottom Line

To track brand mentions AI search tools generate, you don’t need to wait for perfect tooling. A defined question set, tested manually and repeatedly across ChatGPT, Perplexity, and Gemini, is enough to reveal real patterns, where you’re winning, where competitors are beating you, and which of your existing coSntent is already quietly performing well in a channel your dashboards can’t see at all.