How To Track Brand Mentions Across AI Models
The discipline is in how you report their output. Every one of them samples: their own prompt set, their own infrastructure, their own run frequency. Their number is an estimate from a particular vantage point, not a count of what happened.
In most categories the pages that generate answers are review platforms, directories, forum threads, comparison articles and trade publications. Being absent or wrong on those explains far more absences than anything on a brand's own site, and correcting a listing costs an afternoon.
What Padding Looks Like Screenshots of favourable answers with no indication of how many runs produced them. Industry news summaries that could have been written without opening your account. A rising score with no methodology. Traffic charts from unrelated channels included to fill space.
Use the Soft Signals Deliberately Two free signals carry more information than their informality suggests. Add a how did you hear about us question to your enquiry form and read the free text monthly rather than the categories.
The caveat is that most published question sections are marketing in disguise, containing questions no customer has ever asked, phrased to permit a favourable answer. Those get ignored, and they are easy to spot.
Record the conditions alongside the results: which assistant, which model version if visible, whether web access was on, the date and the run number. When a result changes sharply, the conditions log is usually what tells you whether the world changed or your setup did.
Measure Position Change in the Prompt Set This is the closest thing to an output metric that you can genuinely audit, because you own the instrument. Run a fixed prompt set on a fixed schedule under fixed conditions, and track four things:
Control the Session Conditions Personalisation quietly corrupts this. Run from a signed out session, or a fresh session with memory and history disabled, and do not use an account that has been researching your own company all week.
Tracking this is genuinely awkward, and pretending otherwise is how most reporting in this field goes wrong. There is no console. Answers vary between runs. Referral attribution is inconsistent between assistants. Anyone handing you a single confident number has hidden a great deal of variance behind it.
Preference is the wrong word, strictly. These systems do not have taste. They reach for sources that match the shape of the answer being written and that contain claims which can be lifted without distortion, and certain formats do that reliably.
The sustainable version is small and continuous: the prompt set run monthly, listings checked quarterly, a handful of pages updated rather than a burst of new ones, and someone who owns it. That costs less over a year than the three month push and holds its ground. answer engine optimization services
It does not contain a return on investment figure calculated from an assumed conversion rate applied to an estimated mention volume. That calculation looks rigorous and is a chain of guesses, and it will not survive the first person who asks where the first number came from.
Write it once, covering the category question, the problem question, the comparison question, the competitor question and the branded question. Fifty is a workable minimum. Then freeze it, and if you must add prompts later, add them as a separate cohort so the original series stays comparable.
What Matters More Than Format Two things outrank format choice entirely. The first is whether the content can be fetched and read at all, since a page behind a broken crawler rule or dependent on JavaScript is invisible whatever shape it takes.
You also cannot cleanly attribute a purchase to a recommendation the buyer received three weeks earlier in a conversation you never saw. That influence is real, it is often the main value of the channel, and it will not appear in any report you own.
It is also worth recording the reason for every rule you keep. A disallow line with no explanation gets preserved indefinitely through migrations and redesigns because nobody dares remove something they do not understand. A one line comment saying who added it and why turns a permanent mystery into a decision that can be revisited.
Keep a record of what you predicted as well as what you measured. Writing down at the start of a quarter what you expect to move, and then reading it back at the end, is the cheapest way to find out whether your model of this channel is any good. Most teams never do it, which is why the same confident explanations survive for years without ever being tested.
One structural decision saves a lot of trouble later. Keep the raw answers in plain text files named by date, assistant and run number, rather than pasting them into a document that gets reformatted. Six months in you will want to search across every run for the first appearance of a competitor or a source, and a folder of plain files supports that while a slide deck does not.