How Often Should You Re-Test Your AI Visibility
This is closer to public relations than to marketing operations, and it is the skill most teams are furthest from. It is also the one least suited to being learned quickly, which makes it the strongest argument for outside help.
Buying a Score Instead of Evidence A monthly number that rises is easy to present and impossible to audit. The vendor controls the number and the prompt set behind it, and a client has no way to distinguish real improvement from a methodology change.
Ask What They Cannot Measure A competent practitioner will volunteer limitations before you ask. Assistant answers vary between sessions. Referral attribution is inconsistent. Some assistants cannot be measured reliably at all. Sample sizes in the published research are small.
Ahrefs measured the overlap in July 2025 across 15,000 long-tail prompts and four assistants, finding roughly 80 percent of cited pages did not rank for the original query at all. Ranking gets a page considered. It does not reserve a seat.
What to Build and What to Buy Build the prompt set and the measurement habit internally. They are cheap, they depend on knowledge of your customers that no agency has, and owning them means you can audit anyone you hire.
The specific damage is that somebody sees a dip, rewrites a page, sees the number recover for unrelated reasons, and concludes the rewrite worked. That false lesson then gets applied elsewhere. A slower cadence with more runs per prompt is more informative than a faster one with fewer.
The guard against this is boring and effective. Change one substantial thing at a time where you can, record what you did and when, and note the alternative explanations alongside your conclusion. Attribution in this channel is genuinely hard, and a team that admits that will make better decisions than one that produces a confident causal story after every movement.
Ask one final question before signing: what would you tell me if this is not working after six months? The answer reveals whether they have thought about failure, and an agency that has not thought about failure will not recognise it. ai search visibility
What you are looking for is whether the questions sound like a buyer wrote them. If every prompt contains the client's category name phrased the way an internal marketing team would phrase it, they have tested how the brand talks rather than how customers ask.
Keep the brief to something you would be willing to send to three suppliers unchanged. The temptation is to tailor each one, which feels attentive and makes the resulting proposals impossible to compare. Identical briefs produce differences that reflect the agencies rather than the instructions, which is the entire point of asking more than one.
State What You Sell in Concrete Terms Price range, lead time, geography, capacity, what you decline. This feels commercially sensitive and it is the material that makes your pages quotable, so an agency that does not have it will write vague content by necessity.
One scheduling detail improves comparability more than it should. Run on roughly the same date each month rather than whenever somebody remembers. Retrieval behaviour and the freshness of competing sources both vary over a month, and a series taken at irregular intervals introduces variation that looks like a trend.
Log the conditions with every run, including which assistant, which mode, whether web access was enabled and the date. When a result moves sharply, the conditions log is usually what tells you whether the world changed or your setup did.
Retrieval behaviour changes, competitors keep publishing, listings go stale, product details change and reviews accumulate. A position secured once is not held without maintenance, which is the same lesson search taught over twenty years and which is being relearned rather than transferred.
Look at What They Do About Third Party Sources This is where the real work lives and where weak proposals are thinnest. Ask specifically what they will do about the review platforms, directories, forums and comparison articles that assistants actually cite in your category.
Include one deliberately open question at the end, asking what they would do differently from what the brief proposes. A good supplier will disagree with something, and the disagreement is worth more than the rest of the proposal because it shows they read the situation rather than the request. A response that agrees with every assumption in your brief has told you nothing you did not already believe.
Build the run into an existing routine rather than creating a new one. Measurement programmes in this field fail through quiet abandonment rather than through a decision, and a modest set attached to an established monthly process survives far longer than an ambitious one that depends on somebody remembering to start it.
In most categories the pages that generate answers are review platforms, directories, forum threads, comparison articles and trade publications. Being absent or wrong on those explains far more absences than anything on a brand's own site, and correcting a listing costs an afternoon.