How To Tell If Your AI SEO Agency Is Working
Product recommendations are a harder case than service recommendations, because the answer has to be specific enough to act on. A model naming a product is committing to a name, usually a price band and often a comparison, and it needs sources confident enough to support that.
After that, the work is ordinary: accurate structured data, honest comparison content, a steady flow of detailed reviews, and marketplace listings maintained as carefully as your own pages. trusted answer engine optimization agency
Ask what was done, not what happened. If listings were corrected, pages rewritten and outreach attempted, and the numbers are still flat, that is information about the market. If none of it happened, the numbers were never going to move.
Treat your marketplace listings as primary marketing assets rather than as a sales channel afterthought. Check the specifications match your own, that the product name is identical and that the category is right. A listing contradicting your own site creates exactly the inconsistency that stops mentions resolving.
The success measure should include whether the coverage contains a usable descriptive sentence, not only whether it appeared and whether it linked. And the briefing material should lead with specifics rather than with positioning language.
This is the pattern search followed, and there is no obvious reason for it to play out differently here. The advantage of early movement is not that the channel is large yet, it is that the positions are cheap.
One check is worth running independently once a quarter, without telling anyone. Take ten prompts from the agreed set, run them yourself in a signed out session, and compare what you find against the most recent report. Broad agreement is reassuring. A consistent gap in the agency's favour is the single most informative finding available to you, and it is not something a report will ever surface.
Ahrefs measured this in July 2025 across 15,000 long-tail prompts and four assistants, finding roughly 80 percent of cited pages did not rank for the original query, with about 12 percent in the top ten. The overlap is real but partial, which is the worst case for planning: you cannot ignore your rankings and you cannot rely on them either.
A useful way to think about the sequence is that each stage moved a task from the user to the interface. First the fact, then the summary, and now the comparison. Each move removed a reason to visit a website, and each was followed by an industry insisting the change had been overstated. It is reasonable to expect the pattern to continue rather than to stop at a convenient point.
Category Costs Rise as Coverage Fills In Influencing the third party sources assistants cite is easiest while those sources are thin. A category with two mediocre comparison articles is inexpensive to influence. The same category in three years, once somebody has built the definitive resource that every assistant settles on quoting, is not.
An agency doing the work sends these the same day, because they already exist as a by-product of the measurement. One that does not will explain that the platform does not export in that format, or that the data is summarised in the dashboard.
Distinguish between a supplier who is failing and one who is reporting badly, because the remedies differ entirely. Ask for the raw answers and read them yourself before deciding. It is not unusual to find that sound work has been buried under a dashboard nobody understands, and fixing the reporting is far cheaper and less disruptive than replacing a team that is actually doing the job.
A weak brief produces a generic proposal, and a generic proposal produces a generic engagement that spends the first two months discovering things you already knew. The brief is the cheapest lever you have over the quality of the work.
This is the whole argument in one sentence, and it is why the audit is worth running even if you intend to do nothing with the findings for six months. The measurement is cheap. Reconstructing a baseline you never took is impossible.
It is also worth doing while your category is boring. An audit run during a period of stability produces a clean baseline. One run in the middle of a competitor's campaign or immediately after a site migration measures the disruption rather than the position, and you will not know which you have unless you took the earlier reading.
That means the useful ask is not simply for a rating. Prompting customers to say what they used the product for and what situation it suited produces review text that can actually answer a question, which is what gets quoted.
Nor has any of this removed the need for a real product and real customers who will say so. If anything it has increased it, since corroboration from independent sources now feeds directly into whether a machine will recommend you.
When to Change Supplier Three conditions justify it individually. Raw answers cannot be produced on request. The prompt set has been changed without disclosure, which invalidates every comparison in every report you have received. Or two quarters have passed with the agreed inputs completed and no movement on citation presence, accuracy or source coverage.