Building Content That Language Models Quote
Screenshots of favourable answers with no run count, which say nothing about how many attempts produced them. Impressions or traffic from unrelated channels included to fill a report. And activity described in the language of effort, such as ongoing optimisation, with no countable output attached.
Then Measure Again, and Keep Measuring A single snapshot tells you very little. Assistants vary their answers between sessions, between accounts and between model versions, so one run is a sample and not a verdict. Re-run the same prompt set on a fixed schedule and watch the trend rather than any individual answer.
One check is worth running independently once a quarter, without telling anyone. Take ten prompts from the agreed set, run them yourself in a signed out session, and compare what you find against the most recent report. Broad agreement is reassuring. A consistent gap in the agency's favour is the single most informative finding available to you, and it is not something a report will ever surface.
Then re-run the same ten prompts a month later and compare. Attributing improvement to a specific fix is only possible because you recorded the starting point, which is the argument for doing the teardown before the work rather than after it.
An agency doing the work sends these the same day, because they already exist as a by-product of the measurement. One that does not will explain that the platform does not export in that format, or that the data is summarised in the dashboard.
One inversion is worth noticing in your own analytics. The pages that earn citations are frequently not the pages that earn traffic, and teams optimising purely for sessions will deprioritise exactly the specification and comparison content that this channel uses. Keeping a separate note of which pages appear in citation lists prevents a well performing asset being retired because its visit numbers looked unremarkable.
Citation happens at the level of a passage, not a page. A model attaches a source to a specific claim it lifted, which means the real unit of work is a paragraph that stays true and useful once it has been removed from everything around it.
There is a variant of this worth checking separately. Sometimes you appear and the competitor appears above you, which is a different problem from being absent. In that case compare the specificity of the two descriptions rather than the sources: the company described in concrete terms tends to be listed first, because a specific description is easier to justify than a general one.
It is also worth checking which assistant your customers actually use rather than assuming. The answer varies by profession, age and country far more than industry commentary suggests, and several businesses have built measurement programmes around a system their buyers never open. Adding one question to your enquiry form settles it in a fortnight and can redirect the whole effort.
Before leaving, make sure you take the prompt set, the baseline archive and everything published. If those were not yours under the contract, that is a lesson for the next agreement rather than something to negotiate at the exit.
Keep a dated note of what you observed each quarter, including behaviour that later turned out to be temporary. The value is not in the individual observations, most of which expire, but in noticing how fast they expire. A team that has watched three of its confident conclusions become wrong within a year develops the right amount of scepticism about the fourth.
Track three things over time: how often you are named, ai seo services which sources get cited when you are, and which competitors appear alongside you. Movement in the second of those usually predicts movement in the first.
When to Change Supplier Three conditions justify it individually. Raw answers cannot be produced on request. The prompt set has been changed without disclosure, which invalidates every comparison in every report you have received. Or two quarters have passed with the agreed inputs completed and no movement on citation presence, accuracy or source coverage.
This variability is the main practical trap. Testing without web access and concluding you are invisible measures the training corpus rather than current retrieval, and the two can disagree sharply. Record which mode you used with every run.
Two implications follow regardless of which system you are studying. Being findable by the underlying search step is necessary, and being worth quoting once fetched is what decides whether you are used. Almost everything actionable sits in those two requirements.
The other habit worth building is writing down the number rather than the impression. Teams know their typical lead time, their price band and the size of job they decline, and almost never publish any of it, because a range feels like a commitment. It is a commitment, and it is also the only part of the page a machine can use, which makes it the difference between a page that gets cited and one that does not.