AI search
A repeatable method for monitoring AI search visibility
A prompt set, fixed run conditions, a scoring sheet and a correction log: how to measure whether AI assistants mention your store, month over month.
By CartKernel ยท Published
Monitoring AI search visibility is a measurement problem before it is a marketing problem. Answers vary between runs, between accounts and between regions, so a single check tells you almost nothing. What works is a fixed prompt set, run under fixed conditions on a fixed schedule, recorded in a sheet whose columns never change. Then the month-over-month movement means something, because everything except the answer was held constant.
This is the method, including the parts that are unglamorous and the claims it cannot support.
Build a prompt set that mirrors the buying journey
Write between twenty and forty prompts, in the words a shopper would use, spread across four stages. Fewer than twenty and one volatile answer swings the whole score. More than forty and nobody runs it.
| Stage | What the prompt sounds like | What a good result looks like |
|---|---|---|
| Problem | A description of a situation with no product named | Your category is described accurately and your guide is a source |
| Category | A request for options in a product type, with constraints | Your store or products appear among the options |
| Comparison | Two approaches or specifications weighed against each other | Your comparison content is used, with correct facts |
| Brand | Your brand name, alone or with a question | Your description, range, shipping and returns are stated correctly |
Include the constraints your buyers actually apply: budget bands, sizes, compatibility, delivery country, certifications, materials. Include at least three prompts that name your brand directly, because factual errors about your own store are the most damaging and the easiest to fix.
Freeze the set. Adding prompts mid-year makes the trend unreadable. Keep a separate list of new prompts and roll them in at a version boundary you record.
Fix the run conditions
Record and repeat the following for every run, because each of them changes answers:
- The surfaces. Choose the ones your buyers use, typically the AI features in Search, one or two assistants, and one answer engine that shows sources prominently.
- Signed in or signed out. Signed-out, memory-off sessions are the closer approximation of a first-time shopper.
- Location and language. Set them deliberately per market rather than accepting whatever the machine defaults to.
- Date and time window. Run the whole set within one day so the surfaces are comparable to each other.
- Verbatim prompt text. No paraphrasing between runs, ever.
Run the set monthly. Weekly runs mostly measure noise, and quarterly runs miss the change that a release caused.
Record the same fields every time
One row per prompt per surface, with these columns:
- Run date, surface, prompt id.
- Were you mentioned: brand, product, both or neither.
- Were you cited: is one of your URLs listed as a source, and which one.
- Position: roughly where in the answer the mention appears.
- Which other sources were used, in order.
- Factual accuracy: any statement about your brand or products that is wrong.
- A copy of the answer text, saved for later comparison.
The seventh column does more work than it looks like it should. Six months later, the saved text is the only way to tell whether an answer changed because your page changed or because the surface rewrote its style.
Score visibility rather than counting wins
Three numbers, reported per surface and in total:
- Mention rate. The share of prompts where the brand or a product appears at all.
- Citation rate. The share of prompts where one of your URLs is a listed source. This is the number that most closely tracks work you control, because a citation requires a retrievable page that answered the sub-question.
- Accuracy rate. The share of brand-stage prompts with no factual error about the store.
Report them as a trend line with the prompt set version marked. Do not average across surfaces that behave differently, and do not convert them into a single score. The three numbers point at three different pieces of work: presence, page quality and entity data.
Keep a correction log
Every wrong statement about your store goes into a log with the surface, the prompt, the incorrect claim and the likely source. Most errors trace back to something you can fix: an outdated shipping policy page, a discontinued product still described as current, a price that disagrees between the page and the feed, a brand name spelled differently on a third-party profile.
The correction log is the input to the entity consistency audit, which resolves the underlying disagreement rather than the individual answer. Errors that persist across surfaces almost always mean the same wrong fact is published somewhere you control.
Add the server-side half
Prompt monitoring measures the answers. Server data measures the consequence, and each is incomplete alone.
- Referral traffic. Sessions arriving from assistant and answer engine domains, isolated in analytics as their own channel so they stop landing in an undifferentiated bucket. The setup is covered in how do you track traffic from AI search and in GA4 ecommerce tracking.
- Crawler activity. Server logs show which AI crawlers request which URLs and how often, which tells you whether your priority pages are being fetched at all. The technique is in ecommerce log file analysis.
- Access checks. Confirm each crawler you want is permitted in robots.txt and that priority pages render their content in the initial HTML response. A page that needs client-side rendering to show its content may not be retrievable.
Referral volume from these surfaces is typically small relative to search, and the sessions behave differently because the shopper arrives already informed. Judge them on conversion rate and revenue per session rather than on volume.
Turn the log into a work queue
Monitoring that does not produce work is a hobby. After each monthly run, spend an hour turning the sheet into three lists:
- Prompts where a competitor page was cited and you have no equivalent page. These become guide or answer briefs.
- Prompts where you were mentioned without a citation. Usually a page exists but is not retrievable or not answer-first. These become page edits.
- Every entry in the correction log. These become data fixes.
Work the lists by revenue at stake, not by how annoying the answer was. A wrong shipping statement on a brand prompt outranks a missing mention on a low-margin category.
What this method cannot tell you
It cannot tell you market share, because you are sampling a set of prompts you chose. It cannot attribute revenue precisely, because much of the influence happens before any click. It cannot promise that a fix will produce a citation, because retrieval and ranking are not published and they change.
What it can do is show whether the store is trending toward being retrievable, correctly described and quoted, and give a dated record of what changed and when. That is enough to run the work. How individual surfaces select sources is discussed in how does Perplexity choose which stores to cite and how do products appear in Google AI Mode, and the ongoing program sits inside AI search visibility and ChatGPT shopping visibility.