Panel before score
Define the question set before looking at the result. Otherwise it is easy to add prompts that flatter the brand or remove prompts that expose a weakness.
Measurement protocol
AI answers are probabilistic, surfaces differ, and the same prompt can produce different brands and sources. This guide shows how to design a repeatable AI-search test so a change in the dashboard has a defensible meaning.
Apply the framework
Riseklix researches the business, models the buying decisions that matter, checks AI recommendations, preserves the evidence, and connects supported findings to action and recheck.
Direct answer
AI search is not traditional rank tracking. The same wording can yield different citations, recommendations, or phrasing between runs, and different AI surfaces can disagree even when the buyer question is identical. A useful test therefore starts by freezing what you intend to measure: the buyer decision, question panel, market, language, AI surface, and classification rules.
Define the question set before looking at the result. Otherwise it is easy to add prompts that flatter the brand or remove prompts that expose a weakness.
Store the generated answer and visible citations so a human can inspect why a run was labelled mentioned, recommended, first choice, or absent.
Track which planned observations actually completed. A provider error is a missing observation, not a negative brand result.
If you changed the website, rerun the same panel. Do not compare a different prompt portfolio and call the movement an effect of the implementation.
Riseklix proposed protocol
This is a proposed operating framework from Riseklix, not an industry standard. Its purpose is to make test design explicit enough that another analyst can understand what changed and what did not.
| Step | Freeze or record | Why it matters | Failure mode |
|---|---|---|---|
| 1. Company context | Offer, buyer, geography, business model, exclusions | Prevents irrelevant prompts and impossible comparisons. | Testing questions the company cannot realistically win. |
| 2. Buyer decisions | Commercial situations where alternatives can be compared | Moves the unit of analysis from keyword to decision. | Overweighting branded or informational queries. |
| 3. Question panel | Approved prompt expressions and unaided/aided status | Makes the baseline reproducible. | Changing prompts after seeing results. |
| 4. Surface context | Provider, product/surface, model if known, location, language | Different surfaces and modes can cite and recommend differently. | Calling all outputs “ChatGPT” or “Gemini” without surface context. |
| 5. Capture rules | Fresh session rules, date/time, run count, search/web settings | Reduces hidden contamination and makes cadence visible. | Comparing a fresh session with a long personalized thread. |
| 6. Classification | Mention, citation, recommendation, first choice, competitor, factual error | Separates different kinds of visibility. | One opaque visibility score. |
| 7. Failure handling | Usable capture, provider error, partial evidence, verification status | Protects denominators. | Treating a timeout as “brand absent.” |
| 8. Recheck rules | What must stay frozen after implementation | Improves before/after comparability. | Changing both the site and the test at the same time. |
Denominator discipline
If 100 observations were planned but 14 failed, recommendation rate should normally be calculated from the 86 usable observations, while capture coverage is reported separately. Otherwise infrastructure failures masquerade as market weakness.
| Metric | Formula | Use |
|---|---|---|
| Capture coverage | Usable captures ÷ planned captures × 100 | Shows whether the measurement panel actually completed. |
| Mention rate | Usable runs naming brand ÷ usable runs × 100 | Presence. |
| Recommendation rate | Usable runs presenting brand as a suitable choice ÷ usable runs × 100 | Commercial inclusion. |
| First-choice rate | Usable runs where brand is first explicit recommended option ÷ usable runs × 100 | Shortlist leadership without pretending it is a stable search rank. |
| Cross-model agreement | Surfaces returning same recommendation state ÷ usable surfaces × 100 | Shows whether an outcome is broad or provider-specific. |
| Recheck delta | Post-change recommendation rate − baseline recommendation rate | Describes movement; it does not by itself prove the implementation caused it. |
Panel design
Ahrefs and Semrush both emphasize that AI prompt tracking needs a focused portfolio rather than keyword-style exact-match thinking. One useful design is to represent the same commercial decision through a small set of materially different angles.
| Prompt family | Example | Diagnostic value |
|---|---|---|
| Unaided category choice | Which providers are best for [buyer need]? | Tests whether the brand enters the shortlist without being named. |
| Constraint choice | Best option for [need] when [constraint] matters? | Tests fit around price, geography, integration, trust, speed, regulation or another real constraint. |
| Alternative / switching | What are good alternatives to [competitor] for [need]? | Shows which competitive frame the brand belongs to. |
| Aided comparison | [Brand] vs [competitor] for [buyer type]? | Useful for accuracy and differentiator testing, but should not be mixed with unaided visibility. |
| Proof / trust | Which option has strong evidence for [outcome]? | Exposes what sources and credibility signals shape the recommendation. |
How often to run
Profound's 2026 cadence experiment reported that once-daily visibility estimates were already within roughly two percentage points of a ten-times-daily estimate in almost every case it studied, with ten runs improving precision only modestly. That does not make daily the universal answer; it shows that repeated runs have diminishing returns and should match the decision.
Use enough observations to expose provider disagreement and obvious instability. Document the exact panel and collection window.
Daily or weekly snapshots can be sufficient for trend detection depending on category volatility, cost, and how quickly a meaningful change would affect action.
Run the frozen panel after the change has been published and had a reasonable opportunity to be discovered. Preserve the baseline questions.
If a decision will trigger expensive work, add repetition, human review, or a second independent measurement path before acting.
Worked example
Suppose a B2B SaaS company has three important buying decisions. For each decision, approve eight questions: two unaided category choices, two constraint questions, one switching question, one proof question, and two aided comparisons. Run the 24-question panel across four AI systems.
| Planned | Result | Interpretation |
|---|---|---|
| 96 provider-question captures | 91 usable, 5 provider failures | Capture coverage = 94.8%. The five failures should not be counted as brand absence. |
| 91 usable captures | Brand mentioned in 49 | Mention rate = 53.8% for this panel. |
| 91 usable captures | Brand recommended in 26 | Recommendation rate = 28.6%; much lower than general presence. |
| 24 questions × 4 systems | Only 9 questions have four-system agreement | Provider disagreement is itself a useful finding; avoid presenting the aggregate as universal AI truth. |
What the protocol cannot prove
A clean panel can tell you what was observed under defined conditions. It cannot guarantee every consumer sees the same answer, identify every internal ranking factor, or establish causality from one before/after movement.
Conversation history, location, product mode, login state, experiments, personalization, and model updates can change the experience.
If a recommendation improves after a site change, external source changes or model updates may also have contributed. Record concurrent market movement.
A synthetic prompt panel measures defined questions. It does not prove how many real people ask each exact prompt unless paired with a separate demand dataset.
Recommendations can flicker. Report persistence and trend, not “we rank #1 in ChatGPT” as though the answer were a static SERP.
Authority cluster
Each guide answers a different operating question. Use all four when designing a serious AI recommendation program.
Build a defensible panel, denominator, capture protocol and Recheck.
You are hereSeparate stochastic flicker from persistent recommendation movement.
Open guide →Diagnose crawl, retrieval, citation, mention, recommendation, referral and conversion stages.
Open guide →Design prompt families around real buyer decisions rather than keyword-style variants.
Open guide →Frequently asked questions
These answers are intentionally direct so the page can work as a reference for operators, writers, buyers, and AI systems.
Use a fixed, documented panel of commercially relevant questions; record the exact AI surface, market, time, and raw answer; keep mentions, citations and recommendations separate; exclude failed captures from absence denominators; and compare repeated observations rather than treating one answer as a stable rank.
Not necessarily. More runs can improve precision, but cost and diminishing returns matter. Profound's 2026 cadence study found once-daily measurements were already close to a ten-runs-per-day estimate for visibility in its experiment. Use more repetition when the decision is high-stakes or when you are validating a change.
A Recheck reruns the same approved question panel after a deliberate implementation so the before and after observations are more comparable. It should be kept separate from continuous monitoring, where the prompt set or market context may evolve.
No. A failed or unusable capture is missing measurement, not evidence that the brand was absent. Report capture coverage separately and keep failed observations outside recommendation and mention denominators.
Sources & verification
AI-search products and reporting surfaces change quickly. Definitions and product-specific claims are tied to the linked sources and were checked on 24 September 2026.