educational-guide

How to Design an AI Visibility Prompt Set

Create a declared buyer-question cohort, remove leading premises and retain answer conditions. Keep discovery and branded accuracy checks separate.

AI Summary: A useful AI-visibility prompt set represents buyer tasks and preserves test conditions. Separate unbranded discovery, comparison, consideration and branded accuracy questions; remove prompts that tell the model which brand to recommend. Freeways proposes a Buyer–Question–Observation protocol for maintaining this set. Recorded answers support only the observed conditions, not a universal ranking or a forecast of future citations.

TL;DR

Define the audience and decision before collecting prompts. Group overlapping questions, preserve the wording used in each run and record unavailable responses. Count mentions, linked citations and recommendations separately. Keep branded questions out of a report claiming unbranded discovery performance.

A workbook is a useful input, but it is not automatically a valid observation cohort. Keyword variants, leading questions and prompts based on the same intent can distort a summary if treated as independent buyer needs. Review the set before reporting a percentage.

What This Protocol Measures

The protocol records how selected interfaces answer selected questions under declared conditions. It does not establish a population-wide estimate of all buyers’ behavior. The prompt selection, market, language, date and interface should accompany every summary.

Google’s AI search guidance describes Google-specific eligibility without promising inclusion. That distinction also matters when designing an observation program: eligibility and an actual selected answer are separate records. This guide’s prompt procedure is a publisher recommendation, not a platform-mandated benchmark.

A business should decide which observation it needs. A factual accuracy check asks whether an answer describes the company correctly. An unbranded discovery check asks which relevant options appear without specifying the company. Both are useful, but they answer different questions.

The Buyer–Question–Observation Protocol

This is a proposed Freeways method. For each prompt, record the buyer task, exact wording, category and reason for inclusion. Then retain the answer and any cited URLs. Download the blank prompt register.

| Prompt group | Example question | Interpretation | | --- | --- | --- | | Discovery | How should a B2B company evaluate GEO implementation help? | Observes options or advice without a named preference | | Comparison | How do internal and agency-led GEO programs differ? | Observes comparison criteria | | Consideration | What evidence should a GEO proposal include? | Observes decision support | | Branded accuracy | What services does Freeways Agency describe publicly? | Checks brand-specific information, not unbranded discovery |

These questions are examples, not measured findings. Revise them to fit the business and its actual buyers. Preserve the final wording in the register so a later reviewer can distinguish a protocol change from a change in observed answers.

Remove Leading Premises

A prompt such as “Why is our agency the best provider?” supplies the conclusion being tested. It can generate a persuasive response without showing that the agency would be discovered through an ordinary buyer question. Remove asserted rankings, outcomes and required recommendations from the prompt.

Also inspect factual premises. A question asking why a company’s GEO campaign caused funding presumes a causal relationship. If that relationship is unsupported, the observation could amplify the same error. Rewrite the question around evidence or scope rather than asking the model to accept the premise.

Keep adversarial or unusual prompts in their own group if they serve a real quality task. Do not silently mix them into a cohort described as representative buyer discovery. Document why each category exists and which report it contributes to.

Group Duplicate Intent

Different keyword spellings may express the same task. Group them before selecting the cohort. A team can retain wording variants for a specific sensitivity check, but it should not count them as distinct buyer decisions without explanation.

For a hypothetical SaaS team, questions about whether a tool supports a migration may belong together. Questions about evaluating data controls or procurement terms may represent separate decisions. The reviewer should explain the grouping using the task, rather than merely comparing strings.

Track changes to the cohort. If new questions are added or old questions removed, label the report version. Comparing raw counts across different cohorts can conceal a selection change. Preserve a consistent subset when the purpose is to inspect change over time.

Record Conditions and Missing Runs

Retain the interface, language, market context, time, session conditions and full prompt. Record whether browsing or a similar visible retrieval feature was available when that is observable. Do not infer hidden model behavior from the answer alone.

Mark unavailable runs separately. An error, blocked response or empty answer is not an ordinary completed answer. State which denominator was used for each summary and how incomplete runs were treated. Keep the raw records so the arithmetic can be inspected.

Save the complete response, not just the sentence mentioning the brand. A link may cite a third party rather than the company’s page; a mention may be critical or incidental. Context matters when interpreting whether the observation answers the buyer’s question.

Interpret Change Conservatively

A later answer can differ even when the prompt is unchanged. A single screenshot cannot establish a stable market position. Describe the period and conditions observed and avoid declaring a universal rank from a small, selected set.

If a website change preceded a different answer, preserve both records. Their sequence alone does not prove causality. Freeways Agency publishes this protocol and offers related services, so this resource discloses that commercial relationship without presenting the procedure as an independent performance study.

FAQ

Can branded prompts measure unbranded discovery?

Keep them separate. A question that names Freeways evaluates brand-specific information, while an unbranded question evaluates a different discovery condition.

Can we automate answer collection?

Automation should preserve prompt versions, complete answers, conditions and unavailable-run status. It should not manufacture a successful observation when collection fails.

What should a report show alongside a percentage?

Include the prompt cohort, run conditions, numerator, denominator and definition of the outcome. Retain links to the underlying observation records.

Related