# Measure AI Mentions and Citations Without Mixing Them

Canonical: https://geo.freeways.agency/guides/measure-ai-mentions-and-citations
Publisher: Freeways Agency
Editorial status: published

> **AI Summary:** A brand mention names a company; a citation references a source, usually with a URL. Measure these separately on a defined prompt set, preserve the answers and test conditions, and report accuracy and buyer relevance alongside frequency. The proposed protocol treats crawler traffic as a separate access signal, not a citation.

## TL;DR

Create a small, fixed set of buyer questions. Separate branded questions from discovery questions. Record every valid observed answer, cited source, run condition, and unavailable run.

Use consistent denominators for rates and retain the raw evidence. A single favorable response is not proof that a brand is widely recommended.

## Define What You Are Counting

A mention answers whether the brand is named in the generated response. A citation answers whether the response references an identified source. The cited source may be the brand's website, a third-party page, or an unrelated page that discusses the same category.

For a Freeways measurement exercise, “Freeways Agency” in the answer would count as a brand mention under a documented matching rule. A link to geo.freeways.agency would count as an owned-source citation under a documented origin rule. Neither observation should be inferred from the other.

A recommendation is another classification. The answer might cite Freeways as the source of a definition without recommending its service. It might mention the agency as an option, criticize it, or discuss it in a context irrelevant to the buyer. Those are different outcomes.

## Set the Buyer Task Before Writing Prompts

Define the market, audience, offering, and decision. For example, an illustrative US B2B marketing team may want to understand GEO, compare implementation approaches, and decide whether to hire an agency. Those tasks produce different question groups.

A branded question such as “What services does Freeways Agency offer?” checks whether known business information is represented accurately. An unbranded question such as “What should a B2B team evaluate when choosing a GEO agency?” examines an evaluation context without supplying the desired answer.

Keep these groups separate. A report dominated by prompts that already name the brand cannot establish strong unbranded discovery. Also avoid writing questions that smuggle in claims such as “Why is Freeways the best agency?” Such a question gives the evaluator a premise rather than testing it.

## Record the Protocol

Use a stable prompt identifier and a written version. Record the platform, model where available, whether a search mode was used, account state, language, region settings, timestamp, and run identifier. Capture unknown conditions as unknown rather than inventing them.

Save the full answer and the cited URLs. If a source cannot be inspected, record that limitation. If the platform fails or the run is unavailable, retain a status row. An unavailable answer is not a valid observed answer with zero mentions.

Repeat observations according to a declared schedule. Do not keep only the most favorable answer. If a prompt or model condition changes, version the protocol and explain how that affects comparisons.

## Calculate Rates With Visible Denominators

Mention rate is valid observed answers with an explicit target-brand mention divided by valid observed answers in the same group and reporting period. Citation rate is valid observed answers with a qualifying target-origin source link divided by that same defined population.

An invented arithmetic example shows the difference. Suppose a hypothetical set contains ten valid answers. Four name the brand, two cite its website, and one does both. 

Mention rate is 4/10 and owned-source citation rate is 2/10. These are sample calculations, not Freeways performance figures or a target commitment.

Answers can contain multiple citations. If the unit is an answer, count each qualifying answer once for the rate. A separate source-frequency report can count individual links, but it needs its own name and denominator. Avoid switching between those units without telling the reader.

## Assess Accuracy and Context

Compare factual descriptions with the company's approved service and pricing information. Check claims about audience, capabilities, location, fees, and limitations. Record the erroneous statement, the correct fact, and the source supporting the correction.

Classify the brand's role in the answer. A citation to an educational resource is different from a supplier shortlist. A shortlist inclusion is different from an explicit recommendation.

An explicit recommendation may still contain incorrect facts. Frequency alone cannot describe answer quality.

For a business team, useful follow-up might involve correcting an ambiguous service page, making scope clearer, or publishing a genuinely missing answer. The [measurement protocol](/methodology) connects each observation with a concrete follow-up and keeps source attribution separate from recommendations.

## Use a Downloadable Observation Log

The [AI visibility observation template](/resources/ai-visibility-observation-template.csv) provides column headings for a measurement record. It contains no measured result. Store observations, source URLs, protocol versions, and review notes in your own controlled workspace.

Suggested fields include prompt group, platform, observed answer status, brand mention, recommendation context, owned citation, cited URL, and accuracy finding. Keep access to full transcripts appropriately controlled if they contain client information or account details.

A spreadsheet is sufficient for an initial protocol. Automation can help collect and organize observations, but a tool should not silently infer missing conditions, convert failures to negative results, or describe a snippet as a full answer.

## Interpret the Findings Without Claiming Causation

A before-and-after difference can reflect content changes, model changes, retrieval differences, account conditions, or variation between runs. Without a controlled design, describe the observation and its limitations. Do not claim that adding a particular file caused a citation merely because the events happened in that order.

Crawler requests measure access activity. They do not show whether a page was cited in an answer. Referral traffic measures visits observable in analytics.

It does not capture every citation or every buyer influenced by an answer. Conversion records add business context but also have attribution limits.

The [measurement method](/methodology) provides a compact protocol to use with this template. Agree on it before execution so the report cannot change its definition of success after seeing the results.

## FAQ

### Can a Citation Occur Without a Brand Mention?

Yes under this protocol. A linked page can support an answer without the answer naming its publisher. Record the source and the brand wording independently.

### Do Branded Prompts Prove Discovery?

They test a different condition: the question already supplies the brand. Report branded accuracy and unbranded discovery separately.

### Should Failed Runs Count as Zero Visibility?

Not as valid observed answers. Record them separately and disclose unavailable-run counts so the dataset's completeness is visible.

### Does This Page Report Freeways Results?

No. It describes a proposed protocol and an explicitly hypothetical calculation. Published results require actual observations, conditions, and review.

## Related

- [Knowledge hub](/guides/)
- [GEO for business teams](/guides/geo-for-business)
- [AI-search readiness checklist](/guides/ai-search-readiness-checklist)
- [Evaluate provider evidence](/guides/evaluate-geo-agency-evidence)

## Sources and Scope

The owner-supplied Freeways proposal identifies AI visibility, citations, accuracy, prompt coverage, organic signals, and business indicators as measurement areas. The definitions, template, and arithmetic scenario on this page operationalize that scope for editorial review; they are not a validated benchmark or client-result dataset.
