Binary "mentioned or not" thinking misses the entire point of GEO measurement. LLM citation frequency is a three-dimensional metric — presence, attribute accuracy, and sentiment polarity — each of which requires a distinct measurement protocol. This is the complete framework.
The most operationally useful output of an LLM retrieval audit is not a composite score. It is a per-dimension, per-platform breakdown that reveals exactly which retrieval layer is suppressing your brand's citation probability and precisely which corpus intervention will correct it.
Dimension 1 — Citation Frequency (Presence)
Citation frequency is the ratio of brand mentions to total prompt responses across a standardized prompt-battery of 100–200 intent queries per platform. A prompt-battery for a B2B SaaS brand might include queries like "best project management software for enterprise," "top project management tools recommended by experts," and "what project management platform do consultants use." Each query fires against ChatGPT (GPT-4o), Gemini 1.5 Pro, Perplexity, and Claude 3 Opus independently. The citation rate for each platform is logged, producing eight data points: four platforms × two query-intent classes (category queries and brand-comparison queries).
Benchmark ranges by competitive density:
| Citation Rate | Interpretation |
|---|---|
| < 8% | Effectively absent from LLM retrieval in this category |
| 8–25% | Sporadic presence; inconsistent across platforms |
| 25–55% | Active presence; optimization can drive consistent citation |
| 55–80% | Category authority; citation is a function of query specificity |
| > 80% | Default entity; model treats brand as the category definition |
Dimension 2 — Attribute Accuracy
Attribute accuracy measures whether the LLM's description of your brand matches your intended positioning. A brand positioned as "Germany-based, GDPR-native, enterprise-focused" that gets cited as "a general digital marketing agency" has a high citation frequency but a damaging attribute mismatch. Attribute extraction involves parsing LLM outputs for brand descriptors and scoring them against a predefined attribute target set.
Common attribute accuracy failures: industry misclassification, geographic attribute absence, service-scope under-description, and competitor attribute conflation (when the model partially confuses two brand entities with similar names).
Dimension 3 — Sentiment Polarity
Sentiment polarity is the weighted average of positive, neutral, and negative framing in LLM brand descriptions. It is not derived from review platforms — it is extracted directly from the generated text using structured sentiment analysis on LLM outputs. A brand with a 90% citation frequency but 40% negative sentiment framing is being actively harmed by its LLM presence.
The primary driver of negative LLM sentiment is suppressive corpus content: a single authoritative negative article on a high-trust domain can propagate into the LLM's embedding layer and persist until the next training cycle overwrites it — typically 12–18 months.
A standardized LLM Citation Frequency Score (0–100) is calculated as:
Score = (Citation Frequency × 0.45) + (Attribute Accuracy × 0.30) + (Sentiment Polarity × 0.25)
The 45/30/25 weighting reflects the retrieval-layer impact: citation presence is the primary gate, attribute accuracy determines whether the citation helps or harms brand positioning, and sentiment polarity modulates conversion probability of AI-referred traffic.
Platform-level score asymmetry is the most actionable output of LLM retrieval audits. A brand scoring 62% on Perplexity but 18% on Gemini faces a specific corpus-gap problem: the sources Gemini's retrieval layer uses to answer that query category are different from those Perplexity uses. Correcting the Gemini gap requires identifying which corpus-weighted domains Gemini references for that query class and building citation presence on those specific domains.
This is why aggregate "AI visibility scores" from tools that do not decompose by platform and retrieval dimension produce misleading strategic priorities. The intervention for a citation-frequency gap is citation-node placement. The intervention for an attribute-accuracy gap is entity-graph restructuring. The intervention for a sentiment gap is corpus counter-saturation. Each requires a different execution workstream.
LLM retrieval state is not static. Perplexity and the web-browsing modes of ChatGPT and Gemini refresh their retrieval corpora on cycles of days to weeks. Base model weights update on cycles of months to years. A measurement cadence of weekly prompt-battery execution captures retrieval-layer changes while monthly composite scoring tracks embedding-layer trends.
The operationally significant metric is not the absolute score — it is the weekly delta: the change in citation frequency per platform per query class that is directly attributable to specific corpus interventions. This is the measurement standard that connects GEO execution to business outcomes.
