AI Shortlist Score: How We Measure It
The AI Shortlist Score measures how often a brand appears in AI-generated search results across ChatGPT, Perplexity, Gemini, and Claude. This page documents the query set construction, the scoring model, the model panel, and the measurement cadence used in the Q2 2026 State of AI Search Report.
Contents
Query Set Construction
Each brand in the Q2 2026 dataset was measured against a set of 25 queries. Queries are structured across five intent categories, five queries per category. The query set is constructed to reflect how consumers actually search for products in the Beauty and Health & Wellness categories.
Query sets are versioned and immutable. Once a measurement run begins, the query set does not change. This ensures comparability across brands and across measurement windows.
Five intent categories, five queries each
Awareness
Queries where a consumer is learning about a product category for the first time. Example: 'what is a greens powder?' These queries test whether a brand appears at the top of category conversations.
Comparison
Queries where a consumer is evaluating multiple options side by side. Example: 'AG1 vs Seed vs Ritual.' These queries are high-intent and produce competitive shortlists.
Recommendation
Queries where a consumer is asking for a direct product recommendation. Example: 'best protein powder for muscle recovery.' AI engines generate explicit shortlists for these queries.
Authority
Queries where a consumer is evaluating whether a brand is credible, trusted, or endorsed. Example: 'is Momentous NSF certified?' Brand reputation and third-party coverage drive performance here.
Use-case
Queries where a consumer specifies a specific need or constraint. Example: 'protein powder for lactose intolerance.' These queries test whether brand positioning matches consumer need language.
Scoring Model: 0-100 Normalization
The AI Shortlist Score is a normalized 0-100 value. A score of 100 does not mean perfect visibility. It means the highest relative AI presence in the brand's niche in the measured window.
Five steps from query to score
Query execution
Each query in the set is submitted to each model in the panel. Responses are collected independently per model. No response is shared or interpolated across models.
Presence detection
For each response, a presence signal is extracted: was the brand named in the response? Presence is binary per query per model.
Raw score calculation
The raw AI Shortlist Score is the fraction of queries where the brand was present, across the full model panel. A brand present in 20 of 25 queries scores 0.80 raw.
Within-niche normalization
Raw scores are normalized within each niche (Beauty, Health and Wellness) to produce a 0-100 score. The highest raw score in the niche becomes 100. All other brands are scaled proportionally. This removes cross-niche comparison distortion caused by different query set compositions.
AI Shortlist Score
The final AI Shortlist Score is the normalized 0-100 value. A score of 100 means the brand had the highest AI presence in its niche in the Q2 2026 measurement window. A score of 0 means no AI model mentioned the brand across any of the 25 measured queries.
Normalization formula
AI Shortlist Score = (brand raw score / niche max raw score) x 100
Niche max is the highest raw mention fraction observed for any brand in the same niche in the same measurement window. Each measurement window produces its own niche max. Scores are not comparable across windows unless re-normalized against the same max.
Model Panel
The Q2 2026 measurement used four AI models. Each model is queried independently. No response is shared across models. The composite score reflects performance across all four.
ChatGPT
OpenAIGeneral-purpose conversational AI with broad consumer adoption in the health and beauty categories. Query responses are sampled without real-time web retrieval unless browsing mode is active.
Perplexity
Perplexity AIAI search engine with active web retrieval. Perplexity responses cite sources inline. Brand visibility here reflects both training data and active web indexing.
Gemini
Google DeepMindGoogle's conversational AI. Gemini responses reflect Google's training corpus and, in some configurations, integration with Google Search. Included for cross-model consistency measurement.
Claude
AnthropicConversational AI with strong performance on reasoning and comparison queries. Claude responses reflect Anthropic's training data without real-time web retrieval in the standard configuration.
Measurement Cadence
The Q2 2026 report is a point-in-time measurement. All 110 brands were measured in May 2026 using the same query set, the same model panel, and the same presence detection logic.
Modern Discovery Core subscribers receive weekly measurement cadence against their brand and up to five competitors. Each weekly run uses the same query set for comparability. Trend data accumulates across runs, enabling week-over-week delta tracking.
The Q2 public leaderboard is a snapshot. It does not update between quarterly reports. The Q3 2026 report will be published in August 2026.
Design Choices
Why we use presence, not rank
AI-generated responses do not have a stable position ordering equivalent to Google's rank 1-10. Position in a conversational response varies by phrasing, model, and query context. Presence (was the brand named?) is a more reliable and reproducible signal than inferred rank.
Why we normalize within niche, not across the full dataset
Beauty and Health & Wellness have structurally different query behaviors. AI models mention health brands more frequently in recommendation queries than beauty brands. Cross-niche normalization would systematically disadvantage one category. Within-niche normalization preserves the diagnostic value of the score for each category.
Why we use four models, not one
No single AI model represents the full AI search market. ChatGPT, Perplexity, Gemini, and Claude each have distinct citation behaviors, training cutoffs, and audience characteristics. A brand's presence varies significantly across models. A composite score across all four is more stable and more representative than a single-model score.
What the score does not measure
The AI Shortlist Score does not measure sentiment, citation quality, brand position within a response, or share of voice versus specific competitors. It measures one thing: was the brand named in an AI-generated response for this query set? Sentiment analysis, competitive share, and citation attribution are available in Modern Discovery Core.
Changelog
| Date | Update |
|---|---|
| 2026-05-10 | Methodology page published for Q2 2026 State of AI Search Report. Query set: 25 queries per brand, 5 categories, 4 models. 110 brands measured. |
| 2026-07-01 | Scope clarification: the 25-query, 4-model spec on this page is the frozen Q2 2026 report dataset and is never retroactively modified. The Discovery product's current per-entity measurement is 20 queries across six AI surfaces (ChatGPT, Gemini, Claude, Perplexity, Google AI Mode, Google AI Overviews), producing 120 scored responses per entity. Current product methodology: modernai.io/methodology. |
Questions about the methodology? Contact us at info@modernai.io.
See where your brand stands
Enter your brand name for a free AI Shortlist Score from the Q2 2026 dataset. No account required.