ChatGPT crossed 900 million weekly active users in February 2026 (Darkroom Agency, 2026). A December 2025 survey of 1,030 U.S. shoppers found half of them had already bought something after researching it with AI first (Darkroom Agency, 2026; Semrush, 2026). That's the scale AI has reached, and how much it's already shaping what people buy.
The rules just shifted under all of them, without warning. Reddit's share of ChatGPT's citations fell 86 percent in a matter of weeks (Promptwatch, 2026). Not because Reddit got less useful. Because ChatGPT changed how it searches, and nobody outside the platform saw it coming.
That's the pattern underneath everything else this quarter: the tactics people are being sold as universal levers keep turning out to be platform-specific, unproven, or already dead. Here's what actually held up when we checked, and what didn't.
What's Actually Confirmed This Quarter
Watch for case studies without a control group.[^1] Ask any vendor's AI-visibility case study whether they used one. In the one rigorously controlled study we found, the real isolated effect was a fraction of the headline number the raw numbers implied.
A popular technical fix does nothing.[^2] llms.txt, a file vendors sell as an AI-visibility lever, doesn't move citation rates, and multiple studies found AI crawlers barely even request it in the first place. Don't spend engineering time on it.
Standard SEO advice may be working against you.[^3] Adding structured data to your site (schema markup), standard SEO advice for years, may actually be hurting your AI visibility right now instead of helping it. The evidence points the wrong way.
This isn't a one-off.[^4] AI platforms are changing how they decide what to cite quickly and differently from each other, and nobody outside them sees it coming. It's the same story as the Reddit collapse above, not the only example of it.
Your AI-visibility number depends on who's asking.[^5] Researchers found a made-up test brand got recommended by AI far more often once the asking account had personal context on the user, compared to a blank account with none. Every measurement anyone publishes right now, including ours below, is a logged-out baseline. That's real, but it's not the whole picture, and no one's is right now.
For the studies, sample sizes, and exact figures behind the five findings above, see the full research and methodology near the end of this report.
The Wedge Most AI-Visibility Tools Don't Show You
Most tools that measure AI visibility report presence: was your brand mentioned, how often, how high up in the list. That's a real, useful thing to know, and it used to be the whole question, because in classic search the human still made the choice. They saw the list, and they picked.
That's not what's happening anymore. When a buyer asks an AI which vendor to choose, the model doesn't hand back a list for the human to sort through the way Google's ten blue links did. It picks one, or a short list, and states it as an answer. Presence has been demoted from the result to a precondition. A brand can look healthy on a presence metric, mentioned often, cited often, and still be losing every actual recommendation to a competitor, because presence and being the pick are two different outcomes, and most tools only report the first one. The tables below are a presence read, by design. The methodology note attached to them says exactly what they do and don't answer, and why that distinction matters for how you use them.
"The AI doesn't hand your customer a list anymore. It picks one. Showing up is not the same as being picked, and most tools still only measure showing up."
Matt Kott, Founder and CEO, Modern AI
The Category Itself Is Commoditizing. Credibility Is What's Left to Compete On.
The AI-visibility tooling market raised hundreds of millions of dollars and is already consolidating, but 45 percent of marketing leaders still can't accurately measure their own AI visibility.[^6] The market's problem isn't more tools. It's trust: buyers can't tell if any given vendor's number is measuring something real. That's your problem too, not just a vendor's: before you act on any AI-visibility score, including ours, ask exactly how it was measured. If a vendor can't walk you through that, don't act on the number.
The Clock Everyone in This Category Should Be Watching
Starting September 15, 2026, Cloudflare will default to blocking "mixed-use" crawlers (TechCrunch, 2026; effective date per The AI Insider, 2026, reporting Cloudflare's July 1, 2026 policy announcement). Those are bots that blend search, agent, and AI-training functions. The default block hits ad-carrying pages for new customers, new sites of existing customers, and every existing free-tier customer. A site can have a perfect robots.txt file and still get blocked at this layer: this works below the level your site's robots.txt file can control, so passing that check doesn't mean you're safe from it (Cloudflare, 2026). If your visibility checklist stops at robots.txt, it's already missing a real, dated access risk that's about to go live for a meaningful share of the web.
What This Means Going Into the Deep Dive
None of the above is abstract. It's the backdrop the measurement below sits inside: retrieval that moves fast and differs by platform, tactics that keep failing to hold up under real scrutiny, a presence-vs-recommendation gap most tools don't report, and a market getting more crowded and less trustworthy at the same time. The question that actually matters for an operator is not "is AI search changing." It obviously is. The question is what that looks like inside one real market, with real brands, measured the same way, side by side.
Deep Dive: B2B SaaS (110 Brands Measured)
Notion shows up in AI shortlists at a rate 41.5 points higher than its Google organic presence would predict. Smartsheet shows up 21.2 points lower. Same market, same buyer questions, wildly different verdicts depending on which search you trust.
We measured this across the same 50 buyer questions, run through four AI models and Google search, for 110 B2B SaaS brands across two categories.
This is the second report in our State of AI Search series. The first measured DTC consumer brands and found Google equity stopped transferring to AI answers. This one asks the same question of B2B SaaS: is the gap between "ranks well on Google" and "gets recommended by AI" a temporary blip, or a structural shift every operator in this category needs to plan around.
It's structural. Here's the measurement, and here's what changed since our last report that makes the story sharper than it was in Q2.
Your customers are starting to ask ChatGPT which vendor to buy, not Google. Some B2B SaaS brands show up in that answer. Many don't, even ones that rank well on Google. Ranking on Google and getting recommended by AI are now two separate contests, and a brand can win one and lose the other without knowing it. The table below is the list of brands who are, right now, losing a contest they didn't know they were in.
The Measurement, Plainly
We asked 50 real buyer-style questions ("best project management tool for a remote team," "alternatives to Salesforce for a 20-person sales team," and similar) to four AI models: ChatGPT, Claude, Gemini, and Perplexity. Separately, we ran the same 50 questions through Google search and scraped the resulting listicles and comparison pages.
Every brand got two scores, 0 to 100: an AI Shortlist Score for how often it showed up in those AI answers, and a Google SERP Score for how often it showed up in Google's results for the same questions. The delta between them is the story: a positive delta means AI recommends a brand more than its Google ranking predicts, a negative delta means AI is quietly passing over a brand with real Google presence.
The AI Shortlist Score in this report is one measure of AI visibility, not the only one. It tells you how often a brand showed up in an AI answer. It doesn't tell you whether the AI told the buyer to pick that brand first. Those are different questions with different answers. If you want the deeper read on your own brand, whether AI actually recommends you first as the pick, not just mentions you, get your free AI Visibility Snapshot and see how your brand compares on both.
For a CEO or CMO, the point is not the score itself. It is whether AI is starting to move buyer consideration in a direction your current search reporting cannot see.
Table A: The AI Winners
Twenty-two of 110 measured brands clear the floor with a positive delta. The top of the list:
| Brand | AI Score | Google Score | Delta |
|---|---|---|---|
| Notion | 88.9 | 47.4 | +41.5 |
| Asana | 100.0 | 63.2 | +36.8 |
| ClickUp | 71.1 | 36.8 | +34.3 |
| Rippling | 44.4 | 21.1 | +23.3 |
| Confluence | 35.6 | 15.8 | +19.8 |
Notion's case is the one worth sitting with. A 41.5-point gap means AI models are recommending Notion in contexts where its Google organic footprint gives no signal that should happen. Something about how Notion is discussed, cited, or referenced across the sources these models actually pull from (community platforms, comparison content, technical documentation) is doing work that classic SEO never measured.
If a 41.5-point gap is possible for a brand this well known, the obvious next question is where your own brand sits on the same axis, not on Google rank. You can get a free AI Visibility Snapshot for your brand in about 30 seconds, no signup, and see your own AI-versus-Google delta before you look at the laggard side of this measurement below.
Table B: The Quiet Laggards
Twenty-three brands hold real Google presence but get recommended by AI at a materially lower rate. The largest gaps:
| Brand | AI Score | Google Score | Delta |
|---|---|---|---|
| Smartsheet | 15.6 | 36.8 | -21.2 |
| Revenue.io | 4.4 | 23.8 | -19.4 |
| Trello | 33.3 | 52.6 | -19.3 |
| Brevo | 0.0 | 19.0 | -19.0 |
| Airtable | 13.3 | 31.6 | -18.3 |
Trello and Airtable are the two names on this list most operators will recognize immediately, and that's the point. Both rank well on Google. Both are being asked about, by real buyer-style questions, in this exact measurement. Neither is showing up in AI answers at anywhere near the rate their Google presence implies. That's not a visibility problem you'll see in a rank tracker. It's only visible when you measure both surfaces side by side, which is exactly why most SaaS marketing teams don't know it's happening to them yet.
What Changed Since Q2, and Why the Story Is Sharper Now
The old assumption was that strong Google presence would transfer into AI answers. The Q3 data makes that assumption harder to defend.[^7]
What's Actually Earning the Placements in Table A
Winners get there through mentions across the web, not Google rank.[^8]
Read the Full Research
Everything below backs the plain-language findings above: full figures, sample sizes, study design, and citations. Kept out of the main body so a reader moving through this report isn't slowed down by them, kept here in full so anyone who wants to check the work can find it.
[^1] Case studies claiming a big AI-visibility win are overstating it, and even the properly isolated effect is suggestive, not settled. A first-party log study on a single high-traffic domain (hundreds of thousands of YouTube Q&A pages that received a defined bundle of AI-visibility interventions in January 2026) found total ChatGPT referrals to the whole domain grew 5.7 times over the study window. That number is almost entirely platform tailwind: untreated pages on the same domain, used as an on-site control that absorbs the same tailwind, grew 3.5 times over the same period with no intervention applied to them at all. An interrupted time-series model on the weekly treated-to-control ratio, not a simple subtraction of those two multiples, put the actual isolated effect of the intervention at 1.82 times (95 percent CI 1.31 to 2.54). The authors' own conservative placebo-in-time test on that estimate returned p = 0.16, suggestive but not conclusive, given a short and noisy pre-period (Watanabe and Nakayashiki, 2026). If a case study claiming a dramatic AI-visibility lift doesn't disclose a control group, the number in it is not measuring what it claims to measure, and even a study that does isolate the effect properly is currently landing on "probably real, not yet proven," not a clean confirmed multiple.
[^2] llms.txt does not help. The file some vendors still recommend publishing "to help AI crawlers find your content" has now failed to hold up across four independent analyses: two measure whether publishing the file actually changes citation rates, and two measure whether the retrieval agents it's written for even request it in the first place. Neither angle finds anything. A roughly 300,000-domain study found no relationship between publishing the file and being cited (SE Ranking, reported via Inite.ai, 2026; SE Ranking's own primary report was not independently retrieved this cycle). A separate 37,894-domain citation-rate comparison found a null result (Trakkr Research, 2026, Mann-Whitney U test, p = 0.85). A 90-day, 62,100-visit AI-bot server-log study found only 84 requests, 0.1 percent of all AI bot traffic, ever touched the file at all (OtterlyAI, 2026). And first-party server-log data from a live Adobe Experience Manager estate breaks down who's actually requesting the file: the single largest named crawler hitting it is Google's own Googlebot, not an AI retrieval agent, and 92.2 percent of all traffic to the file overall is SEO tooling, monitoring services, and AI-readiness auditors inspecting it rather than the retrieval agents it's supposedly written for; agents verifiably identifiable as large language models account for just 1.1 percent of requests (Longato, 2026). If a vendor pitches llms.txt as a growth lever, ask for the study. There isn't one that shows it works.
[^3] Schema markup shows a real, controlled, negative effect, and it's the opposite of what most guidance says. A matched difference-in-differences study comparing 1,885 pages that added schema against 4,000 similar pages that didn't found a statistically significant decrease in Google AI Overviews citations for the pages that added it (down 4.6 percent) (Ahrefs, Linehan and Guan, 2026). The honest read is narrower than "schema hurts": this specific study only covers pages that were already being cited a lot, so it says nothing about whether schema helps a brand that isn't showing up yet. But "add schema, get cited more" is not a supportable claim on the evidence currently available, and multiple sources pointed the same direction this quarter, against a single, older, secondhand-sourced study claiming a positive effect.
[^4] Retrieval is now platform-specific, and it can move fast, without warning. Reddit's share of ChatGPT's own citations fell from 3.83 percent to 0.52 percent, an 86 percent drop, in weeks, coinciding with ChatGPT sharply increasing how often it runs site-specific searches as part of generating an answer (Promptwatch, 2026). Google's AI Overviews and AI Mode showed no equivalent movement over the same window. Same content, same brands, opposite outcomes, because the two platforms changed different things. Separately, a preregistered study found that simply expanding the underlying retrieval index changed the semantic content of roughly 10 percent of answers, even though overall accuracy across the whole test set barely moved (Ning and Li, 2026). And a third study, tracing citation failures through multi-agent research systems, found that in one evaluated system 84.7 percent of final-answer errors originated in how the system stitched its answer together, not in what it retrieved (Hirsch et al., EMNLP 2026). A brand can be correctly found and still not make it into the answer. None of this is a story about content quality. It's a story about instability in the pipeline between "was this found" and "was this said."
[^5] Personalization has broken the idea of one clean AI-visibility read. Researchers seeded a completely made-up, nonexistent brand into a test Gmail and Google Photos account. That fictional brand got recommended in AI answers 35.7 percent of the time once the account had that context, against roughly 19 to 22 percent for a blank control with no seeding at all (iPullRank, 2026). Whatever a brand's real AI-visibility number is, it is not one number anymore. It depends on who's asking and what that person's own account already knows.
[^6] The category itself, in full: the AI-visibility monitoring category raised well over $300 million in the past year and change (market tracking via Surmado, 2026), and consolidation has already started: Sitecore acquired Scrunch AI in June 2026 (Sitecore, 2026). That usually means a category is heading toward a bundled, low-price feature rather than a standalone purchase. At the same time, Semrush's own published AI Visibility Index found 45 percent of marketing leaders say they cannot accurately measure their brand's visibility in AI answers (Semrush, 2026), and most vendor visibility scores in this category are not independently auditable.
[^7] What changed since Q2: Our first report used the 76 percent figure. As of February 2026, the correlation between Google's top-10 ranked pages and what AI Overviews actually cites has fallen to somewhere between 17 and 38 percent (Ahrefs, 2026, measuring 38 percent across 863,000 keywords; BrightEdge, 2026, measuring 17 percent), down from roughly 76 percent in July 2025. One dated fact worth knowing, separate from what caused the decline: Gemini 3 became the new default model for Google's own AI Overviews on January 27, 2026 (9to5Google, 2026); whether that upgrade specifically drove the citation-correlation drop above isn't independently documented.
Two other findings from our latest research cycle change how confidently anyone should read a chart like Table A or Table B, including ours: the case-study inflation finding [^1] above and the llms.txt null-result finding [^2] above both apply directly to how any AI-visibility chart, including ours, should be read.
[^8] What's actually earning the placements in Table A: none of the winners got there through a Google-ranking proxy. Community platforms (Reddit, forums, third-party comparison content) and consistency of mention across many independent sources are what these models are actually weighting. Google's own top-10 rank was, at best, a weak predictor. Asana and ClickUp both sit at strong AI scores with meaningfully lower Google scores; the AI-side signal is coming from somewhere Google's own algorithm doesn't reward the same way.
Methodology note: we applied a minimum-presence floor of 10.0 on the defining side of each table, so every row in Table A and Table B represents a brand with a measurable position on both surfaces, not statistical noise sorting to the extremes.
Closing
We measured 110 B2B SaaS brands across AI recommendation and Google search, and found the two surfaces disagreeing by double digits for 26 of them, about a quarter. That's not a rounding error. It's the practical shape of the SEO-to-AI-search shift for this specific category, not the DTC consumer market our first report covered.
If you run growth for a B2B SaaS company and you have not checked where your brand sits on the AI side, not your Google rank, you do not know whether AI is creating demand you can prove or quietly sending high-intent buyers to someone else.
Get your free AI Visibility Snapshot →
If you've measured your own AI-versus-Google gap and found something that surprised you, we'd like to hear about it.
Results referenced are based on observed patterns in available data and are not guarantees of performance. Individual outcomes vary based on data quality, implementation, and market conditions. Modern AI recommends independent validation of any metric before business decisions are made.
The brand rankings in this report reflect a single measurement taken during the Q3 2026 window described above: the same 50 buyer-style questions run against four AI models (ChatGPT, Claude, Gemini, and Perplexity) and against Google search, with each brand's AI Shortlist Score and Google SERP Score calculated from that question set and the delta calculated as AI Shortlist Score minus Google SERP Score. These rankings are not a claim about the quality, efficacy, safety, commercial performance, or market standing of any named brand. They report where that brand appeared, and did not appear, across the two channels measured, at this point in time. The measurement is reproducible: the same 50 questions can be run against the same AI models and against Google to check the underlying result for any named brand.
Findings referenced above from third-party researchers and organizations, including Trakkr, Ahrefs, iPullRank, Promptwatch, Cloudflare, Semrush, and Surmado, are attributed to each source by name and, where available, by date. Citing a source's published finding does not imply that the source endorses, sponsors, is affiliated with, or has reviewed Modern AI or this report. These findings reflect the state of research and platform behavior as of the dates cited. AI search and crawler-access behavior can change quickly, and a finding accurate on its publication date may not reflect current conditions at the time this report is read.
