The Answer Layer
Search did not die. It stopped being a list of places to go.
Answer engines split a question into sub-queries, retrieve passages from many sources and write one reply, so organic visibility has stopped being a rank and become a supply problem. A brand's claims must be present, retrievable, consistent, attributable and current wherever models read, and any effect must be proven by repeated sampling, not one run.
Overview
For twenty-five years organic visibility meant one thing: a position in an ordered list of ten links, on a page a person chose to read. That artefact is being replaced by a synthesised answer assembled from many retrievals, most of which never become a visit. This paper sets out what the measured evidence says about that change, what the platforms themselves document, what the academic literature has established and failed to establish, why the tooling category built around the problem is statistically underpowered, and what Nagent has built in DRIS as a response.
- The unit of retrieval: the sub-query, not the query.
- The unit of visibility: a distribution, not a rank.
- The unit of work: an agent team, not a retainer.
The thesis
Organic visibility stopped being a ranking problem and became a supply problem. Everything in this paper follows from that sentence, and every claim under it is sourced.
A search engine used to answer a question by handing back the addresses of documents that might contain the answer. The work of organic marketing was therefore to make one document the most plausible address for a given phrase. Rank, click, visit, convert. The chain had four links and each one could be measured.
An answer engine does something different in kind. It decomposes the question into sub-questions, runs several retrievals in parallel, selects passages from many documents, and writes a single reply in which the brand may appear as a named recommendation, an unnamed implication, a cited link, or nothing at all. Google documents this decomposition as query fan-out and states plainly that both AI Overviews and AI Mode use it to issue "multiple related searches across subtopics and data sources" before composing a response (Google Search Central, AI features and your website, updated 10 December 2025). The chain now has more links, fewer of them are visible, and the last one often does not exist.
Four things follow, and they are the spine of this paper.
First, the click is no longer the measure of the work. Pew Research Center tracked 68,879 real Google searches from 900 United States adults during March 2025 and found that users clicked a traditional result on 8 % of visits where an AI summary appeared against 15 % where none did, and clicked a link inside the summary itself on 1 % of visits (Pew Research Center, 22 July 2025). That is observed browsing behaviour, not a survey and not a keyword panel, which is why it is the most load-bearing number in this paper.
Second, ranking well no longer means being cited. Ahrefs ran 15,000 long-tail queries through Google and Bing and then through four assistants, and found that on average 12 % of the URLs cited by assistants also appeared in Google's top ten for the original query (Ahrefs, 11 August 2025). Inside AI Overviews specifically, the share of cited pages that also ranked in the top ten for the same query measured 38 % across 863,000 keywords and four million AI Overview URLs in March 2026 (Ahrefs via Search Engine Journal, March 2026). A brand can hold position one and remain absent from the answer that replaced position one.
Third, what is measurable is a distribution, not a position. Researchers at the University of St. Gallen ran repeated identical prompts across ChatGPT, Gemini, Google AI Mode and Perplexity and found day-to-day source overlap at Jaccard 0.34 to 0.42, meaning roughly 65 % of cited sources change between consecutive days, with same-day repeated runs showing comparable instability at 0.32 to 0.43 (Schulte, Bleeker and Kaufmann, arXiv:2604.07585, 10 April 2026). SparkToro reached the same conclusion from the other direction with 600 volunteers and 2,961 runs, reporting under a one in a hundred chance that an engine asked the same question repeatedly returns the same list of brands (SparkToro, 28 January 2026). Any product that reports a single AI rank from a single run is reporting noise with a decimal point on it.
Fourth, the remaining traffic is worth more than the traffic that left. Adobe Analytics, working across more than one trillion visits to United States retail sites, found that visitors arriving from AI sources converted 42 % better than non-AI traffic in March 2026, having converted substantially worse a year earlier (Adobe Digital Insights, Q2 2026 AI Traffic Report). The same report puts year-on-year growth in AI-sourced retail traffic at 393 % for the first quarter of 2026. Against that, Semrush measured AI systems at 0.14 % of total web visits across 50,000 sites and seventeen industries for the whole of 2025, with organic search still at 16.04 % (Semrush, traffic channel mix study, 2025). Both are true. The honest reading is that answer engines currently matter far more for influence than for traffic, and that the traffic they do send is qualified.
The work is no longer to rank a page. It is to make a set of claims about a company retrievable, consistent, attributable and fresh across every corpus a model draws on, and then to prove the effect statistically rather than assert it.
That work is too wide for a person and too repetitive for an agency. It spans crawler policy, entity consistency, document structure, third-party presence, review platforms, community surfaces, feed hygiene, licensing posture and a measurement protocol that needs hundreds of samples a week to produce a usable standard error. It is the kind of work that suits a team of AI coworkers with clear authority limits and a human who signs off on anything that changes a live property. That is what DRIS is, and the second half of this paper describes it without claiming more than has been built.
What the machine now does
The unit of retrieval is now the sub-query, and almost nothing downstream survives that change. Before arguing about strategy, it is worth being precise about what the systems actually do, using what their owners document rather than what the trade press infers.
Fan-out
Elizabeth Reid, who runs Google Search, described AI Mode at launch as working by "breaking down your question into subtopics and issuing a multitude of queries simultaneously", with a Deep Search variant that can "issue hundreds of searches, reason across disparate pieces of information, and create an expert-level fully-cited report" (Google, 20 May 2025). In November 2025 Google said the technique had received "a major upgrade" with its newest Gemini model, running more searches and understanding intent well enough to "find new content that it may have previously missed" (Google, 18 November 2025).
The commercial confirmation is in the developer documentation rather than the marketing. Grounding with Google Search in the Gemini API returns a google_search_call object listing the queries the model chose to run, and for the newest Gemini models billing is "per each search query that the model decides to execute" (Google, Gemini API grounding documentation). A platform bills per sub-query because sub-queries are what it issues.
For a brand this is not a technical footnote. It means the keyword a marketing team tracks is not the string the retrieval system runs. It means a page that covers a topic broadly may lose to a page that covers one sub-topic deeply, which is exactly the outcome Reid predicted when she said fan-out "can expose websites that go in more depth on part of a topic, instead of just a webpage that is surface level about the whole topic" (Financial Times interview, April 2025, reported by Search Engine Land). And it means forecasting models built on position and estimated click-through rate are estimating the wrong quantity.
Eligibility
Google's own position on optimisation is deflationary and worth quoting exactly, because a great deal of commercial advice contradicts it. "There are no additional requirements to appear in AI Overviews or AI Mode, nor other special optimizations necessary." A page must be "indexed and eligible to be shown in Google Search with a snippet". There is "no special schema.org structured data that you need to add" (Google Search Central, updated 10 December 2025).
The operational consequence is the reverse of what most sites assume. Because eligibility is snippet eligibility, the snippet controls are the switches: nosnippet, data-nosnippet, max-snippet and noindex. A publisher that set an aggressive max-snippet value to protect content from summarisation has, as a side effect, removed itself from the answer surfaces. Meanwhile Google-Extended, which many teams believe governs AI Overviews, does not: Google states that it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal", and that it covers training and grounding in Gemini Apps and Vertex AI (Google, crawler documentation). Blocking it does not remove a site from AI Overviews. Very few site owners have this the right way round.
Three kinds of robot
The same asymmetry runs through every platform, and it is the single most expensive thing a large organisation currently gets wrong. There are retrieval crawlers, training crawlers and user-initiated fetchers, and they have different consequences.
OpenAI documents OAI-SearchBot as the crawler used "to surface websites in search results in ChatGPT's search features", not used for training, and states that "sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers". GPTBot is the training crawler. ChatGPT-User is user-initiated, and OpenAI notes that "because these actions are initiated by a user, robots.txt rules may not apply" (OpenAI, bot documentation). Perplexity splits the same way, with PerplexityBot for surfacing and linking and Perplexity-User which "generally ignores robots.txt rules" (Perplexity, bot documentation). Meta now runs meta-webindexer for retrieval, distinct from meta-externalagent for training, and says allowing the former "helps us cite and link to your content in Meta AI's responses" (Meta, webmaster crawler documentation).
A legal or security team that files a blanket instruction to block AI bots therefore makes an invisible trade: it retains a modest amount of training corpus control and gives away the retrieval visibility that produces citations, referrals and recommendations. In the audits Nagent has run, this is the most common single cause of a brand being structurally absent from one engine while performing normally in another, and it is usually nobody's fault because no one function owns the file.
Where citations come from
Retrieval supply is concentrated and platform-specific. Similarweb analysed roughly 600,000 United States citation events in January and February 2026 and found ChatGPT drawing 13.15 % of citations from Wikipedia and 11.97 % from Reddit, while Google AI Mode led with Fandom at 7.16 % (Similarweb AI Citation Analysis, 2026). Peec AI, analysing 30 million sources across five surfaces, found Reddit, YouTube, LinkedIn, Wikipedia and Forbes as the overall top five, with Perplexity leaning on Reddit, LinkedIn and G2 for business software questions (Peec AI, 31 March 2026).
Two implications matter. A brand's own site is necessary and insufficient: a study of 102,025 responses across five engines found corporate websites supplying 78 % of citations while "best of" listicles were the single most cited content format at around 21 % (Kumar, arXiv:2606.20065, June 2026, commercial affiliation disclosed). And the supply is unstable by design: Semrush tracked Reddit's share of ChatGPT responses falling from roughly 60 % to roughly 10 % after an undisclosed September 2025 model change, on ChatGPT alone (Semrush, 230,000 prompts, November 2025). A visibility strategy anchored to one third-party platform is a strategy anchored to another company's release notes.
The evidence on clicks
Every independent dataset points the same way, and the disagreements are about size, not direction. What follows is the measured record on clicks, assembled with methods attached, because in this category the method is the argument.
The best evidence is behavioural. Pew Research Center recruited 900 United States adults to a tracked browsing panel and observed 68,879 Google searches during March 2025, of which 12,593 produced an AI summary. Clicks on a traditional result ran at 8 % with a summary present and 15 % without. Clicks on a source link inside the summary ran at 1 %. Browsing sessions ended on the search page 26 % of the time with a summary present against 16 % without (Pew Research Center, 22 July 2025).
Keyword-panel studies find the same direction at different magnitudes. Ahrefs matched 150,000 keywords with AI Overviews against 150,000 without and reported a 34.5 % reduction in click-through rate for the top-ranking page when an AI Overview is present (Ahrefs, 17 April 2025). Amsive, across 700,000 keywords in finance, education, software, healthcare and pets, found an overall decline of 15.49 % on AI Overview keywords, with non-branded terms down 19.98 % and branded terms up 18.68 % (Amsive, 16 April 2025). SISTRIX, on more than 100 million German keywords, measured position-one click-through falling from 27 % to 11 % when an AI Overview appears and estimated 265 million lost organic clicks a month in that market alone (SISTRIX, February 2026 data).
The branded uplift in the Amsive numbers is the detail most summaries drop, and it is the commercially important one. Where a person already knows the company by name, the answer surface appears to help. Where they do not, it substitutes. Everything in the second half of this paper about entity strength rests on that asymmetry.
The number worth reading carefully
The most useful 2026 dataset is Seer Interactive's, because it is longitudinal and because it complicates the story. Across 53 brands, 5.47 million tracked queries and 2.43 billion organic impressions from January 2025 to February 2026, organic click-through where no AI Overview appeared rose from 2.75 % to 3.82 %, while click-through where one appeared fell from 3.19 % to 2.36 %. Where the brand was not cited in the overview, organic click-through declined 67 % over 2025. Being cited delivered 120 % more organic clicks per impression than not being cited, and still underperformed a clean result page by 38 % on informational queries (Seer Interactive, 24 April 2026).
Two cautions belong with that. The widely circulated "61 % drop" from the same research describes a period in which impressions doubled and absolute clicks rose, so the fall is substantially a denominator effect (Search Engine Journal analysis, 26 April 2026). And the series shows a rebound, from a floor of 1.3 % in December 2025 to 2.4 % in February 2026, which is the strongest published evidence that suppression is not monotonic.
Zero click
SparkToro, using Similarweb clickstream data, put United States Google searches ending without a click at 68.01 % for January to April 2026 against 60.45 % in 2024 (SparkToro, 9 June 2026). The same study contains the corrective that most coverage omitted: AI Mode accounted for 0.34 % of search sessions in that window. A surface with a very large user base can still be a rounding error in session share, and the practical consequence is that most answer-engine exposure today still happens inside a classic results page rather than in a conversation.
It is also worth recording that zero-click measurement has been contested since long before AI answers existed, on the grounds that panels miss click-to-call, in-app and map interactions (Search Engine Journal, 29 April 2021). SparkToro itself warns that cross-year comparisons use different panel providers and should be read cautiously. The direction survives the caveats; the precision does not.
Publishers as the leading indicator
Chartbeat data cited by the Reuters Institute put global organic search traffic to more than 2,500 sites down 33 % between November 2024 and November 2025, with the United States figure at 38 %; in the same report, news publishers expected search referrals to fall by more than 40 % over three years, and one fifth expected losses above 75 % (Reuters Institute, 12 January 2026, survey of 280 leaders across 51 countries). Digital Content Next, using first-party analytics from nineteen member publishers, found a median referral decline of 10 % year on year, news brands at 7 % and non-news at 14 % (Digital Content Next via Digiday, August 2025).
The scale asymmetry matters more than either figure. Similarweb measured search referrals to publishers falling by roughly 800 million visits year on year while total AI referral traffic to the same set ran at about 36 million visits a month (Similarweb via Digiday, 2025). Nothing about the new channel replaces the old one by volume. Anyone selling AI visibility as a traffic replacement is selling arithmetic that does not work.
| Finding | Measured value | Method | Source, date |
|---|---|---|---|
| Click on a result, AI summary present | 8 % vs 15 % | 68,879 tracked searches, 900 adults | Pew, Jul 2025 |
| Click on a link inside the summary | 1 % | Same panel | Pew, Jul 2025 |
| Top-position click-through loss | 34.5 % | 300,000 matched keywords | Ahrefs, Apr 2025 |
| Position-one click-through, Germany | 27 % to 11 % | 100m+ keywords, observed | SISTRIX, Feb 2026 |
| Zero-click share, United States | 68.01 % | Clickstream panel, Jan to Apr 2026 | SparkToro, Jun 2026 |
| AI Mode share of search sessions | 0.34 % | Same panel | SparkToro, Jun 2026 |
| Organic search traffic to 2,500 sites | down 33 % | Chartbeat, Nov 2024 to Nov 2025 | Reuters Institute, Jan 2026 |
| AI systems as a share of all web visits | 0.14 % | 50,000 sites, 17 industries, FY2025 | Semrush, 2025 |
| AI-referred retail conversion advantage | 42 % | 1tn+ visits, US retail | Adobe, Mar 2026 |
Two search economies
Search as a business is growing while search as a distribution system is contracting. Both statements are supported by audited data, and conflating them produces most of the bad strategy in this category.
Alphabet reported Search and other revenue growing 17 % year on year in the second quarter of 2026, and Sundar Pichai told investors that AI features are "driving an incremental increase in Search queries overall" (Alphabet, Q2 2026). Comscore measured 76 billion United States desktop searches in the first quarter of 2026, up 10 % against the same quarter of 2024 (Comscore, 2 June 2026). Microsoft said Bing passed a billion monthly active users for the first time in April 2026 (Microsoft Q3 FY2026 earnings, reported 30 April 2026).
At the same time the referral layer is contracting on every independent measure quoted in the previous section. There is no contradiction. A search engine monetises the query; a website monetises the visit. When the engine answers the query itself, the first business improves and the second loses its supply.
The asset being disintermediated is not the search engine. It is the click that used to fall out of it, which is the only part the brand ever owned.
This distinction has a direct consequence for how a board should read the category. A chief financial officer who is told that "search is dying" can check Alphabet's filings in four minutes and dismiss the entire agenda. A chief marketing officer who says instead that the company's owned share of a growing query market is falling, and can show the citation series to prove it, is making a claim that survives scrutiny.
The prediction that should be retired
Gartner stated in February 2024 that "by 2026, traditional search engine volume will drop 25 %, with search marketing losing market share to AI chatbots and other virtual agents" (Gartner, 19 February 2024). No methodology was published with it. The prediction concerned volume rather than referral traffic, and it is routinely requoted as though it concerned traffic. The window has now closed and the measured direction is the opposite: Comscore's 76 billion desktop searches, up 10 % on two years earlier.
Gartner's own later research is closer to the evidence. In January 2026 it reported that only 33 % of consumers say generative assistants match search engines for learning, that 31 % spend more time searching when AI summaries appear against 16 % spending less, and that 31 % consider more product options because of AI overviews against 7 % considering fewer, with the analyst conclusion that marketers "cannot afford to think of AI as a replacement for traditional search" (Gartner, 20 January 2026, two surveys, n=377 and n=365).
Nagent's position is that the 25 % forecast should be cited only as a case study in how this category over-claims. A paper that leans on it loses the analyst-literate reader in the first ten minutes, and that reader is usually the one who controls the budget.
What the platforms say
The platforms publish less guidance than the market assumes, and what they publish contradicts common practice. Here is what each surface actually documents, and where the documentation runs out.
A discipline is usually defined by the platform it serves. Classic search optimisation had twenty years of Google documentation, a webmaster guidelines page, a public spam policy and a named liaison. The answer layer has almost none of that, and the gap has been filled by vendors. It is worth setting out what is actually on the record.
Google's position is that there is no separate discipline. Danny Sullivan, its search liaison, put it as "the acronyms keep changing, but the advice doesn't", adding that structured data "is not structured data and you win AI, it simply supports how systems understand and present content" (Search Off the Record, reported 17 December 2025). In a later episode he addressed the most popular piece of practitioner advice directly: "turn your content into bite-sized chunks, because LLMs like things that are really bite size. So we don't want you to do that" (Search Off the Record, published 8 January 2026).
On traffic, Google's defence is qualitative. Liz Reid wrote that "total organic click volume from Google Search to websites has been relatively stable year-over-year" and that average click quality had increased (Google, 6 August 2025). No series, percentage or methodology accompanied the statement, and every independent dataset in the evidence on clicks points the other way. A strategy paper should quote the claim and note the absence of numbers rather than pretend it was not made.
OpenAI
OpenAI publishes no ranking documentation for general web content. Its guidance amounts to keeping OAI-SearchBot unblocked, using noindex to exclude, and the note that referral URLs carry utm_source=chatgpt.com (OpenAI, publishers and developers FAQ). The exception is shopping, where OpenAI does publish criteria: product results are "selected independently" and not influenced by partnerships, inputs include "structured metadata from first-party and third-party providers", and "merchants are ranked based on factors like availability, price, quality, and whether they are the maker or primary seller of that item" (OpenAI, shopping documentation).
That is the only explicit ranking statement any answer engine has published, and it is feed-based rather than page-based. For a retail or consumer-goods brand it means the product data pipeline, not the blog, is the primary visibility asset.
Anthropic, Microsoft, Perplexity, Meta, Amazon
Anthropic's web search tool makes citation structural rather than optional: citations are always enabled, each returns a URL, title and up to 150 characters of cited text, and the documentation states that "when displaying API outputs directly to end users, citations must be included to the original source". Developers may constrain retrieval with an allowed or blocked domain list (Anthropic, web search tool documentation). The allowed-domain facility is significant and under-discussed: inside an enterprise deployment, the citable universe can be fixed by configuration, which is a visibility channel no amount of public-web optimisation can reach.
Microsoft is, at present, the only platform offering first-party citation reporting. Bing Webmaster Tools added an AI Performance report in February 2026 with total citations, average cited pages and the grounding queries the system used (Microsoft, 10 February 2026), then extended it in June 2026 with citation share, intents, topics and competitive comparison (Microsoft, 16 June 2026). Any measurement programme that ignores it is discarding the only free ground truth in the category.
Perplexity's economics are public in a different way: a 42.5 million dollar pool and an 80 % publisher revenue share for its Comet Plus subscription (Axios, 26 August 2025). Amazon's shopping assistant draws on "Amazon's extensive product catalog, customer reviews, community Q&As, and information from across the web" plus individual shopping history (Amazon, 18 November 2025), which makes listing copy, attributes, questions and reviews the controllable surface rather than anything on the brand's own domain.
Two things the market believes that the evidence does not support
The first is llms.txt. John Mueller of Google stated flatly that "no AI system currently uses llms.txt" (Search Engine Roundtable, June 2025). Ahrefs then measured it: across server logs from 137,210 domains, 28 % had a file returning 200, and only about 1,100 of those files received any traffic at all, with zero requests for files that did not exist (Ahrefs, 15 June 2026). The file is close to free to publish and close to useless to rely on, and a vendor that bills for implementing it is billing for a gesture.
The second is schema as a citation lever. Ahrefs tracked 1,885 pages that added JSON-LD between August 2025 and March 2026 against 4,000 matched controls using difference-in-differences, and found AI Overview citations down 4.6 %, AI Mode up 2.4 % and ChatGPT up 2.2 %, with the latter two indistinguishable from zero (Ahrefs, 11 May 2026). Their own caveat is important and is the reason DRIS still audits structured data: the sample was pages already heavily cited, so schema may still matter for pages that are not being seen at all. Structured data belongs in the retrievability workstream, not in the citation-uplift promise.
How Nagent treats platform guidance. Documented statements from a platform are treated as constraints. Third-party measurement is treated as evidence. Practitioner assertion is treated as a hypothesis to be tested on the tenant's own property with a holdout. DRIS proposals carry which of the three they rest on, in the proposal itself, so that a marketing lead approving an action can see how much is actually known.
What the research says
The academic work says visibility is manipulable, congested, and only weakly connected to revenue. Forty-five studies have now been reviewed in a single survey, and the findings are less flattering to the category than the category admits.
The founding paper is Aggarwal and colleagues at IIT Delhi and Princeton, published at KDD 2024, which introduced generative engine optimisation as a formal problem and the GEO-bench benchmark of 10,000 queries across 25 domains (arXiv:2311.09735). Its lasting contribution is not the headline that visibility can be raised "by up to 40 %" but the metric design. Visibility is defined as a share of the answer's words attributable to a source, with a position-adjusted variant that decays exponentially with sentence position. Being cited late in a long answer is worth measurably less than being cited early, which is a property no rank-based metric can express.
The measured effects per tactic are the part worth carrying into practice. Against a 19.3 % baseline impression, adding quotations produced a 40.9 % relative improvement, adding statistics 32.8 %, citing sources 27.5 %, improving fluency 28.0 %, while keyword stuffing produced a negative effect of 8.3 %. The pattern is consistent and, for once, intuitive: the tactics that add verifiable evidence work, the tactics that add packaging do not.
A much larger factorial design confirmed it. A study of 252,000 paired comparisons across six models and eighteen content factors, with brand anonymisation and counterbalanced ordering to control position bias, found odds ratios above 10,000 for topical relevance, similarly extreme values for list position and for an explicit price, 14.4 and above for a recent timestamp, and odds ratios near 1.0 for formatting-only edits (Vishwakarma and colleagues, arXiv:2605.25517, May 2026, preprint). Confident rather than hedged language mattered, most strongly on Gemini and Claude. Presentation, on its own, did not.
The three findings that should temper any vendor claim
Optimising for citation can cost retrieval. Yonsei University's SAGEO Arena evaluated the full pipeline of retrieval, reranking and generation over 171,003 documents and 2,700 queries, and found that body-text rewrites in the usual generative-optimisation style produced an average rank drop of 4.54 positions and a 9 % fall in hit rate at retrieval, while citation gains at generation were marginal. Structural and metadata work, by contrast, raised hit rate by 22 % (arXiv:2602.12187, February 2026, preprint). Structure gets a document retrieved; body evidence gets it cited; doing only the second can remove it from the candidate set entirely.
The game is congested and close to zero-sum. C-SEO Bench, accepted at the NeurIPS 2025 datasets and benchmarks track, found that most conversational-optimisation methods "are not only largely ineffective but also frequently have a negative impact on document ranking", that traditional optimisation outperformed the dedicated techniques, and that "as we increase the number of C-SEO adopters, the overall gains decrease" (arXiv:2506.11097). A separate study of recommendation dynamics found individual payoff collapsing from 0.802 to 0.007 when many brands optimise simultaneously, while brands that did not participate received nothing at all (arXiv:2606.17443, June 2026, preprint). Both halves matter: participation confers no durable advantage and abstention is fatal.
Citations are not yet demonstrated to produce revenue. A critical survey of 45 studies from Sciences Po grades the field's confidence explicitly: high confidence that documents already in the retrieval context can causally shift citation, low confidence that white-hat interventions durably improve discoverability across platforms, and very low confidence that citations predict clicks, conversions or revenue (Martinez, arXiv:2607.14035, July 2026, preprint). The same survey reports URL-level overlap between Google AI Overview, Gemini and organic results at a Jaccard similarity of 0.11 to 0.18, which is a formal way of saying there is no single ranking to optimise for.
What else the literature establishes
Position bias is real and mechanical: Stanford's "Lost in the Middle", published in TACL, showed that retrieved information placed in the middle of a long context is used measurably less than information at either end (Liu and colleagues, 2024). Brand bias is real: a study across four models found statistically significant favouritism towards global over local brands, and one model recommending luxury brands to high-income contexts 98.88 % of the time (Kamruzzaman and colleagues, University of South Florida). Persuasion does not transfer from human psychology: an EMNLP 2025 paper found social proof consistently raising recommendation rates while scarcity and exclusivity reduced visibility (Filandrianos and colleagues, NTUA), a result echoed in Harvard Business Review research reporting that "only star ratings consistently increased choice in the expected direction" for AI shopping agents, with countdown timers, strike-through pricing and bundles unreliable or negative (Sabbah and Acar, HBR, May 2026).
Verifiability remains weak. Stanford's evaluation of generative search engines found only 51.5 % of generated sentences fully supported by their citations and 74.5 % of citations supporting the sentence they were attached to (Liu, Zhang and Liang, EMNLP 2023). A 2026 study of deep research agents found link validity above 94 % while factual accuracy ran between 39 % and 77 % across frontier models (arXiv:2605.06635). For a brand this is a reputational exposure as much as a visibility one: the answer that names a company may attach that name to a claim the company never made.
And the economics have been modelled. An NBER working paper funded by the Harvard Business School AI Institute argues that an AI platform which "underinternalizes future content reproduction retains too little referral traffic", destabilising the open web "even with truthful content, accurate answers, and rational users" (Chan, NBER w35344, June 2026). A causal estimate using Google's staggered geographic rollout against Wikipedia's language editions, across 161,382 matched article-language pairs, put the traffic effect of AI Overview exposure at roughly minus 15 % (Khosravi and Yoganarasimhan, arXiv:2602.18455). This is the cleanest causal number in the entire literature, and it is negative.
Consultancies and analysts
Four in ten companies already do this work, inside marketing budgets that are shrinking. The consultancies have found the story. The defensible numbers underneath it are fewer than the volume of publication suggests.
The best adoption benchmark is academic rather than commercial. The CMO Survey, run by Duke University's Fuqua School with the American Marketing Association and Deloitte, surveyed 308 marketing leaders at United States for-profit companies in January 2026, 97 % of them vice-president level or above, and found generative engine optimisation already in use at 41.5 % of companies (The CMO Survey, February 2026 edition). The same survey reports artificial intelligence performing 24.2 % of marketing activities, up from 13.1 % in 2024, and marketing budgets at 9.0 % of revenues, the lowest since 2021.
That combination is the commercial fact of the category. The work is being adopted quickly inside budgets that are not growing, which means it is being funded by reallocation. Practitioner reporting bears this out: agencies describe clients expecting roughly half of an existing search retainer to now cover answer-engine work, and one brand reallocating 30 to 35 % of a new product's budget experimentally (Digiday, 2026). Gartner's spend survey of 401 mostly billion-dollar-plus companies puts artificial intelligence at 15.3 % of marketing budgets with only 30 % reporting mature readiness (Gartner, 11 May 2026).
What the consultancies are telling boards
Accenture's consumer research is the largest sample in circulation, at 25,590 consumers across sixteen countries, and its findings are about influence rather than traffic: 71 % expect artificial intelligence to influence at least half their spending decisions within twelve months, 74 % say they would trust a personal agent more than their closest friend to make purchases on their behalf, and only 9 % are comfortable with fully autonomous purchasing (Accenture, 3 June 2026). The formulation that travels best from that work is that brands now have to earn trust twice, once with the person and once with the system that advises them.
McKinsey's European survey of 749 consumers in December 2025 found 38 % using AI tools when researching products and deciding what to buy, with comparison of brands, models, prices and reviews the leading use at 63 %, and framed the current state as decision influence arriving before execution does (McKinsey, 2 March 2026). BCG's index of 238 senior marketing leaders found 67 % expecting high disruption to their category's consumer journey, and sorted seventeen verticals into four exposure archetypes (BCG, 21 January 2026). Bain's much-quoted claim that around 60 % of searches end without onward travel and that organic traffic is down 15 to 25 % rests on 1,117 survey respondents from December 2024 and should be cited with that attached (Bain and Dynata, 19 February 2025).
For business-to-business sellers the sharper numbers come from the review platforms. G2 surveyed 1,076 decision-makers in March 2026 and found 51 % now beginning software research with an assistant more often than with Google, 69 % choosing a different vendor because of AI guidance, 33 % buying from a vendor they had not previously heard of, and 85 % viewing a vendor more favourably when an assistant mentions it, alongside 64 % encountering inaccurate recommendations often or very often (G2, 15 April 2026). TrustRadius, surveying 1,862 technology buyers in January 2026, found 63 % using assistants during the purchase journey, 94 % of those fact-checking the output, and 83 % shortlisting three or fewer products (TrustRadius, 15 July 2026).
Three products make the shortlist and one prompt makes the three. Answer-engine presence has become a qualification gate rather than a brand-awareness nicety.
Where the category was formalised
Gartner published a Market Guide for Answer Engine Visibility Tools on 9 March 2026 (Gartner, paywalled). Whatever its analytical content, the effect is procedural: once a category has a Gartner definition, enterprise procurement has a box to tick, and within weeks several vendors had issued releases announcing their inclusion. That is why the acronym fight matters commercially even though it is intellectually empty. Answer engine optimisation is winning the budget line; generative engine optimisation still leads in search volume, at 4,400 monthly United States searches against 2,400, though the first is falling 33 % year on year and the second rising 26 % (DataForSEO Labs data, August 2026).
Two observations are worth recording for balance. Omnicom's first-quarter 2026 earnings call, read in full, contains no discussion of AI search, search disruption or answer-engine optimisation at all (Omnicom, 28 April 2026). WPP's restructuring announcement in February 2026 attributed its position to a broader change in the commercial ecosystem rather than to search (VideoWeek, 26 February 2026). The largest holding companies are not yet telling investors that this is what is happening to them, which is either a lagging indicator or a caution, and an honest paper should say it could be either.
The vendor landscape
A billion-dollar exit, a wave of venture funding, and a product that is a dashboard. Here is who sells what, and which structural gaps remain open.
The defining transaction was Adobe's acquisition of Semrush, announced on 19 November 2025 at twelve dollars a share against a prior close of 6.89, a premium of roughly 74 % (TechCrunch, 19 November 2025), and completed on 28 April 2026 (Adobe, 28 April 2026). Adobe's own framing was explicit about the reason, citing generative optimisation and folding the asset into its customer experience suite alongside its LLM Optimizer product. Two structural facts follow. First, the public market had under-priced the option by a wide margin. Second, because Semrush had earlier acquired Third Door Media, the trade press that covers this category most heavily is now owned by one of its largest vendors (Search Engine Land, 16 October 2024).
The incumbents have all repositioned. Similarweb launched a generative intelligence toolkit in July 2025 built on clickstream rather than prompt sampling, which makes it the one incumbent measuring a different quantity (Similarweb, 28 July 2025). Conductor describes its product as a system of record for answer-engine optimisation, with an application inside ChatGPT (BusinessWire, 1 April 2026). Botify grounds its version in crawl, Search Console and log files (BusinessWire, 22 October 2025). Yext moved earliest, in March 2025, measuring visibility at store level (Yext, 3 March 2025). Alongside the incumbents sits a cohort of venture-funded specialists selling prompt sampling, sentiment and citation tracking, most of them as a dashboard.
Around the software sits a services layer with the economics of 2015 search consulting. Published agency ranges put answer-engine retainers at 1,500 to 15,000 dollars a month, with full enterprise programmes at 15,000 to 50,000 (Gigawatt Group, 2026 survey of agency pricing). The delivery model is content rewriting, schema implementation, entity work, third-party placement, technical crawlability and a dashboard resold from one of the vendors above. Margin lives in the services. That is why so many tool vendors are moving into content production, and why the tools alone are structurally fragile.
The gaps that are actually open
- Statistical power. Entry tiers built on small prompt sets with a single daily run are underpowered against the sampling requirements established in the literature, which the measurement crisis below sets out. Almost no commercial tool reports a confidence interval.
- The denominator. Every tool asks the customer to hand-pick prompts and then reports share of voice against that arbitrary set. Nobody publishes a weighted, auditable prompt universe for a category.
- Attribution. The category measures mentions and sells them as proxies for outcomes. There is no commercial incrementality testing for answer-engine visibility.
- Independence. Every funded vendor sells both the measurement and the remedy. Two tools can measure the same brand in the same week and return materially different results, which argues for a split between directional and decision-grade measurement.
- Geography and language. Coverage is overwhelmingly United States English, logged out and unpersonalised. No funded product vendor is headquartered in India despite a dense agency market there.
- The action layer. Buyers interviewed by Digiday described the tools as benchmarkers, with no ability to correct a wrong answer, no trending and no sales attribution (Digiday, 1 May 2026). The measurement is a product. The work is still a person.
The last two are the ones Nagent has built against, and the DRIS sections below describe how.
The measurement crisis
Most published AI visibility numbers are one sample from a distribution that changes every day. This is the technical heart of the paper, and the place where a serious product can be separated from a dashboard.
Answer engines are not deterministic. The same question asked twice returns a different set of sources, in a different order, at a different length. The University of St. Gallen study quantified it across four surfaces and four verticals: day-to-day source overlap at Jaccard 0.34 to 0.42, meaning about 65 % of cited sources change between one day and the next, with same-day repeated runs at 0.32 to 0.43, which establishes that the variance is stochastic rather than temporal drift. Brand-level overlap is more stable, at 0.45 to 0.59, because a brand can be mentioned while the page cited to support the mention changes (arXiv:2604.07585, 10 April 2026).
The same paper does something no vendor has done and publishes the sampling requirements. Seven to eight runs per prompt per day are needed for a per-brand estimate with a standard error below 0.10; eight runs at the source level; a ten-day minimum window; and twenty-one to twenty-eight day rolling windows for a brand estimate with a standard error between 0.03 and 0.05. It also reports citation concentration as a Gini coefficient of 0.715 overall, differing by engine, which means cross-engine scores are not comparable without normalisation.
SparkToro arrived at the same place through a public experiment with 600 volunteers running 2,961 trials on twelve prompts: under a one in a hundred chance of the same brand list twice, closer to one in a thousand for the same list in the same order, with lists varying in membership, order and length (SparkToro, 28 January 2026). Agency-side analysis has made the same case against its own commercial interest, noting that panel estimation carries a meaningful margin of error in niche and business-to-business categories, that keyword-to-prompt modelling rests on an assumed conversion factor, and that direct sampling makes no claim about real-world volume at all (Brainlabs, 20 April 2026).
A visibility score without a sample size, a run count and an interval is not a measurement. It is a screenshot of one draw from a distribution, sold monthly.
Four substrates, four different questions
Conflating these is the category's central intellectual error, and it is worth separating them plainly.
- Prompt sampling. Fire a defined prompt set at engines on a schedule and parse the replies. Measures what a model says. Depends entirely on the prompt set, the run count and the model version. Most vendors sit here.
- Inferred prompts. Derive likely prompts from keyword databases and related questions. Higher volume, but the prompts are modelled from search behaviour rather than observed from assistant behaviour.
- Clickstream. Measure actual referral visits from assistants. Answers a different question: what arrives, not what is said. Panel composition drives the error.
- Server logs. Measure which retrieval agents fetch which pages, and whether they can render them. The only substrate that is first-party, complete and free of sampling error.
A defensible programme uses all four and says which one each number comes from. DRIS is built that way: log analysis and crawler probing are ground truth, prompt sampling is a measured estimate with an interval attached, clickstream and analytics are the outcome series, and inferred prompts are used only to construct candidate universes that are then weighted and disclosed.
The denominator problem
Share of voice was a coherent metric when the denominator was a fixed keyword set with known search volumes. In the answer layer the universe of possible phrasings is effectively infinite, so every share-of-voice figure is a share of an arbitrary sample. The honest version requires three disclosures that almost nobody makes: how the prompt universe was constructed, how it is weighted to reflect real demand, and what changed in it between two reporting periods.
There is also exogenous volatility with no marketing cause. When a major model version shipped in September 2025, platform-wide outbound citation volume moved, changing every brand's score for reasons unrelated to anything a brand did (Search Engine Land, 8 June 2026). Any reporting layer that cannot separate a model release from a marketing effect will attribute the platform's behaviour to the team's work, in both directions, and will eventually be caught doing it.
The measurement contract DRIS holds itself to. Every reported metric carries its run count, its window, its engine, its locale and an interval. A change is reported as significant only when it clears the interval. Model-version changes are recorded as annotations on the series. Prompt universes are versioned, and a change to the universe invalidates comparison across the boundary rather than quietly resetting the baseline.
Category exposure
Category decides how much of classic search still carries over, and the spread is nearly two to one. An average across industries is close to useless here. The variance is the finding.
BrightEdge measured the overlap between AI Overview citations and the organic top ten by industry in February 2026, against the same measure a year earlier. Healthcare sat at 24.0 %, business technology at 22.6 %, education at 23.1 %, insurance at 22.4 %, entertainment at 18.5 %, travel at 17.7 %, e-commerce at 13.4 %, finance at 11.3 % and restaurants at 9.3 % (BrightEdge, 12 February 2026).
Read that as an inheritance rate. In healthcare and business technology, roughly a quarter of what appears in the answer comes from the pages that already rank, so a strong classic search position still transfers. In e-commerce and finance, roughly seven citations in eight come from outside the top ten, which means a retailer holding position one for a category term is contributing almost nothing to the answer that has replaced it. The same measurement also shows travel and entertainment moving from near-zero overlap in 2025 to high-teens in 2026, which is a caution against treating any of these as stable constants.
Trigger rates differ as sharply. Seer Interactive, across 5.47 million tracked queries, found AI Overviews present on 36 % of informational queries, 8 % of commercial and 5 % of transactional, while comparison-format queries triggered them 95.4 % of the time and question-format queries 85.9 % (Seer Interactive, April 2026). Comparison is where brands are chosen, and comparison is where the answer layer is nearly total.
How much of the map is still unclaimed
Semrush studied 50,000 brands across 1,094 United States categories, 220,000 domains and 600,000 citations between January and June 2026, defining a clear owner as a brand appearing in four of five representative prompts with a lead of at least five percentage points. Only 15.2 % of topics had a clear owner, 31.2 % had an emerging leader, and 53.7 % were unsettled. Domain authority predicted ownership in 52.5 % of cases, organic traffic in 48.4 % and branded search volume in 55.7 %, which is to say barely better than a coin toss. Where a clear owner existed, it retained first place in 90.4 % of month-on-month comparisons (Semrush, 20 July 2026).
Those two numbers together, 53.7 % unsettled and 90.4 % retention, describe a land grab with a closing window. Positions are mostly unowned, existing search strength does not automatically claim them, and once claimed they are sticky. That is the most commercially useful pair of statistics in the whole evidence base, and it is the reason a brand's first serious programme in this area is worth more than its fifth.
The visibility ladder by brand size
A study of 102,025 responses across five engines placed tier-one global brands at 73 % visibility, mid-market brands at 44 % and niche brands at 11 %, with prompt type driving enormous variation: brand-research prompts above 97 %, comparison above 74 %, discovery at 22.9 % and problem-to-solution prompts at 10.8 % (arXiv:2606.20065, June 2026, commercial affiliation disclosed). The gradient by prompt type is the actionable part. A company that tests only its own name will conclude it is visible. The prompts that matter are the ones where the company is not named, and those are precisely the ones where mid-market and niche brands measure worst.
| Where the exposure sits | Classic search era | Answer layer |
|---|---|---|
| Retail and D2C | Category rank, product page SEO, marketplace listing | Feed completeness, price and availability attributes, review corpus, machine-readable product pages, only 13.4 % citation inheritance from rank |
| FMCG and consumer goods | Brand search, retailer co-op, shelf | Entity facts, ingredient and claim consistency, retailer and review-site agreement, weak AI referral in grocery categories |
| B2B software | Category keyword rank, review-site profile | Shortlist formation in one prompt, review-site and community presence, comparison-page evidence, 22.6 % inheritance from rank |
| Financial services | Rate tables, comparison sites, branded search | Regulated-claim consistency, 11.3 % inheritance, high hallucination exposure on product terms |
| Travel and hospitality | Destination content, OTA placement | Itinerary and planning prompts, freshness sensitivity, inheritance from rank rising quickly year on year |
| Healthcare | Condition content, authority signals | 24.0 % inheritance, strongest carry-over from classic search, highest accuracy and liability exposure |
India and the vernacular front
India is a leading market for this change, not a lagging one, and it is voice-first and vernacular. The measurement category has almost no presence here. The behaviour it measures is further along than in the markets it serves.
Sam Altman said in February 2026, ahead of the India AI Impact Summit, that India has 100 million weekly active ChatGPT users and is OpenAI's second-largest market after the United States, with the largest student user base globally (TechCrunch, 15 February 2026). Google India states that "India represents our largest user base for both voice and visual search globally", and has expanded AI Mode into seven further Indian languages beyond English and Hindi, making India the first country outside the United States to receive conversational voice and camera search (Google India, 8 October 2025).
Distribution is being subsidised at telco scale. Jio users were offered eighteen months of a premium Google AI subscription at no cost from 31 October 2025, a stated value of 35,100 rupees (Jio). A low-priced ChatGPT tier was made free for a year in India. Willingness-to-pay curves that gate adoption in Western markets do not apply here, which means assistant-mediated discovery will reach tier-two and tier-three India faster than it reached mid-market America.
The commerce base underneath is large and growing quickly. Bain, working with Flipkart, put India at 290 to 300 million online shoppers in 2025 with e-retail gross merchandise value of 65 to 66 billion dollars growing 19 to 21 %, and quick commerce at 10 to 11 billion dollars, doubling annually since 2023 and projected to supply 45 to 50 % of incremental e-retail value by 2030 (Bain and Flipkart, 8 April 2026). Amazon India reports more than ten million customers using its shopping assistant since its November 2024 launch there (Amazon India).
Three consequences follow for an Indian brand, and they are different from the Western playbook.
- Questions are spoken, long and code-mixed. AI Mode questions run far longer than a traditional query, and India's queries are disproportionately voice and visual. Content written to match short typed keywords answers a question nobody is asking any more.
- Discovery already happens in chat. Meta-commissioned research reports 72 % of product discovery in India happening on a messaging surface and nearly 60 % of users likely to buy after seeing an offer there (Meta and Ipsos, 13 May 2026, vendor-commissioned). Conversational commerce is not a new behaviour to teach in this market; it is the existing one, now with a model behind it.
- There is no regulatory backstop. The Competition Commission of India published a market study on artificial intelligence and competition in September 2025 and recommended self-regulation with competition-law checks rather than hard rules. Unlike the United Kingdom, where the Competition and Markets Authority has imposed conduct requirements obliging Google to rank organic results including AI Overviews on objective and non-discriminatory criteria (CMA, 17 June 2026), Indian brands have no remedy to appeal to. Visibility here is purely an earned problem.
The final asymmetry is technical. Indian consumer sites lean heavily on application-style front ends that render content in the browser. Retrieval crawlers largely do not execute that code. Adobe's machine-readability scoring of United States retailers found product pages averaging 66 out of 100, the weakest page type, and grocery at 48 (Adobe, 16 April 2026). No equivalent India benchmark has been published. Every audit Nagent has run here suggests the local baseline is worse, and that absence of a benchmark is itself the opportunity.
Commerce and the agent rails
Discovery and checkout are being re-platformed at the same time, and the rails are not settled. For a retail or consumer brand, the product feed is now a visibility asset with a ranking function attached.
The traffic numbers in retail are the most dramatic in the category and the most frequently misused. Adobe Analytics, across more than one trillion visits to United States retail sites, measured AI-referred traffic up 693 % year on year over the 2025 holiday season, 393 % in the first quarter of 2026 and 138 % by May 2026, with a cumulative increase of 1,324 % since October 2024 (Adobe Digital Insights, 10 January 2026; Digital Commerce 360, 17 June 2026). Read against Semrush's 0.14 % base, these are very large multiples of a very small number. Both belong in the same paragraph or neither should be quoted.
The conversion crossover is the durable finding. In March 2025 AI-referred retail visitors converted at roughly half the rate of other traffic. By March 2026 they converted 42 % better, with 37 % higher revenue per visit, and by May the advantage was 54 % (Adobe, Q2 2026 report). Adobe's March 2026 consumer survey of more than 5,000 respondents adds the behavioural detail: 39 % had used artificial intelligence for online shopping, 50 % click the links an assistant provides, 27 % have purchased through one, and 69 % say they are less likely to return what they bought (Adobe, 16 April 2026).
Salesforce, measuring across 1.5 billion shoppers in 89 countries, attributed 262 billion dollars of global online holiday sales in 2025 to AI and agent influence, around 20 % of the total, and reported that shoppers arriving from AI search converted nine times more often than those from social referrals (Salesforce, January 2026). Its forecast for the 2026 holiday season is that 20 % of e-commerce traffic will arrive through AI chat agents (Salesforce, 20 July 2026).
Three competing rails
OpenAI and Stripe launched Instant Checkout and the Agentic Commerce Protocol on 29 September 2025 with Etsy and more than a million Shopify merchants, stating that results are "organic and unsponsored, ranked purely on relevance", with merchants paying a fee on completed purchases (OpenAI, 29 September 2025). By March 2026 OpenAI had moved that flow into its apps surface rather than keeping a standalone in-chat checkout (Digital Commerce 360, 6 March 2026). Google published the Agent Payments Protocol in September 2025 with more than sixty partners, built on signed intent and cart mandates carried over its agent-to-agent and tool protocols (Google Cloud, 16 September 2025), then a Universal Commerce Protocol in January 2026 with Shopify, Etsy, Wayfair, Target and Walmart, and more than twenty endorsers including Flipkart (Google, 11 January 2026). Visa and Mastercard have both shipped agent payment tooling.
Nagent's position is that a brand should not build to one protocol in 2026. What every rail requires is the same underlying asset: a complete, accurate, machine-readable product record with price, availability, specification, primary-seller status and returns policy expressed as data rather than as page furniture. Feed quality is the part that survives whichever protocol wins, and it is currently owned by nobody in most marketing organisations.
The uncomfortable implication for brand marketing. The published ranking criteria that exist are availability, price, quality and primary-seller status. The research on AI shopping agents finds star ratings the only consistently positive conventional lever, with scarcity cues, countdown timers and bundles unreliable or negative. A large share of conventional conversion-rate optimisation was built for a human eye and does not transfer. That is a finding to test rather than to accept, but it should be tested deliberately rather than discovered in a quarter's numbers.
The economics of the corpus
The corpus a model reads is becoming a market, and access is now a commercial decision. Crawl economics, licensing and litigation are converging on a single question: who pays for the information that answers get made from.
Cloudflare published the ratio that framed the debate. Measured across its network, pages crawled per referred visitor ran at roughly 5.4 for Google, 1,091 for OpenAI and 38,065 for Anthropic in July 2025, having been far higher for Anthropic earlier in the year before it added web search (Cloudflare, 2025). In July 2025 it changed the default for new domains to block AI crawlers and introduced a pay-per-crawl marketplace (Cloudflare, 1 July 2025). In September it published a content signals policy adding purpose-scoped directives to robots.txt, distinguishing search indexing, AI input for grounding and retrieval, and AI training, and applied it to a managed file already running on 3.8 million domains (Cloudflare, 24 September 2025).
The nuance in that default matters and is widely missed. The managed policy expresses a preference against training while leaving the grounding case unstated, which means a site can believe it has opted out of AI use while remaining fully available to the retrieval pipeline that produces answers. Crawler policy has become a three-way commercial decision and it is being made, in most companies, by whoever last edited a text file.
Licensing is the second channel and it is mostly private: a few publisher agreements with model developers have reported values, alongside a long list of undisclosed arrangements. Litigation is the third, with publisher actions against model developers continuing to accumulate. The regulatory layer is moving fastest in the United Kingdom, where the Competition and Markets Authority has brought AI Overviews inside a non-discrimination ranking requirement and mandated search data portability.
Reddit is the clearest economic signal. Its reported revenue is overwhelmingly advertising, with data licensing a small and slower-growing line beside it. Being the corpus pays far less than being the destination. For a brand the lesson is not about licensing at all; it is that the platforms which supply most citations are being paid as media businesses, and will make product decisions accordingly.
Behind all of this sits the structural argument from the economics literature in the research section: a system that extracts content without returning enough traffic degrades the supply it depends on, and a model good enough to induce machine-written content homogenises the corpus that trains its successor. Nagent takes no position on how that should be resolved politically. Operationally it means one thing. A brand's own first-party corpus, its documentation, its data, its named experts, its customer evidence, is the only part of the supply chain it controls, and its value rises as the commons is enclosed.
The hypothesis
Organic visibility is now a claim-supply problem, and the unit of work is a system rather than a page. What follows is the argument Nagent is building against. It is stated as a hypothesis because parts of it are not yet proven, and the parts that are not are marked.
Take the evidence above at face value and a coherent picture emerges. Retrieval is decomposed into sub-queries. Selection draws from a candidate pool that overlaps only partially with classic rankings and differs by engine. Composition favours passages that carry checkable evidence, sit early, are recent, and come from sources the system has learned to weight. Presentation returns an answer in which the brand is named, implied or absent, and only occasionally becomes a visit.
If that is the pipeline, then the job of an organic programme is no longer to win a position. It is to ensure that a defined set of claims about a company are, at every stage of that pipeline, present, retrievable, consistent, attributable and current. Nagent calls this claim supply. The five properties are the hypothesis.
- Present. The claim exists somewhere a retrieval agent can reach, on the brand's own property and on the third-party surfaces that specific engine draws from. Absence is the most common failure and the least often diagnosed.
- Retrievable. The document is fetchable by retrieval agents that do not execute client-side code, is structured so that the relevant passage is a coherent unit, and is not excluded by a snippet or crawler directive set for another purpose.
- Consistent. The same fact reads the same way across the site, the review platforms, the directories, the encyclopaedia entry and the community threads. Inconsistency does not average out; it produces hedged answers or a competitor's version.
- Attributable. The claim is backed by something checkable: a figure, a named source, a quotation, a document. This is the one property with direct experimental support, from the KDD paper's tactic tests to the 252,000-comparison factorial study.
- Current. The page carries a recent and genuine update. Seer's study of 47,097 citations found 75 % of cited pages updated within a year, and that freshness came from updating rather than publishing, at 72 % against 42 % (Seer Interactive, 2026).
- Proven statistically. Because the surface is non-deterministic, none of the five can be asserted from a single observation. Every claim of effect requires repeated sampling, a stated window, an interval and, where possible, a holdout.
Why this is a system problem and not a content problem
Each property sits with a different owner in a typical company. Presence on third-party surfaces belongs to communications and to partnerships. Retrievability belongs to engineering and to whoever chose the front-end framework. Consistency belongs to brand and to product marketing and to whoever maintains the retailer feeds. Attribution belongs to content and to the subject-matter experts who do not report to marketing. Currency belongs to nobody, which is why it decays.
Classic search optimisation could survive that fragmentation because the work concentrated in one place: a page, on a domain the marketing team controlled. Claim supply cannot, because the pipeline reads everything. This is the structural reason the agency retainer model strains here. The work is continuous, cross-functional, evidence-heavy and repetitive, and it requires several hundred measurements a week to know whether any of it is working.
A team of AI coworkers with clear authority limits is not a cheaper agency. It is the only shape that matches work which is continuous, cross-functional and statistically demanding at the same time.
What the hypothesis does not claim
It does not claim that answer-engine visibility produces revenue. The strongest published survey grades confidence in that link as very low, and Nagent will not assert what forty-five studies decline to. It does not claim that classic search optimisation is obsolete; in healthcare and business technology roughly a quarter of citations still inherit from the top ten, and in every category indexing remains the entry condition. It does not claim a durable advantage from the tactics themselves, since the congestion results suggest gains decay as adoption rises. And it does not claim that a vendor score is a business metric.
What it does claim is narrower and more defensible. The five properties are measurable. They are addressable. They are currently unmanaged in most companies. And the categories in which they are unclaimed are, by the best available count, a majority.
Nine propositions
The argument, stated formally: nine propositions, each with the evidence that supports it and the condition that would falsify it.
1. The measurable object is a distribution, not a rank
Falsified if engines become deterministic at the answer level.
Evidence. Day-to-day source overlap of 0.34 to 0.42 and same-day overlap of 0.32 to 0.43 across four engines, with a published requirement of seven to eight runs per prompt per day and twenty-one to twenty-eight day windows for a usable standard error. Independently, under a one in a hundred chance of an identical brand list across repeated runs.
Consequence. Reporting must carry run counts and intervals. A month-on-month change inside the interval is not a result, and treating it as one produces false attribution in both directions.
Test. If a platform ships deterministic, cacheable answers for commercial queries, single-sample measurement becomes valid and this proposition fails.
2. Retrieval and citation are separate problems with opposing optimisations
Falsified if structural and body-text work stop trading off.
Evidence. Full-pipeline evaluation over 171,003 documents found body-text rewrites producing an average retrieval rank drop of 4.54 positions and a 9 % fall in hit rate, while structural and metadata work raised hit rate by 22 %, with the great majority of citations still drawn from body text.
Consequence. Optimisation must be staged: structure for retrieval, evidence for citation, measured separately. A single composite score conceals a trade-off that can move the two in opposite directions.
Test. A repeated study on production sites showing no retrieval penalty from evidence-heavy rewriting would collapse the two workstreams into one.
3. Access policy is a marketing decision being made by other functions
Falsified if platforms unify their crawler taxonomies.
Evidence. Three distinct crawler classes with different consequences, documented separately by OpenAI, Google, Perplexity and Meta; an explicit statement that opting out of one retrieval agent removes a site from that engine's answers; user-initiated fetchers that may ignore robots.txt entirely; and a default-block posture applied at network scale to millions of domains.
Consequence. The first deliverable of any programme is an access verdict per engine, per crawler class, with the line of the file that causes it and a named owner for the decision.
Test. A common standard with a single purpose-scoped directive honoured by every major engine would reduce this to a one-time configuration.
4. Evidence beats packaging, and packaging alone does nothing
Falsified if formatting-only interventions start showing effects.
Evidence. Quotation addition at 40.9 % relative improvement, statistics at 32.8 %, source citation at 27.5 %, keyword stuffing at minus 8.3 % in the founding benchmark; odds ratios near 1.0 for formatting-only edits against extreme values for topical relevance, explicit price and recency in a 252,000-comparison factorial design.
Consequence. The content workstream is an evidence-gathering exercise, not a rewriting exercise. Its inputs are figures, named experts, documents and permissions, which is why it stalls without access to the rest of the business.
Test. A large, well-controlled study showing presentation effects independent of evidence would change the work substantially.
5. Consistency across third-party corpora is a first-class channel
Falsified if engines converge on first-party sources.
Evidence. Encyclopaedia and community sources supplying more than a quarter of United States citations on one engine, a different platform leading on another, review platforms dominating business software answers, and a single undisclosed model change cutting one source's share from roughly 60 % to roughly 10 % in weeks.
Consequence. Third-party presence is managed as a portfolio with concentration limits, not as a campaign. Dependence on any single external platform is recorded as a risk with a number attached.
Test. A sustained shift towards first-party and licensed corpora would move the weight of the work back onto owned properties.
6. Freshness is an operating discipline, not a publishing schedule
Falsified if recency weighting is removed.
Evidence. Across 7,683 dated pages and 47,097 citations on three engines, 75 % of cited pages were updated within a year and 88 % within two; among pages with both dates, 72 % were fresh by last update against 42 % by publication date; recency carried an odds ratio of 14.4 and above in the factorial study.
Consequence. A maintenance backlog on existing high-value pages outperforms a new-content calendar, which inverts the standard content plan and the standard agency incentive.
Test. Engines de-weighting recency for evergreen categories would reduce this to a category-specific rule.
7. The field is congested and the advantage goes to the early and the specific
Falsified if late entrants match early ones on equal effort.
Evidence. Gains falling as adopters rise in a peer-reviewed benchmark; individual payoff collapsing from 0.802 to 0.007 under simultaneous optimisation while non-participants received nothing; 53.7 % of categories with no clear owner, and 90.4 % month-on-month retention where an owner exists.
Consequence. Programmes should target unsettled sub-topics rather than contested head terms, and should be resourced to claim and hold rather than to spike.
Test. Category ownership turning over freely month to month would remove the first-mover argument entirely.
8. Answer visibility is an accuracy and liability surface, not only a growth surface
Falsified if citation support approaches completeness.
Evidence. 51.5 % of generated sentences fully supported by their citations and 74.5 % of citations supporting the sentence attached to them; link validity above 94 % alongside factual accuracy between 39 % and 77 % in deep research agents; manipulation studies showing manufactured consensus succeeding at high rates against weaker models.
Consequence. Monitoring must include what is said about a brand, not only whether it is named, with escalation paths for a materially wrong statement about price, safety, availability or regulated claims.
Test. Verifiability rates approaching completeness would reduce this to routine brand monitoring.
9. The work is continuous, so it belongs to a governed agent team
Falsified if the workload proves to be periodic.
Evidence. Sampling requirements of hundreds of runs a week for a usable interval; five properties owned by five different functions; buyers reporting that current tools benchmark but cannot act; agency retainers priced from 1,500 to 50,000 dollars a month for delivery that is largely repetitive.
Consequence. The operating model is a standing team with defined authority per action class, an approval queue for anything that changes a live property, and a record of every action taken and its measured effect.
Test. If a quarterly audit plus a content calendar produced comparable measured outcomes, the continuous model would be unnecessary overhead.
What the work now is
Seven workstreams replace the keyword plan, and only two of them look like content. This is the shape of an organic programme once the five properties are taken seriously.
The practical translation of claim supply is a standing set of workstreams, each with its own inputs, its own owner and its own measurement. They are listed here as work rather than as software so that a marketing leader can see what has to happen whether or not any of it is bought from Nagent.
- Access and eligibility. Crawler policy per engine and per class, snippet directives, rendering behaviour, status codes, network-level rules, feed endpoints. The entry condition for everything else. Owned in DRIS by the Access Prober.
- Corpus mapping. Which sources each engine actually draws on for this category, how concentrated they are, where the brand appears in them, and which of them are controllable, influenceable or neither. Owned in DRIS by the Corpus Mapper.
- Extractability. Whether the relevant passage is a coherent retrievable unit, whether it survives rendering, whether structure supports retrieval, and whether the evidence density is sufficient to be cited once retrieved. Owned in DRIS by the Extractability Analyst.
- Knowledge consistency. The canonical facts about the company, expressed identically wherever they appear, with divergence flagged as a defect rather than a variation. Pricing, positioning, capabilities, leadership, categories, claims. Owned in DRIS by the Knowledge Consistency Auditor.
- Evidence production. Figures, named experts, documents, customer proof, benchmarks and permissions, gathered from the business and placed where the retrieval pipeline can reach them. Shared with CREA, the content lead.
- Third-party presence. Review platforms, directories, community threads, listicles, encyclopaedic entries and earned coverage, managed as a portfolio with concentration limits and disclosure discipline. Shared with MOXA and the communications function.
- Measurement and proof. Prompt universe construction, sampling at the required density, interval reporting, model-version annotation, holdouts where possible, and reconciliation against analytics and first-party platform reporting. Owned in DRIS by the Orchestrator.
Two features of this list are worth naming. The first is how little of it is writing. Evidence production is the only workstream that produces prose, and even there the scarce input is the figure rather than the paragraph. The second is how much of it is recurring verification rather than one-time change. Access policy drifts when a security team ships a rule. Consistency drifts when a retailer updates a listing. Corpus composition drifts when a model version ships. Freshness drifts by definition. A programme that treats any of these as a project will be wrong within a quarter and will not know it.
What this replaces
| Artefact | Classic organic programme | Claim-supply programme |
|---|---|---|
| Target | Keyword list with volumes and difficulty | Versioned, weighted prompt universe with intent classes |
| Primary metric | Average position, organic sessions | Citation and mention rate with intervals, by engine and locale |
| Unit of change | A page | A claim, and every surface that carries it |
| Technical scope | Site speed, indexation, internal links | Retrieval-agent access, rendering, passage structure, feed integrity |
| Off-site work | Link acquisition | Presence and factual agreement across the cited corpus |
| Content cadence | New publishing calendar | Maintenance backlog on high-value pages, evidence first |
| Reporting rhythm | Monthly rank report | Rolling window with significance, plus an action log with measured effect |
| Failure mode | Rankings fall and everyone can see it | Silent absence from a surface nobody is sampling |
DRIS in three parts
DRIS audits, proposes, and then acts within limits a human sets. It has three parts, in that order, because the order is the governance.
DRIS is Nagent's organic visibility AI coworker: the team lead for search and answer-engine work inside a marketing organisation that also contains CREA for content, MOXA for social, HOOK for hooks and briefs, RUPA for graphics, Alpha for advertising video and NIA for paid media, with MIRA as chief of staff. It runs on the Nagent platform, which supplies the control plane, the shared workspace, the memory layer and the governance model. There is no separate DRIS application to administer, because agents are configured and constrained through the control plane rather than through their own interfaces.
The product is organised in three parts, and the decision to order them this way was deliberate.
- Audit and analysis. Establish what is true now: access verdicts per engine, corpus composition, extractability defects, knowledge inconsistencies, and a measured visibility baseline with intervals. Nothing is recommended before this is complete, because most recommendations in this category are guesses about a state nobody has checked.
- Proposals. Every finding becomes a proposal with an expected effect expressed as a range, an evidence grade, the surface it touches, the risk, the measurement plan and the rollback. Proposals are ranked by expected effect against effort and risk, not by how easy they are to produce.
- Actions. Approved proposals are executed within the authority level set for that action class, logged, and measured against the plan attached to them. What the agent may do without asking is a configuration decision, not a product default.

The reason for the sequence is that the category's failure mode is acting before measuring. A vendor dashboard reports a number, an agency proposes work, the work ships, the number moves, and nobody can say whether the movement was the work, a model release or the ordinary day-to-day variance of a surface where roughly 65 % of sources change overnight. DRIS is built so that the third part cannot run ahead of the first.
An agent that acts without a baseline is a faster way to be wrong, and in this category being wrong is nearly invisible for a quarter.
What DRIS does not do
It does not promise a ranking. It does not report a single composite AI visibility score as a headline number, because the literature shows that structure and citation can move in opposite directions and a composite hides it. It does not write content without evidence supplied from the business. It does not publish to a live property at any autonomy level below explicit approval unless a client has deliberately raised it. And it claims revenue attribution only where it can demonstrate it: where a holdout is not possible, the reporting says so.
The agent architecture
Six agents, each owning one question that can be answered wrongly on its own. The division is by failure mode rather than by feature, because the failures are what a programme has to catch.
DRIS is a team rather than a model with a long instruction. The reason is diagnostic. A single agent asked to assess visibility will produce one plausible narrative. Six agents, each responsible for a different class of defect, produce six findings that can contradict each other, and the contradictions are usually where the real problem is. A brand that is retrievable, consistent and well evidenced but still absent is telling a different story from one that is cited constantly with the wrong price attached.
| Agent | Owns the question | What it does | Fails if |
|---|---|---|---|
| Orchestrator, team lead | What should be looked at | Holds the prompt universe and its versions, schedules sampling at the required density, reconciles the four measurement substrates, decides what is significant, sequences the other five, and assembles the period narrative. Reports to MIRA. | A change is declared that sits inside the interval |
| Access Prober | Can the machines reach this at all | Tests retrieval, training and user-initiated agents separately against every relevant property, reads robots directives, snippet controls, network rules, status codes and rendering behaviour, and returns an eligibility verdict per engine with the offending line cited. | An engine-level block persists undetected |
| Corpus Mapper | Where do the answers come from | Builds the cited-source map for the category by engine, measures concentration, locates the brand and its competitors within it, and separates controllable, influenceable and inaccessible sources. Flags dependence above a set threshold as a risk. | A source shift is read as a marketing effect |
| Extractability Analyst | Is the passage usable once found | Assesses rendering survival, passage coherence, structural markup, evidence density and the position of the relevant claim within the document. Reports retrieval effects and citation effects separately, never as one score. | A citation gain is bought with a retrieval loss |
| Knowledge Consistency Auditor | Does the world agree about this company | Compares canonical facts against every surface that carries them, including review platforms, directories, retailer listings, encyclopaedic entries and community threads, and raises divergence as a defect with a named owner and a correction path. | A wrong price or claim circulates uncorrected |
| Report Composer | Can a human act on this | Turns findings into the client-facing audit and the period report, carrying method, sample size, window and interval with every number, and writing the uncertainty in rather than around it. | A reader cannot tell measurement from assertion |
How a cycle runs
- The Orchestrator opens the period against the current prompt universe version and schedules sampling across engines and locales at the run density the measurement contract requires.
- The Access Prober runs first, because every other finding is conditional on eligibility. An engine-level block halts the rest of the analysis for that engine and is escalated immediately rather than reported at period end.
- The Corpus Mapper rebuilds the source map and compares it against the previous period, separating composition change from brand performance change.
- The Extractability Analyst works the pages that the corpus map shows are in or near the candidate set, and only those, since auditing pages no engine reaches is the most common way this work wastes a quarter.
- The Knowledge Consistency Auditor sweeps the canonical fact set across surfaces and raises divergences with a severity based on commercial exposure rather than on how visible the surface is.
- The Report Composer assembles findings, proposals and the measured effect of the previous period's actions into one document, and the Orchestrator signs the period narrative.
Around this sits the wider organic system Nagent has specified, which adds agents for content maintenance, third-party placement, feed integrity and community presence. The six above are the ones that constitute DRIS itself, and the ones a client is buying when they buy it.
Measurement done properly
The measurement contract is the product, and the dashboard is a consequence of it. This section sets out what DRIS samples, how often, and what it refuses to report.
The measurement crisis above set out why almost every published visibility number is one draw from a distribution. This section sets out what Nagent does about it, in specific terms, because this is where a buyer should press hardest on any vendor including this one.
The prompt universe
Prompts are not hand-picked by the client and then treated as the world. The universe is assembled from five sources, each tagged: language taken verbatim from sales calls and support tickets, search demand from the licensed keyword corpus, competitor and category framing, the intent classes that matter commercially, and manual additions from the client. Every prompt carries a weight, an intent class and the version in which it entered. Universes are versioned, and a change to the universe breaks comparison across the boundary explicitly rather than quietly resetting a baseline.
The intent classes matter more than the count. A universe weighted towards brand-research prompts will report high visibility for almost any company with a website, because measured visibility on brand-research prompts runs above 97 % while problem-to-solution prompts run near 11 %. DRIS reports by intent class as standard, and the summary figure is a weighted composite whose weights are visible.
Sampling
The published thresholds are the floor: at least seven to eight runs per prompt per day for brand-level detection, comparable density at source level, a ten-day minimum before any statement is made, and twenty-one to twenty-eight day rolling windows for a durable estimate. DRIS samples across engines and, where the client operates in more than one market, across locales and languages, because the same brand scores differently by language and a single-locale number is not a global number.
Reconciliation
Prompt sampling is one substrate of four. DRIS reconciles it against server-log evidence of which retrieval agents fetched which pages, against first-party platform reporting where it exists, and against the client's own analytics for arriving traffic. Where the substrates disagree, the disagreement is reported rather than resolved silently, because the disagreement is usually informative: heavy crawling with no citations points at extractability, citations with no arriving traffic points at answer sufficiency rather than at a defect.
| Question | Substrate | What it can prove | What it cannot |
|---|---|---|---|
| Can machines reach us | Active probes and server logs | Eligibility, rendering, fetch frequency, agent class | Whether reaching us led to anything |
| Are we in the answer | Prompt sampling with intervals | Mention and citation rate by engine, intent and locale | Real-world prompt volume |
| Are we cited by name | First-party platform reporting | Actual citations and grounding queries on one engine | Coverage of engines that publish nothing |
| Did anyone arrive | Analytics and clickstream | Referred sessions, behaviour, conversion | Influence that never became a visit |
| Did our work cause it | Holdout or staged rollout | Directional effect on the treated set | Effect where a holdout is impossible |
What is refused
DRIS will not report a rank position inside an answer as though it were a stable property. It will not report a week-on-week change that sits inside the interval as a result. It will not attribute a movement to client work in a period where a model version changed on that engine without saying so in the same sentence. And it will not report sentiment with the same confidence as mention detection, because sentiment measurement on this surface is materially noisier and pretending otherwise is how a report becomes fiction with charts.
The action layer
What the agent may do without asking is set per action class, and it is earned. The action layer is where this stops being a report and starts being work, which is exactly why it needs limits.
Nagent's platform runs an earned-autonomy ladder of five rungs with a trust score from 0 to 100, set per agent alongside a risk level. Locked means a human must act on every output. Suggest only allows drafts that are visible and never sent. Execute with approval carries out an action after human sign-off. Audit only runs autonomously with logging for review. Fully autonomous acts alone, within the guardrails set out below. Trust scores drift on evidence and can be automatically downgraded, which is visible as a warning in live operations rather than buried.
For DRIS the ladder is applied by action class rather than to the agent as a whole, because the classes differ enormously in blast radius.
| Action class | Example | Default authority |
|---|---|---|
| Observe | Sampling, crawling, log analysis, corpus mapping | Audit only: runs autonomously, logged |
| Diagnose | Raising a finding, grading evidence, ranking proposals | Audit only: runs autonomously, logged |
| Draft | Proposed copy, structured data, a correction request | Suggest only: visible, never sent |
| Change owned property | Editing a live page, altering markup, changing a feed | Execute with approval: requires named approval |
| Change access policy | Editing robots directives or snippet controls | Execute with approval: approval plus a second owner outside marketing |
| Contact a third party | Correction requests to a review platform or directory | Execute with approval: approval, with the outbound text attached |
| Post in a community | Any participation in a public forum | Locked: human acts, agent prepares only |
The last row is deliberate and is worth defending. Community surfaces supply a large share of citations on some engines, which makes them commercially attractive and ethically hazardous in equal measure. The manipulation literature is explicit that manufactured consensus across apparently independent pages is the single most effective attack on retrieval systems, succeeding at high rates against weaker models. An agent that can post in communities at scale is an agent that can manufacture consensus. Nagent does not ship that capability at autonomy, and a client who asks for it is told why the answer is no.
Guardrails and policies
Beyond the ladder, each agent carries guardrails with a severity, a category and an enforcement action: advisory, queue for approval, or block. The templates that matter most for DRIS are the ones that constrain claims. Every factual statement must cite a source. Pricing statements queue for approval. Brand voice and banned terms are enforced. No promises are made about future features or dates. Execution policies cap runs per hour and set daily and monthly spend limits, with the overrun behaviour set to queue actions for approval rather than to stop the agent mid-cycle.
Approval authority is named rather than generic: specific people hold the final decision for an agent, with a fallback to anyone holding the relevant permission. That matters here because the two highest-risk classes, access policy and third-party correction, usually need an approver outside marketing, and a queue that routes to a role nobody occupies is a queue that quietly becomes an autonomy grant.

Every action DRIS takes is recorded with the proposal that justified it, the person who approved it, and the measurement plan it will be judged against. That record is the difference between a programme and a series of opinions.
The learning loop
Actions feed back. Each approved action carries its measurement plan into the next window, and the measured outcome is written to the agent's memory alongside what was expected. Over a few cycles this produces something the category currently lacks: a tenant-specific record of which interventions moved which metric on which engine, with intervals, which is the only honest basis for prioritising the next set. Where the platform's own evidence contradicts a published best practice, the tenant's evidence wins for that tenant, and the contradiction is surfaced rather than hidden.
DRIS inside the org
DRIS is a team lead, not a tool, because the work crosses four other functions. Claim supply fails at the handovers, which is the argument for putting it inside an organisation rather than beside one.
Recall the five properties from the hypothesis and who owns them in a normal company. Presence sits with communications. Retrievability sits with engineering. Consistency sits with product marketing and commerce operations. Attribution sits with subject-matter experts. Currency sits with nobody. A visibility tool reports on all five and can act on none, which is precisely what buyers told Digiday when they described the tools as benchmarkers.
Inside Nagent's marketing organisation, DRIS reports to MIRA as chief of staff and sits alongside CREA for content, MOXA for social, HOOK for hooks and briefs, RUPA for graphics, Alpha for advertising video and NIA for paid media. Team leads can put work into each other's backlogs. That is not an organisational diagram for its own sake; it is what makes a finding actionable. An extractability defect on a product page becomes a task for content. A knowledge inconsistency on a retailer listing becomes a correction request with named ownership. A third-party presence gap becomes a brief. A citation drop on a comparison prompt becomes a paid-media hypothesis for NIA to test while the organic fix ships.
Humans sit in the same workspace, not outside it. The shared thread is where agents post findings, analyses, plans, drafts and results as typed surfaces, where a person can say "record this as a decision", and where every agent turn carries feedback. The team repository holds the charter, the roster, the memory file, the task list, the pending decisions and the decision log. A marketing leader auditing the programme six months later reads the decision log, not a slide.
The forward deployed layer
Nagent pairs the software with a managed layer of forward deployed marketers who run the first cycles with a client, build the prompt universe from real sales and support language, chase the internal evidence that the content workstream depends on, and hand over once the loop is running. The honest description of the first ninety days is that a person does a lot of the gathering and the agents do all of the sampling, the diagnosis and the drafting. The ratio shifts with each cycle, and the client can see it shift in the action log.
Commercial shape
Three tiers, a cost ceiling enforced in code, and the data subscription included. The pricing is set against the sampling requirement rather than against the competitive set, because the sampling requirement is what actually costs money.
The category prices a dashboard. Entry tiers built on small prompt sets at a single daily run are sold at low monthly prices, enterprise deals are quoted separately, and the services wrapper that makes any of it actionable runs from 1,500 to 50,000 dollars a month (Gigawatt Group, 2026 survey of agency pricing). The structural problem with the entry tiers is not the price; it is that the sampling density they can afford is below the level at which a change can be distinguished from noise, which means the cheap tiers are selling a number that cannot support the decision it is bought for.
DRIS comes in three tiers: Starter, Growth and Enterprise. For what each costs, see the plans on the pricing page. The tiers differ by prompt universe size, engine and locale coverage, sampling density, the number of properties audited, and the action classes enabled. A cost-of-goods ceiling of 15 % is enforced in code rather than managed in a spreadsheet, so a tenant cannot quietly become unprofitable through sampling volume, and an agent cannot spend past its budget without queueing for approval.
Two inclusions matter commercially. The keyword and search data subscription is provided by Nagent rather than billed to the client, which removes the most common hidden cost in comparable programmes. And the platform ships with an agentic customer relationship system, so a client does not need to hold a separate licence to route a finding into a commercial workflow.
- Starter: one market, the core surfaces. A single locale, the principal engines, a bounded prompt universe, full audit and proposals, drafting at suggestion level, and monthly reporting with intervals. Intended for a company establishing a baseline and fixing access and extractability defects.
- Growth: multi-engine, multi-locale, acting. A wider prompt universe with intent weighting, several locales and languages, higher sampling density, third-party consistency auditing across the cited corpus, and the change and correction action classes enabled behind approval.
- Enterprise: governed deployment. Private deployment, named approvers and policy sets per property, integration with existing analytics and platform reporting, holdout design where the estimate supports it, and forward deployed marketers for the first cycles.
What a buyer should ask, of Nagent as much as of anyone
- How many runs per prompt per day, and what is the standard error on the headline figure.
- How was the prompt universe constructed, how is it weighted, and what changed in it since the last report.
- Which engines and locales, and are the scores normalised before they are compared.
- Which numbers come from sampling, which from logs, which from first-party platform reporting, and which from analytics.
- What happens to the series when a model version ships, and is that annotated.
- Which recommendations rest on platform documentation, which on measured research, and which on practitioner assertion.
- What can the system change without a human, and who is the named approver for everything else.
- What was measured after the last set of actions, including the actions that did not work.
Nagent's view is that a vendor unable to answer the first two is selling a screenshot. DRIS is held to the same test, which is why its measurement contract is written down rather than implied.
What the pilot is for
Ten design partners, and a deliberate refusal to publish a before-and-after number yet. The reason for the refusal is the same reason the rest of this paper spends so long on method.
DRIS is running with a design-partner cohort of ten tenants across retail, direct-to-consumer, business software and professional services. The audit and proposal layers are in production. The action layer is in production behind approval for owned-property and correction classes. The measurement contract described above is implemented.
What Nagent is not doing is publishing a headline uplift figure from that cohort, and the reason is worth stating plainly because it cuts against commercial interest. Given that roughly 65 % of cited sources change day to day with no intervention at all, a before-and-after comparison of single-run measurements is statistically a coin toss presented as a result. Almost every published case study in this category, including the impressive ones, is built exactly that way. Two of the most visible carry percentage uplifts with no control period, no run count and no interval.
What the cohort has produced so far is a set of findings about the state of the world that are worth more, at this stage, than an uplift claim.
- Access defects are the most common material finding, and in most cases nobody in marketing knew the rule existed. The rule was usually added for a reason that had nothing to do with visibility.
- Rendering is the second. Content that a human sees and a retrieval agent does not is widespread on application-style front ends, and it is invisible to every conventional site audit.
- Knowledge inconsistency is the most commercially dangerous. Divergent pricing, positioning and capability statements across third-party surfaces produce hedged or wrong answers, and the correction path usually runs through a function that has never been asked before.
- Measured visibility on brand-research prompts tells a company almost nothing. The gap between that figure and the problem-to-solution figure is the actual diagnosis, and it is uncomfortable for most brands the first time they see it.
- The single most valuable early action is rarely content. It is an access fix, a feed correction or a consistency repair, all of which are cheap and none of which an agency retainer is set up to find.
When Nagent does publish measured effects, it will publish them with run counts, windows, intervals, engine and locale, the actions taken, and the actions that produced nothing. The last part is the one that will make the document credible.
Disclosure. Nagent sells the product described in the DRIS sections of this paper. The research in the first half is drawn from third parties and is linked so that any reader can check it. Where a number comes from a vendor with an interest, including Nagent, this paper says so in the sentence that carries the number rather than in a footnote.
The counter-case
Five arguments against everything above, stated as strongly as their proponents state them. A thesis that cannot survive its best objections is a brochure.
One: the traffic is a rounding error, and influence is an unfalsifiable claim
Semrush measured AI systems at 0.14 % of all web visits across 50,000 sites for the whole of 2025, against organic search at 16.04 %. Google AI Mode registered 0.01 %. A programme funded on the argument that answer engines matter is being funded on 0.14 % of traffic and a story about influence that cannot be measured directly. The strongest version of this objection adds that every large percentage in the category is growth from near zero, and that the base effect explains most of the excitement.
Response. The objection is correct about traffic and wrong about the decision. The case does not rest on referral volume; it rests on measured substitution in the surface that used to send the volume, on evidence that the arriving minority converts better, and on the finding that more than half of categories have no clear owner in the answer layer while ownership, once held, persists at above 90 % month on month. A company that waits for the traffic to justify the work will arrive after the positions are held. But the objection does mean the budget should be sized as an option on a shifting channel, not as a replacement for search spend, and anyone selling it as the latter should be refused.
Two: the platforms say there is nothing to optimise
Google states there are no additional requirements and no special structured data; its search liaison says the advice has not changed and has explicitly told sites not to chunk content for models. The most rigorous controlled test of a popular technical tactic, adding structured data, found no uplift on any platform. The honest conclusion might be that this is search optimisation with a new acronym and a higher price.
Response. For the retrieval half, that is very nearly right, and this paper says so: indexing and snippet eligibility remain the entry condition and most technical work is ordinary hygiene done properly. What is new is not a ranking lever but a set of facts about the pipeline: that citation and rank overlap at around 12 % across engines, that cited sources change roughly 65 % day to day, that citation supply is concentrated in third-party platforms a search team has never managed, and that crawler classes now determine eligibility per engine. None of that is addressed by doing classic search optimisation harder. The category's error is inventing new levers; the platforms' framing under-describes new failure modes.
Three: it is congested and zero-sum, so the returns will not last
A peer-reviewed benchmark found gains falling as adoption rises and dedicated techniques often damaging ranking. A recommendation study found individual payoff collapsing from 0.802 to 0.007 when brands optimise simultaneously. If everyone does this, nobody benefits, and the vendors selling it know that.
Response. This is the objection Nagent takes most seriously, and it is why proposition seven is stated as it is. Two things survive it. The same studies show that non-participants received nothing at all, which makes this a cost of staying in the consideration set rather than a source of advantage. And the congestion applies to the manipulable tactics, not to the structural properties: an access fix, a rendering fix or a corrected price on a retailer listing does not decay because a competitor also fixed theirs. A programme weighted towards structure and accuracy is far more robust to congestion than one weighted towards persuasive rewriting.
Four: the measurement is too unstable to manage against
If roughly 65 % of sources change overnight, if the same prompt returns a different brand list nineteen times in twenty, and if two tools measuring the same brand in the same week disagree materially, then perhaps this cannot be managed at all, and the responsible answer is to do good marketing and ignore the scores.
Response. Instability is an argument for more sampling, not for no measurement; the same objection would have retired opinion polling and media panels. The published thresholds exist precisely because the variance has been quantified. What the objection does correctly kill is the monthly single-run report, the composite score without an interval, and the case study built on a before-and-after screenshot. Nagent's answer is to adopt the thresholds and to refuse the reporting formats that cannot meet them, which costs more per tenant and is the reason the pricing is what it is.
Five: the visibility will be bought, not earned
Sponsored placement is arriving in answer surfaces. One platform has already added a crawler for validating advertisements, another has launched agent-payment rails, a third is measuring sponsored chat advertising as a currency. If the answer layer monetises the way search did, organic visibility becomes a smaller and smaller slice of a paid surface, and this entire discipline is a temporary arbitrage.
Response. Probably partly true, and the honest position is to plan for it. Two things argue against the strong form. The only published ranking criteria that exist are stated as independent of partnerships and built on product data, and the feed quality that satisfies them is the same asset a paid placement would need. And the accuracy surface does not monetise: a company still has to ensure that what is said about it is correct, whoever paid for the placement. A programme built on structure, accuracy and evidence keeps its value under a paid regime. A programme built on gaming retrieval does not.
What would prove the thesis wrong. Deterministic answers for commercial queries. A single honoured access standard across engines. Convergence back onto first-party sources. A controlled study showing structural and evidence work producing no durable effect on citation. Or a paid layer large enough that organic citation becomes residual. Each is plausible. Each is stated in the nine propositions as a falsification condition, and each would change what Nagent builds rather than how it describes what it has built.
Evidence ledger
Every load-bearing number in this paper, with its method attached. Evidence type is assigned by how the number was produced, not by who published it. Observed behaviour means tracked user sessions or network measurement. Platform documentation means a statement by the company that operates the surface. Academic includes preprints, marked as such where they have not been peer reviewed. Vendor measurement means a study published by a company selling into this category, including studies whose methods are strong.
| Type | Source | Finding | Method and detail |
|---|---|---|---|
| Observed behaviour | Pew Research, 22 Jul 2025 | Clicks fall by roughly half when an AI summary appears | 8 % against 15 %, with 1 % clicking a link inside the summary and 26 % ending the session. 68,879 tracked searches, 900 United States adults, March 2025. |
| Observed behaviour | SparkToro, 9 Jun 2026 | Zero-click at 68.01 %, AI Mode at 0.34 % of sessions | Clickstream panel, January to April 2026, against 60.45 % in 2024. The same study is the source of the caution that cross-year panels are not directly comparable. |
| Observed behaviour | SparkToro, 28 Jan 2026 | Under one in a hundred chance of the same brand list twice | 600 volunteers, 2,961 runs, twelve prompts, three surfaces. Lists varied in membership, order and length. |
| Observed behaviour | Cloudflare, 2025 | Pages crawled per referred visitor: 5.4, 1,091 and 38,065 | Google, OpenAI and Anthropic respectively, July 2025, measured across the Cloudflare network as HTML requests by user agent against HTML requests carrying that referrer. |
| Platform documentation | Google, 10 Dec 2025 | Query fan-out, and no special optimisation | Both AI Overviews and AI Mode "may use a query fan-out technique". Eligibility is snippet eligibility. No special structured data is required. |
| Platform documentation | Google, crawler documentation | Google-Extended does not affect Search inclusion | It governs training and grounding in Gemini Apps and Vertex AI, and "does not impact a site's inclusion in Google Search nor is it used as a ranking signal". |
| Platform documentation | OpenAI, bot documentation | Three crawler classes with different consequences | OAI-SearchBot for surfacing, GPTBot for training, ChatGPT-User user-initiated. "Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers." |
| Platform documentation | OpenAI, shopping documentation | The only published ranking criteria in the category | Shopping results selected independently of partnerships, ranked on "availability, price, quality, and whether they are the maker or primary seller of that item". |
| Platform documentation | Anthropic, web search tool documentation | Citations are structural, and the citable universe is configurable | Citations always enabled with URL, title and cited text; retrieval constrainable by allowed or blocked domain lists. |
| Platform documentation | Microsoft, 10 Feb and 16 Jun 2026 | The only first-party AI citation reporting | Bing Webmaster Tools AI Performance: total citations, grounding queries, citation share, intents, topics and comparison. |
| Platform documentation | Cloudflare, 24 Sep 2025 | Purpose-scoped crawl signals on 3.8 million domains | Search, AI input and AI training expressed separately, with the managed default leaving grounding unstated. |
| Academic | Aggarwal et al., KDD 2024 | Evidence-adding tactics work, packaging does not | Quotations at 40.9 %, statistics at 32.8 %, source citation at 27.5 % relative improvement; keyword stuffing at minus 8.3 %. GEO-bench, 10,000 queries, 25 domains. |
| Academic | Puerto et al., NeurIPS 2025 | Conversational optimisation is often ineffective and congested | Most dedicated methods damaged ranking; traditional techniques outperformed them; gains fall as adopters rise. |
| Academic | Kim et al., Yonsei, Feb 2026, preprint | Citation optimisation can cost retrieval | Body-text rewrites dropped retrieval rank 4.54 positions and hit rate 9 %; structural work raised hit rate 22 %. 171,003 documents, 2,700 queries. |
| Academic | Schulte et al., St. Gallen, 10 Apr 2026, preprint | Sampling thresholds for a usable estimate | 65 % of cited sources change daily; seven to eight runs per prompt per day, ten-day minimum, twenty-one to twenty-eight day windows. Citation concentration at a Gini of 0.715. |
| Academic | Vishwakarma et al., May 2026, preprint | Relevance, position, price and recency dominate; formatting does not | 252,000 paired comparisons, six models, eighteen factors, brand anonymisation and counterbalanced ordering. |
| Academic | Martinez, Sciences Po, Jul 2026, preprint | Very low confidence that citations predict revenue | Critical survey of 45 studies, with graded confidence and a seven-component decomposition of visibility. Cross-engine URL overlap at Jaccard 0.11 to 0.18. |
| Academic | Liu et al., Stanford, TACL 2024 | Position inside the context changes what gets used | Information in the middle of a long retrieved context is used measurably less than information at either end. |
| Academic | Liu, Zhang and Liang, Stanford, EMNLP 2023 | Half of generated sentences are not fully supported | 51.5 % citation recall and 74.5 % citation precision across four generative search engines. |
| Academic | Chan, NBER, Jun 2026, working paper | The collapse is a market-design failure, not a hallucination problem | A platform that underinternalises future content reproduction retains too little referral traffic, destabilising the ecosystem even with accurate answers. |
| Academic | Khosravi and Yoganarasimhan, Feb 2026, preprint | A causal estimate of minus 15 % traffic | Difference-in-differences on Google's staggered geographic rollout against Wikipedia language editions, 161,382 matched article-language pairs. |
| Academic | Sabbah and Acar, HBR, May 2026 | Conventional persuasion does not transfer to shopping agents | Only star ratings consistently increased choice; scarcity cues, timers, strike-through pricing and bundles were unreliable or negative. |
| Analyst and survey | The CMO Survey, Feb 2026 | 41.5 % of companies already do this work | Duke Fuqua with the American Marketing Association and Deloitte, 308 United States marketing leaders, 97 % vice-president or above, fielded January 2026. |
| Analyst and survey | Gartner, 20 Jan 2026 | Only a third say assistants rival search engines | Two consumer surveys, 377 and 365 respondents. 31 % spend more time searching when summaries appear; 31 % consider more product options. |
| Analyst and survey | Comscore, 2 Jun 2026 | 76 billion United States desktop searches, up 10 % on two years | First quarter 2026 against first quarter 2024. The clearest measured answer to the 25 % volume-decline forecast of February 2024. |
| Analyst and survey | Accenture, 3 Jun 2026 | 71 % expect AI to influence half their spending within a year | 25,590 consumers across sixteen countries. Only 9 % are comfortable with fully autonomous purchasing. |
| Analyst and survey | G2, 15 Apr 2026 | Half of software buyers start with an assistant | 1,076 decision-makers, March 2026. 69 % chose a different vendor because of AI guidance; 33 % bought from a vendor they had not heard of. |
| Analyst and survey | Reuters Institute, 12 Jan 2026 | Organic search traffic to 2,500 sites down 33 % | Chartbeat data, November 2024 to November 2025, United States at 38 %. Survey of 280 media leaders across 51 countries. |
| Vendor measurement | Ahrefs, 11 Aug 2025 | Around 12 % overlap between AI citations and the organic top ten | 15,000 long-tail queries across four assistants against Google and Bing. Perplexity highest at 28.6 %. |
| Vendor measurement | Ahrefs, 11 May 2026 | Adding structured data produced no citation uplift | 1,885 pages against 4,000 matched controls, difference-in-differences. Caveat: the sample was already heavily cited pages. |
| Vendor measurement | Ahrefs, 15 Jun 2026 | llms.txt is published widely and read almost never | Server logs from 137,210 domains: 28 % published a file, around 1,100 received any traffic, and no bot requested a file that did not exist. |
| Vendor measurement | Seer Interactive, 24 Apr 2026 | Being cited is worth 120 % more clicks, and still less than no summary | 53 brands, 5.47 million queries, 2.43 billion impressions. Uncited brands lost 67 % of organic click-through over 2025; the series rebounded from December 2025. |
| Vendor measurement | Seer Interactive, 2026 | Freshness comes from updating, not publishing | 7,683 dated pages, 47,097 citations: 75 % updated within a year; among pages with both dates, 72 % fresh by update against 42 % by publication. |
| Vendor measurement | Semrush, 2025 and Jul 2026 | 0.14 % of visits, and 53.7 % of categories unowned | Channel mix across 50,000 sites and seventeen industries; topic ownership across 50,000 brands, 1,094 categories and 600,000 citations, January to June 2026. |
| Vendor measurement | Adobe, 16 Apr 2026 | AI-referred retail traffic converts 42 % better | More than one trillion visits to United States retail sites and a survey of 5,000 respondents. Product pages score 66 out of 100 on machine readability. |
| Vendor measurement | BrightEdge, 12 Feb 2026 | Citation inheritance from rank varies almost two to one by industry | Healthcare 24.0 %, business technology 22.6 %, e-commerce 13.4 %, finance 11.3 %, restaurants 9.3 %. |
| Vendor measurement | Similarweb and Peec AI, 2026 | Citation supply is concentrated and platform-specific | Around 600,000 United States citation events and, separately, 30 million sources. Encyclopaedic and community platforms lead, differently per engine. |
Appendix: the 90-day sequence (as of October 2026)
This is how a team would run its first ninety days, in the order the evidence says the work should happen, as the sequence stood in October 2026. It holds whether or not the team buys anything from Nagent.
Week 1. Establish access, per engine and per crawler class. Audit robots directives, snippet controls, network-level rules and status codes against retrieval, training and user-initiated agents separately. Identify who owns the file and who changed it last. Expect at least one finding nobody in marketing knew about. Fix nothing yet; record the verdict per engine so the baseline is honest.
Weeks 1 to 2. Build the prompt universe from real language. Take phrasing verbatim from sales calls, support tickets and won and lost deal notes before adding anything from a keyword tool. Classify by intent. Weight it. Version it. Resist the urge to load it with brand-research prompts, which will flatter the first report and mislead the second.
Weeks 2 to 4. Measure to the published threshold, not to a budget. Sample at seven to eight runs per prompt per day across the engines and locales that matter, for at least ten days before any statement is made, and hold the first report until a twenty-one day window exists. Annotate any model release inside the window. Report by intent class.
Weeks 3 to 5. Map the corpus and measure the concentration. Establish which sources the relevant engines actually draw on for this category, where the brand and its competitors sit in them, and how concentrated the supply is. Record dependence on any single external platform as a named risk with a number, because that platform's next release is not in anyone's control.
Weeks 4 to 6. Fix rendering and passage structure before writing anything. Confirm that the content a human sees survives for an agent that does not execute client-side code. Check that the claim a company wants cited sits in a coherent, early, retrievable passage. Structural work raises retrieval hit rate materially; evidence-heavy rewriting on its own can lower it.
Weeks 5 to 8. Repair the canonical facts everywhere they appear. Assemble the canonical fact set, compare it against every third-party surface that carries it, and route corrections with named owners. This is the cheapest high-severity work available and it almost never sits in a search plan.
Weeks 6 to 10. Add evidence to the pages that are already in the candidate set. Figures, named experts, quotations, documents and dates, on the pages the corpus map shows are reachable. Update rather than publish: the citation evidence favours maintained pages over new ones by a wide margin. Auditing or rewriting pages no engine reaches is the most common way a quarter is wasted.
Weeks 8 to 12. Set authority levels and start the action log. Decide per action class what may happen without a human, name the approvers including the one outside marketing for access changes, and start recording every action with its expected effect and its measurement plan. The log, not the dashboard, is what will make the next budget conversation possible.
Day 90. Report what moved, what did not, and what was inconclusive. Publish internally with run counts, windows, intervals and model annotations. Include the actions that produced nothing. A first report that contains no inconclusive results is a first report that has not been measured properly.
Frequently asked questions
Do AI Overviews reduce clicks to websites?
Yes, on every independent measure. Pew Research Center tracked 68,879 real Google searches and found people clicked a traditional result on 8 % of visits with an AI summary against 15 % without, and clicked a link inside the summary 1 % of the time. The studies disagree on the size of the fall, not its direction.
Why is a single AI visibility score unreliable?
Answer engines are not deterministic. Roughly 65 % of cited sources change from one day to the next, and repeated runs of the same prompt rarely return the same brand list. A usable estimate needs seven to eight runs per prompt per day, a ten-day minimum and a twenty-one to twenty-eight day window, reported with an interval.
Does blocking AI crawlers remove a site from AI answers?
It depends on the crawler. Platforms run retrieval crawlers, training crawlers and user-initiated fetchers, with different consequences. OpenAI states that sites opted out of OAI-SearchBot will not be shown in ChatGPT search answers, while blocking Google-Extended does not remove a site from AI Overviews. A blanket block usually trades retrieval visibility for modest training control.
Do llms.txt files and schema markup increase AI citations?
Not as a citation lever, on current evidence. Google has said no AI system currently uses llms.txt, and server logs from 137,210 domains showed almost no traffic to the files. A controlled Ahrefs study found no citation uplift from adding structured data, though only on pages already heavily cited, so structure belongs in retrievability work.
What is claim supply?
Claim supply is Nagent's name for the job an organic programme now has: making sure a defined set of claims about a company are present, retrievable, consistent, attributable and current at every stage of the answer pipeline, on the brand's own site and on the third-party sources each engine draws from, with every effect proven statistically.
What is DRIS?
DRIS is Nagent's organic visibility AI coworker, the team lead for search and answer-engine work. It audits first, turns findings into proposals with an evidence grade, an expected effect and a rollback, and then acts only within the authority a person sets for each action class. Six agents cover orchestration, access, corpus, extractability, consistency and reporting.
Does answer-engine visibility produce revenue?
That link is not yet proven. A survey of 45 studies grades confidence that citations predict clicks, conversions or revenue as very low. What is measured is that AI-referred retail visitors converted 42 % better than other traffic in March 2026, and that most categories still have no clear owner in the answer layer.
Sources
- Google Search Central, AI features and your website, updated 10 December 2025 https://developers.google.com/search/docs/appearance/ai-features
- Pew Research Center, 22 July 2025 https://www.pewresearch.org/short-reads/2025/07/22/google-users-are-less-likely-to-click-on-links-when-an-ai-summary-appears-in-the-results/
- Ahrefs, 11 August 2025 https://ahrefs.com/blog/ai-search-overlap/
- Ahrefs via Search Engine Journal, March 2026 https://www.searchenginejournal.com/google-ai-overview-citations-from-top-ranking-pages-drop-sharply/568637/
- Schulte, Bleeker and Kaufmann, arXiv:2604.07585, 10 April 2026 https://arxiv.org/pdf/2604.07585
- SparkToro, 28 January 2026 https://sparktoro.com/blog/new-research-ais-are-highly-inconsistent-when-recommending-brands-or-products-marketers-should-take-care-when-tracking-ai-visibility/
- Adobe Digital Insights, Q2 2026 AI Traffic Report https://business.adobe.com/resources/sdk/.2026-q2-ai-traffic-report/q2-2026-adi-ai-sourced-traffic-insights.pdf
- Semrush, traffic channel mix study, 2025 https://www.semrush.com/blog/traffic-channel-mix-study/
- Google, 20 May 2025 https://blog.google/products/search/google-search-ai-mode-update/
- Google, 18 November 2025 https://blog.google/products-and-platforms/products/search/gemini-3-search-ai-mode/
- Google, Gemini API grounding documentation https://ai.google.dev/gemini-api/docs/google-search
- Financial Times interview, April 2025, reported by Search Engine Land https://searchengineland.com/google-search-boss-ai-overviews-boost-click-quality-454386
- Google, crawler documentation https://developers.google.com/search/docs/crawling-indexing/google-common-crawlers
- OpenAI, bot documentation https://developers.openai.com/api/docs/bots
- Perplexity, bot documentation https://docs.perplexity.ai/guides/bots
- Meta, webmaster crawler documentation https://developers.facebook.com/docs/sharing/webmasters/web-crawlers/
- Similarweb AI Citation Analysis, 2026 https://aisearch.similarweb.com/blog/most-cited-domains-llms/
- Peec AI, 31 March 2026 https://searchengineland.com/ai-search-engines-cite-reddit-youtube-and-linkedin-most-study-473138
- Kumar, arXiv:2606.20065, June 2026, commercial affiliation disclosed https://arxiv.org/abs/2606.20065v1
- Semrush, 230,000 prompts, November 2025 https://www.semrush.com/blog/most-cited-domains-ai/
- Ahrefs, 17 April 2025 https://ahrefs.com/blog/ai-overviews-reduce-clicks/
- Amsive, 16 April 2025 https://www.amsive.com/insights/seo/google-ai-overviews-new-research-reveals-how-to-navigate-click-drop-off/
- SISTRIX, February 2026 data https://ppc.land/ai-overviews-cost-germany-265m-organic-clicks-monthly-sistrix-data-shows/
- Seer Interactive, 24 April 2026 https://www.seerinteractive.com/insights/aio-impact-on-google-ctr-2026-update
- Search Engine Journal analysis, 26 April 2026 https://www.searchenginejournal.com/ai-overview-ctr-fell-61-but-clicks-didnt-collapse/572993/
- SparkToro, 9 June 2026 https://sparktoro.com/blog/in-2026-less-than-one-third-of-google-searches-still-send-a-click/
- Search Engine Journal, 29 April 2021 https://www.searchenginejournal.com/stop-quoting-zero-click-search-studies/403763/
- Reuters Institute, 12 January 2026, survey of 280 leaders across 51 countries https://reutersinstitute.politics.ox.ac.uk/journalism-media-and-technology-trends-and-predictions-2026
- Digital Content Next via Digiday, August 2025 https://digiday.com/media/google-ai-overviews-linked-to-25-drop-in-publisher-referral-traffic-new-data-shows/
- Similarweb via Digiday, 2025 https://digiday.com/media/in-graphic-detail-ai-platforms-are-driving-more-traffic-but-not-enough-to-offset-zero-click-search/
- Alphabet, Q2 2026 https://blog.google/company-news/inside-google/message-ceo/alphabet-earnings-q2-2026/
- Comscore, 2 June 2026 https://www.comscore.com/Insights/Press-Releases/2026/6/Comscore-s-Q1-2026-AI-Intelligence-Report
- Microsoft Q3 FY2026 earnings, reported 30 April 2026 https://www.searchenginejournal.com/microsoft-says-bing-reached-1b-monthly-active-users/573444/
- Gartner, 19 February 2024 https://www.gartner.com/en/newsroom/press-releases/2024-02-19-gartner-predicts-search-engine-volume-will-drop-25-percent-by-2026-due-to-ai-chatbots-and-other-virtual-agents
- Gartner, 20 January 2026, two surveys, n=377 and n=365 https://www.gartner.com/en/newsroom/press-releases/gartner-survey-finds-only-one-third-of-consumers-say-genai-rivals-search-engines-marketers-must-optimize-for-both-ai-driven-and-traditional-search
- Search Off the Record, reported 17 December 2025 https://searchengineland.com/google-danny-sullivan-seo-for-ai-is-still-seo-466368
- Search Off the Record, published 8 January 2026 https://www.seroundtable.com/google-content-bite-sized-chunks-40728.html
- Google, 6 August 2025 https://blog.google/products-and-platforms/products/search/ai-search-driving-more-queries-higher-quality-clicks/
- OpenAI, publishers and developers FAQ https://help.openai.com/en/articles/12627856-publishers-and-developers-faq
- OpenAI, shopping documentation https://help.openai.com/en/articles/11128490-shopping-with-chatgpt-search
- Anthropic, web search tool documentation https://platform.claude.com/docs/en/agents-and-tools/tool-use/web-search-tool
- Microsoft, 10 February 2026 https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview
- Microsoft, 16 June 2026 https://blogs.bing.com/search/June-2026/New-AI-Visibility-Insights-in-Bing-Webmaster-Tools-Intents-Topics-Citation-Share-Compare
- Axios, 26 August 2025 https://www.axios.com/2025/08/26/perplexity-comet-plus-subscription
- Amazon, 18 November 2025 https://www.aboutamazon.com/news/retail/amazon-rufus-ai-assistant-personalized-shopping-features
- Search Engine Roundtable, John Mueller on llms.txt, June 2025 https://www.seroundtable.com/google-ai-llms-txt-39607.html
- Ahrefs, 15 June 2026 https://ahrefs.com/blog/llmstxt-study/
- Ahrefs, 11 May 2026 https://ahrefs.com/blog/schema-ai-citations/
- Aggarwal and colleagues, GEO: generative engine optimization, KDD 2024, arXiv:2311.09735 https://arxiv.org/abs/2311.09735
- Vishwakarma and colleagues, arXiv:2605.25517, May 2026, preprint https://arxiv.org/abs/2605.25517
- Kim and colleagues, Yonsei, SAGEO Arena, arXiv:2602.12187, February 2026, preprint https://arxiv.org/html/2602.12187v1
- Puerto and colleagues, C-SEO Bench, NeurIPS 2025, arXiv:2506.11097 https://arxiv.org/abs/2506.11097
- Recommendation dynamics under simultaneous optimisation, arXiv:2606.17443, June 2026, preprint https://arxiv.org/abs/2606.17443
- Martinez, arXiv:2607.14035, July 2026, preprint https://arxiv.org/html/2607.14035v1
- Liu and colleagues, Lost in the Middle, TACL 2024 https://aclanthology.org/2024.tacl-1.9/
- Kamruzzaman and colleagues, University of South Florida, brand bias in language models https://arxiv.org/html/2406.13997v1
- Filandrianos and colleagues, NTUA, persuasion and recommendation, EMNLP 2025 https://arxiv.org/abs/2502.01349
- Sabbah and Acar, HBR, May 2026 https://hbr.org/2026/05/research-traditional-marketing-doesnt-work-on-ai-shopping-agents
- Liu, Zhang and Liang, EMNLP 2023 https://aclanthology.org/2023.findings-emnlp.467/
- Deep research agents, link validity and factual accuracy, arXiv:2605.06635 https://arxiv.org/html/2605.06635v1
- Chan, NBER w35344, June 2026 https://www.nber.org/papers/w35344
- Khosravi and Yoganarasimhan, arXiv:2602.18455 https://arxiv.org/abs/2602.18455
- The CMO Survey, February 2026 edition https://cmosurvey.org/wp-content/uploads/2026/04/The_CMO_Survey-Highlights_and_Insights_Report-2026.pdf
- Digiday, 2026 https://digiday.com/marketing/marketers-shift-growing-shares-of-search-spending-to-geo/
- Gartner, 11 May 2026 https://www.gartner.com/en/newsroom/press-releases/2026-05-11-gartner-2026-cmo-spend-survey-finds-cmos-allocate-15-point-3-percent-of-marketing-budgets-to-ai-but-only-30-percent-are-ready-to-scale-ai-capabilities
- Accenture, 3 June 2026 https://www.accenture.com/us-en/insights/consulting/talk-my-ai-agent
- McKinsey, 2 March 2026 https://www.mckinsey.com/capabilities/quantumblack/our-insights/europes-agentic-commerce-moment-decision-influence-is-here-execution-is-coming
- BCG, 21 January 2026 https://www.bcg.com/publications/2026/introducing-the-consumer-ai-disruption-index
- Bain and Dynata, 19 February 2025 https://www.bain.com/insights/goodbye-clicks-hello-ai-zero-click-search-redefines-marketing/
- G2, 15 April 2026 https://www.prnewswire.com/news-releases/new-g2-research-half-of-b2b-software-buyers-now-start-their-research-with-ai-chatbots-302742807.html
- TrustRadius, 15 July 2026 https://www.prnewswire.com/news-releases/trustradius-2026-b2b-buying-disconnect-report-reveals-ai-has-changed-how-buyers-research-but-not-what-they-trust-302825792.html
- Gartner, Market Guide for Answer Engine Visibility Tools, 9 March 2026 (paywalled) https://www.gartner.com/en/documents/7559273
- DataForSEO Labs data, August 2026 https://www.serp-secrets.com/blog/aeo-or-geo-who-decides
- Omnicom, 28 April 2026 https://s201.q4cdn.com/282904488/files/doc_financials/2026/q1/OMC-1Q26-Transcript-04-28-26.pdf
- VideoWeek, 26 February 2026 https://videoweek.com/2026/02/26/wpp-rejects-holdco-label-in-new-ai-driven-strategy/
- TechCrunch, 19 November 2025 https://techcrunch.com/2025/11/19/adobe-to-buy-semrush-for-1-9-billion/
- Adobe, 28 April 2026 https://news.adobe.com/news/2026/04/adobe-completes-semrush-acquisition
- Search Engine Land, 16 October 2024 https://searchengineland.com/semrush-acquires-search-engine-land-447555
- Similarweb, 28 July 2025 https://ir.similarweb.com/news-events/press-releases/detail/125/similarweb-launches-genai-intelligence-toolkit-tracking-visibility-and-traffic-across-ai-chatbots
- BusinessWire, 1 April 2026 https://www.businesswire.com/news/home/20260401188763/en/Conductor-Delivers-Next-Generation-AI-Search-Performance-Introducing-the-Industrys-Only-System-of-Record-for-AEO
- BusinessWire, 22 October 2025 https://www.businesswire.com/news/home/20251022669493/en/Botify-Launches-AI-Visibility-to-Eliminate-Blind-Spots-for-Brands
- Yext, 3 March 2025 https://investors.yext.com/news-events/press-releases/detail/365/introducing-yext-scout-an-ai-search-competitive
- Gigawatt Group, 2026 survey of agency pricing https://gigawattgroup.com/insights/geo-aeo-pricing-models-what-agencies-charge-in-2026/
- Digiday, 1 May 2026 https://digiday.com/marketing/marketers-question-expensive-ai-visibility-tools-as-inconsistent-results-fuel-skepticism/
- Brainlabs, 20 April 2026 https://www.brainlabsdigital.com/ai-visibility-data-accuracy/
- Search Engine Land, 8 June 2026 https://searchengineland.com/ai-share-of-voice-metrics-that-matter-more-479611
- BrightEdge, 12 February 2026 https://www.brightedge.com/resources/weekly-ai-search-insights/ai-overviews-one-year-presence-size-citing
- Semrush, 20 July 2026 https://www.semrush.com/blog/chatgpt-topic-authority-study/
- TechCrunch, 15 February 2026 https://www.techcrunch.com/2026/02/15/india-has-100m-weekly-active-chatgpt-users-sam-altman-says/
- Google India, 8 October 2025 https://blog.google/intl/en-in/products/supercharging-search-for-india-new-languages-in-ai-mode-search-live-debuts/
- Jio, Google AI subscription offer https://www.jio.com/google-gemini-offer/
- Bain and Flipkart, 8 April 2026 https://www.bain.com/insights/how-india-shops-online-2026/
- Amazon India, shopping assistant launch in India https://www.aboutamazon.in/news/retail/rufus-ai-shopping-assistant-launch-in-india
- Meta and Ipsos, 13 May 2026, vendor-commissioned https://about.fb.com/news/2026/05/from-scroll-to-chat-to-cart-trends-reshaping-how-india-shops/
- CMA, 17 June 2026 https://www.gov.uk/government/news/further-cma-action-to-secure-a-fairer-deal-for-businesses-and-improve-google-search-services-in-uk
- Adobe, 16 April 2026 https://business.adobe.com/blog/ai-traffic-surge-retail-sites-not-machine-readable
- Adobe Digital Insights, 10 January 2026 https://business.adobe.com/blog/ai-driven-traffic-surges-across-industries
- Digital Commerce 360, 17 June 2026 https://www.digitalcommerce360.com/2026/06/17/adobe-ai-referred-traffic-to-retail-sites-doubles-in-a-year/
- Salesforce, January 2026 https://www.salesforce.com/news/stories/2025-holiday-shopping-data/
- Salesforce, 20 July 2026 https://www.salesforce.com/blog/holiday-retail-predictions-2026/
- OpenAI, 29 September 2025 https://openai.com/index/buy-it-in-chatgpt/
- Digital Commerce 360, 6 March 2026 https://www.digitalcommerce360.com/2026/03/06/openai-shifts-checkout-plans-agentic-commerce-strategy/
- Google Cloud, 16 September 2025 https://cloud.google.com/blog/products/ai-machine-learning/announcing-agents-to-payments-ap2-protocol
- Google, 11 January 2026 https://blog.google/products/ads-commerce/agentic-commerce-ai-tools-protocol-retailers-platforms/
- Cloudflare, 2025 https://blog.cloudflare.com/crawlers-click-ai-bots-training/
- Cloudflare, 1 July 2025 https://blog.cloudflare.com/content-independence-day-no-ai-crawl-without-compensation/
- Cloudflare, 24 September 2025 https://blog.cloudflare.com/content-signals-policy/
- Seer Interactive, content recency and AI visibility study, 2026 https://www.seerinteractive.com/insights/study-content-recencys-impact-on-ai-visibility-in-2026
Cite this page
Plain:
Nagent AI. The Answer Layer. Nagent thesis series, no. 3. 2026. https://nagent.ai/artefacts/the-answer-layer
BibTeX:
@misc{nagent2026theanswerlayer,
author = {Nagent AI},
title = {The Answer Layer},
series = {Nagent thesis series},
year = {2026},
url = {https://nagent.ai/artefacts/the-answer-layer},
note = {Published 2026-10-04, updated 2026-10-04}
}The direct answer at the top of this page is written to be quoted as one sentence with this URL as its source.
About Nagent
Nagent is Multiplayer AI for end to end growth: a team of AI coworkers and your own people, working together in one workspace across marketing, sales and customer experience. Three things make it different. You approve the AI coworkers' work until they earn the right to act on their own. They carry the work all the way to pipeline and customers, not just content. And where your plan includes it, a Nagent marketer joins your team and owns the number with you. Founded in Bengaluru in 2024, Nagent is an Anthropic partner, holds four filed patents on orchestration and memory, and deploys in the customer's private cloud.
Published 4 October 2026. All rights reserved. Quote with attribution to Nagent AI and a link to this page.
