TL;DR

94% of links cited by AI answer engines come from non-paid media, and 82% from earned coverage — not brand-owned marketing copy. Google's March 2024 core update introduced a site-wide scaled content abuse classifier that can suppress rankings for an entire domain if it contains a large share of low-value pages, regardless of whether humans or AI wrote them. Academic research on generative engine optimization found keyword stuffing had little effect, while adding credible statistics and citations produced measurable gains in visibility.

The bottom line: stop publishing isolated pages at volume; instead build connected topic clusters with clear entity definitions, evidence density, and depth — AI search rewards topical authority, not page count.

Publishing 200 articles a month used to be a defensible SEO strategy. It is no longer a defensible AI-search strategy, and the two are increasingly not the same game. Google's automated content-quality systems and the retrieval layers behind AI Overviews, ChatGPT, and Perplexity are all converging on the same conclusion: volume without connection is noise, and noise gets filtered before it gets cited.

This isn't a hunch. It's documented in the platforms' own guidance and in independent research on how generative engines select sources. The practical implication for B2B teams chasing AI-search visibility is that the unit of competition has changed — it's no longer the page, it's the topic.

Why the content-mill playbook is losing

Google formalized this shift with its March 2024 core update, which folded "helpful content" signals into the core ranking algorithm and introduced a specific spam classification for what it calls scaled content abuse — "many pages generated for the primary purpose of manipulating search rankings and not helping users," regardless of whether the pages were written by people, AI, or some mix of both (Google Search Central, spam policies; Google Search Central Blog, March 2024 core update). Crucially, this is described as a site-wide signal: Google's own documentation frames it as a classifier that scans the overall content of a domain, and a large share of low-value pages can suppress rankings for the genuinely good pages sitting right next to them.

That site-wide framing matters more for AI search than it did for classic SEO, because generative engines aren't just picking a single best-matching URL — they're running a retrieve-then-generate process that pulls passages from whichever sources look reliable enough to ground an answer. Google has also been explicit that AI-generated content isn't penalized for being AI-generated; what gets filtered is content produced "primarily to manipulate rankings" rather than to help a reader, and it applies the same standard whether a human or a model wrote it (Google Search Central Blog, guidance on AI-generated content). A content mill publishing dozens of near-duplicate "X vs Y" pages a week is exactly the pattern this system was built to catch.

What actually earns citations in AI answers

Direct answer: The foundational academic work on this space, Aggarwal et al.'s "GEO: Generative Engine Optimization" (KDD 2024), tested which content interventions actually move the needle on generative-engine visibility. The result that mattered most: keyword stuffing — the classic content-mill tactic — showed little to no effect, while adding credible statistics, citations, and direct quotations produced measurable, domain-dependent gains (arXiv:2311.09735). In other words, the signal generative engines respond to is evidentiary density, not keyword coverage.

A newer factorial study running 252,000 trials across six LLMs reinforces this from a different angle: it found that relevance and context position are the primary determinants of whether a source gets cited at all — meaning a page has to already be a strong, well-targeted candidate before any on-page optimization can help (arXiv, "Think Before Writing"). A 2026 survey of the GEO research literature makes a related point about scale: it synthesizes findings that generative engines increasingly weight passage-level completeness and sourcing quality over raw page count (arXiv, "Optimizing Visibility in Generative Engines").

Independent industry analysis points the same direction. Muck Rack's review of citation patterns across AI answer engines found that roughly 94% of the links cited came from non-paid media and about 82% from earned coverage — third-party analysis and commentary, not brand-owned marketing copy pushed out at volume (Muck Rack, "Where LLMs pull from"). And Backlinko's practitioner-facing GEO guide converges on the same operational advice as the academic work: structure, credibility signals, and topical depth outperform publishing cadence (Backlinko, "Generative Engine Optimization (GEO)").

Search Engine Land's ongoing coverage of AI Overviews adds useful context on stakes: AI Overview visibility rose sharply through mid-2025 before Google pulled back coverage on commercial and navigational queries later in the year, which means the set of queries where AI-search visibility actually matters is narrower and more contested than the "cover everything" content-mill model assumes (Search Engine Land, "Google AI Overviews surged in 2025, then pulled back"). Fewer, better-targeted opportunities favor depth over volume even more.

The five things that actually build topical authority

Direct answer: Google's own self-assessment framework for helpful content asks creators whether they're "producing lots of content on many different topics in hopes that some of it might perform well" — explicitly naming that pattern as a warning sign, not a strategy (Google Search Central, "Creating helpful, reliable, people-first content"). Building the opposite of that pattern comes down to five concrete practices.

1. Connected evidence, not isolated pages. A topic cluster where every page links to and reinforces adjacent pages — a pricing methodology piece that cites the benchmark study, which cites the definitions page, which links back to the implementation guide — reads as a coherent body of work to both crawlers and retrieval systems. A content mill's pages are usually islands: each one optimized for a single keyword, none of them referencing each other because they were assigned to different freelancers on different days.

2. Clear entity definition. Before writing subtopic pages, define the core entities precisely — what exactly do you mean by "topical authority," "GEO," "answer engine," in terms consistent across every page. Entity ambiguity (using a term five different ways across five different articles) is one of the clearest tells of assembly-line production, and it makes it harder for a retrieval system to treat your pages as a single reliable source on the concept.

3. Genuinely useful subtopic coverage. The test is whether a subtopic page would exist if no one were tracking search volume for it — does it answer a real question a practitioner has, or is it a keyword variation of a page you already published ("GEO for startups," "GEO for enterprise," "GEO for SaaS," each saying the same three things with different nouns swapped in)? The research above is consistent on this point: statistics, comparisons, and specific evidence earn citations; reworded restatements don't.

4. Expert review. Someone with real domain knowledge needs to read, correct, and stand behind the content before it publishes — not run it through a single AI pass. This is the "experience" and "expertise" half of E-E-A-T, and it's the hardest thing for a content mill to fake at scale, because genuine review time doesn't compress the way generation time does.

5. Ongoing maintenance. Authority decays. Pages with stale data, dead internal links, or outdated claims signal exactly the kind of "unedited AI, no oversight" pattern that both Google's classifier and generative retrieval systems are tuned to discount. A maintenance cadence — quarterly fact and link audits on the core cluster — is table stakes, not a nice-to-have.

Volume model vs. evidence model

DimensionContent-mill approachEvidence-based topic authority
Unit of productionIndividual page, optimized aloneInterlinked cluster, built as a system
Volume targetAs many pages as budget allowsAs many pages as the topic genuinely needs
DifferentiationKeyword/phrase variationDistinct evidence, data, or use case per page
Review processSingle pass, often uneditedNamed expert review before publish
Internal linkingSparse or templatedDeliberate, bidirectional, topic-mapped
Update cadencePublish and abandonScheduled recheck of facts and links
Google's readScaled content abuse candidatePeople-first, site-wide positive signal
AI-engine readLow evidentiary density, rarely citedHigh passage-level credibility, citable

Where to start

Direct answer: If you're auditing an existing content library, the fastest diagnostic is Google's own question: would this page exist if no one was chasing a keyword with it? Pages that fail that test are the ones to consolidate, rewrite with real evidence, or retire — because under a site-wide quality signal, keeping them isn't neutral. They actively drag down the pages you've built properly.

For new topic clusters, start from the entity and the internal link map before a single article is drafted: define the core concept precisely, map the five to ten subtopics a genuinely informed practitioner would need answered, decide who reviews each one, and set a recheck date before publishing the first page. That's a slower start than assigning fifty briefs to a content mill. It's also the version of the work that both search engines and AI answer engines are now explicitly built to reward.

Sources: