TL;DR
AI Overviews now slash click-through rates on the top organic result by 58% compared to similar queries without one—up from 34.5% just a year earlier, and Pew found organic clicks occur in only 8% of searches with an AI summary versus 15% without. The article's framework scores every SEO candidate on Impact, Confidence, and Effort (1–5 each) using the formula (Impact × Confidence) / Effort, adding Reach for sitewide changes. The single most common mistake is chasing the highest-volume keyword, because volume alone is a weak proxy for traffic when SERP features compress CTR.
The verdict: pull candidates from all sources, filter out low-value items, then score everything with the same formula and work the list top-down—starting with pages already ranking in positions 4–20, which require the least new effort for the most measurable gain. Re-score monthly because rankings, competitors, and AI Overview penetration shift the inputs.
The short answer
Direct answer: Score every SEO candidate — a keyword, a technical fix, a content refresh — on expected impact against required effort, using the same formula for every item so scores are comparable. Work the list top-down, but treat pages ranking in positions 4–20 ("striking distance") as your default starting queue, because they need the least new work to produce the most measurable gain. Re-score monthly, because rankings, competitor moves, and search-result formats (AI Overviews chief among them) all shift the inputs.
That's the whole idea. The rest of this piece is the mechanics: what to measure, how to score it without fooling yourself, and where the model breaks down.
Why "just do the highest-volume keyword" doesn't work anymore
The instinct to chase the biggest search-volume keyword is the single most common prioritization mistake, and it's gotten worse, not better, as Google's result page has changed. Click-through rate for a given position isn't fixed — it depends on what else is on the page. First Page Sage's aggregated CTR data puts the top three organic results at roughly 68.7% of all clicks on a page, but that share is compressing. Ahrefs analyzed 300,000 keywords and found that when an AI Overview appears, click-through rate on the top-ranking organic result fell 58% compared to similar queries without one — up from a 34.5% decline the firm measured a year earlier, meaning the effect is accelerating, not stabilizing. Pew Research Center independently found the same pattern from actual browsing data: people clicked an organic result in just 8% of searches that produced an AI summary, versus 15% for searches that didn't — nearly half the rate. Separately, GrowthSRC's study of 200,000+ keywords found position-1 CTR down roughly 32% year over year, with clicks redistributing to positions 6–10 as more searchers scroll past AI-generated summaries.
None of that means keyword volume is meaningless. It means volume alone is a weak proxy for traffic, and a prioritization model that only asks "how many people search this?" will misrank work. You need a model that also asks how much of that volume will actually reach your site, and how much effort it costs to get there.
Definitions
A shared vocabulary keeps scoring consistent across a team:
- Impact — the estimated effect on organic traffic, conversions, or revenue if the opportunity is executed. Not the ceiling case; a realistic case.
- Effort — the time and resources required: dev hours, writing hours, design, review, and any cross-team dependency.
- Confidence — how certain you are that your impact estimate is right, given data quality and precedent.
- Reach (optional, borrowed from RICE) — how many pages, users, or sessions the change touches, when a single fix applies broadly (e.g., a sitewide template change).
- Striking distance — an opportunity, usually a keyword, where a page already ranks in roughly positions 4–20. The term has no single official definition; SEOTesting and others treat 4–20 as the common working range, with judgment applied at the edges.
The scoring framework
Direct answer: Two frameworks dominate the "how do we compare unlike things" problem, and both were built outside SEO before being adapted into it.
ICE was created by Sean Ellis at GrowthHackers as a lightweight way to rank growth experiments; it scores Impact, Confidence, and Ease (the inverse of effort) — see Growth Method's writeup of the ICE framework's origin. NOVOS applies ICE directly to technical SEO tickets, scoring each ticket 1–5 on all three factors and computing (Impact × Confidence) / Effort.
RICE, developed by Sean McBride on Intercom's growth team and published on Intercom's own blog, adds Reach: (Reach × Impact × Confidence) / Effort. Reach matters most when a single change affects many pages or users at once — a sitewide schema fix, for instance — where ICE alone would understate its value.
A related model worth knowing is PIE (Potential, Importance, Ease), which Chris Goward built at WiderFunnel for prioritizing conversion-rate tests before ICE or RICE existed in most teams' vocabulary — see Growth Method's summary of the PIE framework. PIE's "Importance" is closer to strategic/business value than ICE's "Confidence," which is closer to statistical belief in the estimate. Some SEO teams use PIE-style scoring to decide which pages deserve attention, then ICE to decide which specific fixes on those pages come first.
Adapted for SEO, a working scoring table looks like this:
| Factor | What it measures | 1 (low) | 5 (high) | SEO-specific inputs |
|---|---|---|---|---|
| Impact | Realistic traffic/revenue effect | <5% traffic lift | >20% traffic lift on the affected page(s) | Search volume, current position, CTR-by-position benchmark, page value/conversion rate |
| Effort | Time and coordination cost | <5 hours, one person | >40 hours or multi-team dependency | Dev hours, writing hours, review cycles, external approvals |
| Confidence | Certainty in the impact estimate | Speculative, no precedent | Backed by GSC data or a prior comparable win | Data recency, historical precedent, algorithm volatility in that niche |
| Reach (optional) | How many pages/sessions the change touches | Single page | Sitewide template or pattern | Affected page count, share of total organic sessions |
Formula: Score = (Impact × Confidence × Reach) / Effort — drop Reach and you're back to plain ICE. Score everything on the same 1–5 scale, apply the same formula every time, and rank the resulting list. The number itself is meaningless outside your own backlog; a score of 12 only tells you it beats a score of 6 in the same list.
The process, step by step
- Collect candidates from every source, not just keyword research. Pull from Google Search Console (queries with high impressions, low CTR, or a position that dropped), technical crawls, content audits, competitor gap analysis, and business requests. A backlog built only from one source misses whole categories of opportunity.
- Filter before you score. Discard anything that isn't technically feasible, doesn't align with a real business goal, or duplicates work already in flight. Scoring is expensive in team time; don't spend it on non-starters.
- Score Impact using position-adjusted CTR, not raw volume. Estimate the click delta between current position and target position using a CTR-by-position curve, and be conservative — First Page Sage's benchmarks put position 1 at roughly 28–40% CTR and position 3 around 10%, but AI Overview presence can cut that in half on informational queries, per the Ahrefs and Pew data above. Check whether the query in question tends to trigger an AI Overview before you bank on a big lift.
- Score Effort in the smallest unit you can estimate honestly. "Improve page speed" isn't scoreable; "compress hero images (3 hours), defer non-critical JS (5 hours)" is. Include the dependency tax — anything that needs a developer, designer, or legal sign-off costs more than the raw hours suggest.
- Score Confidence against evidence quality, not optimism. A GSC-verified striking-distance keyword with clear intent match is high confidence. A brand-new content play based on a hunch about a trend is low confidence — that's fine, it just shouldn't rank above sure things without a reason.
- Compute the score and rank the backlog. Sort descending. Then apply judgment: are there sequencing dependencies (fix the technical issue before the content update)? Are there low-scoring quick wins worth doing anyway to build momentum?
- Assign one owner per opportunity and set a review cadence — weekly for status, monthly or quarterly to re-score the backlog as data changes.
- Track predicted impact against actual impact after execution. This is the step most teams skip, and it's the one that makes the next round of scoring more accurate.
Where striking-distance keywords fit
Positions 4–20 are the highest-confidence, lowest-effort segment of most backlogs, because the page is already indexed, already relevant enough to rank, and the fix is usually refinement rather than a new asset. Scale Growth Digital breaks the range into sub-tiers: positions 4–7 are visible but outclicked by the top three and often need refinement, not a rewrite; 8–10 is the page-one threshold, where a small gain moves a page from occasionally seen to consistently clicked; 11–15 rarely gets scrolled to, usually losing on depth or authority rather than relevance; and 16–20 is worth pursuing mainly for high-volume or high-value terms, since it takes longer to move.
Pull this list straight from Search Console — it's free, and SEOTesting's guide to striking-distance keywords walks through filtering by average position and impressions to find it. Weight by monthly volume and by impressions-with-low-CTR (a signal the query is being seen but not clicked), not by position alone. Scale Growth Digital also notes that on-page changes to striking-distance pages tend to show movement within 2–4 weeks and stabilize over 6–8 — useful for setting a review cadence and not calling a change a failure at day 10.
Technical opportunities — crawl errors, broken internal links, missing structured data — should go through the same scoring model, not a separate track. Ahrefs' own guidance on finding SEO opportunities points at internal-link opportunity reports and site audits as sources for exactly this kind of candidate; the framework treats a sitewide fix (high Reach, moderate Effort) the same way it treats a single striking-distance page (high Confidence, low Effort) — by score, not by category.
Limitations
Be honest about what this framework does and doesn't do:
- Impact estimates are inherently uncertain. CTR-by-position benchmarks are aggregates across many sites and queries; your actual click-through will vary by brand strength, SERP features present, and query intent. Treat every impact number as a directional estimate, not a forecast.
- Effort estimates vary by team. The same technical fix can be two hours for one engineering team and two weeks for another, depending on codebase, deploy process, and who else needs to sign off. Scores are only comparable within your own team's estimates — never borrow another company's effort scores.
- The scoring model doesn't replace judgment. Dependencies, sequencing, morale-building quick wins, and non-quantifiable strategic bets (entering a new market, defending a competitive threat) all belong in the decision even when they don't score highest. Use the ranked list as an input to a conversation, not a verdict.
- AI Overviews and other SERP features change the Impact side of the equation faster than most teams update their CTR assumptions. A keyword that would have justified real investment eighteen months ago may now return a much smaller click pool. Re-check your CTR assumptions at least as often as you re-score the backlog.
- Confidence is easy to overstate. Historical precedent from one page or category doesn't guarantee the same result elsewhere; treat "we saw this work once" as moderate, not high, confidence.
Where nqzai fits
nqzai's SEO module builds the striking-distance and opportunity backlog automatically: it pulls current rankings, search performance, and site structure, estimates position-adjusted impact for each candidate keyword or fix, and ranks the resulting list using the same impact-versus-effort logic described above rather than sorting by raw search volume. It surfaces the reasoning behind each score — current position, estimated CTR delta, and what the fix would involve — so a human can check the estimate before committing engineering or writing time to it. It doesn't decide business priorities or replace the judgment calls in step 6 above; it removes the manual work of pulling GSC exports, calculating CTR curves, and re-scoring the backlog every time rankings shift.
FAQ
Do I need RICE, ICE, or PIE — which one should I actually use?
Start with ICE (or its SEO variant) for ranking individual tasks; it's the simplest to score consistently. Add Reach (making it RICE) once you're comparing sitewide technical fixes against single-page content work, where "how many pages does this touch" changes the calculus. PIE is more useful earlier, for deciding which pages or sections deserve attention before you get to task-level scoring.
Are striking-distance keywords always the highest priority?
Usually the highest-confidence, lowest-effort segment, not automatically the highest-impact one. A position-15 keyword with 40 monthly searches will score lower than a genuinely new content opportunity targeting a 5,000-search term, even though the striking-distance keyword is "closer." Score both the same way and let the number decide.
How much should AI Overviews change my prioritization?
Check whether your target query tends to trigger an AI Overview before banking on a CTR estimate from pre-2024 benchmarks. For queries where it does, discount the expected click volume — the Ahrefs and Pew data above both suggest a meaningful cut, not a marginal one.
What if the team disagrees on a score?
Document the disagreement and the reasoning, then let the SEO or growth lead make the call. The point of scoring isn't perfect consensus; it's making the reasoning visible so disagreements are about assumptions, not vibes.
How often should the backlog be re-scored?
Monthly at minimum for active items; quarterly for a full re-evaluation of the scoring model itself, including whether your CTR assumptions still hold given SERP changes.
Does this replace keyword research or technical audits?
No — it's the layer that sits on top of them. You still need the raw candidate list from GSC, crawls, and competitor analysis; the framework is how you decide which of those candidates to act on first.