Search Query Performance: the four-index diagnosis
The only report Amazon gives you that shows the whole funnel, one query at a time, for your ASIN and the market side by side. It is also the report most commonly read backwards.
Contents — 15 sections
- Why SQP is the spine
- The caveat that inverts diagnoses: the percentage columns are not conversion rates
- Exports and preparation
- Scope and data integrity
- The mandatory splits, before any index is read
- The four indices
- Significance floors and the corroboration doctrine
- Structural diagnosis: did the impressions disappear?
- The delta method for before-and-after measurement
- Judge-by windows and the rollback doctrine
- The Search Protection Register
- Advanced reads
- Conversion levers that actually produce rank
- Returns, NCX and the negative-signal chain
- What this means in practice
Search Query Performance is the only report Amazon gives a seller that shows the entire funnel, one query at a time, for your ASIN and for the whole market side by side. Amazon confirmed That second half is what turns it from a report into an instrument. An absolute click-through rate of 0.4% is uninterpretable on its own. A click-through rate of 0.4% on a query where the market runs at 1.1% is a diagnosis, and it points at one specific surface.
Everything else in the discipline is a proxy. Third-party volume estimates are models built on other models. Rank trackers tell you position without telling you what position bought you. Ad reports tell you about the traffic you paid for. SQP is the market's own arithmetic on demand that actually happened, and it is the spine of every measurement decision in the six-layer stack.
It is also the report most commonly read backwards. The export ships with percentage columns that look exactly like conversion rates and are not conversion rates. Analysts who read them as stage rates reach conclusions that are not merely imprecise but inverted, and then spend a quarter fixing the wrong layer. That single error is the most consequential measurement mistake in Amazon SEO, and it is the reason this page spends its second section on arithmetic before it gets to anything strategic.
What follows is the full measurement system: how to pull and prepare the data, the four splits you apply before reading anything, the four indices and what each one localises, the sample floors below which a number is not evidence, and the register that makes before-and-after measurement possible at all.
Why SQP is the spine
Three properties make this dataset structurally different from every other input available to a seller, and each one closes off a class of error that the rest of the industry lives inside.
Because the market columns exist, every performance question becomes a ratio rather than a judgement. You never have to decide whether a 12% cart-add rate is good. You divide it by the market's cart-add rate on the same query and read the answer.
Absolute rates are uninterpretable. Relative rates name the failing layer. Every index in this system is your stage rate divided by the market's stage rate on the same query, computed from raw counts. Parity is 1.0.
The caveat that inverts diagnoses: the percentage columns are not conversion rates
The SQP export contains share and rate columns expressed as percentages. They are normalised against total query volume, not against your own funnel. Observed Read as click-through rate, cart-add rate or conversion rate, they do not merely mislead at the margins. They routinely point at the opposite layer from the one that is failing.
Delete them from the working file before you begin. Keep the raw counts — your impressions, your clicks, your cart adds, your purchases, and the market total for each — and compute every rate yourself. There is no situation in this methodology where an exported percentage column is the right number to read.
A worked illustration
The numbers below are arbitrary and chosen to make the arithmetic legible. They are not drawn from any account.
Take one query. Across the whole market it generates 100,000 impressions, 5,000 clicks, 1,000 cart adds and 500 purchases. So the market clicks at 5.0% of impressions, adds to cart at 20.0% of clicks, and purchases at 50.0% of cart adds.
Your ASIN on that same query records 4,000 impressions, 320 clicks, 70 cart adds and 30 purchases.
| Stage | What the share column shows | Your real stage rate | Market stage rate | Index |
|---|---|---|---|---|
| Impressions | 4.0% share | — | — | 0.04 impression share |
| Clicks | 6.4% share | 320 ÷ 4,000 = 8.0% | 5.0% | 1.60 |
| Cart adds | 7.0% share | 70 ÷ 320 = 21.9% | 20.0% | 1.09 |
| Purchases | 6.0% share | 30 ÷ 70 = 42.9% | 50.0% | 0.86 |
The share column climbs steadily from 4.0% to 7.0% and then dips slightly. Read as a funnel it looks unremarkable, maybe mildly encouraging. It suggests nothing worth acting on.
The real reading is completely different. Your tile beats the market by 60% on this query. Your page holds a small edge. Your offer is a whisker under the action threshold and worth a price check. And the actual problem, the one costing money, is that you are collecting 4% of the impressions on a query you out-click by a wide margin. This is a findability failure on a query you are demonstrably good at, and the correct instrument is indexation, attribute coverage and rank depth, not a single word of copy.
An operator reading the share columns sees a flat funnel and does nothing, or rewrites bullets because that is the default reflex. An operator computing indices from raw counts sees an under-served winner and goes to work on retrieval. Same export, opposite decision.
The two ways it inverts
Small impression share, strong tile
Every share column reads low because the denominator is total query volume. It looks like a whole-funnel problem. It is a findability problem sitting on top of a healthy funnel, and copy work on it is wasted quarter.
Dominant impression share, weak stages
Every share column reads high because you own most of the query. It looks healthy. Your per-stage rates may all sit below market, meaning you are converting the query worse than the average competitor while the report congratulates you.
Both readings are common. Both are produced by the same normalisation. The SQP index calculator on this site takes the raw counts and returns the four indices directly, if you would rather not maintain the spreadsheet.
Exports and preparation
The dataset has a minimum shape below which the method does not work.
Scope and data integrity
Four properties of this export are not documented in the interface, and each one reverses a conclusion if you miss it. Most analysts make at least two of these errors in their first month with the report.
Market price medians blend pack architectures. A four-pack judged against a query dominated by single-unit purchases reads as badly overpriced when its per-unit price is actually winning. Normalise to per-unit before any pricing decision derived from SQP, and confirm pack-size parity before escalating a price-gate finding to a client.
The mandatory splits, before any index is read
Aggregate SQP hides the business. An account-level CTR index is a blend of traffic types that behave nothing like one another, and the blend is dominated by whichever type carries the volume. Four splits, applied in this order, before a single index is computed.
1 · Branded versus generic
Match brand tokens and their misspell variants, and separate the two sides completely. Destination traffic and discovery traffic are different products. In one portfolio we measured branded queries at roughly 5% of volume carrying about 58% of purchases at around 8.5 times the generic click-through rate. Observed That is one account and not a benchmark, but the shape recurs.
Why it changes the decision: a blended CTR index of 0.9 can be a branded index of 3.2 sitting on top of a generic index of 0.35. The blended number says "roughly fine". The split says the tile is failing everywhere it matters and the brand is carrying the account. Every downstream decision, from image testing to spend allocation, differs by side.
2 · Constraint-modified versus commodity
Split generic queries by whether they carry a modifier signalling a deciding constraint: material or quality, origin, size or fitment, audience, use case. A commodity query names the product. A constraint query names the product plus the thing the shopper will not compromise on.
Why it changes the decision: this is where the "our actual business is 2% of our volume" finding lives. A premium or specification-led product judged against commodity queries reads as broken on both price and click rate. Judged against the constraint queries it is built for, the same product often reads as market-leading. The split decides whether you are fixing a listing or fighting on the wrong battlefield.
3 · Head versus family
Group queries by shared root plus modifier class. Year variants, size variants, colour variants and spec variants of one intent are members of a family, not independent queries. Below-floor queries are read only as families: sum the raw counts across the family, then compute the family's indices.
Why it changes the decision: the wrong unit of analysis manufactures noise. A family member with 300 impressions and two clicks produces a CTR index that swings between 0.2 and 2.4 week to week and means nothing. The family it belongs to, with 14,000 impressions, produces a stable number you can act on.
4 · Price-normal versus price-mismatched
Compare your ASIN price to each query family's purchase median, per-unit normalised. Flag any family where the gap exceeds roughly 1.5 times. Model
Why it changes the decision: a price-mismatched family is not a copy problem and cannot be made into one. At large multiples of what a query actually buys, no image and no title recovers the click rate, because the shoppers filtering that query have already excluded your price band. The correct output is battlefield reallocation, not a listing brief.
Classifying the queries you keep
Not every row in the export is the same kind of object. Tag four classes at ingest, because each is judged by a different rule.
| Class | Signature | How it is judged |
|---|---|---|
| Head and body terms | Human-typed, high volume, competitive | Individually. These are the core of impression-share strategy |
| Family members | Year, size, colour or spec variants of one intent | As a group. Individually they are all below floor |
| AI-branch fingerprints | Five or more words, fully grammatical, prepositions intact, tiny volume — machine-formulated retrievals landing in the report Observed | On trend only, never on volume. Tracked as a set; their combined impression share is the closest available proxy for assistant-layer visibility |
| Tile-injected queries | Exact strings authored by attribute cards in AI shopping surfaces — machine-authored, human-clicked, ordinary to elevated volume | Identified by matching harvested card labels. These are disclosed priority branches, so coverage is mandatory Amazon confirmed |
The four indices
Every index is constructed identically: your stage rate divided by the market's stage rate on the same query, from raw counts. Parity is approximately 1.0. Our action threshold is approximately 0.85, chosen to sit outside normal noise at the volumes we typically operate on. Model
| Index | Formula, from raw counts | A weak reading means | Fix family |
|---|---|---|---|
| Impression share | your impressions ÷ query impressions | Findability. Indexation gap, missing gate attribute, or insufficient rank depth | Indexation check, then attribute census, then coverage expansion. Never copy. |
| CTR index | (your clicks ÷ your impressions) ÷ (market clicks ÷ market impressions) | The tile. Main image, item name legibility, price at first glance, star rating, badge | Image test first, it has the largest effect size. Then name clarity. Then price position against the click median |
| Cart-add index | (your cart adds ÷ your clicks) ÷ (market cart adds ÷ market clicks) | The page. Unanswered objections, thin evidence, missing specification | Bullets and image panels, A+ FAQ modules, specification density |
| Purchase index | (your purchases ÷ your cart adds) ÷ (market purchases ÷ market cart adds) | The offer. Closing price, delivery promise, coupon environment, stock reliability | Price gate and delivery depth. Copy cannot fix this. |
Four numbers, four different fixes, and on any given query three of the four are wasted effort. The default industry behaviour, which is to rewrite copy whenever performance dips, is correct roughly a quarter of the time by chance.
Compute all four before choosing an instrument, then fix the upstream-most failing index first. A failing tile starves every layer below it of the data required to evaluate them. You cannot read a cart-add index honestly on a query where you only collect 40 clicks a week because the tile is losing.
The signature reads
Four patterns recur often enough to be recognised on sight.
Significance floors and the corroboration doctrine
No standard for this exists in circulation, so these are ours. They are working floors calibrated to our portfolio: tight enough to exclude noise, loose enough to keep most of the query set actionable. Model
| Read | Minimum sample | Below the floor |
|---|---|---|
| CTR index | ≈2,500 of your own impressions on the query | Roll the query into its family and read the family |
| Cart-add and purchase index | ≈100 of your own clicks on the query | Roll up to family. Never act on a single query |
| Trend claim | 3 consecutive periods | One period is a data point, not a trend |
| Price comparison | Same price tier (±1), pack-size normalised | The comparison is invalid, not merely noisy |
The last row is different in kind from the others. A below-floor CTR index is a weak signal. A price comparison across pack architectures is not a weak signal, it is an arithmetic error dressed as a finding.
A below-floor signal becomes evidence only when an independent signal points the same way: a price gap, a family-level pattern, a delivery differential, a return reason code. Two signals in one direction is a diagnosis. One below-floor number is an anecdote, however confidently it is presented in a meeting.
This matters most when the number is dramatic. A query where your purchase index reads 0.18 on nine clicks will get attention in any review. It should get none, unless something else in the data agrees with it.
How exposed is your catalogue on this?
Twenty checks across the five layers, scored 0–100, returning a ranked issue list by severity — the same audit we run on client accounts. Free, no signup to see your result, and it runs entirely in your browser.
Score my listings →Nothing you answer is transmitted or stored. The written report and the six working templates are the optional email step afterwards.
Structural diagnosis: did the impressions disappear?
Before any four-index decomposition, ask one binary question of the affected queries. It sorts the problem into two classes that share no fixes at all.
Yes, impressions went to zero or near zero
This is structural, not performance. Something removed your eligibility. Work the causes in this order, because this is roughly their frequency order in practice. Observed
Spending into a structural loss buys you paid impressions on a query you are no longer organically eligible for. It masks the problem, inflates the cost of discovering it, and produces a register entry that reads as recovery when nothing recovered.
A query sitting at zero after a structural edit is a candidate orphan until proven otherwise. Confirm the market demand still exists, re-home the term according to the placement hierarchy, and then verify indexation directly rather than assuming the edit took.
No, impressions held but position is worse
This is performance, and the four-index decomposition applies. Identify the stage that fell, fix that stage and only that stage, one variable per measurement window. If rank moved without impressions moving, start with rank drop diagnosis and work the causes in frequency order.
The delta method for before-and-after measurement
Most claimed Amazon SEO results are unfalsifiable: no baseline, no window definition, no statement of what else changed in the account during the period. The delta method is the minimum procedure that produces a verdict rather than a story.
Judge-by windows and the rollback doctrine
Amazon's systems settle at different speeds at different layers, and reading a layer before it has settled is the second most common measurement error after the percentage columns. These are the windows we hold to. Model
| Window | What becomes readable |
|---|---|
| +48 hours | Publication acceptance. Attribute re-pull states — did the change actually persist, or did it silently revert |
| +1 to 2 weeks | Indexation settles. Re-check every pre-export query and run the orphan sweep |
| +2 to 4 weeks | Click-through. Tile and item-name effects become visible |
| +4 to 6 weeks | Impression share and family coverage. This is the verdict window. Judge the change here |
| 2 to 6 months | Assistant-layer effects. Answers sourcing from the listing, review-evidence timelines, then compounding |
The five requirements for any change you ship
Revert only on confirmed indexation loss that cannot be re-homed. A week-one wobble is indexation settling, and a single-keyword dip is noise until the family agrees with it. When you do revert, restore the exact prior version. A partial revert creates an uninterpretable third state that is neither the control nor the test, and it destroys the window for both.
The Search Protection Register
Every rule above depends on a baseline existing. "Archive the export" has proven too vague to be executed consistently, so the baseline is specified as an artefact with a name, a shape and a cadence.
One sheet per brand. One row per protected query or query cluster. Weekly refresh. Append-only history, meaning last week's row is never overwritten, only added to. Built before any title, copy or coverage change ships.
| Column | Why it is there |
|---|---|
| Query or cluster | The unit being protected |
| Family | The group it is read with when it falls below floor |
| Classification | Branded, generic-commodity, or generic-constraint. Determines which rules apply to it |
| Relative commercial value | Volume × price × margin. Decides what a loss actually costs, and therefore how much risk is acceptable on that query |
| Eight-week raw counts | Impressions, clicks, cart adds, purchases — yours and the market's. The four-index baseline |
| The four indices | Computed, not exported |
| Price medians | Market click median and purchase median against your click price and purchase price |
| Paid coverage | Yes or no, plus the campaign. Mandatory |
| Rank position | From your ranking data, so organic position and impression share can be read together |
| Supporting content elements | Which title, bullet or attribute you believe carries this query, so you know what you are risking when you edit it |
| Notes | The confound log for this query. Deals, stock events, competitor moves |
An organic loss with paid coverage running reads completely differently from an organic loss without it. With coverage, your SQP click counts include sponsored clicks, so an organic collapse can be entirely masked until the budget changes. Without coverage, the same number is a clean organic signal. This is the field most often omitted from a register, and it is the one that invalidates the most post-mortems.
Intake feeds, in priority order
SQP is the spine. Everything else supplements it, in this order of trust: SQP, then the advertising Prompts report Observed which is sparse with a short lookback and should be treated as directional only, then ad search-term reports, then Brand Analytics top search terms, then Opportunity Explorer, then third-party master lists and competitor ranking sets, then mining your own detail pages for the questions shoppers ask, then review and Q&A language, then return reason codes.
Third-party volume estimates. They are modelled, unreproducible, and blind to conversational and machine-formulated queries. Third-party lists contribute candidates and competitor coverage. They never contribute volumes, and no index in this system is ever computed from them. See keyword strategy for how candidates from these sources are qualified.
What "protected" means, and how the project is judged
Protected does not mean never change. It means do not change casually. A title already performing strongly is the worst place to test speculative conversational language. Validate a pattern on the tail first, then move it to heroes family by family. The 75/125 title system covers how that sequencing works in practice.
Expansion
Did visibility improve across new need states, use cases, occasions, audiences, comparisons and follow-up questions? Measured on branch-family count and impression share on newly covered queries.
Protection
Did established query visibility, traffic, conversion, advertising performance and sales hold or improve? Measured against the register.
A project that gains conversational visibility while losing high-value search demand is not a success. It is expanded but diluted. The only acceptable outcome is expanded and protected: assistant-layer visibility improves while traditional search performance stays within guardrails or strengthens. Expand meaning, protect demand.
Advanced reads
Three uses of the same export that most operators never run, each of which converts an argument into arithmetic.
Delivery-speed conversion analysis
The highest-leverage under-used report in the suite. SQP can be columned to show clicks and purchases segmented by delivery speed. Amazon confirmed That makes the conversion penalty of a slow delivery promise directly measurable on your own ASIN rather than argued from principle, which is what turns an inventory conversation into a ranking conversation.
It circulates widely and we cannot verify it as a multiple. Directionally it is consistent with what we observe; as a number it is unsourced. The method above makes it unnecessary. Measure the spread on the account's own ASIN and quote that instead of somebody else's statistic. Observed
Branded share as a spend-control instrument
Branded advertising spend is the largest reliably recoverable line in most managed accounts, and SQP supplies the only honest way to test it. Revenue is useless as the judging metric because external traffic, seasonality and deals contaminate it. Branded purchase share in SQP is the correct instrument, because it isolates the question actually being asked: when somebody searches the brand name, are you still capturing the purchase?
Run it as a defined test, never as a standing instruction.
The supporting evidence in circulation amounts to isolated accounts in which a full cut cost only a few points of branded share against material savings. Observed That is a strong argument for running the test and a weak argument for the policy. Note also the structural incentive: branded spend flatters reported ACOS, and any party compensated on ACOS has a reason to defend it. State that conflict openly, then let the test decide. More on the interaction between spend and organic position in PPC and organic rank.
Outlier and category-term positioning for event days
On major deal days shopper behaviour shifts from targeted search to broad discovery browsing, and volume floods into short category terms that are unremarkable the rest of the year. An ASIN already ranked and converting on those terms before the event captures a disproportionate share of the surge. An ASIN that starts competing during the event does not, because rank cannot be built inside a 48-hour window.
Conversion levers that actually produce rank
Conversion is the term the ranking system watches most closely, and the levers that move it are mostly not copy. Ordered by observed effect size and speed. Observed
Delivery promise and inventory depth
The standing policy is 90 or more days of cover in the fulfilment network, never below 75, distributed across five warehouses where the programme allows. Model Low stock silently degrades the delivery promise long before it produces a stockout, and can strip the Prime badge across whole regions with no notification. The badge is tied to processing status rather than physical arrival, so units sitting in a receiving queue during peak are not eligible regardless of when they shipped. Build receiving lag into peak planning, not just transit time. Monthly ZIP-panel audit on Tier 1 ASINs, and an immediate audit on any stock dip.
Price level, and price history
Two distinct mechanisms, and the second one is newer.
Model any increase against 90-day and 365-day history before executing. Repeated movement is read as instability and degrades recommendation quality regardless of where you land. For a market-leading ASIN that is already maxed on rank, coverage, variations and marketplaces, a 5 to 10% increase observed for 14 days against BSR, conversion, rank and velocity is usually the highest-leverage remaining move, and it is fully reversible. It is one move, not a sequence of adjustments.
Reviews as launch physics
Review count and text quality gate both conversion and the evidence layer an assistant reads. At launch the decisive tactic is marketplace-stacked Vine: the same hero child ASIN launched across every active marketplace, roughly 30 units to fulfilment per marketplace, enrolled in each, with reviews consolidating to the single ASIN. Target 90 to 150 reviews at launch across a minimum of two marketplaces. Model The change in click-through and conversion during that window is the whole point: it separates a launch that compounds from one that stalls before PPC has anything to amplify. Full sequence in the launch sequence.
The offer envelope
Returns, NCX and the negative-signal chain
The causal chains in this system run both directions, and this is the direction most agencies never model. A return costs far more than the refund line suggests, because it is a ranking event with a delay.
The chain starts with a claim the product cannot keep: a placeholder attribute value, an overstated line of copy, a wrong fitment entry. What follows is predictable and slow.
| When | What happens |
|---|---|
| Immediately | Return filed. Refund cost absorbed, unit often unsellable |
| Within weeks | A negative review enters the corpus permanently |
| Within weeks | Negative customer experience rate rises. Return-rate badge risk appears on the detail page |
| Ongoing | Assistant answers begin sourcing from the complaint |
| Net effect | Conversion falls, then rank follows |
The original defect was a data-layer decision. The penalty arrives at the content and evidence layers months later, where nobody connects it back. A guessed attribute value is worse than a blank precisely because it passes gates on a promise the product does not keep, and this is where that promise gets paid for. Attribute census discipline is not compliance hygiene. It is return-rate management executed twelve weeks early.
| Signal | What it governs | Standing practice |
|---|---|---|
| Return rate by reason code | The specific defect: fit, expectation, quality, damage, wrong item | Pull monthly per Tier 1 ASIN. Reason codes are a listing brief. "Not as described" and "wrong size" are copy and attribute failures, not product failures |
| Negative customer experience rate | Aggregate dissatisfaction, badge and account-health exposure | Monitor at the same cadence as account health. Treat a rising trend as a stop-the-line event on that ASIN |
| Return-rate badging | A visible conversion penalty on the detail page | Once badged, conversion falls before rank does. The fix must be shipped and the badge earned back, which is slow. Prevention is the only economical strategy |
| Review text themes | The evidence corpus an assistant quotes | Every recurring complaint gets a direct answer on the page, or it becomes the answer the assistant gives |
| Voice-of-customer complaints | Now a direct input to the featured-offer ranking formula | Graded influence rather than binary eligibility. It competes against price rather than gating you out |
Pull return reason codes, map each one to the surface that caused it, fix that surface, answer the objection on the page, re-measure at 90 days. "Wrong size" is an attribute and a size chart. "Not as described" is a copy overclaim and an image gap. "Poor quality" is either a genuine product issue or an expectation the copy set too high. Fitment-driven categories should expect this loop to generate more listing changes than any keyword exercise.
There is a level of claim above which every additional conversion point costs more in returns than it earns in orders. Most listings sit below it. Aggressive copy programmes push through it, and the damage arrives with a lag long enough that nobody attributes it correctly. When conversion rises and return rate rises with it, the copy is writing cheques the product cannot cash, and the ranking system settles the account eventually.
What this means in practice
Five things you can do this week, in order.
The discipline here is unglamorous and it is the difference between a measurement system and a collection of anecdotes. Compute from raw counts, split before you read, respect the floors, judge in the declared window, and never let a single below-floor number set the agenda.
Sources
Primary sources for the confirmed and published claims above. Observations and framework inferences are labelled as such in the text and are not sourced here.
- Amazon Search: The Joy of Ranking Products — Sorokina & Cantú-Paz, SIGIR 2016 — www.amazon.science
- Amazon Science — Semantic product search (KDD 2019) — www.amazon.science
- COSMO — SIGMOD 2024, Amazon Science — www.amazon.science
- Amazon — What is Amazon SEO (official guidance) — sell.amazon.com
- Amazon — Best Sellers Rank — sell.amazon.com
- Amazon — Brand Analytics and Search Query Performance — sell.amazon.com
- Amazon Seller Forums — 250-byte search terms announcement — sellercentral.amazon.com
- Amazon Seller Forums — Featured Offer eligibility update, July 2026 — sellercentral.amazon.com
Four indices, forty ASINs, every week
The method is simple. Running it continuously across a catalogue, with the advertising and listing work attached to whatever it finds, is the job.
Get a free account teardown →