Scenario runbooks
When something breaks, the expensive mistake is not the wrong fix. It is the fast fix — because now the account has a broken listing and a contaminated measurement window.
Contents — 13 sections
- The diagnostic order of operations
- The first question in any drop: did the impressions disappear?
- Indexation verification: the ten-second test
- Runbook A — rank dropped sharply, no obvious cause
- Runbook B — Buy Box or featured offer lost, no notification
- Runbook C — listing suppressed or restricted with no obvious cause
- Runbook D — ads spending, nothing ranking
- Runbook E — incident inside a deal window
- Measurement discipline
- The experimentation protocol
- Constants and thresholds
- Confidence grading and known gaps
- What this means in practice
When something breaks on a listing, the expensive mistake is not the wrong fix. It is the fast fix. A number falls, somebody rewrites a title, and the account now has a broken listing and a contaminated measurement window. Two weeks later nobody can say what caused the drop, because two things changed and only one of them was recorded.
These runbooks exist to slow the first ten minutes down and speed everything after that up. Each one starts from an observable symptom, runs a fixed check sequence, and resolves to the layer where the cause usually lives rather than the layer where the symptom appeared. That distinction is the whole discipline. Every layer of the stack produces symptoms that look like they belong to the layer above it: a Buy Box loss that is really a rejected attribute contribution, an advertising problem that is really a browse-node misassignment, a traffic collapse that is really an orphaned keyword.
Three rules govern everything below. Name the layer before you act. One variable per measurement window. Baseline before change. Break any of them and the account still might recover, but you will not know why, and you will not be able to do it again.
The final section on this page lists the places where this framework is thin or wrong. That is deliberate. A methodology that cannot name its own gaps is not a methodology, it is marketing.
The diagnostic order of operations
This runs ahead of every runbook on the page and every troubleshooting procedure in the wider system. It is five steps and it takes minutes.
A miscategorised ASIN makes every downstream optimisation unmeasurable, not merely suboptimal. Copy tests, image tests, bid changes and price tests run against a wrong browse node produce data that cannot be interpreted and conclusions that will be wrong in a direction you cannot predict. Verifying the data layer is not hygiene. It is the precondition for the validity of every experiment in the account.
The first question in any drop: did the impressions disappear?
Before the four-index work, before the copy conversation, before anyone opens the campaign manager, one question splits the entire investigation in two. Pull search query performance against your last archived export and look at impressions on the affected queries.
Yes — zero or near-zero
This is structural, not performance. The listing has stopped being eligible for that query. No bid change, no budget change and no copy change fixes this class of problem, and every hour spent on them is wasted.
No — impressions held, position is worse
This is performance. Decompose with the four indices and fix the stage that fell, one variable per window. Read the method on the four-index diagnosis.
When the answer is yes, work the causes in this order. It is ordered by how often each turns out to be responsible, not by how interesting it is.
A query at zero impressions after a structural edit is an orphaned term until you prove it is not. Confirm the market demand still exists, re-home the term according to the placement hierarchy, then verify indexation. Do not assume the term died. Assume you dropped it.
Indexation verification: the ten-second test
Three states fail separately and are fixed separately: indexed, ranked, visible. Diagnosing a ranking problem on a term that is not indexed wastes the entire investigation, and it happens constantly.
The test itself is trivial. Search the exact phrase plus the ASIN in the marketplace search bar. If the ASIN returns, the term is indexed. If nothing returns, it is not. That is the whole mechanism, and it is more reliable than any tool that claims to check indexation for you.
One term proves nothing. The sample has to be built deliberately, because each slice catches a different failure mode:
Why backend-only terms must always be in the sample
A backend search-terms field that exceeds its byte limit is rejected in its entirety — not truncated, not partially accepted — and no error is returned to the submitter. Observed The field looks populated in the interface. The listing behaves as though the field is empty, because it is.
Terms that also appear in the title or bullets cannot detect this, because they index from the visible copy regardless of whether the backend entry took. Only a term that exists solely in the backend field can prove the field saved. A head-terms-only sample will pass cleanly on a listing whose entire backend allocation was silently discarded. When that happens, trim to 249 bytes, resubmit once, wait the window and re-run the identical sample.
24 to 48 hours for a field update; up to two weeks for a structural change. Observed Testing inside the window and resubmitting because the result looked wrong turns one change into two variables and destroys the measurement. Inside the window the correct action is to stop and schedule.
Record results verbatim, per term, with timestamp and tester — not summarised. Hold the account state, marketplace and posture constant and write down what they were, because results differ by personalisation. Re-verification always runs the identical list, never a fresh one; a new list cannot tell you whether a fix worked. Full detail on the failure modes sits in indexation troubleshooting, and the structured version of this check is one of the procedures in the operator's kit.
When a term comes back not indexed, diagnose by term type rather than guessing:
| Term type | Most likely cause | Next step |
|---|---|---|
| Newly re-homed | Orphaned — deleted with no destination, or placed on a surface that did not save | Check the term inventory, place it, re-verify. This is the migration failure mode. |
| Backend-only | Byte limit exceeded; the whole entry was rejected with no error | Trim to 249 bytes, resubmit once, re-verify. If still absent, escalate to a catalog check. |
| Long-standing head term now absent | An attribute or product type change altered eligibility, or a contribution reverted | Halt. Resolve the catalog question before any copy work. |
| New branch term on a new ASIN | Never placed, or placed only inside an image | Place it as indexed text. Images do not index as search terms. |
Runbook A — rank dropped sharply, no obvious cause
Symptom. Organic position on one or several important queries falls hard inside a short window. Nothing was announced, no notification arrived, and the page looks normal.
Do not ship a fix until the cause is named. A fix applied to the wrong layer does not fail neutrally — it becomes a second variable inside the measurement window, and now the drop and the fix are entangled. If you genuinely cannot name the cause, the correct action is to change nothing and collect another week of data. That feels like inaction. It is the cheapest thing on this page.
The extended version of this investigation, including the causes that get blamed most and turn out guilty least, is in rank drop diagnosis.
Runbook B — Buy Box or featured offer lost, no notification
Symptom. The featured offer is gone, or a child ASIN has become unbuyable, and no suppression notice or policy message arrived to explain it. Sessions may look normal while units collapse.
The reflex is to check price. Price is the last check in this sequence, not the first, because the silent failures live in the catalog and the commercial ones announce themselves.
If steps one to three come back clean and the commercial inputs are competitive, you are probably looking at a routing problem rather than a data problem. Apply the stop rule: two identical responses through the same channel and you change the channel, not the wording.
Runbook C — listing suppressed or restricted with no obvious cause
Symptom. A listing is suppressed, restricted or under review. The visible page looks compliant. The flat file shows nothing wrong. Appeals come back with the same response.
Enforcement runs against catalog data, not against the visible detail page. Every appeal is judged against a backend record you may never have seen. This is why appeals fail repeatedly while the seller argues, accurately but irrelevantly, about the page.
This framework holds diagnosis and triage for this class of problem. It does not hold the escalation resolutions — the exact routing, the phrasing that moves a case to a team with jurisdiction, the language that unlocks it. That material is not published anywhere and is not discoverable from the outside. Anyone selling you a guaranteed reinstatement script is describing something they cannot possess.
How exposed is your catalogue on this?
Twenty checks across the five layers, scored 0–100, returning a ranked issue list by severity — the same audit we run on client accounts. Free, no signup to see your result, and it runs entirely in your browser.
Score my listings →Nothing you answer is transmitted or stored. The written report and the six working templates are the optional email step afterwards.
Runbook D — ads spending, nothing ranking
Symptom. Campaigns are delivering, spend is landing, impressions and clicks exist, and organic position on the funded queries does not move. Often accompanied by a request to increase the budget.
If the purchase index on a funded query is below the gate, the correct action is to stop funding it and route the work to the offer — price, delivery promise, stock reliability. Copy cannot fix a purchase-index problem and neither can a bid. The full argument is in PPC and organic rank.
Runbook E — incident inside a deal window
Symptom. Something breaks during a deal, a promotional event or a major traffic spike. Stock, price integrity, Buy Box, a suppression, a feed failure. Everyone wants to fix everything at once.
Stabilise, do not optimise. Restore availability, price integrity and the featured offer. Ship no structural change of any kind until the event is over. The window is already contaminated for measurement, so there is nothing to learn from a change made inside it — only damage to accumulate.
Measurement discipline
This section exists because the industry has no shared standard for it. Nearly every case study in circulation is a success story with a headline multiple attached, no control, no window definition and no statement of what else changed in the account during the period. Procedures either bring their own evidentiary standard or they reproduce the same unfalsifiable claims with a different name on them.
The judge-by windows
| Window | What is readable |
|---|---|
| +48 hours | Publication acceptance and attribute re-pull states. Did the change actually take? |
| Weeks 1–2 | Indexation settles. Re-check every pre-export query and run the orphan sweep. |
| Weeks 2–4 | Click-through. Tile and name effects become readable. |
| Weeks 4–6 | Impression share and family coverage. This is the verdict window. Judge the change here. |
| Months 2–6 | Assistant-layer effects — answers sourcing from the listing. Review-evidence timelines, then compounding. |
The five requirements for any change you ship
The rollback doctrine
Reverting is a stronger action than it feels like, and most reverts are made too early, for the wrong reason, and in the wrong way.
Every non-trivial problem the team resolves gets written up in one format and archived with a number: problem · prevailing assumption · actual system behaviour · escalation sequence · resolution. This is what converts solved tickets into institutional capability. Treat it as a deliverable of the work, not as documentation overhead.
The experimentation protocol
Measurement discipline compares a change against its own baseline. Where experiment tooling is available, a change can be tested against a live control instead, and that is the only method in this framework that removes the confound problem entirely. It is also available for a narrower range of questions than most people assume.
Good candidates
Hero image variants. Title formulations within the same identity root. A+ module arrangements. Bullet ordering and framing. The common shape is single-variable, reversible, and conversion-visible.
Poor candidates
Anything whose effect is on retrieval rather than conversion — attribute fills, backend changes, coverage expansion. These change who arrives, not what they do on arrival.
A conversion test cannot see a visibility gain. An attribute census that makes an ASIN eligible for six new query families will show no conversion-rate improvement whatsoever, and may show a decline as broader traffic arrives. Judge eligibility and coverage work on impression share and branch count, never on a conversion experiment. Teams have killed genuinely successful retrieval work on exactly this reading.
For most catalogs, most of the time, the tooling is not available. The fallback is the full measurement discipline above: pre-export baseline, one variable, declared judge-by date, significance floors, confound log, rollback trigger. The fallback is weaker, and it is what most decisions will actually rest on, which is precisely why that discipline is not optional.
Constants and thresholds
Every load-bearing number in one place, so procedures reference rather than restate. The standing column matters as much as the value. Amazon confirmed means it comes from Amazon's own documentation or announcements. Observed means reproducible practitioner testing without official confirmation. Framework standard means it is this system's own operating threshold — a decision rule, not a discovered fact. Needs verification means it rests on limited observation and should be confirmed before it enters anything client-facing.
| Domain | Constant | Value | Standing |
|---|---|---|---|
| Listing | Item name / item highlights | 75 / 125 characters; both index; neither prioritised | Amazon confirmed |
| Listing | Review Listing Changes window | 14 days | Amazon confirmed |
| Listing | Bulk upload application time | ~8 hours | Needs verification |
| Catalog | Contribution scores | 0–100 scale; circulating figures of 30 without Brand Registry, 52 with it, 60 for a Vendor record; data augmenters and internal teams higher and unpublished | Observed — no Amazon primary source; do not quote to a client |
| Copy | Backend allocation | 60% specification · 20% general · 10% own-brand equivalence · 10% complement | Framework standard |
| Copy | Specification density target | ≥8 concrete attributes with units per listing | Framework standard |
| Search query data | Index parity and action threshold | 1.0 is parity; act below ≈0.85 | Framework standard |
| Search query data | Significance floors | ≈2,500 impressions for a click-through index · ≈100 clicks for cart-add or purchase · 3 periods for a trend | Framework standard |
| Search query data | Attribution window | ~24 hours, same-day | Observed |
| PPC | Spend eligibility gate | Purchase index ≥ ~0.9 | Framework standard |
| PPC | Top-of-search prevalence | Best converting in ~70% of brands; best cost-per-acquisition in 40–50% | Observed |
| PPC | Turnaround reallocation | 70–80% of spend to top-of-search; ~80% to hero SKUs per marketplace | Turnaround situations only |
| PPC | New-to-brand pause trigger | Below 10% new-to-brand on vCPM or display | Observed |
| Pricing | Price gate | +10–15% above the market purchase median, tier- and pack-normalised | Framework standard |
| Pricing | Price-increase test | 5–10%, observed 14 days on BSR, conversion rate, rank and velocity | Observed |
| Pricing | Business pricing floor | 5% off retail | Amazon confirmed |
| Inventory | Cover policy | 90+ days target; never below 75; five warehouses where possible | Observed |
| Inventory | Inventory Performance Index | 0–1,000 scale; 450 risk threshold; quarterly checkpoints; optimise 6–8 weeks prior | Amazon confirmed |
| Inventory | Prime-badge postcode audit | 12-postcode panel; pass at 8 of 12 same-day or next-day | Framework standard |
| Inventory | Long-term storage risk | 365 days in fulfilment centres | Amazon confirmed |
| Launch | Review-programme stacking | ~30 units per marketplace on the same hero child; target 90–150 reviews; minimum 2 marketplaces | Observed |
| Deals | Deal architecture | 60-day runway · price discount during · best deal after (~80/20) · 4-day price discount with a 12-hour lightning deal inside | Observed |
| Deals | Eligibility floor | Storefront rating ≥3.5 stars | Amazon confirmed |
| Compliance | Featured offer model | Gate-then-rank moving to rank-only, from July 2026 | Amazon confirmed |
| Compliance | AI-people image disclosure | "contains synthetic performer" in the dc:subject XMP field, pre-upload | Amazon confirmed |
| Compliance | Business delivery standard | 90% threshold from 30 September 2026; 14-day rolling; deactivation risk from 30 October | Amazon confirmed |
| Windows | Judge-by | Indexation +2 weeks · click 2–4 weeks · visibility 4–6 weeks · assistant layer 2–6 months | Framework standard |
Dated compliance items are tracked in the 2026 changelog. The index calculations behind the search-query rows are worked through in the SQP index calculator.
Confidence grading and known gaps
Nothing on this page is worth much if you cannot tell which parts are load-bearing and which are provisional. This section separates them.
What is settled — build here first
Seven findings hold across unrelated accounts, categories and mechanisms, and are treated as settled:
- Structured data beats presentation. What the catalog believes the product is governs more than what the page says it is.
- The title is now two co-equal indexed fields, not one field with a truncation problem. Amazon confirmed See the 75/125 title system.
- Delivery speed is a conversion and ranking lever, not a logistics detail.
- Escalation is a routing problem, not a persistence problem.
- Pricing consistency is an algorithmic signal with memory.
- An increasing share of purchase decisions is made in advance, by a system, against structured data, with the listing absent. Amazon research
- Deal architecture is a system rather than a discount.
A finding is only strengthened by evidence that could have contradicted it. Two observations sharing a mechanism are one observation. Agreement between two accounts in the same category, on the same tactic, in the same window, is a single data point wearing two hats — and treating it as two is how a plausible idea quietly becomes an unexamined rule.
Contested questions, and where this framework stands
| Question | The disagreement | This framework's position |
|---|---|---|
| Branded PPC | Cut it entirely as recoverable waste, versus defend the brand's own search real estate | A test protocol, not a policy. Defined window, branded-share metric, pre-declared rollback. Anyone with a universal answer here has not run the test. |
| How much AI search matters now | Rapid growth in assistant-originated traffic, versus conventional rank still deciding most outcomes | Optimise for AI retrieval because the work is identical to conventional relevance work — never as a displacement of fundamentals. |
| Who owns PPC | One senior operator owning the full P&L including spend, versus specialist ownership | A brand-side default does not transfer to an agency-side team. Decide ownership per engagement before writing it into a procedure. |
| Title change severity | "Discoverability is unchanged", versus "it materially alters content architecture" | Both, on different clocks. No immediate ranking change; meaningful medium-term architectural change. |
Known gaps — where a procedure written today would fail
Being explicit here is not a weakness in the framework. It is the part that makes the rest of it checkable.
When this framework and your own account data disagree, the account data wins, and the disagreement gets logged as a case. This is a prior, not an authority. It is versioned, dated, and expected to be wrong in specific places. The discipline is in noticing where, and writing it down.
What this means in practice
Five things you can do this week, in the order that pays back fastest.
If your diagnosis lands on something structural across a large catalog — orphaned terms after a migration, contribution rejections nobody logged, a browse tree that stopped matching the products in it — that is systems work rather than listing work, and it is what a free account teardown is designed to surface.
Sources
Primary sources for the confirmed and published claims above. Observations and framework inferences are labelled as such in the text and are not sourced here.
- Amazon Search: The Joy of Ranking Products — Sorokina & Cantú-Paz, SIGIR 2016 — www.amazon.science
- Amazon Science — Semantic product search (KDD 2019) — www.amazon.science
- COSMO — SIGMOD 2024, Amazon Science — www.amazon.science
- Amazon — What is Amazon SEO (official guidance) — sell.amazon.com
- Amazon — Best Sellers Rank — sell.amazon.com
- Amazon — Brand Analytics and Search Query Performance — sell.amazon.com
- Amazon Seller Forums — 250-byte search terms announcement — sellercentral.amazon.com
- Amazon Seller Forums — Featured Offer eligibility update, July 2026 — sellercentral.amazon.com
If the diagnosis lands on something structural
Orphaned terms after a migration, contribution rejections nobody logged, a browse tree that stopped matching the products in it — that is systems work, and it is what the teardown is built to surface.
Get a free account teardown →