How LeakStop measures waste
Last updated: August 17, 2026 · The working standard: every number shown to a user is either derived from that user's own data, or labelled an estimate with its method named.
Recoverable
Spend on search terms with 0 conversions that share a theme with other non-converting terms, where that theme never appears in a converting search. That is precisely the waste a negative keyword removes without collateral damage, which is what the word has to mean if the figure is going to survive a sceptical reader.
Two different click floors, because two different claims are being made. A theme is only established by terms carrying at least 5 clicks each — a lone quiet search is not evidence of anything. Once the theme is established, a term joins it from 2 clicks, since a search that happened twice under a proven pattern is the same waste as its heavier siblings. This one is a judgement, and we would rather say so than dress it up. It used to be defended here with recall and precision against a professional's hand-built waste list. We withdrew that in August 2026 after running the check we had skipped: 92.9% of that account's zero-conversion terms were already on their list, so agreeing with it is near-free, and simply sorting the export by cost scores better than our rule on both recall and precision. The floor stays at 2 because a search that happened twice under a proven pattern is the same waste as its heavier siblings — but it is an unvalidated number, and we will say what moved it the day we can measure it forward.
It is measured from the export, never taken as a percentage of it. Waste that occurs once is reported separately as unclustered, because no negative keyword removes a search that happens once, and counting it would inflate the headline with money the reader cannot get back.
Summed per term, never per n-gram. A search sits under several themes at once, and adding up the themes would count the same pound twice and report more recoverable waste than the account spent.
Safe to negative
Every theme carries its converting side as well as its wasted side. A theme that also appears in searches that earned is shown, and never suggested — recommending "free" as a negative while "free consultation" converts is not a wrong estimate, it is advice that costs the reader money.
Where a theme is unsafe but the money is real, the account's own conversions choose the match type rather than whether to report: the suggestion falls back to exact negatives on that cluster's own zero-conversion terms.
Well evidenced
A query performing at exactly the account's own conversion rate still shows zero conversions with probability (1 − rate)^clicks. At a 6.1% account rate that is 64% of the time on seven clicks. So spend counts as well evidenced only once it has enough clicks for a run of zeros to be unlikely at that account's rate: ceil(log 0.05 / log(1 − rate)) — 48 clicks at 6.1%, 116 at 2.55%, 8 at 33%.
This labels strength; it is deliberately not a filter. Measured across our validation corpus, gating on it would delete four of the five findings the product makes, including the one with independent human confirmation — an account whose own analysts hand-excluded the same query family we flag. Their judgement was semantic, not statistical, and a bar that overrules it is quieter, not truer.
What you could have found by sorting
Reported on every audit. We take the terms making up the recoverable estimate and rank them against every zero-conversion row by cost — your view, taken with no click floor and no theme test, so the comparison is deliberately generous to it. Where most of our estimate sits in the rows you were going to read anyway, the report says so instead of selling you something.
On one real account of $342,047 across 54,112 search terms: recoverable $8,285, of which 12% sits in the twenty most expensive zero-conversion rows. Reaching 80% of it means reading 2,773 rows. The median flagged row cost $12; row twenty of the sorted list cost $166.
Across the rest of our validation set the same figure runs 56%, 71%, 75% and 100%. The line we draw at 80% is therefore a judgement that changes real answers, not a safe gap between extremes — which is worth saying plainly, because a threshold nobody admits to is how a measurement quietly becomes a marketing number.
When we decline to judge
If an export records no conversions anywhere, no waste rule speaks. Waste is spend that produced nothing while other spend produced something, and an export with no second half has no comparison to make. Above 100 clicks the report additionally says the likely cause is tracking rather than waste; below it, a quiet month is not evidence of anything. This is why an account can upload real spend and correctly receive no estimate.
Thresholds, and which of them are calibrated
Money floors are expressed as shares of the account's own 30-day spend, never as absolute amounts. An absolute figure carries no currency: a "$500/month" floor is about $3 for a Japanese advertiser, and would fire on everything. Ratios are scale-invariant — which is not the same as calibrated, and we say so.
The savings ledger's baseline is the one constant that has been measured against outcomes: across 2,251 days from three real accounts, each 28-day baseline compared to the 90 days that actually followed came out at a median of 0.84–0.99 and a 90th percentile of 1.07–1.20. Conservatism holds at the middle. The tail is seasonality, which the outlier clipping structurally cannot reach, and a leak sealed just after a seasonal peak can accrue against a baseline well above what that scope went on to spend. That is a limitation, not a footnote.
What acting on this costs, measured
Every suggestion we make is checked against the searches that earned in the same export, and one that would block an earner is shown but never suggested. That check is worth exactly what it sounds like and no more: it is true by construction, since a theme only reaches the suggestion list if nothing it appears in converted. Reporting it as a safety record would be circular.
The honest version looks forward. We took the suggestions produced from one window of a real account and asked what they would have blocked in a later window — a period nothing in the calculation could know about. Across three window pairs our list would have blocked 0%, 7.7% and 17.2% of the conversions that account went on to earn. The naive approach on the same data — take the most expensive searches that never converted and block the phrases inside them — blocked 3.4%, 42.3% and 82.8%.
So: three to five times safer than the obvious method, and not zero. On one pair, acting on the whole list unreviewed would have cost that account five conversions out of twenty-nine. This is one account and the conversion counts are small, so treat it as an order of magnitude rather than a rate. It is also the reason every list here is a list of candidates and the report says "verify before acting" — that sentence is not legal throat-clearing, it is what the measurement says.
What we do not claim
- That the estimate is money you will recover. It is the spend a negative keyword could remove without cutting traffic that earns. Acting on it is your decision and your judgement.
- That silence means a clean account. It usually means the rules found nothing they can stand behind on the data you gave them.
- That our figures replace your own. Everything above is computed from the export you upload, and every input to it is a column you can read yourself.