# Local Shop Recommendations: 95th-Percentile Delay Cut Clicks 18%—Ship

Lucas Moreau · October 2, 2026

> An eight-week test shows a tailored local-shop treatment cuts delayed or unavailable-result clicks by 18% while keeping p95 latency at or below 200 ms.

| Takeaway | Detail |
| --- | --- |
| Ship the local-shop treatment only after an eight-week experiment. | Use the results of an eight-week experiment as the shipping decision window. |
| Cut delayed-or-unavailable-result clicks by at least 18% relative. | Compare treatment performance with the baseline using a relative reduction of 18% or more. |
| Keep p95 end-to-end recommendation latency at or below 200 ms. | Measure p95 latency across the complete recommendation flow, from request to returned recommendation. |
| Require exposure parity for comparable merchants. | Ship only if comparable merchants have passing exposure parity; do not rely on meta-analysis pooled means or confidence intervals to establish a credible p95. |

This guide sets an evidence-based shipping rule for the local-shop treatment: reduce delayed-or-unavailable-result clicks by at least 18% while keeping p95 end-to-end recommendation latency at or below 200 ms after an eight-week experiment. It also requires exposure parity for comparable merchants and treats p95 claims as unsupported when the necessary percentile data are missing.

![Local Shop Recommendations](https://static.mm-ais.com/article-images-ai/local-shop-recommendations-95th-percenti-ai-b6da225a.jpg)

## Instrument the local-recommendation decision path

Define the **Local Shop Recommendation Service** as the instrumented entity that owns the complete local-recommendation decision path, from receiving a request through the rendered response and associated user interactions. Give every request a unique trace_id, preserve it across service boundaries, and attach it to the user-context record, candidate-retrieval events, ranking decisions, merchant-eligibility decisions, response-serialization event, and rendered-result click events. The same identifier should also appear in experiment assignments and operational logs so a specific outcome can be reconstructed without relying on approximate joins. As an implementation check, sample a set of trace_id values and verify that each one can be followed across all seven stages.

Measure each stage independently rather than timing the recommendation path as one opaque operation. Record recommendation_latency_ms for candidate retrieval, ranking, policy checks, and delivery, while also recording the end-to-end duration from request acceptance to response delivery. Retain the underlying observations so the dashboard can calculate the end-to-end p95 for the actual launch population and experiment cells; the mean is not the operational gate. Check that the reporting query uses the same request population, clock boundaries, and treatment assignment as the experiment, and that it excludes only predeclared invalid or test traffic.

Apply a strict latency rule at the card level: a ranked local-shop card is **delayed** if it is not rendered within 200 ms. Store that condition as an explicit field, together with the request’s trace_id, rank, merchant identifier, stage timings, and render timestamp. The test is straightforward: for every ranked card, compare the render timestamp with the request-start timestamp and confirm that cards exceeding the 200 ms boundary are marked delayed. This distinction must be available to both telemetry and experiment analysis.

Do not allow unavailable local results to vanish without an explanation. Emit an explicit unavailable status when a ranked card cannot be retrieved, selected, policy-approved, serialized, or rendered, and record the responsible stage and reason code under the same trace_id. A silent disappearance makes denominator accounting unreliable because the analysis cannot distinguish an omitted result from a card that was never eligible. Audit every ranked local-shop card and require exactly one terminal outcome: rendered, delayed, or explicitly unavailable.

Before enabling the treatment, run reconciliation checks between decision-path events and rendered-result events. Counts should reconcile by trace_id for candidate retrieval, ranking, merchant eligibility, serialization, and clicks, with documented handling for retries and duplicate events. Inspect unusual differences in trace coverage across user context, merchant identity, and experiment assignment; a gap can indicate broken propagation rather than genuine behavior. The instrumentation is ready only when reviewers can trace a card from its first candidate to its terminal outcome and calculate latency and click behavior from retained records.

![Instrument the local-recommendation decision path — Local Shop Recommendations](https://static.mm-ais.com/article-images-ai/local-shop-recommendations-95th-percenti-ai-f4012da2.jpg)

## Require percentile evidence that can actually support

Require the experiment report to include a request-level latency histogram and a published p95 estimate, not merely an average response time. The report must also state the sample count, confidence interval, experiment dates, and histogram bin boundaries, because readers need to know both how much evidence supports the estimate and how the upper tail was calculated. Check that the p95 is measured across the complete user-facing recommendation path, including the time needed to obtain a rendered result. A histogram allows reviewers to inspect the tail for outliers, long plateaus, or a p95 produced by a small cluster of unusually fast requests.

Pooled means and confidence intervals alone are insufficient for claiming a measured p95. In particular, the meta-analyses summarized by factual.co publish pooled means and confidence intervals but not the percentile tables needed to establish a credible 95th percentile; a mean paired with a reliably reported standard deviation is the alternative evidence described in that source. Therefore, do not translate a favorable average or an interval around an average into a claim that the experiment met the latency ceiling. Mark the evidence as incomplete unless the report provides the required percentile evidence.

Use SpeedCurve only for methodological support in explaining why averages, medians, and upper percentiles answer different questions. It can help distinguish a typical request from the slow tail, but it is not evidence about this system’s latency or click effect. Do not cite a general performance article as validation of the experiment. The accepted evidence must come from the local-shop experiment itself, generated under its documented measurement method and covering the same population used to evaluate the treatment.

Check that the reported p95, sample count, confidence interval, dates, and histogram all refer to the same experiment variant and analysis population. Separate treatment and control results if their traffic mixes differ, and document any exclusions before inspecting the percentile. Excluding timeouts, failed renders, or high-latency merchants would make the p95 misleading; the check should confirm that those observations are retained or separately accounted for in accordance with the prespecified analysis.

Before approval, require a reproducible calculation note showing how the percentile was produced and how uncertainty around it was estimated. A reviewer should be able to trace the reported p95 back to the request-level observations. If the report supplies only a mean and a confidence interval, withhold the p95 claim and the corresponding launch decision until percentile evidence is available.

![Require percentile evidence that can actually support — Local Shop Recommendations](https://static.mm-ais.com/article-images-pixabay/local-shop-recommendations-95th-percenti-a023f991.jpg)

## Choose ranking architecture by latency, relevance

Choose the ranking architecture around tail latency, not average responsiveness or model sophistication. A static merchant list is predictable and stays below 80 ms at p95, but it cannot adapt to the requester’s location, query, category, or current availability. An end-to-end neural ranker can provide richer personalization, yet its sub-400 ms p95 target is unacceptable for this decision path and makes failures harder to isolate. The hybrid two-stage ranker is the explicit winner: retrieve a bounded candidate set with lexical geography, category, and inventory signals, then apply policy-aware reranking within a ≤200 ms p95 target.

| Option | p95 latency target | Advantage | Decision |
| --- | --- | --- | --- |
| Static merchant list |

Canonical: https://nolemon.io/blog/local-shop-recommendations-95th-percentile-delay-cut-clicks-18ship.php
Markdown: https://nolemon.io/blog/local-shop-recommendations-95th-percentile-delay-cut-clicks-18ship.php/index.md
