p99 = tail latency: the slowest 1% of requests take 3s. The first move is to measure and break it down, not guess. Compare p50 vs p99 — if p50 is fast but p99 is 3s, the path isn't uniformly slow; something bites only under load or in edge cases (pool saturation, GC pauses, N+1, cold cache, an un-timed external call).
p99 is the , not the average wait — speeding up the average does nothing for the customer stuck at the back.
