A p-value is the probability of seeing a difference this big if the change did nothing — not the probability that the change did nothing.
Flip a fair coin ten times and getting eight heads isn't proof it's rigged — it's how often a fair coin does that anyway. The p-value is that 'how often'.
Reading it backwards — as the chance the feature works — is the single most common statistical error in industry.
Run a test where the variants are genuinely identical and you will still measure a difference, because samples vary. The p-value quantifies that: assuming no real effect, how often would random noise produce a gap at least as large as the one observed? A small p-value means the data are surprising under 'no effect', which is evidence — but it is evidence about the data given a hypothesis, not about the hypothesis given the data. Nearly every misuse in industry comes from swapping those two, which is also why 'p = 0.04' does not mean 'a 96% chance the feature works'.
A p-value is P(data this extreme | no real effect). It is not the probability the null is true, and 1 − p is not the probability your variant wins. The 0.05 threshold is a convention that fixes the false-positive rate at 5% per test, which is why testing twenty metrics finds a false winner roughly two-thirds of the time. Report the confidence interval — it carries the effect size, which the p-value throws away.
P-values and significance tests | AP Statistics | Khan Academy — Khan Academy, 7:58