Work out how many users you need before you start, or you'll spend three weeks proving nothing and call it a null result.
Deciding how many coin flips you need before you start. Ten flips can't detect a slightly biased coin, and no amount of staring at the result will change that.
An underpowered test kills good ideas by producing a null result that gets reported as 'it doesn't work'.
Power is the probability of detecting an effect that is really there. An underpowered test is worse than no test: it produces a non-significant result that gets reported as 'no impact', and a good idea is killed by insufficient sample rather than by evidence. Four quantities are locked together — baseline rate, minimum detectable effect, significance level, and power — so fixing three determines the fourth. That is the calculation you do before launching, and its output is a required sample size and therefore a duration.
Power is P(detect the effect | it exists), conventionally set at 80%. Sample size depends on the baseline rate, the minimum effect you care about, α and power. Halving the detectable effect costs four times the sample. Compute it before launching; an underpowered test's null result means 'we couldn't see it', which is routinely mis-reported as 'it doesn't work'.
Statistical Power, Clearly Explained!!! — StatQuest with Josh Starmer, 8:19