Split users at random, change one thing, and the difference you measure is caused by the change — that last clause is the whole point, and randomisation is what buys it.
Two identical queues, one gets the new till software. Because people were sent to queues at random, any difference in speed is the software.
It's the only routinely available tool that licenses the word 'caused'.
An A/B test is the only routinely available tool that gives a causal answer. Randomly assigning users to control and treatment makes the two groups equivalent in expectation on everything — including the things you never thought to measure — so any difference in outcome is attributable to the change. Everything else about running one is protecting that property: assigning consistently, deciding the sample size in advance, choosing one primary metric before you look, and running for whole weeks so that day-of-week composition matches.
Randomise users into control and treatment, change exactly one thing, and compare a pre-declared primary metric. Randomisation makes the groups comparable on unobserved variables, which is what licenses a causal claim. The discipline is up front: fixed sample size, one primary metric, guardrails, and no peeking — because deciding when to stop after seeing the data invalidates the statistics.
Simple explanation of A/B Testing — codebasics, 5:49