What Is Hypothesis Testing in Statistics
SkillVeris Team
Data Science Team

Hypothesis testing is a formal procedure for deciding whether sample data provides enough evidence to reject a default claim about a population.
In this guide, you'll learn:
- You start with a null hypothesis (no effect) and an alternative hypothesis (an effect), then use data to decide between them.
- The p-value is the probability of seeing data at least this extreme if the null hypothesis were true — small p-values cast doubt on the null.
- You compare the p-value to a significance level, often 0.05, chosen before running the test.
- A Type I error rejects a true null (false positive); a Type II error fails to reject a false null (false negative).
1What Is Hypothesis Testing?
Hypothesis testing is a structured statistical method for deciding whether the data you collected provides enough evidence to reject a default assumption about a population. You state a claim, assume it is false to begin with, and then check whether your sample would be surprising under that assumption. If it would be very surprising, you reject the default.
It is the backbone of scientific and business experimentation, from clinical trials to A/B tests. Rather than eyeballing a difference and guessing whether it is real, hypothesis testing gives a disciplined, probability-based way to separate genuine effects from random noise.
2Null and Alternative Hypotheses
Every test starts with two competing statements. The null hypothesis is the skeptical default — usually 'there is no effect' or 'no difference'. The alternative hypothesis is what you suspect is true — 'there is an effect'. The test is built to challenge the null, not to prove the alternative directly.
- Null hypothesis (H0): the status quo, e.g. the new button colour has no effect on clicks.
- Alternative hypothesis (H1): the claim of an effect, e.g. the new colour changes clicks.
- One-tailed test: the alternative specifies a direction (higher OR lower).
- Two-tailed test: the alternative allows a difference in either direction.
💡You Never Prove the Null
A hypothesis test can only reject the null or fail to reject it. Failing to reject is not proof that the null is true — it simply means you lack sufficient evidence against it.
3The Testing Process
Hypothesis testing follows the same repeatable steps regardless of the specific test. Deciding the rules before you look at the outcome is what keeps the procedure honest.
- State H0 and H1 clearly in terms of a population parameter.
- Choose a significance level (alpha), commonly 0.05, before running the test.
- Pick and run the appropriate test statistic (t-test, chi-square, ANOVA, etc.).
- Compute the p-value from the test statistic.
- Compare the p-value to alpha, then decide to reject or fail to reject H0.
Set Alpha Before You Peek
The significance level is your tolerance for a false alarm, and it must be fixed before you see the results. Choosing alpha after looking at the data — or repeatedly testing until something passes — inflates your false-positive rate and undermines the whole exercise.
4Understanding the P-Value
The p-value is the probability of observing data at least as extreme as yours if the null hypothesis were actually true. A small p-value means your data would be unlikely under the null, which counts as evidence against it. If the p-value falls below your chosen alpha, you reject the null.
- p < alpha (e.g. p < 0.05): reject the null — the result is statistically significant.
- p >= alpha: fail to reject the null — insufficient evidence of an effect.
- A smaller p-value means stronger evidence against the null, not a bigger effect.
⚠️What a P-Value Is Not
The p-value is not the probability that the null hypothesis is true, and it is not the probability your result happened by chance. It is the probability of the data given that the null holds — a subtle but crucial distinction.
5Type I and Type II Errors
Because you decide based on a sample, you can be wrong in two opposite ways. A Type I error is a false positive — rejecting a null that is actually true. A Type II error is a false negative — failing to reject a null that is actually false. There is a trade-off: lowering the risk of one generally raises the risk of the other.
- Type I error (alpha): concluding there is an effect when there is none.
- Type II error (beta): missing a real effect that exists.
- Power (1 - beta): the probability of correctly detecting a real effect.
- Larger samples reduce both error rates and increase power.
Balancing the Two
Setting alpha very low guards against false positives but makes real effects harder to detect, raising Type II errors. The right balance depends on the cost of each mistake — a false positive in a medical screen and a missed disease carry very different consequences.
6Common Mistakes to Avoid
Hypothesis tests are easy to run and easy to misinterpret.
- Treating p >= 0.05 as proof the null is true, rather than absence of evidence.
- Confusing statistical significance with practical importance — a trivial effect can be significant with a huge sample.
- P-hacking: running many tests or peeking repeatedly until something crosses 0.05.
- Choosing alpha or the test direction after seeing the data.
- Ignoring assumptions of the test, such as normality or independence.
7Common Tests and When to Use Them
The general procedure stays the same, but the specific test depends on your data and question. Picking the right one is mostly about the type of variables involved and how many groups you are comparing.
- t-test: compare the means of one or two groups of continuous data.
- ANOVA: compare the means of three or more groups at once.
- Chi-square test: check whether two categorical variables are associated.
- Paired t-test: compare before-and-after measurements on the same subjects.
- Mann-Whitney U: a rank-based alternative when the data is not normal.
8Key Takeaways
Hypothesis testing brings discipline to the question 'is this difference real?'
- You test a null hypothesis of no effect against an alternative of an effect.
- The p-value measures how surprising the data would be if the null were true.
- Compare the p-value to a pre-chosen significance level (often 0.05) to decide.
- Type I errors are false positives; Type II errors are false negatives.
- Significance is not importance, and failing to reject is not proof of the null.
9Frequently Asked Questions
Q: What does a p-value of 0.05 mean? A: It means that if the null hypothesis were true, there would be a 5% chance of observing data at least as extreme as what you collected. If you had set your significance level at 0.05, a p-value at or below it would lead you to reject the null. It does not mean there is a 5% chance the null is true.
Q: What is the difference between the null and alternative hypothesis? A: The null hypothesis is the default claim of no effect or no difference, which the test tries to disprove. The alternative hypothesis is the claim that an effect or difference exists. Evidence against the null is interpreted as support for the alternative, but the test never proves the null.
Q: What are Type I and Type II errors? A: A Type I error is a false positive — rejecting a null hypothesis that is actually true. A Type II error is a false negative — failing to reject a null hypothesis that is actually false. The significance level controls the Type I error rate, while statistical power relates to avoiding Type II errors.
Q: Does a statistically significant result mean the effect is important? A: Not necessarily. With a large enough sample, even a tiny, practically meaningless difference can be statistically significant. Always look at the effect size and a confidence interval alongside the p-value to judge whether the result matters in the real world.
Related Reading
Get The Print Version
Download a PDF of this article for offline reading.
About the Publisher
SkillVeris Team
Data Science Team
Our data team shares real-world analytics, ML, and SQL insights grounded in industry practice.
View all postsRelated Posts
Never miss an update
Get the latest tutorials and guides delivered to your inbox.
No spam. Unsubscribe anytime.