What is hypothesis testing and what are the null and alternative hypotheses?
Understand hypothesis testing, the null and alternative hypotheses, p-values, and significance levels with clear examples for data science interviews.
Expected Interview Answer
Hypothesis testing is a formal statistical procedure that uses sample data to decide between two competing claims about a population, called the null and alternative hypotheses.
The null hypothesis (H0) states there is no effect or no difference and is the default assumed true, while the alternative hypothesis (H1) states there is an effect or difference. You compute a test statistic from the data, convert it to a p-value, and compare that p-value against a significance level such as 0.05. If the p-value is below the threshold you reject the null in favor of the alternative; otherwise you fail to reject the null. You never prove the null true, you only fail to find enough evidence against it.
- Provides an objective decision rule
- Controls the rate of false positives via the significance level
- Separates real effects from random noise
- Standardizes reasoning across experiments
- Underpins A/B testing and scientific claims
AI Mentor Explanation
A third-umpire review is hypothesis testing: the on-field not out is the null hypothesis, assumed to stand unless the replay shows strong evidence. The alternative is out. Only when the video evidence is convincing enough do you overturn the null; weak or inconclusive footage means you fail to reject and the original decision holds.
Step-by-Step Explanation
Step 1
State the hypotheses
Define H0 (no effect) and H1 (an effect), and decide whether the test is one-sided or two-sided.
Step 2
Choose the significance level
Set alpha, commonly 0.05, which is your tolerated probability of a false positive.
Step 3
Compute the test statistic
Calculate the appropriate statistic (t, z, chi-square) from the sample data.
Step 4
Find the p-value
Determine the probability of a result at least as extreme as observed, assuming H0 is true.
Step 5
Make a decision
Reject H0 if the p-value is below alpha; otherwise fail to reject H0 and interpret in context.
What Interviewer Expects
- Clear definitions of null and alternative hypotheses
- Correct meaning of the p-value and significance level
- Understanding of Type I and Type II errors
- Knowing you never prove the null true
- Awareness of one-sided versus two-sided tests
Common Mistakes
- Interpreting the p-value as the probability the null is true
- Saying a non-significant result proves the null hypothesis
- Confusing statistical significance with practical importance
- Setting the alternative hypothesis after peeking at the data
- Ignoring test assumptions such as independence or normality
Best Answer (HR Friendly)
“Hypothesis testing is a structured way to check whether a claim is supported by data. You start by assuming nothing special is happening (the null hypothesis), then see if the evidence is strong enough to conclude something real is going on (the alternative), using a cutoff to avoid being fooled by chance.”
Code Example
from scipy import stats
control = [72, 68, 75, 70, 69, 74, 71]
treatment = [78, 82, 80, 77, 85, 79, 81]
# H0: means are equal, H1: means differ
t_stat, p_value = stats.ttest_ind(control, treatment)
alpha = 0.05
print(f"t-statistic: {t_stat:.3f}, p-value: {p_value:.4f}")
if p_value < alpha:
print("Reject H0: the difference is statistically significant")
else:
print("Fail to reject H0: not enough evidence of a difference")Follow-up Questions
- What is the difference between a Type I and a Type II error?
- What exactly does a p-value represent?
- When would you use a one-tailed versus a two-tailed test?
- What is statistical power and how do you increase it?
- How does the multiple comparisons problem affect hypothesis testing?
MCQ Practice
1. The null hypothesis in a test typically states that:
The null hypothesis is the default claim of no effect that the data must provide evidence against.
2. You reject the null hypothesis when:
A p-value below the chosen significance level indicates the observed result is unlikely under the null, so you reject it.
3. A Type I error occurs when you:
A Type I error is a false positive: rejecting the null when it is actually true, occurring at rate alpha.
Flash Cards
What is the null hypothesis (H0)? — The default claim that there is no effect or no difference, assumed true until evidence contradicts it.
What is the alternative hypothesis (H1)? — The claim that there is a real effect or difference, accepted only if the null is rejected.
What is a p-value? — The probability of observing a result at least as extreme as the data, assuming the null hypothesis is true.
Type I vs Type II error — Type I is rejecting a true null (false positive); Type II is failing to reject a false null (false negative).
Can you prove the null is true? — No. You can only fail to reject it; absence of evidence against it is not proof of it.