Lesson 17 / 25

Hypothesis Tests and Confidence Intervals

Run and interpret common hypothesis tests in R.

Testing claims with data

A hypothesis test asks whether data are consistent with a null hypothesis (no difference, no association). The p-value is the probability of observing results at least as extreme as yours if the null hypothesis were true; a small p-value (commonly below 0.05) is evidence against the null, but it is not the probability that the null is true and says nothing about the size or importance of an effect. Always report effect sizes and confidence intervals alongside p-values. Common tests in R: t.test() for comparing means (one sample, two independent samples with Welch's correction by default, or paired samples); wilcox.test() as a non-parametric alternative; chisq.test() for association between categorical variables in a contingency table; prop.test() for comparing proportions, such as A/B test conversion rates; cor.test() for correlation; and aov() for comparing several group means (ANOVA). Check assumptions (independence, approximate normality, equal variances where needed), beware of running many tests (multiple comparisons inflate false positives, so adjust with p.adjust()), and decide on your analysis before looking at the results.

An A/B test and a two-sample comparison

prop.test for conversion rates; t.test for average order value.

# A/B test: did the new checkout page change the conversion rate?
conversions <- c(new = 312, old = 270)
visitors <- c(new = 4000, old = 4000)
prop.test(conversions, visitors)
# reports both proportions (7.8% vs 6.75%), a 95% confidence interval for the
# difference and a p-value; check whether the interval excludes zero

# average order value for two cities (Welch t-test by default)
pune <- sales$amount_inr[sales$city == "Pune"]
delhi <- sales$amount_inr[sales$city == "Delhi"]
result <- t.test(pune, delhi)
result$conf.int        # confidence interval for the difference in means
result$p.value

# association between city and payment method
tab <- table(sales$city, sales$payment_method)
chisq.test(tab)

# several tests at once: adjust p-values
p.adjust(c(0.01, 0.04, 0.03, 0.20), method = "holm")

A courtroom

The null hypothesis is "innocent until proven guilty". A small p-value means the evidence would be very surprising if the defendant were innocent, but it does not tell you how serious the crime is (effect size).

Quick check: What does a p-value of 0.03 mean?

  • There is a 3% chance the null hypothesis is true
  • If the null hypothesis were true, results at least this extreme would occur about 3% of the time
  • The effect is large
  • The result is certainly important
Answer

If the null hypothesis were true, results at least this extreme would occur about 3% of the time — p-values are computed assuming the null is true; they are not the probability of the null.