Lesson 16 / 25

Descriptive Statistics and Distributions

Summarise data and work with probability distributions and simulation.

Describing data and randomness

Descriptive statistics summarise data: centre (mean(), median()), spread (sd(), var(), IQR(), range()), position (quantile()) and shape (skewness, seen in histograms). Use the median and IQR for skewed data such as incomes or order values, where a few large values pull the mean up. R has a consistent naming scheme for probability distributions: for the normal distribution, dnorm() gives the density, pnorm() the cumulative probability, qnorm() the quantile and rnorm() random draws; the same prefixes work for binom, pois, unif, exp, t, chisq and others. Simulation with sample() and r* functions is a powerful way to understand probability and uncertainty, for example bootstrapping a confidence interval by resampling your data many times. Always call set.seed() before random operations so results are reproducible. The central limit theorem, which you can demonstrate by simulation, explains why sample means are approximately normal for large samples.

Centre and spread

A distribution summarised by its centre and spread; skewed data pulls the mean away from the median.

A right-skewed curve with two vertical markers close together for median and further right for mean, and a bracket showing spread.
Figure 6.1 — Mean, median and spread in a skewed distribution.

Summaries, distributions and a bootstrap

Reproducible random numbers with set.seed.

orders <- c(250, 400, 420, 480, 510, 600, 650, 900, 1200, 9800)

mean(orders)        # 1521: pulled up by one large order
median(orders)      # 555
IQR(orders)
quantile(orders, c(0.1, 0.5, 0.9))

pnorm(1.96)         # about 0.975: P(Z <= 1.96)
qnorm(0.975)        # about 1.96
dbinom(3, size = 10, prob = 0.2)   # P(exactly 3 successes in 10 trials)

set.seed(42)
boot_medians <- replicate(5000, median(sample(orders, replace = TRUE)))
quantile(boot_medians, c(0.025, 0.975))   # 95% bootstrap interval for the median

set.seed(1)
sample_means <- replicate(2000, mean(rexp(50, rate = 1)))
hist(sample_means)  # roughly normal, centred near 1 (central limit theorem)

Always set a seed

Without set.seed(), a simulation or random split gives different numbers each run, making results impossible to reproduce or review. Set it once near the top of the script.

Quick check: Which R function gives the cumulative probability P(X <= x) for a normal distribution?

  • dnorm
  • qnorm
  • pnorm
  • rnorm
Answer

pnorm — p-functions return cumulative probabilities; d is density, q is quantile, r is random draws.