पाठ 16 / 25
Descriptive Statistics and Distributions
Summarise data and work with probability distributions and simulation.
Describing data and randomness
Descriptive statistics summarise data: centre (mean(), median()), spread (sd(), var(), IQR(), range()), position (quantile()) and shape (skewness, seen in histograms). Use the median and IQR for skewed data such as incomes or order values, where a few large values pull the mean up. R has a consistent naming scheme for probability distributions: for the normal distribution, dnorm() gives the density, pnorm() the cumulative probability, qnorm() the quantile and rnorm() random draws; the same prefixes work for binom, pois, unif, exp, t, chisq and others. Simulation with sample() and r* functions is a powerful way to understand probability and uncertainty, for example bootstrapping a confidence interval by resampling your data many times. Always call set.seed() before random operations so results are reproducible. The central limit theorem, which you can demonstrate by simulation, explains why sample means are approximately normal for large samples.
Centre and spread
A distribution summarised by its centre and spread; skewed data pulls the mean away from the median.
Summaries, distributions and a bootstrap
Reproducible random numbers with set.seed.
orders <- c(250, 400, 420, 480, 510, 600, 650, 900, 1200, 9800)
mean(orders) # 1521: pulled up by one large order
median(orders) # 555
IQR(orders)
quantile(orders, c(0.1, 0.5, 0.9))
pnorm(1.96) # about 0.975: P(Z <= 1.96)
qnorm(0.975) # about 1.96
dbinom(3, size = 10, prob = 0.2) # P(exactly 3 successes in 10 trials)
set.seed(42)
boot_medians <- replicate(5000, median(sample(orders, replace = TRUE)))
quantile(boot_medians, c(0.025, 0.975)) # 95% bootstrap interval for the median
set.seed(1)
sample_means <- replicate(2000, mean(rexp(50, rate = 1)))
hist(sample_means) # roughly normal, centred near 1 (central limit theorem)Always set a seed
Without set.seed(), a simulation or random split gives different numbers each run, making results impossible to reproduce or review. Set it once near the top of the script.
त्वरित जाँच: Which R function gives the cumulative probability P(X <= x) for a normal distribution?
- dnorm
- qnorm
- pnorm
- rnorm
Answer
pnorm — p-functions return cumulative probabilities; d is density, q is quantile, r is random draws.