# Descriptive Statistics and Distributions — R Programming

Source: https://www.geekswithgeeks.com/en/r-programming/s-describe

> Summarise data and work with probability distributions and simulation.

## Describing data and randomness

Descriptive statistics summarise data: **centre** (`mean()`, `median()`), **spread** (`sd()`, `var()`, `IQR()`, `range()`), **position** (`quantile()`) and **shape** (skewness, seen in histograms). Use the median and IQR for **skewed** data such as incomes or order values, where a few large values pull the mean up. R has a consistent naming scheme for **probability distributions**: for the normal distribution, `dnorm()` gives the density, `pnorm()` the cumulative probability, `qnorm()` the quantile and `rnorm()` random draws; the same prefixes work for `binom`, `pois`, `unif`, `exp`, `t`, `chisq` and others. **Simulation** with `sample()` and `r*` functions is a powerful way to understand probability and uncertainty, for example bootstrapping a confidence interval by resampling your data many times. Always call **`set.seed()`** before random operations so results are reproducible. The **central limit theorem**, which you can demonstrate by simulation, explains why sample means are approximately normal for large samples.

## Centre and spread

A distribution summarised by its centre and spread; skewed data pulls the mean away from the median.

![A right-skewed curve with two vertical markers close together for median and further right for mean, and a bracket showing spread.](assets/figures/r-programming/section-6-map.svg) — Figure 6.1 — Mean, median and spread in a skewed distribution.

## Summaries, distributions and a bootstrap

Reproducible random numbers with set.seed.

```r
orders <- c(250, 400, 420, 480, 510, 600, 650, 900, 1200, 9800)

mean(orders)        # 1521: pulled up by one large order
median(orders)      # 555
IQR(orders)
quantile(orders, c(0.1, 0.5, 0.9))

pnorm(1.96)         # about 0.975: P(Z <= 1.96)
qnorm(0.975)        # about 1.96
dbinom(3, size = 10, prob = 0.2)   # P(exactly 3 successes in 10 trials)

set.seed(42)
boot_medians <- replicate(5000, median(sample(orders, replace = TRUE)))
quantile(boot_medians, c(0.025, 0.975))   # 95% bootstrap interval for the median

set.seed(1)
sample_means <- replicate(2000, mean(rexp(50, rate = 1)))
hist(sample_means)  # roughly normal, centred near 1 (central limit theorem)
```

## Always set a seed

Without `set.seed()`, a simulation or random split gives different numbers each run, making results impossible to reproduce or review. Set it once near the top of the script.

**Quiz:** Which R function gives the cumulative probability P(X <= x) for a normal distribution?

- [ ] dnorm
- [ ] qnorm
- [x] pnorm
- [ ] rnorm

*Answer:* pnorm. p-functions return cumulative probabilities; d is density, q is quantile, r is random draws.
