# Exploratory Data Analysis — R Programming

Source: https://www.geekswithgeeks.com/en/r-programming/v-eda

> Explore a new dataset systematically: distributions, relationships, missing data and outliers.

## Asking questions of data

**Exploratory data analysis (EDA)** is the iterative process of understanding a dataset before modelling or reporting. A practical checklist: understand **structure** (`glimpse`, `summary`, row counts, keys); check **missing values** per column (`colSums(is.na(df))` or `naniar`), and ask why they are missing; look at each variable's **distribution** (histograms, `count()` for categories), noticing skew, impossible values and **outliers**; examine **relationships** between variables (scatter plots, box plots by group, correlation with `cor()`), remembering that correlation is not causation; look at **changes over time**; and compare **groups**. Record findings and decisions (removing duplicates, capping outliers, excluding test orders) in code, not by editing files by hand, so the analysis is reproducible. Packages such as **skimr** (`skim()`) give compact summaries of every column, and quick plots help more than tables for spotting problems. EDA often sends you back to data cleaning, which is normal.

## A quick EDA pass

Missing values, distributions, outliers and relationships.

```r
library(dplyr)
library(ggplot2)

glimpse(sales)
skimr::skim(sales)                              # one-line summary per column

colSums(is.na(sales))                           # missing values per column
sales |> count(city, sort = TRUE)               # category frequencies

# outliers with the 1.5 x IQR rule
q <- quantile(sales$amount_inr, c(0.25, 0.75), na.rm = TRUE)
iqr <- diff(q)
outliers <- sales |> filter(amount_inr > q[2] + 1.5 * iqr | amount_inr < q[1] - 1.5 * iqr)
nrow(outliers)

# compare groups
ggplot(sales, aes(x = city, y = amount_inr)) +
  geom_boxplot() +
  scale_y_log10()

# relationship between two numeric variables
cor(sales$items, sales$amount_inr, use = "complete.obs")
```

## Plot before you summarise

Very different datasets can share the same mean, variance and correlation (the famous Anscombe's quartet). A quick plot reveals patterns and problems that summary numbers hide.

**Quiz:** What is the main goal of exploratory data analysis?

- [ ] To produce the final published report immediately
- [ ] To train deep learning models
- [x] To understand structure, quality, distributions and relationships in the data before modelling
- [ ] To delete missing data

*Answer:* To understand structure, quality, distributions and relationships in the data before modelling. EDA builds understanding and uncovers data problems early.
