पाठ 15 / 25

Exploratory Data Analysis

Explore a new dataset systematically: distributions, relationships, missing data and outliers.

Asking questions of data

Exploratory data analysis (EDA) is the iterative process of understanding a dataset before modelling or reporting. A practical checklist: understand structure (glimpse, summary, row counts, keys); check missing values per column (colSums(is.na(df)) or naniar), and ask why they are missing; look at each variable's distribution (histograms, count() for categories), noticing skew, impossible values and outliers; examine relationships between variables (scatter plots, box plots by group, correlation with cor()), remembering that correlation is not causation; look at changes over time; and compare groups. Record findings and decisions (removing duplicates, capping outliers, excluding test orders) in code, not by editing files by hand, so the analysis is reproducible. Packages such as skimr (skim()) give compact summaries of every column, and quick plots help more than tables for spotting problems. EDA often sends you back to data cleaning, which is normal.

A quick EDA pass

Missing values, distributions, outliers and relationships.

library(dplyr)
library(ggplot2)

glimpse(sales)
skimr::skim(sales)                              # one-line summary per column

colSums(is.na(sales))                           # missing values per column
sales |> count(city, sort = TRUE)               # category frequencies

# outliers with the 1.5 x IQR rule
q <- quantile(sales$amount_inr, c(0.25, 0.75), na.rm = TRUE)
iqr <- diff(q)
outliers <- sales |> filter(amount_inr > q[2] + 1.5 * iqr | amount_inr < q[1] - 1.5 * iqr)
nrow(outliers)

# compare groups
ggplot(sales, aes(x = city, y = amount_inr)) +
  geom_boxplot() +
  scale_y_log10()

# relationship between two numeric variables
cor(sales$items, sales$amount_inr, use = "complete.obs")

Plot before you summarise

Very different datasets can share the same mean, variance and correlation (the famous Anscombe's quartet). A quick plot reveals patterns and problems that summary numbers hide.

त्वरित जाँच: What is the main goal of exploratory data analysis?

  • To produce the final published report immediately
  • To train deep learning models
  • To understand structure, quality, distributions and relationships in the data before modelling
  • To delete missing data
Answer

To understand structure, quality, distributions and relationships in the data before modelling — EDA builds understanding and uncovers data problems early.