Lesson 21 / 25

Performance and Big Data in R

Speed up R code and handle data larger than memory.

Making R fast enough

R can be slow when used against its grain, but fast when used well. Principles: vectorise (built-in vectorised functions run in compiled C or Fortran); pre-allocate instead of growing objects; avoid unnecessary copies (R uses copy-on-modify, so modifying a column of a large data frame in a loop can copy it repeatedly); and measure with system.time(), the bench package (bench::mark()) and the profvis profiler before optimising. For large tabular data, data.table is extremely fast and memory-efficient for grouping, joining and updating by reference; duckdb and duckplyr run fast SQL-style analytics, including on Parquet files; and arrow queries datasets larger than memory with dplyr verbs, reading only needed columns and row groups. For heavy computation, parallel processing with the future and furrr packages spreads work over cores, and Rcpp lets you write hot loops in C++. Often the biggest win is simply doing less: filter early, select only needed columns, and aggregate in the database.

Faster grouping and larger-than-memory data

data.table, arrow with dplyr and a quick benchmark.

library(data.table)

dt <- fread("data/clickstream.csv")                 # fast CSV reader
dt[, .(sessions = .N, revenue = sum(amount)), by = .(city, device)]   # grouped summary
dt[amount > 1000, segment := "high_value"]           # update by reference, no copy

library(arrow)
library(dplyr)

events <- open_dataset("data/events_parquet/")      # a folder of Parquet files, not loaded into RAM
events |>
  filter(event_date >= as.Date("2026-09-01"), event_type == "purchase") |>
  summarise(revenue = sum(amount), .by = city) |>
  collect()                                           # only the small result enters memory

bench::mark(
  loop = { s <- 0; for (x in dt$amount) s <- s + x; s },
  vectorised = sum(dt$amount),
  check = FALSE
)

Profile before optimising

Guessing where code is slow is usually wrong. profvis::profvis({ ... }) shows exactly which lines take the time, and the fix is often a single vectorised call or an earlier filter.

Quick check: Which approach lets dplyr code query a folder of Parquet files larger than memory?

  • read.csv on each file
  • Converting everything to a matrix
  • arrow::open_dataset with dplyr verbs and collect() at the end
  • Using for loops
Answer

arrow::open_dataset with dplyr verbs and collect() at the end — arrow datasets push filtering and aggregation down to efficient readers before collecting results.