# Performance and Big Data in R — R Programming

Source: https://www.geekswithgeeks.com/en/r-programming/a-performance

> Speed up R code and handle data larger than memory.

## Making R fast enough

R can be slow when used against its grain, but fast when used well. Principles: **vectorise** (built-in vectorised functions run in compiled C or Fortran); **pre-allocate** instead of growing objects; avoid unnecessary copies (R uses **copy-on-modify**, so modifying a column of a large data frame in a loop can copy it repeatedly); and **measure** with `system.time()`, the **bench** package (`bench::mark()`) and the **profvis** profiler before optimising. For large tabular data, **data.table** is extremely fast and memory-efficient for grouping, joining and updating by reference; **duckdb** and **duckplyr** run fast SQL-style analytics, including on Parquet files; and **arrow** queries datasets larger than memory with dplyr verbs, reading only needed columns and row groups. For heavy computation, **parallel processing** with the **future** and **furrr** packages spreads work over cores, and **Rcpp** lets you write hot loops in C++. Often the biggest win is simply doing less: filter early, select only needed columns, and aggregate in the database.

## Faster grouping and larger-than-memory data

data.table, arrow with dplyr and a quick benchmark.

```r
library(data.table)

dt <- fread("data/clickstream.csv")                 # fast CSV reader
dt[, .(sessions = .N, revenue = sum(amount)), by = .(city, device)]   # grouped summary
dt[amount > 1000, segment := "high_value"]           # update by reference, no copy

library(arrow)
library(dplyr)

events <- open_dataset("data/events_parquet/")      # a folder of Parquet files, not loaded into RAM
events |>
  filter(event_date >= as.Date("2026-09-01"), event_type == "purchase") |>
  summarise(revenue = sum(amount), .by = city) |>
  collect()                                           # only the small result enters memory

bench::mark(
  loop = { s <- 0; for (x in dt$amount) s <- s + x; s },
  vectorised = sum(dt$amount),
  check = FALSE
)
```

## Profile before optimising

Guessing where code is slow is usually wrong. `profvis::profvis({ ... })` shows exactly which lines take the time, and the fix is often a single vectorised call or an earlier filter.

**Quiz:** Which approach lets dplyr code query a folder of Parquet files larger than memory?

- [ ] read.csv on each file
- [ ] Converting everything to a matrix
- [x] arrow::open_dataset with dplyr verbs and collect() at the end
- [ ] Using for loops

*Answer:* arrow::open_dataset with dplyr verbs and collect() at the end. arrow datasets push filtering and aggregation down to efficient readers before collecting results.
