# Revision: Cheat Sheet और Self-Check — Apache Spark: DataFrames से Big Data Processing

Source: https://www.geekswithgeeks.com/hi/spark/wrap-revision

> पूरे कोर्स की Spark की अवधारणाओं, APIs और ट्यूनिंग आदतों की दोहराई करें।

## Cheat sheet

**Architecture**: driver योजना बनाता है; executors partitions पर tasks चलाते हैं; action → job → stages (shuffles पर बँटे) → tasks। **API**: SparkSession; select/filter/withColumn/groupBy/agg/join/window वाले DataFrames (immutable); UDFs की जगह built-ins; SQL = DataFrame API। **निष्पादन**: transformations lazy, actions चलाते हैं; narrow बनाम wide; `Exchange` = shuffle; `explain()`; AQE। **Partitions**: ~100-200 MB; `repartition` (shuffle) बनाम `coalesce` (मिलाना); skew ठीक करें (AQE, salting, broadcast)। **Performance**: बार-बार उपयोग होने वाले डेटा को cache, छोटे joins broadcast, Parquet + partition pruning, छोटी files से बचें, कुछ configs ट्यून करें। **Streaming**: असीमित table, trigger, output mode, checkpoint, event-time windows + watermark। **Production**: object storage, YARN/Kubernetes/managed, function-शैली कोड + tests + डेटा-गुणवत्ता जाँचें, Spark UI के लक्षण, Delta/Iceberg, idempotent partition overwrite।

## इंटरव्यू में पूछे जाने वाले सवाल

इन्हें समझाने को तैयार रहें: transformations और actions में अंतर तथा lazy evaluation, shuffle किससे होता है और उसे कैसे घटाएँ, narrow बनाम wide dependencies, repartition बनाम coalesce, skewed join कैसे सँभालेंगे, कब broadcast करें, caching कैसे काम करती है, और Spark UI से धीमे Spark job को कैसे debug करेंगे।

**Quiz:** Job एक महँगे DataFrame पर बिना caching के `df.count()` दो बार बुलाता है। क्या होता है?

- [ ] Spark error देता है
- [ ] दूसरी call मुफ़्त है
- [x] पूरी lineage दो बार गणना होती है
- [ ] नतीजा अपने आप cache हो जाता है

*Answer:* पूरी lineage दो बार गणना होती है. हर action स्रोत से दोबारा गणना करता है जब तक DataFrame cache या persist न हो।

**Quiz:** 2 TB की table और 5 MB की lookup table के बीच join धीमा है। सबसे अच्छा पहला विचार क्या है?

- [x] 5 MB की table broadcast करें
- [ ] 2 TB की table driver पर collect करें
- [ ] सिर्फ़ और shuffle partitions जोड़ें
- [ ] दोनों को CSV में बदलें

*Answer:* 5 MB की table broadcast करें. छोटी ओर को broadcast करने से विशाल table का shuffle बचता है।

**Quiz:** coalesce और repartition के बारे में कौन-सा कथन सही है?

- [ ] repartition कभी डेटा नहीं हिलाता
- [ ] दोनों एक-जैसे हैं
- [ ] coalesce पूरे shuffle के साथ partitions बढ़ा सकता है
- [x] coalesce बिना पूरे shuffle के सिर्फ़ partitions घटाता है; repartition ठीक संख्या के लिए shuffle करता है

*Answer:* coalesce बिना पूरे shuffle के सिर्फ़ partitions घटाता है; repartition ठीक संख्या के लिए shuffle करता है. coalesce घटाने के लिए सस्ता मिलान है; repartition सारा डेटा समान रूप से पुनर्वितरित करता है।
