Apache Spark: DataFrames से Big Data Processing

PySpark के साथ Apache Spark सीखें: architecture, DataFrames, SQL, joins, windows, lazy execution, shuffles, performance ट्यूनिंग, streaming और production deployment।

कोर्स शुरू करें →

पाठ्यक्रम

Spark की बुनियाद

  1. Spark क्या है
  2. Architecture: Driver, Executors, Jobs, Stages, Tasks
  3. Spark चलाना और SparkSession बनाना
  4. RDDs, DataFrames और Spark SQL

DataFrame API

  1. Selecting, Filtering और Columns
  2. Aggregations और Null सँभालना
  3. Joins
  4. Window Functions

Spark SQL और Files

  1. Spark SQL और Temporary Views
  2. File Formats: CSV, JSON, Parquet
  3. Partitioned Output और Partition Pruning

Spark आपका कोड कैसे चलाता है

  1. Transformations, Actions और Lazy Evaluation
  2. Narrow और Wide Transformations: Shuffle
  3. explain() से Plans पढ़ना और Adaptive Query Execution
  4. Partitions: repartition, coalesce और Skew

Performance ट्यूनिंग

  1. Caching और Persistence
  2. Join रणनीतियाँ और Broadcast Joins
  3. Data Layout, छोटी Files और Configuration

Structured Streaming

  1. Structured Streaming मॉडल
  2. Event-Time Windows और Watermarks

Production Spark

  1. Cluster Managers और Platforms
  2. Spark Jobs का Testing
  3. Spark UI से Monitoring और Lakehouse Tables

सब कुछ जोड़ना

  1. केस स्टडी: दैनिक Sales Pipeline
  2. Revision: Cheat Sheet और Self-Check