होम / Apache Spark: DataFrames से Big Data Processing · English
Apache Spark: DataFrames से Big Data Processing
PySpark के साथ Apache Spark सीखें: architecture, DataFrames, SQL, joins, windows, lazy execution, shuffles, performance ट्यूनिंग, streaming और production deployment।
कोर्स शुरू करें →
पाठ्यक्रम Spark की बुनियाद Spark क्या है Architecture: Driver, Executors, Jobs, Stages, Tasks Spark चलाना और SparkSession बनाना RDDs, DataFrames और Spark SQL
DataFrame API Selecting, Filtering और Columns Aggregations और Null सँभालना Joins Window Functions
Spark SQL और Files Spark SQL और Temporary Views File Formats: CSV, JSON, Parquet Partitioned Output और Partition Pruning
Spark आपका कोड कैसे चलाता है Transformations, Actions और Lazy Evaluation Narrow और Wide Transformations: Shuffle explain() से Plans पढ़ना और Adaptive Query Execution Partitions: repartition, coalesce और Skew
Performance ट्यूनिंग Caching और Persistence Join रणनीतियाँ और Broadcast Joins Data Layout, छोटी Files और Configuration
Structured Streaming Structured Streaming मॉडल Event-Time Windows और Watermarks
Production Spark Cluster Managers और Platforms Spark Jobs का Testing Spark UI से Monitoring और Lakehouse Tables
सब कुछ जोड़ना केस स्टडी: दैनिक Sales Pipeline Revision: Cheat Sheet और Self-Check
GeeksWithGeeks — Python, Java, JavaScript, TypeScript, Angular, Node.js, DSA, सिस्टम डिज़ाइन, डेटाबेस और Docker के मुफ़्त, व्यावहारिक ट्यूटोरियल — पाठ या 30 सेकंड के शॉर्ट्स के रूप में, अंग्रेज़ी और हिंदी में।