# Spark चलाना और SparkSession बनाना — Apache Spark: DataFrames से Big Data Processing

Source: https://www.geekswithgeeks.com/hi/spark/sp-install-session

> PySpark को स्थानीय रूप से या Docker में चलाएँ और SparkSession बनाएँ।

## सीखने के लिए local mode काफ़ी है

सीखने के लिए Spark को **local mode** में चलाएँ, जहाँ driver और executors एक process के threads हैं: `pip install pyspark` (Java runtime चाहिए) या official Docker image `apache/spark` उपयोग करें। हर प्रोग्राम **`SparkSession`** बनाकर शुरू होता है, जो DataFrames और SQL का प्रवेश-बिंदु है; `master("local[2]")` का मतलब दो worker threads। Production में आप `spark-submit` से applications को cluster manager को सौंपते हैं, और master setting submit command से आती है, कोड से नहीं। Spark का डिफ़ॉल्ट log output बहुत लंबा होता है, इसलिए सीखते समय log level `ERROR` रखें।

## Session और पहला DataFrame, चलाकर

मैंने यह Apache Spark 4.0.0 (PySpark, local mode, official Docker image में) पर चलाया। Version पंक्ति engine की पुष्टि करती है। छह rows वाला `orders` DataFrame (एक amount जानबूझकर null है) अगले खंडों में उपयोग होता है।

```python
from pyspark.sql import SparkSession, functions as F, Window
spark = (SparkSession.builder.master("local[2]").appName("demo")
         .config("spark.sql.shuffle.partitions", "4").config("spark.ui.enabled", "false").getOrCreate())
spark.sparkContext.setLogLevel("ERROR")
print("version", spark.version)

orders = spark.createDataFrame([
    (1, "asha", "pune", 120.0, "2026-09-01"), (2, "ravi", "delhi", 80.0, "2026-09-01"),
    (3, "asha", "pune", 200.0, "2026-09-02"), (4, "meera", "delhi", 50.0, "2026-09-02"),
    (5, "ravi", "delhi", 300.0, "2026-09-03"), (6, "kiran", "mumbai", None, "2026-09-03"),
], ["order_id", "customer", "city", "amount", "order_date"])
orders = orders.withColumn("order_date", F.to_date("order_date"))

```

Output:

```
version 4.0.0
```

## Cluster को submit करना (उदाहरण)

Production में यही script `spark-submit` से submit होती है। Option मान आपके cluster पर निर्भर हैं।

```bash
spark-submit \
  --master yarn --deploy-mode cluster \
  --num-executors 10 --executor-cores 4 --executor-memory 8g \
  --conf spark.sql.shuffle.partitions=400 \
  daily_orders_job.py --date 2026-10-01
```

**Quiz:** PySpark में DataFrames और SQL का प्रवेश-बिंदु क्या है?

- [ ] pandas.read()
- [ ] SparkPlug
- [ ] DataBase()
- [x] SparkSession

*Answer:* SparkSession. SparkSession पहले बनता है और डेटा पढ़ने व SQL चलाने में उपयोग होता है।
