# Spark Jobs का Testing — Apache Spark: DataFrames से Big Data Processing

Source: https://www.geekswithgeeks.com/hi/spark/prod-testing

> Transformations को functions के रूप में रखें और स्थानीय SparkSession से परखें।

## छोटे functions, छोटा test डेटा

Transformations को ऐसे सादे functions के रूप में लिखें जो DataFrame लें और DataFrame लौटाएँ (`def add_revenue(df): ...`), और पढ़ना-लिखना किनारों पर रखें। फिर हर function का unit-test **स्थानीय** `SparkSession` में कुछ rows से करें (प्रति test run एक साझा fixture, क्योंकि Spark शुरू करना धीमा है) और `collect()` या helper assertions से नतीजे मिलाएँ। Pipeline के भीतर **डेटा-गुणवत्ता जाँचें** जोड़ें: row counts, null दरें, keys की अनोखापन, मानों की सीमाएँ और referential जाँचें, अपेक्षाएँ टूटने पर job विफल करें (या ख़राब rows अलग करें)। यथार्थपूर्ण किनारे के मामलों पर परखें: nulls, दोहराव, ख़ाली input, timezone सीमाएँ।

## स्थानीय session वाला unit test (उदाहरण)

परखा जा रहा function शुद्ध है: DataFrame अंदर, DataFrame बाहर।

```python
import pytest
from pyspark.sql import SparkSession, functions as F

@pytest.fixture(scope="session")
def spark():
    s = SparkSession.builder.master("local[2]").config("spark.ui.enabled", "false").getOrCreate()
    yield s
    s.stop()

def add_revenue(df):
    return df.withColumn("revenue", F.col("qty") * F.col("price"))

def test_add_revenue(spark):
    df = spark.createDataFrame([(2, 10.0), (3, 5.0)], ["qty", "price"])
    out = add_revenue(df).orderBy("qty").collect()
    assert [r["revenue"] for r in out] == [20.0, 15.0]
```

**Quiz:** Transformations को DataFrames के functions के रूप में क्यों लिखें?

- [ ] यह tests की ज़रूरत हटाता है
- [ ] Spark बाक़ी सब मना करता है
- [ ] Functions परिभाषा से तेज़ चलते हैं
- [x] उन्हें छोटे स्थानीय डेटा से unit-test करना आसान है

*Answer:* उन्हें छोटे स्थानीय डेटा से unit-test करना आसान है. शुद्ध functions तर्क को I/O से अलग करते हैं, इसलिए tests को cluster नहीं चाहिए।
