पाठ 17 / 25

विफलता परिदृश्य और Leader Election

देखें कि broker, network link या consumer विफल होने पर क्या होता है।

क्या अपेक्षा करें

Broker बंद: controller हर प्रभावित partition के लिए उसके ISR से नया leader चुनता है; clients metadata ताज़ा करके retry करते हैं, इसलिए छोटा ठहराव दिखता है, डेटा हानि नहीं (ऊपर की टिकाऊ settings के साथ)। Follower पिछड़ता है: पकड़ बनाने तक उसे ISR से हटा दिया जाता है, और acks=all writes कम replicas का इंतज़ार करते हैं, इसीलिए min.insync.replicas मायने रखता है। Network partition: जो clients leader तक नहीं पहुँच पाते वे timeout होते हैं; controller से कटे leaders बदल दिए जाते हैं। Consumer मरता है: session timeout के बाद उसके partitions पुनः बाँटे जाते हैं, और नया मालिक आख़िरी committed offset से फिर शुरू करता है (कुछ records दोबारा प्रोसेस हो सकते हैं)। Producer retries: idempotence दोहराव रोकता है। इन विफलताओं का अभ्यास test cluster में करें, पहली बार किसी incident में नहीं।

विफलता अभ्यास की checklist

हर अभ्यास test परिवेश में चलाएँ और लिखें कि आपने क्या देखा।

Drill                         Expect
stop one broker               leader election within seconds; no data loss; producers retry
stop two brokers (rf=3,isr=2) writes to affected partitions fail loudly; reads may continue
kill a consumer               rebalance; lag rises then recovers; duplicates possible
throttle a follower           shrinks ISR; under-replicated partitions alert fires
restart the whole cluster     rolling restart: one broker at a time, wait for ISR to recover

Rolling restart: एक बार में एक

Brokers को एक बार में एक restart करें और अगले को छूने से पहले तब तक रुकें जब तक under-replicated partitions शून्य न हो जाएँ, वरना आप min.insync.replicas से नीचे जा सकते हैं।

त्वरित जाँच: Broker विफल होता है। नया partition leader कहाँ से आता है?

  • Producer से
  • किसी random consumer से
  • उस partition के in-sync replicas से
  • नया partition बनाया जाता है
Answer

उस partition के in-sync replicas से — सिर्फ़ in-sync replicas योग्य हैं, इसलिए नए leader के पास सभी स्वीकृत records होते हैं।