पाठ 17 / 25
विफलता परिदृश्य और Leader Election
देखें कि broker, network link या consumer विफल होने पर क्या होता है।
क्या अपेक्षा करें
Broker बंद: controller हर प्रभावित partition के लिए उसके ISR से नया leader चुनता है; clients metadata ताज़ा करके retry करते हैं, इसलिए छोटा ठहराव दिखता है, डेटा हानि नहीं (ऊपर की टिकाऊ settings के साथ)। Follower पिछड़ता है: पकड़ बनाने तक उसे ISR से हटा दिया जाता है, और acks=all writes कम replicas का इंतज़ार करते हैं, इसीलिए min.insync.replicas मायने रखता है। Network partition: जो clients leader तक नहीं पहुँच पाते वे timeout होते हैं; controller से कटे leaders बदल दिए जाते हैं। Consumer मरता है: session timeout के बाद उसके partitions पुनः बाँटे जाते हैं, और नया मालिक आख़िरी committed offset से फिर शुरू करता है (कुछ records दोबारा प्रोसेस हो सकते हैं)। Producer retries: idempotence दोहराव रोकता है। इन विफलताओं का अभ्यास test cluster में करें, पहली बार किसी incident में नहीं।
विफलता अभ्यास की checklist
हर अभ्यास test परिवेश में चलाएँ और लिखें कि आपने क्या देखा।
Drill Expect
stop one broker leader election within seconds; no data loss; producers retry
stop two brokers (rf=3,isr=2) writes to affected partitions fail loudly; reads may continue
kill a consumer rebalance; lag rises then recovers; duplicates possible
throttle a follower shrinks ISR; under-replicated partitions alert fires
restart the whole cluster rolling restart: one broker at a time, wait for ISR to recoverRolling restart: एक बार में एक
Brokers को एक बार में एक restart करें और अगले को छूने से पहले तब तक रुकें जब तक under-replicated partitions शून्य न हो जाएँ, वरना आप min.insync.replicas से नीचे जा सकते हैं।
त्वरित जाँच: Broker विफल होता है। नया partition leader कहाँ से आता है?
- Producer से
- किसी random consumer से
- उस partition के in-sync replicas से
- नया partition बनाया जाता है
Answer
उस partition के in-sync replicas से — सिर्फ़ in-sync replicas योग्य हैं, इसलिए नए leader के पास सभी स्वीकृत records होते हैं।