पाठ 7 / 27

Pagination: Offset बनाम Cursor

बड़े collections को सुरक्षित रूप से page करें और जानें कि बदलते डेटा के लिए cursors offsets से बेहतर क्यों हैं।

सब कुछ कभी न लौटाएँ

Collections paginated होने चाहिए, उचित डिफ़ॉल्ट और अधिकतम limit के साथ। Offset pagination (?offset=20&limit=10) सरल है और किसी पन्ने पर कूदने देता है, पर बड़ी tables पर धीमा है (database rows छोड़ता है) और अस्थिर है: दो अनुरोधों के बीच row जुड़े या हटे तो items दो बार दिख सकते हैं या छूट सकते हैं। Cursor (keyset) pagination (?cursor=...) आख़िरी देखे item की ओर इशारा करने वाला अपारदर्शी token लौटाती है, और अगली query उसके "बाद" के items माँगती है; यह बदलते डेटा के लिए तेज़ और स्थिर है, पर पन्ना N पर कूद नहीं सकती। Cursors अपारदर्शी रखें (clients उन्हें parse न करें), next_cursor लौटाएँ (अंत में null), और अनोखी, स्थिर key से क्रम दें।

असली API पर cursor paging, चलाकर

मैंने यह केवल Python standard library से बने छोटे असली HTTP API पर चलाया (पूरा कोड केस स्टडी में)। पन्ना 1 items 1 और 2 और एक next_cursor लौटाता है; उसे देने पर items 3 और 4 मिलते हैं; आख़िरी पन्ना item 5 और null cursor लौटाता है। यहाँ cursor सिर्फ़ base64-encoded JSON है, पर clients को उसे अपारदर्शी मानना चाहिए।

# uses srv, base and call() from the runnable demo in the case study
s, h, b = call("GET", "/v1/items?limit=2"); print(s, b)
s, h, b = call("GET", "/v1/items?limit=2&cursor=" + b["next_cursor"]); print(s, b)
s, h, b = call("GET", "/v1/items?limit=2&cursor=" + b["next_cursor"]); print(s, b)

Output:

200 {'data': [{'id': 1, 'plan': 'pro'}, {'id': 2, 'plan': 'free'}], 'next_cursor': 'eyJhZnRlciI6IDJ9'}
200 {'data': [{'id': 3, 'plan': 'pro'}, {'id': 4, 'plan': 'free'}], 'next_cursor': 'eyJhZnRlciI6IDR9'}
200 {'data': [{'id': 5, 'plan': 'pro'}], 'next_cursor': None}

Offsets क्यों खिसकते हैं, चलाकर

मैंने यह सादा-Python उदाहरण चलाया। पन्ना 1 [1, 2, 3] दिखाता है। फिर आगे एक नई row जुड़ती है। Offset पन्ना 2 (नई सूची की rows 3 से 5) item 3 दोहराता है, जबकि cursor संस्करण 4, 5, 6 के साथ सही जारी रखता है।

import hmac, hashlib, json, time, bisect

rows = list(range(1, 11))
page1 = rows[0:3]
rows = [0] + rows                     # a new row is inserted at the front between requests
offset_page2 = rows[3:6]
cursor_page2 = [r for r in rows if r > page1[-1]][:3]
print(page1, offset_page2, cursor_page2)

Output:

[1, 2, 3] [3, 4, 5] [4, 5, 6]

त्वरित जाँच: Cursor pagination, offset pagination से ज़्यादा स्थिर क्यों है?

  • यह आख़िरी देखे item के बाद आगे बढ़ती है, इसलिए inserts या deletes पन्ने नहीं खिसकाते
  • यह random संख्याएँ उपयोग करती है
  • यह हर row एक साथ लौटाती है
  • Offsets अवैध हैं
Answer

यह आख़िरी देखे item के बाद आगे बढ़ती है, इसलिए inserts या deletes पन्ने नहीं खिसकाते — किसी item से जुड़ी स्थिति सूची में कहीं और हुए बदलावों में भी टिकती है।