Lesson 7 / 27

Pagination: Offset vs Cursor

Page large collections safely and know why cursors beat offsets for changing data.

Never return everything

Collections must be paginated, with a sensible default and a maximum limit. Offset pagination (?offset=20&limit=10) is simple and allows jumping to a page, but it is slow on big tables (the database skips rows) and unstable: if a row is inserted or deleted between two requests, items can be shown twice or skipped. Cursor (keyset) pagination (?cursor=...) returns an opaque token pointing to the last item seen, and the next query asks for items "after" it; it is fast and stable for changing data, but cannot jump to page N. Make cursors opaque (clients must not parse them), return next_cursor (null at the end), and sort by a unique, stable key.

Cursor paging against the real API, run

I ran this against a small real HTTP API built with only the Python standard library (full code in the case study). Page 1 returns items 1 and 2 plus a next_cursor; passing it returns items 3 and 4; the last page returns item 5 and a null cursor. The cursor is just base64-encoded JSON here, but clients should treat it as opaque.

# uses srv, base and call() from the runnable demo in the case study
s, h, b = call("GET", "/v1/items?limit=2"); print(s, b)
s, h, b = call("GET", "/v1/items?limit=2&cursor=" + b["next_cursor"]); print(s, b)
s, h, b = call("GET", "/v1/items?limit=2&cursor=" + b["next_cursor"]); print(s, b)

Output:

200 {'data': [{'id': 1, 'plan': 'pro'}, {'id': 2, 'plan': 'free'}], 'next_cursor': 'eyJhZnRlciI6IDJ9'}
200 {'data': [{'id': 3, 'plan': 'pro'}, {'id': 4, 'plan': 'free'}], 'next_cursor': 'eyJhZnRlciI6IDR9'}
200 {'data': [{'id': 5, 'plan': 'pro'}], 'next_cursor': None}

Why offsets drift, run

I ran this plain-Python example. Page 1 shows [1, 2, 3]. A new row is then inserted at the front. Offset page 2 (rows 3 to 5 of the new list) repeats item 3, while the cursor version continues correctly with 4, 5, 6.

import hmac, hashlib, json, time, bisect

rows = list(range(1, 11))
page1 = rows[0:3]
rows = [0] + rows                     # a new row is inserted at the front between requests
offset_page2 = rows[3:6]
cursor_page2 = [r for r in rows if r > page1[-1]][:3]
print(page1, offset_page2, cursor_page2)

Output:

[1, 2, 3] [3, 4, 5] [4, 5, 6]

Quick check: Why is cursor pagination more stable than offset pagination?

  • It continues after the last seen item, so inserts or deletes do not shift pages
  • It uses random numbers
  • It returns every row at once
  • Offsets are illegal
Answer

It continues after the last seen item, so inserts or deletes do not shift pages — A position anchored to an item survives changes elsewhere in the list.