Skip to content

Instantly share code, notes, and snippets.

@erdaltoprak
Created July 12, 2026 09:31
Show Gist options
  • Select an option

  • Save erdaltoprak/02926a971a2d8c5aa9037e6bf13138f7 to your computer and use it in GitHub Desktop.

Select an option

Save erdaltoprak/02926a971a2d8c5aa9037e6bf13138f7 to your computer and use it in GitHub Desktop.
Fetching All Bookmarks from the X (Twitter) API
# Fetching All Bookmarks from the X (Twitter) API
A practical guide to pulling a user's complete bookmark history via the X API v2,
including pagination, incremental sync, and a few gotchas that aren't obvious from
the docs.
## Prerequisites
- A project with **OAuth 2.0 User Context** access.
- An app with a client ID / client secret and a redirect URI configured.
- The `bookmark.read` and `users.read` (and `tweet.read`) scopes.
- A user access token obtained via the **Authorization Code with PKCE** flow.
> Bookmarks are **only** available through User Context tokens. App-only bearer
> tokens cannot read bookmarks.
## The endpoint
```
GET https://api.x.com/2/users/{user_id}/bookmarks
```
### Key query parameters
| param | purpose |
|--------------------|------------------------------------------------------|
| `max_results` | Page size. Allowed range is `5`–`100`. |
| `pagination_token` | Opaque cursor from the previous response's `meta`. |
| `tweet.fields` | Request full tweet metadata (e.g. `created_at`, `public_metrics`, `note_tweet`). |
| `expansions` | Inline related objects: `author_id`, `attachments.media_keys`, `referenced_tweets.id`. |
| `user.fields` | Extra fields on expanded authors. |
| `media.fields` | Extra fields on expanded media. |
Request as many `tweet.fields` and `expansions` as you need in one call — the API
returns tweet, user, media, and referenced-tweet objects under `includes`, which
you resolve by ID to avoid N+1 follow-up requests.
## Pagination model
The bookmarks endpoint uses **cursor-based pagination**, not offset pagination.
```json
{
"data": [ /* up to max_results tweets, newest first */ ],
"includes": { "users": [...], "media": [...], "tweets": [...] },
"meta": {
"result_count": 5,
"next_token": "abc123..."
}
}
```
The loop:
```python
next_token = None
while True:
params = {"max_results": 100, ...}
if next_token:
params["pagination_token"] = next_token
response = get_bookmarks_page(user_id, access_token, next_token)
tweets = response.get("data", [])
process(tweets)
next_token = response.get("meta", {}).get("next_token")
if not tweets or not next_token:
break
```
### Important behaviors
- **Results are newest → oldest.** The first page contains the most recently
bookmarked tweets.
- **`next_token` is opaque.** Don't parse, store long-term, or reuse across
different accounts. Each cursor is tied to the user and the point-in-time view.
- **No total count is returned.** You cannot know in advance how many bookmarks
exist; you must paginate until `next_token` disappears.
- **There is no `previous_token` for bookmarks** (unlike some other v2 endpoints),
so you can only walk forward through history.
## Rate limits
The bookmarks endpoint shares the **User Context per-user** limit (15 requests /
15 minutes as of writing). With `max_results=100` that's up to ~1500 bookmarks per
15-minute window.
Practical implications:
- Prefer `max_results=100` for backfills to minimize requests.
- Handle `429 Too Many Requests` by reading `x-rate-limit-reset` and sleeping
until that epoch, rather than retrying immediately.
- For an incremental poll where you expect few new bookmarks, a smaller page size
reduces wasted quota — but only use this once you've already done a full backfill.
## Incremental sync (don't re-fetch everything)
Because pages are newest-first, you can stop as soon as you reach a tweet you
already have stored:
```python
latest_stored_id = get_latest_stored_bookmark_id(connection_id)
while next_token is not None:
page = fetch_page(...)
for tweet in page["data"]:
if int(tweet["id"]) == latest_stored_id:
return # reached the tip we already have
upsert(tweet)
next_token = page["meta"].get("next_token")
```
Store the highest tweet ID (or most recent `bookmarked_at`) seen after each sync
and use it as the stop boundary next time. This turns a 100-page backfill into a
1–2 page poll.
## Upsert, don't insert
Tweet metrics (likes, impressions, bookmark count) change over time. Store
bookmarks with a unique constraint on `(user_id, tweet_id)` and use an
`INSERT ... ON CONFLICT DO UPDATE` so re-fetched pages refresh the numbers
instead of being ignored or duplicated.
## Resolving expansions
The API returns denormalized tweet objects in `data` and related objects in
`includes`. Resolve them once per page:
```python
users_by_id = {u["id"]: u for u in includes.get("users", [])}
media_by_key = {m["media_key"]: m for m in includes.get("media", [])}
tweets_by_id = {t["id"]: t for t in includes.get("tweets", [])}
for tweet in data:
author = users_by_id.get(tweet.get("author_id"))
media = [media_by_key[k] for k in (tweet.get("attachments", {}).get("media_keys") or []) if k in media_by_key]
refs = [tweets_by_id[r["id"]] for r in (tweet.get("referenced_tweets") or []) if r.get("id") in tweets_by_id]
```
This gives you a fully hydrated bookmark in one API call per page.
## Token management
- X OAuth 2 access tokens expire (typically 2 hours).
- Store the `refresh_token` securely and refresh before it expires or when an API
call returns `401`.
- Refresh tokens can rotate on each use — always persist the new one returned by
the token endpoint.
## Checklist
- [ ] Use OAuth 2 User Context with PKCE, not app-only bearer tokens.
- [ ] Request `bookmark.read` scope.
- [ ] Set `max_results` to `100` for backfills; lower for incremental polls.
- [ ] Loop on `meta.next_token` until it's absent.
- [ ] Handle `429` by respecting `x-rate-limit-reset`.
- [ ] Persist the latest seen tweet ID to short-circuit future syncs.
- [ ] Upsert on `(user_id, tweet_id)` to keep metrics fresh.
- [ ] Resolve `includes` by ID to avoid follow-up requests.
- [ ] Refresh access tokens proactively using the stored refresh token.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment