Created
July 12, 2026 09:31
-
-
Save erdaltoprak/02926a971a2d8c5aa9037e6bf13138f7 to your computer and use it in GitHub Desktop.
Fetching All Bookmarks from the X (Twitter) API
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| # Fetching All Bookmarks from the X (Twitter) API | |
| A practical guide to pulling a user's complete bookmark history via the X API v2, | |
| including pagination, incremental sync, and a few gotchas that aren't obvious from | |
| the docs. | |
| ## Prerequisites | |
| - A project with **OAuth 2.0 User Context** access. | |
| - An app with a client ID / client secret and a redirect URI configured. | |
| - The `bookmark.read` and `users.read` (and `tweet.read`) scopes. | |
| - A user access token obtained via the **Authorization Code with PKCE** flow. | |
| > Bookmarks are **only** available through User Context tokens. App-only bearer | |
| > tokens cannot read bookmarks. | |
| ## The endpoint | |
| ``` | |
| GET https://api.x.com/2/users/{user_id}/bookmarks | |
| ``` | |
| ### Key query parameters | |
| | param | purpose | | |
| |--------------------|------------------------------------------------------| | |
| | `max_results` | Page size. Allowed range is `5`–`100`. | | |
| | `pagination_token` | Opaque cursor from the previous response's `meta`. | | |
| | `tweet.fields` | Request full tweet metadata (e.g. `created_at`, `public_metrics`, `note_tweet`). | | |
| | `expansions` | Inline related objects: `author_id`, `attachments.media_keys`, `referenced_tweets.id`. | | |
| | `user.fields` | Extra fields on expanded authors. | | |
| | `media.fields` | Extra fields on expanded media. | | |
| Request as many `tweet.fields` and `expansions` as you need in one call — the API | |
| returns tweet, user, media, and referenced-tweet objects under `includes`, which | |
| you resolve by ID to avoid N+1 follow-up requests. | |
| ## Pagination model | |
| The bookmarks endpoint uses **cursor-based pagination**, not offset pagination. | |
| ```json | |
| { | |
| "data": [ /* up to max_results tweets, newest first */ ], | |
| "includes": { "users": [...], "media": [...], "tweets": [...] }, | |
| "meta": { | |
| "result_count": 5, | |
| "next_token": "abc123..." | |
| } | |
| } | |
| ``` | |
| The loop: | |
| ```python | |
| next_token = None | |
| while True: | |
| params = {"max_results": 100, ...} | |
| if next_token: | |
| params["pagination_token"] = next_token | |
| response = get_bookmarks_page(user_id, access_token, next_token) | |
| tweets = response.get("data", []) | |
| process(tweets) | |
| next_token = response.get("meta", {}).get("next_token") | |
| if not tweets or not next_token: | |
| break | |
| ``` | |
| ### Important behaviors | |
| - **Results are newest → oldest.** The first page contains the most recently | |
| bookmarked tweets. | |
| - **`next_token` is opaque.** Don't parse, store long-term, or reuse across | |
| different accounts. Each cursor is tied to the user and the point-in-time view. | |
| - **No total count is returned.** You cannot know in advance how many bookmarks | |
| exist; you must paginate until `next_token` disappears. | |
| - **There is no `previous_token` for bookmarks** (unlike some other v2 endpoints), | |
| so you can only walk forward through history. | |
| ## Rate limits | |
| The bookmarks endpoint shares the **User Context per-user** limit (15 requests / | |
| 15 minutes as of writing). With `max_results=100` that's up to ~1500 bookmarks per | |
| 15-minute window. | |
| Practical implications: | |
| - Prefer `max_results=100` for backfills to minimize requests. | |
| - Handle `429 Too Many Requests` by reading `x-rate-limit-reset` and sleeping | |
| until that epoch, rather than retrying immediately. | |
| - For an incremental poll where you expect few new bookmarks, a smaller page size | |
| reduces wasted quota — but only use this once you've already done a full backfill. | |
| ## Incremental sync (don't re-fetch everything) | |
| Because pages are newest-first, you can stop as soon as you reach a tweet you | |
| already have stored: | |
| ```python | |
| latest_stored_id = get_latest_stored_bookmark_id(connection_id) | |
| while next_token is not None: | |
| page = fetch_page(...) | |
| for tweet in page["data"]: | |
| if int(tweet["id"]) == latest_stored_id: | |
| return # reached the tip we already have | |
| upsert(tweet) | |
| next_token = page["meta"].get("next_token") | |
| ``` | |
| Store the highest tweet ID (or most recent `bookmarked_at`) seen after each sync | |
| and use it as the stop boundary next time. This turns a 100-page backfill into a | |
| 1–2 page poll. | |
| ## Upsert, don't insert | |
| Tweet metrics (likes, impressions, bookmark count) change over time. Store | |
| bookmarks with a unique constraint on `(user_id, tweet_id)` and use an | |
| `INSERT ... ON CONFLICT DO UPDATE` so re-fetched pages refresh the numbers | |
| instead of being ignored or duplicated. | |
| ## Resolving expansions | |
| The API returns denormalized tweet objects in `data` and related objects in | |
| `includes`. Resolve them once per page: | |
| ```python | |
| users_by_id = {u["id"]: u for u in includes.get("users", [])} | |
| media_by_key = {m["media_key"]: m for m in includes.get("media", [])} | |
| tweets_by_id = {t["id"]: t for t in includes.get("tweets", [])} | |
| for tweet in data: | |
| author = users_by_id.get(tweet.get("author_id")) | |
| media = [media_by_key[k] for k in (tweet.get("attachments", {}).get("media_keys") or []) if k in media_by_key] | |
| refs = [tweets_by_id[r["id"]] for r in (tweet.get("referenced_tweets") or []) if r.get("id") in tweets_by_id] | |
| ``` | |
| This gives you a fully hydrated bookmark in one API call per page. | |
| ## Token management | |
| - X OAuth 2 access tokens expire (typically 2 hours). | |
| - Store the `refresh_token` securely and refresh before it expires or when an API | |
| call returns `401`. | |
| - Refresh tokens can rotate on each use — always persist the new one returned by | |
| the token endpoint. | |
| ## Checklist | |
| - [ ] Use OAuth 2 User Context with PKCE, not app-only bearer tokens. | |
| - [ ] Request `bookmark.read` scope. | |
| - [ ] Set `max_results` to `100` for backfills; lower for incremental polls. | |
| - [ ] Loop on `meta.next_token` until it's absent. | |
| - [ ] Handle `429` by respecting `x-rate-limit-reset`. | |
| - [ ] Persist the latest seen tweet ID to short-circuit future syncs. | |
| - [ ] Upsert on `(user_id, tweet_id)` to keep metrics fresh. | |
| - [ ] Resolve `includes` by ID to avoid follow-up requests. | |
| - [ ] Refresh access tokens proactively using the stored refresh token. |
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment