Skip to main content

When to use this

Use this when a catalog spans several pages and failed runs must preserve progress. Stagehand extracts typed records and resolves the Next action; a stable pagination marker proves when to stop. If your application already has reliable product locators or an API, use those for the fixed parts of the job.

Goal and output

Export the Books to Scrape catalog into out/catalog.json. Each entry has a nonempty title, displayed price, and availability. The file also reports pages and count. The job saves validated pages atomically to out/checkpoint.json. The final catalog appears only after pagination ends.

Agent prompt

Export a complete paginated catalog with Stagehand.

Prerequisites and inputs

  • Add BROWSERBASE_API_KEY and OPENAI_API_KEY to .env. The TypeScript runner can also use a local browser by setting BROWSER_ENV=LOCAL.
  • CATALOG_URL defaults to the two-page mystery category on Books to Scrape. MAX_PAGES defaults to 2 and accepts 1 through 100. The default completes a small category. For the full catalog, set CATALOG_URL=https://books.toscrape.com/ and raise MAX_PAGES to 50. NEXT_SELECTOR defaults to li.next a; change it when adapting the target. OUT_DIR defaults to out.
  • TypeScript needs Node.js 22.18+ and pnpm; Python needs Python 3.11+ and uv; Go needs Go 1.26+.

Check out and run

The source is packages/examples/cookbooks/paginated-catalog. Clone this folder with the overview command, then run one version:

How the job works

1

Open the catalog

Launch one cloud browser and navigate to CATALOG_URL. Keep the same page for all extraction and navigation calls.
2

Extract typed records

Check the current URL against a visited set. Extract all visible books into a typed schema with nonempty fields.
3

Resolve pagination

Check NEXT_SELECTOR for a Next link. Its absence ends pagination. If it exists, ask observe() for the action, pass that action to act(), and require navigation. An empty observation while Next exists is an error.
4

Checkpoint progress

Save validated pages to the checkpoint before navigation. On rerun, revisit the last saved URL to find Next without extracting it again. Fail before extracting beyond the total MAX_PAGES budget.
5

Publish the complete catalog

Write the final file atomically after natural completion. Deduplicate identical title, price, and availability tuples. Close Stagehand and the browser.
These excerpts show extraction and the next-page action. The runnable project also extracts records, saves checkpoints, and checks origin, cycles, page limits, and navigation results.

Recover an interrupted export

Run again with the same CATALOG_URL and OUT_DIR. Increase MAX_PAGES when the previous run stopped at its budget. The checkpoint binds progress to the source URL and rejects corrupt records or navigation to another origin. Completed reruns rebuild the final file from the checkpoint without extracting pages again.
Use one writer per output directory. Concurrent runs can overwrite each other’s checkpoint.
Use a new OUT_DIR for a fresh snapshot or changed source. Checkpoints preserve previously extracted records; they do not refresh changing prices. The example deduplicates identical records. For a real catalog, add a product ID to the schema and use it as the deduplication key. The same checkpoint format works across TypeScript, Python, and Go.

Expected result and failure checks

A completed run prints Saved N books from P pages to out/catalog.json. The file contains pages, count, and books; count equals the array length. Set MAX_PAGES=1 against a multipage catalog to confirm the limit fails with a checkpoint and without producing a new final catalog. A one-page target stops when the configured Next marker is absent. A cycle or unchanged URL fails instead of looping. Run credential-free recovery tests with pnpm test, uv run --locked python -m unittest, or go test ./... from the matching language folder.

Adapt to your site

Change CATALOG_URL, NEXT_SELECTOR, and the extraction schema together. Use a stable product ID as the deduplication key when different products can share displayed fields. Keep checkpoint/source validation, the page budget, and atomic final writes. Use a fresh OUT_DIR for each new snapshot. Extraction, observation, and deployment. For a recorded data-collection job, see the showcase.