Skip to main content

When to use this

Use this when extracted page data and downloaded files must enter a storage pipeline. Stagehand supplies typed data from the page; the server process downloads and uploads bytes. Prefer a download API when the site provides one. S3 is optional, so the first run can produce local artifacts.

Goal and output

Read the first page of Books to Scrape in a Browserbase cloud browser. Stagehand extracts book data, and the job downloads the first cover as bytes. It always writes out/catalog.json and out/cover.jpg; when S3_BUCKET is set, Files SDK uploads both objects to that bucket. Bucket credentials stay in the server process, outside the browser.

Agent prompt

Connect Stagehand extraction to Files SDK storage.

Prerequisites and inputs

Use Node.js 22.18+ and pnpm. Add BROWSERBASE_API_KEY and OPENAI_API_KEY to .env. For upload, set S3_BUCKET and credentials supported by the AWS credential chain. AWS_REGION defaults to us-east-1, S3_PREFIX defaults to stagehand-cookbooks/files-to-bucket, and S3_ENDPOINT can target an S3-compatible service. Without S3_BUCKET, the job stops after local output.

Check out and run

The source is packages/examples/cookbooks/files-to-bucket. Clone it with the overview command:

How the job works

1

Extract the catalog

Launch Browserbase and create Stagehand. Open the catalog page and extract() titles and prices into a nonempty schema.
2

Download the cover

Read the first cover URL directly from the product-grid image DOM, then download it with host, time, byte, and JPEG checks.
3

Write and upload artifacts

Write JSON and cover bytes under out/. When S3_BUCKET exists, create a Files SDK s3() adapter and upload the same bytes.
4

Close browser resources

Close Stagehand and the browser in finally blocks. Storage credentials are read only by the server-side adapter.
The extraction runs in the browser. Files SDK receives the validated JSON in the server process:
The runnable folder also reads the cover URL from the DOM, downloads the image, and writes both files locally before optional upload. The S3 adapter’s upload shape follows Files SDK.

Expected result and failure checks

out/catalog.json contains at least one book, and out/cover.jpg has nonzero bytes. Without S3_BUCKET, the CLI reports that upload was skipped. With it, the two objects appear under S3_PREFIX. An invalid page result, failed download, or upload error exits nonzero and still closes the browser.

Download limits

The download runs in the Node process. The example accepts only HTTPS images on books.toscrape.com, rejects redirects, checks JPEG bytes, and stops after 15 seconds or 5 MiB. When adapting the target, change the download allowlist explicitly. Do not fetch arbitrary model-generated URLs from a server with private network access. Browser lifetime is five minutes. S3 credentials stay in the SDK process. Uploads overwrite the configured object keys; use a distinct S3_PREFIX when you need separate snapshots. Local artifacts remain available if an upload fails. Run pnpm test to check the download failure paths without credentials.

Adapt to your site

Change the target URL, schema, download host allowlist, and file-format checks together. Use a distinct S3_PREFIX for each snapshot when you need history. Keep byte and time limits, redirect rejection, and server-side credentials. A valid extension or content-type header alone does not prove a file format. Paginated catalogs, extraction, and Files SDK.