Skip to main content

When to use this

Use this when an agent should choose which of two approved pages to read, then return a structured report. The AI SDK owns the tool loop and Stagehand owns browser extraction. For a fixed two-page scrape, call extraction directly; the agent loop is useful when you need question-driven research.

Goal and output

Ask an AI SDK agent to compare Stagehand act() and extract() by reading their two documentation pages. One tool visits an approved URL and extracts facts in the same operation. The tool serializes concurrent requests because they share one Browserbase page. The final JSON report has an answer and facts attributed to the exact source URLs. The job rejects missing or unapproved sources.

Agent prompt

Build a source-bound AI SDK research agent with Stagehand.

Prerequisites and inputs

Use Node.js 22.18+ and pnpm. Add BROWSERBASE_API_KEY and OPENAI_API_KEY to .env. The AI SDK agent and Stagehand use OpenAI; see model configuration to adapt Stagehand to another provider. The approved URLs are /v4/basics/act and /v4/basics/extract on docs.stagehand.dev.

Check out and run

The source is packages/examples/cookbooks/ai-sdk-research-agent. Clone it with the overview command:

How the job works

1

Open one browser

Launch a Browserbase session and create Stagehand. Keep the same page throughout the tool loop.
2

Read approved sources

readSource checks the URL, then serializes navigation and extraction. It rejects redirects and validates extracted facts.
3

Validate the report

AI SDK returns a structured report. Require both sources to have been read and cited; reject any extra citation before writing the output.
The core pattern is an array of sources and a tool that visits a page and extracts facts:
This excerpt shows the tool flow. The runnable project also serializes page reads and checks the report’s sources.

Expected result and failure checks

The CLI prints JSON with a comparison answer and source entries for both approved pages. A visit request for another URL fails before navigation. Missing facts, a missing page, an incomplete report, or an extra citation fails rather than returning a partial answer. The agent’s model may vary its wording; the source and schema checks are deterministic.

Adapt sources and limits

Edit the sourceUrls array in src/index.ts and set RESEARCH_QUESTION to your question. The tools accept only those exact URLs, and reject redirects and citations outside the allowlist. The agent instructions treat page content as untrusted data. This is a tool navigation policy, not a network firewall for page subresources. The agent has a 10-step limit, a 120-second generation timeout, a 3,000-token output limit, and no automatic model retries. Browser lifetime is five minutes. The job saves the validated report to out/report.json. Citation checks show which sources the agent used; review the answer for factual accuracy before using it downstream.

Adapt to your site

Edit sourceUrls and RESEARCH_QUESTION for your task. If you need more sources, update both the input and report schemas and the completion checks. Keep exact URL checks, tool-loop limits, and post-run citation validation. Review generated facts before using the answer downstream. AI SDK integration, extraction, and model configuration.