2026-06-08 23:06:14 +00:00
2026-06-08 22:56:32 +00:00
2026-06-08 23:06:14 +00:00
2026-06-08 22:50:48 +00:00
2026-06-08 23:05:43 +00:00

seleniumbase-website-scraper

Private A2A agent for compliant, authorized website scraping with SeleniumBase/browser automation where available.

Safety policy

Use this agent only for authorized scraping of public or user-owned sites. The agent must not bypass, solve, defeat, or evade CAPTCHAs, anti-bot protections, login gates, paywalls, robots.txt, ToS restrictions, or access controls. If a CAPTCHA, anti-bot challenge, login wall, paywall, or access-control indicator is detected, the run stops and returns a manual-intervention/blocked status.

The agent does not use stealth/undetected-browser modes, CAPTCHA solvers, proxy rotation, credential injection, or access-control circumvention.

Skills

scrape_website

Inputs:

  • url — absolute http or https URL.
  • selectors — mapping of output field names to CSS selectors.
  • max_pages — maximum same-origin pages to visit.
  • wait_seconds — wait after page open before capture/extraction.
  • output_format — json or csv.
  • headless — toggle browser headless mode.
  • user_agent — optional custom single-line user agent.
  • respect_robots — when true, checks robots.txt before fetch and pagination.
  • rate_limit_seconds — polite delay between pagination requests.
  • screenshot — save PNG screenshots for visited pages.

Outputs:

  • status — ok, invalid_input, blocked_by_robots_txt, manual_intervention_required, browser_runtime_unavailable, or error.
  • items — extracted records.
  • saved_paths — output files under outputs/seleniumbase-website-scraper/.
  • warnings — safety, robots, selector, browser, or runtime warnings.
  • screenshots — screenshot file paths when requested.
  • pages_visited and html_captures.

iterate_scrape_goal (new)

An LLM-powered planner that iterates scrapes until a goal is satisfied or a safety/stop condition occurs. Inputs:

  • goal — natural-language target (e.g., "Collect quotes text and authors from the first 2 pages").
  • url — start page (absolute http/https). Same-origin only for all iterations.
  • selectors — optional initial CSS selector map (defaults to {"title": "title"}).
  • min_items — stop when this many items have been collected.
  • max_iterations — planner budget (default 5, hard-capped at 3 improvements per runtime notes).
  • per_pass_max_pages — pages per scrape pass (default 1, capped by agent config).
  • wait_seconds, output_format, headless, user_agent, respect_robots, rate_limit_seconds, screenshot — mirror scrape_website.

Outputs:

  • status — ok | no_items | blocked_by_robots_txt | manual_intervention_required | browser_runtime_unavailable | error
  • items — accumulated unique items across passes
  • pages_visited — union of pages visited across passes
  • html_captures — captures from passes
  • screenshots — when requested
  • saved_paths — includes an iter-plan manifest under outputs/seleniumbase-website-scraper/
  • warnings — including planner parsing issues, safety stops, or budget stops
  • iterations — number of planner decisions taken
  • decisions — array of {iteration, action, reason, confidence, url?, selectors?}

Behavior and guarantees:

  • Uses platform-provided ctx.llm with strict system prompt that forbids any evasion of CAPTCHAs, anti-bot, logins, paywalls, or robots.txt.
  • Planner can only propose same-origin URLs; cross-origin plans are rejected and end the loop.
  • Each pass runs the deterministic scrape_website skill; stops on any challenge or robots block.
  • Produces a compact decisions trail and saves a small manifest under outputs/seleniumbase-website-scraper/.

Selector syntax

Default extraction returns visible text:

{"title": "h1", "summary": ".article-summary"}

Supported suffixes:

  • ::text — visible text.
  • ::html — inner HTML.
  • ::attr(name) — attribute extraction, e.g. a.card::attr(href).

Special pagination selector field names are not emitted as data fields and are used to find the next page:

  • next
  • _next
  • next
  • pagination_next
  • next_page

Pagination is restricted to the same scheme and host as the start URL.

Example payloads

Simple one-pass scrape:

{
  "url": "https://example.com/articles",
  "selectors": {
    "headline": "h2.article-title",
    "links": "a.article-link::attr(href)",
    "next": "a[rel='next']::attr(href)"
  },
  "max_pages": 3,
  "wait_seconds": 2,
  "output_format": "json",
  "headless": true,
  "respect_robots": true
}

Iterative planner:

{
  "goal": "collect at least 30 product tiles with name, price, link",
  "url": "https://example.com/store",
  "selectors": {"name": ".tile .name", "price": ".tile .price", "href": ".tile a::attr(href)"},
  "min_items": 30,
  "max_iterations": 3,
  "per_pass_max_pages": 1,
  "respect_robots": true
}

Smoke test: quotes.toscrape.com

Goal: "Collect quotes text and authors from the first 2 pages"

  • Expected: status=ok, >= 1 item, <= 2 iterations (often 12), a saved JSON manifest, no screenshots when disabled.
  • Example call (schematic):
{
  "goal": "Collect quotes text and authors from the first 2 pages",
  "url": "https://quotes.toscrape.com/",
  "selectors": {"text": ".quote .text", "author": ".quote .author", "pagination_next": "li.next a::attr(href)"},
  "min_items": 2,
  "max_iterations": 2,
  "per_pass_max_pages": 2,
  "screenshot": false
}

Deployment notes

The agent declares a larger runtime budget because browser work can be CPU/memory intensive. Actual browser execution is performed through the workspace sandbox helper using the SeleniumBase image, so files written under /workspace/outputs/seleniumbase-website-scraper/ are durable workspace outputs.

Description
agent seleniumbase-website-scraper
Readme 993 KiB
Languages
Python 100%