Install
openclaw skills install @jorg18/agentcrawlRead web pages and documentation as clean Markdown with AgentCrawl, locally. Use when you need the content of a URL, several URLs, or a whole docs site; when your built-in web fetch fails, returns raw HTML, or the page needs JavaScript; when a page is too long and you only need one section or the parts about a question; or when you want structured data from a page with CSS selectors and no LLM.
openclaw skills install @jorg18/agentcrawlAgentCrawl turns web pages into Markdown, links and metadata for agents. It runs on the user's machine: no account, no API key, no hosted scraper. It follows robots.txt, refuses private-network targets unless told otherwise, and never returns a bot-check page as if it were content.
This is the official skill, by the AgentCrawl author.
Check first: agentcrawl --version. If it is missing, install it (Python 3.10+):
pipx install "agentcrawl-ai[mcp]" # or: python3 -m pip install "agentcrawl-ai[mcp]"
agentcrawl doctor # says what works and what is missing
Plain HTTP pages need nothing else. For pages that only render with JavaScript, add the browser (about 150 MB, only when needed):
pipx inject agentcrawl-ai playwright && python3 -m playwright install chromium
# with pip: python3 -m pip install "agentcrawl-ai[browser]" && python3 -m playwright install chromium
Register the stdio server in your own client config, keeping every other server as it is:
{ "mcpServers": { "agentcrawl": { "command": "agentcrawl", "args": ["mcp"] } } }
Then use the tools: scrape_url (one URL), scrape_many (several), map_site (list a site's URLs), crawl_site (bounded multi-page read), extract_structured (CSS schema), and search_web when a search engine is configured. The full client-by-client procedure is in INSTALL_FOR_AGENTS.md in the repository.
Without MCP, use the CLI below. Every command prints JSON on stdout.
agentcrawl scrape https://docs.python.org/3/library/json.html
agentcrawl scrape-many https://example.com/a https://example.com/b
agentcrawl scrape-many --file urls.txt
The page text is in the markdown field: pipe through jq -r .markdown to read only that. Add --format links or --format metadata (repeatable) for more fields; --full-page keeps navigation and footers, which are dropped by default.
Do not load a 30,000-token page to answer one question.
agentcrawl scrape https://example.com/guide --format outline # sections with their size in tokens
agentcrawl scrape https://example.com/guide --section s4 # one section, by id or heading text
agentcrawl scrape https://example.com/guide --query "rate limits" --max-tokens 1500
With MCP: scrape_url(url, formats=["outline"]), then scrape_url(url, section="s4"). The page is downloaded once and reused for 10 minutes.
agentcrawl map https://docs.example.com --max-urls 200 # URLs only, from sitemaps and links
agentcrawl crawl https://docs.example.com --max-pages 30 --max-depth 2
agentcrawl crawl https://docs.example.com --query "authentication" --max-pages 30
Always set --max-pages. With --query the crawl reads the most relevant links first and stops once pages stop matching the question. Prefer map and then scrape-many on the URLs you picked over a blind crawl.
agentcrawl extract-css https://shop.example.com/catalog --schema products.json
products.json:
{
"baseSelector": "div.product",
"fields": [
{"name": "title", "selector": "h2", "type": "text"},
{"name": "url", "selector": "a", "type": "attribute", "attribute": "href", "transform": "url"},
{"name": "price", "selector": ".price", "type": "text", "transform": "number"}
]
}
Field types: text, attribute, html, regex, nested, list. Selectors outside the supported subset are rejected instead of silently matching nothing.
agentcrawl search "python json streaming parser" --limit 5
Off by default: it needs a search engine the user configured (AGENTCRAWL_SEARCH_ENGINE=serper plus SERPER_API_KEY). Without one the command says so; then ask the user for a URL or use your own search tool, and read the results with agentcrawl scrape.
A failed page has errors and a metadata.next_step that says what to do (for example: the site blocks automated clients, try the browser, wait for a rate limit, or give up on this URL). Follow next_step instead of retrying the same call. The CLI exits 1 when the result has errors. Use another tool only after AgentCrawl returns a definitive error.
localhost and private addresses are refused on purpose. Allow them only when the user asked for a local dev server: agentcrawl scrape http://127.0.0.1:3000 --allow-private-network.
--max-pages small, do not disable robots.txt (--no-robots) unless the user owns the site or asked for it.