AnyCrawl Skill
AnyCrawl API integration for OpenClaw - Scrape, Crawl, and Search web content with high-performance multi-threaded crawling.
Setup
Method 1: Environment variable (Recommended)
export ANYCRAWL_API_KEY="your-api-key"
Make it permanent by adding to ~/.bashrc or ~/.zshrc:
echo 'export ANYCRAWL_API_KEY="your-api-key"' >> ~/.bashrc
source ~/.bashrc
Get your API key at: https://anycrawl.dev
Method 2: OpenClaw gateway config
openclaw config.patch --set ANYCRAWL_API_KEY="your-api-key"
Functions
1. anycrawl_scrape
Scrape a single URL and convert to LLM-ready structured data.
Parameters:
url (string, required): URL to scrape
engine (string, optional): Scraping engine - "cheerio" (default), "playwright", "puppeteer"
formats (array, optional): Output formats - ["markdown"], ["html"], ["text"], ["json"], ["screenshot"]
timeout (number, optional): Timeout in milliseconds (default: 30000)
wait_for (number, optional): Delay before extraction in ms (browser engines only)
wait_for_selector (string/object/array, optional): Wait for CSS selectors
include_tags (array, optional): Include only these HTML tags (e.g., ["h1", "p", "article"])
exclude_tags (array, optional): Exclude these HTML tags
proxy (string, optional): Proxy URL (e.g., "http://proxy:port")
json_options (object, optional): JSON extraction with schema/prompt
extract_source (string, optional): "markdown" (default) or "html"
Examples:
// Basic scrape with default cheerio
anycrawl_scrape({ url: "https://example.com" })
// Scrape SPA with Playwright
anycrawl_scrape({
url: "https://spa-example.com",
engine: "playwright",
formats: ["markdown", "screenshot"]
})
// Extract structured JSON
anycrawl_scrape({
url: "https://product-page.com",
engine: "cheerio",
json_options: {
schema: {
type: "object",
properties: {
product_name: { type: "string" },
price: { type: "number" },
description: { type: "string" }
},
required: ["product_name", "price"]
},
user_prompt: "Extract product details from this page"
}
})
2. anycrawl_search
Search Google and return structured results.
Parameters:
query (string, required): Search query
engine (string, optional): Search engine - "google" (default)
limit (number, optional): Max results per page (default: 10)
offset (number, optional): Number of results to skip (default: 0)
pages (number, optional): Number of pages to retrieve (default: 1, max: 20)
lang (string, optional): Language locale (e.g., "en", "zh", "vi")
safe_search (number, optional): 0 (off), 1 (medium), 2 (high)
scrape_options (object, optional): Scrape each result URL with these options
Examples:
// Basic search
anycrawl_search({ query: "OpenAI ChatGPT" })
// Multi-page search in Vietnamese
anycrawl_search({
query: "hướng dẫn Node.js",
pages: 3,
lang: "vi"
})
// Search and auto-scrape results
anycrawl_search({
query: "best AI tools 2026",
limit: 5,
scrape_options: {
engine: "cheerio",
formats: ["markdown"]
}
})
3. anycrawl_crawl_start
Start crawling an entire website (async job).
Parameters:
url (string, required): Seed URL to start crawling
engine (string, optional): "cheerio" (default), "playwright", "puppeteer"
strategy (string, optional): "all", "same-domain" (default), "same-hostname", "same-origin"
max_depth (number, optional): Max depth from seed URL (default: 10)
limit (number, optional): Max pages to crawl (default: 100)
include_paths (array, optional): Path patterns to include (e.g., ["/blog/*"])
exclude_paths (array, optional): Path patterns to exclude (e.g., ["/admin/*"])
scrape_paths (array, optional): Only scrape URLs matching these patterns
scrape_options (object, optional): Per-page scrape options
Examples:
// Crawl entire website
anycrawl_crawl_start({
url: "https://docs.example.com",
engine: "cheerio",
max_depth: 5,
limit: 50
})
// Crawl only blog posts
anycrawl_crawl_start({
url: "https://example.com",
strategy: "same-domain",
include_paths: ["/blog/*"],
exclude_paths: ["/blog/tags/*"],
scrape_options: {
formats: ["markdown"]
}
})
// Crawl product pages only
anycrawl_crawl_start({
url: "https://shop.example.com",
strategy: "same-domain",
scrape_paths: ["/products/*"],
limit: 200
})
4. anycrawl_crawl_status
Check crawl job status.
Parameters:
job_id (string, required): Crawl job ID
Example:
anycrawl_crawl_status({ job_id: "7a2e165d-8f81-4be6-9ef7-23222330a396" })
5. anycrawl_crawl_results
Get crawl results (paginated).
Parameters:
job_id (string, required): Crawl job ID
skip (number, optional): Number of results to skip (default: 0)
Example:
// Get first 100 results
anycrawl_crawl_results({ job_id: "xxx", skip: 0 })
// Get next 100 results
anycrawl_crawl_results({ job_id: "xxx", skip: 100 })
6. anycrawl_crawl_cancel
Cancel a running crawl job.
Parameters:
job_id (string, required): Crawl job ID
7. anycrawl_search_and_scrape
Quick helper: Search Google then scrape top results.
Parameters:
query (string, required): Search query
max_results (number, optional): Max results to scrape (default: 3)
scrape_engine (string, optional): Engine for scraping (default: "cheerio")
formats (array, optional): Output formats (default: ["markdown"])
lang (string, optional): Search language
Example:
anycrawl_search_and_scrape({
query: "latest AI news",
max_results: 5,
formats: ["markdown"]
})
Engine Selection Guide
| Engine |
Best For |
Speed |
JS Rendering |
cheerio |
Static HTML, news, blogs |
⚡ Fastest |
❌ No |
playwright |
SPAs, complex web apps |
🐢 Slower |
✅ Yes |
puppeteer |
Chrome-specific, metrics |
🐢 Slower |
✅ Yes |
Response Format
All responses follow this structure:
{
"success": true,
"data": { ... },
"message": "Optional message"
}
Error response:
{
"success": false,
"error": "Error type",
"message": "Human-readable message"
}
Common Error Codes
400 - Bad Request (validation errors)
401 - Unauthorized (invalid API key)
402 - Payment Required (insufficient credits)
404 - Not Found
429 - Rate limit exceeded
500 - Internal server error
API Limits
- Rate limits apply based on your plan
- Crawl jobs expire after 24 hours
- Max crawl limit: depends on credits
Links
1---2name: anycrawl-skill3description: AnyCrawl API integration for OpenClaw - Scrape, Crawl, and Search web content with high-performance multi-threaded crawling.4---5
6# AnyCrawl Skill
7
8AnyCrawl API integration for OpenClaw - Scrape, Crawl, and Search web content with high-performance multi-threaded crawling.
9
10## Setup
11
12### Method 1: Environment variable (Recommended)
13
14```bash
15export ANYCRAWL_API_KEY="your-api-key"
16```
17
18Make it permanent by adding to `~/.bashrc` or `~/.zshrc`:
19```bash
20echo 'export ANYCRAWL_API_KEY="your-api-key"' >> ~/.bashrc
21source ~/.bashrc
22```
23
24Get your API key at: https://anycrawl.dev
25
26### Method 2: OpenClaw gateway config
27
28```bash
29openclaw config.patch --set ANYCRAWL_API_KEY="your-api-key"
30```
31
32## Functions
33
34### 1. anycrawl_scrape
35
36Scrape a single URL and convert to LLM-ready structured data.
37
38**Parameters:**
39- `url` (string, required): URL to scrape
40- `engine` (string, optional): Scraping engine - `"cheerio"` (default), `"playwright"`, `"puppeteer"`
41- `formats` (array, optional): Output formats - `["markdown"]`, `["html"]`, `["text"]`, `["json"]`, `["screenshot"]`
42- `timeout` (number, optional): Timeout in milliseconds (default: 30000)
43- `wait_for` (number, optional): Delay before extraction in ms (browser engines only)
44- `wait_for_selector` (string/object/array, optional): Wait for CSS selectors
45- `include_tags` (array, optional): Include only these HTML tags (e.g., `["h1", "p", "article"]`)
46- `exclude_tags` (array, optional): Exclude these HTML tags
47- `proxy` (string, optional): Proxy URL (e.g., `"http://proxy:port"`)
48- `json_options` (object, optional): JSON extraction with schema/prompt
49- `extract_source` (string, optional): `"markdown"` (default) or `"html"`
50
51**Examples:**
52
53```javascript
54// Basic scrape with default cheerio
55anycrawl_scrape({ url: "https://example.com" })
56
57// Scrape SPA with Playwright
58anycrawl_scrape({
59 url: "https://spa-example.com",
60 engine: "playwright",
61 formats: ["markdown", "screenshot"]
62})
63
64// Extract structured JSON
65anycrawl_scrape({
66 url: "https://product-page.com",
67 engine: "cheerio",
68 json_options: {
69 schema: {
70 type: "object",
71 properties: {
72 product_name: { type: "string" },
73 price: { type: "number" },
74 description: { type: "string" }
75 },
76 required: ["product_name", "price"]
77 },
78 user_prompt: "Extract product details from this page"
79 }
80})
81```
82
83### 2. anycrawl_search
84
85Search Google and return structured results.
86
87**Parameters:**
88- `query` (string, required): Search query
89- `engine` (string, optional): Search engine - `"google"` (default)
90- `limit` (number, optional): Max results per page (default: 10)
91- `offset` (number, optional): Number of results to skip (default: 0)
92- `pages` (number, optional): Number of pages to retrieve (default: 1, max: 20)
93- `lang` (string, optional): Language locale (e.g., `"en"`, `"zh"`, `"vi"`)
94- `safe_search` (number, optional): 0 (off), 1 (medium), 2 (high)
95- `scrape_options` (object, optional): Scrape each result URL with these options
96
97**Examples:**
98
99```javascript
100// Basic search
101anycrawl_search({ query: "OpenAI ChatGPT" })
102
103// Multi-page search in Vietnamese
104anycrawl_search({
105 query: "hướng dẫn Node.js",
106 pages: 3,
107 lang: "vi"
108})
109
110// Search and auto-scrape results
111anycrawl_search({
112 query: "best AI tools 2026",
113 limit: 5,
114 scrape_options: {
115 engine: "cheerio",
116 formats: ["markdown"]
117 }
118})
119```
120
121### 3. anycrawl_crawl_start
122
123Start crawling an entire website (async job).
124
125**Parameters:**
126- `url` (string, required): Seed URL to start crawling
127- `engine` (string, optional): `"cheerio"` (default), `"playwright"`, `"puppeteer"`
128- `strategy` (string, optional): `"all"`, `"same-domain"` (default), `"same-hostname"`, `"same-origin"`
129- `max_depth` (number, optional): Max depth from seed URL (default: 10)
130- `limit` (number, optional): Max pages to crawl (default: 100)
131- `include_paths` (array, optional): Path patterns to include (e.g., `["/blog/*"]`)
132- `exclude_paths` (array, optional): Path patterns to exclude (e.g., `["/admin/*"]`)
133- `scrape_paths` (array, optional): Only scrape URLs matching these patterns
134- `scrape_options` (object, optional): Per-page scrape options
135
136**Examples:**
137
138```javascript
139// Crawl entire website
140anycrawl_crawl_start({
141 url: "https://docs.example.com",
142 engine: "cheerio",
143 max_depth: 5,
144 limit: 50
145})
146
147// Crawl only blog posts
148anycrawl_crawl_start({
149 url: "https://example.com",
150 strategy: "same-domain",
151 include_paths: ["/blog/*"],
152 exclude_paths: ["/blog/tags/*"],
153 scrape_options: {
154 formats: ["markdown"]
155 }
156})
157
158// Crawl product pages only
159anycrawl_crawl_start({
160 url: "https://shop.example.com",
161 strategy: "same-domain",
162 scrape_paths: ["/products/*"],
163 limit: 200
164})
165```
166
167### 4. anycrawl_crawl_status
168
169Check crawl job status.
170
171**Parameters:**
172- `job_id` (string, required): Crawl job ID
173
174**Example:**
175```javascript
176anycrawl_crawl_status({ job_id: "7a2e165d-8f81-4be6-9ef7-23222330a396" })
177```
178
179### 5. anycrawl_crawl_results
180
181Get crawl results (paginated).
182
183**Parameters:**
184- `job_id` (string, required): Crawl job ID
185- `skip` (number, optional): Number of results to skip (default: 0)
186
187**Example:**
188```javascript
189// Get first 100 results
190anycrawl_crawl_results({ job_id: "xxx", skip: 0 })
191
192// Get next 100 results
193anycrawl_crawl_results({ job_id: "xxx", skip: 100 })
194```
195
196### 6. anycrawl_crawl_cancel
197
198Cancel a running crawl job.
199
200**Parameters:**
201- `job_id` (string, required): Crawl job ID
202
203### 7. anycrawl_search_and_scrape
204
205Quick helper: Search Google then scrape top results.
206
207**Parameters:**
208- `query` (string, required): Search query
209- `max_results` (number, optional): Max results to scrape (default: 3)
210- `scrape_engine` (string, optional): Engine for scraping (default: `"cheerio"`)
211- `formats` (array, optional): Output formats (default: `["markdown"]`)
212- `lang` (string, optional): Search language
213
214**Example:**
215```javascript
216anycrawl_search_and_scrape({
217 query: "latest AI news",
218 max_results: 5,
219 formats: ["markdown"]
220})
221```
222
223## Engine Selection Guide
224
225| Engine | Best For | Speed | JS Rendering |
226|--------|----------|-------|--------------|
227| `cheerio` | Static HTML, news, blogs | ⚡ Fastest | ❌ No |
228| `playwright` | SPAs, complex web apps | 🐢 Slower | ✅ Yes |
229| `puppeteer` | Chrome-specific, metrics | 🐢 Slower | ✅ Yes |
230
231## Response Format
232
233All responses follow this structure:
234
235```json
236{
237 "success": true,
238 "data": { ... },
239 "message": "Optional message"
240}
241```
242
243Error response:
244```json
245{
246 "success": false,
247 "error": "Error type",
248 "message": "Human-readable message"
249}
250```
251
252## Common Error Codes
253
254- `400` - Bad Request (validation errors)
255- `401` - Unauthorized (invalid API key)
256- `402` - Payment Required (insufficient credits)
257- `404` - Not Found
258- `429` - Rate limit exceeded
259- `500` - Internal server error
260
261## API Limits
262
263- Rate limits apply based on your plan
264- Crawl jobs expire after 24 hours
265- Max crawl limit: depends on credits
266
267## Links
268
269- API Docs: https://docs.anycrawl.dev
270- Website: https://anycrawl.dev
271- Playground: https://anycrawl.dev/playground