AnyCrawl MCP server
[AnyCrawl](https://anycrawl.dev) MCP Server, Powerful web scraping and crawling for Cursor, Claude, and other LLM clients via the Model Context Protocol (MCP).
How to read this: tool names here are observed from a live tools/list handshake. The Risk label is a heuristic inferred from the tool name (write/destructive verbs), not from executing the tool — a conservative guess, not a verified capability. We never escalate risk from a description. Found one that's wrong? Tell us — we fix on report.
| Tool | Risk | Side effects | Approval |
|---|---|---|---|
| anycrawl_crawl_results Get the results of a completed crawl job.
Retrieve the scraped content and metadata from a completed crawl job.
Usage (parameters):
- job_id: The crawl job ID (string, required)
- skip: Number of results to skip (number, optional, default: 0)
Returns: { status, total, completed, creditsUsed, next?, data[] }
Examples:
- Get all results: { "job_id": "crawl_12345" }
- Paginated results: { "job_id": "crawl_12345", "skip": 50 } | read | false | unknown |
| anycrawl_crawl_status Get the status of a crawl job.
Check progress, completion status, and statistics for an ongoing or completed crawl job.
Usage (parameters):
- job_id: The crawl job ID (string, required)
Returns: { job_id, status, start_time, expires_at, credits_used, total, completed, failed }
Examples:
- Check status: { "job_id": "crawl_12345" } | read | false | unknown |
| anycrawl_cancel_crawl Cancel a running crawl job.
Stop an ongoing crawl job and prevent further processing.
Usage (parameters):
- job_id: The crawl job ID to cancel (string, required)
Returns: { success: boolean, message: string }
Examples:
- Cancel job: { "job_id": "crawl_12345" } | destructive | true | true |
| anycrawl_scrape Scrape a single URL and extract content in selected formats.
Best for: One known page (articles, docs, product pages).
Not recommended for: Multi-page coverage (use anycrawl_crawl) or open-ended discovery (use anycrawl_search).
RECOMMENDED: Use 'playwright' engine for best results with dynamic content and modern websites.
Usage (parameters):
- url: HTTP/HTTPS URL to scrape (string, required)
- engine: 'playwright' | 'cheerio' | 'puppeteer' (required, default: 'playwright')
- proxy: Proxy URL (string, optional)
- formats: Output formats ['markdown'|'html'|'text'|'screenshot'|'screenshot@fullPage'|'rawHtml'|'json'] (optional)
- timeout: Request timeout in ms (number, optional)
- retry: Enable auto-retry on failure (boolean, optional)
- wait_for: Wait in ms for dynamic pages (number, optional)
- include_tags: HTML tags to include (string[], optional)
- exclude_tags: HTML tags to exclude (string[], optional)
- json_options: { schema?, user_prompt?, schema_name?, schema_description? } (optional)
- extract_source: 'html' | 'markdown' (optional)
Returns: { url, status, jobId?, title?, html?, markdown?, metadata?, timestamp? }
Examples:
- Recommended: { "url": "https://example.com", "engine": "playwright" }
- With JSON extraction: { "url": "https://news.ycombinator.com", "engine": "playwright", "formats": ["markdown"], "json_options": { "user_prompt": "Extract titles", "schema_name": "Articles" } }
- With JSON schema extraction: { "url": "https://example.com/article", "engine": "playwright", "json_options": { "schema_name": "Article", "schema_description": "Extract article metadata and content", "schema": { "type": "object", "properties": { "title": { "type": "string" }, "author": { "type": "string" }, "date": { "type": "string" }, "content": { "type": "string" } }, "required": ["title", "content"] } } } | read | false | unknown |
| anycrawl_search Search the web and optionally scrape results.
Best for: Open-ended discovery, finding relevant content.
Not recommended for: Known URLs (use anycrawl_scrape) or comprehensive site coverage (use anycrawl_crawl).
RECOMMENDED: Use limit=5 for balanced performance and cost. Use 'playwright' engine for scraping results.
Usage (parameters):
- query: Search query string (string, required)
- engine: Search engine 'google' (optional, default: 'google')
- limit: Number of results to return (number, optional, default: 5)
- offset: Number of results to skip (number, optional, default: 0)
- pages: Number of search result pages to process (number, optional)
- lang: Language code (string, optional)
- country: Country code (string, optional)
- safeSearch: Safe search level 0-2 (number, optional)
- scrape_options: Options for scraping search results (object, optional)
Returns: Array of search results with optional scraped content
Examples:
- Recommended: { "query": "artificial intelligence news", "limit": 5 }
- With scraping: { "query": "TypeScript tutorials", "limit": 5, "scrape_options": { "formats": ["markdown"], "engine": "playwright" } }
- Localized search: { "query": "machine learning", "lang": "es", "country": "ES", "limit": 5 } | read | false | unknown |
| anycrawl_crawl Crawl an entire website with configurable depth and limits.
Best for: Multi-page coverage, site mapping, content discovery.
Not recommended for: Single pages (use anycrawl_scrape) or open-ended discovery (use anycrawl_search).
RECOMMENDED: Use 'playwright' engine for best results with dynamic content and modern websites.
Usage (parameters):
- url: Starting URL to crawl (string, required)
- engine: 'playwright' | 'cheerio' | 'puppeteer' (required, default: 'playwright')
- max_depth: Maximum crawl depth (number, optional, default: 10)
- limit: Maximum pages to crawl (number, optional, default: 100)
- strategy: Crawl strategy 'all' | 'same-domain' | 'same-hostname' | 'same-origin' (optional, default: 'same-domain')
- include_paths: Path patterns to include (string[], optional)
- exclude_paths: Path patterns to exclude (string[], optional)
- retry: Enable auto-retry on failure (boolean, optional)
- poll_seconds: Polling interval for job status (number, optional)
- poll_interval_ms: Polling interval in milliseconds (number, optional)
- timeout_ms: Job timeout in milliseconds (number, optional)
- scrape_options: Nested scrape options for each page (object, optional)
Returns: { job_id, status, message } for async jobs
Examples:
- Recommended: { "url": "https://example.com", "engine": "playwright", "limit": 50 }
- Deep crawl: { "url": "https://docs.example.com", "engine": "playwright", "max_depth": 5, "limit": 200 }
- Filtered crawl: { "url": "https://blog.example.com", "engine": "playwright", "include_paths": ["/posts/*"], "exclude_paths": ["/admin/*"] } | read | false | unknown |
- repohttps://github.com/any4ai/anycrawl-mcp-server
- adoption6 stars · 2 forks
The access this server can exercise, inferred from its verified tools — not a declared OAuth scope.
Add the “as seen on MCPExplorer” badge to your README.
This is one server. A loadout combines the right servers, governance, and proven plays for a whole job — assembled deliberately, not tool-dumped.
Explore loadouts →