servers / app-crawlio-crawlio-browser
app.crawlio/crawlio-browser MCP server
communitystdiolocaldestructive capablehealthy
Bridge a live Chrome browser to your agent: 147 CDP tools for capture, extraction, and recording.
01Tools · 7
How to read this: tool names here are observed from a live tools/list handshake. The Risk label is a heuristic inferred from the tool name (write/destructive verbs), not from executing the tool — a conservative guess, not a verified capability. We never escalate risk from a description. Found one that's wrong? Tell us — we fix on report.
| Tool | Risk | Side effects | Approval |
|---|---|---|---|
| observe Manage extension-resident observation. Actions: training_start, training_status, training_stop, training_clear, training_artifacts; recording_start, recording_status, recording_stop, recording_clear, recording_artifacts (RecordingBundle aliases); monitor_start, monitor_status, monitor_results, monitor_stop, monitor_clear. Training and monitors continue while MCP is disconnected; stop materializes canonical files and confirmed clear deletes only Chrome-retained records. | unknown | unknown | unknown |
| connect_tab Pin a browser tab for subsequent commands and start CDP capture. Three modes: (1) provide a URL: with the optional tabs grant Crawlio reuses a match, otherwise it creates a fresh owned tab; (2) provide a tabId to adopt a specific existing tab (tabs grant required); (3) no args to discover and pin the active tab (tabs grant required). Pass background:true to avoid stealing the user's active tab/window focus. | unknown | unknown | unknown |
| search Search available commands by keyword — both browser automation (via bridge.send) and Crawlio HTTP endpoints (via crawlio.api). Returns matching command names, descriptions, and parameter schemas. Use this to discover what commands are available before writing execute() code. | read | false | unknown |
| cancel_job Cancel a running background execute job by jobId — terminates its sandbox worker. No-op if the job already finished. | destructive | true | true |
| list_jobs List background execute jobs (running + recently finished). Returns { count, jobs: [{ jobId, status, ageMs, runtimeMs? }] }. | read | false | unknown |
| get_job_result Poll a background execute job by jobId (returned by execute({ background: true })). Returns { status: running|done|error|cancelled, value?, console?, error?, ageMs, runtimeMs? }, or { status: 'not_found' } if unknown/expired (finished jobs are kept ~10 min). phase/percent come from reportPhase(name, percent) calls inside the job; phases is the recent timeline while it is still running. | read | false | unknown |
| execute Execute JavaScript code with access to the browser bridge, Crawlio HTTP client, and smart object.
Use search() first to discover available commands and their parameters.
IMPORTANT WARNINGS:
- smart.screenshot() does NOT exist. For screenshots: bridge.send({ type: 'take_screenshot' }).
- For structured page evidence, prefer smart.extractPage() — runs 7 ops in parallel with typed gaps[].
- capture_page returns a ~1KB shaped summary. For raw data, use stop_network_capture or get_console_logs.
- Use smart.waitForIdle() instead of sleep(). Use smart.scrollCapture() instead of manual scroll loops.
- Scope large snapshots with smart.snapshot({ compact: true, maxDepth: 8, selector: '#main' }); use { interactive: true } for controls only.
- For cross-page navigation, use smart.navigate(url) — never location.href = "..." (breaks CDP).
Available in scope:
- bridge.send(command, timeout?) — send command to browser extension via WebSocket
command must have a `type` field matching a command name (e.g. { type: 'list_tabs' })
- crawlio.api(method, path, body?) — generic HTTP to ControlServer
e.g. await crawlio.api('GET', '/status')
e.g. await crawlio.api('POST', '/start', { url: 'https://example.com' })
e.g. await crawlio.api('POST', '/export', { format: 'zip', destinationPath: '/tmp/site.zip' })
e.g. await crawlio.api('PATCH', '/settings', { settings: { maxConcurrent: 8 } })
Returns { status: number, data: unknown }
- crawlio.getStatus() — shortcut for GET /status
- crawlio.startCrawl(url) — shortcut for POST /start
- crawlio.getEnrichment(url?) — shortcut for GET /enrichment
- crawlio.getCrawledURLs(params?) — shortcut for GET /crawled-urls
- crawlio.postEnrichment(url, data) — shortcut for POST /enrichment/bundle
- sleep(ms) — async wait (max 30s)
- TIMEOUTS — per-command timeout constants
- compileRecording(session, { name, description? }) — compile RecordingSession to SKILL.md
Returns { skillMarkdown, name, pageCount, interactionCount }
- ocrScreenshot(opts?) — extract text from current page via macOS Vision.framework OCR (macOS only)
opts: { fullPage?: boolean, selector?: string }
Returns { regions: [{ text, confidence, bounds }], regionsLimited? }
- saveArtifact(name, contents, opts?) — persist bulk data straight to disk (host-side), bypassing the return value.
Use this for large extractions instead of returning megabytes: accumulate rows in the sandbox, then write them.
name: safe relative path (e.g. 'hfs/cross_ref.ndjson'); contents: string (objects are JSON-stringified);
opts: { append?: boolean } to stream in chunks. Writes under ~/.crawlio/artifacts (CRAWLIO_ARTIFACT_DIR).
Returns { path, bytes }. Ideal pattern: authed in-page fetch via bridge → accumulate → saveArtifact → return { path, bytes }.
- smart — auto-waiting wrappers and framework-specific data accessors:
smart.evaluate(expr) → {result, type} — access .result for value. Never JSON.parse() the return directly.
smart.click(selector, opts?) — poll + click + 500ms settle (accepts CSS or snapshot [ref=X])
smart.type(selector, text, opts?) — poll + type + 300ms settle
smart.navigate(url, opts?) — navigate + 1000ms settle
smart.waitFor(selector, timeout?) — poll until actionable
smart.snapshot(opts?) — capture accessibility snapshot. opts: { interactive?: boolean, compact?: boolean, maxDepth?: number, selector?: string }.
smart.scrollCapture(opts?) — state-aware page scroll with screenshots, stops at page bottom
smart.waitForIdle(timeout?) — wait for DOM mutations to settle (500ms quiet window)
smart.extractPage(opts?) — capture_page + perf + security + fonts + meta + accessibility + mobileReadiness. Returns { capture, performance, security, fonts, meta, accessibility, mobileReadiness, gaps[] }. opts: { trace: true } adds _trace.
smart.comparePages(urlA, urlB, opts?) — navigate to each URL, run extractPage(), return { siteA, siteB, scaffold }. scaffold has dimensions[], sharedFields, missingFields, metrics.
smart.finding({ claim, evidence, sourceUrl, confidence, method, dimension? }) — create validated Finding, accumulate in session. Confidence auto-capped if dimension has active gap with reducesConfidence.
smart.findings() — return all accumulated Finding[] from current session.
smart.clearFindings() — reset accumulated findings and session gaps.
smart.detectTables(opts?) — find repeating data patterns in the page. Native <table> elements first (strategy 'table-rows'), then class-frequency div-soup scan with geometric filters (strategy 'sibling-repeated-blocks'). Returns TableCandidate[] (selector, score, rowCount, sampleText, confidence 0..1, strategy, warnings[]).
smart.extractTable(selector, opts?) — extract structured data from a container. Native tables get thead-derived column names + <name>_url companions; div-soup gets semantic link-first naming. Returns { columns, rows, totalRows, truncated }. opts: { maxRows: 200 }.
smart.detectSections(opts?) — perceive page structure: depth-capped SectionNode tree from semantic tags, ARIA landmarks, data-component/testid attrs, and PascalCase/BEM class hints. Each node: role, source, name, selector (+ selectorVerified), box, inViewport, interactiveCount, textLength, headings, children. Emits SELECTORS, never [ref=eN] handles — compose with bridge.send({ type: 'browser_snapshot', selector }) for fresh interaction refs inside a region. opts: { maxDepth: 2, maxSections: 40 }.
smart.waitForNetworkIdle(opts?) — wait for all network requests to settle (CDP-level, catches fetch/XHR/images/CSS/fonts). Returns { status, elapsed }. opts: { timeout: 15000, idleTime: 500 }.
smart.extractData(opts?) — compound: detectTables + extractTable + JSON-LD. Returns { tables, structuredData, url }.
smart.parseTrackingPixels() — parse captured network data for tracking pixel fires (Facebook, GA4, TikTok, LinkedIn, Pinterest). Returns { totalPixelFires, vendors, pixels, events, unrecognizedTrackingUrls }.
smart.validateTracking() — validate tracking events against per-vendor parameter schemas (Facebook 18 standard + GA4 recommended). Returns { events, issues, errorCount, warningCount, infoCount, isHealthy }. Each issue has severity (error/warning/info), code, message, recommendation, and optional parameter.
smart.inspectDataLayer() — inspect tracker runtime state via CDP (fbq queue, GA4 dataLayer, GTM containers, TikTok ttq). Returns DataLayerState with null for absent trackers. No content script needed.
smart.detectDuplicates() — detect duplicate pixel fires grouped by vendor+pixelId+eventName+URL. Excludes PageView (legitimate SPA behavior). Returns DuplicateCluster[] with count and timestamps.
smart.detectTechnologies(opts?) — detect technologies via fingerprint matching against CDP signals (headers, scripts, JS globals, meta, cookies, URL). Returns TechnographicResult { technologies[], categories, totalDetected, highConfidenceCount, signalsUsed }. Each technology has numeric confidence (0-100, additive), version, matchedSignals[]. opts: { confidenceThreshold: 1 }.
smart.diffSnapshots(before?) — Myers diff current ARIA snapshot against baseline. If before omitted, uses last cached snapshot. Returns { diff, additions, removals, unchanged, changed }.
Framework namespaces (injected based on detected framework):
smart.react.{getVersion,getRootCount,hasProfiler,isHookInstalled}
smart.vue.{getVersion,getAppCount,getConfig,isDevMode}
smart.angular.{getVersion,isDebugMode,isIvy,getRootCount,getState}
smart.svelte.{getVersion,getMeta,isDetected}
smart.redux.{isInstalled,getStoreState}
smart.alpine.{getVersion,getStoreKeys,getComponentCount}
smart.nextjs.{getData,getRouter,getSSRMode,getRouteManifest}
smart.nuxt.{getData,getConfig,isSSR}
smart.remix.{getContext,getRouteData}
smart.gatsby.{getData,getPageData}
smart.shopify.{getShop,getCart}
smart.wordpress.{isWP,getRestUrl,getPlugins} | smart.woocommerce.{getParams}
smart.laravel.{getCSRF} | smart.django.{getCSRF} | smart.drupal.{getSettings}
smart.jquery.{getVersion}
Example (HTTP API):
const { data } = await crawlio.api('GET', '/status');
return data;
Example (browser):
const tabs = await bridge.send({ type: 'list_tabs' }, 5000);
return tabs;
Example (smart — auto-waiting click):
await smart.click('#submit-btn');
return await smart.snapshot();
Example (smart — framework data):
const nextData = await smart.nextjs?.getData();
return { page: nextData?.page, buildId: nextData?.buildId };
Example (session recording + compile):
const s = await bridge.send({ type: 'start_recording', maxDurationSec: 120 });
// ... interact with page ...
const session = await bridge.send({ type: 'stop_recording' });
const skill = compileRecording(session, { name: 'my-flow' });
return skill;
IMPORTANT: Keep scripts fast (<15s). Each smart.click costs ~1-2s. Never loop 5+ clicks — use smart.evaluate to read DOM data in bulk instead.
IMPORTANT: smart.evaluate returns {result, type}. Access .result for the value. Never JSON.stringify inside evaluate then JSON.parse outside — just return objects directly. | write | true | unknown |
02Install & source
npx -y crawlio-browser
npx- repohttps://github.com/Crawlio-app/crawlio-browser
- homepagehttps://docs.crawlio.app/browser-agent/overview
- licenseApache-2.0
- adoption5 stars · 0 forks
05Provenance & freshness
sourcesOfficial MCP Registry [p1]
last_checked2026-08-16 18:23Z
next_check2026-08-16 21:20Z
cadenceevery 3h
verifiedtools_list:passed handshake:passed metadata:passed tools_list:passed handshake:passed metadata:passed tools_list:passed handshake:passed metadata:passed tools_list:passed
index_statusindex — 8 unique facts >= 5
06Badge
Add the “as seen on MCPExplorer” badge to your README.
[](https://mcpexplorer.com/servers/app-crawlio-crawlio-browser)
Next step
This is one server. A loadout combines the right servers, governance, and proven plays for a whole job — assembled deliberately, not tool-dumped.
Explore loadouts →