Docs

MeshArc documentation

Reach every page. Remember every site.

MeshArc is a web scraping API and website crawler that turns URLs into clean Markdown, text, HTML or structured JSON, remembers every site between runs, tells you what changed, and delivers the results to your own data stores. It is a hosted service: a web app at https://mesharc.dev, an API at https://api.mesharc.dev, SDKs for Python and Node, and an MCP server for agents.

Everything about the product is here: how it works, how to get started, every screen, every setting with its default, the connectors and destinations, the API, the SDKs, and what it costs. Current as of 18 September 2026.

What it covers
Getting startedSign up, read a page, watch a site, get the pages somewhere, use the API
Using the appSign-up, workspaces and members, keys, the playground, projects and every project screen, exports, monitor, billing
ConfigurationEvery setting a project, a crawl, a scrape or the playground takes — what it means, its default, what it costs
Connectors and destinationsThe providers, what each needs, how a destination writes rows, the shaping language
The APIAuthentication, conventions, every endpoint, webhooks, error codes, verdicts
SDKs and MCPpip install mesharc, npm install mesharc, the MCP server
Credits, plans and billingWhat a page costs, the plans, limits, overage, payments

How MeshArc thinks#

A project is a site (a seed URL) plus its configuration: scope, sources, schedule, formats, browser steps, tier. It lives in a workspace. A run is one crawl of a project: pages, a stop reason, counts, a credit tally, a sitemap snapshot, a link graph, and a change record against the previous run. A page is one URL in one run: bodies, head metadata, structured fields, a verdict, the engine that read it, the tier, the cost.

The ladder. Every page starts on the cheapest engine that could read it and climbs only when refused. Eight engines, cheapest first:

Engine (wire name)What it isTierCredits
crawler, modern-ua, tlsPlain HTTP with a real browser's TLS fingerprinthttp1
minted (http-jar)The same request carrying cookies a browser earnedhttp2
browserA real Chrome, headed, persistent profilebrowser4
browser-residentialChrome through a residential exitstealth16, +4 per MB through the exit past the first
solverA render plus a paid captcha tokenstealth22 (a widget that passed on a click is billed as a render)
unblockerA rented fetch: someone else's pool and solversstealth40

A refused page climbs one rung; the site's starting rung moves only when three of the last five pages needed it. What a host needs is remembered per domain, across every workspace, for a day — so the platform learns a site once.

The judge. Every response gets a typed verdict: ok, thin (a short page that is still a page — kept, with warnings: ["short"]), shell (an app shell whose content is script — rendered once), blocked_status / blocked_challenge (climb), rate_limited (wait), captcha (solve), geo_block (an exit elsewhere), login_wall (stop), missing (a 404/410, or a 200 that is the site's not-found page — stop). A short page whose markup repeats is what the markup says — shape: listing / table / form — and is ok.

What you pay. A page costs the rung that answered it. A page the site refused, a 404, a soft 404, a page robots.txt kept us from, a cached copy: 0. A render bought to check a short page that turns out unchanged is billed at the http rung. A map costs one credit per sitemap file read. Every job envelope carries creditsUsed; every response carries X-MeshArc-Credits.

Change detection is the product. Two runs of a project give a change record: pages added, removed, modified, same, withheld, unverified, skipped; field changes (title, description, canonical, robots, h1, Open Graph, Twitter card, lang); "the selector broke" signals; volatile blocks filtered. A run that reached under 90 % of the site cannot tell "removed" from "not reached", so removals are withheld and the record says so.

Where the pages go. Webhooks (signed), 19 connectors (object stores, databases, warehouses, vector indexes, streams) bound as destinations with a declarative shape, and exports (JSONL / CSV) of pages, markdown, changes, fields and the sitemap.