Docs · SDKs and MCP

SDKs and the MCP server

Both SDKs are thin: every method is one API call and returns the API's JSON unwrapped, so the API reference is the reference for them too. Both read MESHARC_API_KEY from the environment.

Python#

Python 3.10+, one dependency (httpx), fully typed.

from mesharc import MeshArc, MeshArcError

arc = MeshArc("mesharc_...")                     # or MeshArc() with MESHARC_API_KEY set

# One URL, waited for: the API holds the request open for the page
page = arc.scrape("https://quotes.toscrape.com/js/", config={"render_js": "always"})
print(page["markdown"][:200], page["credits"], page["verdict"])

# Many URLs, no project: grouped by host, fetched in parallel where allowed
batch = arc.scrape(["https://a.com/", "https://b.com/x"], config={"concurrency": 4})
for row in batch["pages"]:
    print(row["url"], row["httpStatus"], row["verdict"], row["credits"])

# What a site declares, before fetching any of it
for u in arc.map("https://docs.example.com"):
    print(u["url"], u["lastmod"])

# A whole site, once, with no project to set up first
job = arc.crawl("https://docs.example.com", limit=200, maxDepth=3,
                scrapeOptions={"formats": ["markdown", "links"]})
for page in job.pages():                         # follows the cursor, keeps up with the crawl
    print(page["url"], page["words"])
print(job.status, job.envelope["counts"], job.envelope["creditsUsed"])
project = job.keep(name="Docs", schedule="weekly")   # ...and keep it, if it is worth watching

# A site watched over time
project = arc.projects.create("https://docs.example.com", name="Docs", schedule="weekly",
                              config={"max_pages": 300, "include_paths": ["/docs/*"]})
run = arc.runs.start(project["id"], wait=True)
record = arc.changes(project["id"])
print(record["change"]["counts"])
hits = arc.search(project["id"], '"free shipping" returns')             # words in the markdown
hits = arc.search(project["id"], "table.pricing td", mode="selector")   # elements in the html
arc.export(project["id"], "pages.jsonl", dataset="markdown")
MethodCalls
extract(url, config, wait, poll, timeout)POST /playground/extract — one URL with every format
scrape(urls, config, webhook_url, formats, wait, …) · scrape_one(url, config, formats, wait, timeout_s, idempotency_key)POST /scrape
crawl(url, wait, poll, timeout, idempotency_key, **opts) → Crawl (.status, .envelope, .refresh(), .wait(), .pages(cursor=…), .page(url), .keep(), .cancel()) · get_crawl(id)POST /crawl, GET /crawl/{id}, …/page, …/keep, DELETE
map(url, **opts) → rows · map_details(url, search, limit, …) → the envelopePOST /map
batch(batch_id, formats, wait, …)GET /scrape/{id}
projects.list() / create(seed, name, schedule, retention, config) / get / update(project_id, **fields) / delete/projects
runs.list(project_id, limit) / get / start(project_id, wait, poll, timeout) / wait / cancel/projects/{id}/runs
pages(project_id, run_id) · page(project_id, url, run_id) · changes(project_id, run_id) · page_diff(project_id, url, run_id) · search(project_id, q, mode, run_id) · recrawl(project_id, urls) · sources(project_id) · export(project_id, path, dataset, fmt, run_id, urls)The project routes
me() · keys() · create_key(name, scopes, projects, expires_in_days, rpm) · revoke_key(key_id) · usage() · monitor() · meta()The workspace routes
close()Close the HTTP client

Errors raise MeshArcError with status, detail, code (validation, rate_limited, plan_limit, not_found, …) and request_id — the id the API put on the response header and in its logs. A job the client stopped waiting for raises MeshArcTimeoutError (also a TimeoutError) with its job_id. Requests time out after timeout seconds (150) and are retried on 429, 502, 503, 504 and network failures when safe to repeat — a GET, a DELETE, or a POST with an idempotency_key= — up to max_retries times (2).

Node#

Node 20+, zero dependencies (native fetch). TypeScript types, ESM and CommonJS builds.

import { MeshArc, MeshArcError } from 'mesharc';

const arc = new MeshArc('mesharc_...');            // or new MeshArc() with MESHARC_API_KEY set

const page = await arc.scrape('https://quotes.toscrape.com/js/', { render_js: 'always' });
console.log((page.markdown as string).slice(0, 200), page.credits, page.verdict);

const batch = await arc.scrape(['https://a.com/', 'https://b.com/x'], { concurrency: 4 });
for (const row of batch.pages) console.log(row.url, row.httpStatus, row.verdict, row.credits);

for (const u of await arc.map('https://docs.example.com')) console.log(u.url, u.lastmod);

const job = await arc.crawl('https://docs.example.com', { limit: 200, maxDepth: 3 });
for await (const p of job.pages()) console.log(p.url, p.words);       // follows the cursor
const kept = await job.keep({ name: 'Docs', schedule: 'weekly' });

const project = await arc.projects.create('https://docs.example.com', { name: 'Docs', schedule: 'weekly' });
const run = await arc.runs.start(project.id, { wait: true });
const record = await arc.changes(project.id);
const hits = await arc.search(project.id, 'table.pricing td', 'selector');
const csv = await (await arc.export(project.id, { dataset: 'pages', format: 'csv' })).text();
MethodCalls
extract(url, config?, {wait, poll, timeout})POST /playground/extract
scrape(urls, config?, {webhookUrl, formats, apiTimeoutS, idempotencyKey, …}) · scrapeOne(url, config?, {…})POST /scrape
crawl(url, {limit, maxDepth, …, wait, idempotencyKey}) → Crawl (.status, .envelope, .refresh(), .wait(), .pages() async iterator, .keep(), .cancel()) · getCrawl(id)/crawl
map(url, {search, limit}) → rows · mapDetails(url, {…}) → the envelopePOST /map
batch(batchId, {formats, wait})GET /scrape/{id}
projects.list() / create(seed, {name, schedule, retention, config}) / get / update(id, fields) / delete · runs.list / get / start(id, {wait}) / wait / cancel/projects
pages, page, changes, pageDiff, `search(projectId, q, 'content''selector', runId?), recrawl, sources, export(projectId, {dataset, format, runId, urls})→Response`
me(), keys(), createKey(name, {scopes, projects, expiresInDays, rpm}), revokeKey, usage(), monitor(), meta()The workspace routes
call(method, path, body?, params?) · raw(...)Any route

Errors throw MeshArcError with status, detail, code and requestId; a job the client stopped waiting for throws MeshArcTimeoutError with its jobId. Requests time out after timeoutMs (150 000) and are retried on 429, 502, 503, 504 and network failures when safe to repeat — a GET, a DELETE, or a POST with an idempotencyKey — up to maxRetries times (2).

The MCP server#

The Python package also ships MeshArc as an MCP server, so Claude Desktop, Claude Code, Cursor and any MCP client can scrape, crawl, map, set up and change projects, and read change records as tools. There is a hosted one at https://mcp.mesharc.dev/mcp -- paste the URL, sign in, approve, and nothing is installed -- or you can run it yourself over stdio with your own key. Either way it sits on top of the Python client, so what an agent gets is what the API gives. A result with many pages in it is an index of the pages read (up to 500, fewer if the budget needs the room), an excerpt of as many as a 60,000-character budget pays for (fifty at most), and a count of each page’s links rather than the links; ask for one page on its own, each body up to 12,000 characters, with get_job(kind="crawl", id=…, url=…), or carry on through a crawl with cursor=. Hosted, a connection approved read-only can read what the workspace has stored — projects, pages, change records, search, get_job — but not fetch a new page or change anything: those tools reach the site or alter the workspace, and most of them spend credits; those nine tools answer read_only until the app is reconnected with write access. Python 3.10+. The MCP page walks through setup for each client and shows real conversations.

pip install "mesharc[mcp]"
MESHARC_API_KEY=mesharc_... mesharc-mcp

# Claude Code
claude mcp add mesharc -e MESHARC_API_KEY=mesharc_... -- mesharc-mcp
ToolWhat it does
scrape_urls(urls, config?, formats?, parse_documents?)Content for a list of pages; PDFs and office files read as text unless told not to
extract_url(url, config?, parse_documents?)One page with every format
map_site(url, search?, limit?)The site's URLs from its sitemap
crawl_site(url, limit?, max_depth?, include_paths?, …)A one-shot crawl, pages returned as they land; the paths are globs (/blog/*)
keep_crawl_as_project(crawl_id, name?, schedule?)Promote a crawl
describe_project_config()Every setting an assistant can set, its meaning and default — read before setting one
list_projects() · get_project(project_id) · create_project(seed, name?, schedule?, config?) · update_project(project_id, name?, schedule?, config?) · start_run(project_id, wait?)Projects, their settings, and runs; a partial config on update merges, and a key that is not a setting is refused by name
list_pages(project_id, run_id?) · get_page(project_id, url, run_id?)A run's pages
get_changes(project_id, run_id?)The change record
search_pages(project_id, q, mode?, run_id?) · recrawl_pages(project_id, urls)Search; re-crawl now

The clients are open source: mesharc-python and mesharc-node.