SDKs and the MCP server
Both SDKs are thin: every method is one API call and returns the API's JSON unwrapped, so the API reference is the reference for them too. Both read MESHARC_API_KEY from the environment.
Python#
Python 3.10+, one dependency (httpx), fully typed.
from mesharc import MeshArc, MeshArcError
arc = MeshArc("mesharc_...") # or MeshArc() with MESHARC_API_KEY set
# One URL, waited for: the API holds the request open for the page
page = arc.scrape("https://quotes.toscrape.com/js/", config={"render_js": "always"})
print(page["markdown"][:200], page["credits"], page["verdict"])
# Many URLs, no project: grouped by host, fetched in parallel where allowed
batch = arc.scrape(["https://a.com/", "https://b.com/x"], config={"concurrency": 4})
for row in batch["pages"]:
print(row["url"], row["httpStatus"], row["verdict"], row["credits"])
# What a site declares, before fetching any of it
for u in arc.map("https://docs.example.com"):
print(u["url"], u["lastmod"])
# A whole site, once, with no project to set up first
job = arc.crawl("https://docs.example.com", limit=200, maxDepth=3,
scrapeOptions={"formats": ["markdown", "links"]})
for page in job.pages(): # follows the cursor, keeps up with the crawl
print(page["url"], page["words"])
print(job.status, job.envelope["counts"], job.envelope["creditsUsed"])
project = job.keep(name="Docs", schedule="weekly") # ...and keep it, if it is worth watching
# A site watched over time
project = arc.projects.create("https://docs.example.com", name="Docs", schedule="weekly",
config={"max_pages": 300, "include_paths": ["/docs/*"]})
run = arc.runs.start(project["id"], wait=True)
record = arc.changes(project["id"])
print(record["change"]["counts"])
hits = arc.search(project["id"], '"free shipping" returns') # words in the markdown
hits = arc.search(project["id"], "table.pricing td", mode="selector") # elements in the html
arc.export(project["id"], "pages.jsonl", dataset="markdown")
| Method | Calls |
|---|---|
extract(url, config, wait, poll, timeout) | POST /playground/extract — one URL with every format |
scrape(urls, config, webhook_url, formats, wait, …) · scrape_one(url, config, formats, wait, timeout_s, idempotency_key) | POST /scrape |
crawl(url, wait, poll, timeout, idempotency_key, **opts) → Crawl (.status, .envelope, .refresh(), .wait(), .pages(cursor=…), .page(url), .keep(), .cancel()) · get_crawl(id) | POST /crawl, GET /crawl/{id}, …/page, …/keep, DELETE |
map(url, **opts) → rows · map_details(url, search, limit, …) → the envelope | POST /map |
batch(batch_id, formats, wait, …) | GET /scrape/{id} |
projects.list() / create(seed, name, schedule, retention, config) / get / update(project_id, **fields) / delete | /projects |
runs.list(project_id, limit) / get / start(project_id, wait, poll, timeout) / wait / cancel | /projects/{id}/runs |
pages(project_id, run_id) · page(project_id, url, run_id) · changes(project_id, run_id) · page_diff(project_id, url, run_id) · search(project_id, q, mode, run_id) · recrawl(project_id, urls) · sources(project_id) · export(project_id, path, dataset, fmt, run_id, urls) | The project routes |
me() · keys() · create_key(name, scopes, projects, expires_in_days, rpm) · revoke_key(key_id) · usage() · monitor() · meta() | The workspace routes |
close() | Close the HTTP client |
Errors raise MeshArcError with status, detail, code (validation, rate_limited, plan_limit, not_found, …) and request_id — the id the API put on the response header and in its logs. A job the client stopped waiting for raises MeshArcTimeoutError (also a TimeoutError) with its job_id. Requests time out after timeout seconds (150) and are retried on 429, 502, 503, 504 and network failures when safe to repeat — a GET, a DELETE, or a POST with an idempotency_key= — up to max_retries times (2).
Node#
Node 20+, zero dependencies (native fetch). TypeScript types, ESM and CommonJS builds.
import { MeshArc, MeshArcError } from 'mesharc';
const arc = new MeshArc('mesharc_...'); // or new MeshArc() with MESHARC_API_KEY set
const page = await arc.scrape('https://quotes.toscrape.com/js/', { render_js: 'always' });
console.log((page.markdown as string).slice(0, 200), page.credits, page.verdict);
const batch = await arc.scrape(['https://a.com/', 'https://b.com/x'], { concurrency: 4 });
for (const row of batch.pages) console.log(row.url, row.httpStatus, row.verdict, row.credits);
for (const u of await arc.map('https://docs.example.com')) console.log(u.url, u.lastmod);
const job = await arc.crawl('https://docs.example.com', { limit: 200, maxDepth: 3 });
for await (const p of job.pages()) console.log(p.url, p.words); // follows the cursor
const kept = await job.keep({ name: 'Docs', schedule: 'weekly' });
const project = await arc.projects.create('https://docs.example.com', { name: 'Docs', schedule: 'weekly' });
const run = await arc.runs.start(project.id, { wait: true });
const record = await arc.changes(project.id);
const hits = await arc.search(project.id, 'table.pricing td', 'selector');
const csv = await (await arc.export(project.id, { dataset: 'pages', format: 'csv' })).text();
| Method | Calls |
|---|---|
extract(url, config?, {wait, poll, timeout}) | POST /playground/extract |
scrape(urls, config?, {webhookUrl, formats, apiTimeoutS, idempotencyKey, …}) · scrapeOne(url, config?, {…}) | POST /scrape |
crawl(url, {limit, maxDepth, …, wait, idempotencyKey}) → Crawl (.status, .envelope, .refresh(), .wait(), .pages() async iterator, .keep(), .cancel()) · getCrawl(id) | /crawl |
map(url, {search, limit}) → rows · mapDetails(url, {…}) → the envelope | POST /map |
batch(batchId, {formats, wait}) | GET /scrape/{id} |
projects.list() / create(seed, {name, schedule, retention, config}) / get / update(id, fields) / delete · runs.list / get / start(id, {wait}) / wait / cancel | /projects |
pages, page, changes, pageDiff, `search(projectId, q, 'content' | 'selector', runId?), recrawl, sources, export(projectId, {dataset, format, runId, urls})→Response` |
me(), keys(), createKey(name, {scopes, projects, expiresInDays, rpm}), revokeKey, usage(), monitor(), meta() | The workspace routes |
call(method, path, body?, params?) · raw(...) | Any route |
Errors throw MeshArcError with status, detail, code and requestId; a job the client stopped waiting for throws MeshArcTimeoutError with its jobId. Requests time out after timeoutMs (150 000) and are retried on 429, 502, 503, 504 and network failures when safe to repeat — a GET, a DELETE, or a POST with an idempotencyKey — up to maxRetries times (2).
The MCP server#
The Python package also ships MeshArc as an MCP server, so Claude Desktop, Claude Code, Cursor and any MCP client can scrape, crawl, map, set up and change projects, and read change records as tools. There is a hosted one at https://mcp.mesharc.dev/mcp -- paste the URL, sign in, approve, and nothing is installed -- or you can run it yourself over stdio with your own key. Either way it sits on top of the Python client, so what an agent gets is what the API gives. A result with many pages in it is an index of the pages read (up to 500, fewer if the budget needs the room), an excerpt of as many as a 60,000-character budget pays for (fifty at most), and a count of each page’s links rather than the links; ask for one page on its own, each body up to 12,000 characters, with get_job(kind="crawl", id=…, url=…), or carry on through a crawl with cursor=. Hosted, a connection approved read-only can read what the workspace has stored — projects, pages, change records, search, get_job — but not fetch a new page or change anything: those tools reach the site or alter the workspace, and most of them spend credits; those nine tools answer read_only until the app is reconnected with write access. Python 3.10+. The MCP page walks through setup for each client and shows real conversations.
pip install "mesharc[mcp]"
MESHARC_API_KEY=mesharc_... mesharc-mcp
# Claude Code
claude mcp add mesharc -e MESHARC_API_KEY=mesharc_... -- mesharc-mcp
| Tool | What it does |
|---|---|
scrape_urls(urls, config?, formats?, parse_documents?) | Content for a list of pages; PDFs and office files read as text unless told not to |
extract_url(url, config?, parse_documents?) | One page with every format |
map_site(url, search?, limit?) | The site's URLs from its sitemap |
crawl_site(url, limit?, max_depth?, include_paths?, …) | A one-shot crawl, pages returned as they land; the paths are globs (/blog/*) |
keep_crawl_as_project(crawl_id, name?, schedule?) | Promote a crawl |
describe_project_config() | Every setting an assistant can set, its meaning and default — read before setting one |
list_projects() · get_project(project_id) · create_project(seed, name?, schedule?, config?) · update_project(project_id, name?, schedule?, config?) · start_run(project_id, wait?) | Projects, their settings, and runs; a partial config on update merges, and a key that is not a setting is refused by name |
list_pages(project_id, run_id?) · get_page(project_id, url, run_id?) | A run's pages |
get_changes(project_id, run_id?) | The change record |
search_pages(project_id, q, mode?, run_id?) · recrawl_pages(project_id, urls) | Search; re-crawl now |
The clients are open source: mesharc-python and mesharc-node.