The API
Base URL https://api.mesharc.dev/api/v1. The OpenAPI spec is at https://api.mesharc.dev/docs. The app uses these same routes; there is no private API behind it.
Authentication#
Authorization: Bearer mesharc_…
A key is issued under Settings → API keys (or POST /me/keys) and shown once. It carries scopes (read < write < admin), optionally a set of projects it may see, an expiry, and a rate limit. The app signs in with a session cookie instead and reaches the API through its own /api/mesharc proxy; the routes are the same.
A route that changes something needs write (the member role); keys, the rate ceiling, deleting a project, webhook secrets, connections and destinations need admin. A session whose email is not verified is refused those with 403 email_unverified, and one in a workspace that requires two-factor authentication without having it on is refused with 403 mfa_required; a session with two-factor on that last showed its code more than ten minutes ago gets 403 mfa_stale from the handful of actions that ask again (POST /auth/2fa/verify with a code clears it). API keys are unaffected by all three.
Conventions#
| Job envelope | Every asynchronous verb answers the same shape: {id, kind, status, progress: {done, total, queued}, creditsUsed, counts, data: [...], next, cursor, stop, request_id, error}. status is queued / running / done / error / cancelled; next is a ready URL for the page after this one. |
| Pagination | ?cursor=&limit= — limit up to 1,000 (25 by default on a crawl's pages; up to 100 without bodies); next carries the cursor. |
| Idempotency | Idempotency-Key: <your key> on POST /scrape, /crawl and /map returns the first answer for 24 hours, scoped to the route, so a retry after a dropped connection never starts a second job. POST /projects/{id}/runs needs no key: a second run while one is queued or running answers 409. |
| Request ids | Every response carries X-Request-Id (yours if you sent one) and the body's request_id; quote it to support. |
| Credits | X-MeshArc-Credits on a scrape's response says what it cost; every envelope carries creditsUsed. |
| Rate limits | X-RateLimit-Limit, -Remaining, -Reset per key; 429 when exceeded. |
| Errors | {"error": "<what happened>", "code": "<what it means>", "request_id": "…"}. Codes: validation (400 / 422), unauthorized (401), plan_limit (402 — a plan that does not include something, or credits spent), forbidden (403), email_unverified (403), not_found (404 — including anything the key may not see; existence is never confirmed), conflict (409), too_large (413), rate_limited (429), unavailable (503), internal (500). |
| Time | ISO 8601 UTC everywhere. |
The verbs#
POST /scrape — URLs in, content out, no project#
{"url": "https://example.com/pricing", "formats": "markdown,links", "timeout": 60,
"config": {"render_js": "auto"}, "scrapeOptions": {"onlyMainContent": true}, "maxCredits": 4}
One url waits for the page and returns it — up to timeout seconds (60 by default, 120 at most), then 202 with the id to poll, the fetch carrying on either way. A list of urls (up to 500) is always 202: grouped by host, fanned out, polled at GET /scrape/{id} (a row per URL with httpStatus, verdict, credits) or announced to webhook_url (batch.finished, signed with the secret returned once). DELETE /scrape/{id} cancels; GET /scrape lists recent ones. formats picks the bodies returned (markdown,text,cleanHtml,rawHtml). maxCredits caps what one page may cost. destination names one of the workspace's destinations — a database, a warehouse, a bucket — and the result is written there as well as returned: a single URL before it answers, so the response carries destinationSync (status, written, skipped, location, error); a batch once, when its last URL is in, reported on GET /scrape/{id}. The destination is checked before anything is queued: another workspace's is a 404, a paused one a 400.
A page row: url, httpStatus, status, verdict, errorCode, shape, warnings, signals, method, tier, credits, billedAs, climbedTo, words, language, markdown, text, cleanHtml, html, links, fields, screenshot, head, reason, crawledAt.
POST /crawl — a one-shot crawl, no project first#
{"url": "https://docs.example.com/", "limit": 200, "maxDepth": 3,
"includePaths": ["/docs/*"], "crawlMode": "sitemap_first", "maxTier": "browser",
"scrapeOptions": {"formats": ["markdown", "links"]},
"webhook": {"url": "https://you.example/hook", "events": ["crawl.completed"], "metadata": {"ref": "job-7"}}}
202 with the envelope; the crawl runs as a hidden, ephemeral project, kept for a day unless you keep it with POST /crawl/{id}/keep. GET /crawl/{id}?cursor=&limit=&formats= pages through its pages oldest first; GET /crawl/{id}/page?url= returns one of them in full (matched under either scheme and with or without a trailing slash); DELETE /crawl/{id} stops it; POST /crawl/{id}/keep {name, schedule, retention} promotes it to a standing project with its run already in place. The short names (limit, maxDepth, includePaths, excludePaths, useSitemap, sitemapMode, crawlMode, allowSubdomains, maxAge, maxTier, delay, concurrency, respectRobots) and scrapeOptions map onto the configuration; config takes any key directly. The webhook's secret comes back once as webhookSecret. destination writes the crawl's pages to one of the workspace's destinations when it is done — every page, as a baseline — and GET /crawl/{id} carries how the write went under destinationSync.
POST /map — the site's URLs from its sitemap#
{"url": "https://www.gov.uk/", "search": "visa", "limit": 500, "sitemapMode": "auto", "timeout": 10}
Synchronous when the tree resolves within timeout (10 s by default, 60 at most), else 202 and GET /map/{id}. Rows: {url, lastmod, changefreq, source, file, section}; totals (files, urls, indexes, blocked), robots, method. limit up to 50,000 (5,000 by default); search filters on the URL. Costs 1 credit per sitemap file read — most sites are one file; gov.uk's 29 are 29.
POST /playground/extract — one URL with every format, for the app#
{url, allow_browser, use_proxy, config} → an id; GET /playground/{id} for the result and its evidence; …/screenshot for the image.
Projects#
| Route | What |
|---|---|
GET /projects · POST /projects {seed, name?, schedule?, retention?, config?} · POST /projects/batch {seeds[], template_project_id?, schedule?, retention?} | List; create one (schedule defaults to manual here — the app sends weekly); create many from a seed list with a template's config |
GET /projects/{id} · PATCH /projects/{id} {name?, schedule?, retention?, config?} · DELETE /projects/{id} (admin) | Read; change anything; delete with its runs and pages |
POST /projects/{id}/config/apply {project_ids[], groups[]} | Copy this project's saved config onto others, by group (scope, fetch, content, output, …) |
POST /projects/run {project_ids[]} | Start several |
POST /projects/{id}/dry-run | Resolve the sources and sample URLs without spending |
GET /projects/{id}/sources · POST …/sources/preview {sitemap_include, sitemap_exclude, include_paths, …} | The seed, sitemap, URL list, feeds and patterns with what the last run found through each; how many URLs a selection keeps |
GET /projects/{id}/discovery · POST …/discovery · GET …/discovery/entries?file=&cursor= | The stored sitemap tree (files, sections, counts, lastmod ranges, the diff against the last snapshot); take a new snapshot (a map, billed); page through a file's entries |
POST /projects/{id}/selectors/match {selectors[]} · POST …/content/preview {config} | Match counts against the last run's stored HTML; the markdown a config would produce for the seed |
Runs and pages#
| Route | What |
|---|---|
GET /projects/{id}/runs?limit=&cursor=&v= · POST /projects/{id}/runs {trigger?} · GET …/runs/{run} · POST …/runs/{run}/cancel | History (25 a page); start (202); one run with its counts, credits, stop reason, engine breakdown, log; cancel |
GET /projects/{id}/pages?run_id=&v=&flag=&cursor=&limit= | Rows of a run (100 a page, 1,000 at most). v filters by verdict; flag=orphan|unlisted|dead_end. The app's query language (status:403 depth:>2) is applied on the rows |
GET /projects/{id}/pages/content?url=&run_id=&format= | One page's bodies, head, fields, versions |
GET /projects/{id}/pages/screenshot?url=&run_id= | The JPEG |
POST /projects/{id}/pages/search {q, mode: content|selector, run_id?} | Which pages say this (words, "phrases") or contain this (CSS / XPath), scanned on the server |
POST /projects/{id}/pages/recrawl {urls[]} | A scoped run of just these, now |
GET /projects/{id}/pages/removed?run_id= | Pages in an earlier run and not this one |
GET /projects/{id}/runs/{run}/links?kind=orphans|unlisted|dead_ends | The link graph's findings |
Changes#
| Route | What |
|---|---|
GET /projects/{id}/changes?run_id= | The change record: {counts: {added, removed, modified, same, withheld, unverified, skipped, fields, selectorBreaks, volatileBlocks}, coverage, baseline, entries[]} |
POST /projects/{id}/changes/{change}/review {reviewed: true} | Mark reviewed |
GET /projects/{id}/changes/page?url=&run_id= | The word-level content diff of one page against the run before |
GET /projects/{id}/changes/markup?url=&run_id= | The markup comparison: DOM diff, noise filter, selector repointing |
Exports#
GET /exports lists what can be exported. GET /projects/{id}/export?dataset=pages|markdown|changes|fields|sitemap&format=jsonl|csv&run_id= streams a dataset; POST /projects/{id}/export {dataset, format, urls[]} restricts it to listed URLs.
Two more datasets are text files, for a model rather than a table: dataset=llms is an llms.txt — the run's usable pages (ok, thin) as links with their descriptions, grouped by first path segment, under the site's title and the home page's description — and dataset=llms-full is the llms-full.txt: every usable page's markdown in one file, each under its title and a Source: line. Both take format=txt (the default for them) and the same urls[] selection.
Connections and destinations#
| Route | What |
|---|---|
GET /connections/providers | The catalogue with each provider's fields, secrets, note and verified |
GET /connections · POST /connections {provider, name, config, secrets} · GET/PATCH/DELETE /connections/{id} · POST /connections/{id}/test | Workspace connections; secrets are never returned |
GET /destinations · POST /destinations · PATCH/DELETE /destinations/{id} · GET …/syncs · POST …/sync | Shared destinations (scope all or projects) and their syncs |
GET /projects/{id}/destinations · POST … · PATCH/DELETE …/{dest} · GET …/{dest}/syncs · POST …/{dest}/sync | A project's own destinations |
POST /destinations/preview {project_id, shape, columns} · GET /destinations/shape-functions | The shaped rows of the last run; the function and operator lists |
What goes in a destination, and how a shape is written, is on the Connectors and destinations page.
Webhooks#
Project webhooks — webhook_url and webhook_events in the config; GET /projects/{id}/webhooks (the secret, the events, recent deliveries), POST …/webhooks/test (a run.finished with test: true), POST …/webhooks/secret (rotate, admin).
Job webhooks — the webhook on POST /crawl and webhook_url on a batch /scrape, with their own secret returned once. Events crawl.started, crawl.page (gathered into one message per 50 pages, or every 5 s), crawl.completed, crawl.failed, scrape.completed, scrape.failed, map.completed, batch.finished; metadata comes back on every message.
Project events — run.finished (status, pages, counts, coverage, stop), page.changed (every page added / modified / removed with words added and removed, the section, whether it was volatile; a modified page only when its edit reaches notify_min_words), page.blocked, field.canonical (canonical, robots and title changes), site.failing / site.recovered, page.orphaned / page.linked, sitemap.added / removed / updated / file_added / file_removed.
The digest — with change_digest: {connection_id, focus} in the config (an llm connection of the workspace; focus is what the reader cares about, e.g. "pricing and plan limits"), a compared run with something in it asks the model for a paragraph on what changed and why it matters, up to five points and a significant flag. It rides on run.finished as data.digest ({summary, points[], significant}, or null), in the run email, in the chat message, and on the change set (GET /projects/{id}/changes → digest). Tokens are charged like enrichment, through your own key.
Chat channels — chat_url in the config is a Slack, Discord, Microsoft Teams (Workflows), Mattermost or Google Chat incoming-webhook URL. When a run finishes with something to say — the same rule as the run email: never for a run that compared clean — the record is posted there in that service's own shape (Block Kit, an embed, an Adaptive Card, or {text}), unsigned, through the same delivery queue and retries as a webhook; site.failing / site.recovered post too. The URL is a secret to the channel, so the API returns it masked (https://hooks.slack.com/services/••••), and a masked value sent back in a PATCH keeps the one stored; send "" to remove it. GET /webhooks lists the channel with its health beside the project's webhook (kind: "chat", service).
The message — POST, JSON {id, event, createdAt, data}, headers X-MeshArc-Event, X-MeshArc-Delivery, X-MeshArc-Signature: sha256=<HMAC-SHA256 of the raw body with the secret>, User-Agent: MeshArc-Webhooks/1.0. Answer 2xx within 10 s.
Retries — no answer, a 5xx, a 408 or a 429 is retried six times over about a day (1 m, 5 m, 30 m, 2 h, 6 h, 12 h), then the delivery is paused. Any other 4xx is your endpoint's answer — a wrong URL, a rejected signature — and the delivery is failed at once. GET /webhooks/deliveries?runId=|projectId=&limit= lists what was sent, attempt by attempt (status, attempts, lastStatus, lastError, nextAt, the log); POST /webhooks/deliveries/{id}/redeliver sends one again now with its attempts kept. GET /webhooks gives every project's webhook with its health.
The workspace#
| Route | What |
|---|---|
GET /me · PATCH /me {max_rpm} (admin) | The workspace, its plan and limits, credits used and remaining, what this key may do; the rate ceiling |
GET /me/usage · GET /me/monitor | Pages per day and this month's pages, runs, blocked share, pages by engine; what is queued and running, lane depths |
GET /me/keys · POST /me/keys {name, scopes[], projects[], expires_in_days, rpm} (admin) · DELETE /me/keys/{id} | Keys; the secret is in the create response only |
GET /me/billing · POST /me/billing/subscribe · /change · /confirm · /card · /cancel · /keep · /credits/buy · PATCH /me/billing/cap · GET /me/billing/invoices · /receipts · /credit-notes/{id} | The plan, the meter, the ledger, the price list; subscribing, changing plan and card through Razorpay Checkout (confirm verifies Checkout's answer), cancelling and undoing it, buying credits, the overage cap; invoices, receipts and credit notes |
GET /me/proxies | The pool with per-exit parity and success |
GET /meta | Verdict meanings, engine costs, the config defaults, the ladder |
/auth/* | POST /auth/signup, /login, /logout, GET /session, POST /switch, /verify/request, /verify, /forgot, /reset, PATCH /profile, POST /password, GET/DELETE /sessions, GET /members, POST /invites, DELETE /invites/{id}, GET /invites/{token}, POST /invites/accept, PATCH/DELETE /members/{user}, PATCH /workspace |
What a page row says#
| Field | Values |
|---|---|
verdict | ok, thin (short but a page; kept), blocked (refused, or a 404 / soft 404 — see errorCode), skipped (under the word floor or outside the language list; the row stays, out of the change record), robots, cached |
errorCode | OK, SHELL, BLOCKED, NOT_FOUND, RATE_LIMITED, CAPTCHA, LOGIN_REQUIRED, GEO_RESTRICTED, TIER_LIMIT (a higher tier than allowed would have read it), BUDGET_EXHAUSTED (the credits the workspace had left could not buy the engine that would have read it), TIMEOUT, NETWORK. A short page is OK with a warning |
shape | listing, table, form, or empty — what a short page's markup says it is |
warnings | short |
signals | Why the judge decided as it did: nav, 30 links, json-ld, open graph, main, no script, app shell, listing, render confirmed, soft 404, host fingerprint, soft404? |
method, tier, climbedTo | The engine that answered and its tier; the highest engine a refused page was tried on |
credits, billedAs | What the page cost; the rung it is priced at when not the one that fetched it |
widget, solved, exitBytes | A captcha widget on the page and what was done; bytes through the residential exit |
foundVia, firstReferer, inSitemap, inLinks, orphan, unlisted | How the crawl reached it and what the link graph says |