Configuration: every setting
A project's config, a one-shot POST /crawl, a POST /scrape (its config and scrapeOptions) and the playground all take the same 65 settings with the same defaults. The Configure screen and the playground rail show exactly these, so a request that names only what differs runs what the screen showed. GET /meta returns the full table.
Keys are given by their API name; the app's label follows in parentheses. Default is what you get when the key is absent.
How a setting reaches the engine#
| Where | How to set it |
|---|---|
| A project | PATCH /projects/{id} with {"config": {...}}, or the Configure screen |
| A one-shot crawl | POST /crawl with config: {...}, or the short names limit, maxDepth, includePaths, excludePaths, useSitemap, sitemapMode, crawlMode, allowSubdomains, maxAge, maxTier, delay, concurrency, respectRobots |
| A scrape | POST /scrape with config: {...}, or scrapeOptions: {formats, actions, actionsOn, actionsPaths, onlyMainContent, includeTags, excludeTags, waitFor, waitMs, viewport, maxTier, jsonSchema, parseDocuments, minWords, languages}, plus maxCredits and formats at the top level |
| The playground | The rail; the request sends what differs from the defaults |
Unknown keys are ignored, not refused, so a newer SDK never fails on an option the API could do without. A value out of range is refused with 400 and the range.
Output formats#
| Key | Meaning | Default |
|---|---|---|
formats | Which bodies a page keeps: markdown (always), text, rawHtml (the page as served, for every page — 5–20× the storage), cleanHtml (after boilerplate, selectors and images, before markdown, without scripts and styles — selectors still resolve), links (every link with anchor text, resolved, typed internal / external / anchor), raw (status, headers, redirect chain and served body), screenshot (a JPEG of the viewport per page, ≤600 kB; forces a render, +1 credit), json (structured fields) | ["markdown", "links"] |
json_schema (JSON schema) | The fields the json format fills from what the page declares — head and Open Graph meta, JSON-LD, microdata, <time>, byline markup: [{name, type, required}], types string / number / date / boolean / url / array. Names match loosely (published_at, datePublished and date are one fact). Unresolved is reported, never guessed | [] |
rows (Read the page as records) | {source, path, selector, columns, key, next, filters}: a page that is a list — a board of roles, a table of recalls, a feed of posts — read as records instead of as one thing. source is json (a dotted path into the body, e.g. data.children[].data), table (the page's markdown tables, headings matched loosely) or list (a CSS selector matching one record). columns are typed as a JSON schema is; key is the column that names a record, or several when no single one does. next: {path, param, max_pages} follows a cursor the body hands out, for a feed nothing links to. filters: {since_days, must_contain, drop} keep records out and are counted. Rows are compared with the previous run's by key: added, changed (one event per column, with both values) and gone — gone withheld when the run reached under 90% of the site | null (off) |
llm_extract (Fill missing fields with a model) | {connection_id, instructions, only_missing, max_pages, max_chars, max_output_tokens}: a Models connection fills the fields the page did not declare, from the markdown, cached by content hash; 1 credit per 1,000 tokens as the pipeline fee, the model bill is yours. only_missing true asks only for unresolved fields; max_pages caps calls per run (1–2,000). max_chars (0 = as much of the page as the model's context holds) and max_output_tokens (0 = a budget that grows with the schema, an array field being a table) are ceilings for a workspace that would rather spend fewer credits on a long page: past max_chars the blocks naming your schema's fields are sent instead of the page's opening, and a reply that runs out of room is cut at its last complete value rather than lost | null (off) |
heading_style | atx (# prefixes) or setext (underlines) | atx |
link_style | inline ([text](url)), reference (footnotes), strip (text only — smallest) | inline |
Content filters#
| Key | Meaning | Default |
|---|---|---|
only_main_content | Drop nav, header, footer and sidebars — the chrome that repeats on every page — so a menu edit is not "the whole site changed" | true |
strip_boilerplate | Cookie banners, newsletter prompts, "was this helpful", share bars, back-to-top — by vendor markup and by shape | true |
strip_base64 | Inline data-URI images out of the markdown (their alt text stays) | true |
include_tags | CSS selectors to keep exclusively; everything outside them is dropped | [] |
exclude_tags | CSS selectors to drop; exclude wins over include | [] |
preserve_tags | Selectors that survive every kind of pruning — a price table in a sidebar the chrome drop would take | [] |
min_words (Minimum word count) | Off at 0. Set a floor and pages under it are recorded as skipped — stored, out of the change record. A page whose markup has a shape (a listing, a table, a form) is never under the floor, and the home page is never skipped. A short page is marked warnings: ["short"] on its row either way (0–5,000) | 0 |
languages | Two-letter codes; a page outside the list is skipped. Detected from the html tag, headers or text; an undetectable one is never skipped | [] |
Crawl scope#
| Key | Meaning | Default |
|---|---|---|
max_pages (Page budget) | The most pages one run fetches; the plan caps it (free 500 · Starter 5,000 · Growth 50,000 · Scale 200,000). A run that stops at its budget with URLs queued reaches only part of the site, and the diff withholds removals | 100 |
max_depth | Link hops from the seed; 0 is the seed alone (0–10) | 3 |
breadth_per_level | Cap on pages fetched at each depth; 0 is none. A cap silently leaves the rest of a level unread, which the link graph reports as reduced coverage | 0 |
respect_robots | Obey Allow / Disallow (for mesharc, else *) and Crawl-delay; disallowed URLs are recorded as robots, never fetched | true |
allow_subdomains | Follow links to other hosts under the same registrable domain | false |
include_paths / exclude_paths | Globs on the URL path (/docs/*); excludes win; empty include means everything | [] |
sitemap_mode (Sitemap source) | auto (robots.txt, the well-known paths, the seed page's link tags; indexes walked), listed (only sitemaps), off | auto |
sitemaps | The sitemap URLs a listed project reads (up to 20, same site) | [] |
sitemap_include / sitemap_exclude | Globs on a child file's URL or section label (*press*) that choose which files a crawl keeps | [] |
crawl_mode (Crawl from) | sitemap_first (the selected sitemap URLs at depth 1, then links — the only mode that can tell an orphan from a page never reached), sitemap_only (exactly the selected URLs, nothing followed), links (seed and links; the sitemap still informs the link graph and the sitemap diff). A project that turns its sitemap off without naming a mode falls back to links | sitemap_first |
sitemap_incremental | A scheduled run fetches only what the sitemap says moved (added, or lastmod newer than the last real fetch) and carries the rest over from the last full run at no fetch. A full run is forced every sitemap_full_every runs, when fewer than 90 % of entries carry lastmod, or when a carried page is over 30 days old | false |
sitemap_full_every | Full run every N runs of an incremental project (1–365) | 10 |
crawl_delay_ms | The pause between requests to one host; a robots.txt crawl-delay raises it, a rising refusal rate slows it, nothing lowers it (0–60,000) | 1000 |
concurrency | Parallel fetches against the host on the http rungs once the first page has settled which rung works; the delay becomes the gap between rounds. A render stays one page at a time; a robots crawl-delay forces 1. 2–4 is safe on sites that do not rate-limit (1–8) | 1 |
url_list / feeds / patterns | Sources beyond the seed, set on the Sources screen (5,000 URLs; 20 feeds; 20 patterns, 5,000 expansions) | [] |
Images#
| Key | Meaning | Default |
|---|---|---|
image_mode | alt writes the alt text in place of the image; link keeps  with the URL made absolute | alt |
image_exclude | Globs on the image path (*/icons/*, *sprite*) | [] |
image_min_width / image_min_height | Drop images the markup declares smaller (tracking pixels, icons); only attributes are read, nothing is fetched | 0 / 0 |
image_max_kb | Drop inline data-URI images above this | 0 |
alt_fallback | When an image has no alt: filename, caption (the figcaption), nearby (surrounding text), none | filename |
Fetch behaviour#
| Key | Meaning | Default |
|---|---|---|
max_tier (Max tier) | How far a refused page may climb; a page starts where the site is known to answer (the cheapest engine, on a site not read before) and climbs only when refused: auto (as high as the plan allows), http (1–2 credits; a walled page is recorded TIER_LIMIT), browser (a render when needed, 4), stealth (a residential exit 16, a solved widget 22, a rented fetch 40). The plan caps it (free and Starter: browser), and on any tier a climb never goes past the credits left | auto for a project made with New project, else browser |
render_js (Render JavaScript) | auto: render only when the http rung comes back an app shell — and a render that only confirms a short page is billed as the http rung. always: every page, 4 credits. never: whatever the server sent (sets max_tier to http) | auto |
googlebot (Ask as Googlebot) | Add Googlebot's identity to the engines this project may use. A claim about who is asking; sites that verify by reverse DNS are not persuaded; never reached for unless on | false |
user_agent | Where the ladder starts: default (a rotating crawler identity), chrome (a real Chrome with its TLS fingerprint), googlebot (turns the Googlebot setting on) | default |
solve_captcha | Have MeshArc's token solver answer a Turnstile / reCAPTCHA widget on a form the steps submit, before the steps run. Stealth tier; 22 credits when a token is bought. A widget on a page you only read needs no solving | false |
max_credits_per_page (Max credits per page) | The most one page may cost; engines dearer than this are never tried, and a page only they could read is TIER_LIMIT at 0. 0: the tier is the ceiling. maxCredits on /scrape (0–100) | 0 |
wait_for_selector | When rendering, wait up to 20 s for this node before reading; read what is there if it never appears | "" |
wait_ms | A flat delay after load on top of the settle; paid on every rendered page (0–60,000) | 0 |
actions (Browser steps) | Steps after load and before reading: click (with repeat: "until_gone" and max for a Load-more button), clickall, each (over every element a selector matches, with nested steps), type, select, press, wait (ms or selector), scroll (bottom ×N or to a selector). Up to 40 leaf steps; a step that fails is noted on the page and the page is read as it stands; a page that takes steps is always rendered | [] |
actions_on (Run steps on) | start (the seed plus the URL list and pattern URLs — a listing whose Load-more reveals the pages to crawl; a consent click the browser remembers), paths (pages matching actions_paths), all (every page — forces a render for the whole run, 4 credits a page) | start |
actions_paths | Path globs for paths | [] |
fetch_timeout_ms | Abort a fetch after this; recorded as a timeout (1,000–120,000) | 30000 |
fetch_retries | Extra tries on a dropped connection or 500 / 502 / 503 / 504 / 400 / 408 / 409 / 430. A 429 is handled by the pacing; a 403 climbs the ladder (0–5) | 2 |
viewport | desktop or mobile for renders | desktop |
use_proxy (Proxy pool for http fetches) | Route the plain http fetches through MeshArc's datacenter proxy pool (each exit is checked against a direct fetch before it is trusted). No change to the price. The residential exit for renders is max_tier: stealth, a separate thing | false |
geo | auto or a two-letter country: requests from that country through the residential exit; a requirement the ladder honours (stealth) | auto |
cache_max_age_hours | Reuse a page any project of the workspace fetched this recently, at 0 credits, instead of fetching (maxAge on the verbs; 0–720) | 0 |
force_cache | Whether a scheduled run may answer from the cache. Off, a scheduled run always looks | false |
Advanced#
| Key | Meaning | Default |
|---|---|---|
Schedule (schedule on the project, not in config) | manual, hourly, daily (03:00 UTC), weekly (Mondays 03:00 UTC). Change detection needs two runs. A 100-page project is about 400 credits a month weekly, 3,000 daily | The app creates projects weekly; POST /projects and one-shot crawls: manual |
Retention (retention on the project) | How long raw HTML is kept: 30d, 90d, 1y, forever. Markdown and hashes are kept regardless | 90d (one-shot crawls 30d) |
webhook_url | Where the project's messages go (http(s)) | "" |
webhook_events | Which of page.changed, run.finished, page.blocked, field.canonical, site.failing, site.recovered, page.orphaned, page.linked, sitemap.added / removed / updated / file_added / file_removed, and the job events | ["page.changed", "run.finished", "field.canonical"] |
notify_min_words | A modified page under this many changed words stays out of page.changed; added and removed pages always count | 5 |
chat_url | A Slack, Discord, Microsoft Teams (Workflows), Mattermost or Google Chat incoming-webhook URL (https). The run record posts there when a run finds something — never when it compared clean — in the service's own shape, plus site.failing / site.recovered. Returned masked; a masked value sent back keeps the one stored, "" removes it. An edit here does not bump the config version | "" |
change_digest | {connection_id, focus} — an llm connection of the workspace and a line on what the reader cares about. After each compared run with something in it, the model writes a paragraph on what changed and why it matters, up to five points and a significant flag, from the changed paragraphs, the added pages' openings, the removed paths and the field diffs. Carried by the run email, the chat message, run.finished and the change set. Charged like enrichment, through your own key | null |
parse_documents | Follow links to PDFs and record each one's text as a document — one markdown block per page, a word count and a hash, compared like a page (2 credits; a rendered one 4; up to 400 pages / 25 MB). Word, PowerPoint and Excel files are recognised and recorded as seen, not read (no reader ships yet). Off, documents are never fetched — one met by content type is recorded as seen and not charged | false |
Legacy aliases#
use_sitemap (true unless sitemap_mode is off), allow_browser (false when render_js is never) and keep_pages are kept in step automatically for older clients.
Things that follow from a setting#
screenshotinformats, or steps withactions_on: all, forcerender_js: always.render_js: neverforcesmax_tier: http.user_agent: googlebotsetsgooglebot: true.crawl_modeother thanlinksneeds a sitemap: withsitemap_mode: offan explicit mode is refused, an implicit one becomeslinks.sitemap_mode: listedneeds at least one URL insitemaps; a sitemap or a source on another host is refused.- A project's
max_pagesis capped by the plan'smax_pages_per_run;max_tierby the plan's tier, on a scrape, the playground and a map as well.autois the plan's tier.