Docs · Configuration

Configuration: every setting

A project's config, a one-shot POST /crawl, a POST /scrape (its config and scrapeOptions) and the playground all take the same 65 settings with the same defaults. The Configure screen and the playground rail show exactly these, so a request that names only what differs runs what the screen showed. GET /meta returns the full table.

Keys are given by their API name; the app's label follows in parentheses. Default is what you get when the key is absent.

How a setting reaches the engine#

WhereHow to set it
A projectPATCH /projects/{id} with {"config": {...}}, or the Configure screen
A one-shot crawlPOST /crawl with config: {...}, or the short names limit, maxDepth, includePaths, excludePaths, useSitemap, sitemapMode, crawlMode, allowSubdomains, maxAge, maxTier, delay, concurrency, respectRobots
A scrapePOST /scrape with config: {...}, or scrapeOptions: {formats, actions, actionsOn, actionsPaths, onlyMainContent, includeTags, excludeTags, waitFor, waitMs, viewport, maxTier, jsonSchema, parseDocuments, minWords, languages}, plus maxCredits and formats at the top level
The playgroundThe rail; the request sends what differs from the defaults

Unknown keys are ignored, not refused, so a newer SDK never fails on an option the API could do without. A value out of range is refused with 400 and the range.

Output formats#

KeyMeaningDefault
formatsWhich bodies a page keeps: markdown (always), text, rawHtml (the page as served, for every page — 5–20× the storage), cleanHtml (after boilerplate, selectors and images, before markdown, without scripts and styles — selectors still resolve), links (every link with anchor text, resolved, typed internal / external / anchor), raw (status, headers, redirect chain and served body), screenshot (a JPEG of the viewport per page, ≤600 kB; forces a render, +1 credit), json (structured fields)["markdown", "links"]
json_schema (JSON schema)The fields the json format fills from what the page declares — head and Open Graph meta, JSON-LD, microdata, <time>, byline markup: [{name, type, required}], types string / number / date / boolean / url / array. Names match loosely (published_at, datePublished and date are one fact). Unresolved is reported, never guessed[]
rows (Read the page as records){source, path, selector, columns, key, next, filters}: a page that is a list — a board of roles, a table of recalls, a feed of posts — read as records instead of as one thing. source is json (a dotted path into the body, e.g. data.children[].data), table (the page's markdown tables, headings matched loosely) or list (a CSS selector matching one record). columns are typed as a JSON schema is; key is the column that names a record, or several when no single one does. next: {path, param, max_pages} follows a cursor the body hands out, for a feed nothing links to. filters: {since_days, must_contain, drop} keep records out and are counted. Rows are compared with the previous run's by key: added, changed (one event per column, with both values) and gone — gone withheld when the run reached under 90% of the sitenull (off)
llm_extract (Fill missing fields with a model){connection_id, instructions, only_missing, max_pages, max_chars, max_output_tokens}: a Models connection fills the fields the page did not declare, from the markdown, cached by content hash; 1 credit per 1,000 tokens as the pipeline fee, the model bill is yours. only_missing true asks only for unresolved fields; max_pages caps calls per run (1–2,000). max_chars (0 = as much of the page as the model's context holds) and max_output_tokens (0 = a budget that grows with the schema, an array field being a table) are ceilings for a workspace that would rather spend fewer credits on a long page: past max_chars the blocks naming your schema's fields are sent instead of the page's opening, and a reply that runs out of room is cut at its last complete value rather than lostnull (off)
heading_styleatx (# prefixes) or setext (underlines)atx
link_styleinline ([text](url)), reference (footnotes), strip (text only — smallest)inline

Content filters#

KeyMeaningDefault
only_main_contentDrop nav, header, footer and sidebars — the chrome that repeats on every page — so a menu edit is not "the whole site changed"true
strip_boilerplateCookie banners, newsletter prompts, "was this helpful", share bars, back-to-top — by vendor markup and by shapetrue
strip_base64Inline data-URI images out of the markdown (their alt text stays)true
include_tagsCSS selectors to keep exclusively; everything outside them is dropped[]
exclude_tagsCSS selectors to drop; exclude wins over include[]
preserve_tagsSelectors that survive every kind of pruning — a price table in a sidebar the chrome drop would take[]
min_words (Minimum word count)Off at 0. Set a floor and pages under it are recorded as skipped — stored, out of the change record. A page whose markup has a shape (a listing, a table, a form) is never under the floor, and the home page is never skipped. A short page is marked warnings: ["short"] on its row either way (0–5,000)0
languagesTwo-letter codes; a page outside the list is skipped. Detected from the html tag, headers or text; an undetectable one is never skipped[]

Crawl scope#

KeyMeaningDefault
max_pages (Page budget)The most pages one run fetches; the plan caps it (free 500 · Starter 5,000 · Growth 50,000 · Scale 200,000). A run that stops at its budget with URLs queued reaches only part of the site, and the diff withholds removals100
max_depthLink hops from the seed; 0 is the seed alone (0–10)3
breadth_per_levelCap on pages fetched at each depth; 0 is none. A cap silently leaves the rest of a level unread, which the link graph reports as reduced coverage0
respect_robotsObey Allow / Disallow (for mesharc, else *) and Crawl-delay; disallowed URLs are recorded as robots, never fetchedtrue
allow_subdomainsFollow links to other hosts under the same registrable domainfalse
include_paths / exclude_pathsGlobs on the URL path (/docs/*); excludes win; empty include means everything[]
sitemap_mode (Sitemap source)auto (robots.txt, the well-known paths, the seed page's link tags; indexes walked), listed (only sitemaps), offauto
sitemapsThe sitemap URLs a listed project reads (up to 20, same site)[]
sitemap_include / sitemap_excludeGlobs on a child file's URL or section label (*press*) that choose which files a crawl keeps[]
crawl_mode (Crawl from)sitemap_first (the selected sitemap URLs at depth 1, then links — the only mode that can tell an orphan from a page never reached), sitemap_only (exactly the selected URLs, nothing followed), links (seed and links; the sitemap still informs the link graph and the sitemap diff). A project that turns its sitemap off without naming a mode falls back to linkssitemap_first
sitemap_incrementalA scheduled run fetches only what the sitemap says moved (added, or lastmod newer than the last real fetch) and carries the rest over from the last full run at no fetch. A full run is forced every sitemap_full_every runs, when fewer than 90 % of entries carry lastmod, or when a carried page is over 30 days oldfalse
sitemap_full_everyFull run every N runs of an incremental project (1–365)10
crawl_delay_msThe pause between requests to one host; a robots.txt crawl-delay raises it, a rising refusal rate slows it, nothing lowers it (0–60,000)1000
concurrencyParallel fetches against the host on the http rungs once the first page has settled which rung works; the delay becomes the gap between rounds. A render stays one page at a time; a robots crawl-delay forces 1. 2–4 is safe on sites that do not rate-limit (1–8)1
url_list / feeds / patternsSources beyond the seed, set on the Sources screen (5,000 URLs; 20 feeds; 20 patterns, 5,000 expansions)[]

Images#

KeyMeaningDefault
image_modealt writes the alt text in place of the image; link keeps ![alt](url) with the URL made absolutealt
image_excludeGlobs on the image path (*/icons/*, *sprite*)[]
image_min_width / image_min_heightDrop images the markup declares smaller (tracking pixels, icons); only attributes are read, nothing is fetched0 / 0
image_max_kbDrop inline data-URI images above this0
alt_fallbackWhen an image has no alt: filename, caption (the figcaption), nearby (surrounding text), nonefilename

Fetch behaviour#

KeyMeaningDefault
max_tier (Max tier)How far a refused page may climb; a page starts where the site is known to answer (the cheapest engine, on a site not read before) and climbs only when refused: auto (as high as the plan allows), http (1–2 credits; a walled page is recorded TIER_LIMIT), browser (a render when needed, 4), stealth (a residential exit 16, a solved widget 22, a rented fetch 40). The plan caps it (free and Starter: browser), and on any tier a climb never goes past the credits leftauto for a project made with New project, else browser
render_js (Render JavaScript)auto: render only when the http rung comes back an app shell — and a render that only confirms a short page is billed as the http rung. always: every page, 4 credits. never: whatever the server sent (sets max_tier to http)auto
googlebot (Ask as Googlebot)Add Googlebot's identity to the engines this project may use. A claim about who is asking; sites that verify by reverse DNS are not persuaded; never reached for unless onfalse
user_agentWhere the ladder starts: default (a rotating crawler identity), chrome (a real Chrome with its TLS fingerprint), googlebot (turns the Googlebot setting on)default
solve_captchaHave MeshArc's token solver answer a Turnstile / reCAPTCHA widget on a form the steps submit, before the steps run. Stealth tier; 22 credits when a token is bought. A widget on a page you only read needs no solvingfalse
max_credits_per_page (Max credits per page)The most one page may cost; engines dearer than this are never tried, and a page only they could read is TIER_LIMIT at 0. 0: the tier is the ceiling. maxCredits on /scrape (0–100)0
wait_for_selectorWhen rendering, wait up to 20 s for this node before reading; read what is there if it never appears""
wait_msA flat delay after load on top of the settle; paid on every rendered page (0–60,000)0
actions (Browser steps)Steps after load and before reading: click (with repeat: "until_gone" and max for a Load-more button), clickall, each (over every element a selector matches, with nested steps), type, select, press, wait (ms or selector), scroll (bottom ×N or to a selector). Up to 40 leaf steps; a step that fails is noted on the page and the page is read as it stands; a page that takes steps is always rendered[]
actions_on (Run steps on)start (the seed plus the URL list and pattern URLs — a listing whose Load-more reveals the pages to crawl; a consent click the browser remembers), paths (pages matching actions_paths), all (every page — forces a render for the whole run, 4 credits a page)start
actions_pathsPath globs for paths[]
fetch_timeout_msAbort a fetch after this; recorded as a timeout (1,000–120,000)30000
fetch_retriesExtra tries on a dropped connection or 500 / 502 / 503 / 504 / 400 / 408 / 409 / 430. A 429 is handled by the pacing; a 403 climbs the ladder (0–5)2
viewportdesktop or mobile for rendersdesktop
use_proxy (Proxy pool for http fetches)Route the plain http fetches through MeshArc's datacenter proxy pool (each exit is checked against a direct fetch before it is trusted). No change to the price. The residential exit for renders is max_tier: stealth, a separate thingfalse
geoauto or a two-letter country: requests from that country through the residential exit; a requirement the ladder honours (stealth)auto
cache_max_age_hoursReuse a page any project of the workspace fetched this recently, at 0 credits, instead of fetching (maxAge on the verbs; 0–720)0
force_cacheWhether a scheduled run may answer from the cache. Off, a scheduled run always looksfalse

Advanced#

KeyMeaningDefault
Schedule (schedule on the project, not in config)manual, hourly, daily (03:00 UTC), weekly (Mondays 03:00 UTC). Change detection needs two runs. A 100-page project is about 400 credits a month weekly, 3,000 dailyThe app creates projects weekly; POST /projects and one-shot crawls: manual
Retention (retention on the project)How long raw HTML is kept: 30d, 90d, 1y, forever. Markdown and hashes are kept regardless90d (one-shot crawls 30d)
webhook_urlWhere the project's messages go (http(s))""
webhook_eventsWhich of page.changed, run.finished, page.blocked, field.canonical, site.failing, site.recovered, page.orphaned, page.linked, sitemap.added / removed / updated / file_added / file_removed, and the job events["page.changed", "run.finished", "field.canonical"]
notify_min_wordsA modified page under this many changed words stays out of page.changed; added and removed pages always count5
chat_urlA Slack, Discord, Microsoft Teams (Workflows), Mattermost or Google Chat incoming-webhook URL (https). The run record posts there when a run finds something — never when it compared clean — in the service's own shape, plus site.failing / site.recovered. Returned masked; a masked value sent back keeps the one stored, "" removes it. An edit here does not bump the config version""
change_digest{connection_id, focus} — an llm connection of the workspace and a line on what the reader cares about. After each compared run with something in it, the model writes a paragraph on what changed and why it matters, up to five points and a significant flag, from the changed paragraphs, the added pages' openings, the removed paths and the field diffs. Carried by the run email, the chat message, run.finished and the change set. Charged like enrichment, through your own keynull
parse_documentsFollow links to PDFs and record each one's text as a document — one markdown block per page, a word count and a hash, compared like a page (2 credits; a rendered one 4; up to 400 pages / 25 MB). Word, PowerPoint and Excel files are recognised and recorded as seen, not read (no reader ships yet). Off, documents are never fetched — one met by content type is recorded as seen and not chargedfalse

Legacy aliases#

use_sitemap (true unless sitemap_mode is off), allow_browser (false when render_js is never) and keep_pages are kept in step automatically for older clients.

Things that follow from a setting#

  • screenshot in formats, or steps with actions_on: all, force render_js: always.
  • render_js: never forces max_tier: http.
  • user_agent: googlebot sets googlebot: true.
  • crawl_mode other than links needs a sitemap: with sitemap_mode: off an explicit mode is refused, an implicit one becomes links.
  • sitemap_mode: listed needs at least one URL in sitemaps; a sitemap or a source on another host is refused.
  • A project's max_pages is capped by the plan's max_pages_per_run; max_tier by the plan's tier, on a scrape, the playground and a map as well. auto is the plan's tier.