Using the app
The app is at https://mesharc.dev. Everything it shows comes from the API; nothing is computed in the browser that the API does not also return.
Accounts and workspaces#
- Sign up (
/signup) with a work email and a password. The first account of a workspace is its owner. A six-digit verification code arrives by mail (30-minute lifetime); until it is entered you can look around, but creating a project, starting a run, issuing a key or inviting someone answers403 email_unverified. The shell shows a banner with a link to enter the code. - Sign in (
/login). Sessions last 30 days and end after 14 idle days; "Keep me signed in" is the default.Profile → Sessionslists them and ends any. An account with two-factor authentication on is asked for a code after the password. - Forgot password (
/forgot) mails a reset link;Profilechanges it while signed in. - Workspaces. An account can belong to several; the sidebar switches.
Settings → Workspacerenames it. - Members and roles (
Settings → Members):owner>admin>member>viewer. A viewer reads everything. A member creates and configures projects, starts runs, scrapes, crawls and maps, reviews changes, tests webhooks and syncs destinations. An admin also issues and revokes API keys, sets the rate ceiling, deletes projects, rotates webhook secrets, manages connections and destinations, invites and removes members, sets roles, renames the workspace and opens the billing portal. The owner is the admin who cannot be removed. Invites go by mail with a link that lives seven days; an invite names the role. - Profile: name, time zone, two-factor authentication, whether you receive run emails (a run record when a project finishes) and the weekly digest (Monday from 08:00 in your time zone: every project's week in one message).
Two-factor authentication#
A second factor is a six-digit code from an authenticator app — Google Authenticator, 1Password, Authy, Microsoft Authenticator, anything that does time-based codes — asked for after the password every time you sign in.
- Turning it on (
Profile → Two-factor authentication → Turn on): enter your password, scan the QR code with the app (or type the key it shows), then enter the code the app displays. Every other session on the account is signed out, since none of them showed a factor. You are then handed ten recovery codes; keep them where you keep passwords, because they are shown only once. - Recovery codes sign you in when the phone is lost. Each works once; the sign-in screen has a "Lost the phone?" link that takes one instead of an app code. After a recovery sign-in the Profile shows how many remain, and
New recovery codesmakes a fresh set of ten (against a current app code) — the old set stops working. - Confirming it's you. With two-factor on, a few actions ask for the code again if the session last showed it more than ten minutes ago: changing the password, turning two-factor off, making new recovery codes, creating an API key, changing a member's role or removing them, and changing the workspace's two-factor rule. A small dialog asks for the code and the action then goes through.
- Turning it off takes your password and a current code. A mail goes to the account's address whenever two-factor is turned on or off, when a recovery code is used, and when new codes are made.
- Resetting a forgotten password on an account with two-factor also asks for an app code or a recovery code: an inbox alone is not enough to take over the account.
- Requiring it for a workspace (
Settings → Workspace → Require two-factor authentication): an owner — who has it on for their own account — can make it a condition of acting. A member without it can still sign in and look around, but creating a project, starting a run, issuing a key or changing anything answers403 mfa_requireduntil it is set up; the shell shows a banner with the link, and the members list marks who has it on. - API keys are unaffected by any of this: a key was made by someone who had already passed the gate, and a program cannot read a phone.
API keys#
A key is shown once when it is created. It carries scopes (read < write < admin, matching the member roles), optionally a set of projects it may see (any other project — or another workspace's — answers 404, never 403), an expiry, and its own rate limit (rpm), reported on every response as X-RateLimit-Limit / -Remaining / -Reset. Keys start with mesharc_. Revoking is immediate.
The workspace's own ceiling is max_rpm (Settings → Rate ceiling, PATCH /me); concurrency scales with it so the rate to a host never exceeds it.
The playground#
One URL through the ladder, with every setting on the left as a request field. Three modes:
| Mode | What happens | API |
|---|---|---|
| Scrape | One page, synchronously; the tabs (Markdown, Text, Clean HTML, Fields, Screenshot, Links, Raw) follow the formats you chose; the evidence panel says which engine answered, the verdict, the credits | POST /scrape |
| Crawl | A one-shot crawl of the site with the page budget and depth from the rail; poll its pages; Keep promotes it to a project | POST /crawl, POST /crawl/{id}/keep |
| Map | The site's URLs from its sitemap tree | POST /map |
The rail's settings are the same 65 controls as a project's Configure screen — see Configuration for each one; the request sends only what differs from the defaults, and the code snippet (curl, Python, Node) is a copy of the request that just ran.
Projects#
The workspace table: every project with its pages, reach, blocked share, pending changes, schedule and last run. Select several and Run. New project takes a seed (a domain or a URL), an optional name and a page budget; a new project runs weekly and the form says what that costs ("about 400 credits a month at 1 a page"). New from a list creates many from seeds with a template project's config.
Each project has these screens.
Overview#
Stat cards (pages tracked, pending review, last-run coverage, credits), "what changed in the last run" with field-change chips and the top changed pages, a health panel (healthy / failing / recovered from the last two runs), and recent runs. Run now starts a run.
Sources#
Where the project's URLs come from:
- The seed and its links.
- The sitemap tree — discovered from robots.txt, the well-known paths and the seed page's link tags, indexes walked to their children; or only the files you list (
sitemap_mode: listed); or off. Every file with its section label, URL count andlastmodrange; a diff against the last snapshot (files and URLs added, removed, updated). Pick the sections a crawl keeps (sitemap_include/sitemap_excludeglobs on the file URL or section) and see how many URLs the selection keeps before a run (Preview). - URL list — up to 5,000 exact URLs, fetched as given, never filtered by the path globs.
- Feeds — up to 20 RSS/Atom feeds, read each run; every entry is queued first.
- Patterns — templates expanded into URLs:
/products/{1..500},/docs/{intro,setup,faq}; capped at 5,000 expansions.
The scope preview shows pages per depth with the excluded and beyond-budget bands, and an estimate (~340 pages · ~1.4M tokens · ~4 min).
Configure#
The whole configuration as controls in six groups — Output formats, Content filters, Crawl scope, Images & media, Fetch behaviour, Advanced — each with a tip that says what the engine does and what it charges; a JSON pane that round-trips to the controls with line-level parse errors; Dry run (resolve sources and sample URLs without spending); a staged-change diff; Apply to… (copy the saved config onto other projects by group); Webhook deliveries (every message sent, attempt by attempt, with resend); the signing secret and a Send test; a chat channel (a Slack, Discord, Teams, Mattermost or Google Chat incoming-webhook URL — the run record posts there when a run finds something, never when it compared clean; the URL is shown masked once saved); the change digest (a model and a focus line); and the danger zone (delete the project). Configuration has every control with its meaning and its default.
Pages#
Every page of a run with its verdict, engine, tier, credits, words, status and how it was found (seed / sitemap / link / URL list / feed / pattern). Filter with the query language (status:403 depth:>2 flag:orphan). Two search modes scan the run on the server: content (words and "phrases" over the markdown) and selector (a CSS selector or XPath over the stored HTML). A selection can be re-crawled now (a scoped run), exported (JSONL / CSV) or excluded (added to exclude_paths). Removed lists pages that were in an earlier run and not this one.
Page detail#
Five tabs: the rendered markdown, the sandboxed HTML, head fields (title, description, canonical, robots, h1, Open Graph, Twitter card, JSON-LD), structured fields with what resolved each one, and versions across runs with the dual-hash diagnostic (content hash vs HTML hash: a page whose HTML changed but whose content did not is a template edit, not a content change). A screenshot when the project takes one.
Changes#
The change record of a run against the one before:
- Field changes — title, description, canonical, robots, h1, Open Graph, Twitter card, lang — with the before and after, ranked.
- Content diff per page — word-level, side by side or unified, a volatile-block filter (dates, counters, "3 comments"), a section rail.
- Markup comparison — a noise filter, the DOM tree, selector-breakage repointing ("
.pricematched 40 pages last run and 0 now; the same nodes now match.product-price"), ignore rules. - The coverage banner: a run that reached under 90 % of the site withholds removals, and the withheld URLs are listed under it.
- Reviewed marks per change.
- The digest, when the project has one (Configure → Advanced → Change digest, a model from Connectors → Models and a line on what you care about): a paragraph on what changed and why it matters, up to five points, and whether any of it is significant — written by the model from the changed paragraphs, the added pages' openings, the removed paths and the field diffs, on your own key. The same paragraph opens the run email and the chat message.
Rows#
For a project that reads its pages as records (Configure → Rows): every record its pages hold, what moved on the last run, and the records a run stopped finding. A record opens its own history — when it was first and last seen, and every column that ever moved, with the run that moved it. A board of roles, a table of recalls, a schedule of charges and a feed of posts are all lists of records, and watching them as pages only ever says "something changed". Exported as the rows and row-events datasets, and carried by the row.changed webhook.
Runs#
History with status, trigger (manual / scheduled / API / recrawl), pages, credits, stop reason and duration; the live log with time and level while a run is on (polled every 2 s); the engine breakdown; grouped errors; Cancel.
Destinations#
The project's destinations and the shared ones that include it, each with its target, mode, columns, shape and sync history ("synced 4 h ago · 84 written, 1,190 skipped"). Sync now writes the last run by hand. How a destination is set up, and how rows are shaped, is under Connectors and destinations.
Connectors#
Workspace-level connections: a provider plus the credentials to reach it, named, and tested before it is trusted. The catalogue (19 providers) with a connect wizard per provider, then per-connection overview, mapping and settings. Secrets are never returned by the API and are encrypted at rest. The Models kind (any OpenAI-compatible or Anthropic chat endpoint on your key) is what a project's LLM enrichment stage uses.
Shared destinations#
One target for many projects: one target that every project (scope all) or a chosen set (scope projects) writes after each of their runs, with site, project_id and project_name as columns so one table holds many sites. A shape (filter / map / drop) previewed against the last run before saving.
Exports#
Datasets streamed on request as JSONL or CSV: pages, markdown, changes, fields, sitemap — per run or per project, or restricted to a selection from the Pages screen. Two text files for a model: llms.txt (the usable pages as links with descriptions, by section — what an assistant reads to know what the site holds) and llms-full.txt (every usable page's markdown in one file, under its title and source URL — the corpus). Also every project's webhook and chat channel with its health from recent deliveries (healthy / failing / paused / idle), and every destination with its last sync.
Monitor#
What is queued and running for the workspace, each lane's queue depth and oldest wait, sparklines of pages per day.
Billing#
The plan and what it includes, this month's meter (credits used of the allowance; the free tier's is a lifetime allowance), the ledger (every charge: a scrape, a map, a crawl's rows per engine with a note when renders were billed at the rung below, model passes, free lines for cached pages), the price list, and the billing portal (invoices, the card on file, plan changes). Credits, plans and billing explains what a page costs and what each plan includes.
Settings#
Members and invites (with whether each has two-factor on, and the workspace's two-factor rule), API keys, the rate ceiling, operator limits (what the plan and the operator allow: projects, pages per run, concurrent jobs, tier, rpm), the ladder (/meta), the proxy pool with per-exit parity and success, and usage by day and by engine with the blocked share.