use cases

What teams build with a web scraper that remembers

Web data is only worth having if it is fresh, accurate and in the right place. Each case below starts from a problem people describe in their own words — a scraper that broke again, an index that lies, a page that vanished — and ends with the configuration that answers it and what comes back. Pick the outcome you need.
60%
of RAG projects that fail after a proof of concept fail on freshness, not retrieval.
10,000
billed failures a month at a 90% success rate on 100,000 requests — with providers that charge per attempt.
8,000+
US federal pages removed or substantially changed since January 2025.
0
credits here for a refused page, a 404, a cached page or a run that compared clean.
02what people say

Eight complaints, from the people who have them.

Read across scraping forums, tool reviews and RAG engineering write-ups this month. Each with its source, and with what this product has for it.

Blocked, then blocked again“It works for ten minutes, then a 403 or a Cloudflare page. Rotating user agents or IPs no longer gets meaningful results.” ScrapingBee, Browserless — scraping challenges 2026→The ladder climbs one rung per refusal; a site learned once starts on the cheap rung next time, for every workspace.
Scrapers rot“Maintenance is the worst part. Building the scraper is easy; keeping it running for months as sites change is the job.” BinaryBits — why scrapers keep breaking→Markdown instead of selectors, and a run that counts selector breaks when a field stops matching.
Paying for garbage“At 100,000 requests a month, a 90% success rate is 10,000 failed requests — and some providers bill every one.” String — 13 scraping APIs’ billing compared→A refused page, a 404, a cached page and a page the sitemap says is unchanged all cost nothing.
Page monitors are page-at-a-time, and noisy“Misses JavaScript content, false alerts, too many alerts. The advice is to start with five to eight monitors.” Adversa — Visualping alternatives→A project is a site or a section; a run that compared clean says nothing; noise is suppressed before the diff.
RAG indexes go stale“60% of projects that fail after a successful proof of concept fail because they cannot maintain data freshness at scale.” tianpan.co — the staleness problem→Content-hash diffs, only the delta pushed to the vector store, and a removed page loses its vectors.
Agents eat context“A scrape that returns 5,000 tokens consumes 5,000 tokens of context. For multi-page research, the context fills up.” dev.to — MCP web scraping for Claude and Cursor→Main-content markdown, structured extraction on the server, and a map of a site for one credit with no page fetched.
Pages disappear and nobody noticed“More than 8,000 US federal web pages removed or substantially modified since January 2025; trackers built by hand from archived sitemaps.” George Mason University — tracking removed government data→A sitemap snapshot per run, removals confirmed twice, and “withheld” for a page the site declined — not “removed”.
Hundreds of sites, on a schedule“The cheap way to watch many sites is the sitemap and its lastmod: pay per change, not per scan.” Apify — sitemap change monitor→Scheduled projects, incremental runs by lastmod, robots crawl-delay honoured, one rate bucket per domain.
03developer-tools companies · support and docs teams

Keep an AI assistant’s knowledge fresh.

The index is a snapshot and the docs keep moving. A deprecated flag keeps being recommended, a deleted page keeps being retrieved, and nobody sees it because a stale embedding scores as well as a fresh one. Full re-indexing is the fix teams reach for, and it is the one they cannot afford to run every night.

the configuration, as sent
{
  "seed": "https://docs.example.com/",
  "schedule": "daily",
  "config": {
    "include_paths": ["/docs/*"],
    "only_main_content": true,
    "sitemap_incremental": true
  }
}
// and a destination on the project:
{
  "provider": "qdrant", "target": "docs",
  "mode": "upsert", "key": "url",
  "columns": ["url", "title", "markdown", "content_hash", "change_class"]
}

Incremental runs read only what the sitemap says moved. The destination is incremental too: only added, modified and removed pages reach the index, chunked and embedded on your own key.

what comes back
  • Only the pages whose content hash moved are re-embedded; a run that compared clean writes nothing.
  • A removed page — gone from the site, confirmed twice — loses its chunks in the index. A page the site merely declined is “withheld”, and stays.
  • A change record you can audit: which pages, which paragraphs, when.
  • llms.txt and llms-full.txt exports, for the assistant that reads a corpus rather than an index.
proofQdrant, Weaviate and pgvector are exercised against live instances in the repository; Pinecone is written to its documentation and awaits an account.
change detectionincremental runsvector destinationsllms.txtown model key
04compliance · legal · public affairs · data journalists

Watch the regulator.

A guidance page changes on a Tuesday and you read about it in a newsletter on Friday. A consultation closes and the page that said so is gone. Page-monitoring tools watch one URL at a time, alert on a rotating banner, and cannot tell a page that was deleted from a page that was blocked for a crawler.

the configuration, as sent
{
  "seed": "https://www.gov.example/",
  "schedule": "daily",
  "config": {
    "crawl_mode": "sitemap_first",
    "include_paths": ["/guidance/*", "/news/*"],
    "sitemap_incremental": true,
    "notify_min_words": 20,
    "chat_url": "https://hooks.slack.com/services/…",
    "change_digest": { "connection_id": "…", "focus": "deadlines, forms, thresholds" }
  }
}

Two sections of a large site, from its sitemap; an incremental run touches only what lastmod says changed. The record goes to the channel the team already watches, with a paragraph on what it means.

what comes back
  • “added” for a page new in the sitemap, “modified” with the paragraphs that moved, “removed” only when two runs in a row agree the page is gone.
  • A field diff for title, canonical, robots and published date — the changes that decide whether a document is the current one.
  • A run that reached under 90% of the site withholds removals and says why, instead of reporting a WAF hiccup as a purge.
  • History and exports: the sitemap, the pages, the changes, per run, for as long as retention keeps them.
proof508,290 gov.uk URLs mapped from 29 sitemap files in one discovery.
sitemap sectionsremoved vs withheldchat channelchange digestexports
05product marketing · founders · product managers

Know when a competitor moves.

Pricing, feature and changelog pages change without an announcement. The tools that watch them alert on the navigation bar and go quiet on the number that mattered, and reading forty diffs to find the one price change is nobody’s Monday.

the configuration, as sent
{
  "seed": "https://competitor.example/",
  "schedule": "daily",
  "config": {
    "include_paths": ["/pricing", "/changelog/*", "/docs/*"],
    "only_main_content": true,
    "json_schema": [
      { "name": "plan_names", "type": "array" },
      { "name": "monthly_prices", "type": "array" }
    ],
    "llm_extract": { "connection_id": "…" },
    "change_digest": { "connection_id": "…", "focus": "prices, plan limits, new features" },
    "chat_url": "https://discord.com/api/webhooks/…"
  }
}

Three sections of one site. The fields are filled by your model only for pages whose content changed; the digest is one paragraph, ranked by what you said you care about.

what comes back
  • “Business went from $16 to $18 a seat; two enterprise features moved down a tier” — the paragraph, then the pages.
  • A field diff on the extracted prices, so the change is a number and not a paragraph to read.
  • A new changelog entry as “added”, with its opening lines.
  • Silence when nothing moved. The channel hears from a project only when there is something to say.
proofLinear’s pricing page through the MCP server: 1 credit, 4 seconds, four tiers.
change digeststructured fieldschat channelnoise suppressionown model key
06SEO and content teams · agencies

Content and SEO operations.

Client sites change under you; a rival’s blog publishes twice a week; a migration drops forty pages and the 404s are found when the traffic is. Crawlers that audit a site once do not remember the site the week before.

the configuration, as sent
{
  "seed": "https://client.example/",
  "schedule": "weekly",
  "config": {
    "include_paths": ["/blog/*"],
    "formats": ["markdown", "links"],
    "chat_url": "https://hooks.slack.com/services/…"
  }
}
// before a migration: run; after cut-over: run again — the record is the audit.

One section, weekly. The head fields — title, description, canonical, robots, h1 — are read on every page and diffed on every run without being asked for.

what comes back
  • Every title, description and canonical that moved, ranked; a canonical or robots change flagged as a ranking event.
  • Every URL now “missing” — a 404 or a soft 404 — listed apart from pages the site refused.
  • An orphan and link-graph report per run: pages nothing links to any more, pages not in the sitemap.
  • A rival’s blog: the new posts as “added”, with their openings, the morning after.
proofhotspotseo.com: the assistant found the blog lives under /blogs/, not /blog/, and the first run read 30 blog pages and nothing else.
head-field diffsmissing pageslink graphsection scopechat channel
07people who work in Claude or Cursor · agent builders

Give your agent the web.

An assistant cannot read a docs page, a pricing page or a supplier’s spec sheet — and when it can, the page eats the context window and the assistant reasons over navigation. Ask it to set up monitoring and it guesses at configuration keys.

the setup, and three things to say
pip install "mesharc[mcp]"
claude mcp add mesharc -e MESHARC_API_KEY=mesharc_… -- mesharc-mcp

# then, in the assistant:
you     Map fastapi.tiangolo.com and tell me which pages are about security.
you     Set up a project for hotspotseo.com that only scrapes the blog, weekly.
you     What changed on it since last week?

Seventeen tools. Mapping, searching, listing and changing settings cost nothing; a page costs what it took to read. Hosted, you approve the app and revoke it like any key; run locally, the key never leaves your machine.

what comes back
  • Main-content markdown, budgeted across the whole answer — a crawl of two hundred pages comes back as an index with excerpts, and any one page on its own on request, so a chat turn stays a chat turn.
  • Structured fields extracted on the server; the assistant sees the JSON, not the page.
  • A project the assistant configured itself, with keys the API validated — a guessed key is refused by name, never dropped.
  • The change record, on demand: “what changed?” is a tool call.
proofSeven conversations, run and recorded, on the MCP page.
MCP serverdescribe_project_configmap_siteextract_urlget_changes
08data engineers · analysts

A site into your warehouse, every week.

A directory, a catalogue, a listings site has to become a table and stay current. The scraper that fills it is a pet: it breaks on a layout change, it re-fetches what did not change, and the rows it writes are whatever it found, not the shape the query wants.

the configuration, as sent
{
  "seed": "https://catalogue.example/",
  "schedule": "weekly",
  "config": { "include_paths": ["/products/*"], "json_schema": [{ "name": "price", "type": "number" }, { "name": "sku", "type": "string" }] }
}
// the destination, with its shape:
{
  "provider": "postgres", "target": "products", "mode": "upsert", "key": "url",
  "columns": ["url", "title", "fields", "content_hash", "crawled_at", "change_class"],
  "shape": {
    "filter": [{ "field": "fields.price", "op": "exists" }],
    "map": [{ "name": "price_cents", "from": "fields.price", "ops": ["number", "mul:100"], "cast": "int" }],
    "drop": ["fields"]
  }
}

Fields come from what the page declares — JSON-LD, Open Graph, microdata — with a model on your key for the rest. The shape is declarative and previewed against the last run before it is saved.

what comes back
  • Rows upserted by URL after every run; only the pages that changed are written.
  • A “change_class” column — added, modified, removed — so the table carries its own history.
  • The changes, fields and sitemap datasets as tables of their own, or as CSV and JSONL.
  • S3, GCS, Azure Blob, Postgres, MySQL, ClickHouse, Mongo, Kafka: exercised against live instances.
destinationsshapingstructured fieldsupsert by URLexports

BigQuery, Snowflake, Redshift and Databricks are written to their documentation and marked unverified in the catalogue until a real account has confirmed them.

09and

Four more, in a line each.

Hiring signals

sales · recruiters · analysts

A project per company on its careers section; a page “added” under /careers/ is the signal, the morning it appears. Roles, teams and locations as fields.

A careers board hosted on Greenhouse or Lever is another host, so another project. Comfortable on Growth; the free plan’s two projects are for trying it.

Catalogue and price watch

retail · marketplaces

Product pages as structured rows, prices diffed as numbers, out-of-stock as a field change. Walled retail sites read at 2 credits a page after one render earns the session.

Most retail sites need the stealth tier — the residential exit and the solver — which comes with Growth and above.

Research archives

researchers · journalists · librarians

A section of a site kept over time: every run’s pages, the sitemap it declared, what appeared and what went, with removals confirmed twice and exports per run.

A corpus for a model

ML teams · evaluation

A whole documentation site as de-duplicated, main-content markdown: llms-full.txt in one file, or JSONL with the head fields and hashes, re-built as the site changes.

10not for

Said plainly, so you do not have to find out.

A one-off scrape you will never repeatThe memory is the product. For a single page, the playground or /scrape answers, and so do several cheaper tools.
Pages behind a sign-inThe crawler reads what a visitor can see. A login wall comes back as blocked, at no cost — and stays that way.
Sites that use hCaptchaDetected and refused. Turnstile and reCAPTCHA widgets are solved on request; hCaptcha is not.
Watching a page every minuteHourly is the finest schedule. A price that moves every ten minutes wants a feed, not a crawler.

Everything else on this page is true of the code today; the numbers under each case come from runs recorded in the repository. Where a destination has not been run against a live account, the case says so. Set a project up in the app, or describe it to your assistant with mesharc-mcp.

Pick a case and set it up.

A free key, a seed URL and a schedule. The first run is the baseline; the second one tells you what changed.