Clean web data for AI, ready to train on
Better data in, better models out
- No boilerplate: navigation, footers, cookie banners and ads are removed, so compute is spent on content.
- No duplicates: every page is content-hashed, so identical pages and unchanged re-crawls never enter the dataset twice.
- Documents included: PDFs, Word files, spreadsheets and slide decks are converted to text and Markdown alongside web pages.
- Provenance on every row: URL, title, language, word count, crawl date and the method that read it — for filtering and for citation.
A corpus you can rebuild, not restart
Export a whole site as llms-full.txt in one file, or as JSONL with head fields and hashes. Re-run on a schedule and export only what changed, so an evaluation set stays current without being rebuilt from nothing.
Responsible collection by default
MeshArc honours robots.txt and crawl-delay, holds one rate limit per domain, and reads only what a public visitor can see. Pages behind a login are never accessed.
Frequently asked questions
What is the best format for LLM training data from the web?
Clean Markdown of the main content, with metadata kept alongside it: it preserves headings, lists and tables while removing the noise that wastes tokens. MeshArc outputs Markdown by default, plus JSON and plain text.
Can I extract structured fields for AI datasets?
Yes. Define a schema — product name, price, specifications — and MeshArc fills it from the page’s own structured data, falling back to your AI model on your own API key.
Build your next dataset in an afternoon
Point it at the sources, pick the formats, export the corpus.