A RAG pipeline that never serves stale answers
The hidden cost of a stale embedding
A stale embedding scores as well as a fresh one. So an assistant keeps recommending a deprecated feature and quoting a deleted page, and nobody notices until a customer does. Re-indexing everything nightly fixes it and costs a fortune in embedding calls.
How MeshArc keeps it fresh
- Incremental crawls: only pages the sitemap marks as changed are fetched again.
- Content-hash diffs: only pages whose content actually changed are re-chunked and re-embedded.
- Safe deletes: a page removed from the site loses its vectors. A page a site temporarily blocked keeps them.
- Your keys, your costs: chunking and embedding run on your own embedding key, with no markup.
Works with your vector database
Qdrant, Weaviate and pgvector, tested against live instances, plus Pinecone. Or export llms.txt and llms-full.txt for assistants that read a whole corpus at once.
Frequently asked questions
What is a RAG pipeline?
Retrieval-augmented generation lets a model answer from your own documents. The pipeline collects those documents, splits and embeds them, and stores them in a vector database for retrieval. MeshArc handles collection and freshness for web sources.
What is llms.txt?
llms.txt is a plain-text file at a site’s root that gives AI assistants a clean map of what the site contains. MeshArc generates llms.txt and llms-full.txt from any crawl.
Fresh answers, lower embedding bills
Re-embed what changed. Leave the rest alone.