Connectors and destinations
A connection is a provider plus the credentials to reach it — a bucket and a key, a Postgres DSN, a Mongo URI. It belongs to the workspace, is named, and is tested before it is trusted (POST /connections/{id}/test). Secrets are kept apart from the rest of the config, encrypted at rest, and never returned by the API.
A destination binds a project — or, shared, every project or a chosen set — to a connection and says what to write: the target, which columns, the write mode, the key, whether every page goes or only the ones the change record says moved, and when. A sync is one destination writing one run, recorded with rows written, rows skipped and why, and the location.
Push is free: no credits are charged for writing.
The connectors#
Verified says whether MeshArc has run it end to end against a live instance; the ones marked account are written to the provider's current documentation and are confirmed with a real account on request.
| Provider | Kind | Fields | Secrets | What it writes | Verified |
|---|---|---|---|---|---|
Any LLM (llm) | model | base_url, model | api_key (optional) | Not a destination: fills a project's JSON schema. OpenAI, Anthropic, Gemini, Mistral, Groq, Ollama, vLLM and most others — the base URL of the chat-completions endpoint, the model name, your key | ✓ |
Amazon S3 and S3-compatible (s3) | object store | bucket, prefix, region, endpoint | access_key_id, secret_access_key | One JSONL file per run under a dated prefix; set endpoint for MinIO, R2 or GCS with HMAC keys | ✓ (MinIO) |
Google Cloud Storage (gcs) | object store | bucket, prefix | access_key_id, secret_access_key | Through the S3-interoperability endpoint with HMAC keys; one JSONL file per run | ✓ |
Azure Blob (azure) | object store | container, prefix | connection_string | One JSONL blob per run under a dated prefix; the container is created if missing | ✓ (Azurite) |
PostgreSQL (postgres) | database | schema | dsn | One row per page version; upsert with ON CONFLICT on the dedupe key; the table is created if missing | ✓ |
MySQL (mysql) | database | database | dsn | One row per page version; INSERT … ON DUPLICATE KEY UPDATE; table created if missing | ✓ |
MongoDB (mongodb) | database | database | uri | One document per page version, markdown inline; upsert on the key | ✓ |
ClickHouse (clickhouse) | database | url, database | user, password | MergeTree ordered by (url, crawled_at); upsert through ReplacingMergeTree; over HTTP | ✓ |
Google BigQuery (bigquery) | warehouse | project, dataset, location | service_account_json | A load job per run; upsert stages then MERGEs. Needs bigquery.tables.create/updateData on the dataset | account |
Snowflake (snowflake) | warehouse | account, user, warehouse, database, schema, role | password or key-pair (PEM) | A temporary table then MERGE; JSON columns as VARIANT | account |
Amazon Redshift (redshift) | warehouse | host, port, database, schema, user | password | A temp table then MERGE on the key; JSON as SUPER; rows over the wire, not S3 COPY | account |
Databricks (databricks) | warehouse | server_hostname, http_path, catalog, schema | access_token | A Delta table in Unity Catalog; upsert stages then MERGE INTO | account |
Pinecone (pinecone) | vector index | cloud, region, namespace, embedding_url, embedding_model, embedding_dims | Pinecone key, embedder key | A serverless index with the embedder's dimensions; one vector per chunk, mapped columns as metadata; a page's chunks replaced by id prefix on upsert | account |
Qdrant (qdrant) | vector index | url, embedding_url, embedding_model, embedding_dims | Qdrant key, embedder key | One point per chunk, mapped columns as payload, upsert per URL | ✓ |
Weaviate (weaviate) | vector index | url, grpc_port, embedding_url, embedding_model, embedding_dims | Weaviate key, embedder key | A collection with self-provided vectors; one object per chunk; a page's chunks replaced on upsert. Cloud or self-hosted | ✓ |
pgvector (pgvector) | vector index | schema, embedding_url, embedding_model, embedding_dims | dsn, embedder key | Chunk, embedding and the mapped columns in one table (vector(dims)); upsert per URL; the extension is created if missing | ✓ |
Kafka (kafka) | stream | bootstrap_servers, security_protocol, sasl_mechanism | SASL username / password | One JSON message per row keyed on the URL, with x-mesharc-run and x-mesharc-change headers; the topic is created if missing; PLAINTEXT or SASL (PLAIN / SCRAM) | ✓ (Redpanda) |
HTTP webhook (webhook) | stream | — | — | Not a connection: configured per project under Configure → Advanced — see webhooks | ✓ |
The catalogue also lists a Local disk connector; it writes to MeshArc's own storage and exists for testing, not for your data.
Vector indexes: MeshArc chunks the markdown into paragraph-aligned pieces of about 350 words with a 40-word overlap; your OpenAI-compatible embeddings endpoint (embedding_url + embedding_model, dimensions declared) embeds them. Each chunk is a point with a stable id per (url, chunk), the vector, and the mapped columns as payload with the chunk text and index — never the whole markdown.
A destination#
| Field | Meaning | Default |
|---|---|---|
connection_id | Which connection | — |
target | A table, a collection, a topic, or a path under the bucket / directory | required |
mode | upsert (on the key) or append. A file or a message per run cannot upsert, so object stores and Kafka are recorded as append | upsert |
key | What makes a row unique: url, url+crawled_at, url+content_hash. The key columns lead the row | url |
columns | Which of MeshArc's columns to write, and in what order (every column is listed under The columns) | the default set |
incremental | true: only pages the change record classed added, modified, removed, with field changes, or a baseline — so a destination receives what moved. false: every page of the run | true |
schedule | after_run (fires when a run finishes) or manual (Sync now) | after_run |
scope | project (one project's), all (every project of the workspace) or projects (a chosen set) — shared destinations carry site, project_id and project_name so one table holds many sites | project |
shape | A filter, added columns and dropped columns — see Shaping | none |
enabled, name | Pause without deleting; a label | true, "" |
Up to 10 destinations per project and 20 shared ones per workspace; a sync writes at most 20,000 rows.
The columns#
| Group | Columns |
|---|---|
| identity | url, site, final_url, crawled_at, run_id, project_id, project_name, content_hash, html_hash |
| content | markdown, text, raw_html, word_count, token_count |
| fields | title, meta_description, canonical_tag, robots_content, h1_tag, og_tag, twitter_tag, redirect_urls, response_status, fields (the structured fields as JSON) |
| diagnostics | block_reason, fetch_ms, depth, links, images, archive_path, language, found_via, first_referer, in_links, in_sitemap, orphan, unlisted |
| change | change_class, added_words, removed_words, changed_fields, reviewed_by |
The default set: url, site, final_url, crawled_at, run_id, project_id, content_hash, markdown, word_count, token_count, title, meta_description, canonical_tag, robots_content, h1_tag, redirect_urls, response_status, block_reason, depth, change_class, added_words, removed_words, changed_fields.
Shaping#
A warehouse has dbt after the load; an application table or a vector index has nothing after the write, so what lands must already be the shape the reader wants. A shape is declarative — no code runs — and lives on the destination, so the same run can land raw in one place and shaped in another. POST /destinations/preview shows the shaped rows of the last run before you save.
{
"filter": [
{"field": "status", "op": "eq", "value": "ok"},
{"field": "path", "op": "glob", "value": "/products/*"},
{"field": "fields.price", "op": "exists"},
{"field": "word_count", "op": "gte", "value": 120}
],
"map": [
{"name": "price_cents", "from": "fields.price", "ops": ["number", "mul:100"], "cast": "int"},
{"name": "category", "from": "url", "ops": ["path_segment:1"]},
{"name": "source", "from": "=mesharc"}
],
"drop": ["raw_html"]
}
- filter — conditions every row must pass, all of them (up to 20). Fields are the row's columns, plus
status,pathandhost(visible to the filter even when not written), plus anything underfields.from the page's structured data. Operators:eq,ne,in,not_in,gt,gte,lt,lte,contains,not_contains,matches(regex),glob,exists,missing,empty,not_empty. - map — columns to add or replace (up to 60), each from a source field (or a literal with
=) through a chain of up to 8 functions and a cast (string,int,float,bool,date,json). Functions:trim,lower,upper,number,mul:N,date,host,path,path_segment:N,regex:pattern,split:sep,slice:a:b,replace:from:to,default:value,len,first,join:sep,json,not_empty. - drop — base columns to leave out after mapping.
What a shape cannot do belongs in the reader's own pipeline; the function set is small on purpose.
One-off jobs#
A destination does not need a project. POST /scrape, POST /crawl and the Playground (Send to) may each name one, and when the job is done its pages are written through that destination's columns, shape and sink — every page, with change_class of baseline, since a one-off has no run before it to compare with. A batch of scrapes is one write, however many hosts it spans. The write is on the destination's own sync history and on the job, under destinationSync.
A one-off crawl writes only where it names. A shared destination scoped to every project (all) does not receive one-off crawls: that scope is about the sites you watch, and a quick crawl turning up in that table would be a surprise. Kept as a project (POST /crawl/{id}/keep), it is written like any other.
Syncs#
GET /destinations/{id}/syncs (and the per-project form) lists every sync: when, which run (and which project's, for a shared destination), rows written, rows skipped and why (filtered out, unchanged, over the row cap), the location written, and an error when the write failed. A failed sync does not fail the run; the Destinations screen shows it and Sync now retries.