Docs · Connectors and destinations

Connectors and destinations

A connection is a provider plus the credentials to reach it — a bucket and a key, a Postgres DSN, a Mongo URI. It belongs to the workspace, is named, and is tested before it is trusted (POST /connections/{id}/test). Secrets are kept apart from the rest of the config, encrypted at rest, and never returned by the API.

A destination binds a project — or, shared, every project or a chosen set — to a connection and says what to write: the target, which columns, the write mode, the key, whether every page goes or only the ones the change record says moved, and when. A sync is one destination writing one run, recorded with rows written, rows skipped and why, and the location.

Push is free: no credits are charged for writing.

The connectors#

Verified says whether MeshArc has run it end to end against a live instance; the ones marked account are written to the provider's current documentation and are confirmed with a real account on request.

ProviderKindFieldsSecretsWhat it writesVerified
Any LLM (llm)modelbase_url, modelapi_key (optional)Not a destination: fills a project's JSON schema. OpenAI, Anthropic, Gemini, Mistral, Groq, Ollama, vLLM and most others — the base URL of the chat-completions endpoint, the model name, your key✓
Amazon S3 and S3-compatible (s3)object storebucket, prefix, region, endpointaccess_key_id, secret_access_keyOne JSONL file per run under a dated prefix; set endpoint for MinIO, R2 or GCS with HMAC keys✓ (MinIO)
Google Cloud Storage (gcs)object storebucket, prefixaccess_key_id, secret_access_keyThrough the S3-interoperability endpoint with HMAC keys; one JSONL file per run✓
Azure Blob (azure)object storecontainer, prefixconnection_stringOne JSONL blob per run under a dated prefix; the container is created if missing✓ (Azurite)
PostgreSQL (postgres)databaseschemadsnOne row per page version; upsert with ON CONFLICT on the dedupe key; the table is created if missing✓
MySQL (mysql)databasedatabasedsnOne row per page version; INSERT … ON DUPLICATE KEY UPDATE; table created if missing✓
MongoDB (mongodb)databasedatabaseuriOne document per page version, markdown inline; upsert on the key✓
ClickHouse (clickhouse)databaseurl, databaseuser, passwordMergeTree ordered by (url, crawled_at); upsert through ReplacingMergeTree; over HTTP✓
Google BigQuery (bigquery)warehouseproject, dataset, locationservice_account_jsonA load job per run; upsert stages then MERGEs. Needs bigquery.tables.create/updateData on the datasetaccount
Snowflake (snowflake)warehouseaccount, user, warehouse, database, schema, rolepassword or key-pair (PEM)A temporary table then MERGE; JSON columns as VARIANTaccount
Amazon Redshift (redshift)warehousehost, port, database, schema, userpasswordA temp table then MERGE on the key; JSON as SUPER; rows over the wire, not S3 COPYaccount
Databricks (databricks)warehouseserver_hostname, http_path, catalog, schemaaccess_tokenA Delta table in Unity Catalog; upsert stages then MERGE INTOaccount
Pinecone (pinecone)vector indexcloud, region, namespace, embedding_url, embedding_model, embedding_dimsPinecone key, embedder keyA serverless index with the embedder's dimensions; one vector per chunk, mapped columns as metadata; a page's chunks replaced by id prefix on upsertaccount
Qdrant (qdrant)vector indexurl, embedding_url, embedding_model, embedding_dimsQdrant key, embedder keyOne point per chunk, mapped columns as payload, upsert per URL✓
Weaviate (weaviate)vector indexurl, grpc_port, embedding_url, embedding_model, embedding_dimsWeaviate key, embedder keyA collection with self-provided vectors; one object per chunk; a page's chunks replaced on upsert. Cloud or self-hosted✓
pgvector (pgvector)vector indexschema, embedding_url, embedding_model, embedding_dimsdsn, embedder keyChunk, embedding and the mapped columns in one table (vector(dims)); upsert per URL; the extension is created if missing✓
Kafka (kafka)streambootstrap_servers, security_protocol, sasl_mechanismSASL username / passwordOne JSON message per row keyed on the URL, with x-mesharc-run and x-mesharc-change headers; the topic is created if missing; PLAINTEXT or SASL (PLAIN / SCRAM)✓ (Redpanda)
HTTP webhook (webhook)stream——Not a connection: configured per project under Configure → Advanced — see webhooks✓

The catalogue also lists a Local disk connector; it writes to MeshArc's own storage and exists for testing, not for your data.

Vector indexes: MeshArc chunks the markdown into paragraph-aligned pieces of about 350 words with a 40-word overlap; your OpenAI-compatible embeddings endpoint (embedding_url + embedding_model, dimensions declared) embeds them. Each chunk is a point with a stable id per (url, chunk), the vector, and the mapped columns as payload with the chunk text and index — never the whole markdown.

A destination#

FieldMeaningDefault
connection_idWhich connection—
targetA table, a collection, a topic, or a path under the bucket / directoryrequired
modeupsert (on the key) or append. A file or a message per run cannot upsert, so object stores and Kafka are recorded as appendupsert
keyWhat makes a row unique: url, url+crawled_at, url+content_hash. The key columns lead the rowurl
columnsWhich of MeshArc's columns to write, and in what order (every column is listed under The columns)the default set
incrementaltrue: only pages the change record classed added, modified, removed, with field changes, or a baseline — so a destination receives what moved. false: every page of the runtrue
scheduleafter_run (fires when a run finishes) or manual (Sync now)after_run
scopeproject (one project's), all (every project of the workspace) or projects (a chosen set) — shared destinations carry site, project_id and project_name so one table holds many sitesproject
shapeA filter, added columns and dropped columns — see Shapingnone
enabled, namePause without deleting; a labeltrue, ""

Up to 10 destinations per project and 20 shared ones per workspace; a sync writes at most 20,000 rows.

The columns#

GroupColumns
identityurl, site, final_url, crawled_at, run_id, project_id, project_name, content_hash, html_hash
contentmarkdown, text, raw_html, word_count, token_count
fieldstitle, meta_description, canonical_tag, robots_content, h1_tag, og_tag, twitter_tag, redirect_urls, response_status, fields (the structured fields as JSON)
diagnosticsblock_reason, fetch_ms, depth, links, images, archive_path, language, found_via, first_referer, in_links, in_sitemap, orphan, unlisted
changechange_class, added_words, removed_words, changed_fields, reviewed_by

The default set: url, site, final_url, crawled_at, run_id, project_id, content_hash, markdown, word_count, token_count, title, meta_description, canonical_tag, robots_content, h1_tag, redirect_urls, response_status, block_reason, depth, change_class, added_words, removed_words, changed_fields.

Shaping#

A warehouse has dbt after the load; an application table or a vector index has nothing after the write, so what lands must already be the shape the reader wants. A shape is declarative — no code runs — and lives on the destination, so the same run can land raw in one place and shaped in another. POST /destinations/preview shows the shaped rows of the last run before you save.

{
  "filter": [
    {"field": "status", "op": "eq", "value": "ok"},
    {"field": "path", "op": "glob", "value": "/products/*"},
    {"field": "fields.price", "op": "exists"},
    {"field": "word_count", "op": "gte", "value": 120}
  ],
  "map": [
    {"name": "price_cents", "from": "fields.price", "ops": ["number", "mul:100"], "cast": "int"},
    {"name": "category", "from": "url", "ops": ["path_segment:1"]},
    {"name": "source", "from": "=mesharc"}
  ],
  "drop": ["raw_html"]
}
  • filter — conditions every row must pass, all of them (up to 20). Fields are the row's columns, plus status, path and host (visible to the filter even when not written), plus anything under fields. from the page's structured data. Operators: eq, ne, in, not_in, gt, gte, lt, lte, contains, not_contains, matches (regex), glob, exists, missing, empty, not_empty.
  • map — columns to add or replace (up to 60), each from a source field (or a literal with =) through a chain of up to 8 functions and a cast (string, int, float, bool, date, json). Functions: trim, lower, upper, number, mul:N, date, host, path, path_segment:N, regex:pattern, split:sep, slice:a:b, replace:from:to, default:value, len, first, join:sep, json, not_empty.
  • drop — base columns to leave out after mapping.

What a shape cannot do belongs in the reader's own pipeline; the function set is small on purpose.

One-off jobs#

A destination does not need a project. POST /scrape, POST /crawl and the Playground (Send to) may each name one, and when the job is done its pages are written through that destination's columns, shape and sink — every page, with change_class of baseline, since a one-off has no run before it to compare with. A batch of scrapes is one write, however many hosts it spans. The write is on the destination's own sync history and on the job, under destinationSync.

A one-off crawl writes only where it names. A shared destination scoped to every project (all) does not receive one-off crawls: that scope is about the sites you watch, and a quick crawl turning up in that table would be a surprise. Kept as a project (POST /crawl/{id}/keep), it is written like any other.

Syncs#

GET /destinations/{id}/syncs (and the per-project form) lists every sync: when, which run (and which project's, for a shared destination), rows written, rows skipped and why (filtered out, unchanged, over the row cap), the location written, and an error when the write failed. A failed sync does not fail the run; the Destinations screen shows it and Sync now retries.