A directory, a catalogue, a listings site has to become a table and stay current. The scraper that fills it is a pet: it breaks on a layout change, it re-fetches what did not change, and the rows it writes are whatever it found, not the shape the query wants.
the configuration, as sent
{
"seed": "https://catalogue.example/",
"schedule": "weekly",
"config": { "include_paths": ["/products/*"], "json_schema": [{ "name": "price", "type": "number" }, { "name": "sku", "type": "string" }] }
}
// the destination, with its shape:
{
"provider": "postgres", "target": "products", "mode": "upsert", "key": "url",
"columns": ["url", "title", "fields", "content_hash", "crawled_at", "change_class"],
"shape": {
"filter": [{ "field": "fields.price", "op": "exists" }],
"map": [{ "name": "price_cents", "from": "fields.price", "ops": ["number", "mul:100"], "cast": "int" }],
"drop": ["fields"]
}
}
Fields come from what the page declares — JSON-LD, Open Graph, microdata — with a model on your key for the rest. The shape is declarative and previewed against the last run before it is saved.