Parable raises $16.5M seed funding

How Parable uses DataFusion to build a best-in-class data stack

Getting work data out of a hundred tools usually means Fivetran or Airbyte, then dbt, then Snowflake. Provider Pools does ingestion, transformation, quality and query in one Apache DataFusion engine instead — about 30x cheaper in year one for a 10,000-person company, with data typed the moment it lands.

Title slide: "A data lake as a pure function of schema", Clinton Robinson, co-founder and CTO, Parable

Every Parable customer runs on 100+ software tools: calendar, email, chat, CRM, HR, tickets, code. To show a leader how work actually gets done, we need the activity in all of them — every employee, every system — modeled and joined into something a person or an agent can query. As Clint put it on his first slide: each customer has 100+ tools, and we have to make it make sense.

Earlier this month our co-founder and CTO, Clint Robinson, walked through how we do that at the Boston Apache DataFusion Meetup. The talk is nine minutes and worth it if you build data infrastructure. This post is the shorter version: why we built it this way, and what it buys us.

The default answer is three products

Ask how to get data out of a hundred SaaS tools and into a warehouse, and you'll hear about a stack: Fivetran or Airbyte to ingest, dbt to transform, Snowflake or Databricks to store and query. Each is excellent at its piece. Together they're how most data teams work.

They didn't fit what we're building:

  • Cost. Ingestion is priced per row, and we need every row of activity. The bill grows with exactly the data we care most about.
  • Engineering. Rows land raw. Typing, merging, tests and monitoring get hand-built per table — one to two engineer-years per customer.
  • Coverage. No connector catalog covers both employee activity and business records. Of the 39 sources we need, Fivetran lacks 7 and Airbyte 13.
  • Fragmentation. Ingestion stops at raw rows. Transformation, quality, storage and query are more products to buy, connect and keep in sync.

That last one matters more than it looks. When every stage is a different vendor, nothing gets cheaper as you grow. You pay each tool's margin on every row, and your engineers maintain the glue between them. There are no economies of scale in glue.

One engine, from source to modeled table

Provider Pools is our answer: ingestion, transformation and query in one engine, delivered as an open data lake in the customer's own cloud. Any tool can read it; there's no lock-in. The engine is Apache DataFusion.

The title of Clint's talk is the whole design: a data lake as a pure function of schema. A connector isn't code. Each API endpoint we read — Slack's user list, say — is described in configuration: how to page through it, which field is the key, which marks a deletion, which column holds a person's email, what quality rules a load has to pass. The engine compiles that description into a DataFusion plan. Nothing in between is hand-written.

A few things fall out of that:

  • Data is typed on arrival. An email isn't a string to us. We keep 60+ semantic types — email, person name, time zone and so on — in one Rust core, and DataFusion runs them during the write. " Clint.R@AskParable.com " from Slack and "CLINT.R@askparable.com" from Workday come out as the same person.
  • The same engine writes and reads. Ingestion and our SQL query service call the same code, and CI checks that they agree. Compare an email to a phone number and the planner refuses.
  • Quality is built in, not bought. Every load is checked before anyone builds on it. If today's load looks statistically off — a spike in empty emails, say — it's held back and not promoted.
  • AI calls are part of the plan. We can ask a model a question about every row in plain SQL: is this Slack account a bot? Those calls are the expensive part, so the plan is cut around them and every answer is logged. If a job crashes after 4,000 model calls, the rerun makes zero.

What it buys us

Here's Provider Pools against an assembled stack of Fivetran or Airbyte, dbt and Snowflake:

  • Peak ingest: 578 MB/s, typed on arrival, against 110–259 MB/s untyped. 2–5x faster.
  • Cost per million rows: $0.06–0.27 against $30–500.
  • Year-one cost for a 10,000-person company: $9–13K against $380–420K. About 30x lower.
  • Customer engineering in year one: 27–36 hours against 2,900–4,300. Roughly 100x less.

The stack figures are published rates, not runs on our data. Our year-one number excludes our own subscription, and theirs leaves out event streams, which would only widen the gap. Take the exact ratios with that in mind; the direction is the point. One engine that's typed, checked and queryable end to end is far cheaper to run than three products and the people who keep them talking.

Today that's 42 connectors and 444 tables. Our target for a new connector is under a day, down from about three.

Why DataFusion

We didn't need to write a query engine. DataFusion is fast, written in Rust, and open at nearly every layer: table formats, functions, optimizer rules, the query plan itself. Provider Pools uses almost all of them, which is how a small team can own the whole path from an API response to a modeled table.

We like it enough that we've started giving back. jev-datafusion is an open-source (Apache 2.0) package that adds four SQL functions — noul, choice, score and ask — that call TypeSafe's Jev model and return a typed judgment for every row of a query: a probability, a label, or a score on a rubric you write. It's how we're exploring Jev inside our own pipelines.

Thanks to the Apache DataFusion community for having us in Boston. If you're building on DataFusion too, we'd love to compare notes.

Watch Clint's talkClint's post on LinkedInjev-datafusion on GitHub

Get started

The data layer for team leaders.