Duckstring
Get your ducks in a row.
pip install duckstringDuckDB Data Engineering.
The world has been upsold on distributed computing. Data volumes into the terabytes can easily be handled on single machines, where DuckDB is at its best. Duckstring is an open source data engineering platform that gives DuckDB the space to do its magic.
Duckstring runs the same on your local machine, a single Cloud box, or coordinating across many compute instances for more complex pipelines, allowing you to start small and scale up whenever you need.
It is opinionated, but flexible:
- Modular Transformations. ↓ Every collection of transformations and business logic is encapsulated into version-controlled Ponds. Each declares their versioned dependencies - the DAG is implied, allowing you to treat transformations like software packages.
- Pull Orchestration. ↓ Orchestration is set from the perspective of the pipeline end. Nothing will execute unless something is actually consuming it.
- Modern Incrementality. ↓ Changes are detected and propagated across Ponds as Z-sets - a data format allowing even bushy joins to skip computing parts of a transformation that could not have changed.
- Data Lives With Code. ↓ The orchestrator and catalog live in Duckstring's Catchment - an environment for coordinating Ponds, their interactions and their execution.
Python Based and CLI-first ↓ Ponds are generic Python, with all the flexibility that provides.
Treat transformations like software packages.
Ponds use strict SemVer conventions. A new major version runs concurrently with the old one, and doesn't start until something is consuming from it. The headache of choreographing complex changes is avoided using the same techniques that allowed open source to flourish.
Upgrading a complex sequence of transformations can paralyze development. Just deploy breaks as a separate major-version Pond, and let them sit unused until consumers upgrade. See Versioning.
Only run what's used.
Push Orchestration sets schedules relative to the start of a pipeline. It throttles work downstream of a slow step, but blindly runs pipelines that have no consumers. Setting schedules instead at the end of a pipeline throttles both upstream and downstream of a slow step, and ensures that only paths with consumers are kept fresh.
No sophisticated prediction of run times is required. Duckstring uses the same scheduling system that keeps modern manufacturing processes humming. See Orchestration, or watch it run in the Playground.
Only compute what's changed.
Incremental approaches are typically hand-rolled and need careful engineering. Duckstring's DBSP-based Trickle package includes all standard distributive and algebraic methods in its query builder, so you don't have to think about it. Using DBSP methods and changed-key detection, built queries are optimized to skip any computation that could not have changed since the previous run. Done well, even massive data transformations can approach streaming-level latencies.
Change processing is Duckstring-managed, so you see the performance without the pain. See Trickles.
Manage code and assets together.
The execution environment is the Catchment, which also governs the catalog. Schemas are strictly defined as being within each major Pond version, and objects and tables are enumerated within them. If you have a Catchment, you have a catalog, governance and lineage.
Keeping track of where data lives need not be a separate task to defining how it's generated. See Catalog.
Everything is a Pond.
There's nothing stopping you taking your existing transformations and calling them a Pond. Import straight SQL, a dbt model or Ibis transformations. You can even simply call an external service, using Duckstring as the scheduler alone.
Every action against a Catchment is CLI-first - ready for all your bots to take control.
Want the platform cloud-hosted for you? Contact: dev@duckstring.com