Skip to main content

Puddle Handle

Every @puddle function receives a Puddle object. It writes sample Source data into puddles/ponds/{source}/data/, where duckstring pond run reads it as if it were the Source's published output.

from duckstring import puddle

@puddle("transactions.transaction")
def transactions(p):
p.write_table(p.con.sql("""
SELECT i AS id, DATE '2026-01-01' + i AS created_at, 1 + i % 10 AS product_id,
1 + i % 5 AS quantity, 1 + i % 3 AS store_id
FROM range(100) t(i)
"""))

Attributes​

AttributeTypeDescription
conduckdb.DuckDBPyConnectionA scratch in-memory DuckDB connection.
pathpathlib.PathThe directory the Puddle writes into, puddles/ponds/{source}/data/. Created when first accessed. Anything written here directly is visible to the local run.
sourcestrThe Source name from the decorator target.
tablestr or NoneThe table name from the decorator target, or None for a whole-Source Puddle.

Methods​

write_table​

p.write_table(relation) -> pathlib.Path
p.write_table(name, relation) -> pathlib.Path

Writes a relation to {path}/{name}.parquet and returns the file path. The one-argument form uses the table named in the decorator target; a whole-Source Puddle must name each table.

ParameterTypeDescription
namestrTable name.
relationDuckDB relation or DataFrameThe data. A pandas DataFrame is converted through con.

Raises ValueError for the one-argument form on a whole-Source Puddle.

write_path​

p.write_path(src) -> None

Copies existing Parquet or CSV files into the Puddle. For a single-table Puddle, everything matched becomes that table. For a whole-Source Puddle, each file becomes a table named after the file.

ParameterTypeDescription
srcpath or strA file path or glob, such as "~/samples/*.parquet".

Raises FileNotFoundError when a glob matches nothing.

catchment​

p.catchment(name=None) -> Catchment

Returns a Catchment client whose default Pond is this Puddle's Source and default table is its target table, sharing con. Use it to sample real data from a running Catchment:

@puddle("products.product")
def product(p):
p.write_table(p.catchment().query("SELECT * FROM product USING SAMPLE 10%"))
ParameterTypeDefaultDescription
namestrthe default CatchmentA registered Catchment name, or a URL.

Objects​

MethodDescription
write_object(name, src)Seeds an Object for the Source from a path (file or directory), bytes, or a binary file-like, so a Ripple reading "{source}.{name}" finds it.
read_object(name)Returns the bytes of a seeded single-file Object.
object_path(name)Returns a local path to a seeded Object, file or directory.