Design¶
Status: Draft · Owner: Jonas Hörsch · Date: 2026-08-05
The authoritative design for datarecord. Usage is how to use the package; these pages are what it is and why it is that way. The docstrings in the source cite them by link rather than restating the argument — a comment that re-argues the design is a defect.
What a data record is¶
Dimensioned attribute data with a declared schema.
A record holds components (named members of a type), groups of them — connections between components and buses being the one every network has — attribute values over both, and the axes those values vary along. A schema declares what may exist; the data says what does.
A record exposes seven things:
record.schema what may exist: the axes, the attributes
record.dims the axes themselves, keyed by dim
record.entity_types members, keyed by entity type
record.groups which tuples exist, keyed by group then component type
record.attributes the values, keyed by attribute name
record.outputs results, keyed by attribute name
record.flags(ctype) which axes an attribute actually uses
That is the Record protocol, and The Record protocol gives it precisely.
A component's entity identifies it across every type: names are unique record-wide, not per type (entity is unique across types).
That is why the values are keyed by attribute and not by type — an attribute row names a component and nothing more, and a component's type is something the record knows about it rather than part of its address.
Record is the one class that answers all of this, and it is the narwhals interface over a fold across layers:
- A tree of layers, each adding a partial record on top of its parent, resolved last-writer-wins. No single directory is the record; the answer is the fold across them.
Revision.recordgives one. - One parquet directory, via
Record.at(uri)— folded over the single layer it is, which degenerates to a scan of it. Not a second implementation. WorkingRecord— aRecordwhose last layer is a staging area, so pending edits read back before anything is written.
RecordLike is the protocol all of these satisfy, and so does a framework's own object — a PyPSA Network presenting itself as a record, without depending on this package at all.
A consumer cannot tell which it holds, so a framework reads a hundred-layer overlay through the same call it would use for a single directory.
Neither the concept nor this package names a modelling framework.
A framework consumes a record, a workflow engine produces one, and neither needs to know how the other works.
datarecord depends only on duckdb, narwhals and pydantic.
Scope¶
- In scope: the
Recordprotocol (the definition) andWorkingRecord; the parquet format that stores it; the schema; overlay resolution and its owner map; the write path; why one directory needs no second implementation. - Out of scope: a non-DuckDB implementation (the protocol permits one, see the protocol names no engine, but only DuckDB-backed ones are provided); concurrent writers to one record; unmaterialised/meta layers.
The pages¶
| page | what it settles |
|---|---|
| The Record protocol | what a consumer codes against, and what it may assume |
| The record format | the parquet directory a record is stored as |
| The schema | what may exist: dims, attributes, patch granularity |
| Layered resolution | a tree of layers, folded last-writer-wins |
| The DuckDB read path | the owner map and how a relation resolves |
| Writing a record | write_record, and what it validates |
WorkingRecord |
editing: staging, committing, reading back |
| Consuming a record | tools, and the seam a framework meets |
| Module layout | where the code lives, and the one-way dependency |
| Open questions | what is deliberately unsettled |