Skip to content

WorkingRecord

WorkingRecord

WorkingRecord(base: RecordLike, con: DuckDBPyConnection)

Bases: Record

A Record whose last source is a staging area, plus the edit surface.

A Record in the type as well as in the fold: what it reads is the data with its pending edits applied, and it reads it by being one layer deeper than its base rather than by overlaying anything of its own. Every read member is inherited unchanged - outputs alone is overridden, results not overlaying - so an edit reads back through the same code path a committed layer would.

Staged rows live in connection-scoped DuckDB tables, the only place a staged row exists: the reads fold them rather than holding a copy, so what is staged is asked of the reads themselves.

The Resolver is fixed at construction and the source list never changes; a set changes what the last source's tables hold, not which sources there are. That is what lets the inherited members stay correct - they cache a key set only where the fold is stable, which a staged source makes it not.

Notes
Source code in src/datarecord/mutable.py
def __init__(self, base: RecordLike, con: DuckDBPyConnection) -> None:
    # `object.__setattr__` throughout: the base is a frozen dataclass, and
    # these are set before `super().__init__` because the `StagedSource`
    # below reads them off `self`.
    object.__setattr__(self, "_layer_id", uuid4())
    object.__setattr__(self, "_staged", {})
    base_cache = _base_resolver(base, con)
    object.__setattr__(self, "_base", base_cache)
    # This record *is* the fold one layer deeper, and that layer is the
    # staging area - so the field the base class holds is that fold, whose
    # last source is the staged layer `commit` writes.
    super().__init__(base_cache.with_source(StagedSource(self, self._layer_id)))

outputs property

outputs: Frames

Staged results, keyed by attribute - what a tool handed back.

Results reach a record through set(..., kind="outputs"), so a tool can solve against this record's pending inputs and attach what it computed without committing first. The base's results are not included: they were computed from inputs these edits may have changed, and results do not overlay, so what is staged is the whole answer.

Keeping them coherent with the inputs is the caller's business - editing an input after attaching results leaves results describing a record that no longer exists, and nothing here silently discards them.

Notes

set

set(
    attribute: str,
    value: Any,
    *,
    entity: Sequence[str] | None = None,
    kind: Literal["inputs", "outputs"] = "inputs",
    indexed_by: str | None = None,
    **dims: Any,
) -> None

Stage an attribute value for a group of components.

A labelled series must say what its index holds: indexed_by="snapshot" names the axis, and entity= implies it where the attribute has exactly one other coordinate. Without either the index is read as entity names. Never inferred from the labels themselves - an axis label may be a string just like a name, so that would make one call mean different things in two records.

value takes five forms: a scalar broadcast to every name, a sequence aligned positionally to names, a mapping keyed by name, a long frame supplying its own keys, and a narwhals expression - which is a function of the current value rather than a value, so it reads before it stages and two such calls compose.

No entity_type parameter: the type is looked up from the entity, so one call may span types and each is validated against its own type's spec. entity=None means every component whose type declares attribute.

Every other coordinate goes through **dims, including a group's - bus="north" for a connection attribute, from=/to= for a corridor. None has a parameter of its own, since which coordinates exist is declared rather than fixed.

kind names the destination in the format's own terms: "outputs" stages into outputs/ instead of inputs/, which is how a tool hands results back. Results use the same long schema; what differs is that they do not overlay.

Two checks are skipped for "outputs", both because a result is not a value the schema governs: the attribute need not be declared, and a result's name need not resolve to a declared member. A solve may produce rows for a component type it derived rather than read - PyPSA's SubNetwork is one - and rejecting those would refuse a legitimate result. An input for an undeclared name stays an error.

Notes
Source code in src/datarecord/mutable.py
def set(
    self,
    attribute: str,
    value: Any,
    *,
    entity: Sequence[str] | None = None,
    kind: Literal["inputs", "outputs"] = "inputs",
    indexed_by: str | None = None,
    **dims: Any,
) -> None:
    """Stage an attribute value for a group of components.

    A labelled series must say what its index holds: `indexed_by="snapshot"`
    names the axis, and `entity=` implies it where the attribute has exactly
    one other coordinate. Without either the index is read as entity names.
    Never inferred from the labels themselves - an axis label may be a string
    just like a name, so that would make one call mean different things in
    two records.

    `value` takes five forms: a scalar broadcast to every name, a sequence
    aligned positionally to `names`, a mapping keyed by name, a long frame
    supplying its own keys, and a narwhals expression - which is a *function
    of the current value* rather than a value, so it reads before it stages
    and two such calls compose.

    No `entity_type` parameter: the type is looked up from the entity,
    so one call may span types and each is validated against its own
    type's spec. `entity=None` means every component whose type
    declares `attribute`.

    Every other coordinate goes through `**dims`, including a group's -
    `bus="north"` for a connection attribute, `from=`/`to=` for a corridor.
    None has a parameter of its own, since which coordinates exist is
    declared rather than fixed.

    `kind` names the destination in the format's own terms:
    `"outputs"` stages into `outputs/` instead of `inputs/`, which is how a
    tool hands results back. Results use the same long schema; what differs
    is that they do not overlay.

    Two checks are skipped for `"outputs"`, both because a result is not a
    value the schema governs: the attribute need not be declared, and a
    result's `name` need not resolve to a declared member. A solve may
    produce rows for a component type it derived rather than read - PyPSA's
    `SubNetwork` is one - and rejecting those would refuse a legitimate
    result. An *input* for an undeclared name stays an error.

    Notes
    -----
    - [entity is unique across types](https://energy-models.github.io/datarecord/design/format/#entity-is-unique-across-types)
    - [outputs](https://energy-models.github.io/datarecord/design/read-path/#outputs)
    - [the shape of an edit](https://energy-models.github.io/datarecord/design/working-record/#the-shape-of-an-edit)
    - [set](https://energy-models.github.io/datarecord/design/working-record/#set)
    - [a derived value](https://energy-models.github.io/datarecord/design/working-record/#an-nwexpr-value-derived-from-the-current-one)
    - [validation](https://energy-models.github.io/datarecord/design/working-record/#validation)
    """
    is_long_frame = _is_frame(value) and _series_index(value) is None
    if is_long_frame:
        lazy = _incoming(value, self.con)
        if "entity_type" in lazy.collect_schema().names():
            msg = (
                f"`set({attribute!r}, <frame>)` was given a `entity_type` "
                f"column; names are unique across every type, so an attribute row "
                f"carries no type and the column would be ignored (https://energy-models.github.io/datarecord/design/format/#entity-is-unique-across-types)"
            )
            raise ValueError(msg)
        if kind == "inputs":
            self._validate_frame(lazy, attribute, dims)
        else:
            self._validate_result(attribute, dims)
        self._stage_long(attribute, lazy, kind, dims)
        return

    if isinstance(value, nw.Expr):
        if kind == "inputs":
            self._validate_dims(dims)
        else:
            self._validate_result(attribute, dims)
        self._stage_derived(attribute, value, entity=entity, kind=kind, **dims)
        return

    axis = self._axis_of(attribute) if kind == "inputs" else None
    if axis is not None:
        self._stage_axis(axis, attribute, value, entity=entity)
        return

    target = (
        list(entity) if entity is not None else self._names_declaring(attribute)
    )
    keys, values, per_dim = normalise_value(
        value,
        target,
        indexed_by=self._series_axis(attribute, value, indexed_by),
    )
    if keys is None:
        keys = target
        if len(values) == 1 and len(keys) > 1:
            values = values * len(keys)
    if kind == "inputs":
        self._validate_dims(dims)
        # One lookup serves both: rejects a name with no member row, and
        # returns the type whose spec is checked (https://energy-models.github.io/datarecord/design/format/#entity-is-unique-across-types).
        for name, ctype in self._resolve_types(keys).items():
            self._validate_attribute(ctype, attribute, dims, name=name)
    else:
        self._validate_result(attribute, dims)

    table = self._ensure(kind, attribute)
    self._stage_rows(attribute, table, keys, values, per_dim, dims)

add

add(ctype: str, frame: Any) -> None

Stage new components from a wide frame.

Splits it: attributes addressed by entity alone stay in dims/entity_type/, varying ones become inputs/ rows. Which is which comes from the schema, so this needs no framework registry.

Not a sequence of set calls: a component exists by virtue of its member row, so staging attribute values for a name no layer declares is what _validate_attribute rejects. Adding a bus with no attributes makes the point - nothing to set, yet the bus must exist.

ctype stays a parameter where set loses it: this is the call that establishes a name's type, so there is nothing yet to look it up in. It is also where uniqueness is enforced.

Notes
Source code in src/datarecord/mutable.py
def add(self, ctype: str, frame: Any) -> None:
    """Stage new components from a wide frame.

    Splits it: attributes addressed by `entity` alone stay in `dims/entity_type/`,
    varying ones become `inputs/` rows. Which is which comes from the
    schema, so this needs no framework registry.

    Not a sequence of `set` calls: a component exists by virtue of its
    member row, so staging attribute values for a name no layer declares is
    what `_validate_attribute` rejects. Adding a bus with no attributes makes
    the point - nothing to `set`, yet the bus must exist.

    `ctype` stays a parameter where `set` loses it: this is the call that
    establishes a name's type, so there is nothing yet to look it up in. It is also where uniqueness is enforced.

    Notes
    -----
    - [entity is unique across types](https://energy-models.github.io/datarecord/design/format/#entity-is-unique-across-types)
    - [add / remove](https://energy-models.github.io/datarecord/design/working-record/#add-remove)
    """
    lazy = _incoming(frame, self.con)
    columns = lazy.collect_schema().names()
    if "entity" not in columns:
        msg = "`add` needs an `entity` column"
        raise ValueError(msg)
    self._require_unique(ctype, lazy)

    declared = self.schema.attributes_for(ctype)
    varying = [c for c in columns if declared.get(c) and declared[c].varying]
    # An attribute addressed by a group belongs to that group's table, not
    # the member frame - putting it there would introduce a column the
    # ancestors' files lack, which then reads as NULL for their rows. It
    # says so by naming the group among its `dims`, which is what replaced
    # a field of its own (https://energy-models.github.io/datarecord/design/record/#connections).
    # Only a group `entity` itself keys: a wide frame adding one component
    # can describe that component's own row of such a group (`bus`, for
    # `connection`), but not a `corridor`, which relates two entities
    # neither is "the" one being added here - that goes through `add_group`.
    by_group: dict[str, list[str]] = {}
    for c in columns:
        if not declared.get(c) or c in varying:
            continue
        for group in self.schema.groups_of(c):
            if "entity" in self.schema.group_key(group):
                by_group.setdefault(group, []).append(c)
    ports = [c for cols in by_group.values() for c in cols]
    member_cols = [c for c in columns if c not in varying and c not in ports]

    rel = as_relation(lazy, self.con)
    # Where a group declares the type axis a component's constant columns are
    # a per-type member file; where none does there is no type to classify
    # into, so they are columns of the entity axis itself and there is no
    # member table (https://energy-models.github.io/datarecord/design/format/#where-a-value-lives).
    typed = self.schema.entity_type_dim is not None
    axis_supplied = {"deleted": lit(False)}  # noqa: FBT003
    if typed:
        axis_supplied["entity_type"] = lit(ctype)
        members = self._ensure(_MEMBERS, ctype)
        self._reject_undeclared(f"add({ctype!r}, ...)", members, member_cols)
        self._release_from_other_types(ctype, rel)
    else:
        self._reject_undeclared(
            f"add({ctype!r}, ...)", self._ensure(_ENTITY_AXIS), member_cols
        )

    # One row per component on the entity axis, replacing any this record
    # already staged for the name: it says the component exists, of what type
    # where there are types, and - untyped - carries its constant columns.
    # `add` after `remove` of the same name is thus one row: the tombstone is
    # deleted, not left to be outranked. Where typed, a second row goes to
    # the member table for those constant columns.
    self._insert(rel, self._ensure(_ENTITY_AXIS), axis_supplied, key=("entity",))
    if typed:
        self._insert(
            rel,
            members,
            {"deleted": lit(False)},  # noqa: FBT003
            key=("entity",),
        )
    for attribute in varying:
        # Always `inputs`: `add` declares components, and a component's
        # attribute values are inputs whatever a later solve produces.
        self._stage_long(
            attribute,
            lazy.select("entity", nw.col(attribute).alias("value")),
            "inputs",
            {},
        )
    # A group's coordinates name the row itself rather than being an
    # attribute of one, so they become that group's row; an attribute
    # addressed by the group rides along
    # (https://energy-models.github.io/datarecord/design/record/#connections).
    for group, group_cols in by_group.items():
        coordinates = self.schema.group_coordinates(group)
        extra = [c for c in group_cols if c not in coordinates]
        self.add_group(
            group,
            lazy.select(
                "entity",
                *(nw.col(c) for c in coordinates if c != "entity"),
                *(nw.col(c) for c in extra),
            ),
        )

remove

remove(ctype: str, names: Sequence[str]) -> None

Stage a tombstone per entity.

Need not enumerate what it deletes: one row per key, and the fold applies it to every attribute. Nor scope it - a component exists or it does not, so a deletion removes it whole.

One row, on the entity axis, which is where membership lives and the only place the fold reads a tombstone from. A member file holds values, never a deleted (_member_columns), so no second write there keeps step with this one. The axis row carries the type only where a group declares the axis; the delete keys on entity alone regardless, so it lands whatever type the add named.

Notes
Source code in src/datarecord/mutable.py
def remove(self, ctype: str, names: Sequence[str]) -> None:
    """Stage a tombstone per entity.

    Need not enumerate what it deletes: one row per key, and the fold
    applies it to every attribute. Nor scope it - a component exists or it
    does not, so a deletion removes it whole.

    One row, on the entity axis, which is where membership lives and the only
    place the fold reads a tombstone from. A member file holds values, never
    a `deleted` (`_member_columns`), so no second write there keeps step
    with this one. The axis row carries the type only where a group declares
    the axis; the delete keys on `entity` alone regardless, so it lands
    whatever type the `add` named.

    Notes
    -----
    - [add / remove](https://energy-models.github.io/datarecord/design/working-record/#add-remove)
    - [where a value lives](https://energy-models.github.io/datarecord/design/format/#where-a-value-lives)
    """
    if self.schema.entity_type_dim is not None:
        self._stage_tombstones(
            _ENTITY_AXIS,
            ("entity_type", "entity"),
            [[ctype, name] for name in names],
            ("entity",),
        )
    else:
        self._stage_tombstones(
            _ENTITY_AXIS, ("entity",), [[name] for name in names], ("entity",)
        )

add_group

add_group(group: str, frame: Any) -> None

Stage rows of one declared group from a frame carrying its coordinates.

The one path every group is added through, a record's connection group included.

No component type, which is no coordinate of a group (https://energy-models.github.io/datarecord/design/format/#where-a-value-lives).

Notes
Source code in src/datarecord/mutable.py
def add_group(self, group: str, frame: Any) -> None:
    """Stage rows of one declared group from a frame carrying its coordinates.

    The one path every group is added through, a record's `connection`
    group included.

    No component type, which is no coordinate of a group
    (https://energy-models.github.io/datarecord/design/format/#where-a-value-lives).

    Notes
    -----
    - [groups](https://energy-models.github.io/datarecord/design/schema/#groups)
    """
    coordinates = self.schema.group_coordinates(group)
    lazy = _incoming(frame, self.con)
    columns = lazy.collect_schema().names()
    for required in coordinates:
        if required not in columns:
            msg = f"`add_group({group!r}, ...)` needs a {required!r} column"
            raise ValueError(msg)
    table = self._ensure(group)
    extra = [c for c in columns if c not in coordinates]
    self._reject_undeclared(f"add_group({group!r}, ...)", table, extra)
    self._insert(
        as_relation(lazy, self.con),
        table,
        {"deleted": lit(False)},  # noqa: FBT003
        key=self.schema.group_key(group),
    )

remove_group

remove_group(
    group: str, keys: Sequence[tuple[Any, ...]]
) -> None

Stage a tombstone per key, over one declared group's group_key.

An into label is no part of a key: the tuple is removed, whatever label it carried.

Notes
Source code in src/datarecord/mutable.py
def remove_group(self, group: str, keys: Sequence[tuple[Any, ...]]) -> None:
    """Stage a tombstone per key, over one declared group's `group_key`.

    An `into` label is no part of a key: the tuple is removed, whatever
    label it carried.

    Notes
    -----
    - [groups](https://energy-models.github.io/datarecord/design/schema/#groups)
    """
    group_key = self.schema.group_key(group)
    self._stage_tombstones(group, group_key, [list(key) for key in keys], group_key)

rollback

rollback() -> None

Clear every staged row without writing.

Notes
Source code in src/datarecord/mutable.py
def rollback(self) -> None:
    """Clear every staged row without writing.

    Notes
    -----
    - [WorkingRecord](https://energy-models.github.io/datarecord/design/working-record/)
    """
    for name in self._staged.values():
        self.con.execute(f"DROP TABLE IF EXISTS {name}")
    self._staged.clear()

commit

commit(target: NewChild) -> Revision
commit(target: Directory) -> None
commit(target: Target) -> Revision | None

Write everything staged and clear it.

Returns:

Type Description
The new child for a `NewChild` target, so the caller can read what it
just wrote without going back to the record table; `None` for a
`Directory`, which belongs to no record. Overloaded on the target, so a
caller committing to a child holds a `Revision` rather than an optional
one - which target it passed is what decides, and it is always literal
at the call site.
The layer lands in the *child*, never in the node that was branched
from - layers are write-once - so it is the returned node that
reads back the edits.
Notes
Source code in src/datarecord/mutable.py
def commit(self, target: Target) -> Revision | None:
    """Write everything staged and clear it.

    Returns
    -------
    The new child for a `NewChild` target, so the caller can read what it
    just wrote without going back to the record table; `None` for a
    `Directory`, which belongs to no record. Overloaded on the target, so a
    caller committing to a child holds a `Revision` rather than an optional
    one - which target it passed is what decides, and it is always literal
    at the call site.

    The layer lands in the *child*, never in the node that was branched
    from - layers are write-once - so it is the returned node that
    reads back the edits.

    Notes
    -----
    - [a layer's data is write-once](https://energy-models.github.io/datarecord/design/layers/#a-layers-data-is-write-once)
    - [committing](https://energy-models.github.io/datarecord/design/working-record/#committing)
    """
    if isinstance(target, NewChild):
        parent = (
            target.record if target.record is not None else self._base_revision()
        )
        child = parent.child()
        # The staged layer's own rows - the fold's last source, a `LayerData`
        # like any other - so only the edits are written, the fold resolving
        # the rest from the parent (https://energy-models.github.io/datarecord/design/working-record/#committing).
        write_record(child.id, self.resolver.sources[-1], self.con)
        self.rollback()
        return child
    # The `Resolver`, not `self`: it is the same `LayerData` a fold answers
    # for everything folded in, which is exactly base-plus-staged flattened
    # (https://energy-models.github.io/datarecord/design/working-record/#reading-with-pending-edits).
    write_record(None, self.resolver, self.con, uri=target.uri)
    self.rollback()
    return None

NewChild dataclass

NewChild(record: Revision | None = None)

Write the staged rows as a new child layer of record.

Only the edits are written; the fold resolves the rest from the parent.

record defaults to the node the WorkingRecord was built over, which is what a caller branching from a revision means every time. Passing one explicitly is for the rarer case of re-parenting the edits elsewhere; a WorkingRecord over a base that is not a layered node (a directory, a framework object) has nothing to default to and must supply it.

Notes

Directory dataclass

Directory(uri: str)

Write a standalone record at uri: staged rows plus what the record already reads, there being no parent to resolve against.

Notes