Skip to content

Schema

Schema

Bases: BaseModel

One record's schema.

Attributes:

Name Type Description
version int

Bumped by any change to the declarations. A reader meeting a version it was not written for should refuse rather than guess.

dimensions dict[str, Dimension]

Every declared axis, keyed by dim name.

attributes dict[str, AttributeSpec]

Attribute -> spec, flat and record-wide. One attribute is one spec and one inputs/<attr>.parquet, so a dtype cannot differ per type.

groups dict[str, Group]

Group name -> which tuples over several dims exist. connection is the one every record with connections declares, and the entity-type axis is the group into that axis over [entity].

traits dict[str, Trait]

Trait -> the attributes it bundles and the entity types carrying them. A vocabulary a consumer dispatches on, declared rather than derived, and the only thing that narrows an attribute to some entity types.

partial frozenset[str] | None

Which dims a layer may patch value by value. None for a record with no layers, since nothing overrides anything. A dim outside it is one a layer owns entirely once it touches it.

meta dict[str, Any]

A framework's own top-level data - network attributes, CRS, free-form metadata. Stored and never interpreted, since none of it describes the dimensioned data.

Notes

results class-attribute instance-attribute

results: dict[str, AttributeSpec] = Field(
    default_factory=dict
)

What a solve computes, keyed like attributes and shaped the same.

Separate because the two are governed differently, not because they are stored differently: a result is written to outputs/<attr>.parquet rather than inputs/, never overlays a parent's, and may name a component the record does not declare. Keeping it out of attributes is what stops it reaching attributes_for, and so add's wide-frame split and the input validation, neither of which a result should meet.

dims property

dims: tuple[str, ...]

Every declared dim, in declaration order - the long schema's dim columns.

Notes

broadcast_dims property

broadcast_dims: tuple[str, ...]

The dims a NULL broadcasts over: every dim but entity, the type axis and a group's key.

A NULL here means "every value of this dim", which the fold expands against the axis. The three exclusions cannot mean that:

  • entity, because a NULL there is a value belonging to no component rather than to all of them. The one dim named literally, being the one every entity-type axis classifies.
  • The entity-type axis, because it inherits entity's exclusion: its labels are a column of dims/entity.parquet, so a NULL there is a component whose type is unknown rather than one of every type. An attribute addressed by the type alone never reaches here: it is a column of the type axis file rather than a long row, so it has no NULL to expand.
  • A coordinate a group is keyed by, because there is no axis to expand against. "Every bus of this component" is the group's rows, not the bus axis - a sparse subset only the group's table knows.

A functional group's into dim is not excluded, though it is one of the group's coordinates: only key addresses a row, so country broadcasts like any other axis.

The complement of this is what Schema requires to be partial: a dim whose values are addressed individually is one a layer patches value by value.

What the varies/broadcast structs have a field per, and what expand_dims joins.

Notes

entity_type_dim property

entity_type_dim: str | None

The dim classifying entity, or None where none does.

The into of the group over entity alone, of which the schema admits at most one.

Notes

entity_types property

entity_types: frozenset[str]

Every declared entity-type label - the types a component may be.

The entity-type axis's enum categories, so a schema declaring it as a plain String has none: the labels are then data rather than declarations, and attributes_for accepts any of them.

Notes

membership_keys property

membership_keys: tuple[str, ...]

The dims addressed per row rather than broadcast: entity and group coords.

A membership key is a coordinate a layer patches one row of at a time - one component, one connection - never "every value" of an axis. It is the non-broadcast, addressable dims: every dim but the broadcast ones and the entity-type axis, which is a column of dims/entity.parquet rather than an addressable coordinate.

These land in the fold key by being membership, not by being partial.

Notes

partial_dims property

partial_dims: tuple[str, ...]

The fold key's dims, in declaration order.

The membership keys plus the broadcast value dims a layer may patch per value (partial). The fold's key is one fixed tuple over all attributes, so it carries every axis any layer may patch by value or by row, not only those some currently declared attribute varies over. An attribute not owned per one of them writes NULL there, the "NULL means all values" rule - which also lets a schema declare an axis before any attribute uses it.

Notes

long_columns property

long_columns: tuple[str, ...]

The long schema's full column set.

The map's column set, which is uniform across attributes because the map is one relation over all of them. An individual file carries only its own attribute's columns (long_columns_for), and union_by_name supplies NULL for the rest when the fold unions them here.

No entity_type: a row here is keyed by entity, and an entity is unique record-wide, so a type column would restate what the entity already says and let the two disagree. Every declared attribute is shaped by long_columns_for rather than by this, where naming both is rejected outright and one addressed by the type alone is a column of the type axis file (attributes_on) rather than a long row at all.

Notes

input_key property

input_key: tuple[str, ...]

Inputs-map key columns, compared NULL-safely when folding.

partial_dims, plus attribute. entity and a group's coordinates are in it as membership keys - a layer may patch one component's value, or one connection's, without restating every other's - and the broadcast partial value dims beside them.

A coordinate an attribute's own file does not carry reads as NULL, which is what makes the key one fixed tuple over attributes whose columns differ.

Notes

input_columns property

input_columns: tuple[str, ...]

The inputs map's full column set.

Notes

coordinates_of

coordinates_of(attribute: str) -> tuple[str, ...]

The dim columns one attribute's rows carry, groups expanded.

One rule resolves a name in dims: it is the dim of that name if one is declared, and otherwise the group of that name expanded to its coordinates. So dims={"connection", "snapshot"} gives ("entity", "bus", "snapshot") where no dim connection exists, and dims={"country"} gives ("country",) - the dim, where a group of that name is shadowed.

Per attribute rather than schema-wide: one file per attribute means one column set per attribute, and an all-NULL entity on a record-level weighting would be a column claiming a component the value has none of.

Notes
Source code in src/datarecord/schema.py
def coordinates_of(self, attribute: str) -> tuple[str, ...]:
    """The dim columns one attribute's rows carry, groups expanded.

    One rule resolves a name in `dims`: **it is the dim of that name if one
    is declared, and otherwise the group of that name expanded to its
    coordinates**. So `dims={"connection", "snapshot"}` gives `("entity",
    "bus", "snapshot")` where no dim `connection` exists, and
    `dims={"country"}` gives `("country",)` - the dim, where a group of that
    name is shadowed.

    Per attribute rather than schema-wide: one file per attribute means one
    column set per attribute, and an all-NULL `entity` on a record-level
    weighting would be a column claiming a component the value has none of.

    Notes
    -----
    - [groups](https://energy-models.github.io/datarecord/design/schema/#groups)
    - [addressing](https://energy-models.github.io/datarecord/design/schema/#addressing-dims-x)
    - [the long schema](https://energy-models.github.io/datarecord/design/format/#the-long-schema)
    """
    spec = self.spec_for(attribute)
    if spec is None:
        return ()
    named: set[str] = set()
    for d in spec.dims:
        group = None if d in self.dimensions else self.groups.get(d)
        named.update(group.coordinates if group is not None else (d,))
    # Declaration order, so every consumer sees one column order.
    return tuple(d for d in self.dims if d in named)

long_columns_for

long_columns_for(attribute: str) -> tuple[str, ...]

One attribute's full long column set, in order - input or result.

An attribute carries the coordinates its dims name and no others, so a record-level weighting has no entity column and a component attribute has no bus.

An attribute neither vocabulary declares is long_columns - every declared dim, the widest shape. That is a schema with no manifest yet, every declared attribute having its own coordinates.

Notes
Source code in src/datarecord/schema.py
def long_columns_for(self, attribute: str) -> tuple[str, ...]:
    """One attribute's full long column set, in order - input or result.

    An attribute carries the coordinates its `dims` name and no others, so a
    record-level weighting has no `entity` column and a component attribute
    has no `bus`.

    An attribute neither vocabulary declares is `long_columns` - every
    declared dim, the widest shape. That is a schema with no manifest yet,
    every declared attribute having its own coordinates.

    Notes
    -----
    - [the long schema](https://energy-models.github.io/datarecord/design/format/#the-long-schema)
    - [results](https://energy-models.github.io/datarecord/design/working-record/#results-through-kindoutputs)
    """
    if self.spec_for(attribute) is None:
        return self.long_columns
    return (*self.coordinates_of(attribute), *LONG_TAIL)

addresses_entity

addresses_entity(attribute: str) -> bool

Whether attribute reaches a component at all.

True where its dims name entity, or a group one of whose coordinates draws on entity. False for an attribute over an axis alone - a snapshot weighting belongs to the record, so no entity type carries it however few traits mention it.

Notes
Source code in src/datarecord/schema.py
def addresses_entity(self, attribute: str) -> bool:
    """Whether `attribute` reaches a component at all.

    True where its `dims` name `entity`, or a group one of whose
    coordinates draws on `entity`. False for an attribute over an axis
    alone - a snapshot weighting belongs to the record, so no entity type
    carries it however few traits mention it.

    Notes
    -----
    - [where a value lives](https://energy-models.github.io/datarecord/design/format/#where-a-value-lives)
    """
    return "entity" in self.coordinates_of(attribute)

attributes_for

attributes_for(ctype: str) -> dict[str, AttributeSpec]

Which attributes entity type ctype carries.

Every attribute addressed by entity that no trait narrows, plus those the traits naming ctype bundle. Untraited is carried by all: writing entity in an attribute's dims is what says it is per component, and declining to bundle it says it is so for every type - the same thing dims={"scenario"} already means along the scenario axis.

Empty for a label no declared entity-type axis lists, which is why callers rejecting an unknown type test entity_types rather than this. A schema declaring no entity type at all carries everything addressed by entity, whatever ctype is asked for.

Notes
Source code in src/datarecord/schema.py
def attributes_for(self, ctype: str) -> dict[str, AttributeSpec]:
    """Which attributes entity type `ctype` carries.

    Every attribute addressed by `entity` that no trait narrows, plus those
    the traits naming `ctype` bundle. Untraited is carried by all: writing
    `entity` in an attribute's `dims` is what says it is per component, and
    declining to bundle it says it is so for every type - the same thing
    `dims={"scenario"}` already means along the scenario axis.

    Empty for a label no declared entity-type axis lists, which is why
    callers rejecting an unknown type test `entity_types` rather than this.
    A schema declaring no entity type at all carries everything addressed
    by `entity`, whatever `ctype` is asked for.

    Notes
    -----
    - [traits](https://energy-models.github.io/datarecord/design/schema/#traits)
    """
    known = self.entity_types
    if known and ctype not in known:
        return {}
    narrowed: set[str] = set()
    names: set[str] = set()
    for trait in self.traits.values():
        # A trait with no `on` narrows nothing: it is a bundle to dispatch
        # on, so its attributes stay carried by every type.
        scoped = {ctype for labels in trait.on.values() for ctype in labels}
        if not scoped:
            continue
        narrowed |= trait.attributes
        if any(ctype in labels for labels in trait.on.values()):
            names |= trait.attributes
    names |= {
        a for a in self.attributes if a not in narrowed and self.addresses_entity(a)
    }
    return {a: self.attributes[a] for a in sorted(names)}

owned_per

owned_per(attribute: str) -> frozenset[str]

Which dims a layer owns attribute per.

Derived rather than declared: AttributeSpec.dims says which axes the attribute may vary over, partial_dims the fold key (membership keys plus the partial value dims), and ownership is their intersection. A dim in dims but not the fold key - a non-partial value axis like timestep - is owned whole, so a patch to one of its values restates the attribute's entire extent along it (_owned_whole).

Notes
Source code in src/datarecord/schema.py
def owned_per(self, attribute: str) -> frozenset[str]:
    """Which dims a layer owns `attribute` per.

    Derived rather than declared: `AttributeSpec.dims` says which axes the
    attribute may vary over, `partial_dims` the fold key (membership keys
    plus the `partial` value dims), and ownership is their intersection. A
    dim in `dims` but not the fold key - a non-`partial` value axis like
    `timestep` - is owned whole, so a patch to one of its values restates the
    attribute's entire extent along it (`_owned_whole`).

    Notes
    -----
    - [partial](https://energy-models.github.io/datarecord/design/schema/#partial-the-granularity-of-an-override)
    - [one fold for every axis](https://energy-models.github.io/datarecord/design/read-path/#one-fold-for-every-axis)
    """
    spec = self.attributes.get(attribute)
    if spec is None:
        return frozenset()
    return spec.dims & set(self.partial_dims)

axis_key

axis_key(dim: str) -> tuple[str, ...]

A dim's axis-table key: (*parents, dim), parents first.

Parents in declaration order, and transitively - a dim within another that is itself within a third is keyed by all three.

Notes
Source code in src/datarecord/schema.py
def axis_key(self, dim: str) -> tuple[str, ...]:
    """A dim's axis-table key: `(*parents, dim)`, parents first.

    Parents in declaration order, and transitively - a dim `within` another
    that is itself `within` a third is keyed by all three.

    Notes
    -----
    - [within](https://energy-models.github.io/datarecord/design/schema/#within-an-axis-inside-an-axis)
    """
    seen = _ancestors(dim, {d: s.within for d, s in self.dimensions.items()})
    return (*(d for d in self.dims if d in seen), dim)

attributes_on

attributes_on(dim: str) -> tuple[str, ...]

Attributes stored as columns of dims/{dim}.parquet.

An attribute addressed by dim alone: a per-country CO2 budget, a snapshot weighting, a per-type icon. AttributeSpec.varying is False for exactly these, and this is the axis-side counterpart of addresses_entity - what dims/entity_type/<Type>.parquet is to a component's constant columns, the axis file is to these.

entity is one of these axes only where no group declares the type axis: with no type to classify a component into there is no member file for its constant columns, so they live on dims/entity.parquet like any other axis's (entity_type_dim). Where a group does declare the axis this returns () for entity - the columns are the component frame's, dims/entity_type/<Type>.parquet, a different destination with a different key.

Keyed off dims rather than coordinates_of, because a group with one coordinate is indistinguishable there: dims={"connection"} over a single bus coordinate also yields ("bus",), and it belongs in the group's file rather than on the bus axis. A group over entity alone is keyed by the group name, not entity, so its into label and any attribute it bundles never match here.

Notes
Source code in src/datarecord/schema.py
def attributes_on(self, dim: str) -> tuple[str, ...]:
    """Attributes stored as columns of `dims/{dim}.parquet`.

    An attribute addressed by `dim` alone: a per-country CO2 budget, a
    snapshot weighting, a per-type icon. `AttributeSpec.varying` is False
    for exactly these, and this is the axis-side counterpart of
    `addresses_entity` - what `dims/entity_type/<Type>.parquet` is to a
    component's constant columns, the axis file is to these.

    `entity` is one of these axes only where no group declares the type
    axis: with no type to classify a component into there is no member file
    for its constant columns, so they live on `dims/entity.parquet` like any
    other axis's (`entity_type_dim`). Where a group *does* declare the axis
    this returns `()` for `entity` - the columns are the *component* frame's,
    `dims/entity_type/<Type>.parquet`, a different destination with a
    different key.

    Keyed off `dims` rather than `coordinates_of`, because a group with one
    coordinate is indistinguishable there: `dims={"connection"}` over a
    single `bus` coordinate also yields `("bus",)`, and it belongs in the
    group's file rather than on the bus axis. A group over `entity` alone is
    keyed by the group name, not `entity`, so its `into` label and any
    attribute it bundles never match here.

    Notes
    -----
    - [where a value lives](https://energy-models.github.io/datarecord/design/format/#where-a-value-lives)
    - [entity types](https://energy-models.github.io/datarecord/design/schema/#entity_type-the-axis-of-kinds)
    """
    if dim not in self.dimensions:
        return ()
    if dim == "entity" and self.entity_type_dim is not None:
        return ()
    return tuple(
        a for a, spec in self.attributes.items() if spec.dims == frozenset({dim})
    )

groups_of

groups_of(attribute: str) -> tuple[str, ...]

Which declared groups address attribute, in declaration order.

An attribute is a connection attribute because its dims name the connection group - not because a separate field says so. That is what lets a second group exist without a second field.

A group a dim shadows is not one of them, dims: [country] naming the axis.

Notes
Source code in src/datarecord/schema.py
def groups_of(self, attribute: str) -> tuple[str, ...]:
    """Which declared groups address `attribute`, in declaration order.

    An attribute is a connection attribute because its `dims` name the
    `connection` group - not because a separate field says so. That is
    what lets a second group exist without a second field.

    A group a dim shadows is not one of them, `dims: [country]` naming the
    axis.

    Notes
    -----
    - [addressing](https://energy-models.github.io/datarecord/design/schema/#addressing-dims-x)
    """
    spec = self.attributes.get(attribute)
    if spec is None:
        return ()
    return tuple(
        g for g in self.groups if g in spec.dims and g not in self.dimensions
    )

group_coordinates

group_coordinates(group: str) -> tuple[str, ...]

One group's columns, or () if it is not declared.

Every column of the group's file, into included. Coordinate names rather than dim names, so two drawing on one axis stay two columns.

Notes
Source code in src/datarecord/schema.py
def group_coordinates(self, group: str) -> tuple[str, ...]:
    """One group's columns, or `()` if it is not declared.

    Every column of the group's file, `into` included. Coordinate names
    rather than dim names, so two drawing on one axis stay two columns.

    Notes
    -----
    - [groups](https://energy-models.github.io/datarecord/design/schema/#groups)
    """
    spec = self.groups.get(group)
    return () if spec is None else spec.coordinates

group_key

group_key(group: str) -> tuple[str, ...]

One group's key columns, or () if it is not declared.

group_coordinates minus into - what the fold keys ownership by and what a tombstone names.

Notes
Source code in src/datarecord/schema.py
def group_key(self, group: str) -> tuple[str, ...]:
    """One group's key columns, or `()` if it is not declared.

    `group_coordinates` minus `into` - what the fold keys ownership by and
    what a tombstone names.

    Notes
    -----
    - [into](https://energy-models.github.io/datarecord/design/schema/#into-a-group-that-classifies)
    - [the owner map](https://energy-models.github.io/datarecord/design/read-path/#owner-map)
    """
    spec = self.groups.get(group)
    return () if spec is None else spec.key

column_type

column_type(column: str) -> DType | None

The declared type for one column, or None if the schema declares none.

Covers the structural columns the format fixes, the declared dims, the attributes an axis file carries as columns (attributes_on), and the owner map's two flag structs, whose fields follow the schema's dims. A narwhals dtype, translated to DuckDB (duck.DuckTypes) only where a caller builds a column of it.

No dim is structural - entity and a group's bus included: each is declared, and typed from that declaration. So an Enum on the entity-type axis pins its vocabulary everywhere the column is built, and an axis a schema happens to call kind is typed no differently.

An attribute addressed by one axis alone is a column rather than a value cell, so this is where its type is read from - cast_declared would otherwise leave an axis file's attribute column as whatever the incoming frame happened to carry. An attribute with any other dims is value_type's, not this: it is a long row's value.

A schema declaring no dims at all is "no manifest yet" rather than a record to fold, and DuckDB has no empty struct - so the flag columns are undeclared there, and a caller building an empty relation falls back to VARCHAR for a map that will never hold a row.

Notes
Source code in src/datarecord/schema.py
def column_type(self, column: str) -> nw.dtypes.DType | None:
    """The declared type for one column, or None if the schema declares none.

    Covers the structural columns the format fixes, the declared dims, the
    attributes an axis file carries as columns (`attributes_on`), and the
    owner map's two flag structs, whose fields follow the schema's dims. A
    narwhals dtype, translated to DuckDB (`duck.DuckTypes`) only where a
    caller builds a column of it.

    No dim is structural - `entity` and a group's `bus` included: each is
    declared, and typed from that declaration. So an `Enum` on the entity-type
    axis pins its vocabulary everywhere the column is built, and an axis a
    schema happens to call `kind` is typed no differently.

    An attribute addressed by one axis alone is a *column* rather than a
    `value` cell, so this is where its type is read from - `cast_declared`
    would otherwise leave an axis file's attribute column as whatever the
    incoming frame happened to carry. An attribute with any other `dims` is
    `value_type`'s, not this: it is a long row's value.

    A schema declaring no dims at all is "no manifest yet" rather
    than a record to fold, and DuckDB has no empty struct - so the flag
    columns are undeclared there, and a caller building an empty relation
    falls back to `VARCHAR` for a map that will never hold a row.

    Notes
    -----
    - [one schema per record](https://energy-models.github.io/datarecord/design/schema/#one-schema-per-record)
    - [the owner map](https://energy-models.github.io/datarecord/design/read-path/#owner-map)
    """
    if column in STRUCTURAL_TYPES:
        return STRUCTURAL_TYPES[column]
    if column in self.dimensions:
        return self.dimensions[column].dtype
    if column in ("varies", "broadcast"):
        return flag_type(self.broadcast_dims) if self.broadcast_dims else None
    spec = self.attributes.get(column)
    if spec is not None and not spec.varying:
        (dim,) = spec.dims
        if column in self.attributes_on(dim):
            return spec.dtype
    return None

spec_for

spec_for(attribute: str) -> AttributeSpec | None

attribute's spec, whether it is an input or a result.

The one lookup that spans both vocabularies, for the questions the long schema asks of a stored attribute regardless of which file holds it - its dtype and its coordinates. Anything governing how an attribute may be written asks attributes or results directly, the two differing exactly there.

Notes
Source code in src/datarecord/schema.py
def spec_for(self, attribute: str) -> AttributeSpec | None:
    """`attribute`'s spec, whether it is an input or a result.

    The one lookup that spans both vocabularies, for the questions the long
    schema asks of a stored attribute regardless of which file holds it -
    its dtype and its coordinates. Anything governing how an attribute may
    be *written* asks `attributes` or `results` directly, the two differing
    exactly there.

    Notes
    -----
    - [the long schema](https://energy-models.github.io/datarecord/design/format/#the-long-schema)
    - [outputs](https://energy-models.github.io/datarecord/design/read-path/#outputs)
    """
    return self.attributes.get(attribute) or self.results.get(attribute)

value_type

value_type(attribute: str) -> DType | None

The value column's type for one attribute, input or result.

No ctype: one attribute is one <kind>/<attr>.parquet with one value column, so the dtype is the attribute's alone. A narwhals dtype, translated to DuckDB (duck.DuckTypes) only where a caller builds a column of it.

Notes
Source code in src/datarecord/schema.py
def value_type(self, attribute: str) -> nw.dtypes.DType | None:
    """The `value` column's type for one attribute, input or result.

    No `ctype`: one attribute is one `<kind>/<attr>.parquet` with one
    `value` column, so the dtype is the attribute's alone. A narwhals
    dtype, translated to DuckDB (`duck.DuckTypes`) only where a caller builds
    a column of it.

    Notes
    -----
    - [the long schema](https://energy-models.github.io/datarecord/design/format/#the-long-schema)
    """
    spec = self.spec_for(attribute)
    return None if spec is None else spec.dtype

types_declaring

types_declaring(attribute: str) -> frozenset[str]

Which entity types carry attribute - what names=None targets.

Empty for a record-level attribute, which no type carries and which therefore targets no names at all, and empty too for a schema declaring no entity-type labels, where the caller has no type vocabulary to enumerate and works from the resolved components instead.

Notes
Source code in src/datarecord/schema.py
def types_declaring(self, attribute: str) -> frozenset[str]:
    """Which entity types carry `attribute` - what `names=None` targets.

    Empty for a record-level attribute, which no type carries and which
    therefore targets no names at all, and empty too for a schema declaring
    no entity-type labels, where the caller has no type vocabulary to
    enumerate and works from the resolved components instead.

    Notes
    -----
    - [set](https://energy-models.github.io/datarecord/design/working-record/#set)
    """
    return frozenset(
        c for c in self.entity_types if attribute in self.attributes_for(c)
    )

compatible_with

compatible_with(other: Schema) -> list[str]

Why layers written under other would not read under self.

Empty when the change is compatible: old layers stay readable and only version moves. The compatible changes are those where NULL already means what the new schema needs it to mean, so the broadcast rule absorbs them without touching a row.

Returns:

Type Description
list of str

One reason per incompatibility, empty if there are none.

Notes
Source code in src/datarecord/schema.py
def compatible_with(self, other: Schema) -> list[str]:
    """Why layers written under `other` would not read under `self`.

    Empty when the change is compatible: old layers stay readable and only
    `version` moves. The compatible changes are those where NULL already
    means what the new schema needs it to mean, so the broadcast rule absorbs
    them without touching a row.

    Returns
    -------
    list of str
        One reason per incompatibility, empty if there are none.

    Notes
    -----
    - [the broadcast rule](https://energy-models.github.io/datarecord/design/record/#the-broadcast-rule)
    - [versioning](https://energy-models.github.io/datarecord/design/schema/#versioning)
    """
    problems = []

    for dim, was in other.dimensions.items():
        now = self.dimensions.get(dim)
        if now is None:
            problems.append(f"dim {dim!r} removed")
            continue
        if now.dtype != was.dtype:
            problems.append(f"dim {dim!r} dtype {was.dtype} -> {now.dtype}")
        if now.within != was.within:
            problems.append(
                f"dim {dim!r} nesting changed; the axis key changes shape"
            )

    # Results version like inputs: a layer's `outputs/<attr>.parquet` is
    # unreadable for the same reasons its `inputs/` counterpart would be.
    for kind, mine, theirs in (
        ("attribute", self.attributes, other.attributes),
        ("result", self.results, other.results),
    ):
        for attr, was_spec in theirs.items():
            now_spec = mine.get(attr)
            if now_spec is None:
                problems.append(f"{kind} {attr!r} removed")
                continue
            if now_spec.dtype != was_spec.dtype:
                problems.append(
                    f"{kind} {attr!r} dtype {was_spec.dtype} -> {now_spec.dtype}"
                )
            narrowed = was_spec.dims - now_spec.dims
            if narrowed:
                problems.append(
                    f"{kind} {attr!r} no longer varies over {sorted(narrowed)}; "
                    f"rows setting those dims have no valid reading"
                )

    # A type losing an attribute is incompatible for the same reason a
    # narrowed `dims` is: its rows are still in the file, now unreadable
    # for that type. Losing a whole type says the same of all of them.
    for ctype in other.entity_types:
        was_attrs = set(other.attributes_for(ctype))
        dropped = sorted(was_attrs - set(self.attributes_for(ctype)))
        if dropped:
            problems.append(
                f"component type {ctype!r} no longer carries {dropped}; "
                f"rows written for it have no valid reading"
            )

    if other.partial is not None and self.partial is not None:
        lost = other.partial - self.partial
        if lost:
            problems.append(
                f"{sorted(lost)} no longer `partial`; a layer that patched one "
                f"value along such an axis is now a partial override of an axis "
                f"owned whole"
            )
    return problems

Dimension

Bases: BaseModel

One axis attribute data may vary over: its shape, not its data.

Not which dims an attribute varies over (AttributeSpec.dims), nor the patch granularity (Schema.partial), nor order - an axis is ordered by its file's row order, undeclared.

Attributes:

Name Type Description
dtype DType

The axis labels' type, as a narwhals dtype instance (nw.String(), nw.Datetime(), ...) - translated to its DuckDB name only where a column of it is built.

within frozenset[str]

Dims this one's labels identify a point only within; transitive.

unit str | None

What this axis's labels measure, if anything - None is undeclared, "" genuinely dimensionless.

description str | None

What the axis is, in prose. Never interpreted.

Notes

AttributeSpec

Bases: BaseModel

What shape one attribute's data may take.

Attributes:

Name Type Description
dtype DType

The value column's type, as a narwhals dtype instance (nw.String(), nw.Datetime(), ...) - translated to its DuckDB name only where a column of it is built.

dims frozenset[str]

Dims this attribute may vary over; a subset of those declared. Varying over nothing is what puts it in dims/entity_type/<Type>.parquet rather than inputs/, so the schema decides the file split.

default Any | None

The value a coordinate no row covers takes.

breakpoints bool

Whether it may carry a piecewise-linear curve.

unit str | None

What the values measure - "MW", "EUR/MWh". Stored and never interpreted; None is undeclared, "" genuinely dimensionless.

description str | None

What the attribute is, in prose. Never interpreted.

Notes

varying property

varying: bool

Whether this attribute's values are long rows rather than a column.

"Varies beyond its address", not "has dims": naming exactly one addressing coordinate is a column on that thing's own table, so dims={"entity"} is a component column and dims={"connection"} a column of the group's table. Anything more is inputs/<attr>.parquet.

A bare bool(dims) was the test before entity was a declared dim, when a component attribute declared none - it would now call every attribute varying and route every constant to inputs/.

Notes