Schema¶
Schema
¶
Bases: BaseModel
One record's schema.
Attributes:
| Name | Type | Description |
|---|---|---|
version |
int
|
Bumped by any change to the declarations. A reader meeting a version it was not written for should refuse rather than guess. |
dimensions |
dict[str, Dimension]
|
Every declared axis, keyed by dim name. |
attributes |
dict[str, AttributeSpec]
|
Attribute -> spec, flat and record-wide. One attribute is one spec and
one |
groups |
dict[str, Group]
|
Group name -> which tuples over several dims exist. |
traits |
dict[str, Trait]
|
Trait -> the attributes it bundles and the entity types carrying them. A vocabulary a consumer dispatches on, declared rather than derived, and the only thing that narrows an attribute to some entity types. |
partial |
frozenset[str] | None
|
Which dims a layer may patch value by value. |
meta |
dict[str, Any]
|
A framework's own top-level data - network attributes, CRS, free-form metadata. Stored and never interpreted, since none of it describes the dimensioned data. |
Notes
results
class-attribute
instance-attribute
¶
results: dict[str, AttributeSpec] = Field(
default_factory=dict
)
What a solve computes, keyed like attributes and shaped the same.
Separate because the two are governed differently, not because they are
stored differently: a result is written to outputs/<attr>.parquet rather
than inputs/, never overlays a parent's, and may name a component the
record does not declare. Keeping it out of attributes is what stops it
reaching attributes_for, and so add's wide-frame split and the input
validation, neither of which a result should meet.
dims
property
¶
Every declared dim, in declaration order - the long schema's dim columns.
Notes
broadcast_dims
property
¶
The dims a NULL broadcasts over: every dim but entity, the type axis and a group's key.
A NULL here means "every value of this dim", which the fold expands against the axis. The three exclusions cannot mean that:
entity, because a NULL there is a value belonging to no component rather than to all of them. The one dim named literally, being the one every entity-type axis classifies.- The entity-type axis, because it inherits
entity's exclusion: its labels are a column ofdims/entity.parquet, so a NULL there is a component whose type is unknown rather than one of every type. An attribute addressed by the type alone never reaches here: it is a column of the type axis file rather than a long row, so it has no NULL to expand. - A coordinate a group is keyed by, because there is no axis to expand against. "Every bus of this component" is the group's rows, not the bus axis - a sparse subset only the group's table knows.
A functional group's into dim is not excluded, though it is one of
the group's coordinates: only key addresses a row, so country
broadcasts like any other axis.
The complement of this is what Schema requires to be partial: a dim
whose values are addressed individually is one a layer patches value by
value.
What the varies/broadcast structs have a field per, and what
expand_dims joins.
Notes
entity_type_dim
property
¶
The dim classifying entity, or None where none does.
The into of the group over entity alone, of which the schema admits
at most one.
Notes
entity_types
property
¶
Every declared entity-type label - the types a component may be.
The entity-type axis's enum categories, so a schema declaring it as a
plain String has none: the labels are then data rather than
declarations, and attributes_for accepts any of them.
Notes
membership_keys
property
¶
The dims addressed per row rather than broadcast: entity and group coords.
A membership key is a coordinate a layer patches one row of at a time -
one component, one connection - never "every value" of an axis. It is
the non-broadcast, addressable dims: every dim but the broadcast ones
and the entity-type axis, which is a column of dims/entity.parquet
rather than an addressable coordinate.
These land in the fold key by being membership, not by being partial.
partial_dims
property
¶
The fold key's dims, in declaration order.
The membership keys plus the broadcast value dims a layer may patch per
value (partial). The fold's key is one fixed tuple over all
attributes, so it carries every axis any layer may patch by value or
by row, not only those some currently declared attribute varies over. An
attribute not owned per one of them writes NULL there, the "NULL means
all values" rule - which also lets a schema declare an axis before any
attribute uses it.
long_columns
property
¶
The long schema's full column set.
The map's column set, which is uniform across attributes because the
map is one relation over all of them. An individual file carries only
its own attribute's columns (long_columns_for), and union_by_name
supplies NULL for the rest when the fold unions them here.
No entity_type: a row here is keyed by entity, and an entity is unique
record-wide, so a type column would restate what the entity already says
and let the two disagree. Every declared attribute is shaped by
long_columns_for rather than by this, where naming both is rejected
outright and one addressed by the type alone is a column of the type
axis file (attributes_on) rather than a long row at all.
input_key
property
¶
Inputs-map key columns, compared NULL-safely when folding.
partial_dims, plus attribute. entity and a group's coordinates are
in it as membership keys - a layer may patch one component's value, or
one connection's, without restating every other's - and the broadcast
partial value dims beside them.
A coordinate an attribute's own file does not carry reads as NULL, which is what makes the key one fixed tuple over attributes whose columns differ.
Notes
input_columns
property
¶
The inputs map's full column set.
Notes
coordinates_of
¶
The dim columns one attribute's rows carry, groups expanded.
One rule resolves a name in dims: it is the dim of that name if one
is declared, and otherwise the group of that name expanded to its
coordinates. So dims={"connection", "snapshot"} gives ("entity",
"bus", "snapshot") where no dim connection exists, and
dims={"country"} gives ("country",) - the dim, where a group of that
name is shadowed.
Per attribute rather than schema-wide: one file per attribute means one
column set per attribute, and an all-NULL entity on a record-level
weighting would be a column claiming a component the value has none of.
Notes
Source code in src/datarecord/schema.py
long_columns_for
¶
One attribute's full long column set, in order - input or result.
An attribute carries the coordinates its dims name and no others, so a
record-level weighting has no entity column and a component attribute
has no bus.
An attribute neither vocabulary declares is long_columns - every
declared dim, the widest shape. That is a schema with no manifest yet,
every declared attribute having its own coordinates.
Notes
Source code in src/datarecord/schema.py
addresses_entity
¶
Whether attribute reaches a component at all.
True where its dims name entity, or a group one of whose
coordinates draws on entity. False for an attribute over an axis
alone - a snapshot weighting belongs to the record, so no entity type
carries it however few traits mention it.
Notes
Source code in src/datarecord/schema.py
attributes_for
¶
attributes_for(ctype: str) -> dict[str, AttributeSpec]
Which attributes entity type ctype carries.
Every attribute addressed by entity that no trait narrows, plus those
the traits naming ctype bundle. Untraited is carried by all: writing
entity in an attribute's dims is what says it is per component, and
declining to bundle it says it is so for every type - the same thing
dims={"scenario"} already means along the scenario axis.
Empty for a label no declared entity-type axis lists, which is why
callers rejecting an unknown type test entity_types rather than this.
A schema declaring no entity type at all carries everything addressed
by entity, whatever ctype is asked for.
Notes
Source code in src/datarecord/schema.py
owned_per
¶
Which dims a layer owns attribute per.
Derived rather than declared: AttributeSpec.dims says which axes the
attribute may vary over, partial_dims the fold key (membership keys
plus the partial value dims), and ownership is their intersection. A
dim in dims but not the fold key - a non-partial value axis like
timestep - is owned whole, so a patch to one of its values restates the
attribute's entire extent along it (_owned_whole).
Source code in src/datarecord/schema.py
axis_key
¶
A dim's axis-table key: (*parents, dim), parents first.
Parents in declaration order, and transitively - a dim within another
that is itself within a third is keyed by all three.
Notes
Source code in src/datarecord/schema.py
attributes_on
¶
Attributes stored as columns of dims/{dim}.parquet.
An attribute addressed by dim alone: a per-country CO2 budget, a
snapshot weighting, a per-type icon. AttributeSpec.varying is False
for exactly these, and this is the axis-side counterpart of
addresses_entity - what dims/entity_type/<Type>.parquet is to a
component's constant columns, the axis file is to these.
entity is one of these axes only where no group declares the type
axis: with no type to classify a component into there is no member file
for its constant columns, so they live on dims/entity.parquet like any
other axis's (entity_type_dim). Where a group does declare the axis
this returns () for entity - the columns are the component frame's,
dims/entity_type/<Type>.parquet, a different destination with a
different key.
Keyed off dims rather than coordinates_of, because a group with one
coordinate is indistinguishable there: dims={"connection"} over a
single bus coordinate also yields ("bus",), and it belongs in the
group's file rather than on the bus axis. A group over entity alone is
keyed by the group name, not entity, so its into label and any
attribute it bundles never match here.
Source code in src/datarecord/schema.py
groups_of
¶
Which declared groups address attribute, in declaration order.
An attribute is a connection attribute because its dims name the
connection group - not because a separate field says so. That is
what lets a second group exist without a second field.
A group a dim shadows is not one of them, dims: [country] naming the
axis.
Notes
Source code in src/datarecord/schema.py
group_coordinates
¶
One group's columns, or () if it is not declared.
Every column of the group's file, into included. Coordinate names
rather than dim names, so two drawing on one axis stay two columns.
Notes
Source code in src/datarecord/schema.py
group_key
¶
One group's key columns, or () if it is not declared.
group_coordinates minus into - what the fold keys ownership by and
what a tombstone names.
Notes
Source code in src/datarecord/schema.py
column_type
¶
The declared type for one column, or None if the schema declares none.
Covers the structural columns the format fixes, the declared dims, the
attributes an axis file carries as columns (attributes_on), and the
owner map's two flag structs, whose fields follow the schema's dims. A
narwhals dtype, translated to DuckDB (duck.DuckTypes) only where a
caller builds a column of it.
No dim is structural - entity and a group's bus included: each is
declared, and typed from that declaration. So an Enum on the entity-type
axis pins its vocabulary everywhere the column is built, and an axis a
schema happens to call kind is typed no differently.
An attribute addressed by one axis alone is a column rather than a
value cell, so this is where its type is read from - cast_declared
would otherwise leave an axis file's attribute column as whatever the
incoming frame happened to carry. An attribute with any other dims is
value_type's, not this: it is a long row's value.
A schema declaring no dims at all is "no manifest yet" rather
than a record to fold, and DuckDB has no empty struct - so the flag
columns are undeclared there, and a caller building an empty relation
falls back to VARCHAR for a map that will never hold a row.
Source code in src/datarecord/schema.py
spec_for
¶
spec_for(attribute: str) -> AttributeSpec | None
attribute's spec, whether it is an input or a result.
The one lookup that spans both vocabularies, for the questions the long
schema asks of a stored attribute regardless of which file holds it -
its dtype and its coordinates. Anything governing how an attribute may
be written asks attributes or results directly, the two differing
exactly there.
Notes
Source code in src/datarecord/schema.py
value_type
¶
The value column's type for one attribute, input or result.
No ctype: one attribute is one <kind>/<attr>.parquet with one
value column, so the dtype is the attribute's alone. A narwhals
dtype, translated to DuckDB (duck.DuckTypes) only where a caller builds
a column of it.
Notes
Source code in src/datarecord/schema.py
types_declaring
¶
Which entity types carry attribute - what names=None targets.
Empty for a record-level attribute, which no type carries and which therefore targets no names at all, and empty too for a schema declaring no entity-type labels, where the caller has no type vocabulary to enumerate and works from the resolved components instead.
Notes
Source code in src/datarecord/schema.py
compatible_with
¶
compatible_with(other: Schema) -> list[str]
Why layers written under other would not read under self.
Empty when the change is compatible: old layers stay readable and only
version moves. The compatible changes are those where NULL already
means what the new schema needs it to mean, so the broadcast rule absorbs
them without touching a row.
Returns:
| Type | Description |
|---|---|
list of str
|
One reason per incompatibility, empty if there are none. |
Notes
Source code in src/datarecord/schema.py
1135 1136 1137 1138 1139 1140 1141 1142 1143 1144 1145 1146 1147 1148 1149 1150 1151 1152 1153 1154 1155 1156 1157 1158 1159 1160 1161 1162 1163 1164 1165 1166 1167 1168 1169 1170 1171 1172 1173 1174 1175 1176 1177 1178 1179 1180 1181 1182 1183 1184 1185 1186 1187 1188 1189 1190 1191 1192 1193 1194 1195 1196 1197 1198 1199 1200 1201 1202 1203 1204 1205 1206 1207 1208 1209 | |
Dimension
¶
Bases: BaseModel
One axis attribute data may vary over: its shape, not its data.
Not which dims an attribute varies over (AttributeSpec.dims), nor the
patch granularity (Schema.partial), nor order - an axis is ordered by its
file's row order, undeclared.
Attributes:
| Name | Type | Description |
|---|---|---|
dtype |
DType
|
The axis labels' type, as a narwhals dtype instance ( |
within |
frozenset[str]
|
Dims this one's labels identify a point only within; transitive. |
unit |
str | None
|
What this axis's labels measure, if anything - |
description |
str | None
|
What the axis is, in prose. Never interpreted. |
AttributeSpec
¶
Bases: BaseModel
What shape one attribute's data may take.
Attributes:
| Name | Type | Description |
|---|---|---|
dtype |
DType
|
The value column's type, as a narwhals dtype instance ( |
dims |
frozenset[str]
|
Dims this attribute may vary over; a subset of those declared. Varying
over nothing is what puts it in |
default |
Any | None
|
The value a coordinate no row covers takes. |
breakpoints |
bool
|
Whether it may carry a piecewise-linear curve. |
unit |
str | None
|
What the values measure - |
description |
str | None
|
What the attribute is, in prose. Never interpreted. |
Notes
varying
property
¶
Whether this attribute's values are long rows rather than a column.
"Varies beyond its address", not "has dims": naming exactly one
addressing coordinate is a column on that thing's own table, so
dims={"entity"} is a component column and dims={"connection"} a
column of the group's table. Anything more is inputs/<attr>.parquet.
A bare bool(dims) was the test before entity was a declared dim,
when a component attribute declared none - it would now call every
attribute varying and route every constant to inputs/.