Skip to content

Record

Record dataclass

Record(resolver: Resolver)

A record's resolved view, as narwhals frames.

The narwhals interface over one Resolver, not a second implementation of it: a member costs what the equivalent Resolver call costs, and flags in particular is free, the owner map having folded it in.

One class for every backing, because a fold over one source degenerates to a scan of it - so a plain parquet directory is read as the one layer it is (at) rather than by a second code path. RecordLike is the protocol a caller annotates against; this is the class a caller constructs.

Notes

over classmethod

over(
    *sources: LayerSource, con: DuckDBPyConnection
) -> Record

A record folded over sources, root first.

The last source's layer_id names the record, which for a layer tree is the revision being resolved, and its schema is the record's - every source in one fold carries the same one, resolved by whoever built them. con is explicit because a source is only obliged to hand over rows - where it reads them is its own business, and the protocol carries no connection.

Source code in src/datarecord/layered/revision.py
@classmethod
def over(cls, *sources: LayerSource, con: DuckDBPyConnection) -> Record:
    """A record folded over `sources`, root first.

    The last source's `layer_id` names the record, which for a layer tree is
    the revision being resolved, and its `schema` is the record's - every
    source in one fold carries the same one, resolved by whoever built them.
    `con` is explicit because a source is only obliged to hand over rows -
    where it reads them is its own business, and the protocol carries no
    connection.
    """
    if not sources:
        msg = "a `Record` needs at least one source to fold"
        raise ValueError(msg)
    return cls(
        Resolver(sources[-1].layer_id, list(sources), con, sources[-1].schema)
    )

at classmethod

at(
    uri: str, con: DuckDBPyConnection | None = None
) -> Record

A plain parquet directory, read as the one layer it is.

Any directory in the layer layout - a single layer of a tree, or a standalone record write_record produced. It needs no revisions row and no tree: being a layer layout is what the fold requires, and that is what a record directory is. Over one source the fold degenerates to a scan, so this costs a scan and not a resolution.

A standalone record carries its own manifest.json, being one whole record rather than a layer of one, and that is what it is read under - so such a directory reads the same through any connection. A single layer directory has none, and is read under the connection's root.

Notes
Source code in src/datarecord/layered/revision.py
@classmethod
def at(cls, uri: str, con: DuckDBPyConnection | None = None) -> Record:
    """A plain parquet directory, read as the one layer it is.

    Any directory in the layer layout - a single layer of a tree, or a
    standalone record `write_record` produced. It needs no `revisions` row
    and no tree: being a layer layout is what the fold requires, and that is
    what a record directory is. Over one source the fold degenerates to a
    scan, so this costs a scan and not a resolution.

    A **standalone** record carries its own `manifest.json`, being one whole
    record rather than a layer of one, and that is what it is read under -
    so such a directory reads the same through any connection. A single
    layer directory has none, and is read under the connection's root.

    Notes
    -----
    - [the record format](https://energy-models.github.io/datarecord/design/format/)
    - [one schema per record](https://energy-models.github.io/datarecord/design/schema/#one-schema-per-record)
    """
    con = con or default_connection()
    # The manifest read needs only the directory URI, not a built source: a
    # standalone record declares its own schema and a bare layer reads the
    # connection root's, and either way the schema must be in hand before the
    # source that carries it is built.
    base = uri if uri.endswith("/") else uri + "/"
    raw = read_json(base + "manifest.json")
    schema = resolve.read_schema(con) if raw is None else Schema.model_validate(raw)
    return cls.over(DirectorySource(base, schema, con), con=con)

groups

groups() -> LazyFrames

Each declared group's rows, keyed by group - one frame each.

Only groups some layer wrote a row of; a declared group nothing populates is absent rather than present-and-empty.

Notes
Source code in src/datarecord/layered/revision.py
@_stable_cache
def groups(self) -> LazyFrames:
    """Each declared group's rows, keyed by group - one frame each.

    Only groups some layer wrote a row of; a declared group nothing
    populates is absent rather than present-and-empty.

    Notes
    -----
    - [groups](https://energy-models.github.io/datarecord/design/schema/#groups)
    """
    groups = tuple(
        g for g in self.resolver.schema.groups if self.resolver.group(g) is not None
    )
    return LazyFrames(groups, self._group_frame)

flags

flags(ctype: str) -> dict[str, Flags]

Straight off the inputs owner map, which folded these in for free.

Notes
Source code in src/datarecord/layered/revision.py
def flags(self, ctype: str) -> dict[str, Flags]:
    """Straight off the `inputs` owner map, which folded these in for free.

    Notes
    -----
    - [the owner map](https://energy-models.github.io/datarecord/design/read-path/#owner-map)
    """
    return self.resolver.attributes_of(ctype)

RecordLike

Bases: Protocol

What a record answers, however it is backed, as narwhals frames.

Read-only: writing is write_record(revision_id, source, con), which takes a LayerData rather than this - a framework's own RecordLike reaches it through the adapter write_record wraps one in.

Notes

schema property

schema: Schema

The record's schema: its dims, its attributes, its patch granularity.

Notes

dims property

dims: Frames

Axis frames, keyed by dim ("scenario" -> dims/scenarios.parquet).

An axis frame is its key column and the attributes addressed by it alone (Schema.attributes_on) - so a per-country CO2 budget or a per-type icon is read from here rather than from attributes, which holds long frames only. A column absent from the frame is one no layer wrote, whose value is that attribute's default.

No classification column: which buses a country holds is the group into it, read from groups.

Notes

entity_types property

entity_types: Frames

Wide member frames, keyed by component type, in member order.

Notes

groups property

groups: Frames

Each declared group's rows, keyed by group - one frame each.

A group declares which tuples over several dims exist - connection over (entity, bus) is the one every record with connections has, and it is one instance rather than a member of its own.

Not split by component type, which is no coordinate of a group.

Notes

attributes property

attributes: Frames

Long input frames, keyed by attribute name - one per file.

Not by component type: one inputs/p_max_pu.parquet holds every type's rows, keyed by entity alone. A row carries no entity_type - entities are unique across every type - so a reader wanting one type joins entity_types on name.

Notes

outputs property

outputs: Frames

Long result frames, keyed by attribute name.

Empty for a record carrying no results.

Unlike its neighbours, this does not overlay on a layered record: a record's results are its own layer's, never a resolution over its ancestors'.

Notes

flags

flags(ctype: str) -> dict[str, Flags]

Every attribute of ctype, mapped to the shape its rows take.

Only attributes with rows are present, so the key set also answers which attributes this type has at all.

Notes
Source code in src/datarecord/record.py
def flags(self, ctype: str) -> dict[str, Flags]:
    """Every attribute of `ctype`, mapped to the shape its rows take.

    Only attributes with rows are present, so the key set also answers
    which attributes this type has at all.

    Notes
    -----
    - [Flags](https://energy-models.github.io/datarecord/design/record/#flags)
    """
    ...

Frames module-attribute

Frames = Mapping[str, 'nw.LazyFrame']

What a Record hands over: named frames, each an unmaterialised plan.

The Mapping ABC, so a plain dict satisfies it as fully as LazyFrames does. A DuckDB-backed record reaches it with nw.from_native(rel), which stays an unexecuted plan.

Notes

LazyFrames

LazyFrames(
    keys: tuple[str, ...], build: Callable[[str], LazyFrame]
)

Bases: Mapping[str, 'nw.LazyFrame']

A Frames whose values are built on __getitem__, not up front.

For a backing where building a frame is itself I/O: read_parquet reads the footer to bind the schema, a round trip per file against a remote record.

Parameters:

Name Type Description Default
keys tuple[str, ...]

Every key this mapping holds, in iteration order. Listing them must be cheap: nothing here builds a frame.

required
build Callable[[str], LazyFrame]

Called with one key to produce its frame, once per __getitem__; memoise inside build if repeated lookups matter.

required
Notes
Source code in src/datarecord/record.py
def __init__(self, keys: tuple[str, ...], build: Callable[[str], nw.LazyFrame]):
    self._keys = keys
    self._build = build

Flags dataclass

Flags(
    varies: frozenset[str],
    broadcast: frozenset[str],
    breakpoints: bool = False,
)

Which axes an attribute's rows actually use, for one component type.

The two sets are not complements: a dim in both means this type's components disagree, which is the instruction to use both containers rather than an ambiguity to resolve. varies | broadcast is the test for whether an attribute touches a dim at all.

Parameters:

Name Type Description Default
varies frozenset[str]

Dims some row of this attribute sets.

required
broadcast frozenset[str]

Dims some row leaves NULL, i.e. "all values of that dim".

required
breakpoints bool

Whether any row carries a breakpoint. Not a dim.

False
Notes