Layered¶
Revision
¶
Bases: BaseModel
A node in the tree of layers.
Each revision adds one parquet layer on top of its parent's; the record it
resolves to is that layer over its ancestors'. The layer location
derives from id via layer_dir, it is not stored.
Holds the connection it was created/loaded with (con), so every method
below reuses it without needing it passed at every call. It is a private
attribute, not a field: it never round-trips through (de)serialization
(e.g. across a Prefect process boundary) - a revision that comes back
without one lazily reattaches the process-level default on next access.
Notes
resolver
property
¶
This record's Resolver, built once from its ancestry.
The ancestry is truncated at the deepest materialised ancestor, so a
deep tree resolves from few entries. Cached rather than rebuilt per
call: layers are write-once, so the only thing that can change
the truncation point is materialise, which clears this itself.
record
property
¶
record: Record
This revision's resolved view, as a Record.
The framework-agnostic view: narwhals frames, with no sign of how many
layers were folded to produce them. resolver remains the
DuckDB-shaped view, which datarecord.tools still builds from.
create
classmethod
¶
Insert a new revision, letting the DB assign the UUID.
Source code in src/datarecord/layered/revision.py
get
classmethod
¶
Load a revision by id.
Source code in src/datarecord/layered/revision.py
child
¶
Branch a new revision off this one.
Any node may be a parent: a layer is write-once, so a base cannot shift under its descendants.
Source code in src/datarecord/layered/revision.py
materialise
¶
Write this node's caches - owner maps and resolved dims.
A policy rather than a lifecycle step, and purely additive: it changes no answer, only how many layers a descendant's read touches. Once these exist, a read stops here instead of walking further up.
Notes
Source code in src/datarecord/layered/revision.py
ancestry
¶
Revision ids along the root->self path, root first.
Notes
Source code in src/datarecord/layered/revision.py
write_record
¶
write_record(
revision_id: UUID | None,
source: LayerData | RecordLike,
con: DuckDBPyConnection,
*,
uri: str | None = None,
) -> None
Write source as revision_id's layer, which must not exist yet.
An existing layer directory is an error rather than an overwrite or a merge, so a whole-record write can never half-replace what a record holds. Keys are looked up one at a time and each file written before the next is built, so a lazily-building source does one read per file rather than one per key up front.
Parameters:
| Name | Type | Description | Default |
|---|---|---|---|
revision_id
|
UUID | None
|
The record whose layer this is; |
required |
uri
|
str | None
|
Write here instead of at the revision's own layer - how a |
None
|
source
|
LayerData | RecordLike
|
The layer's contents: a |
required |
con
|
DuckDBPyConnection
|
Connection to write through. |
required |
Raises:
| Type | Description |
|---|---|
FileExistsError
|
If the layer directory already exists. |
ValueError
|
If a long frame is missing a long-schema column, or the schema declares a key dim no frame carries - either would make the fold misresolve the layer. |
Source code in src/datarecord/layered/write.py
37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 | |