Proposal: one interface, two objects — a source is its own rows, a fold is the resolution¶
Status: Landed · Drafted 2026-09-02 · Landed 2026-09-04
All three commits below landed on
refactor/staging-as-a-layer(cd2ce92,006c159,7f244a9). The authoritative account is the read path, the Record protocol — includingLayerData— writing and the module layout; this page is kept as the argument that led there, and those pages are authoritative where the two disagree.NodeCachelanded asResolver,ResolvedLayeris deleted, the base is aFold(datarecord/layered/fold.py) discovered from the sources, andwrite_recordtakes aLayerData. The untyped-enumerator∅-not-{None}half is the subject of its now-landed successor,a-member-file-holds-values. The "Commit N as landed" notes below record where the implementation departed from this sketch.
One claim: axis(dim) should mean the same thing wherever it is written — "the rows of the thing I am holding". You pick "one layer's contribution" versus "the resolved answer to here" by which object you hold, not by a resolved_ prefix on a method nor by a mode flag. A LayerSource answers for its own layer; a fold answers for everything folded into it; both answer under the same names.
The codebase is already most of the way there and inconsistent about the last step. WorkingRecord does it the clean way — the fold and the staged-layer's-own-rows are two objects (WorkingRecord the Record over a NodeCache, StagedSource the source). ResolvedLayer does it the muddy way — one object that is a fold (a persisted NodeCache) but is typed as a source and put in the fold's own input list, so it answers for its own layer through some methods and for the fold-to-here through others.
This is the prerequisite a-member-file-holds-values names, taken at its root rather than at the symptom. That page proposes adding resolved_axis/resolved_axes beside map_uri — a third folded-answer method under a distinct name, papering the dual role. The root cause is that a fold-result (ResolvedLayer, a persisted NodeCache) is masquerading as a fold-input (a LayerSource); untangle that and no new method name is needed.
What starts it¶
Three abstractions, each with a one-line contract the code states explicitly:
LayerSource— "One layer's own rows, however they are stored" (thesources.pyprotocol docstring).axis"means 'this layer's own rows'". Every implementation obeys it:ParquetLayer,DirectorySource,StagedSource.NodeCache— "A record's resolved view: owner map, dims, schema, and the relations over them" (resolve.py). It holds asourceslist and folds it: theinputsowner map,dims(the entity axis and groups among them), and everyrelation/entity_type_frame/group_frameare computed from those sources.Record/RecordLike— "the narwhals interface over one fold" (record.md). ARecordwraps exactly oneNodeCacheand presents its fold as narwhals frames.
So NodeCache is the fold, and a Record is its public face. The line between "own rows" (LayerSource) and "the fold" (NodeCache) is clean until one object straddles it.
ResolvedLayer is a NodeCache, persisted¶
This is the fact the rest of the page turns on. Revision.materialise() folds a node's sources and writes the result to resolved/ — _fold_kind for the inputs owner map, resolve_coords for the axes (the entity axis and groups among them) (materialise). That is exactly what a NodeCache computes in memory; materialising is persisting a NodeCache's output. ResolvedLayer is the handle that reads it back, and NodeCache seeds from it to skip the re-fold:
NodeCache._fold_map: seed = read(
head.map_uri(kind)
) # a ResolvedLayer's stored owner map
NodeCache.resolve_dims: seed = head.resolved_axes() # a ResolvedLayer's stored dims
So ResolvedLayer and NodeCache are the same thing — the fold to a node — in two tenses: NodeCache computed live, ResolvedLayer computed earlier and read from resolved/. A NodeCache reads a ResolvedLayer as its own prior incarnation, which is the whole point of sources_to_read stopping at a materialised ancestor: its ResolvedLayer already carries what re-folding from the root would recompute.
The defect is that this persisted fold is typed as a LayerSource and put in the sources list — a list of own-layer things. A fold-result is masquerading as a fold-input. Everything below follows from untangling that one type error.
The masquerade shows as a dual role¶
Because the persisted fold sits in the sources list, the head is asked two different questions by two different consumers of the same list — the fold-input question and the fold-result question it should never have had to answer at once:
| consumer | holds the head as | wants | reads |
|---|---|---|---|
resolve_dims seed |
fold seed | resolved axes to here | resolved/dims/ |
_fold_map seed |
fold seed | resolved inputs map to here |
resolved/owner_map/ |
source_for → relation |
a source | the head layer's own winning rows | layers/<id>/ |
source_for(layer_uuid) returns the head ResolvedLayer for a key the head owns, then reads its member with source.attribute(...) — the own-rows accessor. The seed paths read the folded answer. One object, both roles, and the only thing keeping them apart is that the folded answer hides under map_uri (and, in a-member-file-holds-values's sketch, would hide under resolved_axis).
relation already shows the intended shape by accident: for a layer the map names but the truncated source list does not hold, source_for fabricates ParquetLayer(layer_uuid, con) — a plain own-rows source. It wants a source, and where the list does not give it one it builds one. The head is the case where the list gives it a source that is secretly also a fold.
WorkingRecord is the model, and has the latent version of the same seam¶
WorkingRecord keeps the two views as two objects:
- The resolved view (base + staged) is
WorkingRecorditself, aRecord, built asbase_cache.with_source(StagedSource(self, ...))and answered by the inherited fold. - The own-layer view (staged rows alone, for commit) is
StagedSourcethe source, andstaged_only()→_Written, built from the_staged_*helpers.
That is the shape this proposal generalises. But the seam is present here too, quieter: staged_only() returns a _Written, a third type, rather than handing over the StagedSource that already is the staged layer's rows. _Written and StagedSource are two spellings of "the staged layer's own rows" — one shaped for the writer, one for the fold — because write_record's input is spelled as neither the source interface nor the interface a fold answers, but wants what both can give. The write-path half below closes that by naming the interface the two share. Left alone, _Written is the next ResolvedLayer waiting to happen.
The construct¶
A fold and a source answer the same interface; the object decides the meaning. Concretely, four moves: a Fold object holding the fold machinery, a Resolver that is LayerData routing into its Fold, a base discovered from the sources rather than stored, and the write path taking LayerData.
The fold machinery is a Fold object¶
Everything the fold produces and the reads gated by it move into one object, the counterpart to Coords (today's Dims, renamed — it holds axes and groups plus the broadcast/match algebra, which "dims" undersold):
class Fold:
"""A node's resolved view: the folded axes and owner map, and the reads over them."""
schema: Schema
coords: (
Coords # the folded axes and groups (entity among them) + expand/match helpers
)
owner_map: (
DuckDBPyRelation # folded (input_key, layer_uuid, varies/broadcast/breakpoints)
)
# pure-map reads
def attributes(
self,
) -> list[str]: ... # distinct `attribute` (was `attribute_names`)
def owners(
self, attribute: str
) -> DuckDBPyRelation: ... # map rows for one attribute
# map x axis read — here because a Fold has both halves
def flags(
self, entity_type: str | None = None
) -> dict[str, Flags]: ... # was `attributes_of`
@classmethod
def read(
cls, revision_id, con
) -> Fold | None: ... # from resolved/, folds nothing; None if unmaterialised
@classmethod
def compute(
cls, sources, con, schema, *, base: Fold | None
) -> Fold: ... # fold sources over the base
flags takes entity_type=None — whole-record when None, a type when named. The type-scoping (semi-join the map to that type's entities) needs the entity axis, which a Fold has via coords.axes["entity"]; so flags belongs on Fold, the one object holding both the map and the axes, rather than on a map-only OwnerMap that would have to be handed an entities relation from outside. owners(attribute) is the map-only read relation joins against; attributes() is the map-only enumerator. flags_from_rows stays a module function flags calls — it is schema logic, not map logic.
Fold.read and Fold.compute are the two tenses: a materialised node's resolved/ read back, versus a fold computed over sources. The same object shape either way, which is what makes a computed fold able to start from a read one as its base.
NodeCache becomes Resolver, LayerData routing into its Fold¶
The rename says what the class is: it resolves a node's layers into the answer a Record presents. "Cache" named an implementation detail (the frozen-prefix materialisation) as if it were the identity; "resolver" names the job.
class Resolver:
revision_id: UUID
sources: list[LayerSource] # the whole (frozen-truncated) ancestry as sources; no `base` field
con: DuckDBPyConnection
schema: Schema # a plain field (see below), resolved by the caller
@cached_property # stable-gated as `coords` is today
def fold(self) -> Fold: ... # base = deepest materialised source; compute the rest over it
# LayerData surface — routes into `fold`
def axes(self): ... # set(fold.coords.axes)
def axis(self, dim): ... # fold.coords.axes.get(dim)
def attributes(self, kind): ... # fold.attributes() for inputs; own-layer listing for outputs
def attribute(self, name, kind): ... # relation(name)/outputs(name), via fold.owners + source_for
def entity_types(self) / entity_type(self, ct): ... # was entity_types / entity_type_frame
def groups(self) / group(self, g): ... # was (from schema) / group_frame
# routing reads that need the fold's members by name
def owners(self, attribute): ... # fold.owners(attribute)
def flags(self, entity_type=None): ... # fold.flags(entity_type) (was attributes_of)
# extras a single fold-product cannot answer
def source_for(self, layer_uuid) -> LayerSource: ... # dispatch a winning key to its layer
def with_source(self, source) -> Resolver: ... # append one more layer (staging)
@property
def frozen(self) -> bool: ... # all(s.frozen for s in sources) (was `stable`)
inputs (the raw map relation) and entity_axis stay as thin routing properties — fold.owner_map and axis("entity") — because white-box tests and mutable.py read them by those names; they are owners and axis("entity") without the argument.
Resolver.frozen (the LayerData member) is today's NodeCache.stable: all(s.frozen for s in sources), whether the whole source list is write-once. Distinct from frozen_prefix, which counts how far the leading run is frozen to decide the in-memory cache boundary.
schema becomes a plain field (deferred to its own commit): today it is a property returning declared if set else read_schema(con), threaded through with_source. Every fold reads the schema, both construction sites already hold it (the tree node reads the root schema, Record.over has the manifest), and the resolver is lazy, so the fallback buys nothing. Moving read_schema(con) to the one tree-node construction site makes schema a plain field, deletes declared and the property, and lets with_source pass self.schema already resolved.
The base is a Fold, discovered from the sources, not stored¶
The base is not a stored field. sources is the whole ancestry (frozen-truncated for the in-memory cache, unrelated to materialisation) as plain LayerSources, and the fold discovers its base from among them:
def materialised(self, con) -> Fold | None: ... # on LayerSource
# ParquetLayer: Fold.read(self.revision_id, con) — a resolved/ dir, or None
# DirectorySource / StagedSource: None
Resolver.fold walks sources from the deepest, asks each source.materialised(con), and the first hit is the base: everything at or below it is skipped, everything above folds on top via Fold.compute(above, con, schema, base=hit). No materialised ancestor → base=None, fold the whole list.
This is why frozen and materialised are two questions, not one. frozen = "these rows cannot change under a reader" — a source-type flag (False only for StagedSource) gating the in-memory, connection-scoped fold cache (frozen_prefix/_frozen_table). materialised = "this node has a resolved/ directory on disk" — a filesystem fact, set only by an explicit .materialise(), gating the fold-shortening base. Every materialised node is frozen; almost no frozen node is materialised. So the base comes from the deepest materialised source, which the walk finds by testing source.materialised(con), never from "the last frozen source".
ResolvedLayer is deleted. Its fold-role becomes the base Fold.read(...) a source offers via materialised(con); its own-rows role, where source_for reaches a winning layer's rows, becomes a plain ParquetLayer. source_for resolves an owner id to a source and never to the base. The fold-result stops being an entry in the fold-input list.
Sequenced commits¶
This lands as three tested-green commits:
- Read-path split ✅ done — kill the masquerade:
source.materialised(con) -> Fold | None,Resolver.foldtakes its base from the deepest materialised source,sourcesis plainParquetLayers,ResolvedLayerdeleted,NodeCache -> Resolverrename (and.stable->.frozen). This is the piecea-member-file-holds-valuesis blocked on. Foldobject ✅ done — extractcoords+owner_mapand the reads (owners,flags(entity_type=None),attributes, per-type/group frames) intoFold;Resolverroutes into it. Orthogonal to (1); a clean second diff.LayerData+ write path ✅ done — defineLayerData, makeResolverandLayerSourceboth satisfy it (source gainsentity_types/groups/attributesenumerators), rename the fold-named readers to the shared vocabulary,write_recordtakes aLayerData,_Written/staged_only()deleted, both commit callers pass the object they hold.
[!NOTE] Commit 1 as landed (departures from the sketch above, and facts for commits 2–3):
Foldlives in its own moduledatarecord/layered/fold.py, not inresolve.py, to break the import cycle asource.materialised() -> Foldreturn would otherwise create (resolve.pyimportssources.py; a source returning aFoldfromresolve.pywould close the loop). In commit 1 it is a minimal seed struct —axes/groups/entity_typesdicts +owner_map, andFold.read— depending only onduck. Commit 2 moves the reads (owners/flags/attributes) onto it and folds the four dicts intocoords: Coords._base_and_above(sources, con)inresolve.pyis the split helper: it scanssources[:-1]deepest-first for the firstmaterialised(con)and returns(base_fold, sources_above). The last source is never the base — a node resolves from its own layer, never its own cache — which is also what stopsmaterialise(which writesowner_mapbeforedims) from reading its own half-written cache back as a base.- The base seeds the fold as depth 0 of each
fold_axis/fold_inputscall (its resolved, tombstone-free relation prepended), so a layer above still wins per key and its tombstones still evict a base key — behaviourally identical to the oldResolvedLayer-as-first-source seed. This touchesresolve_dims,resolve_groups,_fold_map,entity_type_frame, and_materialise_dims.- The invariant guard
test_resolved_reads_same_as_unresolved(intest_overlay.py) landed first, green before and after: it folds one node through the base and one from the root (via an_UnmaterialisedParquetLayersubclass forcingmaterialised() -> None) and asserts they agree on the owner map, entity axis, and every attribute relation.Commit 2 as landed (the
Foldobject):
Foldgrew into the real object infold.py:schema+axes/groups/entity_typesdicts +owner_map, plus the map readsattributes()(wasattribute_names),owners(attribute), andflags(entity_type=None)(wasattributes_of, now whole-record whenNone).flags_from_rowsmoved tofold.pybeside its only caller.Fold.computeis a resolve.py concern, not aFoldclassmethod. The folding that builds a liveFoldneedsCoords' broadcast algebra and the DuckDB expression machinery, both inresolve.py; relocating them into the lightfold.pywould invert the layering. SoResolver.foldassembles a liveFoldfromdims(the foldedCoords) and_map("inputs")(the folded owner map) — both of which keep their own frozen-scoped connection cache, sofoldre-wraps rather than re-folds and no cache moved.Fold.read(the base tense) stays aFoldclassmethod.- The live
Foldcarriesentity_types={}: it is never another fold's base, and its reads never touch the per-type frames (those go throughResolver.entity_type_frame, which folds per call).Coordsis unchanged —Dimsrenamed, holding axes+groups+the broadcast algebra.materialised(con, schema)gained the schema argument, sinceFold.readneeds it to cast the read-back map; every implementation and_base_and_abovethread it through.Commit 3 as landed (
LayerData+ the write path):
LayerDatalives inrecord.pybesideRecordLike, the neutral protocol module both the write path andmutable.pyalready import. Its enumerators returnIterable[str], notset[str]— aResolver.attributesorders its keys (aRecordover it needs a stable order) while a source's are aset, andIterableis the covariant type both satisfy without either re-wrapping.- Every source carries a
schemareference, so aLayerSourceis aLayerData. The reference is passed in at construction —sources_to_read,source_for,Record.over/Record.atandWorkingRecordall thread the one schema they already hold down into eachParquetLayer/DirectorySource/StagedSource— never re-read from disk. This is the "schemabecomes a plain field" step the commit-2 note deferred:Resolver.schemastops being a lazydeclared or read_schema(con)property and becomes a required field resolved by the caller, and the same reference sits on every source in the fold. A file layer's schema is never read by anything (the writer only ever gets aStagedSourceorResolver), but carrying it is what letscommithandself.resolver.sources[-1]straight towrite_recordwith no cast.- A
LayerSource's new enumerators take noschemaargument and do not filter by declaration —entity_types/groups/attributesare plain directory listings, asaxes()already is. The∅-not-{None}schema-governance the proposal makes an acceptance test of is a property of the objectswrite_recordactually receives (StagedSource,Resolver), whose staged/folded state is already declaration-scoped.Resolverkeepsrelation/outputsas private_relation/_outputsbehind the sharedattribute(name, kind), and returns a non-Nonerelation (an empty one where no layer wrote the attribute) — a valid narrowing ofLayerData.attribute's| None.group_frameandentity_typediffer:entity_type_framerenamed toentity_type(the shared name, no collision), butgroup_framestayed, being theRecord-presentational read (member order, empty→None) thatRecord._group_framealone uses; the sharedgroup(name)is the raw foldResolveralready had.isinstance(x, LayerData)decides the adapter, and@runtime_checkablechecks only member names. ARecordLikeproducer (_NetworkSource) hasdims/entity_typeswhereLayerDatawantsaxes/axis/entity_type/frozen, so it correctly fails the check and routes through_RecordLikeAsLayerDatainwrite.py— a thin adapter reading itsFramesmappings through the enumerate-and-read pairs (as_relationfor the one non-DuckDB backend boundary).write_recordtakesLayerData | RecordLike.commithands the fold's last source straight over:self.resolver.sources[-1](theStagedSource) forNewChild,self.resolverforDirectory, both already aLayerDatawith no cast needed._Written,staged_only(), and the_staged_axes/_staged_entities/_staged_attributesFramesbuilders are deleted;_staged_groupsstays forStagedSource.groups()'s key set.write.pyis raw-relation throughout:_validate_frame/_write_frame/_require_uniquetakeDuckDBPyRelation,nw.concat/.unique/.group_bybecameunion_all_by_name+ a self-join count, andas_relationis reached only inside theRecordLikeadapter.
The interface Resolver and a layer share¶
A Resolver and a LayerSource answer the same questions — "the keys of kind X" and "the rows for one key of kind X" — one over a fold, one over a single layer. That is a shared read protocol; call it LayerData. Each is that protocol plus its own extras.
| the shared question | LayerData member |
a layer answers (own rows) | a Resolver answers (folded) |
|---|---|---|---|
| the schema in force | schema |
— the writer supplies it | schema |
| which axes exist | axes() |
axes() ✓ |
keys of coords.axes |
| one axis's rows | axis(dim) |
axis(dim) ✓ |
coords.axes[dim] |
| which types exist | entity_types() |
new: list member files; ∅ if the schema declares no type axis |
entity_types() ✓ |
| one type's rows | entity_type(name) |
entity_type(name) ✓ |
entity_type_frame → renamed |
| which groups exist | groups() |
new: list group files; only what the schema declares | from schema → made explicit |
| one group's rows | group(name) |
group(name) ✓ |
group_frame → renamed |
| which attributes exist | attributes(kind) |
new: list attr files | attribute_names/output_names → renamed |
| one attribute's rows | attribute(name, kind) |
attribute(name, kind) ✓ |
relation/outputs → renamed |
schema heads the table but governs it: which of the pairs below it are populated depends on what the schema declares — see schema governs which pairs are live.
So LayerData is: schema, frozen, axes/axis, entity_types/entity_type, groups/group, attributes/attribute — the enumerate-and-read pairs (each an enumerate() -> set[str] and a read(key) -> Rel | None) plus the two members that mean the same over a layer and a fold. all_attributes stays source-only (see all_attributes is source-only). A LayerSource gains the three enumerators it lacks (entity_types, groups, attributes — directory listing for a _FileLayer as axes() already is, the staging tables for StagedSource); a Resolver renames its fold-named readers (entity_type_frame → entity_type, relation → attribute, and the coords.axes keys become axes()) so the two answer under one vocabulary.
frozen is on LayerData; all_attributes(kind) is not (source-only, below). frozen is a real property of a fold, not a convention-constant: a Resolver can hold a StagedSource (this is what WorkingRecord is), so it is over a mix of frozen and unfrozen sources, and its frozen is all(s.frozen for s in sources) — whether its whole source list is write-once. (Distinct from frozen_prefix, which counts how far the leading run is frozen to decide materialisation, not whether all of it is.)
Beyond the shared protocol, each keeps what only it has:
LayerSourcealso haslayer_id— the per-source keysource_fordispatches a winninglayer_uuidback through — andurion the file-backed ones.layer_idstays source-only because a fold has no single one: it spans many layers, keyed byrevision_idand routed by the owner map, so "the last layer'slayer_id" is not a valuesource_forcould dispatch on.Resolveralso has the fold machinery a single layer has no answer for, held in itsFoldobject:with_source,source_for,owners(over the one owner map),flags,fold/coords/entity_axis(the folded state), andrevision_idin place of alayer_id. (stableis nowfrozenon the shared protocol; the base is aFolddiscovered from the sources, not a stored field.)
Resolver fulfils LayerData; so does LayerSource. write_record takes a LayerData — a single StagedSource for the NewChild commit, a Resolver for the Directory commit — and _Written is deleted, being the adapter that filled the gap LayerData now closes.
The discipline that keeps this from becoming "one protocol for two unlike things": LayerData holds only the enumerate-and-read members that mean the same for a layer and a fold. A method only one has — all_attributes on a source, with_source on a Resolver — stays off the shared protocol. The shared names are for the shared idea; the rest are each type saying it is not the other.
schema governs which pairs are live, it is not a peer of them¶
The table above lists schema as one shared member among the enumerate-and-read pairs, but it is not their peer — it is what decides which of them exist. a-member-file-holds-values establishes that the entity-type axis is optional: where a schema declares no type axis (entity_type_dim is None), there are no member files, no entity_type column, and a component's attribute values live on dims/entity.parquet directly. So the entity_types()/entity_type(name) pair is not universal — it is populated exactly when the schema declares the axis, and groups() reflects only the groups the schema declares.
LayerData therefore reads through schema, not alongside it: schema is the one member the other pairs are conditional on, and a LayerData for an untyped record is a first-class shape, not the typed shape with types nulled out. The new source enumerators inherit a-member-file-holds-values's own acceptance test on this — the failure mode is an enumerator answering {None} rather than ∅:
entity_types()on a source whose schema declares no type axis returns the empty set, not{None}.entity_type(name)is then never a valid question — a direct call is a caller error, mirroringResolver.entity_type_frameon the same record.groups()returns only the groups the schema declares, and no more.
This is the point where the two proposals are the same work at two altitudes. a-member-file-holds-values is the read/write-path consequence of entity_type being optional, and it names the ResolvedLayer axis-means-one-thing fix as its prerequisite — which is the base/source split above. The two land together: this note gives LayerData the enumerators, and the untyped case is where those enumerators must answer ∅ rather than carry a phantom type.
axis is layer-local on every LayerSource, no exception¶
ResolvedLayer's axis/axes overrides go; it is deleted, and every source inherits the layer-local _FileLayer ones. The folded read moves to the base Fold a source offers via materialised(con), which is not a LayerSource. The test of correctness is a-member-file-holds-values's own: sources.py needs no branch, because a file-backed source no longer has to know whether it is being read as a layer or as a fold.
The writer takes a LayerData, and _Written disappears¶
_Written is not a second shape of StagedSource — it is a RecordLike, because write_record(source: RecordLike) speaks the record dialect and StagedSource speaks the source dialect. _Written is the adapter: staged_only() rebuilds the staged layer's rows as per-kind Frames so the writer can eat what the fold already reads as a source. It exists only because the writer's input was spelled as neither of the two things that can supply it.
But the writer needs exactly LayerData. It uses schema and, per kind, enumerates the keys this thing holds and reads each — which is the enumerate-and-read protocol above, no more. (It never touches flags, and outputs is attributes(kind="outputs").) So write_record takes a LayerData: the NewChild commit hands it the StagedSource, the Directory commit hands it the Resolver, each the object it already is. _Written is deleted, not collapsed — it was the adapter for a gap LayerData closes.
This is why LayerData earns its three new enumerators. The reader never needs them — the fold learns which keys exist from the schema and the owner map, so a LayerSource was point-accessor. The writer must enumerate what a layer actually holds, where the schema over-declares. Adding entity_types/groups/attributes to the source is what lets the reader's interface serve the writer too.
The Directory case proves the interface rather than excepting it. It writes the resolved record — base and staged flattened — which is not one layer's rows, so it cannot be a StagedSource; but "enumerate what I hold, hand each over" is exactly what a fold answers, where "what I hold" means every folded key. That is this proposal's thesis on the write path: one interface, the object deciding whether it means a layer or a fold.
This is the larger of the two halves and the most safely deferred, since it moves write_record's contract and reconciles names across the fold and the source. It belongs to the same idea: _Written is the write-path twin of a resolved_axis prefix — a third spelling of a thing the common interface should already carry.
What this buys¶
No object answers axis two ways. The overload that a-member-file-holds-values has to route around disappears at the source, so that proposal's prerequisite section reduces to "already true".
A file-backed source needs no branch. _FileLayer/ParquetLayer/DirectorySource read whichever columns a file has and never learn which configuration they are in — the strongest single signal the layout is carried by the format, not by the readers.
One vocabulary for "the rows of this thing". axis, entity_type, attribute, group mean "the rows of the object I hold", whether that object is a single layer or a fold. A reader learns the interface once.
The WorkingRecord and resolved-node mechanisms become one pattern. Both are "a fold-object over sources, with a distinguished own-rows source for the layer that is also the fold's boundary". Today they rhyme; after, they are the same construct.
The writer takes LayerData, the interface a source and a Resolver share. write_record's input is the enumerate-and-read surface both fulfil, so both commit callers pass the object they already hold and _Written, the adapter that existed only because the input was spelled as neither, is deleted.
What it costs¶
It reshapes the Resolver's (NodeCache's) source list. The head stops being a LayerSource that is secretly a fold; the base becomes a Fold the deepest materialised source offers via materialised(con), and source_for stops returning it. sources_to_read, the fold construction (now Fold.compute), resolve_coords, and source_for all move together. This is the change a-member-file-holds-values deferred as a prerequisite, taken at its root — the masquerade — rather than papered with a resolved_ prefix.
A shared interface across a source and a fold invites isinstance where a method belongs. The discipline that keeps it honest: if the fold-object needs a method a source does not have (a map has no entity_type; a source has no map), that is the interface telling you they are not the same type — the shared members are the ones that mean "my rows", and the rest stay apart. This proposal is not "one protocol for both"; it is "the same names for the same idea where the idea is shared".
Teaching the writer to take LayerData touches its contract. write_record moves from RecordLike to LayerData, LayerSource gains the three enumerators, and _Written and its staged_only() builder go. _Written has callers in the tests and one in commit, and the Directory path hands over a Resolver under the new contract rather than the record it built. A real change to the write path, not only a read-path cleanup, and the part most safely deferred.
What it settles¶
The base is a Fold, discovered from the sources. It is not a stored field. Fold.read(revision_id, con) reads resolved/ and folds nothing; a source offers it via materialised(con), and Resolver.fold takes the deepest materialised source's Fold as the base for Fold.compute. Same object shape either tense — read or computed — which is what lets a computed fold start from a read one.
LayerData is a Protocol both types satisfy structurally, with layer_id the sole source-only member. LayerSource is LayerData plus layer_id (and uri on file-backed ones) and all_attributes (source-only); Resolver is LayerData plus the fold machinery and revision_id. frozen is on LayerData, not source-only (above). Neither type inherits from LayerData — each satisfies it by shape, as LayerSource implementations already satisfy that protocol.
The untyped case keys on entity_type_dim is None, enum or not. An enumerator answers ∅ exactly when the schema declares no entity-type axis — not when schema.entity_types is empty, which is also true for a declared String-typed axis whose labels are data. A declared-but-non-enum axis is still a declared axis: its member files exist and entity_types() lists them. a-member-file-holds-values leaves whether those two cases should ever converge open; this note does not, keying only on declaration.
What it opens rather than settles¶
How far the name reconciliation reaches. The write-path half already reconciles the rows interface a fold and a source share (entity_type_frame/entity_type, relation/attribute, dims/axes become one vocabulary the writer reads). Left open is whether Record's public narwhals surface joins it — Record.dims returns Frames, not the relation a source's axis does, so it is a presentation over the same idea rather than the idea itself. Whether the resolved-and-own pair should share the outward name too, or only the internal rows interface does, is the larger claim this page does not force.
all_attributes is source-only, for now¶
all_attributes(kind) — a source's <kind>/*.parquet unioned by name, unfolded, values intact — stays a LayerSource member and off LayerData. It reads like the fold's owner map's raw input, and the temptation is to give a Resolver a folded all_attributes so the pair attributes/all_attributes is complete on both. Two facts hold it back:
- No fold-side caller. Both callers read it on a source:
fold_inputson each source to build the map, and theoutputsbulk read on the last layer (sources[-1].all_attributes("outputs")). Nothing asks aResolverfor its ownall_attributes, so a fold-side one would be a method with no reader. - The per-layer aggregate that looks separable is not layer-local. A layer's contribution to the owner map (aggregate
all_attributesto(input_key, layer_uuid, varies/broadcast/breakpoints)) needsexpand_dimsagainst the resolved axes to fan a broadcast NULL out to every value on its axis — a record-wide fact a single layer's files cannot supply. So the clean fault line isall_attributes(raw rows, coords-free, source) | expand+aggregate+recurse (coords, fold), not one row higher where asource.owner_map(coords)would sit.
Whether a later Store-shaped refactor makes a Resolver.owner_map(kind) and an all_attributes folded read worth having — so the two are the same enumerate-and-read pair on both types — is reopened then, not settled here.
How to know it worked¶
ResolvedLayer is gone; the base is a Fold. The single assertion the resolved-node half is downstream of: the persisted fold is a Fold.read(...) the live fold starts from, not a source in its input list. def axis in sources.py is found only on _FileLayer, and nothing folded-answer-shaped is typed as a source.
A resolved record reads the same as an unresolved one. Materialise a node, read a component's value through the resolved node and through the same node read from its own layers unmaterialised; they agree. This is the invariant the whole split protects, and it belongs in the suite before the change, not after — it is what says the base carries the fold correctly once axis stops doing so.
Resolver.sources holds only LayerSources, and the base is a Fold. The masquerade is over exactly when no ResolvedLayer appears in a sources list — every entry is a plain ParquetLayer, and the fold-to-here the head stood for is the Fold the deepest materialised source offers. source_for then always hands back an own-rows source, never the base. Assertable on the source types and on what source_for returns.
sources.py is untouched but for deletion. As in a-member-file-holds-values: if the file-backed sources need no added branch, the layout is carried by the format. A diff that adds a conditional to _FileLayer means the fold is leaking into the source again.
A source's new enumerators answer ∅, not {None}, for an untyped record. entity_types() on a LayerSource whose schema declares no type axis returns the empty set, and groups() only the declared groups — the same ∅-not-{None} check a-member-file-holds-values makes on the owner map. A test that materialises an untyped node and asserts entity_types() == set() on both the base Fold and a ParquetLayer of the head is what says schema governs the pairs rather than the pairs carrying a phantom type.
_Written is gone and write_record took the StagedSource directly — once that half lands. The NewChild commit passes the staged source to the writer, the Directory commit passes the resolved record, and both satisfy one input contract. Until it lands, _Written standing unchanged is the marker that this is the deferred part, not a regression.