Make the data model use Item and List types universally, with nodes authored as rank-polymorphic kernels (#4335)

* Add rank polymorphism node audit classifying all 271 nodes

* Implement StaticType for Item<T>

* Generate Item and mapped List wire variants for nodes declaring an Item<T> primary input

* Migrate nine nodes to Item element-wise kernels, dissolving the blending trait boilerplate

* Document the Item kernel implementation and staging plan

* Route Item<Vector> through TaggedValue::TypeDefault

* Add executor integration tests covering the Item and List wire variants

* Collapse element-wise Item/List wire pairs to the List form for conversion insertion

* Migrate sixteen vector modifier nodes to Item element-wise kernels

* Migrate Sample Image, Extend Image to Bounds, and Dehaze to Item element-wise kernels

* Fix bevel_with_transform test to actually exercise the transform attribute

* Implement From<T> for Item<T>

* Register PromoteNode rank adapters wrapping bare values into Item wires

* Insert PromoteNode adapters for Item/List wire pair fields in the preprocessor

* Define a real promote node backing the PromoteNode registry identifiers

* Zip ranked Item connectors by frame slot in the mapped element-wise variant

* Register ItemToListNode singleton raise adapters

* Resolve Item wires against List connectors by inserting promotion adapters at construction

* Rank the Offset Points distance connector and prove mixed-rank resolution end-to-end

* Implement Clampable for Item and List wires with per-variant clamp bounds

* Rank the Round Corners radius connector, exercising hard bounds on a ranked wire

* Implement ApplyTransform for Item

* Add Item wire implementations to the Transform node, keeping rank-0 chains rank 0

* Detect element-wise nodes by lazy primary connectors declaring Output = Item

* Convert Transform to an Item kernel with ranked parameters, delivering the broadcast milestone

* Rename Apply Transform to Bake Transform, baking item transforms on Vector, DAffine2, and DVec2

* Promote bare wires onto Item connectors at resolution via WrapItemNode adapters

* Rank the numeric, vector, and boolean parameters across the migrated element-wise nodes

* Rank the enum, integer, and seed parameters, registering their rank adapters via a consolidated macro

* Amend the audit with the DashPattern value type resolution

* Migrate the string family to Item element-wise kernels

* Unwrap Item wires into bare legacy connectors at resolution via UnwrapItemNode adapters

* Shadow owned node parameters in bodies instead of mut in signatures

* Migrate the math family and string measure nodes to Item element-wise kernels

* Convert the comparison and clamp nodes to Item kernels, dropping unreachable &str rows

* Flat-map expander kernels returning List under the mapped variant's frame

* Migrate the expander nodes to Item kernels flat-mapping under the frame

* Remove the unused peel_list helper

* Rank the raster adjustment and blending kernels, recontextualizing shader nodes onto an Item stand-in

Migrate the 16 adjustment nodes, Mix, Color Overlay, and Gradient Map from whole-List kernels to rank-0 Item kernels, letting the macro derive the List-mapped (zip) variants. Move the Adjust and Blend per-element seams off List onto the element types (add the Raster<CPU> impls, drop the now-dead List impls).

Shader nodes keep their bodies verbatim: PerPixelAdjust re-emits the identical kernel against a transparent no_std Item stand-in, so every Item<T> connector and .element() call resolves to a zero-cost identity on the GPU while the uniform buffer stays bare repr(C). The macro peels Item off ranked uniform params, wraps the fetched texel and uniforms at the entry point, and unwraps the result. This drops the shader_node/Item incompatibility guard. Register rank adapters for the adjustment enums.

* Update the rank polymorphism roadmap for the landed shader-node and adjustments chunk

* Rename the GPU Item stand-in to ShaderItem, aliased as Item at its shader-node import sites

* Flip the vector shape generators to emit rank-0 Item<Vector>

The shape generators (Rectangle, Circle, Ellipse, Arc, Spiral, Polygon, Star, Arrow, Line, Grid, QR Code) each produced exactly one shape wrapped in a singleton List<Vector>. Emit Item<Vector> directly so they connect to the rank-0 content connector of the migrated Transform node. Downstream List consumers receive the value through the existing Item to List promotion.

Relax the element-wise validation so a `()` (generator) primary may return Item<T> without being element-wise. Adapt the Repeat on Points test, which still takes a List content connector, by raising the generator's Item output through a singleton wrapper node.

* Parse ranked Item<T> parameter defaults against the bare element type

A ranked `Item<T>` parameter's default value is a bare, unranked `T` (promoted to the wire at resolution), but the preprocessor was handed the wrapped `Item<T>` type and could not parse the literal, flooding the console with warnings and dropping the defaults. Key the field's default_type metadata off the peeled element type for concrete ranked parameters, leaving generic `Item<T>` primaries and skip_impl nodes untouched.

* Parse an element-wise primary's scalar default against the bare element type

An element-wise node's primary reports its default_type as the List wire form so an unconnected primary defaults to an empty list. But when the primary carries a scalar `#[default]` (such as Root's radicand), that literal must parse as a bare element, not a List. Key the primary's default_type off the bare element type when it has a Default value source, keeping the List form otherwise.

* Add the DashPattern value type for stroke dash sequences

Introduce a rank-0 DashPattern value type (a Vec<f64> of alternating dash and gap lengths) so a stroke's dash pattern is a single frameable value rather than a rank-1 List<f64>. Register it as an auto-generated TaggedValue variant, parse its default from a comma or space separated string, and register its rank adapters. Not yet wired into the Stroke node.

* Rank the Fill and Stroke nodes element-wise and give Stroke a DashPattern connector

Migrate Fill and Stroke to element-wise Item<V> primaries (over Vector and Graphic element types) via a new element-level VectorItemMut trait, so styling one shape yields one shape and rank is preserved instead of promoting the input to a singleton List and emitting a List. The macro derives the List-mapped variant for genuine collections.

Wire the Stroke dash sequence to the new rank-0 DashPattern value type, collapsing the old content x paint x dash cartesian and dropping the IntoF64Vec trait. Update the stroke properties dash widget, the drawing tool, and graph-operation plumbing to read and write DashPattern, and migrate legacy F64Array, F64, and String dash inputs on document open.

Assign Colors stays a whole-collection node: each element's gradient position depends on its index among all siblings, which the element frame does not expose, so it keeps its List primary and the VectorListIterMut trait.

* Register rank adapters for the ranked Stroke enum parameters

The element-wise Stroke node ranks its align, cap, and paint order parameters as Item<StrokeAlign>, Item<StrokeCap>, and Item<PaintOrder>, but those enums lacked promotion adapters, so a bare default enum value could not be promoted to its Item wire and no Stroke variant resolved ("No construct found for node"). Register their rank adapters alongside StrokeJoin.

* Display Item wires in the Data panel without a List's ID column

Add a TableItemLayout impl for Item<T> and recognize Item wire types when introspecting graph data. An Item holds a single element, so it renders as a one-row table of the element plus its attributes with no leading index column, and it labels as its element type T rather than a List's T[]. Add ItemAttributeValues::get_any for the attribute widget dispatch.

* Register MonitorNode for Item wire types so the Data panel introspects them directly

Graph introspection wraps the inspected output in a generic MonitorNode typed to the wire. Without Item<T> monitor registrations, an Item<Vector> output could only be monitored after an Item to List promotion, so the Data panel captured and displayed a List<Vector> despite the connector being Item<Vector>. Register monitors for the Item types the element-wise nodes emit, and add the matching Data panel downcast entries.

* Color and double Item/List wires and cleave layer-stack connectors in the node graph

* Route wire color and rank through hidden nodes and refresh them on type changes

* Rework the DashPattern connector conversions with element-wise promotion and an explicit reducer node

* Rank the remaining value, context, aggregation, and transform nodes onto Item<T> wires

* Back DashPattern with a List<f64> so the Data panel can introspect its lengths

* Carry a single Item<T> through varargs so the Read context nodes emit Item<T> not List<T>

* Relax rank validation for aggregation shapes, add element adapters, and match variants by fewest promotions

* Rank the remaining bare and unnecessarily-List connectors across the node catalog

* Add Graphic::None and the FillChoice paint value, making colors and gradients plain values

* Rename GradientStops to Gradient and the legacy Gradient/Fill structs to LegacyGradient/LegacyFill

* Restore generator frame-from-params ranking to the roadmap as a planned stage

* Rename the ranked-field adapter identifier from PromoteNode to FieldAdapterNode to reflect its full contract

* Unload only the wires whose displayed style changed when types update

* Peel wire rank in the editor's semantic type checks so rank-0 layers are recognized

* Restore the whole-List Transform variant so rank-1 content wires resolve again

* Register the Item wire forms for the Memoize and Context Modification infrastructure nodes

* Give every ranked connector a field adapter and add numeric cast variants for legacy wires

* Key a ranked param's type default off its Item wire form when no literal default exists

* Inherit the layer's content value when splicing a node into an empty chain

* Migrate stale List-form TypeDefault inputs to the definition's current default

* Generate the mapped wire variant only when the element-wise node has a frame source

* Let a bare wire feed a List connector via a wrap-raise adapter, costed as two rank steps

* Add a zip companion to the whole-List Transform so ranked List parameters pair per slot

* Add the Sum, Average, Minimum, Maximum, Any, and All list reducers

* Convert the measure family to element-wise Item kernels per the audit classification

* Prefer the bare element value over the Item type default so ranked params keep their widgets

* Rename GradientStopsUI to GradientUI

* Split Fill's optional transform into a _has_transform bool and a ranked _transform matrix

* Rename the migration-only OptionalDAffine2 TaggedValue to LegacyOptionalDAffine2

* Flow byte buffers as Item<Resource> instead of List<u8> across the byte nodes

* Macro-generate the list-content wire variant, retiring the hand-written Transform-zip, Area, and Centroid companions

* Let ()-primary generators take ranked params and frame over them via the mapped variant, ranking Circle's radius

* Rank the vector shape generators' params to Item, adding a rank-aware input grab to the introspection harness

* Rank the value, color, and text generator params to Item

* Rank the raster, web-request, and context-reader generator params to Item

* Fix the repeat and brush test wirings left behind by the param-ranking sweeps

* Delete the vestigial Some, Unwrap Option, and Size Of debug nodes

* Delete the Attach Attribute node, folding its role into Write Attribute

* Add the Filter and Sort list companion nodes

* Guard the removed-definition migration swap target with a test

* Add the Box Corners value type in place of the rectangle corner radius list

* Split Text to Vector's per-glyph mode into a Text to Vector Glyphs node

* Rank the Combine Channels node's channel connectors to Item

* Make Map Points an element-wise node

* Delete the deprecated Upload Texture node

* Update the implementation roadmap to reflect the landed stages

* Let monitor introspection read rank-0 wires, locking in the layer coercion promotion path

* Prefer the rank-0 default when disconnecting a rank-capable input

* Make Path Modify an element-wise node

* Wrap node paths in a NodeIdPath newtype so they flow as a single Item

* Give Item<Raster<CPU>> a default so an unconnected Brush background resolves

* Stop the Brush node from setting layer attributes its paint operation doesn't produce

* Present-gate Flatten Path's adopted layer path like its fill and stroke

* Gate carried layer attributes on static column presence, not runtime values

* Give the remaining graphic Item<T> types a default so unconnected primaries resolve

* Dispatch a ranked param's Properties widget from its rank-0 element type

* Make Extract Transform an element-wise node, restoring the Origins to Polyline body

* Rename Flatten Path to Combine Paths

* Stamp Legacy Layer Extend's adopted layer path as a readable NodeIdPath

* Drop the dead List<u8> and List<NodeId> wire rows

* Rank Flatten Graphic's Fully Flatten toggle to Item

* Update the implementation roadmap with the endgame scope

* Make Combine Paths a reducer that collapses the whole frame into one path

* Stop type-converter nodes from carrying the source's unrelated attributes

* Format the Origins to Polyline regression test

* Wrap the Brush node's trace in a BrushTrace newtype so it flows as one value

* Make Switch a framed element-wise select, bundling whole collections

* Widen and align element-type coverage across the list and graphic nodes

* Register the compiler's cache chain pair for every ranked enum and newtype wire

* Fix wire colors for Passthrough outputs, bundled lists, and bools, and widen list wires

* Represent List wire types structurally with Type::List, replacing name-parsed rank promotion

* Treat scope and data fields as environment, rank scope wires as Item, and feed the render boundary through a context vararg

* Delete the vestigial Clone debug node

* Reinstate Upload Texture as an element-wise node and fix the GPU variants' scope executor and rank adapters

* Rename Combine Paths back to Flatten Path, deferring that rename to its own PR

* Deduplicate the promotion adapter registrations into the field adapter macro

* Rank Write Attribute's value connector to Item<AttributeValueDyn>, retiring the UnwrapItem bridge

* Vertical wire styling

* Store the editor layer path attribute as a bare NodeIdPath, not an Item<NodeIdPath>

* Rank Context Modification's features connector to Item<ContextFeatures>, dropping the dead memoize row

* Rank Path Modify's modification parameter to Item<Box<VectorModification>>

* Rename the field adapter node family to input adapter

* Drop the dead bare scalar rows from Context Modification's implementations list

* Move the dynamic executor's test module into its own file

* Drop the registry's unreachable bare rows for Memoize, the cache chain, and ConvertNode

* Materialize stored TaggedValues as ranked Item wires at the source

* Remove the bare-wire promotion and adapter machinery made dead by ranked value materialization

* Plant the input adapter for List-only inputs, composing position conversion from standard rows

* Consolidate Into/Convert conversions into the input adapter umbrella and rename the rank adapter identifiers

* Fix grouped layers gaining a phantom None stack element from the FillChoice default hijacking every List<Graphic> disconnect

* Enforce ranked node inputs in the macro, rejecting bare wire declarations

* Remove the unit Context => () machinery rows, leaving () purely as the no-primary sentinel

* Add a --signatures rank-audit mode to node-docs for the ranked-wire migration

* Remove the node-docs --signatures rank-audit mode now that ranked wires are enforced

* Migrate legacy no-color values on the Black & White, Color Overlay, and Empty Image color inputs

* Rewrite the element-wise accessor wire type at the primary input, not raw index 0

* Register the cache chain for Resource wires, replacing the lone hand-written Monitor row

* Gate the remaining Raster<GPU> registry rows behind the gpu feature

* Let List<DVec2> wires erase to ListDyn for the attribute reader and element counter

* Rename Extract Element to Item at Index, Count Elements to List Length, and Omit Element to Remove at Index

* Store paint picks as plain color/gradient values, removing the FillChoice value type

* Code review restructuring

* Sort by the consumed sort_key attribute or natural element order, adding the Sort Key node

* Remove the new list-combinator and reducer nodes to defer them to a follow-up PR

* Parse Fill and Stroke color defaults through the paint wire's Graphic element

* Emit ranked implementation-row default types structurally so their element TypeIds survive to default-literal parsing

* Exempt the deliberate no-paint choice from the stale List-form TypeDefault migration

* Migrate the legacy 4-input Fill directly to the split has-transform shape

* Upgrade the demo artwork

* Fix the valid AI review findings: Item eq/hash contract, table-era no-paint migration, quantize List rows, and other smaller issues

* Remove the rank polymorphism working documents

* Hash Item attribute values directly instead of debug-formatting them, speeding up cached evaluation

* Replace the data panel's dead bare-wire downcast arms with full coverage of the ranked monitor row types

* Derive PartialEq for Item now that attributes participate in equality

* Extend the data panel's attribute dispatchers with the newly supported scalar and choice enum types

* Add List monitor rows for the framed numeric conversion outputs so inspecting them resolves, with matching data panel arms
This commit is contained in:
Keavon Chambers
2026-09-08 16:03:01 +00:00
committed by Timon
parent 296185b7fc
commit a708a54492
3257 changed files with 766342 additions and 1829 deletions
@@ -0,0 +1,48 @@
[package]
name = "document-container"
description = "Container abstraction for the on-disk side of the .gdd document format"
edition.workspace = true
version.workspace = true
license.workspace = true
authors.workspace = true
[features]
default = []
zip = ["dep:zip"]
xz = ["dep:lzma-rust2", "dep:tar"]
[dependencies]
thiserror = "2.0"
log = { workspace = true }
zip = { workspace = true, optional = true, features = ["deflate-flate2-zlib-rs"], default-features = false}
lzma-rust2 = { workspace = true, optional = true }
tar = { version = "0.4", optional = true, default-features = false }
[target.'cfg(not(target_family = "wasm"))'.dependencies]
mmap-io = { workspace = true }
[target.'cfg(target_family = "wasm")'.dependencies]
web-sys = { workspace = true, features = [
"Navigator",
"DomException",
"Window",
"StorageManager",
"FileSystemCreateWritableOptions",
"FileSystemDirectoryHandle",
"FileSystemFileHandle",
"FileSystemGetFileOptions",
"FileSystemGetDirectoryOptions",
"FileSystemHandle",
"FileSystemHandleKind",
"FileSystemWritableFileStream",
"WritableStream",
"Blob",
] }
js-sys = { workspace = true }
wasm-bindgen = { workspace = true }
wasm-bindgen-futures = { workspace = true }
futures = { workspace = true }
[dev-dependencies]
tempfile = "3"
futures = { workspace = true }
@@ -0,0 +1,114 @@
//! Archive codecs (zip, xz).
//!
//! Each codec streams entries in both directions: writers wrap an `io::Write` sink, and
//! `deserialize` reads from any `io::Read + Seek` source and streams entries into any [`Container`].
#[cfg(any(feature = "xz", feature = "zip"))]
use crate::AsyncContainer;
#[cfg(any(feature = "zip", feature = "xz"))]
use crate::ContainerError;
use crate::Result;
use std::io::{Read, Seek, Write};
/// Hard cap on the total decompressed size a codec will materialize from one archive.
/// Defends against decompression bombs at the cost of refusing legitimately huge archives.
#[cfg(any(feature = "zip", feature = "xz"))]
pub(crate) const MAX_DECOMPRESSED_SIZE: u64 = 4 * 1024 * 1024 * 1024; // 4GB
/// Fold one entry's declared `size` into the running `total` and return it as a `usize` for `write_sized`.
/// Both codecs route entries through here so the decompression-bomb cap and 32-bit-safe conversion live in
/// one place. `write_sized` pre-allocates the declared size, so an over-large one is rejected before that.
#[cfg(any(feature = "zip", feature = "xz"))]
pub(crate) fn checked_entry_size(total: &mut u64, size: u64) -> Result<usize> {
*total = total.saturating_add(size);
if *total > MAX_DECOMPRESSED_SIZE {
return Err(ContainerError::SizeLimitExceeded {
declared: *total,
limit: MAX_DECOMPRESSED_SIZE,
});
}
// `usize` is 32-bit on wasm, so convert fallibly to rule out a silent truncation into a smaller allocation.
usize::try_from(size).map_err(|_| ContainerError::SizeLimitExceeded {
declared: size,
limit: usize::MAX as u64,
})
}
#[cfg(feature = "zip")]
mod zip;
#[cfg(feature = "zip")]
pub use zip::{Zip, ZipWriter};
#[cfg(feature = "xz")]
mod xz;
#[cfg(feature = "xz")]
pub use xz::{Xz, XzWriter};
/// Streaming archive codec. The associated `Writer` type wraps a `Write + Seek` sink (zip needs
/// `Seek` for the central directory; xz doesn't but `Seek` is free on file-like sinks) and
/// accepts entries one at a time. `finish` flushes the codec's trailer and consumes the wrapper.
pub trait Archive {
type Writer<W: Write + Seek>: ArchiveWriter
where
W: Write + Seek;
fn writer<W: Write + Seek>(output: W) -> Result<Self::Writer<W>>;
/// Read entries from `source` and write each into `dest`, streaming so neither the full
/// archive nor the full container ever sits in memory at once.
fn open<R: Read + Seek, C: AsyncContainer>(source: R, dest: &mut C) -> Result<()>;
}
pub trait ArchiveWriter: Sized {
type Sink;
fn write_entry(&mut self, path: &str, bytes: &[u8]) -> Result<()>;
/// Finish the archive and return the underlying sink, for in-memory archives where the caller
/// wants the written bytes (e.g. `Cursor<Vec<u8>>`) back.
fn finish_into(self) -> Result<Self::Sink>;
/// Finish the archive, discarding the sink. For file-backed archives the bytes are already on disk.
fn finish(self) -> Result<()> {
self.finish_into()?;
Ok(())
}
}
/// Archive container formats distinguishable by their leading magic bytes.
#[derive(Copy, Clone, Debug, PartialEq, Eq)]
pub enum ArchiveFormat {
Xz,
Zip,
}
impl ArchiveFormat {
/// Sniff the format from the leading magic bytes: xz streams start with `FD 37 7A 58 5A 00`,
/// zip archives with `50 4B 03 04` (`PK\x03\x04`). Returns `None` for anything else.
pub fn detect(bytes: &[u8]) -> Option<Self> {
if bytes.starts_with(&[0xFD, 0x37, 0x7A, 0x58, 0x5A, 0x00]) {
Some(Self::Xz)
} else if bytes.starts_with(&[0x50, 0x4B, 0x03, 0x04]) {
Some(Self::Zip)
} else {
None
}
}
}
/// Deserialize an archive into `dest`, auto-detecting the format from `bytes`' magic header.
/// Errors if the bytes are neither a recognized xz nor zip archive.
#[cfg(any(feature = "xz", feature = "zip"))]
pub fn open_auto<C: AsyncContainer>(bytes: &[u8], dest: &mut C) -> Result<()> {
let source = std::io::Cursor::new(bytes);
match ArchiveFormat::detect(bytes) {
#[cfg(feature = "xz")]
Some(ArchiveFormat::Xz) => Xz::open(source, dest),
#[cfg(feature = "zip")]
Some(ArchiveFormat::Zip) => Zip::open(source, dest),
#[allow(unreachable_patterns)]
Some(format) => Err(ContainerError::Codec(format!("tried to open {format:?}, but the binary was compiled without the feature enabled "))),
None => Err(ContainerError::Codec("unrecognized archive format (not xz or zip) ".into())),
}
}
@@ -0,0 +1,92 @@
//! Xz-compressed tarball archive codec.
use crate::archive::{Archive, ArchiveWriter, MAX_DECOMPRESSED_SIZE, checked_entry_size};
use crate::{AsyncContainer, ContainerError, Result, validate_path};
use lzma_rust2::{XzOptions, XzReader, XzWriter as InnerXzWriter};
use std::io::{Read, Seek, Write};
pub struct Xz;
/// xz-tar writer. Held as an `Option` so `finish` can take ownership and unwind the layered
/// writers in the right order: drop the tar builder first to flush its trailer, then finish xz.
pub struct XzWriter<W: Write + Seek> {
tar: Option<tar::Builder<InnerXzWriter<W>>>,
}
impl Archive for Xz {
type Writer<W: Write + Seek> = XzWriter<W>;
fn writer<W: Write + Seek>(output: W) -> Result<Self::Writer<W>> {
let xz_writer = InnerXzWriter::new(output, XzOptions::default()).map_err(lzma_err)?;
Ok(XzWriter {
tar: Some(tar::Builder::new(xz_writer)),
})
}
fn open<R: Read + Seek, C: AsyncContainer>(source: R, dest: &mut C) -> Result<()> {
// `take` bounds how many bytes we decompress from the xz stream, but each tar entry's declared
// size is fed to `write_sized`, which pre-allocates from it before reading. Cap the cumulative
// declared size too so a header claiming a huge size can't trigger a giant allocation up front.
let xz_reader = XzReader::new(source, false);
let bounded = xz_reader.take(MAX_DECOMPRESSED_SIZE);
let mut tar_reader = tar::Archive::new(bounded);
let mut total_size = 0u64;
for entry in tar_reader.entries()? {
let mut entry = entry?;
if entry.header().entry_type() != tar::EntryType::Regular {
continue;
}
// Reject non-UTF8 entry names rather than lossily rewriting them, so the path we store matches
// the archive exactly.
let path = entry.path()?;
let path = path.to_str().ok_or_else(|| ContainerError::Codec(format!("tar: non-UTF8 entry name {path:?}")))?.to_owned();
validate_path(&path)?;
let size = checked_entry_size(&mut total_size, entry.size())?;
dest.write_sized_non_blocking(&path, size, &mut |buffer| {
entry.read_exact(buffer).map_err(ContainerError::Io)?;
Ok(())
})?;
}
Ok(())
}
}
impl<W: Write + Seek> ArchiveWriter for XzWriter<W> {
type Sink = W;
fn write_entry(&mut self, path: &str, bytes: &[u8]) -> Result<()> {
validate_path(path)?;
let tar = self.tar.as_mut().ok_or_else(|| ContainerError::Codec("XzWriter already finished".into()))?;
let mut header = tar::Header::new_gnu();
header.set_path(path).map_err(|error| ContainerError::Codec(format!("tar: invalid path {path}: {error}")))?;
header.set_size(bytes.len() as u64);
header.set_mode(0o644);
header.set_cksum();
tar.append(&header, bytes)?;
Ok(())
}
fn finish_into(mut self) -> Result<W> {
self.finish_inner()
}
}
impl<W: Write + Seek> XzWriter<W> {
/// Unwind the layered writers in order (flush the tar trailer, then finish xz) and hand back the
/// innermost sink.
fn finish_inner(&mut self) -> Result<W> {
let mut tar = self.tar.take().ok_or_else(|| ContainerError::Codec("XzWriter already finished".into()))?;
tar.finish()?;
let xz_writer = tar.into_inner()?;
xz_writer.finish().map_err(lzma_err)
}
}
fn lzma_err(error: std::io::Error) -> ContainerError {
ContainerError::Codec(format!("lzma: {error}"))
}
@@ -0,0 +1,76 @@
//! Zip archive codec.
use crate::archive::{Archive, ArchiveWriter, checked_entry_size};
use crate::{AsyncContainer, ContainerError, Result, validate_path};
use std::io::{Read, Seek, Write};
use zip::ZipArchive;
use zip::write::{SimpleFileOptions, ZipWriter as InnerZipWriter};
pub struct Zip;
pub struct ZipWriter<W: Write + Seek> {
inner: InnerZipWriter<W>,
options: SimpleFileOptions,
}
impl Archive for Zip {
type Writer<W: Write + Seek> = ZipWriter<W>;
fn writer<W: Write + Seek>(output: W) -> Result<Self::Writer<W>> {
Ok(ZipWriter {
inner: InnerZipWriter::new(output),
options: SimpleFileOptions::default().compression_method(zip::CompressionMethod::Deflated),
})
}
fn open<R: Read + Seek, C: AsyncContainer>(source: R, dest: &mut C) -> Result<()> {
let mut archive = ZipArchive::new(source).map_err(zip_err)?;
// Zip headers declare each entry's uncompressed size up front; `checked_entry_size` caps the running
// total so a malicious archive can't exhaust memory or disk before any bytes are read.
let mut total_size = 0u64;
for index in 0..archive.len() {
let mut entry = archive.by_index(index).map_err(zip_err)?;
if !entry.is_file() {
continue;
}
let name = entry.name().to_string();
validate_path(&name)?;
let size = checked_entry_size(&mut total_size, entry.size())?;
dest.write_sized_non_blocking(&name, size, &mut |buffer| {
entry.read_exact(buffer).map_err(ContainerError::Io)?;
Ok(())
})?;
}
Ok(())
}
}
impl<W: Write + Seek> ArchiveWriter for ZipWriter<W> {
type Sink = W;
fn write_entry(&mut self, path: &str, bytes: &[u8]) -> Result<()> {
validate_path(path)?;
self.inner.start_file(path, self.options).map_err(zip_err)?;
self.inner.write_all(bytes)?;
Ok(())
}
fn finish_into(self) -> Result<W> {
self.inner.finish().map_err(zip_err)
}
}
fn zip_err(error: zip::result::ZipError) -> ContainerError {
// Preserve a real I/O failure (disk full, etc.) as a structured `Io` so callers can tell it apart from a
// corrupt-archive `Codec` error; only the genuinely archive-format errors collapse into `Codec`.
match error {
zip::result::ZipError::Io(io) => ContainerError::Io(io),
other => ContainerError::Codec(format!("zip: {other}")),
}
}
@@ -0,0 +1,9 @@
//! Container backend implementations.
pub mod memory;
#[cfg(not(target_family = "wasm"))]
pub mod folder;
#[cfg(target_family = "wasm")]
pub mod opfs;
@@ -0,0 +1,191 @@
//! Loose-folder backend.
use crate::{ByteHolder, Container, ContainerError, MmappedBytes, Result, validate_path, validate_prefix};
use mmap_io::mmap::{MemoryMappedFile, MmapMode};
use std::fs::{self, OpenOptions};
use std::io::Write;
use std::path::{Path, PathBuf};
pub struct FolderBackend {
root: PathBuf,
}
impl FolderBackend {
/// Open an existing folder. Errors if `root` is not a directory.
pub fn open(root: impl Into<PathBuf>) -> Result<Self> {
let root = root.into();
if !root.is_dir() {
return Err(ContainerError::NotFound(root.display().to_string()));
}
Ok(Self { root })
}
/// Create the folder if it does not exist, then open it.
pub fn create(root: impl Into<PathBuf>) -> Result<Self> {
let root = root.into();
fs::create_dir_all(&root)?;
Ok(Self { root })
}
pub fn root(&self) -> &std::path::Path {
&self.root
}
fn resolve(&self, path: &str) -> Result<PathBuf> {
validate_path(path)?;
self.reject_symlinked_components(path)?;
Ok(self.root.join(path))
}
/// Reject any existing component along `root/relative` that is a symlink. `validate_path`/`validate_prefix`
/// block `..` and absolute paths, but a symlink stored under the root could still point outside it, so every
/// path that gets joined onto the root must pass through here before it is opened or traversed.
fn reject_symlinked_components(&self, relative: &str) -> Result<()> {
let mut partial = self.root.clone();
for component in Path::new(relative).components() {
partial.push(component);
if let Ok(metadata) = fs::symlink_metadata(&partial)
&& metadata.file_type().is_symlink()
{
return Err(ContainerError::InvalidPath(relative.to_string()));
}
}
Ok(())
}
fn list_filtered(&self, prefix: &str, want_files: bool) -> Result<Vec<String>> {
validate_prefix(prefix)?;
let base = if prefix.is_empty() || prefix == "." {
self.root.clone()
} else {
self.reject_symlinked_components(prefix)?;
self.root.join(prefix)
};
// A missing prefix has no entries; a prefix that names a file is a misuse.
if base.is_file() {
return Err(ContainerError::NotADirectory(prefix.to_string()));
}
if !base.is_dir() {
return Ok(Vec::new());
}
let mut results = Vec::new();
for entry in fs::read_dir(&base)? {
let entry = entry?;
// `file_type` does not follow symlinks, unlike `is_file`/`is_dir`. Skip symlink entries so a
// listing never advertises a path that `resolve` would then reject as a container escape.
let Ok(file_type) = entry.file_type() else { continue };
let matches = if want_files { file_type.is_file() } else { file_type.is_dir() };
if !matches {
continue;
}
let path = entry.path();
let relative = path.strip_prefix(&self.root).map_err(|_| ContainerError::Backend("path escaped root".into()))?;
results.push(relative.to_string_lossy().replace('\\', "/"));
}
Ok(results)
}
}
impl Container for FolderBackend {
fn read(&self, path: &str) -> Result<ByteHolder> {
let full = self.resolve(path)?;
if !full.is_file() {
return Err(ContainerError::NotFound(path.to_string()));
}
// Mmapping a zero-length file is platform-dependent and often fails, so serve empty files as owned
// bytes and reserve mmap for files that actually have content.
if fs::metadata(&full).map(|metadata| metadata.len() == 0).unwrap_or(false) {
return Ok(ByteHolder::Owned(Vec::new()));
}
Ok(ByteHolder::Mmapped(MmappedBytes::new(open_mmap(&full)?)?))
}
fn write(&self, path: &str, bytes: &[u8]) -> Result<()> {
let full = self.resolve(path)?;
if let Some(parent) = full.parent() {
fs::create_dir_all(parent)?;
}
fs::write(&full, bytes)?;
Ok(())
}
fn append(&self, path: &str, bytes: &[u8]) -> Result<()> {
let full = self.resolve(path)?;
if let Some(parent) = full.parent() {
fs::create_dir_all(parent)?;
}
let mut file = OpenOptions::new().create(true).append(true).open(&full)?;
file.write_all(bytes)?;
Ok(())
}
fn write_sized(&self, path: &str, size: usize, fill: &mut dyn FnMut(&mut [u8]) -> Result<()>) -> Result<()> {
if size == 0 {
return self.write(path, &[]);
}
let full = self.resolve(path)?;
if let Some(parent) = full.parent() {
fs::create_dir_all(parent)?;
}
let file = MemoryMappedFile::create_rw(&full, size as u64).map_err(|error| ContainerError::Backend(format!("create_rw {full:?} failed: {error}")))?;
let result = {
let mut slice = file
.as_slice_mut(0, size as u64)
.map_err(|error| ContainerError::Backend(format!("as_slice_mut {full:?} failed: {error}")))?;
fill(slice.as_mut())?;
drop(slice);
file.flush().map_err(|error| ContainerError::Backend(format!("flush {full:?} failed: {error}")))
};
// `create_rw` materializes the full-size file before `fill` runs, so remove it on failure rather
// than leave a zeroed or half-written remnant.
if result.is_err() {
drop(file);
let _ = fs::remove_file(&full);
}
result
}
fn list(&self, prefix: &str) -> Result<Vec<String>> {
self.list_filtered(prefix, true)
}
fn list_dirs(&self, prefix: &str) -> Result<Vec<String>> {
self.list_filtered(prefix, false)
}
fn exists(&self, path: &str) -> bool {
match self.resolve(path) {
Ok(full) => full.is_file(),
Err(_) => false,
}
}
fn remove(&self, path: &str) -> Result<()> {
let full = self.resolve(path)?;
// Idempotent: a missing file is not an error. Directories are left alone.
if full.is_file() {
fs::remove_file(full)?;
}
// TODO: decide if we should remove empty parent directories
Ok(())
}
}
/// Open a memory-mapped read-only view of `path`, trying huge pages first. Callers must ensure `path`
/// is non-empty, since mmapping a zero-length file is platform-dependent.
fn open_mmap(path: &Path) -> Result<MemoryMappedFile> {
match MemoryMappedFile::builder(path).mode(MmapMode::ReadOnly).huge_pages(true).open() {
Ok(file) => Ok(file),
Err(_) => MemoryMappedFile::open_ro(path).map_err(|error| ContainerError::Backend(format!("mmap of {path:?} failed: {error}"))),
}
}
@@ -0,0 +1,89 @@
//! In-memory backend. Useful for tests and as the deserialize target for archive codecs.
use crate::{ByteHolder, Container, ContainerError, Result, validate_path, validate_prefix, with_trailing_slash};
use std::collections::{HashMap, HashSet};
use std::sync::Mutex;
#[derive(Default)]
pub struct MemoryBackend {
files: Mutex<HashMap<String, Vec<u8>>>,
}
impl MemoryBackend {
pub fn new() -> Self {
Self::default()
}
}
impl Container for MemoryBackend {
fn read(&self, path: &str) -> Result<ByteHolder> {
validate_path(path)?;
self.files
.lock()
.unwrap()
.get(path)
.map(|bytes| ByteHolder::Owned(bytes.clone()))
.ok_or_else(|| ContainerError::NotFound(path.to_string()))
}
fn write(&self, path: &str, bytes: &[u8]) -> Result<()> {
validate_path(path)?;
self.files.lock().unwrap().insert(path.to_string(), bytes.to_vec());
Ok(())
}
fn append(&self, path: &str, bytes: &[u8]) -> Result<()> {
validate_path(path)?;
self.files.lock().unwrap().entry(path.to_string()).or_default().extend_from_slice(bytes);
Ok(())
}
fn list(&self, prefix: &str) -> Result<Vec<String>> {
validate_prefix(prefix)?;
let files = self.files.lock().unwrap();
// A prefix that is itself a stored key names a file, not a directory.
if files.contains_key(prefix) {
return Err(ContainerError::NotADirectory(prefix.to_string()));
}
let normalized = with_trailing_slash(prefix);
let results = files.keys().filter(|path| path.starts_with(&normalized) && !path[normalized.len()..].contains('/')).cloned().collect();
Ok(results)
}
fn list_dirs(&self, prefix: &str) -> Result<Vec<String>> {
validate_prefix(prefix)?;
let files = self.files.lock().unwrap();
if files.contains_key(prefix) {
return Err(ContainerError::NotADirectory(prefix.to_string()));
}
let normalized = with_trailing_slash(prefix);
let mut seen = HashSet::new();
let mut results = Vec::new();
for path in files.keys() {
if !path.starts_with(&normalized) {
continue;
}
let remainder = &path[normalized.len()..];
if let Some((segment, _)) = remainder.split_once('/') {
let dir = format!("{normalized}{segment}");
if seen.insert(dir.clone()) {
results.push(dir);
}
}
}
Ok(results)
}
fn exists(&self, path: &str) -> bool {
validate_path(path).is_ok() && self.files.lock().unwrap().contains_key(path)
}
fn remove(&self, path: &str) -> Result<()> {
validate_path(path)?;
// Idempotent: removing a missing path is not an error.
self.files.lock().unwrap().remove(path);
Ok(())
}
}
@@ -0,0 +1,475 @@
//! OPFS (Origin Private File System) backend for browser wasm.
use crate::{AsyncContainer, ByteHolder, ContainerError, Result, validate_path, validate_prefix, with_trailing_slash};
use futures::channel::oneshot;
use js_sys::Uint8Array;
use std::collections::{HashSet, VecDeque};
use std::sync::{Arc, Mutex};
use wasm_bindgen::JsCast;
use wasm_bindgen::prelude::*;
use wasm_bindgen_futures::{JsFuture, spawn_local};
use web_sys::{
Blob, DomException, FileSystemCreateWritableOptions, FileSystemDirectoryHandle, FileSystemFileHandle, FileSystemGetDirectoryOptions, FileSystemGetFileOptions, FileSystemWritableFileStream,
WritableStream,
};
enum Mutation {
Write {
path: String,
bytes: Vec<u8>,
},
Append {
path: String,
bytes: Vec<u8>,
},
Delete {
path: String,
},
/// In-band flush barrier. The worker fulfills the sender once it dequeues this, which (by FIFO
/// order) signals that every mutation enqueued before it has been applied to disk.
Barrier(oneshot::Sender<()>),
}
struct Inner {
directory: FileSystemDirectoryHandle,
/// Paths believed to be on disk, letting `exists_non_blocking` answer without OPFS's missing sync API.
/// Optimistic: a path is inserted when its write is queued and kept even if that write later fails, so it
/// can briefly over-report. Rebuilt from real disk state on the next `open`.
on_disk: HashSet<String>,
queue: VecDeque<Mutation>,
worker_active: bool,
}
pub struct OpfsBackend {
inner: Arc<Mutex<Inner>>,
}
// Safety: only built for browser wasm where JS handles never leave the main thread.
unsafe impl Send for OpfsBackend {}
unsafe impl Sync for OpfsBackend {}
impl OpfsBackend {
/// Open (or create) `directory_name` under the OPFS root.
pub async fn open(directory_name: &str) -> Result<Self> {
let directory = open_directory(directory_name).await.map_err(js_err)?;
let on_disk = enumerate_paths(&directory, "").await.map_err(js_err)?;
Ok(Self {
inner: Arc::new(Mutex::new(Inner {
directory,
on_disk,
queue: VecDeque::new(),
worker_active: false,
})),
})
}
fn directory(&self) -> FileSystemDirectoryHandle {
self.inner.lock().unwrap().directory.clone()
}
/// Wait until all queued non-blocking mutations have hit disk, so awaited reads observe them. Plants a
/// barrier at the back of the queue and awaits the worker reaching it; by FIFO order, everything ahead is
/// then applied. Draining stays the worker's job so it remains the sole mutator.
async fn flush(&self) {
let (sender, receiver) = oneshot::channel();
{
let mut guard = self.inner.lock().unwrap();
guard.queue.push_back(Mutation::Barrier(sender));
kick_worker(&self.inner, &mut guard);
}
// The sender is only dropped without sending if the worker is torn down mid-drain; either way
// there is nothing left to wait for, so a receive error is treated as "already flushed".
let _ = receiver.await;
}
}
impl AsyncContainer for OpfsBackend {
async fn read(&self, path: &str) -> Result<ByteHolder> {
validate_path(path)?;
// Non-blocking writes only land on disk once the queue drains, so flush first and then treat
// disk as authoritative. Draining (rather than folding the queue against an awaited base read)
// avoids racing the background worker, which could otherwise double-apply a queued append.
self.flush().await;
let bytes = read_file(&self.directory(), path).await.map_err(js_err)?;
Ok(ByteHolder::Owned(bytes))
}
async fn write(&self, path: &str, bytes: &[u8]) -> Result<()> {
validate_path(path)?;
// Flush first so this awaited write lands after any non-blocking mutation already queued for the
// same path, then apply directly so the real disk error still propagates to the caller.
self.flush().await;
write_file(&self.directory(), path, bytes).await.map_err(js_err)?;
self.inner.lock().unwrap().on_disk.insert(path.to_string());
Ok(())
}
async fn append(&self, path: &str, bytes: &[u8]) -> Result<()> {
validate_path(path)?;
self.flush().await;
append_file(&self.directory(), path, bytes).await.map_err(js_err)?;
self.inner.lock().unwrap().on_disk.insert(path.to_string());
Ok(())
}
async fn write_sized(&self, path: &str, size: usize, fill: &mut dyn FnMut(&mut [u8]) -> Result<()>) -> Result<()> {
validate_path(path)?;
let mut buffer = vec![0; size];
fill(&mut buffer)?;
self.flush().await;
write_file(&self.directory(), path, &buffer).await.map_err(js_err)?;
self.inner.lock().unwrap().on_disk.insert(path.to_string());
Ok(())
}
async fn list(&self, prefix: &str) -> Result<Vec<String>> {
validate_prefix(prefix)?;
self.flush().await;
list_entries(&self.directory(), prefix, EntryKind::File).await.map_err(|error| list_error(prefix, error))
}
async fn list_dirs(&self, prefix: &str) -> Result<Vec<String>> {
validate_prefix(prefix)?;
self.flush().await;
list_entries(&self.directory(), prefix, EntryKind::Directory).await.map_err(|error| list_error(prefix, error))
}
async fn exists(&self, path: &str) -> bool {
if validate_path(path).is_err() {
return false;
}
self.flush().await;
file_exists(&self.directory(), path).await
}
async fn remove(&self, path: &str) -> Result<()> {
validate_path(path)?;
self.flush().await;
// Idempotent: a missing entry is not an error, matching the other backends.
if let Err(error) = remove_file(&self.directory(), path).await
&& !is_not_found(&error)
{
return Err(js_err(error));
}
self.inner.lock().unwrap().on_disk.remove(path);
Ok(())
}
fn write_non_blocking(&self, path: &str, bytes: &[u8]) -> Result<()> {
validate_path(path)?;
let mut guard = self.inner.lock().unwrap();
guard.on_disk.insert(path.to_string());
guard.queue.push_back(Mutation::Write {
path: path.to_string(),
bytes: bytes.to_vec(),
});
kick_worker(&self.inner, &mut guard);
Ok(())
}
fn append_non_blocking(&self, path: &str, bytes: &[u8]) -> Result<()> {
validate_path(path)?;
let mut guard = self.inner.lock().unwrap();
guard.on_disk.insert(path.to_string());
guard.queue.push_back(Mutation::Append {
path: path.to_string(),
bytes: bytes.to_vec(),
});
kick_worker(&self.inner, &mut guard);
Ok(())
}
fn remove_non_blocking(&self, path: &str) -> Result<()> {
validate_path(path)?;
let mut guard = self.inner.lock().unwrap();
guard.on_disk.remove(path);
guard.queue.push_back(Mutation::Delete { path: path.to_string() });
kick_worker(&self.inner, &mut guard);
Ok(())
}
fn exists_non_blocking(&self, path: &str) -> bool {
validate_path(path).is_ok() && self.inner.lock().unwrap().on_disk.contains(path)
}
}
fn kick_worker(inner: &Arc<Mutex<Inner>>, guard: &mut Inner) {
if guard.worker_active {
return;
}
guard.worker_active = true;
let inner = inner.clone();
spawn_local(drain_queue(inner));
}
/// Apply every queued mutation to disk, in FIFO order, until the queue is empty, then mark the worker
/// idle. Spawned once via [`kick_worker`] and runs as the sole mutator; the awaited read paths wait on
/// a [`Mutation::Barrier`] rather than draining themselves, so there is never more than one drainer.
///
/// Consecutive appends to the same path are coalesced into one combined append. Each `createWritable`
/// on OPFS copies the whole existing file, so appending N frames one at a time is O(N^2) in bytes
/// copied (every per-op autosave queues one append per hot op). Concatenating a run into a single
/// append collapses that to one file copy, while staying byte-identical to applying them in order.
async fn drain_queue(inner: Arc<Mutex<Inner>>) {
loop {
// Pop the front mutation, coalescing a run of same-path appends behind it into one batch.
let (directory, mutation) = {
let mut guard = inner.lock().unwrap();
let Some(mut mutation) = guard.queue.pop_front() else {
guard.worker_active = false;
return;
};
if let Mutation::Append { path, bytes } = &mut mutation {
while let Some(Mutation::Append { path: next_path, .. }) = guard.queue.front()
&& next_path == path
{
let Some(Mutation::Append { bytes: next_bytes, .. }) = guard.queue.pop_front() else {
unreachable!()
};
bytes.extend_from_slice(&next_bytes);
}
}
(guard.directory.clone(), mutation)
};
match mutation {
Mutation::Write { path, bytes } => {
if let Err(error) = write_file(&directory, &path, &bytes).await {
log::error!("OPFS background write for {path} failed: {error:?}");
}
}
Mutation::Append { path, bytes } => {
if let Err(error) = append_file(&directory, &path, &bytes).await {
log::error!("OPFS background append for {path} failed: {error:?}");
}
}
Mutation::Delete { path } => {
// Removal is idempotent, so a missing entry is the expected outcome of a redundant delete, not an error.
if let Err(error) = remove_file(&directory, &path).await
&& !is_not_found(&error)
{
log::error!("OPFS background delete for {path} failed: {error:?}");
}
}
// A receive error on the waiter side just means the reader gave up; nothing to apply.
Mutation::Barrier(sender) => {
let _ = sender.send(());
}
}
}
}
fn js_err(error: JsValue) -> ContainerError {
ContainerError::Backend(format!("{error:?}"))
}
/// Resolve `directory_path` (a `/`-separated relative path) under the OPFS root, creating each
/// segment. OPFS rejects directory names containing `/`, so a multi-segment path like
/// `documents/<id>` must be descended one segment at a time rather than passed whole.
async fn open_directory(directory_path: &str) -> std::result::Result<FileSystemDirectoryHandle, JsValue> {
let storage = web_sys::window().ok_or_else(|| JsValue::from_str("no window"))?.navigator().storage();
let mut current: FileSystemDirectoryHandle = JsFuture::from(storage.get_directory()).await?.dyn_into()?;
for segment in directory_path.split('/').filter(|segment| !segment.is_empty()) {
let options = FileSystemGetDirectoryOptions::new();
options.set_create(true);
current = JsFuture::from(current.get_directory_handle_with_options(segment, &options)).await?.dyn_into()?;
}
Ok(current)
}
/// Descend the `/`-separated path against `root` and return the directory handle plus the final segment.
async fn descend<'a>(root: &FileSystemDirectoryHandle, relative: &'a str, create_dirs: bool) -> std::result::Result<(FileSystemDirectoryHandle, &'a str), JsValue> {
let mut current = root.clone();
let mut segments = relative.split('/').filter(|s| !s.is_empty()).collect::<Vec<_>>();
let file = segments.pop().ok_or_else(|| JsValue::from_str("empty path"))?;
for segment in segments {
let options = FileSystemGetDirectoryOptions::new();
options.set_create(create_dirs);
current = JsFuture::from(current.get_directory_handle_with_options(segment, &options)).await?.dyn_into()?;
}
Ok((current, file))
}
async fn read_file(root: &FileSystemDirectoryHandle, path: &str) -> std::result::Result<Vec<u8>, JsValue> {
let (directory, name) = descend(root, path, false).await?;
let handle: FileSystemFileHandle = JsFuture::from(directory.get_file_handle(name)).await?.dyn_into()?;
let file_value = JsFuture::from(handle.get_file()).await?;
let blob: Blob = file_value.dyn_into()?;
let buffer = JsFuture::from(blob.array_buffer()).await?;
Ok(Uint8Array::new(&buffer).to_vec())
}
async fn write_file(root: &FileSystemDirectoryHandle, path: &str, bytes: &[u8]) -> std::result::Result<(), JsValue> {
let (directory, name) = descend(root, path, true).await?;
let options = FileSystemGetFileOptions::new();
options.set_create(true);
let handle: FileSystemFileHandle = JsFuture::from(directory.get_file_handle_with_options(name, &options)).await?.dyn_into()?;
let writable: FileSystemWritableFileStream = JsFuture::from(handle.create_writable()).await?.dyn_into()?;
let stream: WritableStream = writable.clone().unchecked_into();
// Wrap in an async block so any failure, including a synchronous JS throw from `write_with_js_u8_array`,
// aborts the stream before returning instead of leaving a dangling locked writable.
let write = async {
let array = Uint8Array::from(bytes);
JsFuture::from(writable.write_with_js_u8_array(&array)?).await?;
Ok::<(), JsValue>(())
}
.await;
if let Err(error) = write {
let _ = JsFuture::from(stream.abort()).await;
return Err(error);
}
JsFuture::from(stream.close()).await?;
Ok(())
}
async fn append_file(root: &FileSystemDirectoryHandle, path: &str, bytes: &[u8]) -> std::result::Result<(), JsValue> {
let (directory, name) = descend(root, path, true).await?;
let file_options = FileSystemGetFileOptions::new();
file_options.set_create(true);
let handle: FileSystemFileHandle = JsFuture::from(directory.get_file_handle_with_options(name, &file_options)).await?.dyn_into()?;
// Determine the current end-of-file so we can seek there before writing.
let file_value = JsFuture::from(handle.get_file()).await?;
let blob: Blob = file_value.dyn_into()?;
let offset = blob.size();
// `keepExistingData: true` preserves bytes outside the written range; without it OPFS truncates to the written length.
let writable_options = FileSystemCreateWritableOptions::new();
writable_options.set_keep_existing_data(true);
let writable: FileSystemWritableFileStream = JsFuture::from(handle.create_writable_with_options(&writable_options)).await?.dyn_into()?;
let stream: WritableStream = writable.clone().unchecked_into();
// Wrap in an async block so a synchronous JS throw from `seek_with_f64`/`write_with_js_u8_array`
// aborts the stream instead of bypassing the abort via `?`.
let seek_and_write = async {
JsFuture::from(writable.seek_with_f64(offset)?).await?;
let array = Uint8Array::from(bytes);
JsFuture::from(writable.write_with_js_u8_array(&array)?).await?;
Ok::<(), JsValue>(())
}
.await;
if let Err(error) = seek_and_write {
let _ = JsFuture::from(stream.abort()).await;
return Err(error);
}
JsFuture::from(stream.close()).await?;
Ok(())
}
async fn remove_file(root: &FileSystemDirectoryHandle, path: &str) -> std::result::Result<(), JsValue> {
let (directory, name) = descend(root, path, false).await?;
JsFuture::from(directory.remove_entry(name)).await?;
Ok(())
}
async fn file_exists(root: &FileSystemDirectoryHandle, path: &str) -> bool {
let Ok((directory, name)) = descend(root, path, false).await else {
return false;
};
JsFuture::from(directory.get_file_handle(name)).await.is_ok()
}
#[derive(Clone, Copy)]
enum EntryKind {
File,
Directory,
}
async fn list_entries(root: &FileSystemDirectoryHandle, prefix: &str, want: EntryKind) -> std::result::Result<Vec<String>, JsValue> {
// `.` and the empty string both name the container root.
let prefix = if prefix == "." { "" } else { prefix };
let directory = if prefix.is_empty() {
root.clone()
} else {
let mut current = root.clone();
for segment in prefix.split('/').filter(|s| !s.is_empty()) {
let options = FileSystemGetDirectoryOptions::new();
options.set_create(false);
current = match JsFuture::from(current.get_directory_handle_with_options(segment, &options)).await {
Ok(value) => value.dyn_into()?,
Err(error) if is_not_found(&error) => return Ok(Vec::new()),
Err(error) => return Err(error),
};
}
current
};
let entries = directory.entries();
let iterator: js_sys::AsyncIterator = entries.unchecked_into();
let mut results = Vec::new();
let prefix_with_slash = with_trailing_slash(prefix);
let want_kind = match want {
EntryKind::File => web_sys::FileSystemHandleKind::File,
EntryKind::Directory => web_sys::FileSystemHandleKind::Directory,
};
loop {
let next: js_sys::IteratorNext = JsFuture::from(iterator.next()?).await?.unchecked_into();
if next.done() {
break;
}
let pair: js_sys::Array = next.value().unchecked_into();
let Some(name) = pair.get(0).as_string() else { continue };
let handle: web_sys::FileSystemHandle = pair.get(1).unchecked_into();
if handle.kind() != want_kind {
continue;
}
results.push(format!("{prefix_with_slash}{name}"));
}
Ok(results)
}
/// Walk every file under `prefix` (recursively) and collect their full container paths.
/// Used at open time to populate the in-memory tracking set.
async fn enumerate_paths(root: &FileSystemDirectoryHandle, prefix: &str) -> std::result::Result<HashSet<String>, JsValue> {
let mut paths = HashSet::new();
let mut to_visit = vec![prefix.to_string()];
while let Some(current_prefix) = to_visit.pop() {
for file in list_entries(root, &current_prefix, EntryKind::File).await? {
paths.insert(file);
}
for dir in list_entries(root, &current_prefix, EntryKind::Directory).await? {
to_visit.push(dir);
}
}
Ok(paths)
}
fn is_not_found(error: &JsValue) -> bool {
error.dyn_ref::<DomException>().is_some_and(|error| error.name() == "NotFoundError")
}
/// OPFS raises `TypeMismatchError` when a path segment used as a directory is actually a file.
fn is_type_mismatch(error: &JsValue) -> bool {
error.dyn_ref::<DomException>().is_some_and(|error| error.name() == "TypeMismatchError")
}
/// Map a listing error: a prefix that names a file becomes [`ContainerError::NotADirectory`], anything
/// else passes through as a backend error.
fn list_error(prefix: &str, error: JsValue) -> ContainerError {
if is_type_mismatch(&error) {
ContainerError::NotADirectory(prefix.to_string())
} else {
js_err(error)
}
}
@@ -0,0 +1,512 @@
//! Container abstraction for the on-disk side of the `.gdd` document format.
//!
//! A [`Container`] is a virtual filesystem of named byte payloads.
//! Backends include a loose folder, an in-memory map, and an OPFS-backed wasm store.
//! Archive codecs ([`archive::Zip`], [`archive::Xz`]) round-trip a container's contents
//! through a compressed byte stream.
//!
//! Reads return a [`ByteHolder`], an ownership-carrying handle whose variant depends on
//! how the backend produced the bytes (mmap region, owned vector, external file mmap).
//! [`AsyncContainer`] mirrors [`Container`] for inherently async backends; every sync
//! [`Container`] is reachable from async code via a blanket impl.
pub mod archive;
pub mod backends;
pub enum ByteHolder {
/// Bytes synthesized in memory (decompressed from an archive, produced by serialization).
/// The only variant available on `target_family = "wasm"`.
Owned(Vec<u8>),
/// Bytes mmap'd from a file inside the container.
#[cfg(not(target_family = "wasm"))]
Mmapped(MmappedBytes),
/// Bytes mmap'd from a file outside the container (e.g. a linked resource).
#[cfg(not(target_family = "wasm"))]
External { path: std::path::PathBuf, bytes: MmappedBytes },
}
impl ByteHolder {
pub fn as_slice(&self) -> &[u8] {
match self {
ByteHolder::Owned(bytes) => bytes,
#[cfg(not(target_family = "wasm"))]
ByteHolder::Mmapped(bytes) => bytes.as_ref(),
#[cfg(not(target_family = "wasm"))]
ByteHolder::External { bytes, .. } => bytes.as_ref(),
}
}
/// If the bytes are backed by a real filesystem path, return it. Enables consumers to
/// short-circuit byte copies with `fs::copy` (CoW on supported filesystems).
#[cfg(not(target_family = "wasm"))]
pub fn source_path(&self) -> Option<&std::path::Path> {
match self {
ByteHolder::Owned(_) => None,
ByteHolder::Mmapped(bytes) => Some(bytes.path()),
ByteHolder::External { path, .. } => Some(path),
}
}
#[cfg(target_family = "wasm")]
pub fn source_path(&self) -> Option<&std::path::Path> {
None
}
/// Open an external file and produce a [`ByteHolder::External`] backed by mmap.
#[cfg(not(target_family = "wasm"))]
pub fn open_external(path: impl Into<std::path::PathBuf>) -> Result<Self> {
let path = path.into();
let file = mmap_io::mmap::MemoryMappedFile::open_ro(&path).map_err(|error| ContainerError::Backend(format!("mmap of {path:?} failed: {error}")))?;
let bytes = MmappedBytes::new(file)?;
Ok(ByteHolder::External { path, bytes })
}
}
impl AsRef<[u8]> for ByteHolder {
fn as_ref(&self) -> &[u8] {
self.as_slice()
}
}
impl std::fmt::Debug for ByteHolder {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
match self {
ByteHolder::Owned(bytes) => f.debug_tuple("Owned").field(&format_args!("{} bytes", bytes.len())).finish(),
#[cfg(not(target_family = "wasm"))]
ByteHolder::Mmapped(bytes) => f.debug_tuple("Mmapped").field(&format_args!("{} bytes", bytes.as_ref().len())).finish(),
#[cfg(not(target_family = "wasm"))]
ByteHolder::External { path, bytes } => f.debug_struct("External").field("path", path).field("len", &bytes.as_ref().len()).finish(),
}
}
}
/// Owning wrapper around a memory-mapped file that exposes the mapped region as `&[u8]`.
#[cfg(not(target_family = "wasm"))]
pub struct MmappedBytes(mmap_io::mmap::MemoryMappedFile);
#[cfg(not(target_family = "wasm"))]
impl MmappedBytes {
/// Wrap a mapped file, probing that its region is sliceable so the failure surfaces here rather than
/// later degrading to the `&[]` fallback in [`AsRef::as_ref`], which cannot return an error.
pub fn new(file: mmap_io::mmap::MemoryMappedFile) -> Result<Self> {
let len = file.len();
file.as_slice(0, len)
.map_err(|error| ContainerError::Backend(format!("mmap slice of {:?} failed: {error}", file.path())))?;
Ok(Self(file))
}
pub fn path(&self) -> &std::path::Path {
self.0.path()
}
}
#[cfg(not(target_family = "wasm"))]
impl AsRef<[u8]> for MmappedBytes {
fn as_ref(&self) -> &[u8] {
let len = self.0.len();
match self.0.as_slice(0, len) {
Ok(slice) => slice,
Err(error) => {
log::error!("Failed to obtain mmap slice: {error}");
&[]
}
}
}
}
#[derive(Debug, thiserror::Error)]
pub enum ContainerError {
#[error("path not found: {0}")]
NotFound(String),
#[error("invalid path: {0}")]
InvalidPath(String),
#[error("not a directory: {0}")]
NotADirectory(String),
#[error("I/O error: {0}")]
Io(#[from] std::io::Error),
/// An archive declared more decompressed data (`declared` bytes) than the codec is willing to
/// materialize (`limit` bytes).
#[error("declared decompressed size {declared} exceeds the {limit}-byte limit")]
SizeLimitExceeded { declared: u64, limit: u64 },
/// Failure inside an archive codec (zip, lzma, tar). Wraps a foreign error as text since those
/// error types don't share a common Rust trait we can chain through.
#[error("codec error: {0}")]
Codec(String),
/// Failure inside a storage backend (mmap, OPFS/JS). Wraps a foreign error as text for the same reason.
#[error("backend error: {0}")]
Backend(String),
}
pub type Result<T> = std::result::Result<T, ContainerError>;
/// Normalize a listing prefix to end with a trailing slash (unless it names the container root).
/// The root (empty string or `.`) normalizes to the empty string, so backends concatenate child paths
/// as `prefix/child` without a double, missing, or `./`-rooted slash.
pub(crate) fn with_trailing_slash(prefix: &str) -> String {
if prefix.is_empty() || prefix == "." {
String::new()
} else if prefix.ends_with('/') {
prefix.to_string()
} else {
format!("{prefix}/")
}
}
/// Validate that `path` names a container-safe file: relative, no `.`/`..` segments, no backslashes,
/// no redundant separators (`a//b`, `a/b/`). Dotfile names like `.gitignore` are fine. Requiring a canonical
/// form gives a file one identity across backends, some of which key on the raw string rather than path
/// components. Used by backends to block container escapes and by archive codecs on untrusted entry names.
/// For listing prefixes, which may name the container root, use [`validate_prefix`] instead.
pub fn validate_path(path: &str) -> Result<()> {
let invalid = || ContainerError::InvalidPath(path.to_string());
if path.is_empty() || path.contains('\\') || path.starts_with('/') {
return Err(invalid());
}
// `Path::components` silently folds away `//`, trailing `/`, and interior `.` segments, but backends key
// on the raw string, so a non-canonical path would resolve to one file on a path-joining backend yet a
// different identity on a string-keyed backend. Reject the redundant segments the component loop below
// can't see (it never observes a folded-away `CurDir`/empty segment).
if path.split('/').any(|segment| segment.is_empty() || segment == ".") {
return Err(invalid());
}
// Reject Windows drive-letter prefixes (`C:foo`, `C:/foo`) that platform-agnostic Path doesn't recognize as absolute on Linux.
let bytes = path.as_bytes();
if bytes.len() >= 2 && bytes[1] == b':' && bytes[0].is_ascii_alphabetic() {
return Err(invalid());
}
for component in std::path::Path::new(path).components() {
use std::path::Component;
match component {
Component::Normal(_) => {}
Component::CurDir | Component::ParentDir | Component::Prefix(_) | Component::RootDir => return Err(invalid()),
}
}
Ok(())
}
/// Validate a listing prefix. Same rules as [`validate_path`], except the container root is also a valid
/// prefix, named by either the empty string or `.`. Backends pass this to `list`/`list_dirs`.
pub fn validate_prefix(prefix: &str) -> Result<()> {
if prefix.is_empty() || prefix == "." {
return Ok(());
}
validate_path(prefix)
}
/// Synchronous virtual filesystem of named byte payloads.
pub trait Container {
/// Read the contents of `path` into a [`ByteHolder`].
fn read(&self, path: &str) -> Result<ByteHolder>;
/// Write `bytes` at `path`, creating intermediate directories as needed.
fn write(&self, path: &str, bytes: &[u8]) -> Result<()>;
/// Append `bytes` to the file at `path`, creating it (and any intermediate directories)
/// if it does not yet exist. Equivalent to `write` on a fresh path.
fn append(&self, path: &str, bytes: &[u8]) -> Result<()>;
/// Write `size` bytes whose contents are produced by `fill`.
/// The default implementation allocates and forwards to [`Container::write`];
/// backends that can mmap a writable region may override to fill in place.
fn write_sized(&self, path: &str, size: usize, fill: &mut dyn FnMut(&mut [u8]) -> Result<()>) -> Result<()> {
let mut buffer = vec![0; size];
fill(&mut buffer)?;
self.write(path, &buffer)
}
/// List file entries directly under `prefix` (non-recursive).
/// Returned paths include the `prefix` (e.g. `list("resources")` returns
/// `["resources/abc123", ...]`). A missing prefix yields an empty list; a prefix that names a
/// file (not a directory) is an error.
fn list(&self, prefix: &str) -> Result<Vec<String>>;
/// List subdirectory entries directly under `prefix` (non-recursive).
/// Returned names include the `prefix`, without a trailing slash
/// (e.g. `list_dirs("")` returns `["resources"]`). Same missing/file-prefix semantics as [`Container::list`].
fn list_dirs(&self, prefix: &str) -> Result<Vec<String>>;
/// Whether a file exists at `path`. Directories return `false`.
fn exists(&self, path: &str) -> bool;
/// Remove the file at `path`. Idempotent: removing a missing path succeeds. Directories are never
/// removed (they exist only implicitly as parents of files).
fn remove(&self, path: &str) -> Result<()>;
}
/// Asynchronous virtual filesystem of named byte payloads. Mirrors [`Container`].
///
/// The returned futures are intentionally not `Send`: native uses `block_on` at the save seam
/// and wasm is single-threaded, so neither needs cross-thread futures. Revisit if we ever want
/// to run container I/O on a thread pool.
#[expect(async_fn_in_trait, reason = "see trait docs — Send is not required")]
pub trait AsyncContainer {
async fn read(&self, path: &str) -> Result<ByteHolder>;
async fn write(&self, path: &str, bytes: &[u8]) -> Result<()>;
async fn append(&self, path: &str, bytes: &[u8]) -> Result<()>;
async fn write_sized(&self, path: &str, size: usize, fill: &mut dyn FnMut(&mut [u8]) -> Result<()>) -> Result<()>;
async fn list(&self, prefix: &str) -> Result<Vec<String>>;
async fn list_dirs(&self, prefix: &str) -> Result<Vec<String>>;
async fn exists(&self, path: &str) -> bool;
async fn remove(&self, path: &str) -> Result<()>;
/// Synchronous write. On backends with sync I/O (folder, memory) the write completes durably
/// before return and reports real errors. On OPFS the write is enqueued onto a background task
/// and `Ok` is returned eagerly; a later failure is logged.
fn write_non_blocking(&self, path: &str, bytes: &[u8]) -> Result<()>;
/// Synchronous write. Same eager-enqueue semantics on OPFS as [`write_non_blocking`](Self::write_non_blocking);
/// queued appends preserve order relative to earlier queued writes/appends.
fn write_sized_non_blocking(&self, path: &str, size: usize, fill: &mut dyn FnMut(&mut [u8]) -> Result<()>) -> Result<()> {
let mut buf = vec![0u8; size];
fill(&mut buf)?;
self.write_non_blocking(path, &buf)
}
/// Synchronous append. Same eager-enqueue semantics on OPFS as [`write_non_blocking`](Self::write_non_blocking);
/// queued appends preserve order relative to earlier queued writes/appends.
fn append_non_blocking(&self, path: &str, bytes: &[u8]) -> Result<()>;
/// Synchronous remove. Same semantics as [`write_non_blocking`](Self::write_non_blocking).
fn remove_non_blocking(&self, path: &str) -> Result<()>;
/// Non-blocking existence check. On OPFS this reads from an in-memory tracking set populated by
/// the sync write/remove paths, since the underlying OPFS existence API is async only.
fn exists_non_blocking(&self, path: &str) -> bool;
}
impl<C: Container + ?Sized> AsyncContainer for C {
async fn read(&self, path: &str) -> Result<ByteHolder> {
Container::read(self, path)
}
async fn write(&self, path: &str, bytes: &[u8]) -> Result<()> {
Container::write(self, path, bytes)
}
async fn append(&self, path: &str, bytes: &[u8]) -> Result<()> {
Container::append(self, path, bytes)
}
async fn write_sized(&self, path: &str, size: usize, fill: &mut dyn FnMut(&mut [u8]) -> Result<()>) -> Result<()> {
Container::write_sized(self, path, size, fill)
}
async fn list(&self, prefix: &str) -> Result<Vec<String>> {
Container::list(self, prefix)
}
async fn list_dirs(&self, prefix: &str) -> Result<Vec<String>> {
Container::list_dirs(self, prefix)
}
async fn exists(&self, path: &str) -> bool {
Container::exists(self, path)
}
async fn remove(&self, path: &str) -> Result<()> {
Container::remove(self, path)
}
fn write_non_blocking(&self, path: &str, bytes: &[u8]) -> Result<()> {
Container::write(self, path, bytes)
}
fn write_sized_non_blocking(&self, path: &str, size: usize, fill: &mut dyn FnMut(&mut [u8]) -> Result<()>) -> Result<()> {
Container::write_sized(self, path, size, fill)
}
fn append_non_blocking(&self, path: &str, bytes: &[u8]) -> Result<()> {
Container::append(self, path, bytes)
}
fn remove_non_blocking(&self, path: &str) -> Result<()> {
Container::remove(self, path)
}
fn exists_non_blocking(&self, path: &str) -> bool {
Container::exists(self, path)
}
}
/// Type-erased container that dispatches to one of the in-tree backends.
///
/// `AsyncContainer::read` returns `impl Future`, so `dyn AsyncContainer` is not object-safe.
/// `AnyContainer` is the workaround: `Gdd` holds one of these by value, and the `AsyncContainer`
/// impl forwards to the active variant.
pub enum AnyContainer {
Memory(backends::memory::MemoryBackend),
#[cfg(not(target_family = "wasm"))]
Folder(backends::folder::FolderBackend),
#[cfg(target_family = "wasm")]
Opfs(backends::opfs::OpfsBackend),
}
impl AsyncContainer for AnyContainer {
async fn read(&self, path: &str) -> Result<ByteHolder> {
match self {
Self::Memory(backend) => AsyncContainer::read(backend, path).await,
#[cfg(not(target_family = "wasm"))]
Self::Folder(backend) => AsyncContainer::read(backend, path).await,
#[cfg(target_family = "wasm")]
Self::Opfs(backend) => AsyncContainer::read(backend, path).await,
}
}
async fn write(&self, path: &str, bytes: &[u8]) -> Result<()> {
match self {
Self::Memory(backend) => AsyncContainer::write(backend, path, bytes).await,
#[cfg(not(target_family = "wasm"))]
Self::Folder(backend) => AsyncContainer::write(backend, path, bytes).await,
#[cfg(target_family = "wasm")]
Self::Opfs(backend) => AsyncContainer::write(backend, path, bytes).await,
}
}
async fn append(&self, path: &str, bytes: &[u8]) -> Result<()> {
match self {
Self::Memory(backend) => AsyncContainer::append(backend, path, bytes).await,
#[cfg(not(target_family = "wasm"))]
Self::Folder(backend) => AsyncContainer::append(backend, path, bytes).await,
#[cfg(target_family = "wasm")]
Self::Opfs(backend) => AsyncContainer::append(backend, path, bytes).await,
}
}
async fn write_sized(&self, path: &str, size: usize, fill: &mut dyn FnMut(&mut [u8]) -> Result<()>) -> Result<()> {
match self {
Self::Memory(backend) => AsyncContainer::write_sized(backend, path, size, fill).await,
#[cfg(not(target_family = "wasm"))]
Self::Folder(backend) => AsyncContainer::write_sized(backend, path, size, fill).await,
#[cfg(target_family = "wasm")]
Self::Opfs(backend) => AsyncContainer::write_sized(backend, path, size, fill).await,
}
}
async fn list(&self, prefix: &str) -> Result<Vec<String>> {
match self {
Self::Memory(backend) => AsyncContainer::list(backend, prefix).await,
#[cfg(not(target_family = "wasm"))]
Self::Folder(backend) => AsyncContainer::list(backend, prefix).await,
#[cfg(target_family = "wasm")]
Self::Opfs(backend) => AsyncContainer::list(backend, prefix).await,
}
}
async fn list_dirs(&self, prefix: &str) -> Result<Vec<String>> {
match self {
Self::Memory(backend) => AsyncContainer::list_dirs(backend, prefix).await,
#[cfg(not(target_family = "wasm"))]
Self::Folder(backend) => AsyncContainer::list_dirs(backend, prefix).await,
#[cfg(target_family = "wasm")]
Self::Opfs(backend) => AsyncContainer::list_dirs(backend, prefix).await,
}
}
async fn exists(&self, path: &str) -> bool {
match self {
Self::Memory(backend) => AsyncContainer::exists(backend, path).await,
#[cfg(not(target_family = "wasm"))]
Self::Folder(backend) => AsyncContainer::exists(backend, path).await,
#[cfg(target_family = "wasm")]
Self::Opfs(backend) => AsyncContainer::exists(backend, path).await,
}
}
async fn remove(&self, path: &str) -> Result<()> {
match self {
Self::Memory(backend) => AsyncContainer::remove(backend, path).await,
#[cfg(not(target_family = "wasm"))]
Self::Folder(backend) => AsyncContainer::remove(backend, path).await,
#[cfg(target_family = "wasm")]
Self::Opfs(backend) => AsyncContainer::remove(backend, path).await,
}
}
fn write_non_blocking(&self, path: &str, bytes: &[u8]) -> Result<()> {
match self {
Self::Memory(backend) => AsyncContainer::write_non_blocking(backend, path, bytes),
#[cfg(not(target_family = "wasm"))]
Self::Folder(backend) => AsyncContainer::write_non_blocking(backend, path, bytes),
#[cfg(target_family = "wasm")]
Self::Opfs(backend) => AsyncContainer::write_non_blocking(backend, path, bytes),
}
}
fn append_non_blocking(&self, path: &str, bytes: &[u8]) -> Result<()> {
match self {
Self::Memory(backend) => AsyncContainer::append_non_blocking(backend, path, bytes),
#[cfg(not(target_family = "wasm"))]
Self::Folder(backend) => AsyncContainer::append_non_blocking(backend, path, bytes),
#[cfg(target_family = "wasm")]
Self::Opfs(backend) => AsyncContainer::append_non_blocking(backend, path, bytes),
}
}
fn remove_non_blocking(&self, path: &str) -> Result<()> {
match self {
Self::Memory(backend) => AsyncContainer::remove_non_blocking(backend, path),
#[cfg(not(target_family = "wasm"))]
Self::Folder(backend) => AsyncContainer::remove_non_blocking(backend, path),
#[cfg(target_family = "wasm")]
Self::Opfs(backend) => AsyncContainer::remove_non_blocking(backend, path),
}
}
fn exists_non_blocking(&self, path: &str) -> bool {
match self {
Self::Memory(backend) => AsyncContainer::exists_non_blocking(backend, path),
#[cfg(not(target_family = "wasm"))]
Self::Folder(backend) => AsyncContainer::exists_non_blocking(backend, path),
#[cfg(target_family = "wasm")]
Self::Opfs(backend) => AsyncContainer::exists_non_blocking(backend, path),
}
}
}
#[cfg(test)]
mod tests {
use super::{ContainerError, validate_path, validate_prefix};
#[test]
fn validate_path_accepts_well_formed() {
// `..` and `.` are only rejected as whole path components; as substrings of a name they are fine.
for ok in ["manifest.json", "resources/abc", "a/b/c.bin", "my..file.txt", "..hidden", ".gitignore"] {
assert!(validate_path(ok).is_ok(), "{ok:?} should be accepted");
}
}
#[test]
fn validate_path_rejects_unsafe() {
for bad in ["", "../escape", "a/../b", ".", "./leading", "/abs", "back\\slash", "C:/win", "\\\\?\\unc"] {
let result = validate_path(bad);
assert!(matches!(result, Err(ContainerError::InvalidPath(_))), "{bad:?} should be rejected, got {result:?}");
}
}
#[test]
fn validate_prefix_accepts_root_tokens() {
// The container root is a valid listing prefix, named by either the empty string or `.`.
for root in ["", ".", "resources", "a/b"] {
assert!(validate_prefix(root).is_ok(), "{root:?} should be accepted as a prefix");
}
// Unsafe prefixes are still rejected, same as paths.
assert!(validate_prefix("../escape").is_err());
}
}
@@ -0,0 +1,81 @@
#![cfg(any(feature = "zip", feature = "xz"))]
use document_container::Container;
use document_container::archive::{Archive, ArchiveWriter};
use document_container::backends::folder::FolderBackend;
use document_container::backends::memory::MemoryBackend;
fn entries() -> Vec<(&'static str, &'static [u8])> {
vec![
("manifest.json", br#"{"format":"gdd"}"#),
("document.json", b"{\"registry\":\"...\"}"),
("history.jsonl", b"{\"rev\":1}\n{\"rev\":2}\n"),
("resources/abc123", &[0xDE, 0xAD, 0xBE, 0xEF]),
("resources/xyz789", b"another resource"),
]
}
fn assert_round_trip(restored: &MemoryBackend) {
assert_eq!(restored.read("manifest.json").unwrap().as_slice(), br#"{"format":"gdd"}"#);
assert_eq!(restored.read("document.json").unwrap().as_slice(), b"{\"registry\":\"...\"}");
assert_eq!(restored.read("history.jsonl").unwrap().as_slice(), b"{\"rev\":1}\n{\"rev\":2}\n");
assert_eq!(restored.read("resources/abc123").unwrap().as_slice(), &[0xDE, 0xAD, 0xBE, 0xEF]);
assert_eq!(restored.read("resources/xyz789").unwrap().as_slice(), b"another resource");
}
#[cfg(feature = "zip")]
#[test]
fn zip_round_trip() {
use document_container::archive::Zip;
use std::io::Cursor;
let mut buffer = Cursor::new(Vec::new());
let mut writer = Zip::writer(&mut buffer).unwrap();
for (path, bytes) in entries() {
writer.write_entry(path, bytes).unwrap();
}
writer.finish().unwrap();
let mut restored = MemoryBackend::new();
<Zip as Archive>::open(Cursor::new(buffer.get_ref()), &mut restored).unwrap();
assert_round_trip(&restored);
}
#[cfg(feature = "zip")]
#[test]
fn zip_deserialize_streams_into_folder_backend() {
use document_container::archive::Zip;
use std::io::Cursor;
let mut buffer = Cursor::new(Vec::new());
let mut writer = Zip::writer(&mut buffer).unwrap();
for (path, bytes) in entries() {
writer.write_entry(path, bytes).unwrap();
}
writer.finish().unwrap();
let dir = tempfile::tempdir().unwrap();
let mut restored = FolderBackend::create(dir.path()).unwrap();
<Zip as Archive>::open(Cursor::new(buffer.get_ref()), &mut restored).unwrap();
assert_eq!(restored.read("manifest.json").unwrap().as_slice(), br#"{"format":"gdd"}"#);
assert_eq!(restored.read("resources/abc123").unwrap().as_slice(), &[0xDE, 0xAD, 0xBE, 0xEF]);
}
#[cfg(feature = "xz")]
#[test]
fn xz_round_trip() {
use document_container::archive::Xz;
use std::io::Cursor;
let mut buffer = Cursor::new(Vec::new());
let mut writer = Xz::writer(&mut buffer).unwrap();
for (path, bytes) in entries() {
writer.write_entry(path, bytes).unwrap();
}
writer.finish().unwrap();
let mut restored = MemoryBackend::new();
<Xz as Archive>::open(Cursor::new(buffer.get_ref()), &mut restored).unwrap();
assert_round_trip(&restored);
}
@@ -0,0 +1,178 @@
use document_container::backends::folder::FolderBackend;
use document_container::backends::memory::MemoryBackend;
use document_container::{AnyContainer, Container, ContainerError};
fn run_round_trip<C: Container>(container: C) {
container.write("manifest.json", br#"{"format":"gdd"}"#).unwrap();
container.write("resources/abc123", &[0xDE, 0xAD, 0xBE, 0xEF]).unwrap();
container.write("resources/xyz789", b"another resource").unwrap();
assert!(container.exists("manifest.json"));
assert!(container.exists("resources/abc123"));
assert!(!container.exists("does-not-exist"));
let manifest = container.read("manifest.json").unwrap();
assert_eq!(manifest.as_slice(), br#"{"format":"gdd"}"#);
let blob = container.read("resources/abc123").unwrap();
assert_eq!(blob.as_slice(), &[0xDE, 0xAD, 0xBE, 0xEF]);
let top_level = container.list("").unwrap();
assert!(top_level.iter().any(|p| p == "manifest.json"));
assert!(
top_level.iter().all(|p| !p.starts_with("resources/")),
"list(\"\") must not descend into subdirectories, got {top_level:?}"
);
let mut resources = container.list("resources").unwrap();
resources.sort();
assert_eq!(resources, vec!["resources/abc123".to_string(), "resources/xyz789".to_string()]);
container.remove("resources/abc123").unwrap();
assert!(!container.exists("resources/abc123"));
assert!(matches!(container.read("resources/abc123"), Err(ContainerError::NotFound(_))));
}
#[test]
fn memory_backend_round_trip() {
run_round_trip(MemoryBackend::new());
}
#[test]
fn folder_backend_round_trip() {
let dir = tempfile::tempdir().unwrap();
let backend = FolderBackend::create(dir.path()).unwrap();
run_round_trip(backend);
}
fn run_list_and_remove_semantics<C: Container>(container: C) {
container.write("dir/file", b"x").unwrap();
// Removing a missing path is idempotent.
container.remove("dir/missing").unwrap();
container.remove("dir/file").unwrap();
container.remove("dir/file").unwrap();
// Listing a missing prefix yields an empty list.
assert_eq!(container.list("nonexistent").unwrap(), Vec::<String>::new());
// Listing a prefix that names a file is an error.
container.write("manifest.json", b"{}").unwrap();
assert!(matches!(container.list("manifest.json"), Err(ContainerError::NotADirectory(_))));
assert!(matches!(container.list_dirs("manifest.json"), Err(ContainerError::NotADirectory(_))));
}
#[test]
fn memory_backend_list_and_remove_semantics() {
run_list_and_remove_semantics(MemoryBackend::new());
}
#[test]
fn folder_backend_list_and_remove_semantics() {
let dir = tempfile::tempdir().unwrap();
run_list_and_remove_semantics(FolderBackend::create(dir.path()).unwrap());
}
#[test]
#[cfg(unix)]
fn folder_backend_rejects_symlink_escape() {
let outside = tempfile::tempdir().unwrap();
std::fs::write(outside.path().join("secret"), b"sensitive").unwrap();
let dir = tempfile::tempdir().unwrap();
let backend = FolderBackend::create(dir.path()).unwrap();
// A symlink planted inside the root pointing outside it must not be traversable.
std::os::unix::fs::symlink(outside.path(), dir.path().join("link")).unwrap();
let result = backend.read("link/secret");
assert!(matches!(result, Err(ContainerError::InvalidPath(_))), "symlink escape should be rejected, got {result:?}");
// Listing must reject the symlinked prefix too, not just reads, so it can't leak entry names from outside the root.
assert!(matches!(backend.list("link"), Err(ContainerError::InvalidPath(_))), "list through symlink should be rejected");
assert!(matches!(backend.list_dirs("link"), Err(ContainerError::InvalidPath(_))), "list_dirs through symlink should be rejected");
// A symlink entry directly under root must not appear in listings, since `resolve` would reject reading it.
std::os::unix::fs::symlink(outside.path().join("secret"), dir.path().join("file_link")).unwrap();
std::fs::write(dir.path().join("real"), b"ok").unwrap();
assert_eq!(backend.list("").unwrap(), vec!["real".to_string()], "symlink entries should be omitted from file listings");
assert!(backend.list_dirs("").unwrap().is_empty(), "symlink-to-dir entries should be omitted from dir listings");
}
#[test]
fn folder_backend_rejects_path_traversal() {
let dir = tempfile::tempdir().unwrap();
let backend = FolderBackend::create(dir.path()).unwrap();
for bad in ["../escape", "subdir/../escape", "/abs", "back\\slash", "a//b", "trailing/", "a/./b"] {
let result = backend.write(bad, b"nope");
assert!(matches!(result, Err(ContainerError::InvalidPath(_))), "expected InvalidPath for {bad:?}, got {result:?}");
}
}
fn run_append<C: Container>(container: C) {
// Appending to a non-existent path creates it — same semantics as `OpenOptions::append().create(true)`.
container.append("history.jsonl", b"{\"op\":1}\n").unwrap();
container.append("history.jsonl", b"{\"op\":2}\n").unwrap();
container.append("history.jsonl", b"{\"op\":3}\n").unwrap();
let log = container.read("history.jsonl").unwrap();
assert_eq!(log.as_slice(), b"{\"op\":1}\n{\"op\":2}\n{\"op\":3}\n");
}
#[test]
fn memory_backend_append() {
run_append(MemoryBackend::new());
}
#[test]
fn folder_backend_append() {
let dir = tempfile::tempdir().unwrap();
let backend = FolderBackend::create(dir.path()).unwrap();
run_append(backend);
}
#[test]
fn any_container_dispatches_to_active_variant() {
use document_container::AsyncContainer;
let container = AnyContainer::Memory(MemoryBackend::new());
futures::executor::block_on(async {
container.write("manifest.json", br#"{"format":"gdd"}"#).await.unwrap();
container.append("history.jsonl", b"frame-1\n").await.unwrap();
container.append("history.jsonl", b"frame-2\n").await.unwrap();
assert!(container.exists("manifest.json").await);
let manifest = container.read("manifest.json").await.unwrap();
assert_eq!(manifest.as_slice(), br#"{"format":"gdd"}"#);
let history = container.read("history.jsonl").await.unwrap();
assert_eq!(history.as_slice(), b"frame-1\nframe-2\n");
});
}
#[test]
fn folder_backend_reads_empty_file() {
let dir = tempfile::tempdir().unwrap();
let backend = FolderBackend::create(dir.path()).unwrap();
backend.write("empty.bin", &[]).unwrap();
let read_back = backend.read("empty.bin").unwrap();
assert_eq!(read_back.as_slice(), &[] as &[u8]);
}
#[test]
fn folder_backend_write_sized_fills_via_mmap() {
let dir = tempfile::tempdir().unwrap();
let backend = FolderBackend::create(dir.path()).unwrap();
let payload = b"hello world";
backend
.write_sized("resources/sized", payload.len(), &mut |buffer| {
buffer.copy_from_slice(payload);
Ok(())
})
.unwrap();
let read_back = backend.read("resources/sized").unwrap();
assert_eq!(read_back.as_slice(), payload);
}
@@ -0,0 +1,36 @@
[package]
name = "document-format"
description = "Typed handle for the .gdd document format, sitting over document-graph-storage and document-container"
edition.workspace = true
version.workspace = true
license.workspace = true
authors.workspace = true
[features]
# Runtime bridge: the editor <-> storage conversion methods (stage/commit from a `NodeNetwork`,
# `network_ids`, `declarations`). Off lets a standalone migration tool build without `graph-craft`
# or `core-types`. Forwards to `document-graph-storage/conversion`.
conversion = ["document-graph-storage/conversion", "dep:graph-craft", "dep:core-types"]
# Compressed-archive export/open. Each forwards to the matching `document-container` feature and
# gates that format's `ExportFormat` arm and sink. The folder/memory codec-free core needs neither.
zip = ["document-container/zip"]
xz = ["document-container/xz"]
default = ["conversion", "zip", "xz"]
[dependencies]
document-container = { workspace = true }
document-graph-storage = { workspace = true, default-features = false }
graph-craft = { workspace = true, optional = true }
graphene-resource = { workspace = true }
core-types = { workspace = true, optional = true }
serde = { workspace = true }
serde_json = { workspace = true }
rmp-serde = { workspace = true }
futures = { workspace = true }
thiserror = "2.0"
log = { workspace = true }
[dev-dependencies]
futures = { workspace = true }
graphene-resource = { workspace = true }
tempfile = "3"
@@ -0,0 +1,327 @@
//! Codec for a stream of values. Single-value writes are just streams of length one.
use serde::{Deserialize, Serialize, de::DeserializeOwned};
#[derive(Copy, Clone, Debug, PartialEq, Eq, Serialize, Deserialize)]
pub enum Codec {
/// A single JSON document. `append` to a non-empty buffer errors.
Json,
/// Newline-delimited compact JSON, one value per line.
JsonLines,
/// A single MessagePack blob. `append` to a non-empty buffer errors.
MessagePack,
/// Length-prefixed MessagePack frames: `[u32 big-endian length][MessagePack bytes]` per value.
MessagePackFrames,
}
#[derive(Debug, thiserror::Error)]
pub enum CodecError {
#[error("MessagePack encode error: {0}")]
MessagePackEncode(#[from] rmp_serde::encode::Error),
#[error("MessagePack decode error: {0}")]
MessagePackDecode(#[from] rmp_serde::decode::Error),
#[error("JSON error: {0}")]
Json(#[from] serde_json::Error),
#[error("frame length {0} exceeds u32")]
FrameTooLarge(usize),
#[error("frame length prefix truncated: need 4 bytes, have {0}")]
TruncatedLengthPrefix(usize),
#[error("declared frame length {declared} exceeds remaining buffer ({remaining} bytes)")]
TruncatedFrame { declared: usize, remaining: usize },
#[error("single-value codec cannot append to a non-empty buffer")]
SingleValueAlreadyWritten,
#[error("expected at least one value, got none")]
Empty,
#[error("expected exactly one value, got more")]
ExpectedSingle,
}
impl Codec {
pub fn extension(self) -> &'static str {
match self {
Codec::Json => "json",
Codec::JsonLines => "jsonl",
Codec::MessagePack => "bin",
Codec::MessagePackFrames => "frames",
}
}
/// Append one value to `output` in this codec's framing.
/// Single-value codecs error if `output` is non-empty.
pub fn append<T: Serialize>(self, output: &mut Vec<u8>, value: &T) -> Result<(), CodecError> {
match self {
Codec::Json => {
if !output.is_empty() {
return Err(CodecError::SingleValueAlreadyWritten);
}
serde_json::to_writer_pretty(output, value)?;
Ok(())
}
Codec::JsonLines => {
serde_json::to_writer(&mut *output, value)?;
output.push(b'\n');
Ok(())
}
Codec::MessagePack => {
if !output.is_empty() {
return Err(CodecError::SingleValueAlreadyWritten);
}
rmp_serde::encode::write(output, value)?;
Ok(())
}
Codec::MessagePackFrames => {
let payload = rmp_serde::to_vec(value)?;
let length = u32::try_from(payload.len()).map_err(|_| CodecError::FrameTooLarge(payload.len()))?;
output.extend_from_slice(&length.to_be_bytes());
output.extend_from_slice(&payload);
Ok(())
}
}
}
/// Iterate values from `bytes`. Single-value codecs yield exactly one item;
/// stream codecs yield however many were written.
pub fn iter<'a, T: DeserializeOwned + 'a>(self, bytes: &'a [u8]) -> Box<dyn Iterator<Item = Result<T, CodecError>> + 'a> {
match self {
Codec::Json => {
let single = serde_json::from_slice::<T>(bytes).map_err(CodecError::from);
Box::new(std::iter::once(single))
}
Codec::JsonLines => Box::new(JsonLineIter {
remaining: bytes,
_marker: std::marker::PhantomData,
}),
Codec::MessagePack => {
let single = rmp_serde::from_slice::<T>(bytes).map_err(CodecError::from);
Box::new(std::iter::once(single))
}
Codec::MessagePackFrames => Box::new(MessagePackFrameIter {
remaining: bytes,
_marker: std::marker::PhantomData,
}),
}
}
/// Serialize a single value into a fresh buffer.
pub fn write_single<T: Serialize>(self, value: &T) -> Result<Vec<u8>, CodecError> {
let mut output = Vec::new();
self.append(&mut output, value)?;
Ok(output)
}
/// Deserialize the single value in `bytes`. Errors if zero or more than one value is present.
pub fn read_single<T: DeserializeOwned>(self, bytes: &[u8]) -> Result<T, CodecError> {
let mut iter = self.iter::<T>(bytes);
let first = iter.next().ok_or(CodecError::Empty)??;
match iter.next() {
// A trailing decode error is more informative than `ExpectedSingle`, so surface it.
Some(Err(error)) => return Err(error),
Some(Ok(_)) => return Err(CodecError::ExpectedSingle),
None => {}
}
Ok(first)
}
}
struct JsonLineIter<'a, T> {
remaining: &'a [u8],
_marker: std::marker::PhantomData<fn() -> T>,
}
impl<T: DeserializeOwned> Iterator for JsonLineIter<'_, T> {
type Item = Result<T, CodecError>;
fn next(&mut self) -> Option<Self::Item> {
loop {
if self.remaining.is_empty() {
return None;
}
let (line, tail) = match self.remaining.iter().position(|&byte| byte == b'\n') {
Some(index) => (&self.remaining[..index], &self.remaining[index + 1..]),
None => (self.remaining, &[][..]),
};
self.remaining = tail;
let trimmed = trim_ascii(line);
if trimmed.is_empty() {
continue;
}
return Some(serde_json::from_slice(trimmed).map_err(CodecError::from));
}
}
}
struct MessagePackFrameIter<'a, T> {
remaining: &'a [u8],
_marker: std::marker::PhantomData<fn() -> T>,
}
impl<T: DeserializeOwned> Iterator for MessagePackFrameIter<'_, T> {
type Item = Result<T, CodecError>;
fn next(&mut self) -> Option<Self::Item> {
if self.remaining.is_empty() {
return None;
}
let buffer = std::mem::take(&mut self.remaining);
let Some((length_bytes, tail)) = buffer.split_first_chunk::<4>() else {
return Some(Err(CodecError::TruncatedLengthPrefix(buffer.len())));
};
let length = u32::from_be_bytes(*length_bytes) as usize;
if tail.len() < length {
return Some(Err(CodecError::TruncatedFrame {
declared: length,
remaining: tail.len(),
}));
}
let (frame, after) = tail.split_at(length);
self.remaining = after;
Some(rmp_serde::from_slice(frame).map_err(CodecError::from))
}
}
fn trim_ascii(bytes: &[u8]) -> &[u8] {
let start = bytes.iter().position(|byte| !byte.is_ascii_whitespace()).unwrap_or(bytes.len());
let end = bytes.iter().rposition(|byte| !byte.is_ascii_whitespace()).map(|index| index + 1).unwrap_or(start);
&bytes[start..end]
}
#[cfg(test)]
mod tests {
use super::*;
use serde::{Deserialize, Serialize};
#[derive(Debug, PartialEq, Eq, Serialize, Deserialize)]
struct Frame {
id: u32,
label: String,
}
fn frames() -> [Frame; 3] {
[Frame { id: 1, label: "alpha".into() }, Frame { id: 2, label: "beta".into() }, Frame { id: 3, label: "gamma".into() }]
}
#[test]
fn json_round_trip_single() {
let frame = Frame { id: 7, label: "solo".into() };
let bytes = Codec::Json.write_single(&frame).unwrap();
let decoded: Frame = Codec::Json.read_single(&bytes).unwrap();
assert_eq!(decoded, frame);
}
#[test]
fn json_append_to_non_empty_errors() {
let mut buffer = b"already here".to_vec();
let result = Codec::Json.append(&mut buffer, &Frame { id: 1, label: "x".into() });
assert!(matches!(result, Err(CodecError::SingleValueAlreadyWritten)), "got {result:?}");
}
#[test]
fn message_pack_round_trip_single() {
let frame = Frame { id: 99, label: "blob".into() };
let bytes = Codec::MessagePack.write_single(&frame).unwrap();
let decoded: Frame = Codec::MessagePack.read_single(&bytes).unwrap();
assert_eq!(decoded, frame);
}
#[test]
fn message_pack_append_to_non_empty_errors() {
let mut buffer = vec![0xAB];
let result = Codec::MessagePack.append(&mut buffer, &Frame { id: 1, label: "x".into() });
assert!(matches!(result, Err(CodecError::SingleValueAlreadyWritten)), "got {result:?}");
}
/// A type-erased `serde_json::Value` round-trips through the binary codec: the property postcard
/// could not satisfy (it raises `WontImplement` on self-describing values), which is why the
/// resource/attribute deltas that carry `serde_json::Value` bodies need a self-describing codec.
#[test]
fn message_pack_round_trips_serde_json_value() {
let value = serde_json::json!({ "kind": "embedded", "priority": 1.5, "tags": ["a", "b"] });
let bytes = Codec::MessagePack.write_single(&value).unwrap();
let decoded: serde_json::Value = Codec::MessagePack.read_single(&bytes).unwrap();
assert_eq!(decoded, value);
}
#[test]
fn json_lines_round_trip_and_skip_blanks() {
let frames = [Frame { id: 1, label: "alpha".into() }, Frame { id: 2, label: "beta".into() }];
let mut buffer = Vec::new();
Codec::JsonLines.append(&mut buffer, &frames[0]).unwrap();
buffer.extend_from_slice(b" \n\n");
Codec::JsonLines.append(&mut buffer, &frames[1]).unwrap();
let decoded: Vec<Frame> = Codec::JsonLines.iter(&buffer).collect::<Result<_, _>>().unwrap();
assert_eq!(decoded, frames);
}
#[test]
fn message_pack_frames_round_trip() {
let frames = frames();
let mut buffer = Vec::new();
for frame in &frames {
Codec::MessagePackFrames.append(&mut buffer, frame).unwrap();
}
let decoded: Vec<Frame> = Codec::MessagePackFrames.iter(&buffer).collect::<Result<_, _>>().unwrap();
assert_eq!(decoded, frames);
}
/// A crash mid-append leaves a torn final frame. The length prefix lets us detect that
/// deterministically (declared length exceeds the bytes that actually made it to disk) rather
/// than decoding a partial value into a plausible-but-wrong one.
#[test]
fn message_pack_frames_detect_truncation() {
let mut buffer = Vec::new();
Codec::MessagePackFrames.append(&mut buffer, &Frame { id: 7, label: "ok".into() }).unwrap();
buffer.truncate(buffer.len() - 1);
let last = Codec::MessagePackFrames.iter::<Frame>(&buffer).last().unwrap();
assert!(matches!(last, Err(CodecError::TruncatedFrame { .. })), "got {last:?}");
}
/// A buffer whose first record's length prefix itself is incomplete (fewer than 4 bytes) is
/// reported as a truncated prefix rather than mis-read as a zero-length frame.
#[test]
fn message_pack_frames_detect_truncated_length_prefix() {
let buffer = vec![0x00, 0x00];
let last = Codec::MessagePackFrames.iter::<Frame>(&buffer).last().unwrap();
assert!(matches!(last, Err(CodecError::TruncatedLengthPrefix(2))), "got {last:?}");
}
#[test]
fn write_single_then_read_with_iter_yields_one() {
let frame = Frame { id: 5, label: "one".into() };
for codec in [Codec::Json, Codec::JsonLines, Codec::MessagePack, Codec::MessagePackFrames] {
let bytes = codec.write_single(&frame).unwrap();
let collected: Vec<Frame> = codec.iter(&bytes).collect::<Result<_, _>>().unwrap();
assert_eq!(collected, vec![Frame { id: 5, label: "one".into() }], "codec {codec:?}");
}
}
#[test]
fn read_single_rejects_multi_value_stream() {
let mut buffer = Vec::new();
Codec::JsonLines.append(&mut buffer, &Frame { id: 1, label: "a".into() }).unwrap();
Codec::JsonLines.append(&mut buffer, &Frame { id: 2, label: "b".into() }).unwrap();
let result: Result<Frame, _> = Codec::JsonLines.read_single(&buffer);
assert!(matches!(result, Err(CodecError::ExpectedSingle)), "got {result:?}");
}
#[test]
fn extensions_are_distinct() {
let exts = [
Codec::Json.extension(),
Codec::JsonLines.extension(),
Codec::MessagePack.extension(),
Codec::MessagePackFrames.extension(),
];
let unique: std::collections::HashSet<_> = exts.iter().collect();
assert_eq!(unique.len(), exts.len(), "extensions collide: {exts:?}");
}
}
@@ -0,0 +1,50 @@
//! Unified error type for the `document-format` crate.
//!
//! Every fallible [`crate::Gdd`] method returns [`Result<T>`]. Variants are grouped by failure
//! domain (container I/O, codec, CRDT, format validation, export)
use document_container::ContainerError;
#[cfg(feature = "conversion")]
use document_graph_storage::CommitError;
use document_graph_storage::CrdtError;
use graphene_resource::ResourceHash;
use crate::codec::CodecError;
use crate::io::ReadError;
/// Crate-wide result alias.
pub type Result<T> = std::result::Result<T, Error>;
/// Anything that can go wrong reading, mutating, or exporting a `.gdd` document. Per the format's
/// load-time policy, any unexpected condition is a hard error rather than a silent fallback.
#[derive(Debug, thiserror::Error)]
pub enum Error {
/// Working-copy container I/O failed (read, write, or path validation).
#[error("container error: {0}")]
Container(#[from] ContainerError),
/// A typed payload could not be located or read from the container.
#[error("read error: {0}")]
Read(#[from] ReadError),
/// A payload failed to encode or decode in its recorded codec.
#[error("codec error: {0}")]
Codec(#[from] CodecError),
/// A CRDT operation was rejected while replaying a hot op or moving the undo/redo cursor.
#[error("CRDT error: {0}")]
Crdt(#[from] CrdtError),
/// Staging a runtime snapshot into the session failed (conversion or CRDT apply).
#[cfg(feature = "conversion")]
#[error("commit error: {0}")]
Commit(#[from] CommitError),
/// The manifest's `format` field is not the `.gdd` magic, so this is not a `.gdd` document.
#[error("not a .gdd document (manifest format = {found:?}, expected {expected:?})")]
WrongFormat { found: String, expected: &'static str },
/// The manifest declares a format version newer than this build can open.
#[error("unsupported format version: found {found}, max supported {max_supported}")]
UnsupportedVersion { found: u32, max_supported: u32 },
/// The requested export options are incoherent (e.g. neither registry nor history included).
#[error("invalid export options: {0}")]
InvalidExportOptions(&'static str),
/// An export marked a resource for embedding but its bytes were absent from the byte store.
#[error("embedded resource {0} missing from the byte store")]
MissingResource(ResourceHash),
}
@@ -0,0 +1,310 @@
//! Export: walking the working copy through an archive codec, keeping payloads as-is.
//!
//! [`ExportFormat`] / [`ExportOptions`] are the public settings; the [`Gdd`] export methods drive a
//! [`ExportSink`] (folder / zip / xz) through the manifest → registry → history → resources sequence.
#[cfg(not(target_family = "wasm"))]
use std::path::Path;
use graphene_resource::{LoadResource, Resource, ResourceHash};
use crate::error::Error;
use crate::layout::Layout;
use crate::session_state::SessionState;
use crate::{Gdd, MANIFEST_CODEC, io};
/// Export wrapping. Payloads keep the working copy's recorded per-payload codecs (see
/// [`crate::manifest::PayloadCodecs`]); export does not re-encode.
#[derive(Copy, Clone, Debug)]
pub enum ExportFormat {
/// Copy the working copy to a destination folder.
Folder,
/// Wrap as a `.gdd.zip` archive.
#[cfg(feature = "zip")]
Zip,
/// Wrap as a `.gdd.tar.xz` archive (whole-archive xz via `lzma-rust2`).
#[cfg(feature = "xz")]
Xz,
}
#[derive(Copy, Clone, Debug)]
pub struct ExportOptions {
/// Whether to include the registry snapshot. `false` produces a history-only export, useful
/// for VCS workflows where the diffable `history.jsonl` is the interesting payload and the
/// registry would rewrite whole-file on every retirement. Consumers replay history from an
/// empty registry.
pub include_registry: bool,
/// Whether to include history + hot-log. `false` produces a flat snapshot (registry only),
/// useful for sharing without revealing edit history and for cutting file size.
pub include_history: bool,
/// Materialize every non-`DataSource::Embedded` resource into `resources/<hash>` for portability.
/// Does not mutate the in-memory `Gdd`.
pub embed_all_resources: bool,
}
impl ExportOptions {
/// Returns an error description if the combination is incoherent.
pub fn validate(&self) -> Result<(), &'static str> {
if !self.include_registry && !self.include_history {
return Err("export must include at least one of: registry, history");
}
Ok(())
}
}
impl Default for ExportOptions {
fn default() -> Self {
Self {
include_registry: true,
include_history: true,
embed_all_resources: false,
}
}
}
impl<L: Layout> Gdd<L> {
/// Stream the working copy to `dest` as a folder/zip/xz archive, keeping payload codecs as-is.
/// Does not mutate `self` and does not buffer the export. Native-only (writes a filesystem path).
///
/// # Errors
/// [`Error::InvalidExportOptions`] for incoherent options, [`Error::MissingResource`] if an
/// embedded resource's bytes are absent from `byte_store`.
#[cfg(not(target_family = "wasm"))]
pub async fn export(&self, dest: &Path, format: ExportFormat, options: ExportOptions, byte_store: &dyn LoadResource, legacy_document: Option<&[u8]>) -> Result<(), Error> {
options.validate().map_err(Error::InvalidExportOptions)?;
match format {
ExportFormat::Folder => {
let mut folder = document_container::backends::folder::FolderBackend::create(dest)?;
let mut sink = FolderSink { folder: &mut folder };
self.stream_entries(options, byte_store, &mut sink).await?;
if let Some(legacy) = legacy_document {
sink.write_entry(self.layout.legacy_path(), legacy)?;
}
}
#[cfg(feature = "zip")]
ExportFormat::Zip => {
let file = std::fs::File::create(dest).map_err(document_container::ContainerError::Io)?;
self.export_archive::<document_container::archive::Zip, _>(file, options, byte_store, legacy_document).await?;
}
#[cfg(feature = "xz")]
ExportFormat::Xz => {
let file = std::fs::File::create(dest).map_err(document_container::ContainerError::Io)?;
self.export_archive::<document_container::archive::Xz, _>(file, options, byte_store, legacy_document).await?;
}
}
Ok(())
}
/// In-memory variant of [`export`](Self::export) returning the archive bytes. Available on every
/// target (no `std::fs`) but buffers the whole archive. `legacy_document` is embedded verbatim at
/// [`Layout::legacy_path`]. `ExportFormat::Folder` has no single-file form and is rejected.
#[cfg(any(feature = "zip", feature = "xz"))]
pub async fn export_to_bytes(&self, format: ExportFormat, options: ExportOptions, byte_store: &dyn LoadResource, legacy_document: Option<&[u8]>) -> Result<Vec<u8>, Error> {
options.validate().map_err(Error::InvalidExportOptions)?;
let cursor = std::io::Cursor::new(Vec::new());
let buffer = match format {
ExportFormat::Folder => return Err(Error::InvalidExportOptions("folder export has no single-file byte form")),
#[cfg(feature = "zip")]
ExportFormat::Zip => self.export_archive::<document_container::archive::Zip, _>(cursor, options, byte_store, legacy_document).await?,
#[cfg(feature = "xz")]
ExportFormat::Xz => self.export_archive::<document_container::archive::Xz, _>(cursor, options, byte_store, legacy_document).await?,
};
Ok(buffer.into_inner())
}
/// Stream entries into a fresh `A` archive over `output`, append the optional legacy blob, then
/// finalize and hand back the inner sink. The shared body of both archive export paths.
#[cfg(any(feature = "zip", feature = "xz"))]
async fn export_archive<A, W>(&self, output: W, options: ExportOptions, byte_store: &dyn LoadResource, legacy_document: Option<&[u8]>) -> Result<W, Error>
where
A: document_container::archive::Archive,
W: std::io::Write + std::io::Seek + Send,
A::Writer<W>: ExportSink + document_container::archive::ArchiveWriter<Sink = W>,
{
use document_container::archive::ArchiveWriter;
let mut writer = A::writer(output)?;
self.stream_entries(options, byte_store, &mut writer).await?;
if let Some(legacy) = legacy_document {
ExportSink::write_entry(&mut writer, self.layout.legacy_path(), legacy)?;
}
Ok(writer.finish_into()?)
}
/// Drive a sink through manifest → session → registry → history → resources, one entry at a time,
/// keeping each payload's recorded codec.
async fn stream_entries(&self, options: ExportOptions, byte_store: &dyn LoadResource, sink: &mut dyn ExportSink) -> Result<(), Error> {
use document_container::AsyncContainer;
let codecs = self.manifest.codecs;
sink.write_entry(&io::path_for(self.layout.manifest_basename(), MANIFEST_CODEC), &MANIFEST_CODEC.write_single(&self.manifest)?)?;
// Carry the per-peer cursor + view settings so a `.gdd` reopened elsewhere restores the viewport.
let session_state = SessionState {
peer_id: self.session.peer(),
head_rev: self.session.head_rev(),
last_broadcast_rev: self.session.last_broadcast_rev(),
redo_stack: self.session.redo_stack().to_vec(),
next_node_counter: self.session.next_node_counter(),
view_settings: self.view_settings.clone(),
network_view_settings: self.network_view_settings.clone(),
};
sink.write_entry(&io::path_for(self.layout.session_basename(), codecs.session), &codecs.session.write_single(&session_state)?)?;
let working_copy_hashes: std::collections::HashSet<ResourceHash> = self.resource_hashes().await?.into_iter().collect();
// Resources to embed as bytes: every `Embedded` entry, plus link-only ones when
// `embed_all_resources`. Bytes already in the working copy are written by the copy-through pass
// below, so only the gap is loaded from the byte store here.
let mut export_session = self.session.clone();
let mut hashes_from_store: Vec<ResourceHash> = Vec::new();
let mut links_to_promote: Vec<document_graph_storage::ResourceId> = Vec::new();
for (id, entry) in &export_session.registry().resources {
let Some(hash) = entry.hash else { continue };
let embed = entry.has_embedded_source() || options.embed_all_resources;
if !embed {
continue;
}
if !entry.has_embedded_source() {
links_to_promote.push(*id);
}
if !working_copy_hashes.contains(&hash) {
hashes_from_store.push(hash);
}
}
hashes_from_store.sort_unstable();
hashes_from_store.dedup();
// Fail fast if an embedded resource is missing, then promote link-only sources on the clone so
// the exported registry and history stay consistent. The live `Gdd` is untouched.
let mut embedded_bytes: Vec<(ResourceHash, Resource)> = Vec::new();
for hash in hashes_from_store {
let Some(resource) = byte_store.load(hash).await else {
return Err(Error::MissingResource(hash));
};
embedded_bytes.push((hash, resource));
}
export_session.embed_resource_sources(links_to_promote)?;
if options.include_registry {
// With history, the persisted snapshot is the retired registry and the hot log layers on top
// (mirrors `Session::load`); without history it must be the full working registry, since
// `bootstrap_from_registry` reconstructs the whole document from it alone.
let snapshot = if options.include_history { export_session.retired_registry() } else { export_session.registry() };
sink.write_entry(&io::path_for(self.layout.registry_basename(), codecs.registry), &codecs.registry.write_single(snapshot)?)?;
}
if options.include_history {
let mut buffer = Vec::new();
for delta in export_session.history() {
codecs.history.append(&mut buffer, delta)?;
}
if !buffer.is_empty() {
sink.write_entry(&io::path_for(self.layout.history_basename(), codecs.history), &buffer)?;
}
// Carry the un-retired hot ops alongside history so a document exported mid-interaction (e.g. a save
// during a tool drag) isn't shipped with its pending edits dropped. `open` replays them on top of
// the retired snapshot, same as the working copy does.
let mut hot_buffer = Vec::new();
for hot_op in export_session.hot_log() {
codecs.hot_log.append(&mut hot_buffer, hot_op)?;
}
if !hot_buffer.is_empty() {
sink.write_entry(&io::path_for(self.layout.hot_log_basename(), codecs.hot_log), &hot_buffer)?;
}
}
// Copy bytes the working copy already holds, tracking covered hashes so the embed pass below
// doesn't re-emit them.
let mut emitted = std::collections::HashSet::new();
let resources_dir = self.layout.resources_dir();
if self.working.list_dirs("").await?.iter().any(|d| d == resources_dir) {
let prefix = format!("{resources_dir}/");
for path in self.working.list(resources_dir).await? {
if let Some(hash) = path.strip_prefix(&prefix).and_then(|name| name.parse::<ResourceHash>().ok()) {
emitted.insert(hash);
}
let holder = self.working.read(&path).await?;
// Native `External` (mmap'd) holders copy CoW via a source path; others write bytes.
#[cfg(not(target_family = "wasm"))]
match holder.source_path() {
Some(src_path) => sink.write_entry_from_path(&path, src_path)?,
None => sink.write_entry(&path, holder.as_slice())?,
}
#[cfg(target_family = "wasm")]
sink.write_entry(&path, holder.as_slice())?;
}
}
for (hash, resource) in &embedded_bytes {
if emitted.insert(*hash) {
sink.write_entry(&self.layout.resource_path(hash), resource.as_ref())?;
}
}
Ok(())
}
}
/// Sink an export streams entries into, so one async loop drives folder/zip/xz writes. Archive sinks
/// work on every target; the folder sink and `write_entry_from_path` are native-only. `Send` because
/// `stream_entries` holds `&mut dyn ExportSink` across `.await`s.
pub(crate) trait ExportSink: Send {
fn write_entry(&mut self, path: &str, bytes: &[u8]) -> Result<(), Error>;
/// Copy a file from disk into the sink. Default reads it into memory; the folder sink overrides to
/// `fs::copy` (CoW). Native-only: only reachable for an `External` (mmap'd) holder.
#[cfg(not(target_family = "wasm"))]
fn write_entry_from_path(&mut self, path: &str, src: &std::path::Path) -> Result<(), Error> {
let bytes = std::fs::read(src).map_err(document_container::ContainerError::Io)?;
self.write_entry(path, &bytes)
}
}
#[cfg(not(target_family = "wasm"))]
struct FolderSink<'a> {
folder: &'a mut document_container::backends::folder::FolderBackend,
}
#[cfg(not(target_family = "wasm"))]
impl ExportSink for FolderSink<'_> {
fn write_entry(&mut self, path: &str, bytes: &[u8]) -> Result<(), Error> {
document_container::Container::write(self.folder, path, bytes)?;
Ok(())
}
fn write_entry_from_path(&mut self, path: &str, src: &std::path::Path) -> Result<(), Error> {
document_container::validate_path(path)?;
let dest = self.folder.root().join(path);
if let Some(parent) = dest.parent() {
std::fs::create_dir_all(parent).map_err(document_container::ContainerError::Io)?;
}
std::fs::copy(src, &dest).map_err(document_container::ContainerError::Io)?;
Ok(())
}
}
#[cfg(feature = "zip")]
impl<W: std::io::Write + std::io::Seek + Send> ExportSink for document_container::archive::ZipWriter<W> {
fn write_entry(&mut self, path: &str, bytes: &[u8]) -> Result<(), Error> {
use document_container::archive::ArchiveWriter;
ArchiveWriter::write_entry(self, path, bytes)?;
Ok(())
}
}
#[cfg(feature = "xz")]
impl<W: std::io::Write + std::io::Seek + Send> ExportSink for document_container::archive::XzWriter<W> {
fn write_entry(&mut self, path: &str, bytes: &[u8]) -> Result<(), Error> {
use document_container::archive::ArchiveWriter;
ArchiveWriter::write_entry(self, path, bytes)?;
Ok(())
}
}
@@ -0,0 +1,60 @@
//! Bridge between [`crate::Codec`] and [`document_container::AnyContainer`]. Each payload's codec
//! is known up front (the manifest is always JSON; every other payload's codec is recorded in the
//! manifest), so reads and writes address a fixed `{basename}.{ext}` path without probing.
use document_container::{AnyContainer, AsyncContainer};
use serde::Serialize;
use serde::de::DeserializeOwned;
use crate::{Codec, CodecError};
/// Compose a container path from `basename` and `codec.extension()`.
pub fn path_for(basename: &str, codec: Codec) -> String {
format!("{basename}.{}", codec.extension())
}
#[derive(Debug, thiserror::Error)]
pub enum ReadError {
#[error("file not found for basename {basename:?} with codec {codec:?}")]
NotFound { basename: String, codec: Codec },
#[error("container error: {0}")]
Container(#[from] document_container::ContainerError),
#[error("codec error: {0}")]
Codec(#[from] CodecError),
}
/// Read `{basename}.{ext}` and decode the single value it contains.
pub async fn read_single<T: DeserializeOwned>(container: &AnyContainer, basename: &str, codec: Codec) -> Result<T, ReadError> {
let bytes = read_bytes(container, basename, codec).await?;
Ok(codec.read_single::<T>(bytes.as_slice())?)
}
/// Same as [`read_single`] but yields every value when `codec` is a stream codec.
pub async fn iter<T: DeserializeOwned>(container: &AnyContainer, basename: &str, codec: Codec) -> Result<Vec<T>, ReadError> {
let bytes = read_bytes(container, basename, codec).await?;
Ok(codec.iter::<T>(bytes.as_slice()).collect::<Result<Vec<_>, _>>()?)
}
/// Whether `{basename}.{ext}` exists for the given codec.
pub async fn exists(container: &AnyContainer, basename: &str, codec: Codec) -> bool {
container.exists(&path_for(basename, codec)).await
}
/// Encode `value` with `codec` and write to `{basename}.{ext}`. Synchronous: the write goes through
/// the container's sync write surface (durable on folder/memory, enqueued on OPFS).
pub fn write_single<T: Serialize>(container: &AnyContainer, basename: &str, codec: Codec, value: &T) -> Result<(), crate::Error> {
let bytes = codec.write_single(value)?;
container.write_non_blocking(&path_for(basename, codec), &bytes)?;
Ok(())
}
async fn read_bytes(container: &AnyContainer, basename: &str, codec: Codec) -> Result<document_container::ByteHolder, ReadError> {
let path = path_for(basename, codec);
if !container.exists(&path).await {
return Err(ReadError::NotFound {
basename: basename.to_string(),
codec,
});
}
Ok(container.read(&path).await?)
}
@@ -0,0 +1,51 @@
//! Path layout for a `.gdd` working copy.
//!
//! Layout owns basenames only — the codec choice for each payload is a runtime parameter at the
//! read/write call site. Working-copy creation, exports, and migrations may all hit the same
//! basename with different codecs.
use graphene_resource::ResourceHash;
pub trait Layout {
fn manifest_basename(&self) -> &str;
fn session_basename(&self) -> &str;
fn registry_basename(&self) -> &str;
fn history_basename(&self) -> &str;
fn hot_log_basename(&self) -> &str;
fn resources_dir(&self) -> &str;
fn resource_path(&self, hash: &ResourceHash) -> String;
/// The embedded legacy `.graphite` document, stored verbatim during the dual-write soak so the
/// new format can be validated against (and recovered from) the old one. Dropped once `.gdd`
/// becomes the sole source of truth.
fn legacy_path(&self) -> &str;
}
#[derive(Copy, Clone, Debug, Default)]
pub struct GddV1Layout;
impl Layout for GddV1Layout {
fn manifest_basename(&self) -> &str {
"manifest"
}
fn session_basename(&self) -> &str {
"session"
}
fn registry_basename(&self) -> &str {
"registry"
}
fn history_basename(&self) -> &str {
"history"
}
fn hot_log_basename(&self) -> &str {
"hot-log"
}
fn resources_dir(&self) -> &str {
"resources"
}
fn resource_path(&self, hash: &ResourceHash) -> String {
format!("{}/{hash}", self.resources_dir())
}
fn legacy_path(&self) -> &str {
"legacy.graphite"
}
}
@@ -0,0 +1,344 @@
//! Typed handle for `.gdd` documents.
//!
//! [`Gdd`] owns a [`document_graph_storage::Session`] plus a working-copy [`document_container::AnyContainer`].
//! Mutations flow through `Gdd` to keep the session and the on-disk working copy mirrored.
//! Export is a separate, explicit operation — see [`export::ExportFormat`].
//!
//! See the "On-disk container" section of `node-graph/rfcs/document-format.md` for the format spec.
use std::sync::Arc;
// `Path` and `FolderBackend` are only used by the native-only path-based open/create, so they're
// gated off wasm to avoid unused-import warnings.
#[cfg(not(target_family = "wasm"))]
use std::path::Path;
#[cfg(not(target_family = "wasm"))]
use document_container::backends::folder::FolderBackend;
use document_container::{AnyContainer, AsyncContainer, ByteHolder, ContainerError};
#[cfg(feature = "conversion")]
use document_graph_storage::{CommitError, NodeMetadataSource};
use document_graph_storage::{Delta, HotOp, PeerId, Registry, Session};
#[cfg(feature = "conversion")]
use graphene_resource::LoadResource;
use graphene_resource::ResourceHash;
pub mod codec;
pub mod error;
pub mod export;
pub mod io;
pub mod layout;
pub mod manifest;
pub mod persist;
pub mod resource;
pub mod session_state;
pub use codec::{Codec, CodecError};
pub use error::Error;
pub use export::{ExportFormat, ExportOptions};
pub use io::ReadError;
pub use layout::{GddV1Layout, Layout};
pub use manifest::{Manifest, PayloadCodecs};
pub use resource::ResourceProxy;
pub use session_state::SessionState;
/// The default [`Layout`], so callers write `GddV1` for the common `Gdd<GddV1Layout>` handle.
pub type GddV1 = Gdd<GddV1Layout>;
/// The manifest is always JSON: it is the bootstrap file, read before any other payload's codec is
/// known, so its own codec cannot itself be configurable.
pub const MANIFEST_CODEC: Codec = Codec::Json;
/// Working-copy codecs. The working copy lives in appdata, not under VCS — these defaults
/// optimize for size and write cost. MessagePack is self-describing, so it round-trips the
/// type-erased `serde_json::Value` bodies that resource and attribute deltas carry (a non-self-
/// describing format like postcard cannot). JSON/JSONL is opt-in via `ExportFormat::Folder` for
/// users who want a diffable on-disk representation. Recorded in the manifest at create time and
/// read back on open (see [`manifest::PayloadCodecs`]), so the persist path never probes the filesystem.
pub const DEFAULT_SESSION_CODEC: Codec = Codec::Json;
pub const DEFAULT_REGISTRY_CODEC: Codec = Codec::MessagePack;
pub const DEFAULT_HISTORY_CODEC: Codec = Codec::MessagePackFrames;
pub const DEFAULT_HOT_LOG_CODEC: Codec = Codec::MessagePackFrames;
/// Editor-facing handle. Owns the `Session` and the working-copy container; mutations are mirrored
/// to disk continuously (every retirement appends to the history file and re-snapshots the registry).
///
/// The per-edit persist path (`commit_from_runtime`, `apply_hot_op`, `retire`) is synchronous and
/// read-free: the manifest is cached in memory (so payload codecs need no disk read), and writes go
/// through the container's sync write surface. Only `open` / `create` / `export` are async, since they
/// read.
/// `Clone` shares the working-copy container (`Arc<AnyContainer>`) so a cloned handle reads and writes
/// the *same* on-disk/OPFS working copy — including any writes still queued on the OPFS backend. The
/// `Session` is cloned (a snapshot copy); the container is shared.
#[derive(Clone)]
pub struct Gdd<L: Layout = GddV1Layout> {
pub(crate) session: Session,
pub(crate) working: Arc<AnyContainer>,
pub(crate) layout: L,
/// In-memory copy of the manifest, kept authoritative since `Gdd` is its sole writer. Holds the
/// per-payload codecs so the persist path never probes the filesystem, keeping it fully read-free
/// and synchronous.
pub(crate) manifest: Manifest,
/// Per-peer view settings (PTZ, rulers, etc.), persisted in `session.json` not the registry, so
/// they stay out of the CRDT/history. Opaque to the storage layer; the editor owns the keys/values.
pub(crate) view_settings: std::collections::BTreeMap<String, serde_json::Value>,
/// Per-network view settings (node-graph nav + previewing), keyed by stable [`NetworkId`]. Same per-peer
/// `session.json` treatment as [`view_settings`](Self::view_settings), but scoped per network.
pub(crate) network_view_settings: std::collections::BTreeMap<document_graph_storage::NetworkId, std::collections::BTreeMap<String, serde_json::Value>>,
}
/// Native folder-backed convenience constructors. On wasm the editor builds an OPFS-backed
/// `AnyContainer` itself and uses [`Gdd::open_in`] / [`Gdd::create_in`] directly.
#[cfg(not(target_family = "wasm"))]
impl<L: Layout + Default> Gdd<L> {
/// Open an existing working copy at `path`. Validates the manifest, materializes the session
/// from `registry.bin` (fast path) or by replaying `history.jsonl` (slow path), then applies
/// the persisted hot log on top.
pub async fn open(path: &Path) -> Result<Self, Error> {
let working = AnyContainer::Folder(FolderBackend::open(path)?);
let layout = L::default();
Self::open_in(working, layout).await
}
/// Create a fresh, empty working copy at `path` bound to `peer`. Writes a default manifest
/// and session state; the caller fills in editor metadata via [`Gdd::update_manifest`].
pub async fn create(path: &Path, peer: PeerId, document_uuid: u64, editor_version: String, stdlib_version: String) -> Result<Self, Error> {
let working = AnyContainer::Folder(FolderBackend::create(path)?);
let layout = L::default();
Self::create_in(working, layout, peer, document_uuid, editor_version, stdlib_version).await
}
}
impl<L: Layout> Gdd<L> {
/// Open a `.gdd` from archive bytes (xz or zip, auto-detected) by materializing it into `working`,
/// then opening it as a working copy. The archive is deserialized into an in-memory staging backend
/// (the archive reader is synchronous), then each entry is written into `working` via the sync
/// `write_non_blocking` surface — durable on folder/memory, eagerly enqueued on OPFS. `working` is
/// expected to be a fresh per-document container; entries with colliding paths are overwritten.
#[cfg(any(feature = "zip", feature = "xz"))]
pub async fn open_from_archive(bytes: &[u8], mut working: AnyContainer, layout: L) -> Result<Self, Error> {
document_container::archive::open_auto(bytes, &mut working)?;
Self::open_in(working, layout).await
}
/// Backend-agnostic open. Splits out so tests can supply a [`document_container::backends::memory::MemoryBackend`].
///
/// # Errors
/// [`Error::WrongFormat`] / [`Error::UnsupportedVersion`] if the manifest fails validation, plus
/// the usual [`Error::Read`] / [`Error::Codec`] / [`Error::Crdt`] if a payload is malformed.
pub async fn open_in(working: AnyContainer, layout: L) -> Result<Self, Error> {
let manifest: Manifest = io::read_single(&working, layout.manifest_basename(), MANIFEST_CODEC).await?;
validate_manifest(&manifest)?;
let codecs = manifest.codecs;
let session_state: SessionState = match io::exists(&working, layout.session_basename(), codecs.session).await {
true => io::read_single(&working, layout.session_basename(), codecs.session).await?,
false => SessionState::default(),
};
let has_registry = io::exists(&working, layout.registry_basename(), codecs.registry).await;
let has_history = io::exists(&working, layout.history_basename(), codecs.history).await;
let peer = session_state.peer_id;
let mut session = match (has_registry, has_history) {
(true, true) => {
let registry: Registry = io::read_single(&working, layout.registry_basename(), codecs.registry).await?;
let history = load_history(&working, &layout, codecs.history).await?;
Session::load(peer, registry, history, session_state.head_rev, session_state.redo_stack, session_state.next_node_counter)
}
(true, false) => {
// Registry-only export: synthesize a history that reproduces this state.
let registry: Registry = io::read_single(&working, layout.registry_basename(), codecs.registry).await?;
Session::bootstrap_from_registry(peer, registry)?
}
(false, _) => Session::replay_from_history(peer, load_history(&working, &layout, codecs.history).await?, session_state.next_node_counter)?,
};
// Restore the published frontier (silent/published undo boundary) regardless of which load arm ran.
if let Some(rev) = session_state.last_broadcast_rev {
session.publish_up_to(rev);
}
replay_hot_log(&working, &layout, codecs.hot_log, &mut session).await?;
Ok(Self {
session,
working: Arc::new(working),
layout,
manifest,
view_settings: session_state.view_settings,
network_view_settings: session_state.network_view_settings,
})
}
/// Backend-agnostic create. Records the working-copy default codecs (see `DEFAULT_*_CODEC`) in
/// the manifest and writes each payload with its recorded codec.
pub async fn create_in(working: AnyContainer, layout: L, peer: PeerId, document_uuid: u64, editor_version: String, stdlib_version: String) -> Result<Self, Error> {
let manifest = Manifest::new(document_uuid, editor_version, stdlib_version);
let codecs = manifest.codecs;
io::write_single(&working, layout.manifest_basename(), MANIFEST_CODEC, &manifest)?;
let session_state = SessionState { peer_id: peer, ..Default::default() };
io::write_single(&working, layout.session_basename(), codecs.session, &session_state)?;
let session = Session::with_peer(peer);
io::write_single(&working, layout.registry_basename(), codecs.registry, session.registry())?;
Ok(Self {
session,
working: Arc::new(working),
layout,
manifest,
view_settings: std::collections::BTreeMap::new(),
network_view_settings: std::collections::BTreeMap::new(),
})
}
}
fn validate_manifest(manifest: &Manifest) -> Result<(), Error> {
if manifest.format != manifest::FORMAT_MAGIC {
return Err(Error::WrongFormat {
found: manifest.format.clone(),
expected: manifest::FORMAT_MAGIC,
});
}
if manifest.format_version > manifest::SUPPORTED_FORMAT_VERSION {
return Err(Error::UnsupportedVersion {
found: manifest.format_version,
max_supported: manifest::SUPPORTED_FORMAT_VERSION,
});
}
Ok(())
}
async fn load_history<L: Layout>(working: &AnyContainer, layout: &L, codec: Codec) -> Result<Vec<Delta>, Error> {
if !io::exists(working, layout.history_basename(), codec).await {
return Ok(Vec::new());
}
Ok(io::iter::<Delta>(working, layout.history_basename(), codec).await?)
}
async fn replay_hot_log<L: Layout>(working: &AnyContainer, layout: &L, codec: Codec, session: &mut Session) -> Result<(), Error> {
if !io::exists(working, layout.hot_log_basename(), codec).await {
return Ok(());
}
for hot_op in io::iter::<HotOp>(working, layout.hot_log_basename(), codec).await? {
session.replay_hot_op(hot_op)?;
}
Ok(())
}
impl<L: Layout> Gdd<L> {
pub fn session(&self) -> &Session {
&self.session
}
pub fn can_undo(&self) -> bool {
self.session.can_undo()
}
pub fn can_redo(&self) -> bool {
self.session.can_redo()
}
pub fn registry(&self) -> &Registry {
self.session.registry()
}
/// The in-memory manifest. `Gdd` is its sole writer, so this is authoritative without re-reading
/// disk.
pub fn manifest(&self) -> &Manifest {
&self.manifest
}
pub fn layout(&self) -> &L {
&self.layout
}
/// The per-peer view settings read from `session.json` (PTZ, rulers, overlays, snapping, collapse).
/// Opaque `ui::doc::*` blobs; the editor decodes them. Empty for a fresh document.
pub fn view_settings(&self) -> &std::collections::BTreeMap<String, serde_json::Value> {
&self.view_settings
}
/// The per-network view settings read from `session.json` (node-graph nav + previewing), keyed by
/// [`NetworkId`](document_graph_storage::NetworkId). Opaque `ui::nav::*` / `ui::previewing` blobs the editor decodes.
pub fn network_view_settings(&self) -> &std::collections::BTreeMap<document_graph_storage::NetworkId, std::collections::BTreeMap<String, serde_json::Value>> {
&self.network_view_settings
}
/// Resolve each runtime `network_path` to its stable [`NetworkId`](document_graph_storage::NetworkId), so the
/// editor can key per-network, per-peer view state by a stable id. See [`Session::network_ids`].
#[cfg(feature = "conversion")]
pub fn network_ids<M: NodeMetadataSource>(
&self,
network: &graph_craft::document::NodeNetwork,
metadata: &M,
) -> Result<std::collections::HashMap<Vec<core_types::uuid::NodeId>, document_graph_storage::NetworkId>, CommitError> {
self.session.network_ids(network, metadata)
}
/// Every resource hash referenced by the current registry or anywhere in history, so resource GC keeps
/// redoable/re-undoable interactions' resources (notably proto-node declaration bytes) alive even when an
/// undo has dropped them from the current registry.
pub fn all_referenced_resource_hashes(&self) -> std::collections::HashSet<ResourceHash> {
self.session.all_referenced_resource_hashes()
}
/// Drop the session and return the working-copy container + layout.
/// Intended for test code that needs to reopen against the same container; panics if the container
/// is still shared by a `Gdd` clone (tests don't clone before calling this).
pub fn into_storage(self) -> (AnyContainer, L) {
let working = Arc::try_unwrap(self.working).unwrap_or_else(|_| panic!("into_storage called while the working-copy container is still shared by a Gdd clone"));
(working, self.layout)
}
/// Resolve the proto-node declarations referenced by the registry into a [`document_graph_storage::Declarations`]
/// map, loading each `ProtoNode`'s bytes from `byte_store` (the global cache in the editor, the
/// working-copy container for standalone). Only resources referenced by `Implementation::ProtoNode`
/// are visited, so image/font resources are skipped. Cold-path (open / `to_runtime`); async
/// because resource loads are.
#[cfg(feature = "conversion")]
pub async fn declarations(&self, byte_store: &dyn LoadResource) -> document_graph_storage::Declarations {
use document_graph_storage::Implementation;
let registry = self.session.registry();
let mut declarations = document_graph_storage::Declarations::new();
for node in registry.node_instances.values() {
let Implementation::ProtoNode(id) = node.implementation() else { continue };
if declarations.contains_key(id) {
continue;
}
let Some(hash) = registry.resources.get(id).and_then(|entry| entry.hash) else {
log::error!("Declaration resource {id} has no resolved hash; cannot load ProtoNode");
continue;
};
let Some(resource) = byte_store.load(hash).await else {
log::error!("Declaration bytes for {id} (hash {hash}) missing from byte store");
continue;
};
match document_graph_storage::decode_declaration(resource.as_ref()) {
Ok(proto) => {
declarations.insert(*id, proto);
}
Err(error) => log::error!("Failed to deserialize ProtoNode for {id}: {error}"),
}
}
declarations
}
/// Store the legacy `.graphite` document bytes verbatim inside the working copy (dual-write soak).
/// Synchronous (hot-path safe via `write_non_blocking`): called at the autosave boundary alongside
/// the registry snapshot. The bytes are opaque to `Gdd` — it never deserializes them.
// TODO: Add feature gate for legacy embedding
pub fn store_legacy_document(&self, bytes: &[u8]) -> Result<(), ContainerError> {
self.working.write_non_blocking(self.layout.legacy_path(), bytes)
}
/// Read back the embedded legacy `.graphite` document, if present. The compare-on-open oracle and
/// the recovery fallback both go through here. `None` when no legacy blob was ever written.
pub async fn read_legacy_document(&self) -> Option<ByteHolder> {
self.working.read(self.layout.legacy_path()).await.ok()
}
}
@@ -0,0 +1,63 @@
//! Bootstrap file for a `.gdd` document. Always JSON regardless of payload codec choice.
use serde::{Deserialize, Serialize};
use crate::Codec;
use crate::{DEFAULT_HISTORY_CODEC, DEFAULT_HOT_LOG_CODEC, DEFAULT_REGISTRY_CODEC, DEFAULT_SESSION_CODEC};
/// Magic string carried in [`Manifest::format`] to identify a `.gdd` document.
pub const FORMAT_MAGIC: &str = "gdd";
/// Maximum manifest version this build can open. Bumped when manifest layout changes
/// in a way that older builds can't safely read.
pub const SUPPORTED_FORMAT_VERSION: u32 = 1;
/// The on-disk codec for each working-copy payload, recorded so reads/writes never have to probe
/// the filesystem to discover it. The manifest itself is excluded: it is always JSON, since it must
/// be parsed before any other codec is known.
#[derive(Clone, Copy, Debug, Serialize, Deserialize)]
pub struct PayloadCodecs {
pub registry: Codec,
pub history: Codec,
pub hot_log: Codec,
pub session: Codec,
}
impl Default for PayloadCodecs {
fn default() -> Self {
Self {
registry: DEFAULT_REGISTRY_CODEC,
history: DEFAULT_HISTORY_CODEC,
hot_log: DEFAULT_HOT_LOG_CODEC,
session: DEFAULT_SESSION_CODEC,
}
}
}
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct Manifest {
pub format: String,
pub format_version: u32,
pub editor_version: String,
pub stdlib_version: String,
pub document_id: u64,
/// Codec used for each non-manifest payload on disk. Authoritative — never inferred from which
/// file extension is present.
#[serde(default)]
pub codecs: PayloadCodecs,
}
impl Manifest {
pub fn new(document_id: u64, editor_version: String, stdlib_version: String) -> Self {
Self {
format: FORMAT_MAGIC.to_string(),
format_version: SUPPORTED_FORMAT_VERSION,
document_id,
editor_version,
stdlib_version,
codecs: PayloadCodecs::default(),
}
}
}
@@ -0,0 +1,259 @@
//! The per-edit persist path on the [`Gdd`] handle: stage/retire/commit, hot-log and history
//! append, registry snapshots, session-state and manifest writes, plus view-settings setters.
//! Synchronous and read-free (the manifest is cached on the handle); the container's `*_non_blocking`
//! surface absorbs durability.
use document_container::AsyncContainer;
#[cfg(feature = "conversion")]
use document_graph_storage::NodeMetadataSource;
use document_graph_storage::{HotOp, Rev, TimeStamp};
#[cfg(feature = "conversion")]
use graphene_resource::ResourceStorage;
use crate::error::Error;
use crate::layout::Layout;
use crate::manifest::Manifest;
use crate::session_state::SessionState;
use crate::{Gdd, MANIFEST_CODEC, io};
impl<L: Layout> Gdd<L> {
/// Move the undo cursor back one commit (silent-zone reflog undo) and persist the new cursor. Returns
/// the undone `Rev`. The working registry is rewound in place by the reverse delta, so re-snapshot it
/// (alongside `head`) or a reopen would read a `registry.bin` inconsistent with the persisted cursor.
pub fn undo(&mut self) -> Result<Rev, Error> {
let rev = self.session.undo()?;
self.persist_registry_snapshot()?;
self.persist_session_state()?;
Ok(rev)
}
/// Re-apply the most-recently-undone commit and persist the new cursor and re-snapshotted registry.
pub fn redo(&mut self) -> Result<Rev, Error> {
let rev = self.session.redo()?;
self.persist_registry_snapshot()?;
self.persist_session_state()?;
Ok(rev)
}
/// Edit the cached manifest and persist it. Always JSON, synchronous.
pub fn update_manifest(&mut self, edit: impl FnOnce(&mut Manifest)) -> Result<(), Error> {
edit(&mut self.manifest);
io::write_single(&self.working, self.layout.manifest_basename(), MANIFEST_CODEC, &self.manifest)?;
Ok(())
}
/// Stage a runtime snapshot as hot ops without retiring: diff the runtime against the working
/// registry, append the hot frames (so a crash recovers the work), and persist proto-node
/// declaration bytes. The working registry reflects the edit immediately, but nothing enters durable
/// retired history until [`retire_pending_interaction`](Self::retire_pending_interaction). Staging on
/// every edit while retiring only at interaction boundaries lets several edits coalesce into one retired
/// interaction.
///
/// # Errors
/// [`Error::Commit`] if the runtime diff is rejected by the session. On an [`Error::Container`] /
/// [`Error::Codec`] from persisting the hot frames, the session has already advanced past what the
/// working copy reflects, so the caller should treat the document as needing re-persist.
#[cfg(feature = "conversion")]
pub fn stage_runtime_snapshot<M: NodeMetadataSource>(
&mut self,
network: &graph_craft::document::NodeNetwork,
metadata: &M,
resources: &graphene_resource::ResourceRegistry,
byte_store: &dyn ResourceStorage,
) -> Result<(), Error> {
let (hot_ops, declaration_bytes) = self.session.stage_from_runtime(network, metadata, resources)?;
for hot_op in &hot_ops {
self.append_hot_frame(hot_op)?;
}
// Persist proto-node declaration content to the byte store (the global cache in the editor,
// the working-copy container for standalone export). Content-addressed, so re-storing
// identical bytes on every commit is an idempotent no-op.
for bytes in declaration_bytes.values() {
byte_store.store(bytes);
}
Ok(())
}
/// Retire every pending hot op into durable history as a single interaction (marking the batch's last
/// delta as the interaction boundary), then re-snapshot the registry. One interaction is one undo unit,
/// so the caller invokes this at each undo-step boundary and before any undo/redo. A no-op when there
/// are no pending hot ops.
pub fn retire_pending_interaction(&mut self) -> Result<Vec<Rev>, Error> {
let Some(up_to) = self.session.hot_log().iter().map(|hot_op| hot_op.timestamp).max() else {
return Ok(Vec::new());
};
self.retire_inner(up_to, true)
}
/// Commit a runtime snapshot as one complete interaction: stage it, then immediately retire it into
/// durable history. Convenience for callers that produce a whole interaction atomically (tests, and any
/// one-shot commit). Equivalent to [`stage_runtime_snapshot`](Self::stage_runtime_snapshot) followed
/// by [`retire_pending_interaction`](Self::retire_pending_interaction).
#[cfg(feature = "conversion")]
pub fn commit_from_runtime<M: NodeMetadataSource>(
&mut self,
network: &graph_craft::document::NodeNetwork,
metadata: &M,
resources: &graphene_resource::ResourceRegistry,
byte_store: &dyn ResourceStorage,
) -> Result<Vec<Rev>, Error> {
self.stage_runtime_snapshot(network, metadata, resources, byte_store)?;
self.retire_pending_interaction()
}
/// Apply a hot op from the broadcast stream, appending one frame to the hot log.
///
/// # Errors
/// Returns [`Error::Crdt`] if the op is rejected by the session, or an [`Error::Container`] /
/// [`Error::Codec`] if persisting the hot frame fails. On a persist failure the session has already
/// advanced past what the working copy reflects, so the caller should treat the document as needing
/// re-persist (mirrors [`stage_runtime_snapshot`](Self::stage_runtime_snapshot)).
pub fn apply_hot_op(&mut self, op: HotOp) -> Result<(), Error> {
self.session.apply_hot_op(op.clone())?;
self.append_hot_frame(&op)?;
Ok(())
}
/// Persist freshly-staged hot ops and immediately retire them into durable history. Appends each
/// hot frame (so a crash before retirement still recovers the work), then retires up to the last
/// staged timestamp, which drains exactly these ops and re-snapshots the registry. Returns the
/// retired `Rev`s. A no-op when nothing was staged.
pub(crate) fn append_and_retire(&mut self, hot_ops: &[HotOp], interaction_end: bool) -> Result<Vec<Rev>, Error> {
let Some(last) = hot_ops.last() else { return Ok(Vec::new()) };
for hot_op in hot_ops {
self.append_hot_frame(hot_op)?;
}
self.retire_inner(last.timestamp, interaction_end)
}
/// Encode the history deltas identified by `revs` and append them to the history file. `revs` comes
/// from `Session::retire` in append order, which is a valid replay order, so a direct per-rev lookup
/// preserves replay order without scanning the whole history.
fn append_history_deltas(&mut self, revs: &[Rev]) -> Result<(), Error> {
let mut buffer = Vec::new();
for &rev in revs {
let Some(delta) = self.session.delta(rev) else {
log::error!("Retired rev {rev:?} missing from history; skipping its history frame");
continue;
};
self.manifest.codecs.history.append(&mut buffer, delta)?;
}
self.working.append_non_blocking(&io::path_for(self.layout.history_basename(), self.manifest.codecs.history), &buffer)?;
Ok(())
}
/// Set a local annotation (e.g. a commit message) on an existing retired delta and re-persist it.
/// Unlike the per-interaction marker written inline at retire, this targets an already-written delta, so
/// the whole history file is rewritten in topological order. O(history) — fine for occasional user
/// labeling, not for per-interaction marking (which uses the inline path). No-op if `rev` is unknown.
pub fn annotate_delta(&mut self, rev: Rev, key: &str, value: serde_json::Value) -> Result<(), Error> {
if self.session.annotate_delta(rev, key, value) {
self.rewrite_history()?;
}
Ok(())
}
/// Rewrite the entire history file from the in-memory session. `history()` yields deltas in
/// topological (append) order, which is a valid replay order, so no separate sort is needed.
fn rewrite_history(&mut self) -> Result<(), Error> {
let mut buffer = Vec::new();
for delta in self.session.history() {
self.manifest.codecs.history.append(&mut buffer, delta)?;
}
self.working.write_non_blocking(&io::path_for(self.layout.history_basename(), self.manifest.codecs.history), &buffer)?;
Ok(())
}
fn persist_session_state(&mut self) -> Result<(), Error> {
let state = SessionState {
peer_id: self.session.peer(),
head_rev: self.session.head_rev(),
last_broadcast_rev: self.session.last_broadcast_rev(),
redo_stack: self.session.redo_stack().to_vec(),
next_node_counter: self.session.next_node_counter(),
view_settings: self.view_settings.clone(),
network_view_settings: self.network_view_settings.clone(),
};
io::write_single(&self.working, self.layout.session_basename(), self.manifest.codecs.session, &state)?;
Ok(())
}
/// Re-snapshot the materialized working registry to `registry.bin`. `Session::load` trusts the stored
/// registry to match the persisted `head`, so any cursor move (undo/redo) that rewinds the working
/// registry without retiring must re-persist it or a reopen would read a registry inconsistent with
/// `head`. Synchronous and hot-path-safe (`write_non_blocking`).
fn persist_registry_snapshot(&mut self) -> Result<(), Error> {
io::write_single(&self.working, self.layout.registry_basename(), self.manifest.codecs.registry, self.session.registry())?;
Ok(())
}
/// Replace the per-peer view settings and persist them to `session.json`. Called by the editor when
/// the viewport or a document-level toggle changes; never enters the registry, history, or CRDT.
pub fn set_view_settings(&mut self, view_settings: std::collections::BTreeMap<String, serde_json::Value>) -> Result<(), Error> {
self.view_settings = view_settings;
self.persist_session_state()
}
/// Advance the published frontier to `rev` and persist it to `session.json`, so the silent/published
/// undo boundary survives a reopen. Called by the (future) broadcast transport as commits are shared.
pub fn publish_up_to(&mut self, rev: document_graph_storage::Rev) -> Result<(), Error> {
self.session.publish_up_to(rev);
self.persist_session_state()
}
/// Replace the per-network view settings and persist them to `session.json`. Per-peer, per-network; never
/// enters the registry, history, or CRDT.
pub fn set_network_view_settings(
&mut self,
network_view_settings: std::collections::BTreeMap<document_graph_storage::NetworkId, std::collections::BTreeMap<String, serde_json::Value>>,
) -> Result<(), Error> {
self.network_view_settings = network_view_settings;
self.persist_session_state()
}
fn append_hot_frame(&mut self, op: &HotOp) -> Result<(), Error> {
let mut buffer = Vec::new();
self.manifest.codecs.hot_log.append(&mut buffer, op)?;
self.working.append_non_blocking(&io::path_for(self.layout.hot_log_basename(), self.manifest.codecs.hot_log), &buffer)?;
Ok(())
}
/// Working-copy checkpoint: promote hot ops with timestamp `≤ up_to` into retired deltas,
/// append them to the history file, rewrite the hot log with remaining (unretired) ops, and
/// re-snapshot the registry. Synchronous.
pub fn retire(&mut self, up_to: TimeStamp) -> Result<Vec<Rev>, Error> {
self.retire_inner(up_to, false)
}
/// `interaction_end`: mark the batch's last delta as an interaction boundary (one undo unit) before its
/// history frame is written, so the marker persists on reopen without a later frame rewrite.
fn retire_inner(&mut self, up_to: TimeStamp, interaction_end: bool) -> Result<Vec<Rev>, Error> {
let new_revs = self.session.retire(up_to)?;
// Mark before `append_history_deltas` so the on-disk frame carries the boundary.
if interaction_end && let Some(&last) = new_revs.last() {
self.session.mark_interaction_end(last);
}
if !new_revs.is_empty() {
self.append_history_deltas(&new_revs)?;
}
// Rewrite hot log with whatever survived retirement.
let mut hot_buffer = Vec::new();
for hot_op in self.session.hot_log() {
self.manifest.codecs.hot_log.append(&mut hot_buffer, hot_op)?;
}
self.working
.write_non_blocking(&io::path_for(self.layout.hot_log_basename(), self.manifest.codecs.hot_log), &hot_buffer)?;
self.persist_registry_snapshot()?;
self.persist_session_state()?;
Ok(new_revs)
}
}
@@ -0,0 +1,159 @@
//! Resource I/O on the [`Gdd`] handle: the content-addressed byte store that backs raster images,
//! fonts, embedded WASM, and proto-node declarations. Registration goes through the session as an
//! `AddResource` delta; the bytes live in the working copy's `resources/<hash>` directory.
#[cfg(not(target_family = "wasm"))]
use std::path::Path;
use std::sync::Arc;
use document_container::{AnyContainer, AsyncContainer, ByteHolder, ContainerError};
use graphene_resource::ResourceFuture;
use graphene_resource::{LoadResource, Resource, ResourceHash, ResourceStorage};
use crate::Gdd;
use crate::error::Error;
use crate::layout::Layout;
impl<L: Layout> Gdd<L> {
pub async fn read_resource(&self, hash: &ResourceHash) -> Result<ByteHolder, ContainerError> {
self.working.read(&self.layout.resource_path(hash)).await
}
/// Register a resource under `id` and store its bytes. Commits an `AddResource` delta (a single
/// `DataSource::Embedded` source resolved to the content hash) through the session so the registry
/// records the resource and the entry replicates, then writes the bytes into the working copy's
/// content-addressed store. The caller owns `id` allocation.
pub fn add_resource(&mut self, id: document_graph_storage::ResourceId, bytes: &[u8]) -> Result<(), Error> {
let hash = ResourceHash::from(bytes);
self.working.write_non_blocking(&self.layout.resource_path(&hash), bytes)?;
let hot_ops = self.session.stage_embedded_resource(id, hash)?;
self.append_and_retire(&hot_ops, false)?;
Ok(())
}
/// Like [`add_resource`](Self::add_resource) but copies the bytes from a filesystem `src` rather
/// than buffering them. Folder backends use `fs::copy` (CoW on supported filesystems); other
/// backends fall back to read-then-write. Native-only: there is no filesystem source path on wasm.
#[cfg(not(target_family = "wasm"))]
pub fn add_resource_from_path(&mut self, id: document_graph_storage::ResourceId, hash: ResourceHash, src: &Path) -> Result<(), Error> {
let dest_path = self.layout.resource_path(&hash);
if let AnyContainer::Folder(folder) = self.working.as_ref() {
let full = folder.root().join(&dest_path);
if let Some(parent) = full.parent() {
std::fs::create_dir_all(parent).map_err(ContainerError::Io)?;
}
std::fs::copy(src, &full).map_err(ContainerError::Io)?;
} else {
let bytes = std::fs::read(src).map_err(ContainerError::Io)?;
// The folder fast path trusts the caller's hash to avoid reading the file; here we've read
// the bytes anyway, so verify the hash matches and flag a content-addressing bug in debug.
debug_assert_eq!(hash, ResourceHash::from(bytes.as_slice()), "add_resource_from_path hash does not match the file at {src:?}");
self.working.write_non_blocking(&dest_path, &bytes)?;
}
let hot_ops = self.session.stage_embedded_resource(id, hash)?;
self.append_and_retire(&hot_ops, false)?;
Ok(())
}
pub async fn has_resource(&self, hash: &ResourceHash) -> bool {
self.working.exists(&self.layout.resource_path(hash)).await
}
pub fn remove_resource(&self, hash: &ResourceHash) -> Result<(), ContainerError> {
self.working.remove_non_blocking(&self.layout.resource_path(hash))
}
/// Enumerate every resource currently in the working copy. Paths that don't parse as a
/// `ResourceHash` (foreign files dropped into the resources directory) are silently skipped.
pub async fn resource_hashes(&self) -> Result<Vec<ResourceHash>, ContainerError> {
let dir = self.layout.resources_dir();
if !self.working.list_dirs("").await?.iter().any(|d| d == dir) {
return Ok(Vec::new());
}
let entries = self.working.list(dir).await?;
let prefix = format!("{dir}/");
let mut hashes = Vec::with_capacity(entries.len());
for entry in entries {
let Some(name) = entry.strip_prefix(&prefix) else { continue };
if let Ok(hash) = name.parse::<ResourceHash>() {
hashes.push(hash);
}
}
Ok(hashes)
}
pub fn resource_proxy(&self) -> ResourceProxy<L>
where
L: Clone,
{
ResourceProxy(self.working.clone(), self.layout.clone())
}
}
impl<L: Layout + Send + Sync> LoadResource for Gdd<L> {
fn load(&self, hash: ResourceHash) -> ResourceFuture<'_> {
Box::pin(async move {
let bytes = self.working.read(&self.layout.resource_path(&hash)).await.ok()?;
Some(Resource::new(bytes))
})
}
}
pub struct ResourceProxy<T: Layout>(Arc<AnyContainer>, T);
impl<L: Layout + Send + Sync> LoadResource for ResourceProxy<L> {
fn load(&self, hash: ResourceHash) -> ResourceFuture<'_> {
Box::pin(async move {
let bytes = self.0.read(&self.1.resource_path(&hash)).await.ok()?;
Some(Resource::new(bytes))
})
}
}
impl<L: Layout + Send + Sync> ResourceStorage for Gdd<L> {
fn store(&self, data: &[u8]) -> ResourceHash {
let hash = ResourceHash::from(data);
if let Err(error) = self.working.write_non_blocking(&self.layout.resource_path(&hash), data) {
log::error!("ResourceStorage::store failed for {hash}: {error}");
}
hash
}
fn contains(&self, hash: &ResourceHash) -> bool {
self.working.exists_non_blocking(&self.layout.resource_path(hash))
}
fn garbage_collect(&self, used: &[ResourceHash]) {
// `garbage_collect` is synchronous but listing resources is async, so the native path blocks on
// it. That's unavailable on wasm (single-threaded; `block_on` would deadlock). The editor never
// uses `Gdd` as the runtime `ResourceStorage` on wasm (it GCs the app-global cache instead), so
// this is an unreachable configuration there rather than a missing feature.
#[cfg(target_family = "wasm")]
{
let _ = used;
log::error!("ResourceStorage::garbage_collect is not supported for Gdd on wasm");
}
#[cfg(not(target_family = "wasm"))]
{
let kept: std::collections::HashSet<&ResourceHash> = used.iter().collect();
let hashes = match futures::executor::block_on(self.resource_hashes()) {
Ok(hashes) => hashes,
Err(error) => {
log::error!("Failed to list resources during garbage_collect: {error}");
return;
}
};
for hash in hashes {
if kept.contains(&hash) {
continue;
}
if let Err(error) = self.working.remove_non_blocking(&self.layout.resource_path(&hash)) {
log::error!("ResourceStorage::garbage_collect failed to remove {hash}: {error}");
}
}
}
}
}
@@ -0,0 +1,42 @@
//! Persistent cursor state for the local peer. Separate from [`crate::Manifest`] because the
//! manifest describes document identity (what this document *is*), while [`SessionState`]
//! describes where the local peer's cursor sits inside it.
//!
//! Lives in `session.json`. Rewritten on retirement.
use document_graph_storage::{NetworkId, PeerId, Rev};
use serde::{Deserialize, Serialize};
use std::collections::BTreeMap;
#[derive(Clone, Debug, Default, Serialize, Deserialize)]
pub struct SessionState {
/// This peer's identity for the document, stable per (device, document). Per-peer rather than
/// document identity, so it lives with the cursor here, not in the manifest. Used for CRDT
/// tiebreaking and minting peer-scoped IDs.
#[serde(default)]
pub peer_id: PeerId,
/// Local-chain cursor. Points at the most recently applied retired delta, or `None` on an empty
/// document (no commits yet).
#[serde(default)]
pub head_rev: Option<Rev>,
/// Published frontier: the latest retired commit broadcast to a peer. Commits after it are silently
/// rewritable on undo; commits at or before it are published. `None` until broadcast transport lands.
#[serde(default, skip_serializing_if = "Option::is_none")]
pub last_broadcast_rev: Option<Rev>,
/// Revs the user has undone past, so redo survives a reopen. (The legacy `VecDeque` redo history
/// is not persisted, so within the shadow phase this is strictly more capable than the live editor.)
#[serde(default)]
pub redo_stack: Vec<Rev>,
/// Shared-monotonic counter feeding `Document::next_node_id`. Persisted so reopens don't
/// collide on minted IDs.
#[serde(default)]
pub next_node_counter: u64,
/// Per-peer view settings (PTZ, rulers, overlays, snapping, panel collapse). Local to the viewer,
/// so kept out of the CRDT/history. Editor owns the keys/values (opaque `ui::doc::*` blobs).
#[serde(default, skip_serializing_if = "BTreeMap::is_empty")]
pub view_settings: BTreeMap<String, serde_json::Value>,
/// Per-network view settings (node-graph nav + previewing), keyed by the stable storage [`NetworkId`].
/// Per-peer like [`view_settings`](Self::view_settings); opaque `ui::nav::*` / `ui::previewing` blobs.
#[serde(default, skip_serializing_if = "BTreeMap::is_empty")]
pub network_view_settings: BTreeMap<NetworkId, BTreeMap<String, serde_json::Value>>,
}
@@ -0,0 +1,722 @@
// Exercises the runtime-conversion bridge and the zip archive round-trip, so it only compiles with
// both features. A minimal-dependency build (e.g. `--no-default-features`) skips it entirely.
#![cfg(all(feature = "conversion", feature = "zip"))]
use document_container::AnyContainer;
use document_container::backends::memory::MemoryBackend;
use document_format::{Codec, Error, GddV1, GddV1Layout, Layout, Manifest, io, manifest};
use document_graph_storage::{HotOp, Network, NetworkId, PeerId, ROOT_NETWORK, RegistryDelta, TimeStamp};
fn empty_container() -> AnyContainer {
AnyContainer::Memory(MemoryBackend::new())
}
/// A resource byte store for export calls. Empty unless a test pre-populates it; only consulted when
/// `embed_all_resources` is set.
fn empty_byte_store() -> graph_craft::application_io::resource::HashMapResourceStorage {
graph_craft::application_io::resource::HashMapResourceStorage::new()
}
/// A one-node network referencing `id` via a `TaggedValue::Resource` input. Conversion only snapshots
/// resources the network references, so a resource needs a referencing node to survive into storage.
fn network_referencing_resource(id: graphene_resource::ResourceId) -> graph_craft::document::NodeNetwork {
use graph_craft::ProtoNodeIdentifier;
use graph_craft::document::value::TaggedValue;
use graph_craft::document::{DocumentNode, DocumentNodeImplementation, NodeId, NodeInput, NodeNetwork};
NodeNetwork {
nodes: [(
NodeId(0),
DocumentNode {
inputs: vec![NodeInput::value(TaggedValue::Resource(id), false)],
implementation: DocumentNodeImplementation::ProtoNode(ProtoNodeIdentifier::new("graphene_core::ops::identity::IdentityNode")),
..Default::default()
},
)]
.into_iter()
.collect(),
..Default::default()
}
}
#[test]
fn create_in_round_trips_empty_document() {
futures::executor::block_on(async {
let container = empty_container();
let created = match GddV1::create_in(container, GddV1Layout, PeerId(7), 0xFEED, "editor-x".into(), "stdlib-x".into()).await {
Ok(gdd) => gdd,
Err(error) => panic!("create_in failed: {error:?}"),
};
let (working, layout) = created.into_storage();
let reopened = match GddV1::open_in(working, layout).await {
Ok(gdd) => gdd,
Err(error) => panic!("open_in failed: {error:?}"),
};
assert_eq!(reopened.session().peer(), PeerId(7));
assert!(reopened.registry().node_instances.is_empty());
assert!(reopened.registry().networks.is_empty());
});
}
#[test]
fn open_in_rejects_wrong_format_magic() {
futures::executor::block_on(async {
let container = empty_container();
let layout = GddV1Layout;
let mut bogus = Manifest::new(0xC0DE, "ed".into(), "std".into());
bogus.format = "not-gdd".into();
io::write_single(&container, layout.manifest_basename(), Codec::Json, &bogus).unwrap();
match GddV1::open_in(container, layout).await {
Err(Error::WrongFormat { .. }) => {}
Ok(_) => panic!("expected WrongFormat, got Ok"),
Err(other) => panic!("expected WrongFormat, got {other:?}"),
}
});
}
#[test]
fn manifest_returns_what_create_in_wrote() {
futures::executor::block_on(async {
let gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(13), 0xC0FFEE, "ed-1.2".into(), "std-0.7".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
assert_eq!(gdd.session().peer(), PeerId(13));
let manifest = gdd.manifest();
assert_eq!(manifest.document_id, 0xC0FFEE);
assert_eq!(manifest.editor_version, "ed-1.2");
assert_eq!(manifest.stdlib_version, "std-0.7");
assert_eq!(manifest.format, manifest::FORMAT_MAGIC);
});
}
#[test]
fn update_manifest_changes_visible_after_reopen() {
futures::executor::block_on(async {
let mut gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(1), 0xAB, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
gdd.update_manifest(|m| m.editor_version = "ed-NEW".into())
.unwrap_or_else(|error| panic!("update_manifest failed: {error:?}"));
let (working, layout) = gdd.into_storage();
let reopened = GddV1::open_in(working, layout).await.unwrap_or_else(|error| panic!("open_in failed: {error:?}"));
let manifest = reopened.manifest();
assert_eq!(manifest.editor_version, "ed-NEW");
});
}
#[test]
fn apply_hot_op_persists_to_hot_log_and_survives_reopen() {
futures::executor::block_on(async {
let mut gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(5), 0xDEAD, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
// AddNetwork on the root network. Idempotent at apply, so two hot ops applied in sequence
// produces one network in the registry.
let hot_op = HotOp {
op: RegistryDelta::AddNetwork {
id: ROOT_NETWORK,
network: Network::default(),
},
timestamp: TimeStamp { counter: 1, peer: PeerId(5) },
};
gdd.apply_hot_op(hot_op).unwrap_or_else(|error| panic!("apply_hot_op failed: {error:?}"));
assert!(gdd.registry().networks.contains_key(&ROOT_NETWORK), "hot op should have created the root network in memory");
let (working, layout) = gdd.into_storage();
let reopened = GddV1::open_in(working, layout).await.unwrap_or_else(|error| panic!("open_in failed: {error:?}"));
assert!(reopened.registry().networks.contains_key(&ROOT_NETWORK), "hot op should have been replayed from the hot log on reopen");
});
}
#[test]
fn retire_moves_eligible_hot_ops_to_history_and_keeps_rest() {
futures::executor::block_on(async {
let mut gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(5), 0xDEAD, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
// Two hot ops: one with low timestamp (will retire), one with high (will stay).
let early = HotOp {
op: RegistryDelta::AddNetwork {
id: ROOT_NETWORK,
network: Network::default(),
},
timestamp: TimeStamp { counter: 1, peer: PeerId(5) },
};
let late = HotOp {
op: RegistryDelta::AddNetwork {
id: NetworkId(42),
network: Network::default(),
},
timestamp: TimeStamp { counter: 10, peer: PeerId(5) },
};
gdd.apply_hot_op(early).unwrap();
gdd.apply_hot_op(late).unwrap();
assert_eq!(gdd.session().hot_log().len(), 2);
// Retire only up to timestamp 5 → drains the early op, leaves the late one.
let cutoff = TimeStamp { counter: 5, peer: PeerId(5) };
gdd.retire(cutoff).unwrap_or_else(|error| panic!("retire failed: {error:?}"));
assert_eq!(gdd.session().hot_log().len(), 1, "late hot op should still be in hot log");
assert_eq!(gdd.session().history().count(), 1, "early hot op should be in retired history");
// Reopen and confirm survival: hot log has the late op (replayed), history has the early op.
let (working, layout) = gdd.into_storage();
let reopened = GddV1::open_in(working, layout).await.unwrap_or_else(|error| panic!("open_in failed: {error:?}"));
assert!(reopened.registry().networks.contains_key(&ROOT_NETWORK), "retired op's effect should be in registry");
assert!(reopened.registry().networks.contains_key(&NetworkId(42)), "hot op's effect should be replayed");
assert_eq!(reopened.session().history().count(), 1);
assert_eq!(reopened.session().hot_log().len(), 1);
});
}
/// The published frontier (`last_broadcast_rev`, the silent/published undo boundary) persists in
/// `session.json` and is restored on reopen.
#[test]
fn last_broadcast_rev_persists_across_reopen() {
futures::executor::block_on(async {
let mut gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(5), 0xDEAD, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
// Retire one op so there is a real retired rev to mark as published.
let op = HotOp {
op: RegistryDelta::AddNetwork {
id: ROOT_NETWORK,
network: Network::default(),
},
timestamp: TimeStamp { counter: 1, peer: PeerId(5) },
};
gdd.apply_hot_op(op).unwrap();
let retired = gdd.retire(TimeStamp { counter: 1, peer: PeerId(5) }).unwrap_or_else(|error| panic!("retire failed: {error:?}"));
let published = *retired.last().expect("one retired rev");
assert_eq!(gdd.session().last_broadcast_rev(), None, "nothing is published before publish_up_to");
gdd.publish_up_to(published).unwrap_or_else(|error| panic!("publish_up_to failed: {error:?}"));
assert_eq!(gdd.session().last_broadcast_rev(), Some(published));
let (working, layout) = gdd.into_storage();
let reopened = GddV1::open_in(working, layout).await.unwrap_or_else(|error| panic!("open_in failed: {error:?}"));
assert_eq!(reopened.session().last_broadcast_rev(), Some(published), "published frontier should survive reopen");
});
}
#[test]
fn export_folder_round_trips_through_open() {
use document_format::{ExportFormat, ExportOptions};
futures::executor::block_on(async {
let gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(3), 0xAB, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
let dir = tempfile::tempdir().unwrap();
let dest = dir.path().join("export");
gdd.export(&dest, ExportFormat::Folder, ExportOptions::default(), &empty_byte_store(), None)
.await
.unwrap_or_else(|error| panic!("export failed: {error:?}"));
// Payloads keep the working-copy codecs: registry is MessagePack (`.bin`), manifest is JSON.
assert!(dest.join("registry.bin").exists());
assert!(dest.join("manifest.json").exists());
assert!(dest.join("session.json").exists());
assert!(!dest.join("hot-log.bin").exists());
assert!(!dest.join("hot-log.frames").exists());
// And the export is itself openable.
let reopened = GddV1::open(&dest).await.unwrap_or_else(|error| panic!("open failed: {error:?}"));
assert_eq!(reopened.session().peer(), PeerId(3));
});
}
#[test]
fn export_zip_round_trips_via_deserialize() {
use document_container::archive::{Archive, Zip};
use document_format::{ExportFormat, ExportOptions};
futures::executor::block_on(async {
let gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(4), 0xCD, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
let dir = tempfile::tempdir().unwrap();
let dest = dir.path().join("doc.gdd.zip");
gdd.export(&dest, ExportFormat::Zip, ExportOptions::default(), &empty_byte_store(), None)
.await
.unwrap_or_else(|error| panic!("export failed: {error:?}"));
let bytes = std::fs::read(&dest).unwrap();
let mut restored = document_container::backends::memory::MemoryBackend::new();
Zip::open(std::io::Cursor::new(&bytes), &mut restored).unwrap();
use document_container::Container;
assert!(restored.exists("manifest.json"));
assert!(restored.exists("registry.bin"));
assert!(restored.exists("session.json"));
assert!(!restored.exists("hot-log.frames"));
});
}
#[test]
fn export_rejects_invalid_options() {
use document_format::{ExportFormat, ExportOptions};
futures::executor::block_on(async {
let gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(1), 0xEF, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
let dir = tempfile::tempdir().unwrap();
let dest = dir.path().join("nope");
let options = ExportOptions {
include_registry: false,
include_history: false,
embed_all_resources: false,
};
match gdd.export(&dest, ExportFormat::Folder, options, &empty_byte_store(), None).await {
Err(Error::InvalidExportOptions(_)) => {}
Ok(_) => panic!("expected InvalidOptions, got Ok"),
Err(other) => panic!("expected InvalidOptions, got {other:?}"),
}
});
}
#[test]
fn resource_round_trip_add_read_remove() {
use graphene_resource::{ResourceHash, ResourceId};
futures::executor::block_on(async {
let mut gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(99), 0xCAFE, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
let payload = b"deadbeef cafe babe";
let hash = ResourceHash::from(&payload[..]);
let id = ResourceId::new();
assert!(!gdd.has_resource(&hash).await);
gdd.add_resource(id, payload).unwrap_or_else(|error| panic!("add_resource failed: {error:?}"));
assert!(gdd.has_resource(&hash).await);
let read_back = gdd.read_resource(&hash).await.unwrap();
assert_eq!(read_back.as_slice(), payload);
// The registry records the resource (entry keyed by id, resolved to the content hash).
let entry = gdd.registry().resources.get(&id).expect("registry records the added resource");
assert_eq!(entry.hash, Some(hash));
let hashes = gdd.resource_hashes().await.unwrap();
assert_eq!(hashes, vec![hash]);
gdd.remove_resource(&hash).unwrap();
assert!(!gdd.has_resource(&hash).await);
});
}
#[test]
fn resource_survives_reopen() {
use graphene_resource::{ResourceHash, ResourceId};
futures::executor::block_on(async {
let mut gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(7), 0xC0DE, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
let payload = b"persistent bytes";
let hash = ResourceHash::from(&payload[..]);
let id = ResourceId::new();
gdd.add_resource(id, payload).unwrap();
let (working, layout) = gdd.into_storage();
let reopened = GddV1::open_in(working, layout).await.unwrap_or_else(|error| panic!("open_in failed: {error:?}"));
assert!(reopened.has_resource(&hash).await);
assert_eq!(reopened.read_resource(&hash).await.unwrap().as_slice(), payload);
// The registry entry replicated through the history file and survives reopen.
let entry = reopened.registry().resources.get(&id).expect("reopened registry records the resource");
assert_eq!(entry.hash, Some(hash));
});
}
#[test]
fn resource_from_path_uses_fs_copy_on_folder_backend() {
use document_container::AnyContainer;
use document_container::backends::folder::FolderBackend;
use graphene_resource::{ResourceHash, ResourceId};
futures::executor::block_on(async {
// Need a folder-backed working copy to exercise the fs::copy path.
let working_dir = tempfile::tempdir().unwrap();
let working = AnyContainer::Folder(FolderBackend::create(working_dir.path()).unwrap());
let mut gdd = GddV1::create_in(working, GddV1Layout, PeerId(1), 0xAB, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
// Source file outside the working copy.
let payload = b"external resource bytes";
let src_dir = tempfile::tempdir().unwrap();
let src_path = src_dir.path().join("blob");
std::fs::write(&src_path, payload).unwrap();
let hash = ResourceHash::from(&payload[..]);
let id = ResourceId::new();
gdd.add_resource_from_path(id, hash, &src_path)
.unwrap_or_else(|error| panic!("add_resource_from_path failed: {error:?}"));
assert!(gdd.has_resource(&hash).await);
assert_eq!(gdd.read_resource(&hash).await.unwrap().as_slice(), payload);
});
}
#[test]
fn export_carries_resources() {
use document_format::{ExportFormat, ExportOptions};
use graphene_resource::{ResourceHash, ResourceId};
futures::executor::block_on(async {
let mut gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(2), 0xBC, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
let payload = b"exported resource";
let hash = ResourceHash::from(&payload[..]);
let id = ResourceId::new();
gdd.add_resource(id, payload).unwrap();
let dir = tempfile::tempdir().unwrap();
let dest = dir.path().join("export");
gdd.export(&dest, ExportFormat::Folder, ExportOptions::default(), &empty_byte_store(), None).await.unwrap();
let resource_file = dest.join("resources").join(format!("{hash}"));
assert!(resource_file.exists(), "exported resource file should exist at {resource_file:?}");
assert_eq!(std::fs::read(&resource_file).unwrap(), payload);
});
}
/// `embed_all_resources` makes a link-only resource self-contained: the bytes (which live only in
/// the byte store, not the working copy) are written into the export, the exported registry's chain
/// gains a leading `Embedded` source ahead of the original `Url`, and the export reopens with both.
#[test]
fn embed_all_resources_materializes_link_only_resource() {
use document_format::{ExportFormat, ExportOptions};
use document_graph_storage::NoMetadata;
use graph_craft::application_io::resource::ResourceStorage;
use graphene_resource::{DataSource, ResourceHash, ResourceId, ResourceRegistry};
futures::executor::block_on(async {
let mut gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(8), 0xF00D, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
// A resource whose only source is a URL, resolved to a hash. The bytes live solely in the
// byte store; the working copy never holds them.
let payload = b"bytes behind a url";
let hash = ResourceHash::from(&payload[..]);
let byte_store = empty_byte_store();
byte_store.store(payload);
let mut resources = ResourceRegistry::new();
let id = ResourceId::new();
resources.push_source_back(&id, DataSource::Url("https://example.com/r.bin".parse().unwrap()));
resources.resolve(&id, hash);
gdd.commit_from_runtime(&network_referencing_resource(id), &NoMetadata, &resources, &byte_store)
.unwrap_or_else(|error| panic!("commit_from_runtime failed: {error:?}"));
// The working copy holds no resource bytes (URL source, nothing embedded yet).
assert!(!gdd.has_resource(&hash).await);
let dir = tempfile::tempdir().unwrap();
let dest = dir.path().join("embedded");
gdd.export(
&dest,
ExportFormat::Folder,
ExportOptions {
embed_all_resources: true,
..Default::default()
},
&byte_store,
None,
)
.await
.unwrap_or_else(|error| panic!("export failed: {error:?}"));
// Bytes materialized into the export.
let resource_file = dest.join("resources").join(format!("{hash}"));
assert!(resource_file.exists(), "embedded resource bytes should be written to {resource_file:?}");
assert_eq!(std::fs::read(&resource_file).unwrap(), payload);
// Reopen the export: the registry chain now leads with Embedded, keeping the URL as fallback,
// and the bytes are resolvable from the export itself with no byte store.
let reopened = GddV1::open(&dest).await.unwrap_or_else(|error| panic!("open export failed: {error:?}"));
assert!(reopened.has_resource(&hash).await, "embedded bytes should be resolvable from the export");
let entry = reopened.registry().resources.get(&id).expect("resource entry survived export");
assert_eq!(entry.hash, Some(hash));
let embedded = serde_json::to_value(DataSource::Embedded).unwrap();
let url = serde_json::to_value(DataSource::Url("https://example.com/r.bin".parse().unwrap())).unwrap();
let chain: Vec<_> = entry.sources.iter().map(|(_, value)| value.source.clone()).collect();
assert_eq!(chain, vec![embedded, url], "Embedded leads the chain, URL kept as fallback");
});
}
/// A plain export (no `embed_all_resources`) still materializes the bytes of an already-`Embedded`
/// resource, pulling from the byte store when the working copy doesn't hold them (the editor case
/// where bytes live in the app-global cache, not the per-document working copy).
#[test]
fn export_materializes_embedded_resource_from_byte_store() {
use document_format::{ExportFormat, ExportOptions};
use document_graph_storage::NoMetadata;
use graph_craft::application_io::resource::ResourceStorage;
use graphene_resource::{DataSource, ResourceHash, ResourceId, ResourceRegistry};
futures::executor::block_on(async {
let mut gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(9), 0xBEEF, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
// An Embedded resource whose bytes live only in the byte store, not the working copy.
let payload = b"embedded bytes in the cache";
let hash = ResourceHash::from(&payload[..]);
let byte_store = empty_byte_store();
byte_store.store(payload);
let mut resources = ResourceRegistry::new();
let id = ResourceId::new();
resources.push_source_back(&id, DataSource::Embedded);
resources.resolve(&id, hash);
gdd.commit_from_runtime(&network_referencing_resource(id), &NoMetadata, &resources, &byte_store)
.unwrap_or_else(|error| panic!("commit_from_runtime failed: {error:?}"));
assert!(!gdd.has_resource(&hash).await, "bytes should not be in the working copy");
let dir = tempfile::tempdir().unwrap();
let dest = dir.path().join("plain");
// Default options: embed_all_resources is false.
gdd.export(&dest, ExportFormat::Folder, ExportOptions::default(), &byte_store, None)
.await
.unwrap_or_else(|error| panic!("export failed: {error:?}"));
let resource_file = dest.join("resources").join(format!("{hash}"));
assert!(resource_file.exists(), "embedded resource bytes should be pulled from the store into {resource_file:?}");
assert_eq!(std::fs::read(&resource_file).unwrap(), payload);
});
}
/// A document with un-retired hot ops (e.g. a freshly converted document saved without an interaction
/// boundary, or a save mid-drag) must export losslessly: the hot log travels alongside history and the
/// reopened document reflects the staged edits. Regression for the converted-artwork "no history"
/// export, which previously failed at `embed_resource_sources` and fell back to the legacy blob.
#[test]
fn export_round_trips_unretired_hot_ops() {
use document_graph_storage::NoMetadata;
use graph_craft::application_io::resource::ResourceStorage;
use graphene_resource::{DataSource, ResourceId, ResourceRegistry};
futures::executor::block_on(async {
let mut gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(3), 0xF15E, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
let payload = b"hot resource bytes";
let hash = graphene_resource::ResourceHash::from(&payload[..]);
let byte_store = empty_byte_store();
byte_store.store(payload);
let mut resources = ResourceRegistry::new();
let id = ResourceId::new();
resources.push_source_back(&id, DataSource::Embedded);
resources.resolve(&id, hash);
// Stage into the working copy without retiring, leaving the conversion in the hot log.
gdd.stage_runtime_snapshot(&network_referencing_resource(id), &NoMetadata, &resources, &byte_store)
.unwrap_or_else(|error| panic!("stage_runtime_snapshot failed: {error:?}"));
assert!(!gdd.session().hot_log().is_empty(), "staging without retiring should leave hot ops");
assert!(gdd.registry().resources.contains_key(&id), "working registry should reflect the staged resource");
// Export to an archive and reopen it from scratch: the staged (hot) state must survive.
let archive = gdd
.export_to_bytes(document_format::ExportFormat::Zip, document_format::ExportOptions::default(), &byte_store, None)
.await
.unwrap_or_else(|error| panic!("export_to_bytes failed: {error:?}"));
let reopened = GddV1::open_from_archive(&archive, empty_container(), GddV1Layout)
.await
.unwrap_or_else(|error| panic!("open_from_archive failed: {error:?}"));
assert!(reopened.registry().resources.contains_key(&id), "the staged resource must survive export and reopen");
});
}
#[test]
fn open_in_rejects_future_format_version() {
futures::executor::block_on(async {
let container = empty_container();
let layout = GddV1Layout;
let mut future_version = Manifest::new(0xC0DE, "ed".into(), "std".into());
future_version.format_version = manifest::SUPPORTED_FORMAT_VERSION + 1;
io::write_single(&container, layout.manifest_basename(), Codec::Json, &future_version).unwrap();
match GddV1::open_in(container, layout).await {
Err(Error::UnsupportedVersion { .. }) => {}
Ok(_) => panic!("expected UnsupportedVersion, got Ok"),
Err(other) => panic!("expected UnsupportedVersion, got {other:?}"),
}
});
}
#[test]
fn create_in_records_default_codecs_in_manifest() {
futures::executor::block_on(async {
let gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(1), 0xAB, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
let codecs = gdd.manifest().codecs;
assert_eq!(codecs.registry, Codec::MessagePack);
assert_eq!(codecs.history, Codec::MessagePackFrames);
assert_eq!(codecs.hot_log, Codec::MessagePackFrames);
assert_eq!(codecs.session, Codec::Json);
});
}
/// The `RegisterPeer` op auto-emitted on the first commit rides the hot-op pipeline through
/// persistence and retirement, so the `peer_users` mapping survives a reopen.
#[test]
fn first_commit_registers_peer_and_survives_reopen() {
use document_graph_storage::{NoMetadata, UserId};
use graph_craft::application_io::resource::HashMapResourceStorage;
use graph_craft::document::{DocumentNode, DocumentNodeImplementation, NodeInput, NodeNetwork};
use graph_craft::{ProtoNodeIdentifier, concrete};
use graphene_resource::ResourceRegistry;
futures::executor::block_on(async {
let mut gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(21), 0xAB, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
let network = NodeNetwork {
exports: vec![NodeInput::node(core_types::uuid::NodeId(0), 0)],
nodes: [(
core_types::uuid::NodeId(0),
DocumentNode {
inputs: vec![NodeInput::import(concrete!(u32), 0)],
implementation: DocumentNodeImplementation::ProtoNode(ProtoNodeIdentifier::new("graphene_core::ops::identity::IdentityNode")),
..Default::default()
},
)]
.into_iter()
.collect(),
..Default::default()
};
gdd.commit_from_runtime(&network, &NoMetadata, &ResourceRegistry::new(), &HashMapResourceStorage::new())
.unwrap_or_else(|error| panic!("commit_from_runtime failed: {error:?}"));
assert_eq!(gdd.registry().peer_users.get(&PeerId(21)), Some(&UserId(21)), "first commit registers the peer");
let (working, layout) = gdd.into_storage();
let reopened = GddV1::open_in(working, layout).await.unwrap_or_else(|error| panic!("open_in failed: {error:?}"));
assert_eq!(reopened.registry().peer_users.get(&PeerId(21)), Some(&UserId(21)), "registration survives reopen");
});
}
#[test]
fn persist_path_writes_at_manifest_declared_codec_paths() {
// The manifest declares the on-disk codec for each payload; the persist path must write at the
// extension that codec implies, and reopen (which reads the codec from the manifest) must find them.
futures::executor::block_on(async {
use document_container::AsyncContainer;
let mut gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(5), 0xDEAD, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
let hot_op = HotOp {
op: RegistryDelta::AddNetwork {
id: ROOT_NETWORK,
network: Network::default(),
},
timestamp: TimeStamp { counter: 1, peer: PeerId(5) },
};
gdd.apply_hot_op(hot_op).unwrap_or_else(|error| panic!("apply_hot_op failed: {error:?}"));
let (working, layout) = gdd.into_storage();
// Defaults: hot log is MessagePackFrames (.frames), manifest is always JSON.
assert!(working.exists(&io::path_for(layout.hot_log_basename(), Codec::MessagePackFrames)).await);
assert!(working.exists(&io::path_for(layout.manifest_basename(), Codec::Json)).await);
let reopened = GddV1::open_in(working, layout).await.unwrap_or_else(|error| panic!("open_in failed: {error:?}"));
assert!(reopened.registry().networks.contains_key(&ROOT_NETWORK));
});
}
/// Complete declaration round-trip through the byte store: committing a runtime network with a
/// proto-node persists its `ProtoNode` content into a `ResourceStorage`, and resolving declarations
/// back through that store reconstructs the proto-node identifier in `to_runtime`. This is the
/// editor-shaped path (declaration bytes live in the resource store, not the Gdd container).
#[test]
fn declarations_round_trip_through_byte_store() {
use document_graph_storage::NoMetadata;
use graph_craft::application_io::resource::HashMapResourceStorage;
use graph_craft::document::{DocumentNode, DocumentNodeImplementation, NodeInput, NodeNetwork};
use graph_craft::{ProtoNodeIdentifier, concrete};
use graphene_resource::ResourceRegistry;
const PROTO: &str = "graphene_core::ops::identity::IdentityNode";
futures::executor::block_on(async {
let network = NodeNetwork {
exports: vec![NodeInput::node(core_types::uuid::NodeId(0), 0)],
nodes: [(
core_types::uuid::NodeId(0),
DocumentNode {
inputs: vec![NodeInput::import(concrete!(u32), 0)],
implementation: DocumentNodeImplementation::ProtoNode(ProtoNodeIdentifier::new(PROTO)),
..Default::default()
},
)]
.into_iter()
.collect(),
..Default::default()
};
let mut gdd = GddV1::create_in(empty_container(), GddV1Layout, PeerId(1), 0xAB, "ed".into(), "std".into())
.await
.unwrap_or_else(|error| panic!("create_in failed: {error:?}"));
// Commit: declaration bytes flow into the byte store, not the Gdd container.
let byte_store = HashMapResourceStorage::new();
gdd.commit_from_runtime(&network, &NoMetadata, &ResourceRegistry::new(), &byte_store)
.unwrap_or_else(|error| panic!("commit_from_runtime failed: {error:?}"));
// Resolve declarations back through the store and convert to a runtime network.
let declarations = gdd.declarations(&byte_store).await;
assert_eq!(declarations.len(), 1, "expected one proto-node declaration resolved from the byte store");
let (converted, _entries) = gdd.registry().to_runtime_with_metadata(&declarations).unwrap_or_else(|error| panic!("to_runtime failed: {error:?}"));
let node = converted.nodes.values().next().expect("converted network has the node");
match &node.implementation {
DocumentNodeImplementation::ProtoNode(identifier) => assert_eq!(identifier.as_str(), PROTO, "proto-node identifier survived the byte-store round-trip"),
other => panic!("expected a ProtoNode implementation, got {other:?}"),
}
});
}
@@ -0,0 +1,27 @@
[package]
name = "document-graph-storage"
description = "Provides a delta based graph representation used in the Graphite file format"
edition.workspace = true
version.workspace = true
license.workspace = true
authors.workspace = true
[features]
conversion = ["dep:graph-craft", "dep:core-types"]
default = ["conversion"]
[dependencies]
graph-craft = { workspace = true, optional = true }
core-types = { workspace = true, optional = true }
graphene-resource = { workspace = true }
thiserror = { workspace = true }
serde = { workspace = true }
serde_json = { workspace = true }
blake3 = { workspace = true }
rustc-hash = { workspace = true }
rmp-serde = { workspace = true }
[dev-dependencies]
graph-craft = { workspace = true, features = ["loading"] }
core-types = { workspace = true }
@@ -0,0 +1,71 @@
use crate::TimeStamp;
use serde::{Deserialize, Serialize};
use std::collections::BTreeMap;
/// Attribute keys. Glob-import (`use crate::attr::*`) at conversion sites.
///
/// `ui::*` keys are namespaced per CRDT design so each value gets its own LWW timestamp. Per-input
/// keys live on `Node.inputs_attributes[i]`; per-network keys live on `Network.attributes`.
pub mod attr;
/// A type-erased attribute value paired with the timestamp at which it was last set.
#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
pub struct Value {
pub value: serde_json::Value,
pub timestamp: TimeStamp,
}
impl Value {
pub fn new(value: serde_json::Value, timestamp: TimeStamp) -> Self {
Self { value, timestamp }
}
}
pub type Attributes = BTreeMap<String, Value>;
/// Write helpers for `Attributes`.
pub trait AttributesWrite {
/// Inserts a JSON value under `key`.
fn set(&mut self, key: &str, value: serde_json::Value, timestamp: TimeStamp);
/// Serializes `value` and inserts it under `key`.
fn set_serialized<T: serde::Serialize>(&mut self, key: &str, value: &T, timestamp: TimeStamp) -> Result<(), serde_json::Error> {
self.set(key, serde_json::to_value(value)?, timestamp);
Ok(())
}
/// Inserts only when `value != default`, so the read side falls back to the same default.
fn set_if_not_default<T: serde::Serialize + PartialEq>(&mut self, key: &str, value: &T, default: &T, timestamp: TimeStamp) -> Result<(), serde_json::Error> {
if value != default {
self.set_serialized(key, value, timestamp)?;
}
Ok(())
}
}
impl AttributesWrite for Attributes {
fn set(&mut self, key: &str, value: serde_json::Value, timestamp: TimeStamp) {
self.insert(key.to_string(), Value { value, timestamp });
}
}
/// Typed read helpers for `Attributes`.
pub trait AttributesRead {
/// Deserializes the value under `key`, or `None` if missing or undecodable.
fn get_typed<T: serde::de::DeserializeOwned>(&self, key: &str) -> Option<T>;
/// Same as `get_typed`, falling back to `default`.
fn get_or<T: serde::de::DeserializeOwned>(&self, key: &str, default: T) -> T {
self.get_typed(key).unwrap_or(default)
}
/// Same as `get_typed`, falling back to `T::default()`.
fn get_or_default<T: serde::de::DeserializeOwned + Default>(&self, key: &str) -> T {
self.get_typed(key).unwrap_or_default()
}
}
impl AttributesRead for Attributes {
fn get_typed<T: serde::de::DeserializeOwned>(&self, key: &str) -> Option<T> {
self.get(key).and_then(|v| serde_json::from_value(v.value.clone()).ok())
}
}
@@ -0,0 +1,67 @@
pub mod node {
pub const CALL_ARGUMENT: &str = "call_argument";
pub const VISIBLE: &str = "visible";
pub const SKIP_DEDUPLICATION: &str = "skip_deduplication";
pub const REFLECTION_METADATA: &str = "reflection_metadata";
pub const ORIGINAL_NODE_ID: &str = "original_node_id";
pub mod input {
pub const IMPORT_TYPE: &str = "import_type";
pub mod ui {
pub const NAME: &str = "ui::name";
pub const DESCRIPTION: &str = "ui::description";
pub const WIDGET_OVERRIDE: &str = "ui::widget_override";
/// Prefix for `InputPersistentMetadata::data` entries. Full key: `ui::data::<sub_key>`.
pub const DATA_PREFIX: &str = "ui::data::"; // TODO: Remove and make runtime strongly typed again
}
}
pub mod ui {
pub const POSITION: &str = "ui::position";
pub const IS_LAYER: &str = "ui::is_layer";
pub const DISPLAY_NAME: &str = "ui::display_name";
pub const LOCKED: &str = "ui::locked";
pub const PINNED: &str = "ui::pinned";
pub const OUTPUT_NAMES: &str = "ui::output_names";
pub const REFERENCE: &str = "ui::reference"; // TODO: Remove?
}
}
pub mod session {
pub mod network {
pub const PREVIEWING: &str = "ui::previewing";
// TODO: Remove these graph ui nav-specific attributes
pub const NAV_PTZ: &str = "ui::nav::ptz";
pub const NAV_TRANSFORM: &str = "ui::nav::transform";
pub const NAV_WIDTH: &str = "ui::nav::width";
}
pub mod doc {
// Document-level editor chrome, stored in `Registry.attributes` (document scope). Each setting is
// its own key so concurrent edits to one don't clobber another.
pub const PTZ: &str = "ui::ptz";
pub const RENDER_MODE: &str = "ui::render_mode";
pub const OVERLAYS: &str = "ui::overlays";
pub const RULERS_VISIBLE: &str = "ui::rulers_visible";
pub const SNAPPING: &str = "ui::snapping";
pub const COLLAPSED: &str = "ui::collapsed";
}
}
pub mod registry {
pub const EXPORTED_NODES: &str = "exported_nodes";
}
pub mod network {
/// Whole-map LWW of a network's `scope_injections` (`key -> (storage NodeId, Type)`), stored as a
/// serialized blob so its shape can evolve (e.g. dropping the `Type`) without a model change. The
/// node references use stable storage IDs, resolved back to runtime-local IDs on conversion.
pub const SCOPE_INJECTIONS: &str = "scope_injections";
}
pub mod delta {
/// Marks the last delta of a user interaction, so the undo cursor steps per-interaction, not per-delta.
pub const INTERACTION_END: &str = "interaction_end";
}
@@ -0,0 +1,226 @@
use crate::{Attributes, AttributesWrite, Network, NetworkId, Node, NodeId, NodeInput, PeerId, ResourceEntry, ResourceId, Rev, SourceKey, TimeStamp, UserId, Value, attr, compute_rev};
use graphene_resource::ResourceHash;
use serde::{Deserialize, Serialize};
/// Content-addressed delta: `id` is `blake3_128(parents, author, timestamp, delta_type)`.
///
/// `reverse` is state-dependent undo bookkeeping (it captures pre-state at the moment the forward
/// op was applied), so it's serialized for storage but excluded from the identity hash — two peers
/// observing the same forward delta against different local states would otherwise compute
/// different Revs for the same logical op.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct Delta {
pub id: Rev,
/// Primary parent; `None` for the root delta.
pub parent: Option<Rev>,
pub author: PeerId,
pub timestamp: TimeStamp,
pub kind: RegistryDelta,
pub reverse: RegistryDelta,
/// Local, mutable annotations on this commit (interaction-end marker, future commit messages / labels).
/// Deliberately excluded from `compute_rev`: relabeling a commit must not change its content-addressed
/// identity, and two peers annotating the same op differently must still dedup to one `Rev`.
#[serde(default, skip_serializing_if = "Attributes::is_empty")]
pub attributes: Attributes,
}
impl Delta {
pub fn new(parent: Option<Rev>, author: PeerId, timestamp: TimeStamp, kind: RegistryDelta, reverse: RegistryDelta) -> Self {
let id = compute_rev(parent, author, timestamp, &kind);
Self {
id,
parent,
author,
timestamp,
kind,
reverse,
attributes: Attributes::default(),
}
}
/// Build a merge delta joining `tips` into one node. See [`RegistryDelta::Merge`] for the semantics.
pub fn merge(tips: impl IntoIterator<Item = Rev>, author: PeerId, timestamp: TimeStamp) -> Self {
let mut parents: Vec<Rev> = tips.into_iter().collect();
parents.sort_unstable();
parents.dedup();
let parent = parents.first().copied();
let extra_parents = parents.split_first().map(|(_, rest)| rest.to_vec()).unwrap_or_default();
let kind = RegistryDelta::Merge { extra_parents };
let id = compute_rev(parent, author, timestamp, &kind);
Self {
id,
parent,
author,
timestamp,
reverse: kind.clone(),
kind,
attributes: Attributes::default(),
}
}
/// Every parent: the primary `parent` (absent for the root) plus a merge's `extra_parents`.
pub fn all_parents(&self) -> impl Iterator<Item = Rev> + '_ {
let extras = match &self.kind {
RegistryDelta::Merge { extra_parents } => extra_parents.as_slice(),
_ => &[],
};
self.parent.into_iter().chain(extras.iter().copied())
}
/// Mark this delta as the last op of a user interaction, so the undo cursor treats it as a checkpoint.
pub fn mark_interaction_end(&mut self, timestamp: TimeStamp) {
self.attributes.set(attr::delta::INTERACTION_END, serde_json::Value::Bool(true), timestamp);
}
pub fn is_interaction_end(&self) -> bool {
self.attributes.get(attr::delta::INTERACTION_END).is_some_and(|marker| marker.value == serde_json::Value::Bool(true))
}
/// The content-addressed `Rev` this delta's identity fields hash to. Equals `id` for a delta built
/// via `new`/`merge`; differs only if `id` was tampered with or the hash derivation changed.
pub fn recomputed_id(&self) -> Rev {
compute_rev(self.parent, self.author, self.timestamp, &self.kind)
}
/// Whether `id` matches the recomputed content hash. `Delta` deserializes without checking this
/// (the hash is not cheap over a large history); callers verify explicitly when they don't trust
/// the source via [`Session::verify_history`].
pub fn has_valid_id(&self) -> bool {
self.id == self.recomputed_id()
}
}
/// Op payload. Timestamps live on the wrapping `Delta` — one per delta, applied to all LWW-eligible
/// writes within. See `notes/document-format-collaboration.md`.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub enum RegistryDelta {
AddNode {
id: NodeId,
node: Node,
},
/// `snapshot` lets the reverse `AddNode` rebuild without reading the (already-removed) node from
/// the registry, mirroring `RemoveNetwork`.
RemoveNode {
id: NodeId,
snapshot: Node,
},
ChangeNodeInput {
id: NodeId,
index: u32,
new_input: NodeInput,
},
ChangeNodeAttribute {
id: NodeId,
delta: AttributeDelta,
},
ChangeNodeInputAttribute {
id: NodeId,
index: u32,
delta: AttributeDelta,
},
/// LWW per slot. `export == None` removes the slot.
SetNetworkExport {
id: NetworkId,
index: u32,
export: Option<NodeInput>,
},
/// Per-network attribute change, LWW per key. Mirrors `ChangeDocumentAttribute`.
ChangeNetworkAttribute {
id: NetworkId,
delta: AttributeDelta,
},
AddNetwork {
id: NetworkId,
network: Network,
},
/// `snapshot` lets the reverse delta rebuild without re-walking history.
RemoveNetwork {
id: NetworkId,
snapshot: Network,
},
/// Register a whole resource entry at once. Overwrites any existing entry for `id`; the reverse
/// of `RemoveResource`, the way `AddNetwork` pairs with `RemoveNetwork`.
AddResource {
id: ResourceId,
entry: ResourceEntry,
},
/// LWW on a resource's resolved content hash. Creates the resource entry if absent.
/// Concurrent resolves agree by construction (the hash is content-derived), so LWW is safe.
SetResourceHash {
id: ResourceId,
hash: Option<ResourceHash>,
},
/// Remove a whole resource entry. `snapshot` is the state of the resource before it was removed.
RemoveResource {
id: ResourceId,
snapshot: ResourceEntry,
},
/// Add (or LWW-overwrite) one entry in a resource's source fallback chain. The source body is
/// type-erased; `key` carries the fractional priority + peer that order it. Add-wins: concurrent
/// adds at distinct keys all survive. Creates the resource entry if absent.
AddSource {
id: ResourceId,
key: SourceKey,
source: serde_json::Value,
},
/// Remove one entry from a resource's source chain. LWW against the entry's timestamp.
RemoveSource {
id: ResourceId,
key: SourceKey,
},
/// Append-only registration of a device's `PeerId` against its owning `UserId`.
/// First write wins; conflicting re-registration errors. Duplicate identical registration
/// is a no-op. Not LWW — the mapping is forever.
RegisterPeer {
peer: PeerId,
user: UserId,
},
ChangeDocumentAttribute {
delta: AttributeDelta,
},
/// Joins divergent history tips into one shared node. A registry no-op on replay (it only collapses
/// tips so `head` stays a single `Rev`); the joined tips are `Delta::parent` (the lowest `Rev`) plus
/// these `extra_parents` (sorted). Identity is the parent set alone, so two peers merging the same
/// tips mint the identical delta and it dedups.
Merge {
extra_parents: Vec<Rev>,
},
// Allow for future delta types without a model change
Other(serde_json::Value),
}
/// `value: None` means remove. The timestamp comes from the wrapping `Delta`.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct AttributeDelta {
pub key: String,
pub value: Option<serde_json::Value>,
}
pub(crate) fn reverse_attribute_delta(delta: &AttributeDelta, attributes: &Attributes) -> AttributeDelta {
AttributeDelta {
key: delta.key.clone(),
value: attributes.get(&delta.key).map(|previous| previous.value.clone()),
}
}
pub(crate) fn apply_attribute_delta(delta: AttributeDelta, timestamp: TimeStamp, force: bool, attributes: &mut Attributes) {
let AttributeDelta { key, value } = delta;
match value {
Some(value) => match attributes.entry(key) {
std::collections::btree_map::Entry::Occupied(mut entry) => {
if force || timestamp > entry.get().timestamp {
entry.insert(Value { value, timestamp });
}
}
std::collections::btree_map::Entry::Vacant(entry) => {
entry.insert(Value { value, timestamp });
}
},
None => {
let should_remove = force || attributes.get(&key).is_none_or(|existing| timestamp > existing.timestamp);
if should_remove {
attributes.remove(&key);
}
}
}
}
@@ -0,0 +1,425 @@
use std::collections::HashSet;
use crate::{AttributeDelta, NetworkId, Node, NodeId, Registry, RegistryDelta, ResourceEntry, ResourceId};
/// Collect a `HashSet` walk (difference/intersection) into ascending order. The sets iterate in
/// random order, so sorting keeps `compute_deltas` emitting a deterministic delta sequence.
fn sorted<'a, T: Ord + Copy + 'a>(ids: impl Iterator<Item = &'a T>) -> Vec<T> {
let mut ids: Vec<T> = ids.copied().collect();
ids.sort_unstable();
ids
}
/// Minimal set of deltas to transform `from` into `to`.
///
/// Emits timestamp-less op shapes; the caller (`Document::commit_local` or equivalent) wraps each
/// in a `Delta` with a fresh clock tick.
pub fn compute_deltas(from: &Registry, to: &Registry) -> Vec<RegistryDelta> {
let mut deltas = Vec::new();
let from_network_ids: HashSet<NetworkId> = from.networks.keys().copied().collect();
let to_network_ids: HashSet<NetworkId> = to.networks.keys().copied().collect();
// AddNetwork before any AddNode that references it. `HashSet` difference/intersection iterate in
// random order, so every set walk below is sorted to keep the emitted delta sequence (and thus the
// resulting `Rev` chain) deterministic across runs.
for network_id in sorted(to_network_ids.difference(&from_network_ids)) {
deltas.push(RegistryDelta::AddNetwork {
id: network_id,
network: to.networks[&network_id].clone(),
});
}
let from_node_ids: HashSet<NodeId> = from.node_instances.keys().copied().collect();
let to_node_ids: HashSet<NodeId> = to.node_instances.keys().copied().collect();
for node_id in sorted(from_node_ids.difference(&to_node_ids)) {
deltas.push(RegistryDelta::RemoveNode {
id: node_id,
snapshot: from.node_instances[&node_id].clone(),
});
}
for node_id in sorted(to_node_ids.difference(&from_node_ids)) {
deltas.push(RegistryDelta::AddNode {
id: node_id,
node: to.node_instances[&node_id].clone(),
});
}
for node_id in sorted(from_node_ids.intersection(&to_node_ids)) {
let from_node = &from.node_instances[&node_id];
let to_node = &to.node_instances[&node_id];
// No `ChangeImplementation` op; the only path is remove + re-add. Same for input-count and
// containing-network changes (a moved node has no in-place op either). `inputs_attributes` is
// checked too: the per-slot loops below `zip` only the shared prefix, so a length change there
// must force a remove + re-add rather than silently dropping the extra slots.
let structural_change = !nodes_have_same_implementation(from_node, to_node) || from_node.inputs.len() != to_node.inputs.len() || from_node.network != to_node.network;
if structural_change {
deltas.push(RegistryDelta::RemoveNode {
id: node_id,
snapshot: from_node.clone(),
});
deltas.push(RegistryDelta::AddNode { id: node_id, node: to_node.clone() });
continue;
}
// Compare by value, ignoring the per-slot timestamp. Timestamps are derived from the diff
// (assigned by the caller via clock.tick), not part of the diff itself: a slot whose value
// is unchanged but whose timestamp differs should not emit a delta.
for (input_idx, (from_slot, to_slot)) in from_node.inputs.iter().zip(&to_node.inputs).enumerate() {
if from_slot.input != to_slot.input {
deltas.push(RegistryDelta::ChangeNodeInput {
id: node_id,
index: input_idx as u32,
new_input: to_slot.input.clone(),
});
}
}
for delta in compute_attribute_deltas(&from_node.attributes, &to_node.attributes) {
deltas.push(RegistryDelta::ChangeNodeAttribute { id: node_id, delta });
}
for (input_idx, (from_input, to_input)) in from_node.inputs.iter().zip(&to_node.inputs).enumerate() {
for delta in compute_attribute_deltas(&from_input.attributes, &to_input.attributes) {
deltas.push(RegistryDelta::ChangeNodeInputAttribute {
id: node_id,
index: input_idx as u32,
delta,
});
}
}
}
for network_id in sorted(from_network_ids.difference(&to_network_ids)) {
deltas.push(RegistryDelta::RemoveNetwork {
id: network_id,
snapshot: from.networks[&network_id].clone(),
});
}
for network_id in sorted(from_network_ids.intersection(&to_network_ids)) {
let from_network = &from.networks[&network_id];
let to_network = &to.networks[&network_id];
let max_len = from_network.exports.len().max(to_network.exports.len());
for slot_idx in 0..max_len {
let from_slot = from_network.exports.get(slot_idx);
let to_slot = to_network.exports.get(slot_idx);
let from_target = from_slot.and_then(|s| s.target.as_ref());
let to_target = to_slot.and_then(|s| s.target.as_ref());
if from_target != to_target {
deltas.push(RegistryDelta::SetNetworkExport {
id: network_id,
index: slot_idx as u32,
export: to_target.cloned(),
});
}
}
// Per-network attributes.
for delta in compute_attribute_deltas(&from_network.attributes, &to_network.attributes) {
deltas.push(RegistryDelta::ChangeNetworkAttribute { id: network_id, delta });
}
}
// Document-level attributes (`ui::doc::*`, format version, ...).
for delta in compute_attribute_deltas(&from.attributes, &to.attributes) {
deltas.push(RegistryDelta::ChangeDocumentAttribute { delta });
}
compute_resource_deltas(from, to, &mut deltas);
deltas
}
/// Diff the resource store, emitting whole-entry add/remove for resources that appear or vanish and
/// fine-grained hash/source ops for resources present in both. Value-only: per-entry and per-source
/// timestamps are derived by the caller, so an unchanged resource emits nothing.
fn compute_resource_deltas(from: &Registry, to: &Registry, deltas: &mut Vec<RegistryDelta>) {
let from_ids: HashSet<ResourceId> = from.resources.keys().copied().collect();
let to_ids: HashSet<ResourceId> = to.resources.keys().copied().collect();
for id in sorted(from_ids.difference(&to_ids)) {
deltas.push(RegistryDelta::RemoveResource {
id,
snapshot: from.resources[&id].clone(),
});
}
for id in sorted(to_ids.difference(&from_ids)) {
deltas.push(RegistryDelta::AddResource { id, entry: to.resources[&id].clone() });
}
for id in sorted(from_ids.intersection(&to_ids)) {
diff_resource_entry(id, &from.resources[&id], &to.resources[&id], deltas);
}
}
/// Per-entry diff for a resource present in both registries: hash change, then source chain
/// additions/changes/removals.
fn diff_resource_entry(id: ResourceId, from: &ResourceEntry, to: &ResourceEntry, deltas: &mut Vec<RegistryDelta>) {
if from.hash != to.hash {
deltas.push(RegistryDelta::SetResourceHash { id, hash: to.hash });
}
for (key, _) in &from.sources {
if to.source(key).is_none() {
deltas.push(RegistryDelta::RemoveSource { id, key: *key });
}
}
// Compare source bodies only; the per-source timestamp is derived from the diff, not part of it.
for (key, to_source) in &to.sources {
if from.source(key).is_none_or(|from_source| from_source.source != to_source.source) {
deltas.push(RegistryDelta::AddSource {
id,
key: *key,
source: to_source.source.clone(),
});
}
}
}
fn nodes_have_same_implementation(a: &Node, b: &Node) -> bool {
use crate::Implementation::*;
match (&a.implementation, &b.implementation) {
(ProtoNode(a_id), ProtoNode(b_id)) => a_id == b_id,
(Network(a_id), Network(b_id)) => a_id == b_id,
_ => false,
}
}
fn compute_attribute_deltas(from: &crate::Attributes, to: &crate::Attributes) -> Vec<AttributeDelta> {
let mut deltas = Vec::new();
for key in from.keys() {
if !to.contains_key(key) {
deltas.push(AttributeDelta { key: key.clone(), value: None });
}
}
// Compare by `value` only; the per-entry `timestamp` is derived from the diff, not part of it.
for (key, to_value) in to {
if from.get(key).is_none_or(|from_value| from_value.value != to_value.value) {
deltas.push(AttributeDelta {
key: key.clone(),
value: Some(to_value.value.clone()),
});
}
}
deltas
}
#[cfg(test)]
mod tests {
use super::*;
use crate::{Attributes, ExportSlot, Network, Node, NodeInput, TimeStamp};
#[test]
fn test_compute_deltas_empty() {
let registry = Registry::default();
let deltas = compute_deltas(&registry, &registry);
assert_eq!(deltas.len(), 0, "No deltas should be generated for identical registries");
}
/// The emitted delta sequence must not depend on `HashMap`/`HashSet` iteration order, which varies
/// per run and per compiler version. Building the same registry repeatedly (each `HashMap` gets a
/// fresh random seed) must yield identical `AddNode` order, since the diff sorts its set walks.
#[test]
fn compute_deltas_emits_nodes_in_deterministic_order() {
let make_registry = || {
let mut registry = Registry::default();
registry.networks.insert(NetworkId(0), Network::default());
for node_id in [50, 3, 17, 999, 1, 42, 8, 256, 100, 7] {
registry.node_instances.insert(NodeId(node_id), Node::dummy());
}
registry
};
let empty = Registry::default();
let add_node_ids = |registry: &Registry| -> Vec<NodeId> {
compute_deltas(&empty, registry)
.into_iter()
.filter_map(|delta| match delta {
RegistryDelta::AddNode { id: node_id, .. } => Some(node_id),
_ => None,
})
.collect()
};
let expected = vec![1, 3, 7, 8, 17, 42, 50, 100, 256, 999].into_iter().map(NodeId).collect::<Vec<_>>();
for _ in 0..16 {
assert_eq!(add_node_ids(&make_registry()), expected, "AddNode order must be deterministic (ascending)");
}
}
#[test]
fn test_compute_deltas_add_node() {
let from = Registry::default();
let mut to = from.clone();
let node = Node::dummy();
to.node_instances.insert(NodeId(42), node);
let deltas = compute_deltas(&from, &to);
assert_eq!(deltas.len(), 1);
assert!(matches!(deltas[0], RegistryDelta::AddNode { id: NodeId(42), .. }));
}
/// A change in `inputs_attributes` length is structural: the per-slot diff only `zip`s the shared
/// prefix, so it must force a remove + re-add rather than dropping the extra attribute slots.
#[test]
fn compute_deltas_treats_inputs_attributes_length_change_as_structural() {
// Same implementation/inputs/network in both registries; only `inputs_attributes` length differs.
let base = Node::dummy();
let mut from = Registry::default();
from.node_instances.insert(NodeId(42), base.clone());
let mut to = from.clone();
to.node_instances.get_mut(&NodeId(42)).unwrap().inputs.push(crate::InputSlot {
input: NodeInput::Import { index: 0 },
timestamp: TimeStamp::ORIGIN,
attributes: Attributes::new(),
});
let deltas = compute_deltas(&from, &to);
assert!(
deltas.iter().any(|delta| matches!(delta, RegistryDelta::RemoveNode { id: NodeId(42), .. })) && deltas.iter().any(|delta| matches!(delta, RegistryDelta::AddNode { id: NodeId(42), .. })),
"an inputs_attributes length change must emit RemoveNode + AddNode, got {deltas:?}"
);
}
#[test]
fn test_compute_deltas_change_network_attribute() {
use crate::{AttributesWrite, TimeStamp};
let mut from = Registry::default();
from.networks.insert(NetworkId(0), Network::default());
let mut to = from.clone();
to.networks
.get_mut(&NetworkId(0))
.unwrap()
.attributes
.set("ui::nav::width", serde_json::json!(640.0), TimeStamp::ORIGIN);
let deltas = compute_deltas(&from, &to);
assert_eq!(deltas.len(), 1, "a changed per-network attribute must emit one delta");
assert!(
matches!(&deltas[0], RegistryDelta::ChangeNetworkAttribute { id: NetworkId(0), delta } if delta.key == "ui::nav::width"),
"expected ChangeNetworkAttribute for ui::nav::width, got {:?}",
deltas[0]
);
}
#[test]
fn test_compute_deltas_remove_node() {
let mut from = Registry::default();
let node = Node::dummy();
from.node_instances.insert(NodeId(42), node);
let to = Registry::default();
let deltas = compute_deltas(&from, &to);
assert_eq!(deltas.len(), 1);
assert!(matches!(deltas[0], RegistryDelta::RemoveNode { id: NodeId(42), .. }));
}
#[test]
fn test_compute_deltas_modify_attribute() {
let mut from = Registry::default();
let mut node = Node::dummy();
let stamp = |counter: u64| TimeStamp { counter, peer: crate::PeerId(0) };
node.attributes.insert(
"test".to_string(),
crate::Value {
value: serde_json::json!("old"),
timestamp: stamp(0),
},
);
from.node_instances.insert(NodeId(42), node);
let mut to = from.clone();
to.node_instances.get_mut(&NodeId(42)).unwrap().attributes.insert(
"test".to_string(),
crate::Value {
value: serde_json::json!("new"),
timestamp: stamp(1),
},
);
let deltas = compute_deltas(&from, &to);
assert_eq!(deltas.len(), 1);
assert!(matches!(
&deltas[0],
RegistryDelta::ChangeNodeAttribute { id: NodeId(42), delta: AttributeDelta { key, value: Some(_) } } if key == "test"
));
}
/// Document-level attributes (the `Registry.attributes` bucket) must diff into
/// `ChangeDocumentAttribute` deltas, so a document-scoped attribute change reaches the commit path.
/// (Per-peer `ui::doc::*` view settings live in `session.json`, not here.)
#[test]
fn test_compute_deltas_document_attribute() {
let stamp = |counter: u64| TimeStamp { counter, peer: crate::PeerId(0) };
let from = Registry::default();
let mut to = from.clone();
to.attributes.insert(
"doc::test_attribute".to_string(),
crate::Value {
value: serde_json::json!("value"),
timestamp: stamp(1),
},
);
let deltas = compute_deltas(&from, &to);
assert_eq!(deltas.len(), 1);
assert!(matches!(
&deltas[0],
RegistryDelta::ChangeDocumentAttribute { delta: AttributeDelta { key, value: Some(_) } } if key == "doc::test_attribute"
));
}
#[test]
fn test_compute_deltas_network_changes() {
let make_slot = |id: u64| ExportSlot {
target: Some(NodeInput::Node { id: NodeId(id), index: 0 }),
timestamp: TimeStamp::ORIGIN,
};
let mut from = Registry::default();
from.networks.insert(
NetworkId(0),
Network {
exports: vec![make_slot(1), make_slot(2)],
..Default::default()
},
);
let mut to = from.clone();
to.networks.get_mut(&NetworkId(0)).unwrap().exports.push(make_slot(3));
let deltas = compute_deltas(&from, &to);
// Only slot 2 changed (added). Slots 0 and 1 are unchanged so they don't emit ops.
assert_eq!(deltas.len(), 1);
assert!(matches!(
&deltas[0],
RegistryDelta::SetNetworkExport {
id: NetworkId(0),
index: 2,
export: Some(NodeInput::Node { id: NodeId(3), .. }),
..
}
));
}
}
@@ -0,0 +1,451 @@
use crate::{
CrdtError, Delta, ExportSlot, History, HotOp, LamportClock, MAX_EXPORT_SLOTS, NetworkId, NodeId, NodeInput, PeerId, Registry, RegistryDelta, ResourceEntry, Rev, SourceValue, TimeStamp,
apply_attribute_delta, reverse_attribute_delta,
};
#[derive(Clone, Debug)]
pub struct Document {
/// Working registry: retired state with the current hot ops applied on top. This is what live
/// reads and `registry()` observe, and what undo/redo force-apply against.
pub(crate) working_registry: Registry,
/// Live broadcast stream, applied to the `working_registry` on receive, GC'd at retirement.
/// Persisted for crash recovery so in-flight unretired work survives editor restarts.
pub(crate) hot_log: Vec<HotOp>,
/// The registry as of the last retirement, with no un-retired hot ops applied. Retirement computes
/// each delta's `reverse` against this (so LWW reverses capture the true pre-op value, not the
/// hot-polluted working state) and advances it, stamping fields at the fresh `T_retire`. Kept equal
/// to `registry` *by value* whenever the hot log is empty (undo/redo resync it after moving the
/// cursor), but field timestamps can differ: retirement bumps the snapshot's to `T_retire` while the
/// working registry keeps the staging-time timestamps. Benign while the local monotonic clock makes
/// new edits win
pub(crate) retired_snapshot: Registry,
/// User's cursor in their local history chain. `None` on an empty document (no commits yet).
pub(crate) head: Option<Rev>,
/// Retired delta DAG in topological (append) order. See [`History`](crate::History).
pub(crate) history: History,
/// Revs undone past (most-recent last), so `redo` can re-apply them. Local-view state the DAG can't
/// recover (a parent may have several children). A new edit while non-empty clears it.
pub(crate) redo_stack: Vec<Rev>,
pub(crate) clock: LamportClock,
pub(crate) peer: PeerId,
/// Latest retired commit on the local chain that has been broadcast to at least one peer.
/// Commits after this can be rewritten silently; commits at or before this are published
/// and require forward reverse-delta ops to undo. `None` means nothing broadcast yet.
pub(crate) last_broadcast_rev: Option<Rev>,
/// Shared-monotonic counter feeding `next_node_id`. Bumped on every mint regardless of which
/// peer is calling; collision avoidance comes from hashing `(self.peer, counter)`, so two peers
/// reading the same counter still produce distinct IDs.
pub(crate) next_node_counter: u64,
}
impl Document {
/// Mint a fresh `NodeId` scoped to this document's peer. The 64-bit ID is `blake3(peer, counter)`
/// truncated; the counter is shared across peers and persisted with the document.
pub fn next_node_id(&mut self) -> NodeId {
self.next_node_counter += 1;
let bytes = rmp_serde::to_vec(&(self.peer, self.next_node_counter)).expect("(PeerId, counter) must serialize");
let digest = blake3::hash(&bytes);
let mut truncated = [0u8; 8];
truncated.copy_from_slice(&digest.as_bytes()[..8]);
NodeId(u64::from_le_bytes(truncated))
}
pub(crate) fn restore_node_from_history(&mut self, target: RegistryTarget, node_id: NodeId) -> Result<(), CrdtError> {
let delta = self
.find_in_ancestry(|d| matches!(d.reverse, RegistryDelta::AddNode { id, .. } if id == node_id))
.ok_or(CrdtError::NodeNotInHistory(node_id))?;
self.revert_delta(target, delta)
}
pub(crate) fn restore_network_from_history(&mut self, target: RegistryTarget, network_id: NetworkId) -> Result<(), CrdtError> {
// Find the Delta whose forward op removed this network. Its `reverse` is `AddNetwork`,
// which is what we want to re-apply.
let delta = self
.find_in_ancestry(|d| matches!(d.reverse, RegistryDelta::AddNetwork { id, .. } if id == network_id))
.ok_or(CrdtError::NetworkNotInHistory(network_id))?;
self.revert_delta(target, delta)
}
/// Search every delta reachable from `head` (following all parents, including a merge's
/// `extra_parents`) for the first matching `predicate`, breadth-first. Resurrection needs full
/// ancestry reachability, so a node added only on a merged-in branch is still found.
fn find_in_ancestry(&self, predicate: impl Fn(&Delta) -> bool) -> Option<Delta> {
let mut queue: std::collections::VecDeque<Rev> = self.head.into_iter().collect();
let mut seen: std::collections::HashSet<Rev> = self.head.into_iter().collect();
while let Some(rev) = queue.pop_front() {
let Some(delta) = self.history.get(rev) else { continue };
if predicate(delta) {
return Some(delta.clone());
}
for parent in delta.all_parents() {
if seen.insert(parent) {
queue.push_back(parent);
}
}
}
None
}
/// Apply a delta's `reverse` as the new forward op (silent-zone undo). Force-applied: structural
/// ops are idempotent, and LWW arms assign the reverse value unconditionally even though it carries
/// the same timestamp as the forward op it undoes.
pub(crate) fn revert_delta(&mut self, target: RegistryTarget, mut delta: Delta) -> Result<(), CrdtError> {
for parent in delta.all_parents() {
if !self.history.contains(parent) {
return Err(CrdtError::NotFoundInHistory(parent));
}
}
std::mem::swap(&mut delta.kind, &mut delta.reverse);
self.apply_op_with(target, delta.kind, delta.timestamp, ApplyMode::Force)
}
/// Apply a live broadcast op. Updates the registry via LWW and appends to the hot log.
/// Doesn't touch history or `head` — hot ops are transient.
pub fn apply_hot_op(&mut self, hot_op: HotOp) -> Result<(), CrdtError> {
self.apply_op(hot_op.op.clone(), hot_op.timestamp)?;
self.hot_log.push(hot_op);
Ok(())
}
/// Replay a hot op recovered from persisted state. Idempotent on structural ops so that
/// re-applying an op whose effect is already reflected in the registry is a no-op rather
/// than an error.
pub fn replay_hot_op(&mut self, hot_op: HotOp) -> Result<(), CrdtError> {
self.apply_op_idempotent(hot_op.op.clone(), hot_op.timestamp)?;
self.hot_log.push(hot_op);
Ok(())
}
/// Apply a retired commit. Idempotent on structural ops (AddNode/AddNetwork on existing
/// targets, Remove on missing ones) since hot ops already produced the structural state.
/// The point is to bump field timestamps to T_retire via the LWW arms.
pub fn apply_delta(&mut self, delta: Delta) -> Result<(), CrdtError> {
for parent in delta.all_parents() {
if !self.history.contains(parent) {
return Err(CrdtError::NotFoundInHistory(parent));
}
}
self.apply_op_idempotent(delta.kind.clone(), delta.timestamp)?;
self.history.push(delta);
Ok(())
}
/// The registry an apply reads and writes, resolved from the explicit [`RegistryTarget`].
fn registry_mut(&mut self, target: RegistryTarget) -> &mut Registry {
match target {
RegistryTarget::Working => &mut self.working_registry,
RegistryTarget::Snapshot => &mut self.retired_snapshot,
}
}
fn registry_ref(&self, target: RegistryTarget) -> &Registry {
match target {
RegistryTarget::Working => &self.working_registry,
RegistryTarget::Snapshot => &self.retired_snapshot,
}
}
/// New local/remote op against the working registry: add ops error on duplicate targets and
/// `Change*` ops error on a missing target, while remove ops no-op when the target is already
/// absent; LWW arms keep the newer-timestamp value (strict `>`). The common entry point for edits.
pub(crate) fn apply_op(&mut self, op: RegistryDelta, timestamp: TimeStamp) -> Result<(), CrdtError> {
self.apply_op_with(RegistryTarget::Working, op, timestamp, ApplyMode::Live)
}
/// Replay/retire against the working registry: structural ops skip duplicate/missing targets (the
/// state is already present from hot ops or a prior snapshot); LWW arms still gate on strict `>`.
pub(crate) fn apply_op_idempotent(&mut self, op: RegistryDelta, timestamp: TimeStamp) -> Result<(), CrdtError> {
self.apply_op_with(RegistryTarget::Working, op, timestamp, ApplyMode::Idempotent)
}
/// Silent-zone undo/redo rewind against the working registry: structural ops are idempotent, and
/// LWW arms assign unconditionally. We own the single-writer chain here, so the precomputed reverse
/// (undo) or forward (redo) value is authoritative even though its timestamp ties what it replaces.
pub(crate) fn force_apply_op(&mut self, op: RegistryDelta, timestamp: TimeStamp) -> Result<(), CrdtError> {
self.apply_op_with(RegistryTarget::Working, op, timestamp, ApplyMode::Force)
}
pub(crate) fn apply_op_with(&mut self, target: RegistryTarget, op: RegistryDelta, timestamp: TimeStamp, mode: ApplyMode) -> Result<(), CrdtError> {
// Advance the local clock past every observed op, including ones that subsequently no-op or
// error. Observation is about causality knowledge, not about whether the op took effect.
self.clock.observe(timestamp);
// Structural ops skip (rather than error) on duplicate/missing targets when not a fresh edit;
// LWW arms assign unconditionally only under `Force`.
let idempotent = mode != ApplyMode::Live;
let force = mode == ApplyMode::Force;
// Resurrect any concurrently-removed targets the op references before binding the registry
// (resurrection re-borrows `self` via history), so the mutation below holds one `registry` ref.
self.ensure_referenced_exist(target, &op)?;
let registry = self.registry_mut(target);
match op {
RegistryDelta::AddNode { id, node } => {
if registry.node_instances.contains_key(&id) {
if idempotent {
// Hot ops already created this node; skip rather than error.
return Ok(());
}
return Err(CrdtError::NodeAlreadyExists(id));
}
registry.node_instances.insert(id, node);
}
RegistryDelta::RemoveNode { id, .. } => {
registry.node_instances.remove(&id);
}
RegistryDelta::ChangeNodeInput { id, index, new_input } => {
let node = registry.node_instances.get_mut(&id).ok_or(CrdtError::TargetNodeDoesNotExist(id))?;
let input = node.inputs.get_mut(index as usize).ok_or(CrdtError::InputIndexOutOfBounds(index as usize))?;
if force || timestamp > input.timestamp {
input.input = new_input;
input.timestamp = timestamp;
}
}
RegistryDelta::ChangeNodeAttribute { id, delta } => {
let node = registry.node_instances.get_mut(&id).ok_or(CrdtError::TargetNodeDoesNotExist(id))?;
apply_attribute_delta(delta, timestamp, force, &mut node.attributes);
}
RegistryDelta::ChangeNodeInputAttribute { id, index, delta } => {
let node = registry.node_instances.get_mut(&id).ok_or(CrdtError::TargetNodeDoesNotExist(id))?;
let input = node.inputs.get_mut(index as usize).ok_or(CrdtError::InputIndexOutOfBounds(index as usize))?;
apply_attribute_delta(delta, timestamp, force, &mut input.attributes);
}
RegistryDelta::SetNetworkExport { id, index, export } => {
let net = registry.networks.get_mut(&id).ok_or(CrdtError::NetworkDoesNotExist(id))?;
let slot_idx = index as usize;
if slot_idx >= net.exports.len() {
if slot_idx >= MAX_EXPORT_SLOTS {
return Err(CrdtError::ExportSlotOutOfBounds(index));
}
net.exports.resize(
slot_idx + 1,
ExportSlot {
target: None,
timestamp: TimeStamp::ORIGIN,
},
);
}
let existing = &mut net.exports[slot_idx];
if force || timestamp > existing.timestamp {
existing.target = export;
existing.timestamp = timestamp;
}
}
RegistryDelta::AddNetwork { id, network: contents } => {
if registry.networks.contains_key(&id) {
if idempotent {
return Ok(());
}
return Err(CrdtError::NetworkAlreadyExists(id));
}
registry.networks.insert(id, contents);
}
RegistryDelta::RemoveNetwork { id, .. } => {
registry.networks.remove(&id);
}
RegistryDelta::ChangeNetworkAttribute { id, delta } => {
let net = registry.networks.get_mut(&id).ok_or(CrdtError::NetworkDoesNotExist(id))?;
apply_attribute_delta(delta, timestamp, force, &mut net.attributes);
}
RegistryDelta::SetResourceHash { id, hash } => {
let entry = registry.resources.entry(id).or_default();
if force || timestamp > entry.hash_timestamp {
entry.hash = hash;
entry.hash_timestamp = timestamp;
}
}
RegistryDelta::AddSource { id, key, source } => {
let entry = registry.resources.entry(id).or_default();
let value = SourceValue { source, timestamp };
if force { entry.force_set_source(key, value) } else { entry.set_source(key, value) }
}
RegistryDelta::RemoveSource { id, key } => {
if let Some(entry) = registry.resources.get_mut(&id) {
if force {
entry.force_remove_source(&key);
} else {
entry.remove_source(&key, timestamp);
}
}
}
RegistryDelta::AddResource { id, entry } => {
registry.resources.insert(id, entry);
}
RegistryDelta::RemoveResource { id, .. } => {
registry.resources.remove(&id);
}
RegistryDelta::RegisterPeer { peer, user } => match registry.peer_users.get(&peer) {
Some(existing) if *existing != user => return Err(CrdtError::PeerRegistrationConflict(peer)),
Some(_) => {}
None => {
registry.peer_users.insert(peer, user);
}
},
RegistryDelta::ChangeDocumentAttribute { delta } => {
apply_attribute_delta(delta, timestamp, force, &mut registry.attributes);
}
// Merge is a structural sync point only; it mutates no registry state.
RegistryDelta::Merge { .. } | RegistryDelta::Other(_) => {}
}
Ok(())
}
/// Resurrect (from history) any nodes/networks an op references that were concurrently removed, so
/// the op applies against a consistent registry. Cascading: a node's owning network is restored
/// before the node. No-op for ops that reference nothing absent.
fn ensure_referenced_exist(&mut self, target: RegistryTarget, op: &RegistryDelta) -> Result<(), CrdtError> {
match op {
RegistryDelta::AddNode { node, .. } => self.ensure_network_exists(target, node.network())?,
RegistryDelta::ChangeNodeInput { id, new_input, .. } => {
if let NodeInput::Node { id: referenced, .. } = new_input {
self.ensure_node_exists(target, *referenced)?;
}
self.ensure_node_exists(target, *id)?;
}
RegistryDelta::ChangeNodeAttribute { id, .. } | RegistryDelta::ChangeNodeInputAttribute { id, .. } => self.ensure_node_exists(target, *id)?,
RegistryDelta::SetNetworkExport {
id: network, export: export_target, ..
} => {
if let Some(NodeInput::Node { id: referenced, .. }) = export_target {
self.ensure_node_exists(target, *referenced)?;
}
self.ensure_network_exists(target, *network)?;
}
RegistryDelta::ChangeNetworkAttribute { id: network, .. } => self.ensure_network_exists(target, *network)?,
_ => {}
}
Ok(())
}
fn ensure_node_exists(&mut self, target: RegistryTarget, node_id: NodeId) -> Result<(), CrdtError> {
if !self.registry_ref(target).node_instances.contains_key(&node_id) {
self.restore_node_from_history(target, node_id)?;
}
Ok(())
}
fn ensure_network_exists(&mut self, target: RegistryTarget, network_id: NetworkId) -> Result<(), CrdtError> {
if !self.registry_ref(target).networks.contains_key(&network_id) {
self.restore_network_from_history(target, network_id)?;
}
Ok(())
}
/// Compute the inverse of `delta` against the registry named by `target`. Retirement passes
/// [`RegistryTarget::Snapshot`] so LWW reverses (export target, inputs, attributes, resource hash)
/// capture the true pre-op value rather than the hot-polluted working state.
pub(crate) fn compute_reverse_delta(&self, target: RegistryTarget, delta: &RegistryDelta) -> Result<RegistryDelta, CrdtError> {
let registry = self.registry_ref(target);
Ok(match delta {
RegistryDelta::AddNode { id, node } => RegistryDelta::RemoveNode { id: *id, snapshot: node.clone() },
RegistryDelta::RemoveNode { id, snapshot } => RegistryDelta::AddNode { id: *id, node: snapshot.clone() },
&RegistryDelta::ChangeNodeInput { id, index: input_idx, .. } => {
let node = registry.node_instances.get(&id).ok_or(CrdtError::TargetNodeDoesNotExist(id))?;
let slot = node.inputs().get(input_idx as usize).ok_or(CrdtError::InputIndexOutOfBounds(input_idx as usize))?;
RegistryDelta::ChangeNodeInput {
id,
index: input_idx,
new_input: slot.input.clone(),
}
}
&RegistryDelta::ChangeNodeAttribute { id, ref delta } => {
let node = registry.node_instances.get(&id).ok_or(CrdtError::TargetNodeDoesNotExist(id))?;
RegistryDelta::ChangeNodeAttribute {
id,
delta: reverse_attribute_delta(delta, node.attributes()),
}
}
&RegistryDelta::ChangeNodeInputAttribute { id, index, ref delta } => {
let node = registry.node_instances.get(&id).ok_or(CrdtError::TargetNodeDoesNotExist(id))?;
let input = node.inputs().get(index as usize).ok_or(CrdtError::InputIndexOutOfBounds(index as usize))?;
RegistryDelta::ChangeNodeInputAttribute {
id,
index,
delta: reverse_attribute_delta(delta, &input.attributes),
}
}
&RegistryDelta::SetNetworkExport { id, index, .. } => {
// If the network is absent the forward op will resurrect it; the reverse is "set the export to None"
// since pre-forward there was no export to point at.
let export_target = registry.networks.get(&id).and_then(|net| net.exports.get(index as usize)).and_then(|s| s.target.clone());
RegistryDelta::SetNetworkExport { id, index, export: export_target }
}
RegistryDelta::AddNetwork { id, network } => RegistryDelta::RemoveNetwork { id: *id, snapshot: network.clone() },
&RegistryDelta::RemoveNetwork { id, ref snapshot } => RegistryDelta::AddNetwork { id, network: snapshot.clone() },
&RegistryDelta::ChangeNetworkAttribute { id, ref delta } => {
let current = registry.networks.get(&id).map(|net| &net.attributes).ok_or(CrdtError::NetworkDoesNotExist(id))?;
RegistryDelta::ChangeNetworkAttribute {
id,
delta: reverse_attribute_delta(delta, current),
}
}
RegistryDelta::ChangeDocumentAttribute { delta } => RegistryDelta::ChangeDocumentAttribute {
delta: reverse_attribute_delta(delta, &registry.attributes),
},
// Registrations are append-only and not user-undoable; reverse is the same op,
// which applies as a no-op on the already-registered PeerId.
&RegistryDelta::RegisterPeer { peer, user } => RegistryDelta::RegisterPeer { peer, user },
&RegistryDelta::SetResourceHash { id, .. } => RegistryDelta::SetResourceHash {
id,
hash: registry.resources.get(&id).and_then(|entry| entry.hash),
},
&RegistryDelta::AddSource { id, key, .. } => match registry.resources.get(&id).and_then(|entry| entry.source(&key)) {
// The slot already held a source: undo restores it.
Some(existing) => RegistryDelta::AddSource {
id,
key,
source: existing.source.clone(),
},
// The slot was empty: undo removes what this op added.
None => RegistryDelta::RemoveSource { id, key },
},
&RegistryDelta::RemoveSource { id, key } => match registry.resources.get(&id).and_then(|entry| entry.source(&key)) {
Some(existing) => RegistryDelta::AddSource {
id,
key,
source: existing.source.clone(),
},
// Nothing to restore; reverse is a no-op removal.
None => RegistryDelta::RemoveSource { id, key },
},
&RegistryDelta::AddResource { id, .. } => match registry.resources.get(&id) {
// Overwrote an existing entry: undo restores it.
Some(existing) => RegistryDelta::AddResource { id, entry: existing.clone() },
// Created a new entry: undo removes what this op added (snapshot is empty since there was nothing prior).
None => RegistryDelta::RemoveResource {
id,
snapshot: ResourceEntry::default(),
},
},
&RegistryDelta::RemoveResource { id, .. } => {
let snapshot = registry.resources.get(&id).cloned().unwrap_or_default();
RegistryDelta::AddResource { id, entry: snapshot }
}
RegistryDelta::Merge { extra_parents } => RegistryDelta::Merge { extra_parents: extra_parents.clone() },
&RegistryDelta::Other(_) => RegistryDelta::Other(serde_json::Value::Null),
})
}
}
/// Which of a [`Document`]'s two registries an apply targets: the working copy (retired state plus
/// live hot ops) or the retired snapshot (retired deltas only). Retirement targets the snapshot so
/// reverses capture pre-op values; the hot path and undo/redo target the working copy.
#[derive(Clone, Copy, PartialEq, Eq)]
pub(crate) enum RegistryTarget {
Working,
Snapshot,
}
/// How [`Document::apply_op_with`] resolves structural collisions and LWW timestamp ties.
#[derive(Clone, Copy, PartialEq, Eq)]
pub(crate) enum ApplyMode {
/// Fresh local/remote edit: structural ops error on duplicate/missing targets; LWW uses strict `>`.
Live,
/// Replay/retire: structural ops skip duplicate/missing targets; LWW still uses strict `>`.
Idempotent,
/// Silent-zone undo/redo rewind: structural ops are idempotent and LWW arms assign unconditionally.
Force,
}
@@ -0,0 +1,569 @@
use std::collections::HashMap;
use core_types::Context;
use core_types::uuid::NodeId as RuntimeNodeId;
use graph_craft::concrete;
use graph_craft::document::value::TaggedValue;
use graph_craft::document::{DocumentNode, DocumentNodeImplementation, NodeInput as GraphCraftNodeInput, NodeNetwork};
use serde::Serialize;
use crate::attr::*;
use crate::metadata_source::{NoMetadata, NodeMetadataSource};
use crate::{AttributesWrite, ExportSlot, Implementation, InputSlot, Network, NetworkId, Node, NodeId, NodeInput, PeerId, ProtoNode, ROOT_NETWORK, Registry, ResourceHash, ResourceId, TimeStamp};
fn map_serialization_error(key: &str) -> impl FnOnce(serde_json::Error) -> ConversionError + '_ {
move |e| ConversionError::SerializationError(format!("{key}: {e:?}"))
}
/// Path to a node, used to mint stable global IDs by hashing.
///
/// Hashing uses blake3 truncated to 64 bits with the document's `PeerId` mixed in, so two peers
/// converting runtime states that happen to share local IDs (e.g. both editors seeded the same
/// UUID RNG) still produce distinct global IDs. Determinism: same `(peer, path, local_id)` always
/// yields the same global ID, so a peer re-converting its own runtime state preserves IDs.
#[derive(Clone, Debug, PartialEq, Eq, Serialize)]
struct NodePath {
path: Vec<(RuntimeNodeId, NetworkId)>,
local_id: RuntimeNodeId,
}
impl NodePath {
fn root(node_id: RuntimeNodeId) -> Self {
Self { path: vec![], local_id: node_id }
}
fn nested(parent_path: &NodePath, parent_node_id: RuntimeNodeId, network_id: NetworkId, local_id: RuntimeNodeId) -> Self {
let mut path = parent_path.path.clone();
path.push((parent_node_id, network_id));
Self { path, local_id }
}
fn to_global_id(&self, peer: PeerId) -> NodeId {
let bytes = rmp_serde::to_vec(&(peer, self)).expect("NodePath must serialize");
let digest = blake3::hash(&bytes);
let mut truncated = [0u8; 8];
truncated.copy_from_slice(&digest.as_bytes()[..8]);
NodeId(u64::from_le_bytes(truncated))
}
/// Stable id of the network owned by the node at this path, derived purely from the (structural)
/// path and peer so it reproduces across `to_runtime` -> `from_runtime` round trips rather than
/// depending on traversal order. A domain tag keeps it from colliding with this node's own
/// `to_global_id`. The root network is `ROOT_NETWORK` and never goes through here.
fn owned_network_id(&self, peer: PeerId) -> NetworkId {
let bytes = rmp_serde::to_vec(&("network", peer, self)).expect("NodePath must serialize");
let digest = blake3::hash(&bytes);
let mut truncated = [0u8; 8];
truncated.copy_from_slice(&digest.as_bytes()[..8]);
NetworkId(u64::from_le_bytes(truncated))
}
}
#[derive(Debug, thiserror::Error)]
pub enum ConversionError {
#[error("Failed to serialize value: {0}")]
SerializationError(String),
#[error("Unsupported node implementation type")]
UnsupportedImplementation,
#[error("Invalid network structure: {0}")]
InvalidNetwork(String),
#[error("Index {0} exceeds the storage format's u32 range")]
IndexOverflow(usize),
}
/// Graph-only conversion (no editor metadata). Use [`Registry::from_runtime_with_metadata`] for
/// editor round-trips.
impl TryFrom<&NodeNetwork> for Registry {
type Error = ConversionError;
/// Test/utility entry point: scopes IDs under `PeerId(0)`. Real editor conversions go through
/// `from_runtime_with_metadata` and pass the document's actual peer.
fn try_from(node_network: &NodeNetwork) -> Result<Self, Self::Error> {
Registry::from_runtime_with_metadata(node_network, &NoMetadata, &graphene_resource::ResourceRegistry::new(), PeerId(0))
}
}
/// Proto-node declaration bytes extracted during conversion, keyed by content hash, for the caller
/// to persist into its byte store.
pub type DeclarationBytes = HashMap<ResourceHash, Vec<u8>>;
/// A `from_runtime` conversion result: the reference-only [`Registry`] plus the proto-node
/// declaration *bytes* it extracted, keyed by content hash. `document-graph-storage` doesn't own a byte
/// store, so the caller (the `Gdd`) persists these into its content store; the registry only holds
/// the `ResourceId`/`ResourceHash` references.
pub struct RuntimeConversion {
pub registry: Registry,
pub declaration_bytes: DeclarationBytes,
/// Each network's runtime `metadata_path` mapped to its stable storage `NetworkId`, for associating
/// per-network, per-peer view state (`session.json`) without re-deriving ids.
pub network_ids: HashMap<Vec<RuntimeNodeId>, NetworkId>,
}
impl RuntimeConversion {
/// Rebuild the [`Declarations`](crate::Declarations) map (`ResourceId` → [`ProtoNode`]) from the
/// extracted bytes, for callers that keep the bytes in hand instead of routing them through a
/// byte store (tests, the round-trip CLI). Editor/`Gdd` paths persist the bytes and resolve via
/// their byte store instead.
pub fn declarations(&self) -> Result<crate::Declarations, ConversionError> {
self.declaration_bytes
.iter()
.map(|(hash, bytes)| {
let proto = decode_declaration(bytes).map_err(|error| ConversionError::SerializationError(format!("declaration {hash}: {error}")))?;
Ok((ResourceId::from_hash(hash), proto))
})
.collect()
}
}
/// Encode a [`ProtoNode`] declaration to its content-addressed bytes: through a self-describing
/// `serde_json::Value` (so serde aliases keep working and the on-disk shape stays migratable), then
/// rmp-serialized (which encodes the intermediate `Value` compactly). Paired with [`decode_declaration`].
pub fn encode_declaration(proto: &ProtoNode) -> Result<Vec<u8>, String> {
let value = serde_json::to_value(proto).map_err(|error| error.to_string())?;
rmp_serde::to_vec(&value).map_err(|error| error.to_string())
}
/// Decode a [`ProtoNode`] declaration from the bytes [`encode_declaration`] produced.
pub fn decode_declaration(bytes: &[u8]) -> Result<ProtoNode, String> {
let value: serde_json::Value = rmp_serde::from_slice(bytes).map_err(|error| error.to_string())?;
serde_json::from_value(value).map_err(|error| error.to_string())
}
impl Registry {
/// Convenience wrapper returning only the registry (declaration bytes discarded). For callers
/// that don't persist a byte store — e.g. the graph-only `TryFrom` and value-comparison tests.
pub fn from_runtime_with_metadata<M: NodeMetadataSource>(node_network: &NodeNetwork, metadata: &M, resources: &graphene_resource::ResourceRegistry, peer: PeerId) -> Result<Self, ConversionError> {
Ok(Self::convert_from_runtime(node_network, metadata, resources, peer)?.registry)
}
/// Full conversion: returns the registry and the extracted declaration bytes for the caller to
/// persist. See [`RuntimeConversion`].
pub fn convert_from_runtime<M: NodeMetadataSource>(
node_network: &NodeNetwork,
metadata: &M,
resources: &graphene_resource::ResourceRegistry,
peer: PeerId,
) -> Result<RuntimeConversion, ConversionError> {
let mut registry = Registry::default();
let mut ctx = ConversionContext {
declaration_ids: HashMap::new(),
declaration_bytes: HashMap::new(),
network_ids: HashMap::new(),
metadata,
peer,
};
convert_network(node_network, ROOT_NETWORK, None, &[], &mut registry, &mut ctx)?;
// Only snapshot resources the network actually references. The runtime resource cache also keeps
// resources alive across undo (so legacy redo can restore them), so it can contain orphans whose
// node was removed by an undo. Snapshotting those would re-introduce an `AddResource` on the next
// diff and let an undone resource resurface as a phantom edit. Declaration resources are added
// separately by `convert_network` and are always referenced, so they're unaffected by this filter.
let referenced = collect_referenced_resources(node_network);
convert_resources(resources, &referenced, peer, &mut registry)?;
Ok(RuntimeConversion {
registry,
declaration_bytes: ctx.declaration_bytes,
network_ids: ctx.network_ids,
})
}
}
/// Snapshot the runtime [`ResourceRegistry`](graphene_resource::ResourceRegistry) into the storage
/// [`ResourceStore`](crate::ResourceStore). Each source's chain position becomes a fractional
/// [`Priority`](crate::Priority) (index-as-priority preserves order); the `DataSource` body is
/// stored type-erased as `serde_json::Value` so its on-disk shape can migrate freely. All
/// timestamps are `ORIGIN`, since this is a bootstrap snapshot, not an edit.
fn convert_resources(resources: &graphene_resource::ResourceRegistry, referenced: &std::collections::HashSet<ResourceId>, peer: PeerId, registry: &mut Registry) -> Result<(), ConversionError> {
for id in resources.ids() {
if !referenced.contains(&id) {
continue;
}
let Some(info) = resources.info(&id) else { continue };
let mut entry = crate::ResourceEntry {
hash: info.hash.copied(),
hash_timestamp: TimeStamp::ORIGIN,
..Default::default()
};
for (position, source) in info.sources.iter().enumerate() {
let key = crate::SourceKey {
priority: crate::Priority::new(position as f64).expect("enumerate index is finite"),
peer,
};
let body = serde_json::to_value(source).map_err(|error| ConversionError::SerializationError(error.to_string()))?;
entry.set_source(
key,
crate::SourceValue {
source: body,
timestamp: TimeStamp::ORIGIN,
},
);
}
registry.resources.insert(id, entry);
}
Ok(())
}
/// Collect the `ResourceId`s referenced by `TaggedValue::Resource` inputs anywhere in the network
/// (recursively through nested networks). These are the resources the document actually uses; the
/// runtime cache may hold more (history-retained orphans) that shouldn't be snapshotted into storage.
fn collect_referenced_resources(network: &NodeNetwork) -> std::collections::HashSet<ResourceId> {
let mut referenced = std::collections::HashSet::new();
collect_referenced_resources_inner(network, &mut referenced);
referenced
}
fn collect_referenced_resources_inner(network: &NodeNetwork, referenced: &mut std::collections::HashSet<ResourceId>) {
for export in &network.exports {
collect_input_resource(export, referenced);
}
for node in network.nodes.values() {
for input in &node.inputs {
collect_input_resource(input, referenced);
}
if let DocumentNodeImplementation::Network(nested) = &node.implementation {
collect_referenced_resources_inner(nested, referenced);
}
}
}
fn collect_input_resource(input: &GraphCraftNodeInput, referenced: &mut std::collections::HashSet<ResourceId>) {
if let GraphCraftNodeInput::Value { tagged_value, .. } = input
&& let TaggedValue::Resource(id) = &**tagged_value
{
referenced.insert(*id);
}
}
/// Register a proto-node declaration as a content-addressed resource: a single `DataSource::Embedded`
/// source resolved to `hash`. The bytes themselves are persisted by the caller's byte store.
fn register_declaration_resource(registry: &mut Registry, id: ResourceId, hash: ResourceHash, peer: PeerId) {
registry.resources.insert(id, crate::ResourceEntry::embedded(hash, peer, TimeStamp::ORIGIN));
}
struct ConversionContext<'m, M: NodeMetadataSource + ?Sized> {
/// Cache from proto-node identifier to its derived `ResourceId`, so repeated proto-nodes reuse
/// one id without re-serializing. (Identical content hashes to the same id anyway; this just
/// skips the work.)
declaration_ids: HashMap<String, ResourceId>,
/// Extracted declaration content keyed by hash, handed back for the caller's byte store.
declaration_bytes: DeclarationBytes,
/// Maps each network's runtime `metadata_path` to its stable storage `NetworkId`, so the caller can
/// associate per-network, per-peer view state (in `session.json`) with networks without re-deriving ids.
network_ids: HashMap<Vec<RuntimeNodeId>, NetworkId>,
metadata: &'m M,
peer: PeerId,
}
fn convert_network<M: NodeMetadataSource + ?Sized>(
node_network: &NodeNetwork,
network_id: NetworkId,
parent_path: Option<&NodePath>,
metadata_path: &[RuntimeNodeId],
registry: &mut Registry,
ctx: &mut ConversionContext<'_, M>,
) -> Result<(), ConversionError> {
for (runtime_node_id, doc_node) in &node_network.nodes {
let node_path = child_path(parent_path, network_id, *runtime_node_id);
let global_id = node_path.to_global_id(ctx.peer);
let location = NodeLocation {
network_id,
parent_path,
metadata_path,
runtime_node_id: *runtime_node_id,
};
let mut node = convert_node(doc_node, location, registry, ctx)?;
node.attributes.set(node::ORIGINAL_NODE_ID, serde_json::json!(runtime_node_id.0), TimeStamp::ORIGIN);
registry.node_instances.insert(global_id, node);
}
let exports = node_network
.exports
.iter()
.map(|export| {
Ok(ExportSlot {
target: Some(convert_input(export, parent_path, network_id, ctx.peer)?),
timestamp: TimeStamp::ORIGIN,
})
})
.collect::<Result<Vec<_>, ConversionError>>()?;
let mut attributes = crate::Attributes::new();
write_ui_network_attributes(&mut attributes, ctx.metadata, metadata_path, TimeStamp::ORIGIN)?;
write_scope_injections(&mut attributes, node_network, parent_path, network_id, ctx.peer, TimeStamp::ORIGIN)?;
registry.networks.insert(network_id, Network { exports, attributes });
ctx.network_ids.insert(metadata_path.to_vec(), network_id);
Ok(())
}
/// Serialize a network's `scope_injections` onto its attributes as one whole-map LWW blob, remapping
/// each runtime-local node reference to its stable storage global ID so the reference survives a
/// round trip even if runtime IDs are later reshuffled.
fn write_scope_injections(
attributes: &mut crate::Attributes,
node_network: &NodeNetwork,
parent_path: Option<&NodePath>,
network_id: NetworkId,
peer: PeerId,
timestamp: TimeStamp,
) -> Result<(), ConversionError> {
if node_network.scope_injections.is_empty() {
return Ok(());
}
let stored: HashMap<String, (NodeId, core_types::Type)> = node_network
.scope_injections
.iter()
.map(|(key, (runtime_id, ty))| {
let storage_id = child_path(parent_path, network_id, *runtime_id).to_global_id(peer);
(key.clone(), (storage_id, ty.clone()))
})
.collect();
attributes
.set_serialized(network::SCOPE_INJECTIONS, &stored, timestamp)
.map_err(map_serialization_error(network::SCOPE_INJECTIONS))
}
fn child_path(parent_path: Option<&NodePath>, network_id: NetworkId, local_id: RuntimeNodeId) -> NodePath {
match parent_path {
None => NodePath::root(local_id),
Some(parent) => NodePath::nested(parent, parent.local_id, network_id, local_id),
}
}
/// Where a node sits in both the storage tree (`network_id`, `parent_path`) and the runtime tree
/// (`metadata_path`, `runtime_node_id`). `metadata_path` is the chain of runtime IDs from the root
/// down to (but not including) this node.
struct NodeLocation<'a> {
network_id: NetworkId,
parent_path: Option<&'a NodePath>,
metadata_path: &'a [RuntimeNodeId],
runtime_node_id: RuntimeNodeId,
}
fn convert_node<M: NodeMetadataSource + ?Sized>(doc_node: &DocumentNode, location: NodeLocation<'_>, registry: &mut Registry, ctx: &mut ConversionContext<'_, M>) -> Result<Node, ConversionError> {
let NodeLocation {
network_id,
parent_path,
metadata_path,
runtime_node_id,
} = location;
let node_path = child_path(parent_path, network_id, runtime_node_id);
let timestamp = TimeStamp::ORIGIN;
let mut inputs = Vec::with_capacity(doc_node.inputs.len());
for (input_index, input) in doc_node.inputs.iter().enumerate() {
let mut input_attrs = convert_input_attributes(input)?;
write_ui_input_attributes(&mut input_attrs, ctx.metadata, metadata_path, runtime_node_id, input_index, timestamp)?;
inputs.push(InputSlot {
input: convert_input(input, parent_path, network_id, ctx.peer)?,
timestamp,
attributes: input_attrs,
});
}
// For nested networks, append this node onto the metadata path.
let mut extended_path = Vec::new();
let child_metadata_path = if matches!(doc_node.implementation, DocumentNodeImplementation::Network(_)) {
extended_path.extend_from_slice(metadata_path);
extended_path.push(runtime_node_id);
extended_path.as_slice()
} else {
metadata_path
};
let implementation = convert_implementation(&doc_node.implementation, &node_path, child_metadata_path, registry, ctx)?;
// Defaults match `DocumentNode::default()`; `to_runtime` rehydrates absent keys from the same defaults.
let mut attributes = crate::Attributes::new();
attributes
.set_if_not_default(node::CALL_ARGUMENT, &doc_node.call_argument, &concrete!(Context), timestamp)
.map_err(map_serialization_error(node::CALL_ARGUMENT))?;
attributes
.set_if_not_default(node::VISIBLE, &doc_node.visible, &true, timestamp)
.map_err(map_serialization_error(node::VISIBLE))?;
attributes
.set_if_not_default(node::SKIP_DEDUPLICATION, &doc_node.skip_deduplication, &false, timestamp)
.map_err(map_serialization_error(node::SKIP_DEDUPLICATION))?;
write_ui_attributes(&mut attributes, ctx.metadata, metadata_path, runtime_node_id, timestamp)?;
Ok(Node {
implementation,
inputs,
attributes,
network: network_id,
})
}
fn write_ui_attributes<M: NodeMetadataSource + ?Sized>(
attributes: &mut crate::Attributes,
metadata: &M,
metadata_path: &[RuntimeNodeId],
runtime_node_id: RuntimeNodeId,
timestamp: TimeStamp,
) -> Result<(), ConversionError> {
if let Some(position) = metadata.position(metadata_path, runtime_node_id) {
attributes
.set_serialized(node::ui::POSITION, &position, timestamp)
.map_err(map_serialization_error(node::ui::POSITION))?;
}
// Bool flags are only emitted when true; absence reads as false.
for (key, value) in [
(node::ui::IS_LAYER, metadata.is_layer(metadata_path, runtime_node_id)),
(node::ui::LOCKED, metadata.locked(metadata_path, runtime_node_id)),
(node::ui::PINNED, metadata.pinned(metadata_path, runtime_node_id)),
] {
if value {
attributes.set(key, serde_json::Value::Bool(true), timestamp);
}
}
if let Some(name) = metadata.display_name(metadata_path, runtime_node_id)
&& !name.is_empty()
{
attributes.set(node::ui::DISPLAY_NAME, serde_json::Value::String(name.to_string()), timestamp);
}
// One whole-vec attribute; per-slot LWW would be overkill for rename-on-output.
let output_names = metadata.output_names(metadata_path, runtime_node_id);
if !output_names.is_empty() {
attributes
.set_serialized(node::ui::OUTPUT_NAMES, &output_names, timestamp)
.map_err(map_serialization_error(node::ui::OUTPUT_NAMES))?;
}
Ok(())
}
fn write_ui_network_attributes<M: NodeMetadataSource + ?Sized>(attributes: &mut crate::Attributes, metadata: &M, network_path: &[RuntimeNodeId], timestamp: TimeStamp) -> Result<(), ConversionError> {
if let Some(reference) = metadata.reference(network_path) {
attributes.set(node::ui::REFERENCE, serde_json::Value::String(reference.to_string()), timestamp);
}
Ok(())
}
/// Empty strings (the runtime's "unset" sentinel) and absent values are both skipped.
/// `input_data` entries each get their own `ui::input_data::<sub_key>` attribute for per-key LWW.
fn write_ui_input_attributes<M: NodeMetadataSource + ?Sized>(
attributes: &mut crate::Attributes,
metadata: &M,
metadata_path: &[RuntimeNodeId],
runtime_node_id: RuntimeNodeId,
input_index: usize,
timestamp: TimeStamp,
) -> Result<(), ConversionError> {
let non_empty_string = |key: &'static str, value: Option<&str>, attributes: &mut crate::Attributes| {
if let Some(value) = value.filter(|s| !s.is_empty()) {
attributes.set(key, serde_json::Value::String(value.to_string()), timestamp);
}
};
non_empty_string(node::input::ui::NAME, metadata.input_name(metadata_path, runtime_node_id, input_index), attributes);
non_empty_string(node::input::ui::DESCRIPTION, metadata.input_description(metadata_path, runtime_node_id, input_index), attributes);
non_empty_string(node::input::ui::WIDGET_OVERRIDE, metadata.widget_override(metadata_path, runtime_node_id, input_index), attributes);
for (sub_key, value) in metadata.input_data(metadata_path, runtime_node_id, input_index) {
attributes.set(&format!("{prefix}{sub_key}", prefix = node::input::ui::DATA_PREFIX), value, timestamp);
}
Ok(())
}
fn convert_input(input: &GraphCraftNodeInput, parent_path: Option<&NodePath>, network_id: NetworkId, peer: PeerId) -> Result<NodeInput, ConversionError> {
Ok(match input {
GraphCraftNodeInput::Node { node_id, output_index } => NodeInput::Node {
id: child_path(parent_path, network_id, *node_id).to_global_id(peer),
index: (*output_index).try_into().map_err(|_| ConversionError::IndexOverflow(*output_index))?,
},
GraphCraftNodeInput::Value { tagged_value, exposed } => {
let value = serde_json::to_value(&**tagged_value).map_err(|e| ConversionError::SerializationError(format!("{e:?}")))?;
NodeInput::Value { value, exposed: *exposed }
}
GraphCraftNodeInput::Scope(s) => NodeInput::Scope(s.clone()),
GraphCraftNodeInput::Import { import_index, .. } => NodeInput::Import {
index: (*import_index).try_into().map_err(|_| ConversionError::IndexOverflow(*import_index))?,
},
GraphCraftNodeInput::Reflection(_) => NodeInput::Reflection,
// GPU-specific; not modeled in the Registry format.
GraphCraftNodeInput::Inline(_) => return Err(ConversionError::UnsupportedImplementation),
})
}
fn convert_input_attributes(input: &GraphCraftNodeInput) -> Result<crate::Attributes, ConversionError> {
let mut attributes = crate::Attributes::new();
let timestamp = TimeStamp::ORIGIN;
match input {
GraphCraftNodeInput::Import { import_type, .. } => {
attributes
.set_serialized(node::input::IMPORT_TYPE, import_type, timestamp)
.map_err(map_serialization_error(node::input::IMPORT_TYPE))?;
}
GraphCraftNodeInput::Reflection(metadata) => {
attributes
.set_serialized(node::REFLECTION_METADATA, metadata, timestamp)
.map_err(map_serialization_error(node::REFLECTION_METADATA))?;
}
_ => {}
}
Ok(attributes)
}
fn convert_implementation<M: NodeMetadataSource + ?Sized>(
implementation: &DocumentNodeImplementation,
current_node_path: &NodePath,
child_metadata_path: &[RuntimeNodeId],
registry: &mut Registry,
ctx: &mut ConversionContext<'_, M>,
) -> Result<Implementation, ConversionError> {
Ok(match implementation {
DocumentNodeImplementation::ProtoNode(identifier) => {
let identifier_str = identifier.as_str().to_string();
// Reuse a previously-converted proto-node's id; identical content hashes to the same id
// anyway, so this only skips re-serializing.
if let Some(id) = ctx.declaration_ids.get(&identifier_str) {
return Ok(Implementation::ProtoNode(*id));
}
let proto = ProtoNode {
identifier: identifier_str.clone(),
attributes: Default::default(),
};
// Content-address the declaration: serialize, hash, derive a deterministic id.
let bytes = encode_declaration(&proto).map_err(|error| ConversionError::SerializationError(format!("proto-node {identifier_str}: {error}")))?;
let hash = ResourceHash::from(bytes.as_slice());
let id = ResourceId::from_hash(&hash);
register_declaration_resource(registry, id, hash, ctx.peer);
ctx.declaration_bytes.insert(hash, bytes);
ctx.declaration_ids.insert(identifier_str, id);
Implementation::ProtoNode(id)
}
DocumentNodeImplementation::Network(nested_network) => {
// Stable, traversal-order-independent id derived from the owning node's path, so a
// `to_runtime` -> `from_runtime` round trip reproduces the same `NetworkId` (and thus the
// same node-path hashes underneath it).
let nested_network_id = current_node_path.owned_network_id(ctx.peer);
convert_network(nested_network, nested_network_id, Some(current_node_path), child_metadata_path, registry, ctx)?;
Implementation::Network(nested_network_id)
}
// TODO: Support Extract in the Registry format.
DocumentNodeImplementation::Extract => return Err(ConversionError::UnsupportedImplementation),
})
}
@@ -0,0 +1,180 @@
//! Retired delta history: the durable, append-only DAG of committed deltas.
//!
//! [`History`] owns the deltas in topological order (every parent precedes its children) plus an
//! index from [`Rev`] to position for O(1) lookup. The order is a valid replay order, so it is what
//! gets serialized to the on-disk history file and what [`crate::Session::replay_from_history`]
//! consumes. Retired commits have a single writer in every regime (solo editing, or leader-ordered
//! collaboration where the leader serializes retired commits), so appending preserves the order by
//! construction. The only operation that introduces out-of-order deltas is [`merge`](History::merge),
//! which re-sorts the combined set into the canonical order to restore the invariant.
use std::collections::HashMap;
use crate::{AttributesWrite, CrdtError, Delta, Rev, TimeStamp};
#[derive(Clone, Debug, Default)]
pub struct History {
/// Deltas in topological order. Mutated only via [`push`](Self::push).
deltas: Vec<Delta>,
/// `Rev` to its position in `deltas`. Kept in sync with `deltas` by every mutator.
index: HashMap<Rev, usize>,
}
impl History {
pub fn new() -> Self {
Self::default()
}
/// Build from deltas already in topological order (the on-disk load path), indexing them in place.
pub fn from_ordered(deltas: Vec<Delta>) -> Self {
let index = deltas.iter().enumerate().map(|(position, delta)| (delta.id, position)).collect();
Self { deltas, index }
}
pub fn get(&self, rev: Rev) -> Option<&Delta> {
self.index.get(&rev).map(|&position| &self.deltas[position])
}
pub fn contains(&self, rev: Rev) -> bool {
self.index.contains_key(&rev)
}
pub fn len(&self) -> usize {
self.deltas.len()
}
pub fn is_empty(&self) -> bool {
self.deltas.is_empty()
}
/// Append a delta after its parents, keeping `deltas` and `index` in sync. A duplicate `Rev`
/// (idempotent re-apply) overwrites the existing entry in place rather than appending, so the
/// order and index are unchanged.
pub fn push(&mut self, delta: Delta) {
if let Some(&position) = self.index.get(&delta.id) {
self.deltas[position] = delta;
return;
}
self.index.insert(delta.id, self.deltas.len());
self.deltas.push(delta);
}
/// Deltas in topological order (a valid replay order).
pub fn iter(&self) -> impl Iterator<Item = &Delta> + '_ {
self.deltas.iter()
}
/// Absorb `incoming` (dedup by `Rev`) and canonically re-sort the whole combined history.
///
/// The sort is deterministic (topological, ties broken by `Rev`), so two peers that absorb the same
/// delta set produce byte-identical history, not merely two different valid orderings. This is the
/// history-convergence mechanism: arrival order is erased. Callers update the registry separately
/// (LWW apply is commutative, so the registry converges regardless of order).
pub fn merge(&mut self, incoming: impl IntoIterator<Item = Delta>) {
for delta in incoming {
self.push(delta);
}
self.canonical_sort();
}
/// Re-order `deltas` into the canonical topological order and rebuild the index: parents precede
/// children, and among deltas whose parents are all emitted the lowest `Rev` goes first. O(V + E).
fn canonical_sort(&mut self) {
// Unsatisfied in-history parent count per delta, plus reverse edges to decrement as parents emit.
let mut pending_parents: HashMap<Rev, usize> = HashMap::with_capacity(self.deltas.len());
let mut children: HashMap<Rev, Vec<Rev>> = HashMap::new();
for delta in &self.deltas {
let in_history_parents = delta.all_parents().filter(|parent| self.index.contains_key(parent)).count();
pending_parents.insert(delta.id, in_history_parents);
for parent in delta.all_parents() {
if self.index.contains_key(&parent) {
children.entry(parent).or_default().push(delta.id);
}
}
}
// Ready set as a min-heap on `Rev` (via `Reverse`) so ties resolve deterministically.
let mut ready: std::collections::BinaryHeap<std::cmp::Reverse<Rev>> = pending_parents.iter().filter(|(_, count)| **count == 0).map(|(rev, _)| std::cmp::Reverse(*rev)).collect();
let mut order: Vec<Rev> = Vec::with_capacity(self.deltas.len());
while let Some(std::cmp::Reverse(rev)) = ready.pop() {
order.push(rev);
for child in children.get(&rev).into_iter().flatten() {
let count = pending_parents.get_mut(child).expect("child is in history");
*count -= 1;
if *count == 0 {
ready.push(std::cmp::Reverse(*child));
}
}
}
// `order` is a permutation of the existing revs, so reorder `deltas` to match and rebuild the index.
let mut by_rev: HashMap<Rev, Delta> = self.deltas.drain(..).map(|delta| (delta.id, delta)).collect();
self.index.clear();
for (position, rev) in order.iter().enumerate() {
if let Some(delta) = by_rev.remove(rev) {
self.index.insert(*rev, position);
self.deltas.push(delta);
}
}
}
/// The current tips: revs that no other delta lists as a parent (the divergent heads). A linear
/// history has exactly one tip; concurrent branches have several. Sorted ascending for determinism.
pub fn tips(&self) -> Vec<Rev> {
let referenced: std::collections::HashSet<Rev> = self.deltas.iter().flat_map(|delta| delta.all_parents()).collect();
let mut tips: Vec<Rev> = self.deltas.iter().map(|delta| delta.id).filter(|rev| !referenced.contains(rev)).collect();
tips.sort_unstable();
tips
}
/// Mark a retired delta as the end of a user interaction. Mutates only the delta's attributes
/// (excluded from its `Rev`), so the index stays valid. Returns whether the delta was found.
pub fn mark_interaction_end(&mut self, rev: Rev, timestamp: TimeStamp) -> bool {
match self.index.get(&rev) {
Some(&position) => {
self.deltas[position].mark_interaction_end(timestamp);
true
}
None => false,
}
}
/// Set a local annotation attribute (e.g. a commit message) on a retired delta in place. Excluded
/// from the delta's `Rev`, so identity and the index are unchanged. Returns whether the delta was found.
pub fn annotate(&mut self, rev: Rev, key: &str, value: serde_json::Value, timestamp: TimeStamp) -> bool {
match self.index.get(&rev) {
Some(&position) => {
self.deltas[position].attributes.set(key, value, timestamp);
true
}
None => false,
}
}
/// Test-only mutable access to the first stored delta, for corrupting it to exercise `verify`.
#[cfg(test)]
pub(crate) fn first_mut(&mut self) -> Option<&mut Delta> {
self.deltas.first_mut()
}
/// Verify the two stored invariants for history loaded from an untrusted source: every delta's
/// content-addressed `id` matches its recomputed hash, and the deltas are in topological order
/// (each delta's in-history parents precede it). Returns the first violation found.
pub fn verify(&self) -> Result<(), CrdtError> {
let mut seen: std::collections::HashSet<Rev> = std::collections::HashSet::with_capacity(self.deltas.len());
for delta in &self.deltas {
let expected = delta.recomputed_id();
if delta.id != expected {
return Err(CrdtError::RevMismatch { stored: delta.id, expected });
}
for parent in delta.all_parents() {
if self.index.contains_key(&parent) && !seen.contains(&parent) {
return Err(CrdtError::NotFoundInHistory(parent));
}
}
seen.insert(delta.id);
}
Ok(())
}
}
@@ -0,0 +1,133 @@
use crate::RegistryDelta;
use serde::{Deserialize, Serialize};
/// Stable, document-scoped identity for a node. Minted as a truncated `blake3(peer, counter)` so it
/// reproduces across `to_runtime` -> `from_runtime` round trips. Used purely as an opaque key.
#[derive(Copy, Clone, Debug, PartialEq, Eq, PartialOrd, Ord, Hash, Default, Serialize, Deserialize)]
#[serde(transparent)]
pub struct NodeId(pub u64);
/// Stable identity for a node network. `ROOT_NETWORK` for the renderable graph; nested networks get a
/// path-derived ID via `owned_network_id`. Used purely as an opaque key.
#[derive(Copy, Clone, Debug, PartialEq, Eq, PartialOrd, Ord, Hash, Default, Serialize, Deserialize)]
#[serde(transparent)]
pub struct NetworkId(pub u64);
impl std::fmt::Display for NodeId {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(f, "{}", self.0)
}
}
impl std::fmt::Display for NetworkId {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(f, "{}", self.0)
}
}
/// Content-addressed identity for a `Delta`.
/// 128-bit blake3 truncation: comfortable collision headroom for any plausible document lifetime
/// without being adversarial-grade. Same delta content always produces the same `Rev`. Non-zero so
/// `Option<Rev>` (a missing/root parent) is niche-optimized to the same size as a bare `Rev`.
#[derive(Copy, Clone, Debug, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize)]
#[serde(transparent)]
pub struct Rev(pub std::num::NonZeroU128);
impl Rev {
/// Wrap a raw value, or `None` if it is zero.
pub fn new(value: u128) -> Option<Self> {
std::num::NonZeroU128::new(value).map(Self)
}
pub fn get(self) -> u128 {
self.0.get()
}
}
impl std::fmt::Display for Rev {
fn fmt(&self, f: &mut std::fmt::Formatter<'_>) -> std::fmt::Result {
write!(f, "{}", self.0)
}
}
/// Root network ID. The renderable graph lives in `networks[&ROOT_NETWORK]`.
pub const ROOT_NETWORK: NetworkId = NetworkId(0);
/// Upper bound on a network's export slot count, guarding `SetExport` against a malicious or corrupted
/// slot index forcing an unbounded `exports` allocation.
pub(crate) const MAX_EXPORT_SLOTS: usize = 1 << 16;
/// Per-device identity. Stable per `(device, document)`. Used for CRDT tiebreaking and `NodeId`
/// scoping. Globally unique across all peers ever in a document.
#[derive(Copy, Clone, Debug, PartialEq, Eq, PartialOrd, Ord, Hash, Default, Serialize, Deserialize)]
#[serde(transparent)]
pub struct PeerId(pub u64);
/// Per-human identity. Stable across devices (one user, many devices). Used for identity display
/// and undo-chain walking. Derived from `PeerId` via `Registry.peer_users`.
#[derive(Copy, Clone, Debug, PartialEq, Eq, PartialOrd, Ord, Hash, Default, Serialize, Deserialize)]
#[serde(transparent)]
pub struct UserId(pub u64);
/// Lamport timestamp with a peer-ID tiebreak. Higher counter wins; ties broken by peer.
#[derive(Copy, Clone, Debug, PartialEq, Eq, PartialOrd, Ord, Hash, Default, Serialize, Deserialize)]
pub struct TimeStamp {
pub counter: u64,
pub peer: PeerId,
}
impl TimeStamp {
/// Pre-edit origin. Used by initial `from_runtime` conversion before any edits have happened.
pub const ORIGIN: Self = TimeStamp { counter: 0, peer: PeerId(0) };
}
#[derive(Copy, Clone, Debug, Serialize, Deserialize)]
pub struct LamportClock {
pub(crate) counter: u64,
peer: PeerId,
}
impl LamportClock {
pub fn new(peer: PeerId) -> Self {
Self { counter: 0, peer }
}
/// Mints a fresh local timestamp.
pub fn tick(&mut self) -> TimeStamp {
self.counter += 1;
TimeStamp {
counter: self.counter,
peer: self.peer,
}
}
/// Advances past an incoming op so future local ticks are causally later.
pub fn observe(&mut self, incoming: TimeStamp) {
self.counter = self.counter.max(incoming.counter);
}
}
/// Hash the identity-bearing fields of a `Delta` with blake3 and truncate to 128 bits.
///
/// A [`RegistryDelta::Merge`] is addressed by its sorted parent set alone (author and timestamp
/// excluded), so two peers merging the same tips mint the identical `Rev` and dedup. Every other
/// delta hashes `(parent, author, timestamp, kind)`.
pub(crate) fn compute_rev(parent: Option<Rev>, author: PeerId, timestamp: TimeStamp, delta_type: &RegistryDelta) -> Rev {
let mut hasher = blake3::Hasher::new();
let bytes = match delta_type {
RegistryDelta::Merge { extra_parents } => {
let mut parents: Vec<Rev> = parent.into_iter().chain(extra_parents.iter().copied()).collect();
parents.sort_unstable();
parents.dedup();
rmp_serde::to_vec(&("merge", parents)).expect("Merge identity fields must serialize")
}
_ => rmp_serde::to_vec(&(parent, author, timestamp, delta_type)).expect("Delta identity fields must serialize"),
};
hasher.update(&bytes);
let digest = hasher.finalize();
let mut truncated = [0u8; 16];
truncated.copy_from_slice(&digest.as_bytes()[..16]);
// A 128-bit blake3 truncation is zero with probability 2^-128 (never in practice); map it to 1 so
// the non-zero invariant is total rather than relying on a panic that can't realistically fire.
Rev::new(u128::from_le_bytes(truncated)).unwrap_or(Rev(std::num::NonZeroU128::MIN))
}
@@ -0,0 +1,42 @@
pub use graphene_resource::{ResourceHash, ResourceId};
pub mod attributes;
pub mod crdt;
pub mod delta;
pub mod document;
pub mod history;
pub mod ids;
pub mod model;
pub mod registry;
pub mod resources;
pub mod session;
#[cfg(any(feature = "conversion", test))]
pub mod from_runtime;
#[cfg(any(feature = "conversion", test))]
pub mod metadata_source;
#[cfg(any(feature = "conversion", test))]
pub mod to_runtime;
pub use attributes::*;
pub use crdt::*;
pub use document::*;
pub use history::History;
pub use ids::*;
pub use model::*;
pub use registry::*;
pub use resources::*;
pub use session::*;
#[cfg(any(feature = "conversion", test))]
pub use from_runtime::{RuntimeConversion, decode_declaration, encode_declaration};
#[cfg(any(feature = "conversion", test))]
pub use metadata_source::{InputMetadataEntry, NetworkMetadataEntry, NoMetadata, NodeMetadataEntry, NodeMetadataSource, Position};
#[cfg(any(feature = "conversion", test))]
pub use to_runtime::Declarations;
#[cfg(test)]
mod tests {
mod crdt;
mod round_trip;
}
@@ -0,0 +1,130 @@
//! Lets `from_runtime` read editor-side per-node metadata without depending on the editor crate.
//! The editor implements this on `NodeNetworkInterface`; tests pass [`NoMetadata`].
//!
//! `network_path` is the chain of runtime local `NodeId`s from the root down to (but not including)
//! the queried node, matching `NodeNetworkInterface::node_metadata(node_id, network_path)`.
use std::collections::HashMap;
use core_types::uuid::NodeId as RuntimeNodeId;
/// One node's editor-side metadata, produced by `Registry::to_runtime_with_metadata`.
#[derive(Clone, Debug, PartialEq)]
pub struct NodeMetadataEntry {
pub network_path: Vec<RuntimeNodeId>,
pub local_id: RuntimeNodeId,
pub position: Option<Position>,
pub is_layer: bool,
pub display_name: Option<String>,
pub locked: bool,
pub pinned: bool,
/// Always sized to match the runtime node's `inputs.len()`; absent slots use `Default`. The rebuild
/// returns an error if this length does not match the node's input count.
pub input_metadata: Vec<InputMetadataEntry>,
pub output_names: Vec<String>,
}
impl NodeMetadataEntry {
pub fn is_empty(&self) -> bool {
self.position.is_none()
&& !self.is_layer
&& self.display_name.is_none()
&& !self.locked
&& !self.pinned
&& self.output_names.is_empty()
&& self.input_metadata.iter().all(InputMetadataEntry::is_empty)
}
}
/// Per-network metadata (navigation, previewing). Separate from `NodeMetadataEntry` since these are
/// properties of a network, not of any node.
#[derive(Clone, Debug, Default, PartialEq)]
pub struct NetworkMetadataEntry {
/// Owning-node chain from the root to (and including) the node containing this network.
/// Empty = root network.
pub network_path: Vec<RuntimeNodeId>,
/// Stable storage id of this network. Lets the editor associate per-network, per-peer view state
/// (node-graph nav + previewing, in `session.json`) with a network across reparenting.
pub network_id: crate::NetworkId,
/// Matches the runtime's `NodeNetworkPersistentMetadata::reference` — definition lineage tag.
pub reference: Option<String>,
}
impl NetworkMetadataEntry {
pub fn is_empty(&self) -> bool {
self.reference.is_none()
}
}
/// Per-input editor metadata. Mirrors `InputPersistentMetadata` but wraps strings in `Option` so
/// unset (`""` on the runtime side) is distinguishable from an explicit empty string.
#[derive(Clone, Debug, Default, PartialEq)]
pub struct InputMetadataEntry {
pub input_name: Option<String>,
pub input_description: Option<String>,
pub widget_override: Option<String>,
/// Reassembled from `ui::input_data::<sub_key>` attributes.
pub input_data: HashMap<String, serde_json::Value>,
}
impl InputMetadataEntry {
pub fn is_empty(&self) -> bool {
self.input_name.is_none() && self.input_description.is_none() && self.widget_override.is_none() && self.input_data.is_empty()
}
}
/// Editor-side metadata source. Methods default to "no data" so implementors only override what
/// they carry. Returns are JSON-shaped where the underlying types live editor-side (PTZ, etc.).
pub trait NodeMetadataSource {
fn position(&self, _network_path: &[RuntimeNodeId], _local_id: RuntimeNodeId) -> Option<Position> {
None
}
fn is_layer(&self, _network_path: &[RuntimeNodeId], _local_id: RuntimeNodeId) -> bool {
false
}
fn display_name(&self, _network_path: &[RuntimeNodeId], _local_id: RuntimeNodeId) -> Option<&str> {
None
}
fn locked(&self, _network_path: &[RuntimeNodeId], _local_id: RuntimeNodeId) -> bool {
false
}
fn pinned(&self, _network_path: &[RuntimeNodeId], _local_id: RuntimeNodeId) -> bool {
false
}
/// Empty vec = no overrides. Stored as a single `ui::output_names` attribute (whole-vec LWW).
fn output_names(&self, _network_path: &[RuntimeNodeId], _local_id: RuntimeNodeId) -> Vec<String> {
Vec::new()
}
fn input_name(&self, _network_path: &[RuntimeNodeId], _local_id: RuntimeNodeId, _input_index: usize) -> Option<&str> {
None
}
fn input_description(&self, _network_path: &[RuntimeNodeId], _local_id: RuntimeNodeId, _input_index: usize) -> Option<&str> {
None
}
fn widget_override(&self, _network_path: &[RuntimeNodeId], _local_id: RuntimeNodeId, _input_index: usize) -> Option<&str> {
None
}
/// Returns owned to stay object-safe. Each entry is stored as `ui::input_data::<key>` for per-key LWW.
fn input_data(&self, _network_path: &[RuntimeNodeId], _local_id: RuntimeNodeId, _input_index: usize) -> HashMap<String, serde_json::Value> {
HashMap::new()
}
fn reference(&self, _network_path: &[RuntimeNodeId]) -> Option<&str> {
None
}
}
/// No-op metadata source. Use when there's nothing to attach (synthetic networks, CLI tools).
pub struct NoMetadata;
impl NodeMetadataSource for NoMetadata {}
/// Unified storage-side position. The valid variants depend on `attr::node::ui::IS_LAYER`:
/// layers use `Absolute` or `Stack`; non-layer nodes use `Absolute` or `Chain`.
#[derive(Copy, Clone, Debug, PartialEq, Eq, serde::Serialize, serde::Deserialize)]
pub enum Position {
Absolute([i32; 2]),
Chain,
Stack(u32),
}
@@ -0,0 +1,189 @@
use crate::{Attributes, NetworkId, NodeId, ResourceId, TimeStamp, attributes_value_equal};
use serde::{Deserialize, Serialize};
use std::borrow::Cow;
#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
pub struct Node {
pub(crate) implementation: Implementation,
pub(crate) inputs: Vec<InputSlot>,
pub(crate) attributes: Attributes,
pub(crate) network: NetworkId,
}
impl Node {
pub fn implementation(&self) -> &Implementation {
&self.implementation
}
pub fn inputs(&self) -> &[InputSlot] {
&self.inputs
}
pub fn attributes(&self) -> &Attributes {
&self.attributes
}
pub fn network(&self) -> NetworkId {
self.network
}
/// True if both nodes agree on every value-bearing field, ignoring slot/attribute timestamps.
pub fn value_equal(&self, other: &Self) -> bool {
if self.implementation != other.implementation || self.network != other.network {
return false;
}
if self.inputs.len() != other.inputs.len() {
return false;
}
if !self
.inputs
.iter()
.zip(&other.inputs)
.all(|(a, b)| a.input == b.input && attributes_value_equal(&a.attributes, &b.attributes))
{
return false;
}
attributes_value_equal(&self.attributes, &other.attributes)
}
#[cfg(test)]
pub(crate) fn dummy() -> Self {
Self {
implementation: Implementation::ProtoNode(ResourceId::new()),
inputs: vec![],
attributes: Attributes::new(),
network: crate::ROOT_NETWORK,
}
}
}
/// One positional input. The timestamp drives LWW on concurrent `ChangeNodeInput` ops targeting
/// the same `(node_id, input_idx)`.
#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
pub struct InputSlot {
pub input: NodeInput,
pub timestamp: TimeStamp,
pub attributes: Attributes,
}
#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
pub enum NodeInput {
Node {
id: NodeId,
index: u32,
},
Value {
value: serde_json::Value,
exposed: bool,
},
Scope(Cow<'static, str>),
Import {
index: u32,
},
/// Marker; the `DocumentNodeMetadata` lives in `inputs_attributes`.
Reflection,
Other,
}
#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
pub enum Implementation {
/// References a proto-node declaration resource (see [`ProtoNode`]); the binding to content lives
/// in `Registry.resources` like any other resource.
ProtoNode(ResourceId),
Network(NetworkId),
}
#[derive(Clone, Debug, Default, PartialEq, Serialize, Deserialize)]
pub struct Network {
pub exports: Vec<ExportSlot>,
/// Per-network `ui::*` state (navigation, previewing). Separate from `Node.attributes` so
/// view-state edits LWW independently.
pub attributes: Attributes,
}
impl Network {
/// True if both networks agree on every value-bearing field, ignoring slot/attribute timestamps.
pub fn value_equal(&self, other: &Self) -> bool {
// Compare slot targets index-by-index, treating out-of-range slots as `None`. A `SetExport(None)`
// truncation leaves a trailing empty slot (a tombstone in the CRDT state) that is value-equal to
// the slot being absent, so trailing `None`s must not count as drift. Mirrors `compute_deltas`
// (emits nothing for them) and `to_runtime` (drops them).
let max_len = self.exports.len().max(other.exports.len());
for slot_idx in 0..max_len {
let self_target = self.exports.get(slot_idx).and_then(|slot| slot.target.as_ref());
let other_target = other.exports.get(slot_idx).and_then(|slot| slot.target.as_ref());
if self_target != other_target {
return false;
}
}
attributes_value_equal(&self.attributes, &other.attributes)
}
}
/// One positional export slot. `target == None` marks an empty/removed slot. Timestamp drives LWW
/// on concurrent `SetExport` ops.
#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
pub struct ExportSlot {
pub target: Option<NodeInput>,
pub timestamp: TimeStamp,
}
/// Content of a proto-node declaration. Stored as a content-addressed resource (serialized bytes
/// keyed by `ResourceHash`, held by the `Gdd` byte store) and referenced from
/// `Implementation::ProtoNode(ResourceId)`. `document-graph-storage` itself only holds the reference.
#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
pub struct ProtoNode {
pub identifier: String,
pub attributes: Attributes,
}
#[cfg(test)]
mod tests {
use super::*;
use crate::TimeStamp;
fn target_slot(node_id: u64) -> ExportSlot {
ExportSlot {
target: Some(NodeInput::Node { id: NodeId(node_id), index: 0 }),
timestamp: TimeStamp::ORIGIN,
}
}
fn empty_slot() -> ExportSlot {
ExportSlot {
target: None,
timestamp: TimeStamp { counter: 5, peer: crate::PeerId(1) },
}
}
/// A `SetExport(None)` truncation leaves a trailing empty slot. Such a network is value-equal to
/// the same network without that slot, so the soak oracle doesn't false-report drift.
#[test]
fn trailing_empty_export_slot_is_value_equal() {
let compact = Network {
exports: vec![target_slot(1), target_slot(2)],
..Default::default()
};
let with_trailing_empty = Network {
exports: vec![target_slot(1), target_slot(2), empty_slot()],
..Default::default()
};
assert!(compact.value_equal(&with_trailing_empty));
assert!(with_trailing_empty.value_equal(&compact));
}
/// A `None` slot *between* live targets is a real value difference (a hole), not a trailing
/// tombstone, so it must still count as drift.
#[test]
fn interior_empty_export_slot_is_not_value_equal() {
let dense = Network {
exports: vec![target_slot(1), target_slot(2)],
..Default::default()
};
let with_hole = Network {
exports: vec![target_slot(1), empty_slot(), target_slot(2)],
..Default::default()
};
assert!(!dense.value_equal(&with_hole));
}
}
@@ -0,0 +1,153 @@
use crate::{Attributes, Network, NetworkId, Node, NodeId, PeerId, ResourceId, ResourceStore, SourceKey, TimeStamp, UserId};
use serde::{Deserialize, Serialize};
use std::collections::HashMap;
#[derive(Clone, Debug, Default, PartialEq, Serialize, Deserialize)]
pub struct Registry {
pub node_instances: HashMap<NodeId, Node>,
pub networks: HashMap<NetworkId, Network>,
/// Content-addressable resources (images, fonts, eventually proto-node declarations) referenced
/// by `ResourceId`. See [`ResourceStore`].
pub resources: ResourceStore,
/// Append-only mapping from per-device `PeerId` to per-human `UserId`.
/// Registered by each device's first contribution via `RegistryDelta::RegisterPeer`.
pub peer_users: HashMap<PeerId, UserId>,
pub attributes: Attributes,
}
impl Registry {
/// True if both registries agree on every value-bearing field, ignoring per-slot and
/// per-attribute timestamps. Mirrors `compute_deltas`'s value-only semantics, so unchanged
/// state at a stamped slot doesn't count as drift. `peer_users` is excluded: it isn't diffed by
/// `compute_deltas` (the mapping is injected on the commit path via `RegisterPeer`, never by a
/// fresh `from_runtime` conversion), so a committed registry and a fresh conversion legitimately
/// differ there without it counting as drift.
pub fn value_equal(&self, other: &Self) -> bool {
if !resources_value_equal(&self.resources, &other.resources) {
return false;
}
if !attributes_value_equal(&self.attributes, &other.attributes) {
return false;
}
if self.node_instances.len() != other.node_instances.len() {
return false;
}
for (id, node) in &self.node_instances {
let Some(other_node) = other.node_instances.get(id) else { return false };
if !node.value_equal(other_node) {
return false;
}
}
if self.networks.len() != other.networks.len() {
return false;
}
for (id, network) in &self.networks {
let Some(other_network) = other.networks.get(id) else { return false };
if !network.value_equal(other_network) {
return false;
}
}
true
}
/// True if the relative timestamp order on every shared timestamped slot agrees across
/// the two registries. Catches LWW-bookkeeping bugs that `value_equal` deliberately ignores.
///
/// For every pair of shared keys (a, b), checks that `self[a].cmp(self[b])` and
/// `other[a].cmp(other[b])` are compatible: `Equal` on either side is always compatible;
/// otherwise both sides must agree on direction. Equality on one side imposes no order, so
/// a registry with all-equal timestamps trivially passes against any other.
///
/// Slots present in only one registry are skipped. O(N²) in the number of shared timestamped
/// slots; intended for debug-only use.
pub fn order_consistent(&self, other: &Self) -> bool {
let self_stamps = collect_timestamps(self);
let other_stamps = collect_timestamps(other);
let shared: Vec<(TimestampKey, TimeStamp, TimeStamp)> = self_stamps.into_iter().filter_map(|(key, ts)| other_stamps.get(&key).map(|other_ts| (key, ts, *other_ts))).collect();
for i in 0..shared.len() {
for j in (i + 1)..shared.len() {
let self_order = shared[i].1.cmp(&shared[j].1);
let other_order = shared[i].2.cmp(&shared[j].2);
use std::cmp::Ordering::*;
let compatible = matches!((self_order, other_order), (Equal, _) | (_, Equal) | (Less, Less) | (Greater, Greater));
if !compatible {
return false;
}
}
}
true
}
}
pub(crate) fn attributes_value_equal(a: &Attributes, b: &Attributes) -> bool {
if a.len() != b.len() {
return false;
}
a.iter().all(|(key, value)| b.get(key).is_some_and(|other| value.value == other.value))
}
/// Value-level resource comparison: same resolved hashes and same source chains (keyed by
/// `SourceKey`, comparing source bodies), ignoring LWW timestamps. Mirrors `attributes_value_equal`.
pub(crate) fn resources_value_equal(a: &ResourceStore, b: &ResourceStore) -> bool {
if a.len() != b.len() {
return false;
}
a.iter().all(|(id, entry)| {
b.get(id).is_some_and(|other| {
entry.hash == other.hash
&& entry.sources.len() == other.sources.len()
&& entry.sources.iter().all(|(key, value)| other.source(key).is_some_and(|other_value| value.source == other_value.source))
})
})
}
/// Stable identity for any timestamped slot in a `Registry`. Used by `order_consistent`.
#[derive(Clone, Debug, PartialEq, Eq, Hash, PartialOrd, Ord)]
enum TimestampKey {
NodeInput(NodeId, usize),
NodeInputAttribute(NodeId, usize, String),
NodeAttribute(NodeId, String),
NetworkExport(NetworkId, usize),
NetworkAttribute(NetworkId, String),
DocumentAttribute(String),
ResourceHash(ResourceId),
ResourceSource(ResourceId, SourceKey),
}
fn collect_timestamps(registry: &Registry) -> HashMap<TimestampKey, TimeStamp> {
let mut out = HashMap::new();
for (node_id, node) in &registry.node_instances {
for (i, slot) in node.inputs.iter().enumerate() {
out.insert(TimestampKey::NodeInput(*node_id, i), slot.timestamp);
for (key, value) in &slot.attributes {
out.insert(TimestampKey::NodeInputAttribute(*node_id, i, key.clone()), value.timestamp);
}
}
for (key, value) in &node.attributes {
out.insert(TimestampKey::NodeAttribute(*node_id, key.clone()), value.timestamp);
}
}
for (network_id, network) in &registry.networks {
for (i, slot) in network.exports.iter().enumerate() {
out.insert(TimestampKey::NetworkExport(*network_id, i), slot.timestamp);
}
for (key, value) in &network.attributes {
out.insert(TimestampKey::NetworkAttribute(*network_id, key.clone()), value.timestamp);
}
}
for (key, value) in &registry.attributes {
out.insert(TimestampKey::DocumentAttribute(key.clone()), value.timestamp);
}
for (id, entry) in &registry.resources {
out.insert(TimestampKey::ResourceHash(*id), entry.hash_timestamp);
for (source_key, source_value) in &entry.sources {
out.insert(TimestampKey::ResourceSource(*id, *source_key), source_value.timestamp);
}
}
out
}
@@ -0,0 +1,297 @@
use crate::{PeerId, TimeStamp};
use graphene_resource::{ResourceHash, ResourceId};
use serde::{Deserialize, Serialize};
use std::collections::HashMap;
/// Ordering key for an entry in a resource's source chain: fractional `priority`, with `peer` as
/// the tiebreak so concurrent insertions at the same priority converge deterministically.
#[derive(Copy, Clone, Debug, PartialEq, Eq, PartialOrd, Ord, Hash, Serialize, Deserialize)]
pub struct SourceKey {
pub priority: Priority,
pub peer: PeerId,
}
/// One entry in a resource's source chain. The `source` body is type-erased (`serde_json::Value`)
/// so the on-disk `DataSource` shape can evolve through migrations without the storage layer
/// committing to a Rust enum; `timestamp` drives LWW on re-setting this same entry.
#[derive(Clone, Debug, PartialEq, Serialize, Deserialize)]
pub struct SourceValue {
pub source: serde_json::Value,
pub timestamp: TimeStamp,
}
/// A single content-addressable resource: an ordered, conflict-mergeable chain of fallback sources
/// plus the resolved content hash. The source chain is an add-wins ordered set (concurrent
/// additions all survive); the hash is last-writer-wins (concurrent resolves of the same logical
/// resource agree by construction, since the hash is content-derived).
#[derive(Clone, Debug, Default, PartialEq, Serialize)]
pub struct ResourceEntry {
/// Fallback chain kept sorted by `SourceKey`, so iteration yields highest-priority first.
pub sources: Vec<(SourceKey, SourceValue)>,
pub hash: Option<ResourceHash>,
pub hash_timestamp: TimeStamp,
}
impl<'de> Deserialize<'de> for ResourceEntry {
fn deserialize<D: serde::Deserializer<'de>>(deserializer: D) -> Result<Self, D::Error> {
// The `binary_search`-based accessors require `sources` sorted by `SourceKey` with unique keys.
// On-disk data (older writers, hand edits) can't be trusted to preserve either, so re-sort and
// collapse any duplicate keys, keeping the higher-timestamp value (LWW).
#[derive(Deserialize)]
struct Raw {
sources: Vec<(SourceKey, SourceValue)>,
hash: Option<ResourceHash>,
hash_timestamp: TimeStamp,
}
let Raw { mut sources, hash, hash_timestamp } = Raw::deserialize(deserializer)?;
sources.sort_by_key(|(a, _)| *a);
sources.dedup_by(|(later_key, later_value), (kept_key, kept_value)| {
// `dedup_by` keeps the first of each run; sorting is stable, so resolve duplicates by LWW.
if later_key != kept_key {
return false;
}
if later_value.timestamp > kept_value.timestamp {
*kept_value = later_value.clone();
}
true
});
Ok(Self { sources, hash, hash_timestamp })
}
}
impl ResourceEntry {
/// A resource backed by a single `DataSource::Embedded` fallback resolved to `hash`. Both the
/// source entry and the resolved hash carry `timestamp` so later LWW writes order against it.
/// The bytes themselves are persisted separately by the caller's byte store.
pub fn embedded(hash: ResourceHash, peer: PeerId, timestamp: TimeStamp) -> Self {
let embedded = serde_json::to_value(graphene_resource::DataSource::Embedded).expect("DataSource::Embedded serializes");
let priority = Priority::new(0.).expect("0. is finite");
let sources = vec![(SourceKey { priority, peer }, SourceValue { source: embedded, timestamp })];
Self {
sources,
hash: Some(hash),
hash_timestamp: timestamp,
}
}
/// The source body and timestamp stored under `key`, if any.
pub fn source(&self, key: &SourceKey) -> Option<&SourceValue> {
self.sources.binary_search_by(|(candidate, _)| candidate.cmp(key)).ok().map(|index| &self.sources[index].1)
}
/// Insert or LWW-overwrite the entry at `key`. A re-set at an existing key wins only if `value`'s
/// timestamp is strictly newer; a fresh key is inserted in sorted position.
pub fn set_source(&mut self, key: SourceKey, value: SourceValue) {
match self.sources.binary_search_by(|(candidate, _)| candidate.cmp(&key)) {
Ok(index) => {
if value.timestamp > self.sources[index].1.timestamp {
self.sources[index].1 = value;
}
}
Err(index) => self.sources.insert(index, (key, value)),
}
}
/// Like [`set_source`](Self::set_source) but assigns unconditionally (silent-zone rewind), where the
/// precomputed reverse/forward value is authoritative even if its timestamp ties what it replaces.
pub fn force_set_source(&mut self, key: SourceKey, value: SourceValue) {
match self.sources.binary_search_by(|(candidate, _)| candidate.cmp(&key)) {
Ok(index) => self.sources[index].1 = value,
Err(index) => self.sources.insert(index, (key, value)),
}
}
/// Remove the entry at `key` if its timestamp is strictly older than `timestamp` (LWW). Returns
/// whether anything was removed.
pub fn remove_source(&mut self, key: &SourceKey, timestamp: TimeStamp) -> bool {
match self.sources.binary_search_by(|(candidate, _)| candidate.cmp(key)) {
Ok(index) if timestamp > self.sources[index].1.timestamp => {
self.sources.remove(index);
true
}
_ => false,
}
}
/// Like [`remove_source`](Self::remove_source) but removes unconditionally (silent-zone rewind).
pub fn force_remove_source(&mut self, key: &SourceKey) -> bool {
match self.sources.binary_search_by(|(candidate, _)| candidate.cmp(key)) {
Ok(index) => {
self.sources.remove(index);
true
}
_ => false,
}
}
/// True if the chain already carries a `DataSource::Embedded` source. Decodes each source body into
/// `DataSource` so a shape change in the serialized form can't slip an embedded source past detection.
pub fn has_embedded_source(&self) -> bool {
self.sources.iter().any(|(_, value)| {
matches!(
serde_json::from_value::<graphene_resource::DataSource>(value.source.clone()),
Ok(graphene_resource::DataSource::Embedded)
)
})
}
/// A `SourceKey` ordered strictly ahead of every current source, so an inserted entry becomes the
/// highest-precedence fallback.
pub fn highest_precedence_key(&self, peer: PeerId) -> SourceKey {
let min_priority = self.sources.first().map(|(key, _)| key.priority.value()).unwrap_or(0.);
SourceKey {
priority: Priority::new(min_priority - 1.).expect("finite priority minus one is finite"),
peer,
}
}
}
/// All resources referenced by the document, keyed by stable per-document [`ResourceId`]. Replicates
/// through the normal CmRDT path; bytes live in content-addressed storage keyed by [`ResourceHash`].
pub type ResourceStore = HashMap<ResourceId, ResourceEntry>;
/// Fractional priority for ordering a resource's source chain. New sources are inserted by picking
/// a value strictly between two neighbors, so concurrent insertions elsewhere never collide; an
/// exact tie between two peers inserting at the same gap is broken by `PeerId` in [`SourceKey`].
/// `f64` precision is ample for the short fallback chains resources carry in practice.
#[derive(Copy, Clone, Debug, Serialize, Deserialize)]
#[serde(try_from = "f64")]
pub struct Priority(f64);
impl Priority {
/// Rejects non-finite input. The field is private and deserialization routes through here, so a
/// `Priority` is always finite, keeping its `Ord`/`Hash`/`Eq` agreement sound.
pub fn new(value: f64) -> Result<Self, NonFinitePriority> {
if value.is_finite() { Ok(Self(value)) } else { Err(NonFinitePriority(value)) }
}
pub fn value(self) -> f64 {
self.0
}
}
impl TryFrom<f64> for Priority {
type Error = NonFinitePriority;
fn try_from(value: f64) -> Result<Self, Self::Error> {
Self::new(value)
}
}
/// A [`Priority`] was constructed from a `NaN` or infinite value.
#[derive(Debug, thiserror::Error)]
#[error("priority must be finite, got {0}")]
pub struct NonFinitePriority(pub f64);
// `total_cmp` drives `Ord`, `Hash`, and `Eq` together so `Priority` is a sound `BTree`/`Hash` key:
// a derived `PartialEq` would disagree with this ordering on `-0.0` and `NaN`.
impl PartialEq for Priority {
fn eq(&self, other: &Self) -> bool {
self.cmp(other) == std::cmp::Ordering::Equal
}
}
impl Eq for Priority {}
impl Ord for Priority {
fn cmp(&self, other: &Self) -> std::cmp::Ordering {
self.0.total_cmp(&other.0)
}
}
impl PartialOrd for Priority {
fn partial_cmp(&self, other: &Self) -> Option<std::cmp::Ordering> {
Some(self.cmp(other))
}
}
impl std::hash::Hash for Priority {
fn hash<H: std::hash::Hasher>(&self, state: &mut H) {
self.0.to_bits().hash(state);
}
}
#[cfg(test)]
mod tests {
use super::*;
#[test]
fn priority_rejects_non_finite() {
assert!(Priority::new(f64::NAN).is_err());
assert!(Priority::new(f64::INFINITY).is_err());
assert!(Priority::new(-1.5).is_ok(), "negative finite priorities are valid");
}
/// Deserialization routes through `Priority::new`, so a non-finite value on disk is rejected rather
/// than silently producing an unsound map key. MessagePack (the storage format) can carry a
/// non-finite `f64`, unlike JSON, so this guards the real round-trip path.
#[test]
fn priority_deserialize_validates_finiteness() {
let finite = rmp_serde::to_vec(&3.5_f64).unwrap();
assert!(rmp_serde::from_slice::<Priority>(&finite).is_ok());
let non_finite = rmp_serde::to_vec(&f64::INFINITY).unwrap();
assert!(rmp_serde::from_slice::<Priority>(&non_finite).is_err(), "a non-finite priority on disk must be rejected");
}
/// `ResourceEntry`'s accessors rely on `sources` being sorted by `SourceKey`. Deserializing an
/// out-of-order chain (older writer, hand-edited file) must restore the invariant rather than leave
/// `binary_search` to silently misbehave.
#[test]
fn deserialize_sorts_sources() {
let source = |priority: f64| {
(
SourceKey {
priority: Priority::new(priority).expect("finite"),
peer: PeerId(1),
},
SourceValue {
source: serde_json::json!(priority),
timestamp: TimeStamp::ORIGIN,
},
)
};
// Serialize a deliberately unsorted chain through the raw shape, then deserialize as `ResourceEntry`.
let unsorted = serde_json::json!({
"sources": [source(2.), source(0.), source(1.)],
"hash": null,
"hash_timestamp": TimeStamp::ORIGIN,
});
let entry: ResourceEntry = serde_json::from_value(unsorted).expect("deserialize");
let priorities: Vec<f64> = entry.sources.iter().map(|(key, _)| key.priority.value()).collect();
assert_eq!(priorities, vec![0., 1., 2.], "sources must be sorted by SourceKey after deserialization");
}
/// Duplicate keys on disk violate the `binary_search` uniqueness invariant. Deserialization must
/// collapse them, keeping the higher-timestamp value (LWW).
#[test]
fn deserialize_dedups_sources_by_lww() {
let key = SourceKey {
priority: Priority::new(1.).expect("finite"),
peer: PeerId(1),
};
let entry = |counter: u64, body: &str| {
(
key,
SourceValue {
source: serde_json::json!(body),
timestamp: TimeStamp { counter, peer: PeerId(1) },
},
)
};
let with_duplicates = serde_json::json!({
"sources": [entry(5, "newer"), entry(1, "older")],
"hash": null,
"hash_timestamp": TimeStamp::ORIGIN,
});
let resource: ResourceEntry = serde_json::from_value(with_duplicates).expect("deserialize");
assert_eq!(resource.sources.len(), 1, "duplicate keys must collapse to one entry");
assert_eq!(resource.sources[0].1.source, serde_json::json!("newer"), "the higher-timestamp value must win");
}
}
@@ -0,0 +1,573 @@
#[cfg(any(feature = "conversion", test))]
use crate::NodeMetadataSource;
#[cfg(any(feature = "conversion", test))]
use crate::from_runtime;
use crate::{ApplyMode, Delta, Document, History, LamportClock, NetworkId, NodeId, PeerId, Registry, RegistryDelta, RegistryTarget, ResourceEntry, Rev, TimeStamp, UserId};
use graphene_resource::{ResourceHash, ResourceId};
use serde::{Deserialize, Serialize};
use std::collections::{HashMap, HashSet};
/// A live editing session over a `Document`. Owns the document plus runtime collaboration
/// state that isn't persisted (currently just peer heartbeat tracking).
#[derive(Clone, Debug)]
pub struct Session {
pub(crate) document: Document,
/// Each peer's `retirement_tip` as reported by their most recent heartbeat. Drives
/// leader-eligibility computation (lowest PeerId among peers whose tip matches the session max).
#[expect(dead_code, reason = "Populated once heartbeat/leader-election transport lands; held now so the field and constructors are in place.")]
remote_tips: HashMap<PeerId, Rev>,
}
impl Session {
/// Mints a fresh `PeerId` from the process-wide UUID generator and wraps an empty `Document`.
/// Two peers in the same process will collide (the generator is seeded once); use `with_peer`
/// in tests where determinism matters.
#[cfg(any(feature = "conversion", test))]
pub fn new() -> Self {
Self::with_peer(PeerId(core_types::uuid::generate_uuid()))
}
/// Construct a session bound to a specific `PeerId`. Used by tests; production code wants
/// `Session::new`.
pub fn with_peer(peer: PeerId) -> Self {
Self {
document: Document {
working_registry: Registry::default(),
retired_snapshot: Registry::default(),
history: History::new(),
hot_log: Vec::new(),
head: None,
redo_stack: Vec::new(),
clock: LamportClock::new(peer),
peer,
last_broadcast_rev: None,
next_node_counter: 0,
},
remote_tips: HashMap::new(),
}
}
pub fn peer(&self) -> PeerId {
self.document.peer
}
pub fn registry(&self) -> &Registry {
&self.document.working_registry
}
/// The registry after applying retired history only, without the unretired hot tail. Persisted as the
/// snapshot alongside `history` + hot log so a reopen restores the same retired-then-hot layering.
pub fn retired_registry(&self) -> &Registry {
&self.document.retired_snapshot
}
/// Diff the current registry against a fresh conversion of `network`, then commit each emitted
/// op as its own `Delta` on the local chain. One `clock.tick()` per op (strictly causal within
/// a commit). Returns the new `Rev`s in commit order (empty if nothing changed) plus the
/// proto-node declaration bytes the conversion extracted, keyed by content hash, for the caller
/// to persist into its byte store (`document-graph-storage` itself is byte-unaware).
///
/// Stages the diff as hot ops rather than retired deltas: each op is applied to the registry and
/// pushed onto the hot log. The caller persists the returned hot frames and then calls `retire`
/// to promote them into durable history.
#[cfg(any(feature = "conversion", test))]
pub fn stage_from_runtime<M: NodeMetadataSource>(
&mut self,
network: &graph_craft::document::NodeNetwork,
metadata: &M,
resources: &graphene_resource::ResourceRegistry,
) -> Result<(Vec<HotOp>, from_runtime::DeclarationBytes), CommitError> {
let conversion = Registry::convert_from_runtime(network, metadata, resources, self.document.peer)?;
let ops = crate::delta::compute_deltas(&self.document.working_registry, &conversion.registry);
let hot_ops = self.stage_ops(ops)?;
Ok((hot_ops, conversion.declaration_bytes))
}
/// Resolve each runtime `network_path` to its stable [`NetworkId`] for this document's peer, so the
/// caller can key per-network, per-peer view state (`session.json`) by a stable id. Derived from the
/// network structure alone; resources/declarations are irrelevant to the ids.
#[cfg(any(feature = "conversion", test))]
pub fn network_ids<M: NodeMetadataSource>(&self, network: &graph_craft::document::NodeNetwork, metadata: &M) -> Result<HashMap<Vec<core_types::uuid::NodeId>, NetworkId>, CommitError> {
let conversion = Registry::convert_from_runtime(network, metadata, &graphene_resource::ResourceRegistry::new(), self.document.peer)?;
Ok(conversion.network_ids)
}
/// Register a content-addressed resource as a single `DataSource::Embedded` source resolved to
/// `hash`, staged as one `AddResource` hot op. The caller owns `id` allocation, persists the
/// returned hot frame, retires, and persists the bytes into its byte store separately.
pub fn stage_embedded_resource(&mut self, id: ResourceId, hash: ResourceHash) -> Result<Vec<HotOp>, CrdtError> {
let entry = ResourceEntry::embedded(hash, self.document.peer, self.document.clock.tick());
self.stage_ops([RegistryDelta::AddResource { id, entry }])
}
/// Commit an `AddSource(Embedded)` retired delta for each given resource, making it the highest-
/// precedence fallback. Skips resources that already have an `Embedded` source or no longer exist.
/// Used on a throwaway session clone at export time so the exported registry and history agree;
/// callers must guarantee the bytes are available in the export's resource store.
pub fn embed_resource_sources(&mut self, ids: impl IntoIterator<Item = ResourceId>) -> Result<Vec<Rev>, CrdtError> {
let embedded = serde_json::to_value(graphene_resource::DataSource::Embedded).expect("DataSource::Embedded serializes");
let mut ops = Vec::new();
for id in ids {
let Some(entry) = self.document.working_registry.resources.get(&id) else { continue };
if entry.has_embedded_source() {
continue;
}
let key = entry.highest_precedence_key(self.document.peer);
ops.push(RegistryDelta::AddSource { id, key, source: embedded.clone() });
}
// These are retired deltas, so `commit_ops` advances the retired snapshot and history. The working
// registry sits at `retired_snapshot + hot tail`, so mirror each committed delta onto it with its own
// timestamp rather than cloning the snapshot over it, which would discard any unretired hot-zone edits.
let revs = self.commit_ops(ops, false)?;
for &rev in &revs {
let Some(delta) = self.document.history.get(rev) else { continue };
let (kind, timestamp) = (delta.kind.clone(), delta.timestamp);
self.document.apply_op_idempotent(kind, timestamp)?;
}
Ok(revs)
}
/// Apply each op as a hot op with a freshly-ticked timestamp, returning the staged frames in
/// order. Each tick is strictly later than the last, so the final frame carries the latest
/// timestamp, which is what the caller passes to `retire`.
///
/// The peer's first contribution is preceded by a `RegisterPeer` op, so the device's
/// `PeerId → UserId` mapping is established (and, under causal delivery, observed by other peers)
/// before any of its edits. A no-op batch doesn't register — registration rides a real edit.
fn stage_ops(&mut self, ops: impl IntoIterator<Item = RegistryDelta>) -> Result<Vec<HotOp>, CrdtError> {
let mut pending: Vec<RegistryDelta> = ops.into_iter().collect();
if pending.is_empty() {
return Ok(Vec::new());
}
if !self.document.working_registry.peer_users.contains_key(&self.document.peer) {
let user = UserId(self.document.peer.0);
pending.insert(0, RegistryDelta::RegisterPeer { peer: self.document.peer, user });
}
let mut staged = Vec::with_capacity(pending.len());
for op in pending {
let hot_op = HotOp {
op,
timestamp: self.document.clock.tick(),
};
self.document.apply_hot_op(hot_op.clone())?;
staged.push(hot_op);
}
Ok(staged)
}
/// Wrap each op as a `Delta`, apply it, and chain it onto the local history. One tick per op.
///
/// Operates on the *retired snapshot*: reverses are computed against and forward ops applied to it,
/// so each `reverse` captures the true pre-op value rather than the hot-polluted working state. The
/// working registry already reflects these ops (they were staged as hot ops before retirement, or
/// equal the snapshot when there are none), so it is left untouched.
///
/// `idempotent`: pass `true` when the snapshot already reflects the op (retirement of an already-
/// applied hot op) so duplicate structural inserts no-op rather than error.
fn commit_ops(&mut self, ops: impl IntoIterator<Item = RegistryDelta>, idempotent: bool) -> Result<Vec<Rev>, CrdtError> {
let target = RegistryTarget::Snapshot;
let ops = ops.into_iter();
let mut produced = Vec::with_capacity(ops.size_hint().0);
for op in ops {
// A new edit abandons any undone-forward branch: those revs stay in the DAG but are no
// longer reachable via redo. (Mirrors the legacy editor clearing its redo history on
// commit.) Done on the first real op so a no-op commit doesn't silently disable redo.
if produced.is_empty() {
self.document.redo_stack.clear();
}
let reverse = self.document.compute_reverse_delta(target, &op)?;
let timestamp = self.document.clock.tick();
let parent = self.document.head;
let author = self.document.peer;
let delta = Delta::new(parent, author, timestamp, op, reverse);
let rev = delta.id;
// `parent` is `None` for the root commit; otherwise it must already be in history.
if let Some(parent) = parent
&& !self.document.history.contains(parent)
{
return Err(CrdtError::NotFoundInHistory(parent));
}
let mode = if idempotent { ApplyMode::Idempotent } else { ApplyMode::Live };
self.document.apply_op_with(target, delta.kind.clone(), delta.timestamp, mode)?;
self.document.history.push(delta);
self.document.head = Some(rev);
produced.push(rev);
}
Ok(produced)
}
/// Wrap an already-materialized snapshot. Trusts `registry` to match `history`; advances the
/// clock past every observed timestamp but does not re-apply ops. `history` is taken in on-disk
/// (topological) order.
pub fn load(peer: PeerId, registry: Registry, history: Vec<Delta>, head: Option<Rev>, redo_stack: Vec<Rev>, next_node_counter: u64) -> Self {
let mut clock = LamportClock::new(peer);
for delta in &history {
clock.observe(delta.timestamp);
}
Self {
document: Document {
// The persisted snapshot is the retired state; hot ops (replayed by the caller after
// `load`) build the working registry on top, leaving `retired_snapshot` at retired.
retired_snapshot: registry.clone(),
working_registry: registry,
history: History::from_ordered(history),
hot_log: Vec::new(),
head,
redo_stack,
clock,
peer,
last_broadcast_rev: None,
next_node_counter,
},
remote_tips: HashMap::new(),
}
}
/// Rebuild the registry from scratch by applying every delta in causal order.
/// `deltas` must be in causal order (every parent before its children).
pub fn replay_from_history(peer: PeerId, deltas: impl IntoIterator<Item = Delta>, next_node_counter: u64) -> Result<Self, CrdtError> {
let mut session = Self::with_peer(peer);
session.document.next_node_counter = next_node_counter;
for delta in deltas {
let rev = delta.id;
session.document.apply_op_idempotent(delta.kind.clone(), delta.timestamp)?;
session.document.history.push(delta);
session.document.head = Some(rev);
}
// Pure retired-delta replay: no hot ops, so the working registry is fully retired.
session.document.retired_snapshot = session.document.working_registry.clone();
Ok(session)
}
/// Apply a hot op without going through the broadcast stream.
pub fn apply_hot_op(&mut self, hot_op: HotOp) -> Result<(), CrdtError> {
self.document.apply_hot_op(hot_op)
}
/// Replay a persisted hot op. Idempotent on structural ops, suitable for crash recovery
/// where the registry may already reflect the op's effect from a prior retired snapshot.
pub fn replay_hot_op(&mut self, hot_op: HotOp) -> Result<(), CrdtError> {
self.document.replay_hot_op(hot_op)
}
/// Integrate `incoming` retired deltas from another branch and emit a [`RegistryDelta::Merge`]
/// joining the resulting tips, returning the new merge `Rev` (or `None` if `incoming` adds nothing).
/// Applies each incoming op to the registry, then hands the set to [`History::merge`]. Incoming
/// deltas must arrive in causal order.
pub fn merge(&mut self, incoming: impl IntoIterator<Item = Delta>) -> Result<Option<Rev>, CrdtError> {
let mut absorbed: Vec<Delta> = Vec::new();
for delta in incoming {
if self.document.history.contains(delta.id) {
continue;
}
self.document.apply_op_idempotent(delta.kind.clone(), delta.timestamp)?;
absorbed.push(delta);
}
if absorbed.is_empty() {
return Ok(None);
}
self.document.history.merge(absorbed);
let tips = self.document.history.tips();
let timestamp = self.document.clock.tick();
let merge = Delta::merge(tips, self.document.peer, timestamp);
let merge_rev = merge.id;
// The merge's parents are the current tips, so it sorts last: `push` preserves the canonical
// order without re-sorting the whole history.
self.document.history.push(merge);
self.document.head = Some(merge_rev);
// Merge runs with an empty hot log; keep the retired snapshot in step with the working registry.
self.document.retired_snapshot = self.document.working_registry.clone();
Ok(Some(merge_rev))
}
/// Promote hot ops with timestamp `≤ up_to` into retired deltas, re-applied with fresh
/// retirement timestamps so LWW arms bump field timestamps to `T_retire`.
///
/// Today: one retired delta per hot op. Coarsening is a future step.
pub fn retire(&mut self, up_to: TimeStamp) -> Result<Vec<Rev>, CrdtError> {
let mut drained = Vec::new();
let mut remaining = Vec::with_capacity(self.document.hot_log.len());
for hot_op in self.document.hot_log.drain(..) {
if hot_op.timestamp <= up_to {
drained.push(hot_op);
} else {
remaining.push(hot_op);
}
}
self.document.hot_log = remaining;
self.commit_ops(drained.into_iter().map(|hot_op| hot_op.op), true)
}
/// Mark a retired delta as the end of a user interaction, so the undo cursor treats it as a checkpoint.
/// Called once per interaction by the editor-facing commit path (not by resource/internal commits).
pub fn mark_interaction_end(&mut self, rev: Rev) {
let timestamp = self.document.clock.tick();
self.document.history.mark_interaction_end(rev, timestamp);
}
/// Low-level: set a local annotation attribute (e.g. a commit message) on a retired delta in place.
/// Excluded from the delta's content-addressed `Rev`, so identity is unchanged. Returns whether the
/// delta was found. The `Gdd` layer re-persists the affected history frame after calling this.
pub fn annotate_delta(&mut self, rev: Rev, key: &str, value: serde_json::Value) -> bool {
let timestamp = self.document.clock.tick();
self.document.history.annotate(rev, key, value, timestamp)
}
/// Whether there is a retired commit at `head` that can be undone in the silent zone (a commit
/// after `last_broadcast_rev`). `head == 0` is the empty history; published commits aren't
/// silently undoable (that needs a forward reverse-delta op, deferred until transport lands).
///
/// The earliest interaction (the document's loaded/created base) is *not* undoable: undoing it would
/// rewind into the pre-base state, which legacy never offers (opening a document gives an empty undo
/// history). We detect "head is on the earliest interaction" by walking `head`'s interaction back along
/// first-parents and checking whether it bottoms out at the root with no earlier interaction boundary to
/// land on. If so, there is nothing before this interaction to undo to, so undo is disabled.
pub fn can_undo(&self) -> bool {
let Some(head) = self.document.head else { return false };
if self.document.last_broadcast_rev == Some(head) {
return false;
}
self.interaction_start_parent(head).is_some()
}
/// Walk the interaction containing `rev` back along first-parents to its first delta, returning the
/// rev the cursor would rest on after undoing this interaction, or `None` if that is the root (the
/// earliest interaction, which is not undoable). Mirrors the boundary condition in [`undo`](Self::undo):
/// stop when the parent is an `interaction_end` boundary or the root.
fn interaction_start_parent(&self, rev: Rev) -> Option<Rev> {
let mut current = rev;
loop {
let parent = self.document.history.get(current)?.parent?;
if self.document.history.get(parent).is_some_and(|d| d.is_interaction_end()) {
return Some(parent);
}
current = parent;
}
}
pub fn can_redo(&self) -> bool {
!self.document.redo_stack.is_empty()
}
/// Silent-zone undo of one *interaction*: revert deltas walking `head` back along first-parents until
/// it reaches the previous interaction boundary (a delta marked `interaction_end`) or the empty root. One
/// interaction spans several deltas (one `commit_from_runtime` batch), so undo reverts the whole run,
/// not a single delta — matching the legacy per-interaction undo granularity. The undone interaction's
/// `head` rev is pushed onto the redo stack. Reflog semantics: the DAG is never rewritten.
pub fn undo(&mut self) -> Result<Rev, CrdtError> {
if !self.can_undo() {
return Err(CrdtError::NothingToUndo);
}
let checkpoint = self.document.head.ok_or(CrdtError::NothingToUndo)?;
// Revert this interaction's last delta, then keep going back until `head` rests on the previous
// interaction's boundary (its `interaction_end` delta) or the root.
loop {
let rev = self.document.head.ok_or(CrdtError::NothingToUndo)?;
let delta = self.document.history.get(rev).ok_or(CrdtError::NotFoundInHistory(rev))?.clone();
let parent = delta.parent;
self.document.revert_delta(RegistryTarget::Working, delta)?;
self.document.head = parent;
match parent {
None => break,
Some(parent) if self.document.history.get(parent).is_some_and(|d| d.is_interaction_end()) => break,
Some(_) => {}
}
}
// Undo runs with an empty hot log, so keep the retired snapshot in lockstep with the rewound
// working registry (the next interaction's reverses are computed against it).
self.document.retired_snapshot = self.document.working_registry.clone();
self.document.redo_stack.push(checkpoint);
Ok(checkpoint)
}
/// Redo the most-recently-undone interaction: re-apply every delta from the current `head` forward to
/// (and including) the checkpoint rev, advancing `head` to it. Collects the forward span by walking
/// parents back from the checkpoint to `head` (the chain is linear in the silent solo zone).
pub fn redo(&mut self) -> Result<Rev, CrdtError> {
let checkpoint = self.document.redo_stack.pop().ok_or(CrdtError::NothingToRedo)?;
let mut forward = Vec::new();
let mut cursor = Some(checkpoint);
while cursor != self.document.head {
let Some(rev) = cursor else { break };
let delta = self.document.history.get(rev).ok_or(CrdtError::NotFoundInHistory(rev))?.clone();
cursor = delta.parent;
forward.push(delta);
}
// Force-apply so each forward value wins the LWW tie against the reverse that undo force-applied
// at the same timestamp. Symmetric with `revert_delta`.
for delta in forward.into_iter().rev() {
self.document.force_apply_op(delta.kind.clone(), delta.timestamp)?;
}
self.document.head = Some(checkpoint);
// Redo runs with an empty hot log; keep the retired snapshot in lockstep with the working registry.
self.document.retired_snapshot = self.document.working_registry.clone();
Ok(checkpoint)
}
/// Build a synthetic linear history whose replay reproduces `registry`. Each op gets a
/// freshly-ticked clock timestamp and chains to the previous op's `Rev`.
pub fn bootstrap_from_registry(peer: PeerId, registry: Registry) -> Result<Self, CrdtError> {
let ops = crate::delta::compute_deltas(&Registry::default(), &registry);
let mut session = Self::with_peer(peer);
session.commit_ops(ops, false)?;
// No hot ops on this path, so the working registry must mirror the freshly-built snapshot.
session.document.working_registry = session.document.retired_snapshot.clone();
Ok(session)
}
/// Retired deltas in append order, which is a valid replay order (parents before children).
pub fn history(&self) -> impl Iterator<Item = &Delta> + '_ {
self.document.history.iter()
}
/// The retired delta for `rev`, or `None` if it isn't in history. O(1) lookup, for callers that
/// already hold the revs they want (e.g. persisting a freshly-retired batch) and don't need a scan.
pub fn delta(&self, rev: Rev) -> Option<&Delta> {
self.document.history.get(rev)
}
/// Verify the retired history loaded from an untrusted source: content-addressed ids match their
/// recomputed hashes, and the deltas are topologically ordered. See [`History::verify`].
pub fn verify_history(&self) -> Result<(), CrdtError> {
self.document.history.verify()
}
/// Every resource hash referenced by the current registry *or* anywhere in history. Undo removes a
/// interaction's `AddResource` from the working registry, so a redoable (or re-undoable) interaction's
/// resources no longer appear in `registry().resources` even though redo still needs them. Resource GC
/// must keep this whole set alive, not just the current head's, or undo then redo loses declaration
/// bytes. Walks current resources plus each delta's `AddResource`/`RemoveResource` snapshot.
pub fn all_referenced_resource_hashes(&self) -> HashSet<ResourceHash> {
let mut hashes: HashSet<ResourceHash> = self.document.working_registry.resources.values().filter_map(|entry| entry.hash).collect();
for delta in self.document.history.iter() {
match &delta.kind {
RegistryDelta::AddResource { entry, .. } => hashes.extend(entry.hash),
RegistryDelta::RemoveResource { snapshot, .. } => hashes.extend(snapshot.hash),
_ => {}
}
}
hashes
}
pub fn hot_log(&self) -> &[HotOp] {
&self.document.hot_log
}
pub fn head_rev(&self) -> Option<Rev> {
self.document.head
}
/// The latest retired commit broadcast to at least one peer. Commits after it are silently
/// rewritable; commits at or before it are published. `None` until broadcast transport lands.
pub fn last_broadcast_rev(&self) -> Option<Rev> {
self.document.last_broadcast_rev
}
/// Advance the published frontier to `rev` as commits are broadcast. The frontier is monotonic, so
/// this only moves it forward (never back to `None`). Set by the (future) broadcast transport;
/// persisted in `session.json` so the silent/published boundary survives a reopen.
pub fn publish_up_to(&mut self, rev: Rev) {
self.document.last_broadcast_rev = Some(rev);
}
/// Test-only: every retired delta, cloned, for feeding one session's branch into another's `merge`.
#[cfg(test)]
pub(crate) fn cloned_deltas(&self) -> Vec<Delta> {
self.document.history.iter().cloned().collect()
}
/// Test-only: commit a single op as a retired delta on the local chain, returning the result so a
/// test can observe a resurrection failure (e.g. `NotFoundInHistory`).
#[cfg(test)]
pub(crate) fn commit_op_for_test(&mut self, op: RegistryDelta) -> Result<(), CrdtError> {
self.commit_ops(std::iter::once(op), false).map(|_| ())
}
pub fn redo_stack(&self) -> &[Rev] {
&self.document.redo_stack
}
pub fn next_node_counter(&self) -> u64 {
self.document.next_node_counter
}
}
/// Errors from `Session::commit_from_runtime`.
#[cfg(any(feature = "conversion", test))]
#[derive(Debug, thiserror::Error)]
pub enum CommitError {
#[error("Failed to convert runtime network: {0}")]
Conversion(#[from] from_runtime::ConversionError),
#[error("Failed to apply commit: {0}")]
Crdt(#[from] CrdtError),
}
#[cfg(any(feature = "conversion", test))]
impl Default for Session {
fn default() -> Self {
Self::new()
}
}
/// One live op in the hot zone. Carries only enough to drive live LWW; no parents (transient),
/// no Rev (not content-addressed in the durable DAG). GC'd at retirement.
#[derive(Clone, Debug, Serialize, Deserialize)]
pub struct HotOp {
pub op: RegistryDelta,
pub timestamp: TimeStamp,
}
#[derive(Debug, thiserror::Error)]
pub enum CrdtError {
#[error("Target node {0} does not exist")]
TargetNodeDoesNotExist(NodeId),
#[error("Network {0} does not exist")]
NetworkDoesNotExist(NetworkId),
#[error("Input index {0} out of bounds")]
InputIndexOutOfBounds(usize),
#[error("Export slot index {0} out of bounds")]
ExportSlotOutOfBounds(u32),
#[error("Delta {0} not found in history")]
NotFoundInHistory(Rev),
#[error("No history entry resurrects node {0}")]
NodeNotInHistory(NodeId),
#[error("No history entry resurrects network {0}")]
NetworkNotInHistory(NetworkId),
#[error("Nothing to undo")]
NothingToUndo,
#[error("Nothing to redo")]
NothingToRedo,
#[error("Node {0} already exists")]
NodeAlreadyExists(NodeId),
#[error("Network {0} already exists")]
NetworkAlreadyExists(NetworkId),
/// PeerId is already registered to a different UserId.
#[error("Peer {0:?} is already registered to a different user")]
PeerRegistrationConflict(PeerId),
#[error("Delta stored under {stored} hashes to {expected}")]
RevMismatch { stored: Rev, expected: Rev },
}
@@ -0,0 +1,881 @@
use core_types::uuid::NodeId as RuntimeNodeId;
use graph_craft::ProtoNodeIdentifier;
use graph_craft::concrete;
use graph_craft::document::{DocumentNode, DocumentNodeImplementation, NodeInput, NodeNetwork};
use crate::InputSlot;
use crate::{Delta, Document, HotOp, Network, NetworkId, NoMetadata, Node, NodeId, PeerId, ROOT_NETWORK, RegistryDelta, RegistryTarget, Session, TimeStamp};
fn fresh_document(peer: PeerId) -> Document {
Session::with_peer(peer).document
}
fn remove_node_op(node_id: NodeId) -> RegistryDelta {
// The snapshot only matters for reverse computation; this op is used to test a no-op removal on an
// absent node, so a placeholder node is fine.
let snapshot = Node::dummy();
RegistryDelta::RemoveNode { id: node_id, snapshot }
}
/// Commit a single op to a document as a retired delta. Mints a fresh timestamp, links to
/// current head, applies, records in history, advances head.
fn commit_op(document: &mut Document, op: RegistryDelta) {
let reverse = document.compute_reverse_delta(RegistryTarget::Working, &op).expect("compute_reverse_delta failed");
let timestamp = document.clock.tick();
let delta = Delta::new(document.head, document.peer, timestamp, op, reverse);
let rev = delta.id;
document.apply_delta(delta).expect("apply_retired_delta failed");
document.head = Some(rev);
}
/// Every applied op must advance the local clock past the op's timestamp, so any subsequent
/// local tick is causally later than what we just observed. Locks in the invariant that
/// `apply_op` calls `clock.observe`, regardless of which apply entry point was used.
#[test]
fn apply_hot_op_advances_clock_past_observed_timestamp() {
let mut document = fresh_document(PeerId(1));
assert_eq!(document.clock.counter, 0);
let observed = TimeStamp { counter: 42, peer: PeerId(2) };
let hot_op = HotOp {
op: remove_node_op(NodeId(99)),
timestamp: observed,
};
document.apply_hot_op(hot_op).expect("RemoveNode on absent node is a no-op, not an error");
assert!(
document.clock.counter >= observed.counter,
"clock counter {} did not advance past observed counter {}",
document.clock.counter,
observed.counter
);
let next = document.clock.tick();
assert!(
next.counter > observed.counter,
"next tick {} must be strictly later than the observed timestamp {}",
next.counter,
observed.counter
);
}
/// `next_node_id` must never repeat across successive calls on the same document. The blake3 output
/// space is enormous, so any collision in a small loop is a counter-bumping bug, not a hash
/// collision.
#[test]
fn next_node_id_is_unique_within_a_document() {
let mut document = fresh_document(PeerId(1));
let mut seen = std::collections::HashSet::new();
for _ in 0..1000 {
let id = document.next_node_id();
assert!(seen.insert(id), "next_node_id repeated after {} calls", seen.len());
}
}
/// Two peers reading the same shared counter must produce different `NodeId`s. This is the whole
/// reason the counter can be shared across peers instead of being per-peer.
#[test]
fn next_node_id_differs_across_peers_at_same_counter() {
let mut document_a = fresh_document(PeerId(1));
let mut document_b = fresh_document(PeerId(2));
let id_a = document_a.next_node_id();
let id_b = document_b.next_node_id();
assert_ne!(id_a, id_b, "peer-scoping is broken: two peers minted the same NodeId at counter 1");
}
fn tiny_network() -> NodeNetwork {
NodeNetwork {
exports: vec![NodeInput::node(RuntimeNodeId(0), 0)],
nodes: [(
RuntimeNodeId(0),
DocumentNode {
inputs: vec![NodeInput::import(concrete!(u32), 0)],
implementation: DocumentNodeImplementation::ProtoNode(ProtoNodeIdentifier::new("graphene_core::ops::identity::IdentityNode")),
..Default::default()
},
)]
.into_iter()
.collect(),
..Default::default()
}
}
/// `verify_history` passes on a normally built history and flags a delta whose content-addressed
/// `id` no longer matches its identity fields (corrupt or crafted history).
#[test]
fn verify_history_detects_rev_mismatch() {
let resources = graphene_resource::ResourceRegistry::new();
let mut session = Session::with_peer(PeerId(1));
session.stage_from_runtime(&tiny_network(), &NoMetadata, &resources).expect("stage failed");
let last_timestamp = session.hot_log().last().expect("staged a hot op").timestamp;
session.retire(last_timestamp).expect("retire failed");
session.verify_history().expect("a freshly built history must validate");
// Tamper one delta's stored id so it no longer matches its content hash.
session.document.history.first_mut().expect("history is non-empty").id = crate::Rev::new(0xdead_beef).unwrap();
assert!(matches!(session.verify_history(), Err(crate::CrdtError::RevMismatch { .. })), "a tampered delta id must be flagged");
}
/// History iteration emits parents before children and is a pure function of the delta set: two
/// sessions independently built from the same network produce byte-identical history order. (The
/// append-order invariant guarantees this directly, with no separate topological sort.)
#[test]
fn history_is_causal_and_deterministic() {
let resources = graphene_resource::ResourceRegistry::new();
let build = || {
let mut session = Session::with_peer(PeerId(1));
session.stage_from_runtime(&tiny_network(), &NoMetadata, &resources).expect("stage failed");
let last_timestamp = session.hot_log().last().expect("staged at least one hot op").timestamp;
session.retire(last_timestamp).expect("retire failed");
session
};
let session_a = build();
let session_b = build();
let order_a: Vec<crate::Rev> = session_a.history().map(|delta| delta.id).collect();
let order_b: Vec<crate::Rev> = session_b.history().map(|delta| delta.id).collect();
assert!(order_a.len() > 1, "expected a multi-delta history to make ordering meaningful");
assert_eq!(order_a, order_b, "same delta set must serialize in the same order");
// Every parent that's part of this history precedes its child.
let position: std::collections::HashMap<crate::Rev, usize> = order_a.iter().enumerate().map(|(i, rev)| (*rev, i)).collect();
for delta in session_a.history() {
for parent in delta.all_parents() {
if let Some(parent_pos) = position.get(&parent) {
assert!(*parent_pos < position[&delta.id], "parent {parent} must precede child {} in order", delta.id);
}
}
}
}
fn set_document_attribute(key: &str, value: u32) -> RegistryDelta {
RegistryDelta::ChangeDocumentAttribute {
delta: crate::AttributeDelta {
key: key.to_string(),
value: Some(serde_json::json!(value)),
},
}
}
/// Two peers that each integrate the other's concurrent branch converge to byte-identical history:
/// the merge commit is parent-set-addressed (same `Rev` on both) and the canonical sort erases the
/// arrival-order difference. Exercises `Session::merge`, the `Merge` variant, and `canonical_sort`.
#[test]
fn merge_converges_to_identical_history() {
// Shared base commit, then a concurrent edit on each peer's own clone of that base.
let mut session_a = Session::with_peer(PeerId(1));
session_a.commit_op_for_test(set_document_attribute("compute::base", 0)).expect("base commit");
let mut session_b = session_a.clone();
session_a.commit_op_for_test(set_document_attribute("compute::a", 1)).expect("A edit");
session_b.commit_op_for_test(set_document_attribute("compute::b", 2)).expect("B edit");
// Cross-merge: feed each peer the other's full delta set. The shared base dedups by `Rev`.
let deltas_a = session_a.cloned_deltas();
let deltas_b = session_b.cloned_deltas();
let merge_a = session_a.merge(deltas_b).expect("merge into A failed").expect("A produced a merge");
let merge_b = session_b.merge(deltas_a).expect("merge into B failed").expect("B produced a merge");
assert_eq!(merge_a, merge_b, "same tips must mint the identical parent-set-addressed merge commit");
let order_a: Vec<crate::Rev> = session_a.history().map(|d| d.id).collect();
let order_b: Vec<crate::Rev> = session_b.history().map(|d| d.id).collect();
assert_eq!(order_a, order_b, "both peers must converge to byte-identical history order");
assert_eq!(session_a.head_rev(), session_b.head_rev(), "both peers land on the same merge head");
}
/// Resurrection must reach into a merged-in branch: a network added then removed on the other peer's
/// branch lives only under the merge's secondary parent, so a `SetNetworkExport` targeting it after
/// the merge can only restore it by traversing all ancestors (not the primary-parent chain).
#[test]
fn resurrection_reaches_across_a_merge() {
let network_id = NetworkId(7);
// Shared base, then peer B adds and removes network 7 on its own branch.
let mut session_a = Session::with_peer(PeerId(1));
session_a.commit_op_for_test(set_document_attribute("compute::base", 0)).expect("base commit");
let mut session_b = session_a.clone();
session_a.commit_op_for_test(set_document_attribute("compute::a", 1)).expect("A edit");
session_b
.commit_op_for_test(RegistryDelta::AddNetwork {
id: network_id,
network: Network::default(),
})
.expect("B AddNetwork");
session_b
.commit_op_for_test(RegistryDelta::RemoveNetwork {
id: network_id,
snapshot: Network::default(),
})
.expect("B RemoveNetwork");
// A merges B's branch: 7's AddNetwork now lives only under the merge's secondary parent.
session_a.merge(session_b.cloned_deltas()).expect("merge failed");
// A SetNetworkExport on 7 must resurrect it by walking into the merged-in branch. Before the
// all-ancestors fix this failed with NetworkNotInHistory (the primary-parent walk missed B's branch).
session_a
.commit_op_for_test(RegistryDelta::SetNetworkExport {
id: network_id,
index: 0,
export: None,
})
.expect("resurrection must find the AddNetwork on the merged-in branch");
}
/// Committing the same NodeNetwork twice must produce zero history entries on the second commit.
/// Without value-only diffing in compute_deltas, the second commit would emit spurious
/// ChangeNodeInput / ChangeNodeAttribute ops because self.registry has real timestamps while the
/// freshly-built `to` registry has TimeStamp::ORIGIN.
#[test]
fn stage_from_runtime_is_idempotent_for_unchanged_network() {
let mut session = Session::with_peer(PeerId(1));
let network = tiny_network();
let resources = graphene_resource::ResourceRegistry::new();
let (first, _) = session.stage_from_runtime(&network, &NoMetadata, &resources).expect("first stage failed");
assert!(!first.is_empty(), "first stage should produce at least one hot op for the initial network");
let (second, _) = session.stage_from_runtime(&network, &NoMetadata, &resources).expect("second stage failed");
assert_eq!(second.len(), 0, "second stage of unchanged network produced {} spurious hot ops: {:?}", second.len(), second);
}
/// The peer's first contribution prepends a `RegisterPeer` op (establishing its `UserId` mapping);
/// later contributions don't re-register, and a no-op batch registers nothing.
#[test]
fn first_contribution_registers_the_peer() {
let mut session = Session::with_peer(PeerId(7));
let resources = graphene_resource::ResourceRegistry::new();
assert!(session.registry().peer_users.is_empty(), "no registration before any contribution");
let (first, _) = session.stage_from_runtime(&tiny_network(), &NoMetadata, &resources).expect("first stage failed");
let registrations = first.iter().filter(|hot_op| matches!(hot_op.op, RegistryDelta::RegisterPeer { .. })).count();
assert_eq!(registrations, 1, "exactly one RegisterPeer on first contribution");
assert!(matches!(first[0].op, RegistryDelta::RegisterPeer { .. }), "RegisterPeer must precede the edit ops");
assert_eq!(session.registry().peer_users.get(&PeerId(7)), Some(&crate::UserId(7)), "peer mapped to its UserId");
// A second, distinct contribution must not re-register.
let mut other_network = tiny_network();
other_network.exports.clear();
let (second, _) = session.stage_from_runtime(&other_network, &NoMetadata, &resources).expect("second stage failed");
assert!(
!second.iter().any(|hot_op| matches!(hot_op.op, RegistryDelta::RegisterPeer { .. })),
"already-registered peer must not re-register"
);
// A no-op batch (re-staging an already-converged network) registers nothing on a fresh peer:
// registration rides a real edit, never a lone op.
let mut fresh = Session::with_peer(PeerId(8));
fresh.stage_from_runtime(&tiny_network(), &NoMetadata, &resources).expect("seed stage failed");
let peers_before = fresh.registry().peer_users.clone();
let (empty, _) = fresh.stage_from_runtime(&tiny_network(), &NoMetadata, &resources).expect("no-op stage failed");
assert!(empty.is_empty(), "an unchanged re-stage must produce no hot ops");
assert_eq!(fresh.registry().peer_users, peers_before, "a no-op batch must not add a registration");
}
/// A SetExport against a removed network must restore the network from history rather than error.
#[test]
fn set_export_resurrects_absent_network() {
let mut document = fresh_document(PeerId(1));
let network_id = NetworkId(7);
commit_op(
&mut document,
RegistryDelta::AddNetwork {
id: network_id,
network: Network::default(),
},
);
commit_op(
&mut document,
RegistryDelta::RemoveNetwork {
id: network_id,
snapshot: Network::default(),
},
);
assert!(!document.working_registry.networks.contains_key(&network_id), "network should be removed before the resurrection test");
commit_op(
&mut document,
RegistryDelta::SetNetworkExport {
id: network_id,
index: 0,
export: None,
},
);
assert!(document.working_registry.networks.contains_key(&network_id), "SetExport should have resurrected the network");
}
/// Cascading resurrection: bringing a node back must also restore its owning network when absent.
#[test]
fn add_node_resurrects_owning_network() {
use crate::Node;
let mut document = fresh_document(PeerId(1));
let network_id = NetworkId(7);
let node_id = NodeId(42);
commit_op(
&mut document,
RegistryDelta::AddNetwork {
id: network_id,
network: Network::default(),
},
);
commit_op(
&mut document,
RegistryDelta::RemoveNetwork {
id: network_id,
snapshot: Network::default(),
},
);
let node = Node { network: network_id, ..Node::dummy() };
commit_op(&mut document, RegistryDelta::AddNode { id: node_id, node });
assert!(
document.working_registry.networks.contains_key(&network_id),
"AddNode should have cascaded a resurrection of the owning network"
);
assert!(document.working_registry.node_instances.contains_key(&node_id), "the node itself should also be present");
}
/// Reverting the same removal twice (the moral equivalent of two peers concurrently resurrecting
/// the same node) must not error on the second apply. Today the second revert hits
/// `apply_op(AddNode, false)` against a present node and returns `NodeAlreadyExists`.
#[test]
fn concurrent_resurrection_via_revert_is_idempotent() {
use crate::Node;
let mut document = fresh_document(PeerId(1));
let network_id = NetworkId(7);
let node_id = NodeId(42);
commit_op(
&mut document,
RegistryDelta::AddNetwork {
id: network_id,
network: Network::default(),
},
);
let node = Node { network: network_id, ..Node::dummy() };
commit_op(&mut document, RegistryDelta::AddNode { id: node_id, node: node.clone() });
commit_op(&mut document, RegistryDelta::RemoveNode { id: node_id, snapshot: node });
assert!(!document.working_registry.node_instances.contains_key(&node_id), "node should be removed before the resurrection test");
document.restore_node_from_history(RegistryTarget::Working, node_id).expect("first resurrection should succeed");
assert!(document.working_registry.node_instances.contains_key(&node_id), "first resurrection should bring the node back");
let second = document.restore_node_from_history(RegistryTarget::Working, node_id);
assert!(second.is_ok(), "second resurrection of an already-present node should be a no-op, got {second:?}");
}
/// History-based resurrection must work when the matching delta is the *root* commit. The history
/// walk used to drop the root (its empty parent list short-circuited the iterator before yielding
/// it), so a node removed by the very first commit could not be restored.
#[test]
fn restore_node_from_root_commit() {
use crate::Node;
let mut document = fresh_document(PeerId(1));
let node_id = NodeId(42);
let node = Node::dummy();
// Seed the working state so the root commit can remove the node (its reverse is the `AddNode` the
// resurrection looks for). This `RemoveNode` is the only commit, so the match sits at the root.
document.working_registry.networks.insert(ROOT_NETWORK, Network::default());
document.retired_snapshot.networks.insert(ROOT_NETWORK, Network::default());
document.working_registry.node_instances.insert(node_id, node.clone());
document.retired_snapshot.node_instances.insert(node_id, node.clone());
commit_op(&mut document, RegistryDelta::RemoveNode { id: node_id, snapshot: node });
assert!(!document.working_registry.node_instances.contains_key(&node_id), "node should be removed by the root commit");
document
.restore_node_from_history(RegistryTarget::Working, node_id)
.expect("resurrection from the root commit should succeed");
assert!(document.working_registry.node_instances.contains_key(&node_id), "node must be restored from the root commit");
}
/// Erroring ops still bump the clock: we observed the timestamp on the wire, the fact that the
/// op was rejected locally doesn't unobserve it.
#[test]
fn apply_op_advances_clock_even_when_op_errors() {
let mut document = fresh_document(PeerId(1));
let observed = TimeStamp { counter: 17, peer: PeerId(2) };
let failing_op = RegistryDelta::ChangeNodeInput {
id: NodeId(7),
index: 0,
new_input: crate::NodeInput::Import { index: 0 },
};
let result = document.apply_op(failing_op, observed);
assert!(result.is_err(), "op targeting a nonexistent node should be rejected");
assert!(document.clock.counter >= observed.counter, "clock should advance on observation even when the op errors");
}
// --- Resource CRDT semantics ---
use crate::{Priority, RegistryDelta as RD, ResourceHash, ResourceId, SourceKey};
fn source_key(priority: f64, peer: u64) -> SourceKey {
SourceKey {
priority: Priority::new(priority).expect("test priorities are finite"),
peer: PeerId(peer),
}
}
fn ts(counter: u64, peer: u64) -> TimeStamp {
TimeStamp { counter, peer: PeerId(peer) }
}
/// Two peers concurrently add a source to the same resource at distinct priorities. Both survive
/// (add-wins union), ordered by priority.
#[test]
fn concurrent_source_adds_at_distinct_priorities_both_survive() {
let mut document = fresh_document(PeerId(1));
let id = ResourceId::new();
document
.apply_op(
RD::AddSource {
id,
key: source_key(0.5, 1),
source: serde_json::json!("embedded"),
},
ts(1, 1),
)
.unwrap();
document
.apply_op(
RD::AddSource {
id,
key: source_key(0.75, 2),
source: serde_json::json!("url"),
},
ts(1, 2),
)
.unwrap();
let entry = document.working_registry.resources.get(&id).expect("resource entry exists");
assert_eq!(entry.sources.len(), 2, "both concurrent additions survive");
// The chain iterates in priority order.
let bodies: Vec<_> = entry.sources.iter().map(|(_, v)| v.source.clone()).collect();
assert_eq!(bodies, vec![serde_json::json!("embedded"), serde_json::json!("url")]);
}
/// Re-adding the same source key is LWW on its timestamp: a later write wins, an earlier one is ignored.
#[test]
fn same_source_key_is_last_writer_wins() {
let mut document = fresh_document(PeerId(1));
let id = ResourceId::new();
let key = source_key(0.5, 1);
document
.apply_op(
RD::AddSource {
id,
key,
source: serde_json::json!("old"),
},
ts(5, 1),
)
.unwrap();
// Earlier timestamp: ignored.
document
.apply_op(
RD::AddSource {
id,
key,
source: serde_json::json!("stale"),
},
ts(2, 1),
)
.unwrap();
// Later timestamp: wins.
document
.apply_op(
RD::AddSource {
id,
key,
source: serde_json::json!("new"),
},
ts(9, 1),
)
.unwrap();
let entry = document.working_registry.resources.get(&id).unwrap();
assert_eq!(entry.source(&key).unwrap().source, serde_json::json!("new"));
}
/// SetResourceHash is LWW on the hash; a later resolve wins, an earlier one is ignored.
#[test]
fn register_resource_hash_is_last_writer_wins() {
let mut document = fresh_document(PeerId(1));
let id = ResourceId::new();
let hash_a = ResourceHash::from(&b"alpha"[..]);
let hash_b = ResourceHash::from(&b"beta"[..]);
document.apply_op(RD::SetResourceHash { id, hash: Some(hash_a) }, ts(5, 1)).unwrap();
document.apply_op(RD::SetResourceHash { id, hash: Some(hash_b) }, ts(2, 1)).unwrap();
assert_eq!(document.working_registry.resources.get(&id).unwrap().hash, Some(hash_a), "earlier resolve must not clobber later one");
document.apply_op(RD::SetResourceHash { id, hash: Some(hash_b) }, ts(9, 1)).unwrap();
assert_eq!(document.working_registry.resources.get(&id).unwrap().hash, Some(hash_b), "later resolve wins");
}
/// The reverse delta of a RemoveSource restores the prior source body, and applying op-then-reverse
/// round-trips the source chain.
#[test]
fn remove_source_reverse_restores_prior() {
let mut document = fresh_document(PeerId(1));
let id = ResourceId::new();
let key = source_key(0.5, 1);
commit_op(
&mut document,
RD::AddSource {
id,
key,
source: serde_json::json!("kept"),
},
);
// Compute the reverse while the body is still present, then apply the removal.
let reverse = document.compute_reverse_delta(RegistryTarget::Working, &RD::RemoveSource { id, key }).unwrap();
match &reverse {
RD::AddSource { source, .. } => assert_eq!(*source, serde_json::json!("kept"), "reverse of removal re-adds the body"),
other => panic!("expected AddSource reverse, got {other:?}"),
}
document.apply_op(RD::RemoveSource { id, key }, ts(5, 1)).unwrap();
assert!(document.working_registry.resources.get(&id).unwrap().sources.is_empty(), "source removed");
// Applying the reverse restores the chain.
document.apply_op(reverse, ts(6, 1)).unwrap();
assert_eq!(document.working_registry.resources.get(&id).unwrap().source(&key).unwrap().source, serde_json::json!("kept"));
}
/// AddSource on a fresh slot reverses to a RemoveSource; on an occupied slot it restores the prior body.
#[test]
fn add_source_reverse_depends_on_prior_state() {
let mut document = fresh_document(PeerId(1));
let id = ResourceId::new();
let key = source_key(0.5, 1);
// Fresh slot: reverse removes.
let reverse_fresh = document
.compute_reverse_delta(
RegistryTarget::Working,
&RD::AddSource {
id,
key,
source: serde_json::json!("first"),
},
)
.unwrap();
assert!(matches!(reverse_fresh, RD::RemoveSource { .. }), "reverse of add-to-empty is remove, got {reverse_fresh:?}");
// Occupy the slot, then reverse of a new add restores the existing body.
document
.apply_op(
RD::AddSource {
id,
key,
source: serde_json::json!("existing"),
},
ts(1, 1),
)
.unwrap();
let reverse_overwrite = document
.compute_reverse_delta(
RegistryTarget::Working,
&RD::AddSource {
id,
key,
source: serde_json::json!("overwrite"),
},
)
.unwrap();
match reverse_overwrite {
RD::AddSource { source, .. } => assert_eq!(source, serde_json::json!("existing"), "reverse restores prior body"),
other => panic!("expected AddSource reverse, got {other:?}"),
}
}
// --- compute_deltas resource diffing ---
use crate::{ResourceEntry, ResourceStore, SourceValue};
fn entry_with_source(priority: f64, peer: u64, body: serde_json::Value, hash: Option<ResourceHash>) -> ResourceEntry {
ResourceEntry {
sources: vec![(source_key(priority, peer), SourceValue { source: body, timestamp: ts(1, peer) })],
hash,
hash_timestamp: ts(1, peer),
}
}
fn registry_with_resources(resources: ResourceStore) -> crate::Registry {
crate::Registry { resources, ..Default::default() }
}
/// An unchanged resource store produces zero deltas, even when timestamps differ (value-only diff).
#[test]
fn compute_deltas_ignores_unchanged_resources() {
let id = ResourceId::new();
let hash = ResourceHash::from(&b"img"[..]);
let mut from = ResourceStore::new();
from.insert(id, entry_with_source(0.0, 1, serde_json::json!("embedded"), Some(hash)));
// Same value, different timestamps: must not count as a change.
let mut to = ResourceStore::new();
let mut to_entry = entry_with_source(0.0, 1, serde_json::json!("embedded"), Some(hash));
to_entry.hash_timestamp = ts(99, 2);
to_entry.sources.iter_mut().for_each(|(_, v)| v.timestamp = ts(99, 2));
to.insert(id, to_entry);
let deltas = crate::delta::compute_deltas(&registry_with_resources(from), &registry_with_resources(to));
assert!(deltas.is_empty(), "unchanged resource (value-equal) produced deltas: {deltas:?}");
}
/// Adding, changing, and removing resources each produce the matching delta, and applying the diff
/// transforms `from` into a registry value-equal to `to`.
#[test]
fn compute_deltas_diffs_resources_and_round_trips() {
let kept = ResourceId::new();
let removed = ResourceId::new();
let added = ResourceId::new();
let hash_old = ResourceHash::from(&b"old"[..]);
let hash_new = ResourceHash::from(&b"new"[..]);
let mut from = ResourceStore::new();
from.insert(kept, entry_with_source(0.0, 1, serde_json::json!("embedded"), Some(hash_old)));
from.insert(removed, entry_with_source(0.0, 1, serde_json::json!("gone"), None));
let mut to = ResourceStore::new();
// `kept`: hash changes and a second source is added.
let mut kept_entry = entry_with_source(0.0, 1, serde_json::json!("embedded"), Some(hash_new));
kept_entry.set_source(
source_key(1.0, 1),
SourceValue {
source: serde_json::json!("url"),
timestamp: ts(1, 1),
},
);
to.insert(kept, kept_entry);
// `added`: brand new resource.
to.insert(added, entry_with_source(0.0, 1, serde_json::json!("fresh"), None));
let deltas = crate::delta::compute_deltas(&registry_with_resources(from.clone()), &registry_with_resources(to.clone()));
// A brand-new resource is a single whole-entry AddResource, never a fan-out of per-source ops.
let added_deltas: Vec<_> = deltas.iter().filter(|d| matches!(d, RD::AddResource { id, .. } if *id == added)).collect();
assert_eq!(added_deltas.len(), 1, "adding a resource should produce exactly one AddResource delta, got {added_deltas:?}");
assert!(
!deltas.iter().any(|d| matches!(d, RD::AddSource { id, .. } | RD::SetResourceHash { id, .. } if *id == added)),
"a brand-new resource must not emit per-source or hash ops"
);
// The removed resource is a single whole-entry RemoveResource.
assert_eq!(
deltas.iter().filter(|d| matches!(d, RD::RemoveResource { id, .. } if *id == removed)).count(),
1,
"removing a resource should produce exactly one RemoveResource delta"
);
// Apply the diff to a document seeded with `from`, then check it matches `to` by value.
let mut document = fresh_document(PeerId(1));
document.working_registry = registry_with_resources(from);
for op in deltas {
let timestamp = document.clock.tick();
document.apply_op(op, timestamp).expect("apply resource delta");
}
assert!(
document.working_registry.value_equal(&registry_with_resources(to)),
"applying the resource diff did not reproduce the target registry"
);
}
/// Resource GC must keep an undone interaction's resources alive: undo removes a interaction's `AddResource`
/// from the working registry, but redo still needs those bytes. `all_referenced_resource_hashes` must
/// therefore report history-referenced resources even after they leave the current registry, so the
/// editor's GC "used" set doesn't evict them between an undo and a redo.
#[test]
fn all_referenced_resource_hashes_survives_undo() {
use crate::ResourceId;
let mut session = Session::with_peer(PeerId(1));
let resources = graphene_resource::ResourceRegistry::new();
// Base interaction: the first interaction is intentionally not undoable (the mount-base floor), so commit a
// network first. Undoing the later resource interaction then lands on this base rather than the root.
session.stage_from_runtime(&tiny_network(), &NoMetadata, &resources).expect("stage base");
let base_up_to = session.hot_log().last().expect("staged base").timestamp;
let base_revs = session.retire(base_up_to).expect("retire base");
session.mark_interaction_end(*base_revs.last().expect("one base delta"));
// Second interaction: add a resource and mark the retired delta as a interaction boundary.
let hash = ResourceHash::from(&b"declaration-bytes"[..]);
let id = ResourceId::new();
let hot_ops = session.stage_embedded_resource(id, hash).expect("stage resource");
let up_to = hot_ops.last().expect("staged one op").timestamp;
let revs = session.retire(up_to).expect("retire");
session.mark_interaction_end(*revs.last().expect("one retired delta"));
assert!(session.registry().resources.contains_key(&id), "resource is present after the interaction");
assert!(session.all_referenced_resource_hashes().contains(&hash));
// Undo the interaction: the resource leaves the working registry but stays in history.
session.undo().expect("undo");
assert!(!session.registry().resources.contains_key(&id), "undo drops the resource from the working registry");
assert!(
session.all_referenced_resource_hashes().contains(&hash),
"the undone interaction's resource must still be reported so GC keeps its bytes for redo"
);
}
/// A commit that produces no deltas must not touch the redo stack. Redo is only abandoned by a real
/// new edit; a no-op commit (here `embed_resource_sources` over an empty id set) leaving it cleared
/// would silently disable redo after an undo.
#[test]
fn no_op_commit_preserves_redo_stack() {
let mut session = Session::with_peer(PeerId(1));
let resources = graphene_resource::ResourceRegistry::new();
// Base interaction (the non-undoable mount floor), then a second interaction to undo onto it.
session.stage_from_runtime(&tiny_network(), &NoMetadata, &resources).expect("stage base");
let base_up_to = session.hot_log().last().expect("staged base").timestamp;
let base_revs = session.retire(base_up_to).expect("retire base");
session.mark_interaction_end(*base_revs.last().expect("one base delta"));
let hash = ResourceHash::from(&b"declaration-bytes"[..]);
let id = ResourceId::new();
let hot_ops = session.stage_embedded_resource(id, hash).expect("stage resource");
let up_to = hot_ops.last().expect("staged one op").timestamp;
let revs = session.retire(up_to).expect("retire");
session.mark_interaction_end(*revs.last().expect("one retired delta"));
session.undo().expect("undo");
assert!(session.can_redo(), "undo must populate the redo stack");
// A commit over no resources produces no deltas; redo must survive it.
session.embed_resource_sources(std::iter::empty::<ResourceId>()).expect("no-op embed");
assert!(session.can_redo(), "a no-op commit must not clear the redo stack");
}
/// `embed_resource_sources` commits its `AddSource` deltas as retired, then mirrors them onto the
/// working registry. With unretired hot ops present it must keep the hot-zone edits (export of a
/// mid-interaction document is lossless) rather than clobbering the working registry with the snapshot.
#[test]
fn embed_resource_sources_preserves_unretired_hot_ops() {
let mut session = Session::with_peer(PeerId(1));
let resources = graphene_resource::ResourceRegistry::new();
// Retire a base so the network's nodes live in the retired snapshot.
session.stage_from_runtime(&tiny_network(), &NoMetadata, &resources).expect("stage base");
let base_up_to = session.hot_log().last().expect("staged base").timestamp;
session.retire(base_up_to).expect("retire base");
// Stage an embedded resource without retiring, leaving it in the hot log (the working registry now
// holds it, the retired snapshot does not).
let hash = ResourceHash::from(&b"hot-resource"[..]);
let id = ResourceId::new();
session.stage_embedded_resource(id, hash).expect("stage resource");
assert!(!session.hot_log().is_empty(), "staging should leave unretired hot ops");
assert!(session.registry().resources.contains_key(&id), "working registry should hold the hot resource");
session.embed_resource_sources(std::iter::empty::<ResourceId>()).expect("embed tolerates a non-empty hot log");
// The hot-zone resource survives in the working registry (not reset to the snapshot), and the hot log
// is untouched so a later retire still promotes it.
assert!(session.registry().resources.contains_key(&id), "hot resource must survive the embed");
assert!(!session.hot_log().is_empty(), "embed must not drain the hot log");
}
/// A delta's `Rev` is content-addressed, so two byte-equal deltas must hash identically regardless
/// of the order their attributes were inserted. This guards the `Attributes` map staying canonically
/// ordered (`BTreeMap`): a hash-randomized map would give the same logical delta different `Rev`s.
#[test]
fn add_node_rev_is_independent_of_attribute_insertion_order() {
use crate::{AttributesWrite, Implementation, Value};
let keys = ["ui::position", "ui::display_name", "ui::locked", "ui::pinned", "call_argument", "context_features"];
// Fixed implementation so the two nodes differ only in attribute insertion order.
let implementation = Implementation::ProtoNode(ResourceId::new());
let make_node = |insertion_order: &[&str]| {
let mut attributes = crate::Attributes::new();
for &key in insertion_order {
attributes.set(key, serde_json::json!(key), TimeStamp::ORIGIN);
}
let mut input_attributes = crate::Attributes::new();
for &key in insertion_order {
input_attributes.insert(key.to_string(), Value::new(serde_json::json!(key), TimeStamp::ORIGIN));
}
let inputs = vec![InputSlot {
input: crate::NodeInput::Import { index: 0 },
timestamp: TimeStamp::ORIGIN,
attributes: input_attributes,
}];
Node {
implementation: implementation.clone(),
inputs,
attributes,
network: ROOT_NETWORK,
}
};
let forward: Vec<&str> = keys.to_vec();
let reversed: Vec<&str> = keys.iter().rev().copied().collect();
let parent = crate::Rev::new(1);
let author = PeerId(7);
let timestamp = TimeStamp { counter: 42, peer: PeerId(7) };
let delta_forward = Delta::new(
parent,
author,
timestamp,
RegistryDelta::AddNode {
id: NodeId(9),
node: make_node(&forward),
},
RegistryDelta::AddNode {
id: NodeId(9),
node: make_node(&forward),
},
);
let delta_reversed = Delta::new(
parent,
author,
timestamp,
RegistryDelta::AddNode {
id: NodeId(9),
node: make_node(&reversed),
},
RegistryDelta::AddNode {
id: NodeId(9),
node: make_node(&reversed),
},
);
assert_eq!(delta_forward.id, delta_reversed.id, "Rev must not depend on attribute insertion order");
}
@@ -0,0 +1,781 @@
use std::borrow::Cow;
use std::collections::HashMap;
use core_types::context::ContextDependencies;
use core_types::uuid::NodeId;
use graph_craft::document::{DocumentNode, DocumentNodeImplementation, NodeInput, NodeNetwork};
use graph_craft::graphene_compiler::Compiler;
use graph_craft::{ProtoNodeIdentifier, Type, concrete};
use crate::{NetworkId, NodeMetadataSource, PeerId, Position, Registry};
/// Helper function to verify a NodeNetwork can be compiled successfully.
/// Note: This only works for complete networks with all inputs resolved.
/// Test networks with Import inputs will fail compilation (which is expected).
fn verify_network_compiles(network: &NodeNetwork) -> Result<(), String> {
let compiler = Compiler {};
compiler
.compile_single(network.clone(), &graph_craft::proto::Registry::new())
.map_err(|e| format!("Compilation failed: {:?}", e))?;
Ok(())
}
/// Convert a runtime network to a storage `Registry`, returning the declarations alongside it.
/// Proto-node declaration content is no longer stored in the registry (it lives in a byte store);
/// these tests have no byte store, so they keep the extracted bytes in hand and rebuild a
/// `Declarations` map for the back-conversion.
fn to_registry(network: &NodeNetwork) -> (Registry, crate::Declarations) {
let conversion = Registry::convert_from_runtime(network, &crate::NoMetadata, &Default::default(), PeerId(0)).expect("Failed to convert NodeNetwork to Registry");
let declarations = conversion.declarations().expect("rebuild declarations");
(conversion.registry, declarations)
}
/// A one-node network whose single node references `id` via a `TaggedValue::Resource` input, so
/// `convert_resources` (which only snapshots network-referenced resources) carries the resource.
fn network_referencing_resource(id: graphene_resource::ResourceId) -> NodeNetwork {
network_referencing_resources(&[id])
}
/// A network with one node per resource, each referencing its resource via a `TaggedValue::Resource`
/// input, so all listed resources are network-referenced and survive conversion.
fn network_referencing_resources(ids: &[graphene_resource::ResourceId]) -> NodeNetwork {
use graph_craft::document::value::TaggedValue;
let nodes = ids
.iter()
.enumerate()
.map(|(i, id)| {
(
NodeId(i as u64),
DocumentNode {
inputs: vec![NodeInput::value(TaggedValue::Resource(*id), false)],
implementation: DocumentNodeImplementation::ProtoNode(ProtoNodeIdentifier::new("graphene_core::ops::identity::IdentityNode")),
..Default::default()
},
)
})
.collect();
NodeNetwork { nodes, ..Default::default() }
}
fn create_simple_network() -> NodeNetwork {
NodeNetwork {
exports: vec![NodeInput::node(NodeId(1), 0)],
nodes: [
(
NodeId(0),
DocumentNode {
inputs: vec![NodeInput::import(concrete!(u32), 0), NodeInput::import(concrete!(u32), 1)],
implementation: DocumentNodeImplementation::ProtoNode(ProtoNodeIdentifier::new("graphene_core::structural::ConsNode")),
..Default::default()
},
),
(
NodeId(1),
DocumentNode {
inputs: vec![NodeInput::node(NodeId(0), 0)],
implementation: DocumentNodeImplementation::ProtoNode(ProtoNodeIdentifier::new("graphene_core::ops::AddPairNode")),
..Default::default()
},
),
]
.into_iter()
.collect(),
..Default::default()
}
}
/// Creates a network with a nested sub-network
fn create_nested_network() -> NodeNetwork {
// Create a simple inner network
let inner_network = NodeNetwork {
exports: vec![NodeInput::node(NodeId(10), 0)],
nodes: [(
NodeId(10),
DocumentNode {
inputs: vec![NodeInput::import(concrete!(u32), 0)],
implementation: DocumentNodeImplementation::ProtoNode(ProtoNodeIdentifier::new("graphene_core::ops::identity::IdentityNode")),
..Default::default()
},
)]
.into_iter()
.collect(),
..Default::default()
};
// Create outer network that uses the inner network
NodeNetwork {
exports: vec![NodeInput::node(NodeId(1), 0)],
nodes: [
(
NodeId(0),
DocumentNode {
inputs: vec![NodeInput::import(concrete!(u32), 0)],
implementation: DocumentNodeImplementation::Network(inner_network),
..Default::default()
},
),
(
NodeId(1),
DocumentNode {
inputs: vec![NodeInput::node(NodeId(0), 0)],
implementation: DocumentNodeImplementation::ProtoNode(ProtoNodeIdentifier::new("graphene_core::ops::identity::IdentityNode")),
..Default::default()
},
),
]
.into_iter()
.collect(),
..Default::default()
}
}
#[test]
fn test_simple_round_trip() {
let original_network = create_simple_network();
// Convert to Registry
let (registry, declarations) = to_registry(&original_network);
// Convert back to NodeNetwork
let (converted_network, _) = registry.to_runtime_with_metadata(&declarations).expect("Failed to convert Registry back to NodeNetwork");
// Verify structure is preserved
assert_eq!(converted_network.nodes.len(), original_network.nodes.len(), "Node count should be preserved");
assert_eq!(converted_network.exports.len(), original_network.exports.len(), "Export count should be preserved");
// Verify exports reference the correct nodes
match (&original_network.exports[0], &converted_network.exports[0]) {
(
NodeInput::Node {
node_id: orig_id,
output_index: orig_idx,
},
NodeInput::Node {
node_id: conv_id,
output_index: conv_idx,
},
) => {
assert_eq!(orig_id, conv_id, "Export should reference the same node");
assert_eq!(orig_idx, conv_idx, "Export output index should match");
}
_ => panic!("Exports should both be Node inputs"),
}
// Verify node implementations are preserved
for (node_id, orig_node) in &original_network.nodes {
let conv_node = converted_network.nodes.get(node_id).expect("Node should exist after round-trip");
match (&orig_node.implementation, &conv_node.implementation) {
(DocumentNodeImplementation::ProtoNode(orig_ident), DocumentNodeImplementation::ProtoNode(conv_ident)) => {
assert_eq!(orig_ident.as_str(), conv_ident.as_str(), "ProtoNode identifier should be preserved");
}
_ => panic!("Implementation type should be preserved"),
}
// Verify input count is preserved
assert_eq!(conv_node.inputs.len(), orig_node.inputs.len(), "Input count should be preserved");
}
}
#[test]
fn test_nested_network_round_trip() {
let original_network = create_nested_network();
// Convert to Registry
let (registry, declarations) = to_registry(&original_network);
// Convert back to NodeNetwork
let (converted_network, _) = registry.to_runtime_with_metadata(&declarations).expect("Failed to convert Registry back to NodeNetwork");
// Verify structure is preserved
assert_eq!(converted_network.nodes.len(), original_network.nodes.len(), "Node count should be preserved");
// Find the node with nested network
let orig_nested_node = original_network.nodes.get(&NodeId(0)).expect("Node 0 should exist");
let conv_nested_node = converted_network.nodes.get(&NodeId(0)).expect("Node 0 should exist after round-trip");
// Verify nested network is preserved
match (&orig_nested_node.implementation, &conv_nested_node.implementation) {
(DocumentNodeImplementation::Network(orig_inner), DocumentNodeImplementation::Network(conv_inner)) => {
assert_eq!(orig_inner.nodes.len(), conv_inner.nodes.len(), "Inner network node count should be preserved");
assert_eq!(orig_inner.exports.len(), conv_inner.exports.len(), "Inner network export count should be preserved");
}
_ => panic!("Nested network should be preserved"),
}
}
#[test]
fn test_registry_structure() {
let network = create_simple_network();
let (registry, _declarations) = to_registry(&network);
assert!(registry.resources.len() >= 2, "Should have proto-node declaration resources");
assert!(!registry.networks.is_empty(), "Should have at least one network");
let root_network = registry.networks.get(&crate::ROOT_NETWORK).expect("Root network should exist");
assert_eq!(root_network.exports.len(), network.exports.len(), "Export count should match");
// Exports are first-class slots, no synthetic identity nodes in node_instances.
for slot in &root_network.exports {
assert!(slot.target.is_some(), "Round-tripped exports should have a target");
}
}
#[test]
fn test_nested_network_flattening() {
let network = create_nested_network();
let registry = Registry::try_from(&network).expect("Failed to convert to Registry");
// Outer network has 2 nodes, one of which contains a nested network with 1 node.
// No more identity-node padding, so node_instances has exactly the real nodes.
let expected_nodes = 3;
assert_eq!(
registry.node_instances.len(),
expected_nodes,
"Registry should have exactly {} nodes, found {}",
expected_nodes,
registry.node_instances.len()
);
// Two networks: root (ROOT_NETWORK) and nested (1).
assert!(registry.networks.len() >= 2, "Should have at least 2 networks (root + nested)");
}
#[test]
fn test_metadata_preservation() {
// Create a network with nodes that have non-default metadata
let network = NodeNetwork {
exports: vec![NodeInput::node(NodeId(1), 0)],
nodes: [
(
NodeId(0),
DocumentNode {
inputs: vec![NodeInput::import(concrete!(f64), 0), NodeInput::import(Type::Generic(Cow::Borrowed("T")), 1)],
implementation: DocumentNodeImplementation::ProtoNode(ProtoNodeIdentifier::new("test::NodeWithMetadata")),
call_argument: concrete!(String),
visible: false, // Non-default value
skip_deduplication: true, // Non-default value
..Default::default()
},
),
(
NodeId(1),
DocumentNode {
inputs: vec![NodeInput::node(NodeId(0), 0)],
implementation: DocumentNodeImplementation::ProtoNode(ProtoNodeIdentifier::new("test::OutputNode")),
call_argument: concrete!((u32, u32)),
..Default::default()
},
),
]
.into_iter()
.collect(),
..Default::default()
};
// Convert to Registry and back
let (registry, declarations) = to_registry(&network);
let (converted, _) = registry.to_runtime_with_metadata(&declarations).expect("Failed to convert back to NodeNetwork");
// Verify call_argument is preserved
let orig_node_0 = network.nodes.get(&NodeId(0)).unwrap();
let conv_node_0 = converted.nodes.get(&NodeId(0)).unwrap();
assert_eq!(orig_node_0.call_argument, conv_node_0.call_argument, "call_argument for node 0 should be preserved");
let orig_node_1 = network.nodes.get(&NodeId(1)).unwrap();
let conv_node_1 = converted.nodes.get(&NodeId(1)).unwrap();
assert_eq!(orig_node_1.call_argument, conv_node_1.call_argument, "call_argument for node 1 should be preserved");
// Verify context_features is not stored
assert_eq!(
conv_node_0.context_features,
ContextDependencies::default(),
"context_features should resolve at compile, not round-trip"
);
// Verify visible is preserved
assert_eq!(orig_node_0.visible, conv_node_0.visible, "visible should be preserved");
// Verify skip_deduplication is preserved
assert_eq!(orig_node_0.skip_deduplication, conv_node_0.skip_deduplication, "skip_deduplication should be preserved");
// Verify import_type is preserved for Import inputs
match (&orig_node_0.inputs[0], &conv_node_0.inputs[0]) {
(NodeInput::Import { import_type: orig_type, .. }, NodeInput::Import { import_type: conv_type, .. }) => {
assert_eq!(orig_type, conv_type, "import_type for first import should be preserved (f64)");
}
_ => panic!("First input should be Import"),
}
match (&orig_node_0.inputs[1], &conv_node_0.inputs[1]) {
(NodeInput::Import { import_type: orig_type, .. }, NodeInput::Import { import_type: conv_type, .. }) => {
assert_eq!(orig_type, conv_type, "import_type for second import should be preserved (generic T)");
}
_ => panic!("Second input should be Import"),
}
}
#[test]
fn test_demo_artwork_round_trip() {
use graph_craft::util::{DEMO_ART, load_from_name};
// Test each demo artwork
for artwork_name in DEMO_ART {
println!("Testing artwork: {}", artwork_name);
let original_network = load_from_name(artwork_name);
// Convert to Registry
let (registry, declarations) = to_registry(&original_network);
// Convert back to NodeNetwork
let (converted_network, _) = registry
.to_runtime_with_metadata(&declarations)
.unwrap_or_else(|e| panic!("Failed to convert {} back to NodeNetwork: {:?}", artwork_name, e));
// Basic structural checks
assert_eq!(original_network.nodes.len(), converted_network.nodes.len(), "{}: Node count should be preserved", artwork_name);
assert_eq!(original_network.exports.len(), converted_network.exports.len(), "{}: Export count should be preserved", artwork_name);
// Verify each node's metadata is preserved
for (node_id, orig_node) in &original_network.nodes {
let conv_node = converted_network
.nodes
.get(node_id)
.unwrap_or_else(|| panic!("{}: Node {:?} should exist after round-trip", artwork_name, node_id));
// Check metadata fields
assert_eq!(
orig_node.call_argument, conv_node.call_argument,
"{}: call_argument should be preserved for node {:?}",
artwork_name, node_id
);
assert_eq!(
orig_node.context_features, conv_node.context_features,
"{}: context_features should be preserved for node {:?}",
artwork_name, node_id
);
assert_eq!(orig_node.visible, conv_node.visible, "{}: visible should be preserved for node {:?}", artwork_name, node_id);
assert_eq!(
orig_node.skip_deduplication, conv_node.skip_deduplication,
"{}: skip_deduplication should be preserved for node {:?}",
artwork_name, node_id
);
// Check input count
assert_eq!(
orig_node.inputs.len(),
conv_node.inputs.len(),
"{}: Input count should be preserved for node {:?}",
artwork_name,
node_id
);
}
// Verify the converted demo artwork can be compiled (demo artworks are complete networks)
verify_network_compiles(&converted_network).unwrap_or_else(|e| panic!("{}: Converted artwork should compile successfully: {}", artwork_name, e));
println!("✓ {} passed", artwork_name);
}
}
/// Per-node UI state used by the in-test metadata source. Keyed by `(network_path, local_id)`.
#[derive(Clone, Debug, Default, PartialEq)]
struct UiState {
position: Option<Position>,
is_layer: bool,
display_name: Option<String>,
locked: bool,
pinned: bool,
}
/// In-test `NodeMetadataSource` backed by a `HashMap` keyed on the full `(network_path, local_id)`
/// addressing the editor would use.
struct TestMetadata {
entries: HashMap<(Vec<NodeId>, NodeId), UiState>,
}
impl TestMetadata {
fn new() -> Self {
Self { entries: HashMap::new() }
}
fn insert(&mut self, network_path: &[NodeId], local_id: NodeId, state: UiState) {
self.entries.insert((network_path.to_vec(), local_id), state);
}
fn get(&self, network_path: &[NodeId], local_id: NodeId) -> Option<&UiState> {
self.entries.get(&(network_path.to_vec(), local_id))
}
}
impl NodeMetadataSource for TestMetadata {
fn position(&self, network_path: &[NodeId], local_id: NodeId) -> Option<Position> {
self.get(network_path, local_id).and_then(|s| s.position)
}
fn is_layer(&self, network_path: &[NodeId], local_id: NodeId) -> bool {
self.get(network_path, local_id).is_some_and(|s| s.is_layer)
}
fn display_name(&self, network_path: &[NodeId], local_id: NodeId) -> Option<&str> {
self.get(network_path, local_id).and_then(|s| s.display_name.as_deref())
}
fn locked(&self, network_path: &[NodeId], local_id: NodeId) -> bool {
self.get(network_path, local_id).is_some_and(|s| s.locked)
}
fn pinned(&self, network_path: &[NodeId], local_id: NodeId) -> bool {
self.get(network_path, local_id).is_some_and(|s| s.pinned)
}
}
/// Round-trips a nested network with editor metadata: layer + absolute position on one node,
/// node-in-chain on another, layer-in-stack inside a nested network. Asserts every entry comes
/// back unchanged and addressed by the correct `(network_path, local_id)`.
#[test]
fn test_ui_metadata_round_trip() {
let network = create_nested_network();
let mut metadata = TestMetadata::new();
// Root-network node 0 (the one with a nested network): a layer at an absolute position with
// a display name. Editor `network_path` for root-network nodes is empty.
metadata.insert(
&[],
NodeId(0),
UiState {
position: Some(Position::Absolute([3, 5])),
is_layer: true,
display_name: Some("Outer layer".into()),
locked: true,
pinned: false,
},
);
// Root-network node 1: a plain node in a chain.
metadata.insert(
&[],
NodeId(1),
UiState {
position: Some(Position::Chain),
..Default::default()
},
);
// Nested-network node 10 (lives under node 0): a layer in a stack.
metadata.insert(
&[NodeId(0)],
NodeId(10),
UiState {
position: Some(Position::Stack(7)),
is_layer: true,
..Default::default()
},
);
let conversion = Registry::convert_from_runtime(&network, &metadata, &Default::default(), PeerId(0)).expect("Failed to convert to Registry with metadata");
let declarations = conversion.declarations().expect("rebuild declarations");
let registry = conversion.registry;
let (converted, entries) = registry.to_runtime_with_metadata(&declarations).expect("Failed to convert Registry back with metadata");
// Graph structure still round-trips.
assert_eq!(converted.nodes.len(), network.nodes.len());
// Three entries — one per node we attached metadata to.
assert_eq!(entries.len(), 3, "expected 3 metadata entries, got {}: {entries:#?}", entries.len());
// Look entries back up by their address so we don't rely on emission order.
let lookup: HashMap<(Vec<NodeId>, NodeId), &crate::NodeMetadataEntry> = entries.iter().map(|e| ((e.network_path.clone(), e.local_id), e)).collect();
let root_layer = lookup.get(&(vec![], NodeId(0))).expect("entry for root-network layer node missing");
assert_eq!(root_layer.position, Some(Position::Absolute([3, 5])));
assert!(root_layer.is_layer);
assert_eq!(root_layer.display_name.as_deref(), Some("Outer layer"));
assert!(root_layer.locked);
assert!(!root_layer.pinned);
let root_node = lookup.get(&(vec![], NodeId(1))).expect("entry for root-network chain node missing");
assert_eq!(root_node.position, Some(Position::Chain));
assert!(!root_node.is_layer);
let nested_layer = lookup.get(&(vec![NodeId(0)], NodeId(10))).expect("entry for nested layer-in-stack missing");
assert_eq!(nested_layer.position, Some(Position::Stack(7)));
assert!(nested_layer.is_layer);
}
/// A runtime `ResourceRegistry` (source chain + resolved hash) survives conversion into the storage
/// `Registry`: source bodies are preserved in priority order and the hash carries through.
#[test]
fn resources_round_trip_through_from_runtime() {
use graphene_resource::{DataSource, ResourceHash, ResourceId, ResourceRegistry};
let mut resources = ResourceRegistry::new();
let id = ResourceId::new();
// Two sources in chain order: an embedded fallback then a URL.
resources.push_source_back(&id, DataSource::Embedded);
resources.push_source_back(&id, DataSource::Url("https://example.com/img.png".parse().unwrap()));
let hash = ResourceHash::from(&b"image bytes"[..]);
resources.resolve(&id, hash);
// The resource must be referenced by a node to be snapshotted: `convert_resources` only carries
// resources the network uses (orphans in the runtime cache, e.g. retained across undo, are dropped).
let network = network_referencing_resource(id);
let registry = Registry::from_runtime_with_metadata(&network, &crate::NoMetadata, &resources, PeerId(7)).expect("from_runtime failed");
let entry = registry.resources.get(&id).expect("resource entry present in storage registry");
assert_eq!(entry.hash, Some(hash), "resolved hash carried through");
assert_eq!(entry.sources.len(), 2, "both sources carried through");
// The chain iterates in priority order; decode bodies back to DataSource to compare.
let decoded: Vec<DataSource> = entry.sources.iter().map(|(_, v)| serde_json::from_value(v.source.clone()).expect("source body decodes")).collect();
assert_eq!(decoded, vec![DataSource::Embedded, DataSource::Url("https://example.com/img.png".parse().unwrap())]);
// All source keys carry the document peer.
assert!(entry.sources.iter().all(|(key, _)| key.peer == PeerId(7)), "source keys scoped to the document peer");
}
/// Full resource round-trip: a runtime `ResourceRegistry` converted into storage and back is equal
/// to the original (source chains in order, resolved hashes preserved).
#[test]
fn resource_registry_round_trips_runtime_to_storage_to_runtime() {
use graphene_resource::{DataSource, ResourceHash, ResourceId, ResourceRegistry};
let mut original = ResourceRegistry::new();
// A resolved resource with a two-entry fallback chain.
let image = ResourceId::new();
original.push_source_back(&image, DataSource::Embedded);
original.push_source_back(&image, DataSource::Url("https://example.com/img.png".parse().unwrap()));
original.resolve(&image, ResourceHash::from(&b"image bytes"[..]));
// An unresolved resource (sources but no hash yet).
let font = ResourceId::new();
original.push_source_back(
&font,
DataSource::Font {
family: "Inter".into(),
style: Some("Bold".into()),
},
);
// Both resources must be referenced by a node to be snapshotted (see `convert_resources`).
let network = network_referencing_resources(&[image, font]);
let registry = Registry::from_runtime_with_metadata(&network, &crate::NoMetadata, &original, PeerId(3)).expect("from_runtime failed");
let restored = registry.to_resource_registry().expect("to_resource_registry failed");
// Compare the two document resources specifically; the referencing nodes' proto-node declarations
// also become resources in the registry, so the restored set is a superset of `original`.
for id in [image, font] {
assert_eq!(
restored.info(&id).map(|info| info.sources),
original.info(&id).map(|info| info.sources),
"sources for {id:?} did not survive the round-trip"
);
assert_eq!(
restored.info(&id).and_then(|info| info.hash.copied()),
original.info(&id).and_then(|info| info.hash.copied()),
"resolved hash for {id:?} did not survive the round-trip"
);
}
}
/// A resource present in the runtime cache but not referenced by any node is *not* snapshotted into the
/// storage registry. This is the orphan case: undoing an image paste removes the node but the runtime
/// keeps the resource alive for redo, so a later diff must not see the orphan as a new `AddResource`
/// (which would resurface the undone paste as a phantom interaction). Regression guard for that divergence.
#[test]
fn unreferenced_runtime_resource_is_not_snapshotted() {
use graphene_resource::{DataSource, ResourceHash, ResourceId, ResourceRegistry};
let referenced = ResourceId::new();
let orphan = ResourceId::new();
let mut resources = ResourceRegistry::new();
for id in [referenced, orphan] {
resources.push_source_back(&id, DataSource::Embedded);
resources.resolve(&id, ResourceHash::from(&b"bytes"[..]));
}
// Only `referenced` is wired to a node; `orphan` lingers in the cache (as it would after an undo).
let network = network_referencing_resource(referenced);
let registry = Registry::from_runtime_with_metadata(&network, &crate::NoMetadata, &resources, PeerId(1)).expect("from_runtime failed");
assert!(registry.resources.contains_key(&referenced), "the network-referenced resource must be snapshotted");
assert!(!registry.resources.contains_key(&orphan), "the unreferenced (orphan) resource must not be snapshotted");
}
/// A node-input `TaggedValue::F64` must survive the storage round-trip bit-exact. Inputs are stored as a
/// self-describing `serde_json::Value` (encoded with the registry's MessagePack codec), so this guards
/// against any precision loss in the f64 -> serde_json::Number -> f64 path for a value with a full
/// 17-significant-digit mantissa.
#[test]
fn node_input_f64_round_trips_bit_exact() {
use graph_craft::document::value::TaggedValue;
// A value whose exact f64 bits matter: 1/3-ish with a non-terminating binary expansion.
let precise = 107.33334350585939_f64;
let network = NodeNetwork {
nodes: [(
NodeId(0),
DocumentNode {
inputs: vec![NodeInput::value(TaggedValue::F64(precise), false)],
implementation: DocumentNodeImplementation::ProtoNode(ProtoNodeIdentifier::new("graphene_core::ops::identity::IdentityNode")),
..Default::default()
},
)]
.into_iter()
.collect(),
..Default::default()
};
let (registry, declarations) = to_registry(&network);
let (converted, _) = registry.to_runtime_with_metadata(&declarations).expect("to_runtime");
let input = &converted.nodes.get(&NodeId(0)).expect("node 0").inputs[0];
let NodeInput::Value { tagged_value, .. } = input else {
panic!("expected a value input, got {input:?}")
};
let TaggedValue::F64(actual) = &**tagged_value else {
panic!("expected F64, got {:?}", tagged_value)
};
assert_eq!(actual.to_bits(), precise.to_bits(), "f64 node input drifted: {actual} != {precise}");
}
/// Two storage nodes in one network carrying the same `ORIGINAL_NODE_ID` both map to one runtime ID.
/// Conversion must reject this rather than silently collapse them and drop a node.
#[test]
fn duplicate_runtime_node_id_is_rejected() {
use crate::AttributesWrite;
use crate::TimeStamp;
use crate::to_runtime::ConversionError;
let (mut registry, declarations) = to_registry(&create_simple_network());
// Force both root-network nodes onto the same runtime ID.
for node in registry.node_instances.values_mut() {
node.attributes.set(crate::attr::node::ORIGINAL_NODE_ID, serde_json::json!(7), TimeStamp::ORIGIN);
}
let error = registry.to_runtime_with_metadata(&declarations).expect_err("duplicate runtime ID must error");
assert!(
matches!(error, ConversionError::DuplicateRuntimeNodeId { runtime_id: 7, .. }),
"expected DuplicateRuntimeNodeId, got {error:?}"
);
}
/// A node input referencing a node in a different network can't be remapped to a valid local runtime
/// ID, so conversion must reject it rather than emit a dangling reference.
#[test]
fn cross_network_reference_is_rejected() {
use crate::to_runtime::ConversionError;
use crate::{Network, NodeInput};
let (mut registry, declarations) = to_registry(&create_simple_network());
// `create_simple_network` wires one node's input to another, both in the root network. Find the
// referenced storage ID, then move that node into a fresh second network so the reference crosses
// a network boundary.
let referenced_storage_id = registry
.node_instances
.values()
.flat_map(|node| node.inputs())
.find_map(|slot| match slot.input {
NodeInput::Node { id: node_id, .. } => Some(node_id),
_ => None,
})
.expect("simple network has a node-to-node reference");
let other_network = NetworkId(999);
registry.networks.insert(other_network, Network::default());
registry.node_instances.get_mut(&referenced_storage_id).expect("referenced node exists").network = other_network;
let error = registry.to_runtime_with_metadata(&declarations).expect_err("cross-network reference must error");
assert!(matches!(error, ConversionError::CrossNetworkReference { .. }), "expected CrossNetworkReference, got {error:?}");
}
/// A network's `scope_injections` (key -> (NodeId, Type)) must survive a storage round trip, with the
/// node reference resolved back to the same runtime-local ID it pointed at originally.
#[test]
fn scope_injections_round_trip() {
let mut network = create_simple_network();
network.scope_injections.insert("editor-api".to_string(), (NodeId(0), concrete!(u32)));
let (registry, declarations) = to_registry(&network);
let (converted, _) = registry.to_runtime_with_metadata(&declarations).expect("to_runtime");
let (node_id, ty) = converted.scope_injections.get("editor-api").expect("scope injection must survive the round trip");
assert_eq!(*node_id, NodeId(0), "the injection's node reference must resolve back to its original runtime ID");
assert_eq!(*ty, concrete!(u32), "the injection's type must be preserved");
}
/// A stored scope injection whose node reference no longer resolves (node removed, or moved to another
/// network) must error rather than emit an injection pointing at a nonexistent runtime node.
#[test]
fn dangling_scope_injection_is_rejected() {
use crate::AttributesWrite;
use crate::TimeStamp;
use crate::to_runtime::ConversionError;
let (mut registry, declarations) = to_registry(&create_simple_network());
// Store an injection pointing at a storage ID that no node carries, leaving the reference dangling
// while the rest of the graph stays valid. The root network is whichever one holds the nodes.
let root_network_id = registry.node_instances.values().next().expect("simple network has nodes").network();
let injections: HashMap<String, (crate::NodeId, Type)> = [("editor-api".to_string(), (crate::NodeId(u64::MAX), concrete!(u32)))].into_iter().collect();
registry
.networks
.get_mut(&root_network_id)
.expect("root network exists")
.attributes
.set_serialized(crate::attr::network::SCOPE_INJECTIONS, &injections, TimeStamp::ORIGIN)
.expect("serialize injections");
let error = registry.to_runtime_with_metadata(&declarations).expect_err("dangling scope injection must error");
assert!(matches!(error, ConversionError::DanglingScopeInjection { .. }), "expected DanglingScopeInjection, got {error:?}");
}
#[test]
fn cyclic_network_reference_is_rejected() {
use crate::to_runtime::ConversionError;
use crate::{Implementation, Network, Node};
// A runtime `NodeNetwork` embeds children by value and so can't be cyclic; the cycle only exists
// in the storage form, where networks reference each other by `NetworkId`. Build it directly:
// the root network holds a node whose implementation is the child network, whose own node points
// back at the root, closing the loop.
let child_network_id = NetworkId(1);
let mut registry = Registry::default();
registry.networks.insert(crate::ROOT_NETWORK, Network::default());
registry.networks.insert(child_network_id, Network::default());
registry.node_instances.insert(
crate::NodeId(0),
Node {
implementation: Implementation::Network(child_network_id),
inputs: Vec::new(),
attributes: crate::Attributes::default(),
network: crate::ROOT_NETWORK,
},
);
registry.node_instances.insert(
crate::NodeId(1),
Node {
implementation: Implementation::Network(crate::ROOT_NETWORK),
inputs: Vec::new(),
attributes: crate::Attributes::default(),
network: child_network_id,
},
);
let error = registry.to_runtime_with_metadata(&crate::Declarations::new()).expect_err("cyclic network reference must error");
assert!(matches!(error, ConversionError::CyclicNetwork(_)), "expected CyclicNetwork, got {error:?}");
}
@@ -0,0 +1,380 @@
use std::borrow::Cow;
use std::collections::HashMap;
use core_types::memo::MemoHash;
use core_types::uuid::NodeId as RuntimeNodeId;
use graph_craft::document::value::TaggedValue;
use graph_craft::document::{DocumentNode, DocumentNodeImplementation, NodeInput as GraphCraftNodeInput, NodeNetwork};
use graph_craft::{ProtoNodeIdentifier, Type, concrete};
use rustc_hash::{FxHashMap, FxHashSet};
use crate::attr::*;
use crate::metadata_source::{InputMetadataEntry, NetworkMetadataEntry, NodeMetadataEntry};
use crate::{AttributesRead, Implementation, NetworkId, Node, NodeId, NodeInput, Position, ProtoNode, ROOT_NETWORK, Registry, ResourceId};
#[derive(Debug, thiserror::Error)]
pub enum ConversionError {
#[error("Network {0} not found")]
NetworkNotFound(NetworkId),
#[error("Node {0} not found")]
NodeNotFound(NodeId),
#[error("ProtoNode declaration {0} not found in provided declarations")]
DeclarationNotFound(ResourceId),
#[error("Deserialization error: {0}")]
DeserializationError(String),
#[error("Network {network} has two nodes mapping to runtime ID {runtime_id}")]
DuplicateRuntimeNodeId { network: NetworkId, runtime_id: u64 },
#[error("Network {network} references node {referenced}, which lives in a different network")]
CrossNetworkReference { network: NetworkId, referenced: NodeId },
#[error("Scope injection {key:?} in network {network} references node {referenced}, which is missing or in a different network")]
DanglingScopeInjection { network: NetworkId, key: String, referenced: NodeId },
#[error("Network {0} is reachable from itself through nested implementations, forming a cycle")]
CyclicNetwork(NetworkId),
}
/// Resolved proto-node declarations, keyed by the `ResourceId` that `Implementation::ProtoNode`
/// references. The caller resolves these from its byte store (`ResourceId` → `ResourceHash` →
/// stored `ProtoNode` bytes) before converting, since `document-graph-storage` holds only references.
pub type Declarations = std::collections::HashMap<ResourceId, ProtoNode>;
impl Registry {
/// Returns the network plus per-node metadata entries (one per node carrying any `ui::*` attribute).
pub fn to_runtime_with_metadata(&self, declarations: &Declarations) -> Result<(NodeNetwork, Vec<NodeMetadataEntry>), ConversionError> {
let (network, node_entries, _) = self.to_runtime_with_full_metadata(declarations)?;
Ok((network, node_entries))
}
/// Like `to_runtime_with_metadata` but also returns per-network entries (navigation, previewing).
/// Used by the editor's full-rebuild path.
pub fn to_runtime_with_full_metadata(&self, declarations: &Declarations) -> Result<(NodeNetwork, Vec<NodeMetadataEntry>, Vec<NetworkMetadataEntry>), ConversionError> {
let mut node_metadata = Some(Vec::new());
let mut network_metadata = Some(Vec::new());
// Group nodes by their owning network in one pass, so each `convert_network` call (one per
// network, including nested ones) takes its node list by lookup instead of rescanning the whole
// flat `node_instances` map, which would be quadratic on graphs with many networks.
let mut nodes_by_network: FxHashMap<NetworkId, Vec<(NodeId, &Node)>> = FxHashMap::default();
for (&global_id, node) in &self.node_instances {
nodes_by_network.entry(node.network).or_default().push((global_id, node));
}
let context = ConversionContext {
registry: self,
declarations,
nodes_by_network,
};
// Reject cycles up front so the recursive conversion below can assume the network reference
// graph is acyclic and never blow the stack on a self-referential `Implementation::Network`.
detect_network_cycle(&context, ROOT_NETWORK)?;
let network = convert_network(&context, ROOT_NETWORK, &[], &mut node_metadata, &mut network_metadata)?;
Ok((network, node_metadata.expect("seeded above"), network_metadata.expect("seeded above")))
}
/// Rebuild the runtime [`ResourceRegistry`](graphene_resource::ResourceRegistry) from the stored
/// `resources`. Each entry's source chain is restored in priority order (the chain is kept
/// sorted by key) with bodies decoded from their type-erased `serde_json::Value` form back to
/// `DataSource`; the resolved hash, if any, is restored last. Inverse of `convert_resources` in
/// `from_runtime`.
pub fn to_resource_registry(&self) -> Result<graphene_resource::ResourceRegistry, ConversionError> {
let mut registry = graphene_resource::ResourceRegistry::new();
for (id, entry) in &self.resources {
for (_, source) in &entry.sources {
let decoded: graphene_resource::DataSource = serde_json::from_value(source.source.clone()).map_err(|error| ConversionError::DeserializationError(error.to_string()))?;
registry.push_source_back(id, decoded);
}
if let Some(hash) = entry.hash {
registry.resolve(id, hash);
}
}
Ok(registry)
}
}
/// Immutable shared context threaded through the recursive conversion. `nodes_by_network` is the
/// one-pass grouping of `registry.node_instances` by owning network, so each network's nodes are an
/// O(1) lookup rather than a full rescan.
struct ConversionContext<'a> {
registry: &'a Registry,
declarations: &'a Declarations,
nodes_by_network: FxHashMap<NetworkId, Vec<(NodeId, &'a Node)>>,
}
/// Converts a single network. Recurses through `Implementation::Network` owning nodes.
///
/// **ID remapping:** Registry uses globally hashed IDs; runtime networks need local IDs. We pull
/// the original local ID from `attr::ORIGINAL_NODE_ID` on each node and on each `NodeInput::Node`
/// reference. References only point within the same network, so per-network lookup suffices.
///
/// **Exports:** the storage-side `Vec<ExportSlot>` is sparse (`None` slots are valid). Compacted
/// here into the runtime's dense `Vec<NodeInput>` — slot stability is a storage-side concern.
///
/// `metadata_path` is the owning-node chain naming *this* network (empty for the root).
/// Walk the network reference graph (edges are `Implementation::Network` references between a
/// network and the networks its nodes embed) and reject any cycle, so the recursive `convert_network`
/// can't recurse forever and overflow the stack. Iterative DFS with an explicit stack and a gray set
/// for the active path; a child already on the active path is a back edge, i.e. a cycle.
fn detect_network_cycle(context: &ConversionContext, root: NetworkId) -> Result<(), ConversionError> {
// Networks reachable from `root` that referenced networks, used by an embedded node, are pushed in
// reverse so the natural processing order matches a recursive walk. `Enter`/`Leave` frames let us
// maintain the gray (active-path) set with an explicit stack.
enum Frame {
Enter(NetworkId),
Leave(NetworkId),
}
let mut stack = vec![Frame::Enter(root)];
let mut on_path: FxHashSet<NetworkId> = FxHashSet::default();
let mut fully_explored: FxHashSet<NetworkId> = FxHashSet::default();
while let Some(frame) = stack.pop() {
match frame {
Frame::Leave(network_id) => {
on_path.remove(&network_id);
fully_explored.insert(network_id);
}
Frame::Enter(network_id) => {
if fully_explored.contains(&network_id) {
continue;
}
if !on_path.insert(network_id) {
return Err(ConversionError::CyclicNetwork(network_id));
}
stack.push(Frame::Leave(network_id));
for &(_, node) in context.nodes_by_network.get(&network_id).map(Vec::as_slice).unwrap_or_default() {
if let Implementation::Network(child) = node.implementation {
stack.push(Frame::Enter(child));
}
}
}
}
}
Ok(())
}
fn convert_network(
context: &ConversionContext,
network_id: NetworkId,
metadata_path: &[RuntimeNodeId],
node_collector: &mut Option<Vec<NodeMetadataEntry>>,
network_collector: &mut Option<Vec<NetworkMetadataEntry>>,
) -> Result<NodeNetwork, ConversionError> {
let network = context.registry.networks.get(&network_id).ok_or(ConversionError::NetworkNotFound(network_id))?;
if let Some(collector) = network_collector.as_mut() {
collector.push(extract_network_metadata(&network.attributes, metadata_path, network_id));
}
let mut nodes: FxHashMap<RuntimeNodeId, DocumentNode> = FxHashMap::default();
for &(global_id, node) in context.nodes_by_network.get(&network_id).map(Vec::as_slice).unwrap_or_default() {
let local_id = node.attributes.get(node::ORIGINAL_NODE_ID).and_then(|v| v.value.as_u64()).unwrap_or(global_id.0);
let runtime_id = RuntimeNodeId(local_id);
if let Some(collector) = node_collector.as_mut()
&& let Some(entry) = extract_ui_metadata(node, metadata_path, runtime_id)
{
collector.push(entry);
}
let doc_node = convert_node(context, node, metadata_path, runtime_id, node_collector, network_collector)?;
// Two storage nodes resolving to the same runtime ID would silently collapse into one on
// insert, dropping a node from the reconstructed graph.
if nodes.insert(runtime_id, doc_node).is_some() {
return Err(ConversionError::DuplicateRuntimeNodeId {
network: network_id,
runtime_id: local_id,
});
}
}
// Input attributes aren't round-tripped for exports — Reflection/Import inputs don't appear there.
let empty_attrs = crate::Attributes::new();
let exports: Vec<GraphCraftNodeInput> = network
.exports
.iter()
.filter_map(|slot| slot.target.as_ref())
.map(|input| convert_input(context.registry, network_id, input, &empty_attrs))
.collect::<Result<Vec<_>, _>>()?;
let scope_injections = read_scope_injections(context.registry, network_id, &network.attributes)?;
Ok(NodeNetwork {
exports,
nodes,
scope_injections,
generated: false,
})
}
/// Rebuild a network's `scope_injections` from its serialized attribute blob, resolving each stored
/// storage node ID back to its runtime-local ID. Mirrors `from_runtime::write_scope_injections`.
fn read_scope_injections(registry: &Registry, network_id: NetworkId, attributes: &crate::Attributes) -> Result<FxHashMap<String, (RuntimeNodeId, Type)>, ConversionError> {
let Some(stored) = attributes.get_typed::<HashMap<String, (NodeId, Type)>>(network::SCOPE_INJECTIONS) else {
return Ok(FxHashMap::default());
};
stored
.into_iter()
.map(|(key, (storage_id, ty))| {
// The injection must point at a node in this same network, like any `NodeInput::Node`.
let referenced = registry.node_instances.get(&storage_id).filter(|node| node.network == network_id);
let Some(referenced) = referenced else {
return Err(ConversionError::DanglingScopeInjection {
network: network_id,
key,
referenced: storage_id,
});
};
let local_id = referenced.attributes.get(node::ORIGINAL_NODE_ID).and_then(|v| v.value.as_u64()).unwrap_or(storage_id.0);
Ok((key, (RuntimeNodeId(local_id), ty)))
})
.collect()
}
/// Returns `None` when the node has no `ui::*` attributes at all so callers don't end up with
/// empty entries for unconverted-from-runtime nodes. `input_metadata` is always sized to match
/// `node.inputs.len()` for a strict slot-by-slot rebuild; empty slots use `InputMetadataEntry::default()`.
fn extract_ui_metadata(node: &crate::Node, network_path: &[RuntimeNodeId], local_id: RuntimeNodeId) -> Option<NodeMetadataEntry> {
let position: Option<Position> = node.attributes.get_typed(node::ui::POSITION);
let is_layer = node.attributes.get_or(node::ui::IS_LAYER, false);
let display_name: Option<String> = node.attributes.get_typed(node::ui::DISPLAY_NAME);
let locked = node.attributes.get_or(node::ui::LOCKED, false);
let pinned = node.attributes.get_or(node::ui::PINNED, false);
let output_names: Vec<String> = node.attributes.get_or_default(node::ui::OUTPUT_NAMES);
let input_metadata: Vec<InputMetadataEntry> = node.inputs.iter().map(|slot| &slot.attributes).map(extract_input_metadata).collect();
let entry = NodeMetadataEntry {
network_path: network_path.to_vec(),
local_id,
position,
is_layer,
display_name,
locked,
pinned,
input_metadata,
output_names,
};
(!entry.is_empty()).then_some(entry)
}
fn extract_network_metadata(attributes: &crate::Attributes, network_path: &[RuntimeNodeId], network_id: NetworkId) -> NetworkMetadataEntry {
NetworkMetadataEntry {
network_path: network_path.to_vec(),
network_id,
reference: attributes.get_typed(node::ui::REFERENCE),
}
}
/// Reassembles `input_data` by scanning every attribute under `ui::input_data::` and stripping the prefix.
fn extract_input_metadata(attributes: &crate::Attributes) -> InputMetadataEntry {
let input_data: HashMap<String, serde_json::Value> = attributes
.iter()
.filter_map(|(key, value)| key.strip_prefix(node::input::ui::DATA_PREFIX).map(|sub_key| (sub_key.to_owned(), value.value.clone())))
.collect();
InputMetadataEntry {
input_name: attributes.get_typed(node::input::ui::NAME),
input_description: attributes.get_typed(node::input::ui::DESCRIPTION),
widget_override: attributes.get_typed(node::input::ui::WIDGET_OVERRIDE),
input_data,
}
}
fn convert_node(
context: &ConversionContext,
node: &crate::Node,
metadata_path: &[RuntimeNodeId],
runtime_node_id: RuntimeNodeId,
node_collector: &mut Option<Vec<NodeMetadataEntry>>,
network_collector: &mut Option<Vec<NetworkMetadataEntry>>,
) -> Result<DocumentNode, ConversionError> {
let inputs = node
.inputs
.iter()
.map(|slot| convert_input(context.registry, node.network, &slot.input, &slot.attributes))
.collect::<Result<Vec<_>, _>>()?;
// Defaults must match `DocumentNode::default()` (and the `set_if_not_default` calls in `from_runtime`).
Ok(DocumentNode {
inputs,
call_argument: node.attributes.get_or(node::CALL_ARGUMENT, concrete!(core_types::Context)),
implementation: convert_implementation(context, &node.implementation, metadata_path, runtime_node_id, node_collector, network_collector)?,
visible: node.attributes.get_or(node::VISIBLE, true),
skip_deduplication: node.attributes.get_or(node::SKIP_DEDUPLICATION, false),
// Regenerated during compilation; not stored.
context_features: Default::default(),
original_location: Default::default(),
})
}
fn convert_input(registry: &Registry, network_id: NetworkId, input: &NodeInput, input_attributes: &crate::Attributes) -> Result<GraphCraftNodeInput, ConversionError> {
Ok(match input {
NodeInput::Node { id: node_id, index: output_index } => {
let referenced = registry.node_instances.get(node_id).ok_or(ConversionError::NodeNotFound(*node_id))?;
// Runtime references are local to one network. A cross-network reference would remap to a
// local ID that doesn't exist in the current runtime network, so reject it.
if referenced.network != network_id {
return Err(ConversionError::CrossNetworkReference {
network: network_id,
referenced: *node_id,
});
}
let local_id = referenced.attributes.get(node::ORIGINAL_NODE_ID).and_then(|v| v.value.as_u64()).unwrap_or(node_id.0);
GraphCraftNodeInput::Node {
node_id: RuntimeNodeId(local_id),
output_index: *output_index as usize,
}
}
NodeInput::Value { value, exposed } => {
let tagged_value: TaggedValue = serde_json::from_value(value.clone()).map_err(|e| ConversionError::DeserializationError(format!("TaggedValue: {e:?}")))?;
GraphCraftNodeInput::Value {
tagged_value: MemoHash::new(tagged_value),
exposed: *exposed,
}
}
NodeInput::Scope(s) => GraphCraftNodeInput::Scope(s.clone()),
NodeInput::Import { index: import_idx } => GraphCraftNodeInput::Import {
import_type: input_attributes.get_or(node::input::IMPORT_TYPE, Type::Generic(Cow::Borrowed("T"))),
import_index: *import_idx as usize,
},
NodeInput::Reflection => GraphCraftNodeInput::Reflection(
input_attributes
.get_typed(node::REFLECTION_METADATA)
.ok_or_else(|| ConversionError::DeserializationError("Missing reflection_metadata in input_attributes".to_string()))?,
),
NodeInput::Other => return Err(ConversionError::DeserializationError("Cannot convert NodeInput::Other to a runtime input".to_string())),
})
}
fn convert_implementation(
context: &ConversionContext,
implementation: &Implementation,
parent_metadata_path: &[RuntimeNodeId],
owning_runtime_id: RuntimeNodeId,
node_collector: &mut Option<Vec<NodeMetadataEntry>>,
network_collector: &mut Option<Vec<NetworkMetadataEntry>>,
) -> Result<DocumentNodeImplementation, ConversionError> {
Ok(match implementation {
Implementation::ProtoNode(id) => {
let proto = context.declarations.get(id).ok_or(ConversionError::DeclarationNotFound(*id))?;
DocumentNodeImplementation::ProtoNode(ProtoNodeIdentifier::with_owned_string(proto.identifier.clone()))
}
Implementation::Network(net_id) => {
let mut child_path = Vec::with_capacity(parent_metadata_path.len() + 1);
child_path.extend_from_slice(parent_metadata_path);
child_path.push(owning_runtime_id);
DocumentNodeImplementation::Network(convert_network(context, *net_id, &child_path, node_collector, network_collector)?)
}
})
}