Skip to content

WIP: UC Volume raw file sync persistence prototype - #14

Draft
nkarpov wants to merge 1 commit into
mainfrom
wip/fs-persistence-prototype
Draft

nkarpov wants to merge 1 commit into
mainfrom
wip/fs-persistence-prototype

Conversation

@nkarpov

@nkarpov nkarpov commented Feb 27, 2026

Copy link
Copy Markdown
Owner

Summary

This PR captures the current raw-file UC Volume persistence experiment as a reference implementation.

Included

  • terminal type persistence contract in type.json
  • session env wiring for persistence config
  • shared bootstrap restore/sync helpers using databricks fs
  • claude/codex persistence config + launch integration
  • app resource/env wiring for AGENT_STATE_VOLUME

Why WIP / not ready to merge

We are seeing practical concerns with this approach:

  1. Startup latency is high

    • restore is synchronous and per-file
    • large trees (notably .claude/plugins) make CLI startup visibly slow
  2. Operational complexity and correctness risk

    • many individual copy operations
    • harder to reason about consistency and failure behavior
  3. Platform constraint remains

    • no native POSIX mount at /Volumes/... in app container
    • no usable FUSE path (/dev/fuse unavailable)
    • without a native mount, this approach may be fundamentally constrained

Current recommendation

Keep this PR as reference for discussion/iteration only.
Do not merge until we decide whether to:

  • reduce scope and aggressively trim synced files, or
  • switch to a different persistence strategy better aligned with runtime constraints.

Add a terminal-type persistence contract and shared bootstrap helpers that restore/sync agent files to UC Volume storage through databricks fs.

This is an exploratory implementation intended for validation and discussion, not for immediate merge.
@nkarpov nkarpov added the WIP Work in progress label Feb 27, 2026
@nkarpov

nkarpov commented Feb 27, 2026

Copy link
Copy Markdown
Owner Author

Additional context from our investigation:

We also evaluated a snapshot-based strategy (single snapshot.tar.gz restore/sync) before switching to raw files.

Why we did not settle on snapshots:

  1. Poor debuggability / observability

    • hard to inspect what changed between syncs
    • difficult to inspect specific agent artifacts directly in UC Volume
  2. Coarse-grained updates

    • small edits require rewriting/reuploading the whole snapshot
    • less efficient for frequently-changing session state
  3. Recovery and corruption blast radius

    • one damaged/missing snapshot affects all persisted state for that type
    • less graceful than file-level partial recovery

So while snapshots reduce operation count, they trade away inspectability and incremental behavior in ways that hurt this use case.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

WIP Work in progress

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant