One storage service¶
Scope: runtime (persistence/, audit/, governed_memory/, orchestration/, artifacts/,
fleet/, _workspace_runtime, serve, CLI), schema (the storage block)
Design reference: §14 (runtime architecture), §8.3 (append-only audit). Supersedes the
per-caller resolution introduced piecemeal since postgres-backend.md.
Status: draft
Goal¶
One service that answers "which store do I use, and how do I open it" for every component, so a
workspace's storage: block means the same thing everywhere and cannot be bypassed.
Non-goals¶
- Not a new backend. SQLite and Postgres, as today. What changes is who decides.
- Not an ORM or a repository layer. Components keep their own table modules and queries; the service hands out a resolved engine, not a data-access API.
- Not changing the fleet store's separation. Enrollment credentials stay in their own database by design; that separation becomes explicit and configured rather than hardcoded.
- Not automatic data migration. Moving existing rows is a separate, opt-in command (see "Open questions").
The problem: six resolutions, four of them wrong¶
| store | constructed at | honours storage.runtime? |
|---|---|---|
| serve main store | create_store(path, workspace.raw) |
yes |
swarmkit orchestrator |
_cmd_orchestrator._resolve_saga_store_url |
yes |
swarmkit pipeline * |
_cmd_pipeline.py:29 — hardcoded sqlite:///…/store.sqlite |
no |
| audit provider | audit/_store.py:165 — _resolve_backend(root), no workspace config |
no |
| governed memory | governed_memory/_store.py:86 — hardcoded sqlite:///…/store.sqlite |
no |
| checkpoints | _workspace_runtime.py:355 — hardcoded SqliteSaver |
no |
Observed on a workspace declaring postgres for all three storage blocks, with the URL confirmed
present in the environment: Postgres held twelve tables at zero rows, while the local
store.sqlite held the entire operational history — 11 sagas, 28 pipeline artifacts, 18 usage
records, and a governed-memory store with its reconciliation change-log. audit_events and
governed_memory did not exist in Postgres at all, because those two providers had never opened a
connection to it.
Three compounding defects, all silent:
- Hardcoded paths. Three components never consult configuration at all.
- A resolver called without the config.
audit/_store.pypasses noworkspace_raw, so it can only see environment variables — and the environment path keys offSWARMKIT_STORE_BACKEND, so aSWARMKIT_STORE_URLbeginningpostgresql://is ignored on its own. A URL names its backend unambiguously; requiring a second variable to believe it is a trap. - Schema promises nothing reads.
storage.audit.backendandstorage.checkpoints.backendare enum-validated (sqlite | postgres, plusagtfor audit) and read by nothing.storage.checkpoints.backend: postgresis not merely unread but unimplementable today —langgraph.checkpoint.postgresis not an installed dependency.swarmkit validateaccepts all of it.
The severity ordering is not the obvious one. A saga in the wrong database is recoverable — re-run the pipeline. Governed memory is not: it is the one store whose value is cumulative, and it was writing to a file no other process opens. A learning loop whose output is unreachable has not learned anything shareable.
The service¶
class StorageService:
"""The single source of truth for where data lives. Nothing else constructs a store."""
@classmethod
def for_workspace(cls, root: Path, workspace_raw: Any = None) -> StorageService: ...
# One resolved target per kind, with its source recorded for the startup report.
def target(self, kind: StoreKind) -> StoreTarget: ... # (backend, url, source)
# Engines are cached per URL: one pool per database, not one per component.
def engine(self, kind: StoreKind) -> Engine: ...
def store(self) -> Store: ... # jobs, conversations, usage
def saga_store(self) -> SqlSagaStore: ...
def artifact_store(self) -> ArtifactStore: ...
def audit_provider(self) -> AuditProvider: ...
def memory_store(self) -> GovernedMemoryStore: ...
def checkpointer(self) -> Any: ...
def membership_store(self) -> MembershipStore: ...
def report(self) -> list[str]: ... # one line per kind, for startup logging
StoreKind is an enum — runtime | audit | checkpoints | artifacts | memory | saga | fleet — so a
component asks for a kind, never a path.
Resolution, once¶
Per kind, in order:
SWARMKIT_STORE_BACKEND/SWARMKIT_STORE_URL(orDATABASE_URL) — process-wide override.storage.<kind>when that kind has its own block (audit,checkpoints).storage.runtime— the workspace default for every SQL store.- SQLite under
{workspace}/.swarmkit/.
A URL implies its backend. postgresql://… resolves to postgres without
SWARMKIT_STORE_BACKEND. Setting only a URL is the common case and currently the silent one.
A declared backend that cannot be honoured raises, consistent with the rule established in
1.127.0: refusing to start beats writing to a different database than the one configured. That
covers checkpoints.backend: postgres while langgraph-checkpoint-postgres is absent — with a
message naming the missing extra rather than a generic failure.
Engines are shared¶
Today each component calls make_engine independently, so one workspace on Postgres opens a
connection pool per component. The service caches by resolved URL, so components sharing a target
share a pool. This is a side benefit, not the motivation — but it is the reason engine(kind) is
on the service rather than each component holding its own.
It reports what it chose¶
Startup logs one line per kind:
storage: runtime=postgres (workspace.yaml) audit=postgres (workspace.yaml)
checkpoints=sqlite (default) memory=postgres (workspace.yaml)
artifacts=database→postgres saga=postgres (workspace.yaml)
fleet=sqlite (separate by design)
The absence of this is why the bug survived: every symptom was an empty screen, which reads as "nothing ran" rather than "your data is elsewhere".
It notices the split it just fixed¶
On start, if a kind resolves to a non-SQLite target and a workspace-local SQLite file for that kind exists with rows, warn — naming both locations and the row count. Anyone upgrading into this change has data in the old place, and a silent cutover would look exactly like data loss.
Migration of call sites¶
Every construction site in the table above is replaced by a service call. The hardcoded ones lose
their paths entirely; SqliteStore(workspace_path) stays only as an internal constructor the
service uses, not as a public entry point.
swarmkit pipeline and swarmkit orchestrator build the service from the resolved workspace, so a
CLI invocation and serve agree by construction rather than by both remembering to.
A lint-style test asserts no module outside persistence/ contains sqlite:/// or calls
make_engine — the mechanical guard against this regressing, since it regressed three times.
API shape¶
# serve
storage = StorageService.for_workspace(workspace_path, runtime.workspace.raw)
app.state.storage = storage
app.state.store = storage.store()
app.state.saga_store = storage.saga_store()
# CLI — same two lines, so they cannot diverge
storage = StorageService.for_workspace(workspace, resolve_workspace(workspace).raw)
storage:
runtime: { backend: postgres, url: "${SWARMKIT_STORE_URL}" } # default for every SQL store
audit: { backend: postgres, retention_days: 90 } # inherits runtime's url
checkpoints: { backend: sqlite } # opt a kind back out
Schema change: storage.<kind>.url becomes optional everywhere and inherits storage.runtime.url
when absent — repeating the same URL three times is what made the config look honoured.
Test plan¶
- Unit. Resolution precedence per kind; a URL alone implies its backend; a per-kind block
overrides
runtime; an unimplementable backend raises with the missing extra named; engines are shared per URL and distinct across URLs. - The regression that started this. A workspace declaring Postgres resolves every kind to Postgres — asserted per kind, so a newly added store that forgets the service fails the test.
- The mechanical guard. No
sqlite:///literal and nomake_enginecall outsidepersistence/. - Split detection. A non-SQLite target plus a populated local SQLite file warns, naming both.
- Full pipeline. Live
swarmkit serve+swarmkit orchestrator+swarmkit pipeline emitagainst Postgres, asserting rows land in Postgres and the workspace SQLite stays empty — the exact scenario that failed.
Demo plan¶
just demo-storage — one workspace, two runs:
storage.runtime: postgres, run a topology and a pipeline stage from the CLI, showpipeline_saga,audit_eventsandgoverned_memoryall populated in Postgres and.swarmkit/store.sqliteabsent.- Point
storage.checkpointsatpostgreswithout the extra installed and show the startup refusal naming it, rather than a silent SQLite fallback.
Plus the startup report block, which is the artefact an operator actually reads.
Resolved (implementation, 1.130.0)¶
All four open questions were answered while building this. Two of them differently from the way they were framed.
-
Migrating existing data — built.
swarmkit storage migratecopies the local SQLite rows into the configured Postgres: additive, idempotent (ON CONFLICT DO NOTHING, so a re-run after a partial failure resumes rather than duplicates), and it never deletes the source. The runbook isdocs/site/reference/storage.md. The split-brain warning stays regardless — it is what tells you a migration is needed. -
agtas an audit backend — dropped from the enum. Nothing implemented it; selecting it wrote SQLite. Leaving a third unimplemented value in place would repeat the exact mistake this note exists to fix. -
Postgres checkpointing — added as the
[postgres]extra, andcheckpointsbecame the one store that does not inheritstorage.runtime. This changed during implementation: with inheritance,storage.runtime.backend: postgrespromoted the checkpointer too and the workspace failed at startup on a driver nobody asked for. SettingSWARMKIT_STORE_URLdid the same thing through the environment branch — caught by the demo, not by a test, which is the argument for demos. A LangGraph component with its own install requirement has to be opt-in, so only an explicitstorage.checkpointsblock moves it. -
Fleet separation — reverted to following
storage.runtime. The note leaned toward always separate, and the implementation started there. That was wrong: design 19 Q4 already decided the membership store follows the main backend, and the shippedcreate_membership_storeimplements it. Regressing a shipped, design-referenced behaviour for tidiness is a worse trade than the consistency was worth. It keeps its own SQLite file, not its own backend.
Two things this note did not anticipate¶
-
${VAR}was never expanded in storage URLs. The workspace schema documented${ENV_VAR} interpolationonstorage.runtime.urland nothing in the runtime implemented it, sourl: ${SWARMKIT_STORE_URL}— the form every deployment doc uses — reached SQLAlchemy as those literal characters. Expansion now happens in the service, which is the only place that reads these URLs. -
storage.artifactswas unusable. The runtime has read it since the bundled orchestrator shipped, but the schema'sadditionalProperties: falserejected the very key the code looked for, so declaring it failed validation. Now declared.
The operator surface¶
The report is the point, not a side effect: a misconfigured workspace looked identical to an
empty one, and no amount of correct resolution fixes that on its own. It is printed at serve
startup, available as swarmkit storage status and swarmkit system, served at GET /storage and
GET /system, and rendered on the web UI's System page — reachable from the screen that is
empty. GET /system also lists the workspace's own workspace.env.yaml properties, read from the
file on every request so a parameter a new feature adds appears without registration, and the
runtime environment variables from a curated registry (never os.environ, which holds every API
key the process uses). Values are masked by declaration — the reserved secrets: key — with a
name heuristic as the fallback for workspaces written before it existed.