Working with Documents
Documents are Primitive's local-first collaborative storage. A document is a container that holds your data models — synced across devices, shared with other users, and available offline. This guide covers document concepts, data modeling, and CRUD operations.
Document Concepts
Private by Default
Documents belong to a user and are private until explicitly shared. Sharing grants another user a permission level:
| Permission | Can View | Can Edit | Can Share | Can Delete |
|---|---|---|---|---|
reader | Yes | |||
read-write | Yes | Yes | ||
owner | Yes | Yes | Yes | Yes |
One carve-out: a read-write editor can also turn link access on or off — per-user grants stay owner-only.
Real-Time Sync
When multiple users have access to the same document, changes sync instantly. Simultaneous edits merge cleanly and automatically, with no conflicts.
Offline Access
Data lives in a local database on the device. Your app works without a network connection — changes queue and sync when connectivity returns.
Connectivity and network mode are separate things. The mode is what your app asked for — auto by default, or pinned with goOffline() / goOnline(). Losing the network in auto pauses the connection and reconnects on its own when the network returns; the mode stays auto throughout. Branch your offline UI on isOnline, not on the mode:
const { mode, isOnline, reason } = client.getNetworkStatus();
client.on("networkMode", ({ isOnline }) => showOfflineBanner(!isOnline));let status = client.networkStatus // .mode, .isOnline, .reason
Task {
for await event in client.stream(for: NetworkModeEvent.self) {
showOfflineBanner(!event.isOnline)
}
}The same event fires for both kinds of transition; reason says which ("user_set", "connectivityLost" or "connectivityRestored").
Opening a document that needs the server — waitForLoad: "network", or the local-else-network mode with no local copy yet — fails fast with a coded error when the network path isn't available (for example DOCUMENT_UNAVAILABLE_OFFLINE while offline, or NETWORK_TIMEOUT when the wait for server state runs out its budget) rather than hanging or returning an empty document. Branch on the error's code and retry the open when connectivity returns.
Local-Only Documents
A document created with localOnly: true never reaches the server: the client drops every edit before it's queued to send — this session, any later session, and even after the local copy is evicted and reloaded from its stored metadata. Use it for drafts and on-device data that has no reason to touch the server.
const { metadata } = await client.documents.create({
title: "Draft",
localOnly: true,
});
await client.documents.open(metadata.documentId, {
waitForLoad: "local",
enableNetworkSync: false,
});let result = try await client.documents.create(
options: CreateDocumentOptions(title: "Draft", localOnly: true)
)
let documentId = result.metadata?["documentId"]?.stringValue
if let documentId {
_ = try await client.documents.open(
documentId,
options: OpenDocumentOptions(waitForLoad: .local, enableNetworkSync: false)
)
}Open a local-only document locally, with network sync turned off. In JavaScript, asking one to wait on the network instead — waitForLoad: "network", or leaving enableNetworkSync at its default of true — fails fast with LOCAL_ONLY_UNSUPPORTED_OPTION rather than hanging on a server that has no record of the document.
Size Guidelines
An ordinary document works best under ~10 MB — for most apps, thousands of records. It is held whole in memory on both ends, so its size is what it costs to open, sync and keep resident. For data you expect to outgrow that — multi-year records, imports — create a large document instead of splitting it across documents: it is chosen at creation, never migrated, and not every client can open one.
Defining Models
Models define the typed shape of the data your app stores in documents — like Task, Project, or Contact. All document interaction happens through models: you author them in one TOML file, codegen produces the typed classes/structs, and every read and write goes through those types.
| Web (Vue) | iOS (SwiftUI) | |
|---|---|---|
| Schema file | src/models/models.toml | Sources/<App>/Models/models.toml |
| Codegen | pnpm codegen (run after editing) | Automatic on build (./run-ios.sh regenerates first; swift build runs the SPM plugin) |
| Output | *.generated.ts classes + @/models barrel | PrimitiveModel structs in Models/Generated/ |
The TOML dialect is identical on both platforms — the same schema file works for web and iOS clients sharing an app.
Documents vs. databases
This models-and-codegen loop is how documents work. Server-side databases are different: your app reaches them through a server function, which reads and writes records through a handle typed from the database type's schema in your config tree — see Working with Databases.
Why TOML + Codegen
Defining models in TOML, with codegen producing the TypeScript classes, gives you:
- Reviewable diffs — your data model lives in one config file, versioned alongside your code. A schema change is one diff, not many.
- Strong typing for free — the generator emits typed
*Attrsinterfaces, model classes, and traversal methods for declared relationships. You can't typo a field name. - Auto-registration — the generated
src/models/index.tsbarrel registers every model withjs-baoexactly once at app startup. Nothing to remember to wire up. - Sync validation at boot — the barrel checks that
models.tomland the generated classes match. If they're out of sync, the app throws at startup and tells you to re-run codegen.
The Authoring Loop
A typical model change is three small steps: edit the TOML, regenerate, use the types.
# 1. Edit src/models/models.toml
# 2. Regenerate:
pnpm codegen
# 3. Import from @/models and use like any other class# 1. Edit Sources/<App>/Models/models.toml
# 2. Regenerate — codegen is wired into both build paths:
./run-ios.sh # regenerates, then builds + launches (Xcode path)
swift build # the SPM plugin regenerates automatically
# 3. Use the generated model statics (Task.query / save(in:))On web, the codegen script runs npx js-bao-codegen-v2 -i src/models/models.toml -o src/models under the hood. On iOS, swift-bao-codegen writes into Models/Generated/ — committed, and regenerated by every build path, so a schema change is a diff you commit with the change.
Never edit generated files
The codegen output — *.generated.ts + the src/models/index.ts barrel on web, Models/Generated/*.swift on iOS — is overwritten on each run. Make all changes in the TOML. For custom behavior on top of a generated type: on web, define free functions in src/lib/; on iOS, add a companion extension file alongside (not inside) Models/Generated/, e.g. Models/TaskRecord+Extensions.swift — the codegen sweep only touches files carrying its generated banner, so extensions survive each regen.
iOS: every build path regenerates
Editing models.toml and building is enough, whichever way you build: swift build runs the codegen plugin, and the Xcode app target runs the same tool from a pre-build phase, so Xcode's Run button, xcodebuild and CI all compile the current schema. The phase takes models.toml as its input, so builds that don't touch the schema skip it.
Adding a new model is the exception: the new file has to be added to the Xcode project before it compiles, and xcodegen can only add a file that already exists. Run ./run-ios.sh, which does both in order (codegen, then xcodegen generate), or run the pair yourself — bash scripts/codegen.sh then bash scripts/regenerate-project.sh. Until then an Xcode build fails naming the file instead of quietly leaving the type out; its pre-build phase has written the file by then, so scripts/regenerate-project.sh alone is enough after that failure.
Authoring models.toml
Each model is a top-level [models.<name>] block. Fields go under [models.<name>.fields.<fieldName>]. Relationships go under [models.<name>.relationships.<relName>].
Field option names use snake_case in TOML (the loader maps them to camelCase at runtime).
[models.todos.fields.id]
type = "id"
auto_assign = true
indexed = true
[models.todos.fields.title]
type = "string"
indexed = true
[models.todos.fields.completed]
type = "boolean"
default = false
[models.todos.fields.priority]
type = "number"
default = 0
[models.todos.fields.due_date]
type = "date"
[models.todos.fields.tags]
type = "stringset"
max_count = 10Field Types
| Type | TypeScript | Description | Common Options |
|---|---|---|---|
id | string | Unique identifier | auto_assign = true, indexed |
string | string | Text value | indexed, default, unique |
number | number | Numeric value | indexed, default |
boolean | boolean | True/false | default |
date | string | ISO-8601 timestamp string | indexed |
stringset | StringSet | Set-of-strings (tags, categories); cannot be unique | max_count |
Field Options
[models.tasks.fields.id]
type = "id"
auto_assign = true # the client assigns a ULID at save
indexed = true # required for fast lookup by id
[models.tasks.fields.email]
type = "string"
unique = true # single-field uniqueness — enables upsertOn
indexed = true
[models.tasks.fields.priority]
type = "number"
default = 0
[models.tasks.fields.tags]
type = "stringset"
max_count = 10Reserved Field Names
A model may not declare a field named type, nor any field whose name starts with _. Both collide with the storage engine's own columns: type is the internal _type column, which holds the model name, so a filter on a declared type field would match the model name instead of the value the record stores — while the record still projects the value you saved. id is not reserved; it's the record's primary key.
Codegen refuses a schema that declares one, naming the block and writing nothing:
[models.todos.fields.type]: Field 'type' is reserved (maps to internal _type column in queries)A server function push carries models/models.toml inside the version, and refuses it there for the same reason, with a 400 carrying the same sentence.
The record routes refuse the key itself, whether or not a schema declares it. POST .../documents/:documentId/records/:model, PATCH .../records/:model/:recordId and POST .../records/bulk answer 400 RESERVED_FIELD_NAME and write nothing; a bulk blob is all-or-nothing, so one bad record refuses the whole body. So primitive documents records save <document-id> <model> --data '{"type":"expense"}' is an error rather than a record no filter can find.
Already declaring type?
A schema the server already stored still parses, and stored records keep their values, so nothing already stored is lost. You meet the refusal the next time you run codegen or push — and straight away on any record write that carries the key. Rename the field (kind, category, status) and regenerate — then migrate the data separately, because renaming the declaration moves no stored value. Records written under the old name still carry it (primitive documents records query <document-id> <model> --json shows it), so copy each one to the new field before dropping the old name from your code.
Unique Constraints
Two ways to enforce uniqueness:
Single-field uniqueness — set unique = true on the field. This also enables upsertOn for that field on save.
[models.users.fields.email]
type = "string"
unique = true
indexed = trueComposite (multi-field) uniqueness — declare a named constraint with [[models.<name>.unique_constraints]].
[[models.categories.unique_constraints]]
name = "name_parentId"
fields = ["name", "parentId"]The constraint name is what you pass to upsertByUnique / findByUnique at runtime.
Stringset fields cannot be unique. unique applies to scalar fields, and a composite constraint may name only scalar fields. No writer can build a consistent key from a set of strings, so the platform refuses the declaration wherever a schema is declared: defineModelSchema, loadSchemaFromTomlString, every generated model barrel and a JsBaoClient given schemaToml throw UniqueStringsetError, and codegen (JavaScript and Swift) writes nothing, naming the model and the field:
Model "posts": field "tags" is a stringset and cannot be unique. A unique constraint applies to scalar fields only.A compound constraint that names one reads Model "posts": unique constraint "title_tags" names the stringset field "tags", which cannot be unique. followed by the same rule. primitive config push refuses the tree in its preflight, before it changes anything, when models/models.toml declares one; a function push and the database type routes answer 400 UNIQUE_ON_STRINGSET. In Swift, TomlSchemaLoader throws .uniqueOnStringset, and a PrimitiveSchema built in code registers but its first write throws a JsBaoError with code .invalidArgument and the same sentence.
Already declaring one?
Remove unique = true from the stringset field, or the stringset field from the compound constraint, and regenerate. An app whose generated models declare the combination throws at startup until you regenerate. A constraint a document already recorded is ignored by the server — records sharing a member all save, and reads and writes to every other field are unchanged — so nothing stored is lost.
Relationships
Declare relationships in TOML and codegen emits typed traversal methods on the generated interfaces.
# Author hasMany Posts
[models.authors.relationships.posts]
type = "hasMany"
model = "posts"
related_id_field = "authorId"
order_by_field = "createdAt"
order_direction = "DESC"
# Post refersTo Author
[models.posts.relationships.author]
type = "refersTo"
model = "authors"
related_id_field = "authorId"After pnpm codegen, the generated interfaces include typed traversal methods:
const author = await Author.find(authorId);
if (!author) return;
// hasMany: author.posts() returns a PaginatedResult — rows are on `.data`
const posts = await author.posts();
const firstPost = posts.data[0];
// refersTo: post.author() returns the parent record (or null)
const backRef = await firstPost.author();guard let author = try Author.find(authorId) else { return }
// hasMany: author.posts() returns a plain array, ordered per the relationship
let posts = try author.posts()
guard let firstPost = posts.first else { return }
// refersTo: post.author() returns the parent record (or nil)
let backRef = try firstPost.author()For many-to-many links, use hasManyThrough. It traverses a join model — a record that points at each side — so a post can carry many tags and a tag can belong to many posts:
# Post hasManyThrough Tags (via the postTags join model)
[models.posts.relationships.tags]
type = "hasManyThrough"
model = "tags"
join_model = "postTags"
join_model_local_field = "postId"
join_model_related_field = "tagId"join_model_local_field is the join record's field pointing back at the source; join_model_related_field points at the target. Add join_model_order_by_field / join_model_order_direction to order the join records (they default to id ascending). Traversal accepts an optional page size and cursor, paging through the join records the same way Sorting and Pagination pages a query:
const post = await Post.find(postId);
if (!post) return;
// First page of this post's tags, ordered by the join model.
const page1 = await post.tags({ limit: 20 });
const firstTag = page1.data[0];
// Carry the cursor forward for the next page.
let rows = page1.data;
if (page1.nextCursor) {
const page2 = await post.tags({ limit: 20, afterCursor: page1.nextCursor });
rows = page2.data;
}guard let post = try Post.find(postId) else { return }
// Every tag linked to this post, ordered by the join model.
let allTags = try post.tags()
// Or page the join leg with a cursor.
let page1 = try post.tags(limit: 20)
let firstTag = page1.data.first
if let cursor = page1.nextCursor {
let page2 = try post.tags(limit: 20, afterCursor: cursor)
_ = page2
}Iterating on the Schema
You're free to evolve the schema as the app grows. A few rules of thumb:
- Add new fields freely — a document's data isn't a rigid schema, so adding an optional field is a no-op for existing records.
- Adding
defaultdoesn't backfill — existing records keep their absent values;defaultonly applies to records created after the change. Read sites should treat the field as optional. - Don't remove fields — the underlying document data is still there. Mark unused fields with a TOML comment instead, and stop reading them.
- Renaming a field is a breaking change — pick a new field with a new name, write to both during a migration window, then stop writing the old one. Add
defaultonly for the new field. - Adding a
unique = trueconstraint to an existing field can fail at save time if existing records collide. Inspect the data before tightening uniqueness.
The index.ts barrel will throw at startup if models.toml and the generated classes drift apart. If you see that error, run pnpm codegen and commit the regenerated files.
Registering a Model at Runtime
Most apps know every model they use before the client starts, and the models list the client is built with covers them. When the shape arrives later — a plugin that brings its own type, a tenant-supplied schema, an import that defines its own columns — register the model on the running client instead:
import { defineModelSchema, createModelClass } from "js-bao";
const Tag = createModelClass({
schema: defineModelSchema({
name: "tag",
fields: {
id: { type: "id", autoAssign: true, indexed: true },
name: { type: "string", indexed: true },
},
}),
});
await client.registerModel(Tag);
// Documents the client already has open answer for the new model straight away.
const { data } = await Tag.query({});What the call does and what it promises:
- Documents already open are covered. The model is initialized for each of them, so a save against a document opened long before the registration works, and a
query()returns the records those documents already hold. - It is per client. The model becomes visible on the client you called it on and on no other. Each client that needs it registers its own class.
- Calling it twice is fine. A repeat registration of the same class resolves without redoing the work or disturbing the rows.
- It is retryable. If a document's initialization fails, the error names that document, the model stays registered, and calling
registerModelagain finishes only the documents that were left incomplete.
It is refused, with a message naming what to change, on a destroyed client, for a class with no name, for a different class under a name the client already holds, and for a class that belongs to another live client.
CRUD Operations
Create
A record is created and saved locally first, then synced in the background.
const task = new Task({
title: "Review pull request",
priority: 2,
dueDate: new Date().toISOString(),
});
await task.save();let task = try Task(
id: UUID().uuidString,
title: "Review pull request",
priority: 2,
dueDate: ISO8601DateFormatter().string(from: Date())
).save(in: documentId)
_ = taskIn single-document mode a JavaScript save() targets the active document; otherwise pass { targetDocument }. Choosing a document for a save never changes what a query sees — reads always span every open document, so scope them with { documents } (see Read) when more than one document is open.
Read
Find by id, query with filters, get the first match, or count.
// Find one by id
const task = await Task.find("task-id");
// Query with filters — returns a PaginatedResult; rows are on `.data`
const urgent = await Task.query({ priority: { $gte: 2 }, completed: false });
const rows = urgent.data;
// First match (with a sort)
const topTask = await Task.queryOne({ completed: false }, { sort: { priority: -1 } });
// Count
const remaining = await Task.count({ completed: false });// Find one by id
let task = try Task.find("task-id")
// Query with filters
let urgent = try Task.query(["priority": ["$gte": 2], "completed": false])
// First match (with a sort)
let topTask = try Task.query(
["completed": false],
options: QueryOptions(sort: ["priority": -1])
).first
// Count
let remaining = try Task.count(["completed": false])Update
const task = await Task.find(taskId);
if (task) {
task.completed = true;
await task.save();
}if var task = try Task.find(taskId) {
task.completed = true
// Only `completed` is written — fields you didn't assign are left alone,
// so another device's concurrent edit to them survives. Assign the result
// back: it's the saved record, with no pending changes left to re-write.
task = try task.save(in: documentId)
}Both look the record up first, then apply the change.
Delete
const task = await Task.find(taskId);
if (task) {
await task.delete();
}if let task = try Task.find(taskId) {
try task.delete(in: documentId)
}Upsert by Natural Key
Save-or-update by a unique field (such as email) without knowing the existing record's id. The field must have a single-field unique constraint.
const user = new AppUser({ email: "alice@example.com", name: "Alice" });
// Creates a new record, or merges into the existing one with that email.
await user.save({ upsertOn: "email" });let user = AppUser(
id: UUID().uuidString,
email: "alice@example.com",
name: "Alice"
)
// "email" must have a single-field unique constraint in models.toml.
// On merge, the returned record carries the existing record's id.
let resolved = try user.save(in: documentId, upsertOn: "email")Upsert by Named Unique Constraint
Save-or-update by a named constraint — single- or multi-field — declared with [[models.<name>.unique_constraints]]. Use this for composite keys, or any time you match by a constraint other than a single unique = true field. The match values come from the record's constraint fields, so every constraint field must be set; pass targetDocument so a new record has a home if none matches.
// "name_parentId" is the named constraint declared in the schema
// ([[models.categories.unique_constraints]] name = "name_parentId",
// fields = ["name", "parentId"]).
const category = await Category.upsertByUnique(
"name_parentId", // the constraint NAME — not the field list
["Work", "root"], // lookup values, in the constraint's field order
{ name: "Work", parentId: "root", color: "blue" },
{ targetDocument: documentId }, // required if a new record is created
);let category = Category(
id: UUID().uuidString,
name: "Work",
parentId: "root",
color: "blue"
)
// "name_parentId" is the named constraint declared in models.toml. `mode`
// defaults to .either (create-or-update); pass .mustExist / .mustNotExist to
// require a single path.
let resolved = try category.upsertByUnique("name_parentId", in: documentId)iOS semantics
Task.query(...), queryOne, count, aggregate, Task.find(_:), and Task.findAll() are synchronous throws (no await) — they read the in-process document data and span every open document by default; scope with QueryOptions(documents: [docId]). Writes target one document — save(in:) inserts or updates in place and throws; save(in:upsertOn:) matches by unique field; delete(in:) throws only if the document isn't open. Writes are local-first: visible to local reads on the next line, synced to peers in the background.
Saving an existing record writes only the fields you assigned since you read it, so two devices editing different fields of the same record merge instead of overwriting each other. Inserting a record that isn't in the document yet writes every field. save(in:) returns the record as saved — re-read from the document, with no pending changes left — so a field another device changed while you held your copy, and any schema defaults filled in on insert, are already there. Assign it back (task = try task.save(in: docId)) if you keep using the value afterwards, or call discardChanges() to drop pending edits without writing them. A record you read and didn't change has nothing to write, so saving it into another document that already holds it does nothing; call markAllChanged() first when you mean to copy the whole record over.
Querying
Operators
| Operator | Description | Example |
|---|---|---|
$eq / $ne | Equals / not equals | { status: { $ne: "deleted" } } |
$gt / $gte / $lt / $lte | Comparisons | { priority: { $gte: 2 } } |
$in / $nin | In / not in array | { status: { $in: ["pending", "active"] } } |
$startsWith / $endsWith | String prefix / suffix | { name: { $startsWith: "Project" } } |
$containsText | Case-insensitive contains | { title: { $containsText: "urgent" } } |
$contains | StringSet contains value | { tags: { $contains: "tutorial" } } |
$exists | Field exists | { dueDate: { $exists: true } } |
A server-side query — the records endpoints or a server function's handle — caps a single $in / $nin list at 1,000 values, counted per list, and refuses a longer one with 400 and code: "QUERY_IN_LIST_TOO_LARGE" naming the field and the count. The local replica evaluates a filter in the browser and has no such bound, so a query that runs offline can still be refused on the server; keep a list inside the cap either way.
The filter is bounded as a whole as well: one statement may bind 100 parameters, which is the platform's limit per statement inside a document's SQLite, and the filter shares that budget with the model name and the cursor's sort values. Two shapes cost one parameter however big they are — an $in list, and a group of equality branches on the same fields, so an $or of a hundred { accountId, month } pairs binds one parameter and pages, sorts and counts like any other filter. A filter with no such group, an $and of a hundred conditions for instance, is refused with 400 and code: "QUERY_FILTER_TOO_MANY_BINDS" naming the count and the 100-parameter bound rather than failing inside SQL; narrow it, fold repeated conditions into an $in or an $or of equality branches, or split the query.
Absent fields
Absent fields (#3166)
$ne, $nin and { field: null } match records that never wrote the field; equality, the comparison operators and $in match only records that carry it. To keep excluding the missing case, put null in the $nin list: { deleted: { $nin: [null, true] } }. Full semantics, including $exists and how a stored null differs by path, are in the Agent Guide to Primitive Documents' Absent fields section (AGENT_GUIDE_TO_PRIMITIVE_DOCUMENTS.md#absent-fields, fetched via primitive guides get documents).
Logical Operators
Combine conditions with $or / $and. The filter shape is identical across clients (a dictionary/object).
const result = await Task.query({
$or: [
{ priority: 3 },
{ dueDate: { $lt: new Date().toISOString() } },
],
});let result = try Task.query([
"$or": [
["priority": 3],
["dueDate": ["$lt": "2026-06-02T00:00:00Z"]],
],
])Sorting and Pagination
Pass a sort and a page size, then carry the cursor forward.
const page1 = await Task.query(
{ completed: false },
{ limit: 20, sort: { priority: -1 } },
);
if (page1.nextCursor) {
const page2 = await Task.query(
{ completed: false },
{ limit: 20, sort: { priority: -1 }, uniqueStartKey: page1.nextCursor },
);
return page2.data;
}let page1 = try Task.queryPaged(
["completed": false],
options: QueryOptions(sortOrder: [("priority", -1)], limit: 20)
)
if let cursor = page1.nextCursor {
let page2 = try Task.queryPaged(
["completed": false],
options: QueryOptions(sortOrder: [("priority", -1)], limit: 20, cursor: cursor)
)
_ = page2
}In Swift, use sortOrder (an ordered list) so the cursor is stable across pages.
Sorting on a field some records don't have
Records are schemaless, so the records of one model need not all carry the field you sort on — and a record can carry it as null. Absent and null are one value for ordering: a record that never wrote the field sorts exactly where a record holding null sorts. Those records come first in an ascending sort ({ priority: 1 }) and last in a descending sort ({ priority: -1 }), which is where SQLite puts nulls in an ORDER BY. The rule is identical on the server and in the JS and Swift clients, so the same records page in the same order everywhere — and the same holds for a database's records (sorting on a field some records don't have).
Paging over such a set visits every matching record exactly once, in both sort directions and both paging directions. Ties are broken by id, which every record has, so page boundaries are stable even when many records share a value (or share having none). hasMore: true always comes with a cursor to follow, so "page until the cursor runs out" and "page until hasMore is false" are the same loop.
Cursors stay opaque — a base64 token, never something to parse or construct. A cursor issued before this ordering was specified keeps working unchanged, so a client holding one does not need to restart its walk.
Aggregations
Group-by with count / avg / sum / min / max, an optional pre-filter, sort, and limit.
const stats = await Task.aggregate({
groupBy: ["category"],
operations: [
{ type: "count" },
{ type: "avg", field: "priority" },
{ type: "sum", field: "estimatedHours" },
],
filter: { completed: false },
sort: { field: "count", direction: -1 },
limit: 10,
});
// Grouping by a stringset field counts per member value (facet):
const tagCounts = await Task.aggregate({
groupBy: ["tags"],
operations: [{ type: "count" }],
});
// Group by whether the set contains a value (membership):
const urgentSplit = await Task.aggregate({
groupBy: [{ field: "tags", contains: "urgent" }],
operations: [{ type: "count" }],
});let stats = try Task.aggregate(AggregateOptions(
groupBy: ["category"],
operations: [
AggregateOperation(type: .count),
AggregateOperation(type: .avg, field: "priority"),
AggregateOperation(type: .sum, field: "estimatedHours"),
],
filter: ["completed": false],
sort: AggregateSort(field: "count", direction: -1),
limit: 10
))
// Grouping by a stringset field counts per member value (facet):
let tagCounts = try Task.aggregate(AggregateOptions(
groupBy: ["tags"],
operations: [AggregateOperation(type: .count)]
))
// Group by whether the set contains a value (membership) — rows carry
// a "has_tags_urgent" key of "true" / "false":
let urgentSplit = try Task.aggregate(AggregateOptions(
groupBy: [.stringSetMembership(field: "tags", contains: "urgent")],
operations: [AggregateOperation(type: .count)]
))A grouped aggregation with exactly one operation returns that operation's bare value per group — for count, sum, avg, min and max alike — so read it as result[group]. Two or more operations keep their count / sum_estimatedHours-style keys. The rule is the same on every surface: the JavaScript model statics above, the document HTTP aggregate endpoint and the primitive documents records aggregate CLI (both of which wrap it in the usual { "result": … } envelope), database aggregates, and connectDoDb bindings.
When you group by a stringset field (like tags), each member value becomes its own group — a tag-facet count. To split records by whether the set contains one specific value instead, use a membership groupBy entry (the urgentSplit call above). One stringset facet field is allowed per aggregation.
An aggregation spans every open document by default. documents narrows it to the ones you name — a single id or a list, with an explicit empty list matching nothing — which is the same option, with the same meaning, that query() and count() take. It is how one call reports across several large documents, and how a model with rows in both an ordinary and a large document is scoped to one kind.
// Every large document that is open, in one statement.
const everyAccount = await Task.aggregate({
groupBy: ["category"],
operations: [{ type: "count" }, { type: "sum", field: "estimatedHours" }],
});
// Narrowed to the documents you name — a single id is accepted too.
const twoAccounts = await Task.aggregate({
groupBy: ["category"],
operations: [{ type: "count" }],
documents: accountLedgers,
});
// `documents` means what it means on `query()`: an explicit empty list
// matches nothing, rather than falling back to everything.
const nothing = await Task.aggregate({
groupBy: ["category"],
operations: [{ type: "count" }],
documents: [],
});// Every open document the shared model spans, in one statement.
let everyAccount = try Task.aggregate(AggregateOptions(
groupBy: ["category"],
operations: [
AggregateOperation(type: .count),
AggregateOperation(type: .sum, field: "estimatedHours"),
]
))
// Narrowed to the documents you name.
let twoAccounts = try Task.aggregate(AggregateOptions(
groupBy: ["category"],
operations: [AggregateOperation(type: .count)],
documents: accountLedgers
))
// An explicit empty list matches nothing, rather than falling back to
// everything — the same rule `QueryOptions(documents: [])` follows.
let nothing = try Task.aggregate(AggregateOptions(
groupBy: ["category"],
operations: [AggregateOperation(type: .count)],
documents: []
))Loading Related Data
Pass include in a query to load related records alongside the results, instead of running a follow-up query for each row. Loading the parent a record points to (refersTo) and the children that point back at it (hasMany):
// refersTo — each Post's Author (the `authorId` FK lives on Post):
const posts = await Post.query({}, {
include: [{ model: "authors", type: "refersTo", sourceField: "authorId", as: "author" }],
});
for (const post of posts.data as (Post & { _related?: { author?: Author } })[]) {
console.log(post.title, post._related?.author?.name);
}
// hasMany — every Post that points back at each Author, newest first:
const authors = await Author.query({}, {
include: [{ model: "posts", type: "hasMany", foreignKey: "authorId", as: "posts", sort: { createdAt: -1 }, limit: 10 }],
});
for (const author of authors.data as (Author & { _related?: { posts?: Post[] } })[]) {
console.log(author.name, (author._related?.posts ?? []).length);
}// refersTo — each Post's Author (the `authorId` FK lives on Post):
let posts = try Post.query(include: [Post.includeAuthor()])
for post in posts {
let author = post.relatedAuthor // Author?
print(post.title, author?.name ?? "—")
}
// hasMany — every Post that points back at each Author, newest first:
let authors = try Author.query(include: [Author.includePosts(limit: 10)])
for author in authors {
let authored = author.relatedPosts // [Post]
print(author.name, authored.count)
}Reacting to Changes
Data can change from sync (another user edited it). Subscribe to keep your UI current — the callback fires on both local and synced remote writes. Always release the subscription when you're done.
const unsubscribe = Task.subscribe(() => {
// re-query and update your UI
});
// later, when you no longer need updates:
unsubscribe();let unsubscribe = Task.subscribe {
// re-query and update your UI
}
// later, when you no longer need updates:
unsubscribe()Most apps don't call subscribe directly in views — each starter template ships a framework helper that wraps it: useJsBaoDataLoader (Vue composable) and BaoDataLoader (SwiftUI loader in PrimitiveApp). Both handle document readiness, debounced reloads, and a loaded/empty/loading phase:
const {
data: todos,
initialDataLoaded,
showSkeleton, // gate loading UI on this: true while the document opens, suppressed on fast warm reloads
reload,
} = useJsBaoDataLoader<{ items: TodoItem[]; total: number }>({
subscribeTo: [TodoItem],
queryParams: computed(() => ({ listId: props.listId, showCompleted })),
documentReady,
async loadData(queryParams) {
const { listId, showCompleted } = queryParams ?? {};
const query = showCompleted ? { listId } : { listId, completed: false };
const result = await TodoItem.query(query, { sort: { order: 1 } });
return { items: result.data, total: result.data.length };
},
});struct TodoListView: View {
@EnvironmentObject var appState: MyAppState
@StateObject private var loader = BaoDataLoader<[TodoItem]>()
var body: some View {
Group {
// Render through `loader.phase`, not `loader.data ?? []`.
// `?? []` collapses "not yet loaded" with "loaded, empty",
// flashing the empty state for ~50ms on every appearance.
switch loader.phase {
case .loading: ProgressView()
case .empty: Text("No todos yet")
case .loaded(let todos): List(todos) { /* row */ }
}
}
// Bind once, from a plain `.task`. Don't conditionally bind on
// doc readiness — set `loader.documentReady` instead: the loader
// fires its initial load when it flips to true, and resets
// `initialDataLoaded` if the doc closes.
.task {
loader.documentReady = appState.selectedDocId != nil
loader.bind(
client: appState.client,
subscribeTo: [.onModel(subscribe: TodoItem.subscribe)]
) { _ in
TodoItem.findAll().sorted { $0.sortOrder < $1.sortOrder }
}
}
.onChange(of: appState.selectedDocId) { _, id in
loader.documentReady = id != nil
}
}
}For code that subscribes to client events directly on iOS: prefer for await event in client.stream(for: SomeEvent.self) inside a .task — the subscription lives as long as the loop, so there is nothing to store or cancel. When you need a callback instead (an ObservableObject wiring itself up, say), client.observeOnMainActor(SomeEvent.self) { ... } runs the handler on the main actor; store the returned EventSubscription on a property or it's dropped immediately, use [weak self] in the closure, and cancel on teardown with sub?.cancel(). When the handler needs the event's own timing (a debug timeline, a latency report), the withDelivery: form passes a second argument alongside the event: delivery.emittedAt is the emit instant — a Date() taken inside the plain handler also includes the hop onto the main actor — and delivery.sequence orders events reliably.
For a single occurrence rather than a stream of them, try await client.nextEvent(SomeEvent.self, timeout: 30) { $0.documentId == docId } returns the first event the predicate accepts. It throws JsBaoError with code .unavailable if the timeout elapses first, and CancellationError if the calling task is cancelled. An event that already fired is not redelivered, so start the wait before triggering the work you're waiting on.
An open document that stops syncing tells you. The documentSyncStateChanged event reports error when a document's sync handshake goes unanswered for its whole budget (10 seconds by default), and synced once it catches up again — the client keeps retrying and rebuilds its connection on its own in between, so all the app has to do is show and clear a banner:
client.on("documentSyncStateChanged", ({ documentId, state }) => {
if (state === "error") {
// The document's sync handshake went unanswered for its whole budget
// (10 s by default). Repeats on each timeout while the document stays
// behind; the client keeps retrying on its own.
ui.markStale(documentId);
} else if (state === "synced") {
// Fires as remote updates are applied — and once more when a document
// that reported an error catches up, so the warning can be cleared.
ui.markSynced(documentId);
}
});// Drop this in a SwiftUI `.task`: the loop runs for as long as the view is on
// screen and unsubscribes when it goes away.
for await event in client.stream(for: DocumentSyncStateChangedEvent.self) {
switch event.state {
case "error":
// The document's sync handshake went unanswered for its whole budget
// (10 s by default). Repeats on each timeout while the document stays
// behind; the client keeps retrying on its own.
ui.markStale(event.documentId)
case "synced":
// Fires as remote updates are applied — and once more when a document
// that reported an error catches up, so the warning can be cleared.
ui.markSynced(event.documentId)
default:
break
}
}Creating and Opening Documents
Create a document with documents.create() — it returns the new document's metadata, including the documentId everything else keys off:
const { metadata } = await client.documents.create({
title: "New Project",
tags: ["workspace"],
});
await client.documents.open(metadata.documentId);let result = try await client.documents.create(
options: CreateDocumentOptions(title: "New Project", tags: ["workspace"])
)
// The generated id lives inside the returned metadata.
let documentId = result.metadata?["documentId"]?.stringValue
if let documentId {
// Open before querying or writing — a write to an unopened document throws.
_ = try await client.documents.open(documentId)
}Create-then-requery is safe (JavaScript)
me.ownedDocuments reads a server index that is only eventually consistent, so for a short window after create() a raw server read can omit the document you just made. The JavaScript client bridges that window — it merges a just-created owned document into the me.ownedDocuments result from its local cache — so a list that rebuilds itself on navigation still includes the new document. (In Swift, guard a freshly-created document against a reconcile prune with a pendingCreateIds set instead.)
A document must be open before you can read or write the data inside it — queries and saves only see open documents:
await client.documents.open(documentId);
const result = await Task.query({}, { documents: documentId });_ = try await client.documents.open(documentId)
let result = try Task.query([:], options: QueryOptions(documents: [documentId]))The document is ready as soon as open() resolves — await it before the first query and show a loading state until then. open() is idempotent, so re-opening an already-open document is a no-op. A write to a document that is not open throws rather than being applied, so open a just-created document before writing to it.
Only keep open the documents you actually need — the ones you're querying or want real-time updates from. Every open document is synced continuously, and queries span all open documents by default, so an unneeded open document costs sync traffic and widens query results. Close documents you're done with:
// Close and stop syncing
await client.documents.close(documentId);
// Close and remove the local cached copy (safe: skipped if not fully synced)
const { evicted } = await client.documents.close(documentId, {
evictLocal: true,
});
if (!evicted) {
// Server was not fully in sync — local copy was retained
}// Close and stop syncing
await client.documents.close(documentId)
// Close and remove the local cached copy. `close` returns a
// CloseDocumentResult — inspect `.evicted` when eviction matters.
let result = await client.documents.close(
documentId,
options: CloseDocumentOptions(evictLocal: true)
)
if !result.evicted {
// Server was not fully in sync — the local copy was retained.
}Ensuring Exactly One Document with Aliases
Some apps need exactly one document of a given kind — a personal app's "the user's data". Creating it with an alias guarantees one and only one exists, even when several devices race to create it: getOrCreateWithAlias resolves-or-creates atomically:
const result = await client.documents.getOrCreateWithAlias({
title: "My Data",
alias: { scope: "user", aliasKey: "default-doc" },
});
await client.documents.open(result.documentId);let result = try await client.documents.getOrCreateWithAlias(
options: GetOrCreateWithAliasOptions(
alias: AliasRef(scope: .user, aliasKey: "default-doc"),
title: "My Data"
)
)
_ = try await client.documents.open(result.documentId)Why getOrCreateWithAlias?
Splitting this into a resolve followed by a create looks fine but has a race: two devices onboarding at the same moment both fall into the create branch and you lose one of the docs. getOrCreateWithAlias is a single atomic server-side upsert.
Opening Documents on iOS
On iOS, the canonical place to open documents is your PrimitiveAppState subclass — open in connectClient():
@MainActor
final class MyAppState: PrimitiveAppState {
override func connectClient() async {
await super.connectClient()
guard let client else { return }
// Pre-register models so every open document is mirrored into the
// client's shared store (and listed in the Debug Inspector).
client.registerModels([TaskRecord.self])
let result = try? await client.documents.getOrCreateWithAlias(
alias: DocumentAlias(scope: .user, aliasKey: "library"),
title: "Library"
)
guard let id = result?.documentId else { return }
await selectDocumentAwaiting(id)
}
}There is no per-document model binding: reads go through the model's statics (TaskRecord.query(...), cross-document by default — scope to one document with QueryOptions(documents: [documentId])), and writes go through record instances (try TaskRecord(...).save(in: documentId)). registerModels([...]) at connect time is optional but recommended — the facade also lazily registers a model on its first read. For per-document setup beyond models, override the onDocumentOpened(doc:documentId:) hook.
Common Document Usage Patterns
A document is both the unit of sync and the unit of sharing, so "how many documents should my app have?" follows from how the data is shared. Three shapes cover most apps:
One document per user. All of a user's data lives in a single private document, resolved with an alias at sign-in (above) and kept open for the whole session. The UI never mentions documents — users sign in and see their data. Personal task managers, journals, habit trackers.
One document per sharing context. Anything shared as a unit — a project, a company's books, a household shopping list — gets its own document, shared with exactly the people in that context. Users work in one at a time: list them with the me listing APIs, open the selected one, and close it when they switch.
Many documents, queried together. Apps like chat read across several sharing contexts at once — each channel is its own document, all open simultaneously. Tag documents at creation (tags: ["channel"]) so you can fetch the set with a tag-filtered me.ownedDocuments, then open each one. A model query then spans all open documents. To share a set of documents as one unit, use Collections rather than sharing each individually.
Whether a feature belongs in documents at all — versus a server-side database — is its own decision: see Choosing Your Data Model.
Sharing Documents
Documents are private until you grant a permission level (reader | read-write | owner, see the table above) to a user, an email, or a group:
// By user ID
await client.documents.updatePermissions(documentId, {
userId: "user-abc",
permission: "read-write",
});
// By email — works whether or not the recipient is a member yet
await client.documents.updatePermissions(documentId, {
email: "colleague@example.com",
permission: "read-write",
});
// With a group
await client.documents.grantGroupPermission(documentId, {
groupType: "team",
groupId: "engineering",
permission: "read-write",
});// By user ID
_ = try await client.documents.updatePermissions(
documentId: documentId,
params: .user("user-abc", permission: "read-write")
)
// By email — works whether or not the recipient is a member yet
_ = try await client.documents.updatePermissions(
documentId: documentId,
params: .email("colleague@example.com", permission: "read-write")
)
// With a group
_ = try await client.documents.grantGroupPermission(
documentId: documentId,
params: GrantGroupPermissionParams(groupType: "team", groupId: "engineering", permission: "read-write")
)Share by Email
The most common case — you have a colleague's email but don't know (or care) whether they've signed up yet. If the email belongs to an existing user, access is granted immediately. If not, the share waits and applies automatically when they sign up — see Invitations for how that resolution works. Repeated shares to the same recipient are idempotent — the latest permission wins.
Batch shares can mix user IDs and emails:
await client.documents.updatePermissions(documentId, {
permissions: [
{ userId: "user-abc", permission: "read-write" },
{ email: "alice@example.com", permission: "reader" },
{ email: "bob@example.com", permission: "read-write" },
],
});_ = try await client.documents.updatePermissions(
documentId: documentId,
params: .batch([
PermissionAssignment(userId: "user-abc", permission: "read-write"),
PermissionAssignment(email: "alice@example.com", permission: "reader"),
PermissionAssignment(email: "bob@example.com", permission: "read-write"),
])
)To have the platform email the recipient, pass sendEmail: true along with a documentUrl, and make sure your app's base URL is configured — the server uses them to compose the share and accept links.
Share with a Group
Grant document access to everyone in a group (the grantGroupPermission call in the example above). When group membership changes, document access updates automatically.
Who Has Access? (Members + Pending)
Sharing UIs typically show two sections: people who currently have access, and people who've been invited but haven't signed up yet:
// Current members (accepted permission grants)
const members = await client.documents.getPermissions(documentId);
// Pending email invites on this document
const pending = await client.documents.listPendingInvitations(documentId);// Current members (accepted permission grants)
let members = try await client.documents.getPermissions(documentId: documentId)
// Pending email invites on this document
let pending = try await client.documents.listPendingInvitations(documentId: documentId)Removing Someone
One call handles both "currently has access" and "invited but hasn't signed up yet" — pass an email and the server removes the matching member or cancels the pending grant, whichever exists:
await client.documents.removePermission(documentId, userId);
await client.documents.removePermission(documentId, { email: "alice@example.com" });try await client.documents.removePermission(documentId: documentId, .userId(userId))
try await client.documents.removePermission(documentId: documentId, .email("alice@example.com"))Use the email form whenever you don't want to think about whether the target has signed up yet.
Anyone with the Link
Instead of naming each recipient, turn on a shared access level that applies to anyone who has the document — any signed-in user of your app who has the document ID resolves to at least that level without an explicit grant. Set it to reader or read-write:
// Any signed-in app user who has the document ID can now read it, with no
// explicit grant. The level is a floor: a user who already has a higher
// grant keeps it.
await client.documents.setLinkAccess(documentId, "reader");
// Read the current state — the level, plus who last changed it and when.
const state = await client.documents.getLinkAccess(documentId);
console.log(state.linkAccess); // "reader" | "read-write" | null
// Turn it off. Anyone relying on the link loses access immediately;
// explicit grants are untouched.
await client.documents.clearLinkAccess(documentId);// Any signed-in app user who has the document ID can now read it, with no
// explicit grant. The level is a floor: a user who already has a higher
// grant keeps it.
_ = try await client.documents.setLinkAccess(documentId: documentId, level: .reader)
// Read the current state — the level, plus who last changed it and when.
let state = try await client.documents.getLinkAccess(documentId: documentId)
print(state.linkAccess as Any) // .reader | .readWrite | nil
// Turn it off. Anyone relying on the link loses access immediately;
// explicit grants are untouched.
_ = try await client.documents.clearLinkAccess(documentId: documentId)The link level is a floor, not a cap: it only raises a user's access. Someone with a higher explicit grant keeps it, and the floor applies to reading and writing document data — never to sharing or managing the document, so a link viewer can't reshare it or change permissions. Only a document owner, app owner, or read-write editor can turn link access on or off.
A document a user reached only through the link doesn't appear in their shared-documents list — track it in your app if you want a "recently opened" affordance. getLinkAccess reports the current level along with who last changed it and when.
Clearing link access (or lowering the level) takes effect immediately: anyone relying on the link loses access right away, while explicit grants stay intact.
Collections
Sharing scales one document at a time until access maps to a set of documents — a project's files, a course's materials, a client folder. A collection bundles documents so they can be shared as a unit: a permission granted on the collection applies to every document currently in it and to any document added later, additively and without touching each document's own grants.
Finding Documents a User Can Access
There is no single "my documents" list. A user reaches documents through four distinct paths, and you query each one separately — combine them in your UI as needed:
1. Documents they own
Created by the user, or had ownership transferred to them.
// Paginated page — the unified { items, nextCursor } envelope, same shape as
// sharedDocuments():
const page = await client.me.ownedDocuments({
tag: "channel",
returnPage: true,
});
const { items, nextCursor } = page;
// (Without `returnPage`, the JS client returns a flat `DocumentInfo[]` for
// convenience: `const owned = await client.me.ownedDocuments({ tag: "channel" })`.)// A typed `[DocumentInfo]` — each row carries title, permission, tags, …
let owned = try await client.me.ownedDocuments(tag: "channel")
for doc in owned {
print(doc.title, doc.permission)
}2. Documents shared directly with them
Non-owner permission grants — group and collection shares are not here. Each row carries the base document fields plus the share extras (permission, grantedBy, source).
const { items, nextCursor } = await client.me.sharedDocuments({
tag: "channel",
limit: 50,
});
for (const doc of items) {
// Each row carries the base document fields (title, createdAt, …) plus the
// share extras (permission, grantedBy, source).
console.log(doc.title, doc.permission, doc.grantedBy);
}
// `nextCursor` is an opaque pagination token — pass it back as `cursor` for the next page.
if (nextCursor) {
const next = await client.me.sharedDocuments({ cursor: nextCursor });
return next;
}let page = try await client.me.sharedDocuments(limit: 50, tag: "channel")
for share in page.items {
// Each row nests the base document fields (title, createdAt, …) under
// `.document`, alongside the share extras (grantedBy, source).
print(share.document.title, share.document.permission, share.grantedBy)
}
// `nextCursor` is an opaque pagination token — pass it back as `cursor` for the next page.
if let cursor = page.nextCursor {
_ = try await client.me.sharedDocuments(cursor: cursor)
}3. Documents shared via a group
Listed through the group the user belongs to.
const documents = await client.groups.listDocuments("team", "engineering");let documents = try await client.groups.listDocuments(groupType: "team", groupId: "engineering")4. Documents shared via a collection
Listed through a collection the user is a member of.
const { items, nextCursor } = await client.collections.listDocuments(collectionId, {
limit: 50,
});let page = try await client.collections.listDocuments(
collectionId: collectionId,
options: PaginationOptions(limit: 50)
)
let items = page.itemsownedDocuments and sharedDocuments accept tag / limit / cursor for filtering and pagination.
In both clients, ownedDocuments is local-first: when the client already has owned documents cached locally it returns those right away and refreshes from the server in the background, so the next call is fresh. When a screen must show a server-fresh list, ask for one — waitForLoad: "network" in JavaScript, MeOwnedDocumentsOptions(waitForLoad: .network) in Swift. serverTimeoutMs bounds that wait; exceeding it raises a list-timeout error rather than silently returning stale rows.
Keeping a Cached Combination Fresh
Re-querying all four paths on every navigation is the simplest approach, and the right default until it's actually too slow. An app that instead caches the combination — a document switcher that stays mounted, say — keeps it current with syncMetadata() and two events, rather than re-walking every listing on each navigation.
await client.syncMetadata({ scope: "all", background: true });
client.on("documentMetadataChanged", ({ documentId, action, source }) => {
if (action === "deleted" || action === "evicted") removeFromSwitcher(documentId);
else refreshRow(documentId);
});
client.on("permission", ({ documentId, permission }) => {
updatePermissionBadge(documentId, permission);
});try await client.syncMetadata(options: SyncMetadataOptions(background: true))
Task {
for await event in client.stream(for: DocumentMetadataChangedEvent.self) {
if event.action == "deleted" || event.action == "evicted" {
removeFromSwitcher(event.documentId)
} else {
refreshRow(event.documentId)
}
}
}
Task {
for await event in client.stream(for: PermissionEvent.self) {
updatePermissionBadge(event.documentId, event.permission)
}
}syncMetadata() refreshes the local metadata index from the server without opening every document — call it on a timer, or when the switcher becomes visible, to pick up changes that happened while the app was closed. documentMetadataChanged then keeps that index current going forward, firing for title, tag, thumbnail, and permission-cache changes on any document the client knows about. action: "deleted" is the "the document is gone" signal — a peer revoking your access and a peer hard-deleting the document collapse to the same event, with metadata empty. action: "evicted" is a separate, local-only signal (freeing local storage, say) — the document may still be reachable, there's just no cached copy until the next sync. Drop the row from the switcher on either action; the permission event fires whenever the caller's access level for a document changes, independent of the other metadata fields.
Neither event covers a collection membership change — a document added to a collection the user belongs to doesn't fire documentMetadataChanged or permission for that new grant. Refresh collections.listDocuments() on the same timer or navigation trigger you use for syncMetadata(); there's no event to wait on for that path.
Keep the cache keyed by documentId, storing only the fields the switcher renders. Patch or remove a row on documentMetadataChanged, update its permission on the permission event, and re-fetch the collection-derived rows on your timer — that keeps a document switcher accurate without re-walking all four listings on every navigation.
Deciding What to Show
Having four access paths means "everything I can access" is a query result, not necessarily what the app should show. For some apps a user's raw access and the set the app presents are the same thing, and rendering the four listings directly is the right, simplest choice. For others, the app wants a smaller, definite set it controls — two patterns build on the root document, a document automatically created and opened for every signed-in user, exactly one per user, that can never be shared or deleted:
- A curated index. Store document references — an id, plus whatever a switcher renders — as a field on the root document, and drive navigation from that list instead of the union of the four access paths. Now "the documents this user works with" is a list the app and user curate, independent of what the server currently reports the user can reach.
- An app-side acceptance flow. Sharing a document grants access immediately — the platform has no accept/reject step a recipient controls. An app that wants one layers it on top: track an accepted set of document ids on the root document, show a newly-shared document as pending until its id is in that set, and add to the set when the user accepts.
Neither is universal — skip both when shared documents should simply appear, or when access is meant to flow entirely through group or collection membership.
Document Blobs
Documents can carry binary files — images, PDFs, attachments — reached through the document's blob context, with uploads, downloads, and listing hanging off it. Document blobs inherit the sharing of their document: anyone who can read the document can read its blobs, anyone with write access can add or delete them, and there's no separate permission system to manage. See Blobs and Files for uploading, displaying, downloading, listing, and offline caching, and for how document blobs compare to general-purpose blob buckets.
Document Thumbnails and Metadata
Documents carry presentation fields you can update at any time.
await client.documents.update(documentId, {
title: "Q2 Planning",
thumbnailBlobId: blobId, // a blob you uploaded
metadata: { color: "blue", tags: ["plan", "q2"] }, // ≤4KB JSON, replace semantics
});_ = try await client.documents.update(
documentId: documentId,
data: UpdateDocumentData(
title: "Q2 Planning",
thumbnailBlobId: .value(blobId), // a blob you uploaded
metadata: ["color": "blue", "tags": ["plan", "q2"]] // ≤4KB JSON, replace semantics
)
)thumbnailBlobId points at a blob you've already uploaded; the platform makes the thumbnail readable to anyone with access to the document. metadata is a small JSON blob (4KB cap) for UI hints. To clear either field, JavaScript passes null; Swift uses thumbnailBlobId: .clear and metadata: .null. Omitting a field leaves it unchanged.
Document Access Requests
When a user has a document link (or ID) but no permission, they can request access — the Google-Docs-style "Request access" flow. A 403 from client.documents.get(documentId) carries a DOC_ACCESS_DENIED code, and its body includes a canRequestAccess hint when the document accepts requests:
import { JsBaoApiError, JsBaoNetworkError } from "js-bao-wss-client";
try {
await client.documents.get(documentId);
} catch (err) {
if (err instanceof JsBaoApiError && err.code === "DOC_ACCESS_DENIED" && err.body?.details?.canRequestAccess) {
// Show a "Request access" button
} else if (err instanceof JsBaoNetworkError) {
// The request never reached the server — offer a retry
}
}do {
_ = try await client.documents.get(documentId: documentId)
} catch let err as HttpError where err.status == 403 && err.serverCode == "DOC_ACCESS_DENIED" {
// err.body holds the JSON details — check it for "canRequestAccess": true,
// then show a "Request access" button
}JsBaoApiError covers every non-2xx server response. When the request never reaches the server at all — the device is offline, DNS fails, the connection is refused — the client throws JsBaoNetworkError instead. A network failure is retryable, so branch on instanceof JsBaoNetworkError (or the isJsBaoNetworkError guard) to decide whether to retry, rather than parsing the message.
The full flow — submit a request (permission is required), then an owner lists and approves:
// A user with the link requests access
await client.documents.requestAccess(documentId, {
permission: "read-write",
message: "Please add me to this doc",
});
// An owner lists pending requests and approves one
const requests = await client.documents.listAccessRequests(documentId);
await client.documents.approveAccessRequest(documentId, requestId);// A user with the link requests access
_ = try await client.documents.requestAccess(
documentId: documentId,
options: RequestAccessOptions(permission: .readWrite, message: "Please add me to this doc")
)
// An owner lists pending requests and approves one
let requests = try await client.documents.listAccessRequests(documentId: documentId)
_ = try await client.documents.approveAccessRequest(documentId: documentId, requestId: requestId)Owners can also deny, optionally pointing the requester at the document's canonical URL:
await client.documents.denyAccessRequest(documentId, requestId, {
documentUrl: "https://myapp.example/docs/sales-handbook",
});_ = try await client.documents.denyAccessRequest(
documentId: documentId,
requestId: requestId,
options: DenyAccessRequestOptions(documentUrl: "https://myapp.example/docs/sales-handbook")
)The requester gets an email with the outcome either way. To show pending requests in an owner's UI, call listAccessRequests (shown in the flow above) when the view opens.
Behavior worth knowing: requests expire after 30 days, re-requesting while a request for the same document is pending updates that pending request in place, and a resolved request can't be re-resolved.
Moving Documents Between Environments
Export and import are operator commands, not application code — they run from the primitive CLI with an admin login:
# One user's documents into their own directory, then restore them elsewhere
primitive documents export-all --user-id <user-id> --output ./export-of-one-user
primitive documents import ./export-of-one-user --owner user@example.com--overwrite merges; it does not replace. The import uploads the exported state and the server applies it with Y.applyUpdate, so keys only one side holds survive, and a key both sides set resolves by Yjs's own conflict rule — neither the exported nor the existing value is guaranteed to win. One export is installed rather than merged: a large document (documentFormat: 2) travels as a chain that replaces a document's whole history, so the server installs it only into a document holding no records yet — --overwrite will not merge one into a document that already has content.
Document ids are preserved, with one exception: a root document — the single document the server mints per user and points AppUser.rootDocId at. Its export is applied to the target user's root document (chosen from --owner, or the owner the export recorded), and never restored under the exported id, so an app ends up with exactly one root document per user:
- No root document yet:
importcreates it through the same get-or-create sign-in uses and applies the exported state to it. - One already there: merging into it needs
--overwrite; without it the export is skipped, naming the existing root id. - State, blobs and user-scoped aliases land under the target root id; the root keeps its own id, title, tags and permissions.
- To migrate several users, export each into their own directory and import one directory per run — the recorded owner email maps each root to its owner. Two root exports resolving to the same user fail the run before anything is written.
Bundle layout
Every export, single-document or export-all, writes one subdirectory per document under documents/<document-id>/. export-all additionally writes a top-level manifest.json naming the run's document ids; a single export does not, and import falls back to discovering the documents/ subdirectories directly when there is no manifest to read:
<output>/
manifest.json # export-all only
collections.json # collections export only
documents/
<document-id>/
metadata.json
permissions.json
document.yjs # an ordinary document
chain.json # a large document (documentFormat: 2), in place of document.yjs
snapshot/manifest.json # ...when the chain has a base
snapshot/<model>/<n>.ndjson.gz
epochs/<epoch>.yjs
current.yjs
blobs/index.json
blobs/<blob-id>.bin| File | Fields | Restored by import? |
|---|---|---|
manifest.json (export-all only) | version, exportedAt, sourceAppId, documentCount, documents (the exported ids) | Read to find which documents/<id> directories belong to the run; the file itself is not installed |
metadata.json | documentId, title, tags, documentFormat (present and 2 only on a large document), createdAt, createdBy, aliases (aliasScope, aliasKey — user-scoped only) | title and tags when the document is created (an existing document keeps its own); aliases per --aliases skip|overwrite |
permissions.json | An array of { email, permission, grantedAt?, source?: "invitation", status?: "pending" } | No — exported for reference only. The importing admin becomes the new owner, and sharing is managed from the target app |
document.yjs | The document's full state, one Yjs update | Yes — merged into the target document with Y.applyUpdate |
chain.json (large document only) | version, documentFormat: 2, documentId, epoch (the open epoch), base (epoch, buildId, rows, or null when the chain has none yet), overlays (epoch, sealedAt, oldest first), current (epoch, file: "current.yjs") | Yes — read to drive the install, which refuses up front if a file it names is missing |
snapshot/manifest.json | The base snapshot's manifest, with each chunk's path renumbered to address its file under snapshot/<model>/<n>.ndjson.gz; a chunk entry also carries model, rows, bytes and sha256 | Yes, with its chunk files, when chain.base is set |
epochs/<epoch>.yjs, current.yjs | Each sealed overlay after the base, and the still-open epoch's overlay — one Yjs update per file | Yes |
blobs/index.json | An array of { blobId, filename, contentType, sha256, numBytes }, one entry per file under blobs/ | Yes — each entry is re-uploaded under its original blobId from the matching blobs/<blob-id>.bin |
primitive collections export writes one more file into the same directory, collections.json, so a single directory is a whole app's migration rather than only its documents. It is read by primitive collections import and by nothing else — documents import ignores it, and collections import ignores everything else in the directory. Run documents import first: collections.json refers to documents by the ids that import preserves. See Moving Collections Between Apps for what it carries and what it does not.
The two manifests in a bundle encode their sha256 differently, and only one reader has to reconcile them: blobs/index.json's sha256 is base64, matching how the client computes and checks a blob's checksum; snapshot/manifest.json's per-chunk sha256 is hex, a SHA-256 of that chunk's stored (gzip) bytes. Both name a SHA-256 digest — decode (or re-encode) before comparing one against the other.
Large Documents
A large document removes the ordinary document's size ceiling: records live in a persisted local database and in the server's own table, and only recent changes are held as a document in memory. One is validated at 2 GB, and neither end ever materializes all of it. Reads, queries, live collaboration, permissions and offline writes work exactly as they do on an ordinary document.
Which clients can open one. A large document needs a local database that outlives the session, so it is opened by Node clients (and any host that provides a persistent SQLite store), by a Swift client whose store is on disk — the default, storageConfig: .sqlite(directory:); a client built with .memory fails the open with a typed FORMAT2_STORAGE_UNAVAILABLE error — and by a browser whose client is configured with the durable engine, databaseConfig: { type: "opfs", options: { workerURL } } (desktop Chrome, Firefox and Safari). Without that configuration a browser's engine holds data only for the life of the page, so opening a large document there fails immediately with a typed FORMAT2_STORAGE_UNAVAILABLE error rather than opening a copy that a reload would throw away. Ordinary documents are unaffected under either configuration. Under opfs one store holds every large document a signed-in user has open on that app and server, so a query(), queryOne(), count() or aggregate() on a model spans all of them in one call — with sort, limit, cursors and include intact — and documents narrows it to the ones you name. Two users on the same browser, or two apps, get two stores that share nothing. The one scope that cannot be read in a single call is a model holding rows in an ordinary document and a large one at once: those live in different local stores, so the call is refused with FORMAT2_QUERY_SCOPE and you scope it with documents to one kind.
Ask for one at creation:
primitive documents create "Ledger" --largeIn code, pass documentFormat: 2 to documents.create (JavaScript) or CreateDocumentOptions(title:, documentFormat: 2) (Swift); DocumentInfo.documentFormat says which kind an existing document is.
const { metadata } = await client.documents.create({
title: "Ledger",
documentFormat: 2,
});
// The metadata the create resolves with reports the format it asked for,
// before the server commit has even landed.
const format: number | undefined = metadata.documentFormat;
await client.documents.open(metadata.documentId);
// And the server agrees once the create has committed.
const info = await client.documents.get(metadata.documentId);
const stored: number | undefined = info.documentFormat;let result = try await client.documents.create(
options: CreateDocumentOptions(title: "Ledger", documentFormat: 2)
)
let documentId = result.metadata?["documentId"]?.stringValue
// The format the create recorded locally, before the server commit lands.
let format = result.metadata?["documentFormat"]?.numberValue
_ = format
if let documentId {
// Open before querying or writing, as with any document.
_ = try await client.documents.open(documentId)
}documentFormat means the same thing on every create route — documents.create, documents.createWithAlias and documents.getOrCreateWithAlias — on both clients, and so do title, tags and metadata: each one is applied when the call creates the document. So an app that keeps every document as a large document can make a tagged one behind an alias in a single call, instead of resolving the alias, creating, and writing the tags separately — which reopens the race the idempotent route exists to close:
const result = await client.documents.getOrCreateWithAlias({
title: "Ledger",
alias: { scope: "user", aliasKey: "ledger" },
tags: ["ledger"],
documentFormat: 2,
});
// `created` is true only on the call that made it; `documentFormat` is `2`
// either way. A stated format the existing document does NOT have is refused
// with `DOCUMENT_FORMAT_MISMATCH` rather than quietly handing back a document
// of the other kind.
const created: boolean = result.created;
const format: number | undefined = result.documentFormat;
await client.documents.open(result.documentId);let result = try await client.documents.getOrCreateWithAlias(
options: GetOrCreateWithAliasOptions(
alias: AliasRef(scope: .user, aliasKey: "ledger"),
title: "Ledger",
tags: ["ledger"],
documentFormat: 2
)
)
// `created` is true only on the call that made it; `documentFormat` is `2`
// either way. A stated format the existing document does NOT have fails with
// `HttpError.serverCode == "DOCUMENT_FORMAT_MISMATCH"` rather than quietly
// handing back a document of the other kind.
let created: Bool = result.created
let format: Int? = result.documentFormat
_ = created
_ = format
_ = try await client.documents.open(result.documentId)Omitting the option — on any of the three — creates an ordinary document, and the request is byte-for-byte what it was before the option existed. A value that is neither 1 nor 2 is refused before anything is created; on JavaScript that is a JsBaoError with code INVALID_ARGUMENT. And localOnly: true cannot be combined with documentFormat: 2 (a large document's records live in a store the server's room opens, which a local-only document never reaches): both clients refuse that combination at create, before any local state is written, with code LOCAL_ONLY_UNSUPPORTED_OPTION.
getOrCreateWithAlias is the one route that can be handed a format and find a document already there, so a stated documentFormat is read twice over: applied when the call creates the document, and otherwise taken as a statement about the document the alias already names. A stored format that differs is refused with DOCUMENT_FORMAT_MISMATCH (409) naming the document and both formats, and creates nothing — handing an app that asked for a large document an ordinary one is the mistake the route exists to prevent. One that agrees is echoed back with created: false. State no format and an existing binding is answered exactly as it always was.
One thing to know if you call these routes over HTTP or build a body by hand: all three refuse a top-level body key they do not read, with 400 VALIDATION_FAILED and a details entry naming it (value.<key> is not allowed). That includes name, parentId and a flat scope/aliasKey pair, which the spec has long declared but no handler has ever read — previously they were accepted and ignored, so a typo'd option looked like a success. The typed clients send only keys the routes read, so this reaches you only through a hand-built body.
The choice is made at creation and is never migrated: a document created without it stays an ordinary document, and one created with it cannot be converted back. Pick it for data you expect to accumulate — multi-year records, imports — rather than splitting that data across documents. Past creation it is the same document API — open it, then read and write through the same model classes:
await client.documents.open(documentId);
const imported = new Task({ title: "Imported row", priority: 0 });
await imported.save({ targetDocument: documentId });
const pending = await Task.query({ completed: false }, { documents: documentId });_ = try await client.documents.open(documentId)
let imported = try Task(
id: UUID().uuidString,
title: "Imported row",
priority: 0
).save(in: documentId)
_ = imported
let pending = try Task.query(
["completed": false],
options: QueryOptions(documents: [documentId])
)Two limits are worth knowing before you choose:
Offline writes are bounded by a window — 7 days by default, and the window is your app's to set:
largeDocumentWindowDaysinapp.toml's[app]section (orPUT /settings), anywhere from 1 to 14 days. Delete the line and the app goes back to the deployment's window. A client that has been away longer than the window goes read-only: it keeps serving reads, and rejects local writes until it syncs; it becomes writable again as soon as it does. The same window decides how long the server keeps the change archives a returning client catches up across, so it is one number, not two that can disagree.A refused write is reported the same way on both clients. Every local write the window refuses raises
document:write-refused— payloadDocumentWriteRefusedEvent, carryingdocumentId,model,recordIdanderror— and, wherever the call can throw, throws that same error too. The code isDOCUMENT_OFFLINE_WINDOW_EXPIREDon both clients:DocumentOfflineWindowErroron JavaScript (save()anddelete()),JsBaoError(.documentOfflineWindowExpired)on Swift (create,update,save,upsert,addMember,removeMember;delete(id:)and the record field setters cannot throw and have only the event). Either wayerror.detailscarriesdocumentId,lastSyncAt,windowDaysandoverdueMs. So subscribe once and handle the refusal in one place, on either client; the throw stays the call site's signal where there is a call site. Nothing is raised for an accepted write, for an ordinary document, or for a refusal of another kind.tsclient.on("document:write-refused", (event) => { // `error.code` is `DOCUMENT_OFFLINE_WINDOW_EXPIRED`, and `error.details` // carries `documentId`, `lastSyncAt`, `windowDays` and `overdueMs`. console.warn( `write refused on ${event.documentId}: ${event.model}/${event.recordId}`, event.error.code, event.error.details ); });swiftlet refused = client.observeOnMainActor(DocumentWriteRefusedEvent.self) { event in // `event.error.code == .documentOfflineWindowExpired`, and `details` // carries `documentId`, `lastSyncAt`, `windowDays` and `overdueMs`. print( "write refused on \(event.documentId):", event.model, event.recordId, event.error.code.rawValue ) }Composite field values are rejected. A large document's fields hold plain JSON values; nested collaborative types (rich text, nested maps and arrays) are refused at write time instead of being silently flattened.
Two read-path behaviors are specific to large documents, because their records come from the local database rather than from the in-memory document:
- A field read can be one update behind for a moment. When a peer's change arrives, reading a field on a model instance you are already holding may return the previous value until that change has been folded into the local database.
awaitthe client's projection barrier if you need the settled value —find()andquery()always agree with each other. - A model built from an id alone has its values after its first
await.new Model({ id })returns immediately with schema defaults; the record's stored values are there after the firstfind()orsave(). Such an instance is not treated as a new record: saving it patches only the fields you set and leaves every other stored field alone.
If a change cannot be written into the local database — a full disk, a read-only or damaged store — the document stops answering rather than serving rows that are quietly out of date: reads and writes then fail with a typed FORMAT2_FOLD_BROKEN error until the document is reconnected and catches up.
Reopening a large document
Reopening a large document costs what changed since you last had it open, not what it holds. Its local store remembers the overlay state it last folded, so a reopen — after a close or a reload, on the OPFS engine and on Node alike — folds and projects only the records that changed in the overlay since the last fold. When nothing changed it folds and projects nothing at all, whatever the document's size: the open costs about what loading the document's local copy costs. A record created or re-created since the last fold is rebuilt from its model's overlay in memory before it is folded; every other record costs only itself.
In a browser, a tab running a newer version of your app refuses to open a large document through the worker an older tab of the app started, with a typed FORMAT2_STORE_OUTDATED error ("the tab leading this store runs an older version of the app; reload it or close it"). That window only exists right after you deploy: reloading or closing the older tab ends it, and the document's local data is untouched either way. An older tab joining a newer worker is unaffected.
Removing a large document's local data
An ordinary close keeps everything a large document stores locally — its records, its unacknowledged writes, the marks that say its query tables are current, and the projected query rows — so the next open re-projects nothing and is as cheap as a reload. A closed document is not part of an unscoped query(), count() or aggregate(): those read only the documents that are open. In a browser a close does not release the browser worker that holds the store, and neither does hiding the tab. What releases it is client.destroy(), the page's freeze or pagehide (a real close, or the browser freezing the page), and an evict that empties the pool.
Three doors remove what a large document stores locally — its records, its member index, the writes the server has not acknowledged yet, the marks that say its query tables are current, and the projected query rows themselves:
await client.documents.evict(documentId);
await client.closeDocument(documentId, { evictLocal: true });
await client.logout({ wipeLocal: true });The rules are the same on every client:
- A document open in another tab is never taken from under it. An evict of a large document another tab still has open is refused for that document's local data: that tab keeps reading and writing, its unacknowledged writes stay in the log, and one warning line names the document and the reason. The document is still evicted as far as the evicting client is concerned; its store is removed by a later evict or wipe, once no tab holds it.
- Unacknowledged writes are not dropped silently. An evict without
forceis refused when the store still holds writes the server has not confirmed, and throws the same "has unsynced local changes (use force to override)" error an ordinary document's does.documents.evict(id, { force: true })removes them. An evicting close skips the eviction on the same grounds. - A document evicted while it is open stays editable. Its store is removed when it is closed — whatever that close's own options say — so an eviction never interrupts work in progress.
- A wipe reaches documents this session never opened.
logout({ wipeLocal: true })removes the signing-out user's large documents whether or not they were opened in this session, including a document that is open at the moment of the logout, and including a store left behind by a version of the client that could not remove one. A wipe of one user's data never touches another user's, another app's or another server's. - The disk is reclaimed. In a browser, once a large document's local store holds nothing else, the OPFS directory holding it is removed as well; opening the document again fetches it from the server.
Opening one in several tabs
A user's large documents share one browser store, and that store's worker allows one connection to its database — so add brokerURL beside workerURL in databaseConfig.options to let more than one tab of the same app share it. The first tab to open a large document becomes the leader for that user, however many documents are open, and holds the connection directly; a later tab talks to that same worker through a port a small broker hands over. Saves, reads and queries behave the same in every tab — a save committed in one tab is visible to find() and query() in another as soon as it settles. Closing the leader tab hands the connection to another open tab automatically, in one handover that carries every open large document, with nothing pending lost. Without brokerURL configured, a second tab opening the same large document is refused with a typed FORMAT2_WORKER_OPEN_FAILED error rather than silently taking over the first tab's connection.
When the platform and your client disagree about the format
Both clients record a document's format on their own local metadata row at create time and open the document by it; the platform resolves the format independently and refuses to serve a document the two do not agree about, rather than answering in the wrong shape. Two refusals arrive at the handshake and they are not the same thing:
CLIENT_UPGRADE_REQUIREDsays this client BUILD cannot read any large document. The connection has nothing left to do, so the platform closes the socket (close code 4426). Upgrade the client library.DOCUMENT_FORMAT_MISMATCHsays this client opened ONE document as the wrong format. The connection is untouched — every other document on it keeps syncing — and that one document is not served. An open that was waiting on the network rejects with the typed error (JsBaoErrorwithcode: "DOCUMENT_FORMAT_MISMATCH"anddetails { documentId, declared, actual }on JavaScript;JsBaoError(code: .documentFormatMismatch)with the same details on Swift). An open served from cache, or a document that was already open, is closed under your app instead and the typeddocument:format-mismatchevent carriesdocumentId,declared,actualand that same error. Local rows and writes the server has not acknowledged are kept; nothing is evicted, including for a document you opened withretainLocal: false. The refusal then stands for the rest of that open cycle: the document sends no handshake and no update, and reopening it without closing it first throws the same error rather than asking the platform again. Closing the document or evicting it ends the refusal.
Diagnose before you act. declared is what your client believed and actual is what the platform resolved, and the server logs Document format disagreement with both plus the document id, so which side is stale is a question with an answer. If the platform is right and this client's row is stale — an export and import, an old build — evicting the document and reopening it clears the stale row. If the platform is WRONG, eviction is not a repair: it deletes the row that carries the declaration, and the next open declares nothing and may be served the wrong format silently. Such a document needs the mis-pinned format repaired on the platform side, not cleared on the client.
A records request can state the same belief. Every documents/{documentId}/records… route takes documentFormat=1|2 as a query parameter, the CLI verbs take --document-format <1|2>, and a server function states it once on the typed handle as ctx.doc(documentId, { documentFormat }). A request whose stated format disagrees with the platform's is answered 409 with code: "DOCUMENT_FORMAT_MISMATCH" before any table is read or written; a request that states nothing is answered exactly as before.
On Swift
A Swift client reports a base load through document:snapshot-load (DocumentSnapshotLoadEvent, phases started, progress, model and loaded, with mode: "load"). A model with members in both an ordinary and a large document must scope the read to one kind, or the call is refused with FORMAT2_QUERY_SCOPE: query and count take QueryOptions(documents:) and aggregate takes AggregateOptions(documents:), which means the same thing — a list of documents, with an explicit empty list matching nothing. On a model bound to one document the option narrows rather than replaces, so naming another document answers nothing. A server that refuses the client's formats fails the open with CLIENT_UPGRADE_REQUIRED (a ConnectionErrorEvent on the handshake; the client does not reconnect); a server that refuses ONE document's format answers DOCUMENT_FORMAT_MISMATCH instead — the socket stays up, the awaiting open throws JsBaoError(code: .documentFormatMismatch) with documentId, declared and actual in details, a document with no open waiting is closed under the app with its store intact (a retainLocal: false open included), and either way DocumentFormatMismatchEvent (document:format-mismatch) carries the same fields. The refusal stands for the rest of that open cycle: nothing is handshaken or sent for the document, and openDocument throws the same error until closeDocument or documents.evict ends it. Read "When the platform and your client disagree about the format" above before evicting anything. The document follows the room's epoch seals in place — no reload and no download; the YDocument handle is replaced at each seal, so read records through the model facade — and reloads from the newest base only when the chain cannot be trusted or a replay would drop a delete (mode stays "load"; FORMAT2_RELOAD_REQUIRED refuses writes while that reload is pending, FORMAT2_FOLD_BROKEN when the merged view is untrustworthy). Writes made offline are judged against the sealed chain when the client returns: clearly older writes and writes onto deleted records are dropped, ambiguous ones stand, and each verdict reaches the app as DocumentOfflineWritesResolvedEvent (documentOfflineWritesResolved: per write an outcome of dropped or kept-ambiguous and a reason of outdated, record-deleted, in-window, unverifiable or bulkIngest); a relaunch adopts the previous instance's unacknowledged writes, and a bulk load (below) is crossed without reopening. Past the app's offline write window the document is read-only, and every refused write is reported the same way here as on the JavaScript client: it emits DocumentWriteRefusedEvent (document:write-refused), and create, update, save, upsert and the string-set writes additionally throw JsBaoError(.documentOfflineWindowExpired) (DOCUMENT_OFFLINE_WINDOW_EXPIRED, with lastSyncAt, windowDays and overdueMs in details); delete(id:) and the field setters cannot throw and have only the event. The event is delivered after the document's operation lock is released, so a handler may read any document — including the refusing one. A sync restores writes. A device that cannot hold the whole document names the models to load in JsBaoClientOptions(largeDocumentStorage: LargeDocumentStorageOptions(capability:models:)): the capped load is durable, a read of a model left out throws .format2ModelNotHydrated, and a device that cannot hold even those is refused with .format2StorageUnavailable (reason over-quota) before any chunk is fetched. documents.evict and logout(wipeLocal: true) remove a large document's local tables and unacknowledged writes along with everything else.
Inspecting snapshot builds
A large document's server-side table is periodically written out as a base snapshot that a cold client loads before catching up on recent changes. Every build is checked against the table it was built from — each chunk read back and its rows compared with the table's — before it is registered, and a build that fails is never handed to a client: it is recorded as failed with a reason and retried under a new id. Each build names its source — builder, import or ingest — and, for a bulk load, the ingestSessionId that produced it. Any reader of the document, or an app admin, can list the builds and read a verification result:
primitive documents snapshots list <document-id>
primitive documents snapshots get <document-id> <build-id>The same page is GET /app/{appId}/api/documents/{documentId}/snapshots, and one build is .../snapshots/{buildId}; add --json to get it as the CLI receives it. A build's verification.state is passed once every check ran, failed with a code (SNAPSHOT_VERIFY_DIGEST, SNAPSHOT_VERIFY_COVERAGE) and a reason naming the chunk or the record when one did not, and timings.totals says where the build's time went per stage.
The room seals the open epoch on its own when the overlay passes 1 MiB encoded or 32,768 Yjs items (overwritten values included), no sooner than 10 seconds after it opened unless the overlay reaches three times either limit, and when the epoch is a week old. It builds a base on its own every eight seals, or an hour after a seal no base covers yet, rather than on every seal: under edits spread across a large document every build rewrites the whole base. To get one now — before a cold-load measurement, an audit or a migration, or after a burst of writes that a returning client would otherwise have to fold — ask for one:
primitive documents snapshots build <document-id>
primitive documents snapshots build <document-id> --waitIt seals the epoch the document is writing to and starts the base build that seal arms, then prints the epoch it sealed and the build id, so documents snapshots get can follow it. --wait polls until the build has verified, exiting 0 when it passed and 1 when it failed; --timeout <seconds> exits 124 and Ctrl-C exits 130, both leaving the build running — giving up on watching is not giving up on the build. --json prints the answer as received, with the build the wait settled on beside it.
The route is POST /app/{appId}/api/documents/{documentId}/snapshots, and client.documents.snapshots.build(documentId) is the same call from the JS client. Two answers are worth telling apart:
- if the open epoch carried nothing there is nothing to snapshot, so it is not sealed: the answer is
{ sealed: false, reason: "empty", coveringBuildId }, naming the completed base that already describes the document — andnullwhen none does yet, either because the document has never been snapshotted or because the base for its most recent seal is still being built; - asking again within a minute of the open epoch is refused with 429
SNAPSHOT_TOO_SOONand adetails.retryAfterMs, which is what stops a caller sealing in a loop. The epoch is left open.
Unlike the reads above, asking for a snapshot needs the permission a bulk load needs — a document grant at read-write or above, or an app admin — because it changes what the server stores for the document. A reader, a link-access holder and a caller of another app are refused.
To check the whole document rather than one build — that the base plus its overlays really is the table the server holds — run the audit:
primitive documents snapshots audit <document-id>It rebuilds the chain locally with the client's own loader and fold and compares it with the table, reporting per model what was compared and what differs. It needs both halves of what that takes: the app admin arm to export the chain, and a reader of the document — a console admin, or a read grant on the document — to read the table it is compared with. A caller with only the first is told which grant is missing before anything is exported. Exit codes: 0 the table is the chain, 1 a divergence (named in the report), 2 no verdict — an ordinary document, a document with no completed base, a chain the server will not export, a document that changed during every attempt, or a read the audit could not finish (a refused page, a grant withdrawn mid-run, an artifact that would not download). A write landing while the audit runs is not corruption: the audit notices and retries (--retries, default 3) rather than reporting it, so run it when the document is quiet.
Bulk-Loading a Large Document
A large document (documentFormat: 2) can hold far more records than writing them one at a time is a sensible way to load. Refreshing a dataset or mass-correcting records through ordinary saves forces a seal and an archive every 1 MiB of overlay, with a base build every eight seals, and leaves collaborative history nobody asked for. A bulk load is the other path: the rows go in as one artifact, nothing is visible until one atomic swap, and connected clients converge onto the result instead of reloading the document.
Like export and import, it is an operator command rather than application code:
# A directory of per-model line files — one record per line, either
# `<id><tab><merge patch JSON>` or one JSON object carrying its own `id`:
# records/note.ndjson
# records/tag.ndjson.gz
primitive documents ingest <document-id> --input ./records -y
# An export re-loads as it stands
primitive documents export <app-id> <document-id> --output ./export
primitive documents ingest <document-id> --input ./export -y
# Watch a session that is running, or one started with --no-wait, and see
# where its time went: per state, per stage, and its throughput
primitive documents ingests get <document-id> <session-id>Two input layouts are accepted. A plain directory holds one <model>.ndjson (or .ndjson.gz) per model, one record per line. An export directory — what primitive documents export wrote, either at the directory itself or under documents/<document-id>/ — is read as it stands, so a document can be re-loaded from its own export; the overlays and current.yjs beside the snapshot are named on stderr and skipped. A directory that is neither, that is both, or that is some other document's export is refused, naming what it held.
Each line is an RFC 7396 merge patch over the record: a value replaces a field, null unsets it, a StringSet field takes a whole array ([] is an empty set, null removes it), and {"_deleted": true} alone deletes the record. Only the document's existing models and fields are accepted, and every line is checked against them before a session is opened — the first failure names the file and the line. It replaces records in a live document and cannot be undone, so it confirms unless -y, and exits 0 completed, 1 failed or refused, 124 on --timeout or on a session read that kept failing, 130 on Ctrl-C (in all three of those the session keeps running). A session read that fails while waiting is retried on a backoff that doubles from 1 second to 15, reported per attempt on stderr and reset by any read that answers; only after ten minutes of unbroken silence does it stop watching, and it then says the load is still running and names documents ingests get rather than reporting a failure.
documents ingests get shows a session's progress through uploading → committed → validating → staging → applying → finalizing → registering → complete. On a very large document, registering removes the tables the swap moved aside a bounded amount at a time, over as many alarms as it needs, so it is normal for a session to sit there for a while: the removed figure in the progress row climbing between reads is how you tell it is advancing rather than stuck. If the bound turns out to be wrong for a document, the session stops with INGEST_REGISTER_STALLED, naming the table it could not finish — the bulk load itself has already landed and the document is correct, and the next ingest on that document clears the remains before it starts. A read that cannot be answered says why rather than failing generically: DOCUMENT_UNAVAILABLE (503) means the document was momentarily unreachable — its object was reset, evicted or overloaded — and the same read a moment later usually answers; INGEST_SESSION_READ_FAILED (500) means the session ledger itself could not be read, and whether asking again helps is left to the server log. A session that simply does not exist is INGEST_SESSION_NOT_FOUND (404).
documents ingests get also reports where a session's time went in two ways, and they are not the same kind of number. The per-stage milliseconds (timings.totals.manifestMs and its eight siblings) are wall time, and the server brackets them by yielding at both ends of every stage so the timer reads the work it encloses rather than the 0 a Durable Object's frozen clock would give it; timings.totals.bracketedTicks says how many of the session's ticks managed that, and a run where it is less than timings.ticks should be read as degraded rather than mixed in with a fully bracketed one. Bracketed or not, these are never CPU time: they include the yield's own cost and any I/O inside the bracket.
timings.totals.counters is the half that needs no clock. Per stage (manifest, read, decode, copy, apply, reconcile, swap, seal, register) it carries statements, which is exact everywhere; rowsRead and rowsWritten, taken from the database cursor's own counters and reported as null where there are none rather than as a zero that would read as a measurement; and rowsReturned, which is a different figure again — an aggregate answers one row after reading a million. Inside each stage, byTarget splits the same figures by the table each statement touched (records, members, claims, log, bookkeeping, other), which is what tells you whether an apply stage is spending itself on the records fold, on StringSet index maintenance, or on unique-constraint claims.
The JS client carries thin wrappers over the same routes — client.documents.ingests.create, uploadChunk, commit, abort, get and list, plus client.documents.snapshots.list and get — for a server-side job that needs to drive or watch one (client 3.3.0 or newer, which needs js-bao 0.9.0 or newer).
Writing a large dataset as a series of atomic records/bulk batches instead — primitive documents records bulk, or POST .../documents/{documentId}/records/bulk — is the other way to load one. Unlike ingest it goes through the epoch overlay like any other write, so a load of that kind seals an epoch every 1 MiB (at most every 10 seconds, sooner past 3 MiB) and moves every connected client each time; ingest seals once. It is the right tool for batches an application writes as it runs, and a batch that is refused says whether it is worth sending again. On a large document, a batch refused because the document's object was momentarily unreachable (reset, evicted or overloaded mid-write) answers 503 DOCUMENT_UNAVAILABLE with "The document is momentarily unreachable; retry the write", and the same batch a moment later usually lands. 500 INTERNAL_ERROR means the write failed for a reason the server kept in its log, so retrying it alone is unlikely to help. A 400 or a 409 CONDITION_NOT_MET is the batch itself — a malformed blob, a validation failure, or a precondition that no longer holds — and will fail the same way unchanged. Ordinary documents answer exactly the codes they always did on this route.
Connected clients do not reload when the swap lands. They keep serving reads throughout, re-fetch only the chunks the artifact touched, refold their own recent writes on top, and report the result through document:snapshot-load with mode: "converge" and chunksReused, the number of chunks it kept. A Swift client crosses the load without reopening: it records the base discontinuity, rebuilds from the next base (mode stays "load") and reports the verdicts on its own writes through documentOfflineWritesResolved with reason bulkIngest. A write still unacknowledged when the load landed is classified rather than replayed blindly: one on a record the load deleted is dropped and surfaced through documentOfflineWritesResolved with reason: "bulkIngest", and one on a record it modified is applied and surfaced as ambiguous with the same reason.
Next Steps
- Choosing Your Data Model — When to use documents vs. databases
- Invitations — App membership, and how shares to not-yet-users resolve
- Working with Databases — Server-side structured storage
- Blobs and Files — Binary file storage, document-scoped and general-purpose