Skip to content

Tasks and Requests ​

Every server function runs under one of two runtimes: the request runtime, which answers inside the call, or the task runtime, which starts a run that can sleep for hours or days. This page covers the two runtimes and how a task run behaves — steps, resets, budgets, run status and run error codes. How a caller picks one is on Invoking and Triggers.

Runtimes ​

A function is a function. What changes is the runtime the platform runs it under, and nothing in its config says which — the caller picks at each call, and each door has its own answer:

DoorRuntime
functions.invoke / POST functions/{key}request
A webhook deliveryrequest
functions.start / POST functions/{key}/starttask
ctx.functions.start from inside a functiontask
A cron firetask

Every cron fire starts a task run, and every webhook delivery runs under the request runtime. That is a property of the door, not of the function. A schedule is waiting for nothing, so there is nothing for a bounded invocation to answer; a provider is waiting inside the request for the function's answer, and a task run has none to give there — it returns a run id and finishes later. Neither is a dead end: a webhook-fired invocation can start a task run, the same way any other caller can.

  • the request runtime runs the code inside the call and hands back its result. It has to finish inside a request.
  • the task runtime starts a run, hands back a run id immediately, and lets the code sleep for hours or days in between.

So the file below takes both verbs, and nothing in it says so — that is the point:

toml
# primitive/dev/functions/order-sync.toml
[function]
key = "order-sync"
entry = "functions/order-sync/index.ts"
access = "true"

ctx.runtime ​

Every invocation carries the runtime it is running under, beside the trigger that fired it:

ts
export default defineFunction(async (input, ctx) => {
  // "request" | "task" — set by the platform from the runner in use.
  console.log(ctx.runtime, ctx.trigger.kind);
});

ctx.trigger.kind says which door fired the invocation; ctx.runtime says whether the platform ran the code inside the request or as a task run. They are different questions: a manual invocation can be either, while every other door has the one runtime the table gives it.

ctx.runId, ctx.sliceId and ctx.trigger.runKey ​

A function can read which run it is, and which slice of that run it is running in:

ts
export default defineFunction(async (input, ctx, step) => {
  // The run this invocation belongs to — `null` for an `invoke`, which writes
  // no run row at all.
  ctx.runId;
  // The slice of that run. A task run is a sequence of slices: it changes when
  // the run hibernates and wakes, so it is how you tell a resume from a first
  // pass. `null` under the request runtime — a request invocation is not a
  // slice.
  ctx.sliceId;
});
ctx.runIdctx.sliceId
invoke over HTTPnullnull
A webhook deliverythe fire's runnull
start, a cron fire, a nested startthe runthe slice

ctx.sliceId is the same identifier primitive functions runs and the run status route's slice.sliceId report, so a log line carrying it lines up with what an operator sees.

When a start was coalesced by a run key — start with a runKey, ctx.functions.start(key, input, { runKey }), or an admin's system start — ctx.trigger.runKey carries it, on the http, function and manual arms. It is null when the start named none and on a request invocation, which coalesces nothing. A trigger root's run key is its own run id, which ctx.runId already says, so the webhook and cron arms do not carry one.

assertRuntime ​

Some code only makes sense one way round. A function that sleeps for a day is not something anyone should be calling with invoke and waiting on. Say so in the code, and the platform refuses the call:

ts
import { assertRuntime, defineFunction } from "primitive-functions";

// At MODULE SCOPE: every invocation under the wrong runtime is refused before
// a line of the handler runs.
assertRuntime("task");

export default defineFunction(async (input: { since: string }, ctx, step) => {
  await step.sleep("cooling-off", "24 hours");
  return { sweptAt: Date.now() };
});

It works inside the handler too, and inside a step.do body, if only part of the code needs the guarantee:

ts
export default defineFunction(async (input, ctx, step) => {
  if (input.slow) assertRuntime("task");
  // …
});

A mismatch throws a typed FunctionRuntimeError (name is "FUNCTION_RUNTIME_REFUSED", with runtime and required on it) and the platform settles the invocation as status: "failed" with errorCode: "FUNCTION_RUNTIME_REFUSED" and a message naming both runtimes and the verb that takes the one it needs:

json
{
  "status": "failed",
  "errorCode": "FUNCTION_RUNTIME_REFUSED",
  "error": "FUNCTION_RUNTIME_REFUSED: this function is running under the request runtime and requires the task runtime. Start it as a task: functions.start(...) from a client, or POST .../functions/{key}/start."
}

The same settlement on a task run: the run ends failed, its errorCode is FUNCTION_RUNTIME_REFUSED, and status.error.message leads with the code. Either way the invocation log records it.

There is no declaration in code or config beyond this. The assertion is the lock, and it is the only one: mode and durable are not keys of a function, and nothing in its TOML chooses a runtime.

step under each runtime ​

The handler's third argument, step, is there under both runtimes, so one body runs either way. Where they genuinely differ, the request runtime says so rather than pretending:

Task runtimeRequest runtime
step.doRuns once and the result is memoized; a replayed slice does not re-run it.Runs the body inline. Nothing is persisted and nothing is replayed.
step.sleep / step.sleepUntilHibernates. Nothing runs and nothing is billed while it waits; hours and days are fine.Really waits, when the remaining request budget covers it. When it does not, the invocation fails with errorCode: "FUNCTION_RUNTIME_REFUSED", naming the step, the duration and the budget left.
step.waitForEventWaits for the event.Fails with FUNCTION_RUNTIME_REFUSED: nothing can deliver an event to an invocation that has to return now.

It is the same code an assertRuntime refusal settles with, and for the same reason: this runtime cannot host this code. The thrown error's name is STEP_NOT_AVAILABLE in every step refusal, so a body that catches it keeps catching it. A duration the platform cannot read ("3 fortnights") is an ordinary throw naming the grammar — never a silent "done".

This is the rule on every request-runtime door, not just the HTTP one: a webhook delivery answers FUNCTION_RUNTIME_REFUSED for a wait it cannot cover. The remedy is always the same — start the function as a task instead.

The handler with step:

ts
import { defineFunction } from "primitive-functions";

export default defineFunction(async (input: { orderId: string }, ctx, step) => {
  const charged = await step.do("charge", async () => {
    return ctx.api.documents.get({ documentId: input.orderId });
  });

  // Hibernate. Nothing is running — and nothing is billed — while it waits.
  await step.sleep("cooling-off", "24 hours");

  return { orderId: charged.documentId, confirmedAt: Date.now() };
});

Task runs ​

A task run is a sequence of slices: the engine runs the handler, hibernates it at a sleep, and wakes it again. The sections below cover starting one from code, keeping its steps safe to replay, what it may spend, and reading what it did.

Starting a task run from a function ​

ctx.functions.start starts another function as a task run from inside a running one — from an invocation, from a task run's slice, and from a webhook- or cron-fired run alike:

ts
const started = await ctx.functions.start("digest-page", { userIds }, {
  runKey: `digest-${date}-${page}`,
});
// { runId, runKey, instanceId, status: "running" }

It answers the same start envelope an HTTP start answers, so the run id polls and terminates like any other. Repeating a runKey replays: you get existing: true and the run that already exists, which is how a run key becomes a singleton — including when that run failed. A run key says "this run, once", not "this run, until it succeeds", so a retry is a new key with the same input.

There is no nested invoke. Running a function inside your own invocation is ordinary code, and code composes by being imported — put the shared logic in a module both functions import. ctx.functions.start always starts a task run, whatever the callee is; a callee that must not run that way says so with assertRuntime("request"), and the refusal settles the child run FUNCTION_RUNTIME_REFUSED.

The callee's access expression is not evaluated. Your own gate was the authorization, and your code chose the callee, so a function can start one the invoking member could not call themselves. Decide who may reach the tree in the access expression of the function at the top of it.

Who the run belongs to. Every run in an invocation tree — the root and every nested start under it — is keyed by the original external initiator: the member who made the HTTP call, or the app's system principal when a trigger fired it. So the member who started a tree polls and terminates every run in it, another member gets the routes' 404, and a trigger-rooted tree is an admin's to inspect.

Which document the run is keyed under. A task run is addressed by a context document plus its run key, so a nested run always has one, resolved in this order:

  1. the contextDocId you pass;
  2. else the parent run's context — so a whole tree shares one;
  3. else the root caller's root document, which is the default a member's own start uses;
  4. else fn:<the root function's id> — a synthetic context, the same shape a cron-fired run's cron:<triggerId> has. This is what a trigger-rooted tree gets, and it is why a run key there is a singleton per root function across fires rather than per member.

Rule 2 reads the parent's run row. If that row has expired or been deleted mid-tree, resolution falls through to rules 3 and 4 — a rule, not an accident. The value comes back on the run's status response, and you need it to terminate a trigger-rooted child by key (?contextDocId=fn:<id>).

Both runKey and contextDocId are bounded at 100 characters.

Inside the child, ctx.trigger.kind is "function" and carries the parent function's key and its run id (null when the parent was a request invocation, which writes no run row). ctx.user names the tree's root caller, and ctx.user is null when a trigger rooted the tree — nobody called it.

Nesting is bounded at 4 levels. A root is level 0, so four nested starts fit below it; the fifth answers 409 FUNCTION_NEST_DEPTH_EXCEEDED. The depth travels on the platform's own invocation credential, so it is not something your code can raise, and the ceiling is what stops a function that starts itself from running away.

Each run is listed under its own function — primitive functions runs digest-page — with FIRED BY function and a PARENT column naming the function that started it.

Iterate every user ​

App-wide iteration is a cursor plus a nested start: one page per slice, one child run per page. Three functions, and you can copy them as they are.

The trigger. A function on a schedule (each cron fire is a short task run of its own), which starts the orchestrator with a run key derived from the date — so a second fire the same day replays rather than starting a second tree, and the cron's own overlapPolicy bounds this parent.

toml
# functions/nightly-digest.toml
[function]
key = "nightly-digest"
entry = "functions/nightly-digest.ts"
access = "false"

[[function.triggers.cron]]
name = "nightly"
cron = "0 3 * * *"
overlapPolicy = "skip"
ts
// functions/nightly-digest.ts
export default async function (input, ctx) {
  const date = new Date().toISOString().slice(0, 10);
  return ctx.functions.start("digest-all-users", { date }, {
    runKey: `digest-${date}`,
  });
}

The orchestrator. A function, started as a task run, that takes one page per slice and hands each to a child. Both children declare assertRuntime("task") at module scope: neither is anything but a background run, and ctx.functions.start is the only way in. (The assertion is the authors' choice, not a requirement — ctx.functions.start starts any callee as a task run.) Both declare access = "false", because neither is ever invoked over HTTP: a nested start does not consult the callee's gate, so a closed gate costs the tree nothing and keeps the children off the member-callable surface.

toml
# functions/digest-all-users.toml
[function]
key = "digest-all-users"
entry = "functions/digest-all-users.ts"
access = "false"
ts
// functions/digest-all-users.ts
import { assertRuntime } from "primitive-functions";

// This orchestrator is only ever a background run.
assertRuntime("task");

export default async function (input, ctx, step) {
  let cursor: string | undefined;
  let page = 0;
  let users = 0;
  while (true) {
    const taken = await step.do(`page-${page}`, () =>
      ctx.users.list({ limit: 50, cursor })
    );
    users += taken.items.length;
    await step.do(`start-${page}`, () =>
      ctx.functions.start(
        "digest-page",
        { userIds: taken.items.map((u) => u.userId) },
        { runKey: `digest-${input.date}-${page}` }
      )
    );
    cursor = taken.nextCursor;
    page += 1;
    if (!taken.hasMore || !cursor) break;
    await step.sleep(`next-${page}`, "1 second");
  }
  return { pages: page, users };
}

The pages are fetched inside step.do, with an explicit cursor, and that is not a style choice: a memoized step does not advance an iterator, so ctx.users.iterate would replay pages it had already taken the moment the run resumed after the sleep. Use iterate in an invocation or a single-slice loop; page by hand across slices.

The page worker. One task run per page, so each page's budget is its own.

toml
# functions/digest-page.toml
[function]
key = "digest-page"
entry = "functions/digest-page.ts"
access = "false"
ts
// functions/digest-page.ts
import { assertRuntime, pMap } from "primitive-functions";

assertRuntime("task");

export default async function (input, ctx, step) {
  return step.do("digest", () =>
    pMap(input.userIds, (userId) => ctx.users.send(userId, { digest: true }), {
      concurrency: 8,
    })
  );
}

The budget arithmetic. A task slice may make 10 000 subrequests (a request invocation 128) and 10 minutes of wall-clock budget per slice, refreshed by the platform as it goes so the clock is continuous for up to 12 hours. At one call per user, a page of 50 fits with room for the start and the cursor read — thousands of times over. Neither ceiling binds a page of sends: it is seconds of work, and a single step.do may run for up to a whole slice budget. What does bind is the yield at half the subrequest ceiling (see Budgets in a task run): a worker that has made 5 000 counted calls in one slice ends the slice at its next step boundary, so keep pages bounded and let the child runs do the fanning out. pMap's default concurrency of 8 bounds how many are in flight, not how many run in total.

Inside the children, ctx.user is null and ctx.trigger.kind is "function" — nobody called this tree, the schedule fired it. Every run in it is keyed under fn:<nightly-digest's id>, which the status response reports and which you pass as ?contextDocId= to terminate a child by key.

Partial failure is per page. A page that throws is one failed run in primitive functions runs digest-page, beside its completed siblings; the other pages are unaffected.

Retrying a page takes a NEW run key. A run key is a singleton, and that is all it is: starting digest-page again with the key the failed run already holds answers 200 {existing: true} describing that failed run, and nothing runs a second time. Give the retry its own key and the same input:

ts
await ctx.functions.start(
  "digest-page",
  { userIds },
  { runKey: `digest-${date}-${page}-retry-${attempt}` }
);

Same context document, so the retry is keyed under the same tree and shows up beside the run it replaces. A child that ends without anybody polling it settles its own row as it ends, like any other task run — see When a run's status settles.

Reporting back ​

A run never broadcasts over the WebSocket, so decide up front how its result reaches whoever cares. Three channels, and they compose:

  • The run's output. Return a value and it is recorded on the run's own record by the same settlement that records its status, so it stays readable for the life of the run row — 45 days — whether or not the engine still holds the instance. client.functions.getStatus(runId).output and waitFor(runId).output answer it, and so do primitive functions runs wait <function-id> <run-id> and primitive functions runs <function-id> --json. The runs TABLE does not print it: an output can be a megabyte, so the table says where to find it.

    Outputs over roughly 10 000 characters are stored by reference and answered whole; the split is invisible to a reader, except that a runs page holding many large outputs may be shorter than the limit you asked for and continue with a cursor. If the platform could not store a large output when the run settled, the record carries a preview marked outputTruncated and is completed from the engine on the next read while the engine still holds the run; once it has forgotten the instance, the preview is what is left.

  • Server state. Write database rows, documents or notifications — on the app's authority, attributed to the caller — inside step.do. This is the natural shape for a job that only needs to leave results behind.

  • A live push. ctx.users.send reaches a client that is connected right now and nobody else — pair it with state or a notification if the message has to survive the client being away.

Step discipline ​

The engine re-executes your handler from the top on every resume and replays each completed step.do from its stored result instead of running it again. Two rules follow, and both matter:

  • Put every side effect inside step.do. A charge, an email, a write made outside a step runs again on every resume. Inside one it runs exactly once, and its return value is what later resumes see.
  • Do not branch on anything nondeterministic between steps. Date.now(), Math.random() or a fresh read at the top level can take a different path after a resume, and the engine then looks for steps that are not there. Read such values inside a step, so the value is stored with it.

Locks from a function ​

A function takes a named lock through the platform's own lock API, and a task run has one thing to get right: acquire live on every slice, outside step.do.

ts
import { defineFunction } from "primitive-functions";

export default defineFunction(async (input: { userId: string }, ctx, step) => {
  const key = `portfolio-import:${input.userId}`;

  // Live, on every slice. `ctx.runId` is the owner: presenting it again from a
  // later slice of the SAME run re-takes the lease with a fresh handle.
  const got = await ctx.api.locks.tryAcquire({
    body: { key, ttlMs: 120_000, owner: ctx.runId },
  });
  if (!got.acquired) {
    // Somebody else's run holds it — `got.owner` says whose.
    return { skipped: true, heldBy: got.owner };
  }

  try {
    await step.do("import", async () => {
      // …the work that must not overlap…
    });
  } finally {
    // Live too: the handle belongs to this slice.
    await ctx.api.locks.release({
      body: { key, handle: { handleId: got.handle.handleId } },
    });
  }
});

A memoized acquire — one inside step.do — would hand a resumed slice the handle a previous slice was given, and that handle no longer works. That is why both the acquire and the release are live.

owner: ctx.runId is what makes the re-acquire succeed rather than collide with the run's own earlier slice. When several triggers are coalesced into one run by a run key, name the work that key stands for (ctx.trigger.runKey ?? ctx.runId) instead. Never a static string: two unrelated runs presenting one would re-enter each other's hold. See Re-taking Your Own Lease for the full rule.

What resets a running task ​

A long task run can be interrupted and started again by the platform, and it is worth knowing exactly what does that and what does not.

A config push never resets a running task. A run is pinned to the content hash of the version that started it, and it reloads that same immutable bundle at every wake. Pushing a new version mid-run creates a new version and points the function at it for the next call; the run in flight goes on executing the code it started on. Nothing in config push, config push --prune, functions archive or functions activate reaches a run that is already going.

A platform deploy does. The platform's own code — the primitive-functions SDK and the runtime around your handler — is not pinned: it is whatever is deployed at the moment your run wakes. Your handler executes inside the platform's own runtime, so when the platform is redeployed that runtime is reset under every slice that is executing, and the work in flight stops where it is. An isolate eviction does the same — the runtime may reset for memory, and a run sees the identical interruption. Neither is anything your code did.

What the engine then does. The reset does not lose the run. About five minutes later the engine dispatches your handler again, and:

  • your handler runs from the top;
  • every step.do that had already completed is replayed from its stored result — its body does not run again;
  • the step that was executing when the reset landed runs again, as a further attempt of that step.

That last line is the one with a consequence. The re-run is charged to the step's own retry budget — the engine's default is 5 attempts per step — so a step that lowers retries.limit spends the difference on platform deploys it did not ask for, and a step whose budget the reset exhausts does not run again: the run fails with the engine's own internal-error message (WorkflowInternalError: Attempt failed due to internal workflows error), and the platform records it as ENGINE_CODE_UPDATED — the deploy that ended the attempt, not the engine's own report of it. The message is unchanged: the code says what happened, the message says what the engine said. An eviction that runs a budget out the same way is recorded as ENGINE_ISOLATE_EVICTED, and a teardown the platform could not attribute to either — it compares the build the run's previous slice opened under with the one it wakes on, and two equal or missing identifiers prove nothing — keeps ENGINE_INTERNAL_ERROR, which is also what a run carries when no reset preceded the failure at all. If a step is short and idempotent, leave its retries alone; if you have lowered them, raise the limit by one for the deploys.

What your code may SEE when it is an eviction. A deploy takes the runtime away without warning your handler — the slice simply stops. An eviction can arrive the other way round, and this is the shape worth recognising: the object your handler is running on goes away while it keeps executing, so every step.do it makes from then on — including the ones in its own error handling — rejects immediately, in no time at all, with the runtime's own sentence:

Connection closed: this Durable Object instance is no longer active.

That rejection is not your step failing, and the run it belongs to is not over. The platform reads it as the teardown it is: the slice is recorded as reset, the row is left running, and the engine dispatches your handler again about five minutes later.

The consequence is for anything you wrote in a catch. Clean-up there — rolling something back, marking a job abandoned, writing a failure row of your own — runs against a run that will resume, so whatever it undoes, the replay will be doing again:

ts
try {
  await step.do("import", async () => { /* … */ });
} catch (err) {
  // This runs on an eviction too, and the run is not finished.
  // Anything recorded here is contradicted by the resumed run.
  await step.do("mark-abandoned", async () => { /* … */ });
  throw err;
}

Two rules follow. Keep such clean-up idempotent, the same way step bodies have to be — the resumed run may reach it again. And do not treat it as the run's ending: what the run finally settled as is on the run's own row, and a run that resumes and completes reads completed with its output even if the platform had already recorded a failure for the teardown. The row records the real outcome, not the interrupted one.

What it means for your code. Exactly what step discipline already asks, and a reset is the reason it is not optional: a side effect outside step.do runs again on a reset, and one inside a step that the reset interrupted may run twice — once in the attempt that was cut off, once in the attempt that replaces it. Make each step's body safe to repeat (an upsert rather than an insert, an idempotency key on a charge), and keep the value later steps read as the step's return value rather than something it wrote to a variable.

How to see it. A run that was reset and then finished reads completed with no failure and no error code, so the count is where the history lives:

bash
# the RESETS column: how many of this run's slices the platform tore down
primitive functions runs <function-id>

# which step was interrupted, and which ones replayed
primitive functions runs steps <function-id> <run-id>

# what the run printed, per slice
primitive functions logs <function-key> --run <run-id>

client.functions.getStatus(runId).slice carries the same three fields — resets, lastResetAt and lastResetCause (code-updated when the platform build changed underneath the run, evicted when the runtime said the isolate was reset, unknown when the platform saw a teardown it could not attribute). Treat the SDK as a stable interface rather than a frozen artifact, and treat a reset as ordinary: on a busy platform a run of several hours will meet one.

Budgets in a task run ​

The ceilings below apply per engine invocation, not per run. A run is a series of slices: the engine runs your handler from the top until it returns or reaches a sleep that has not finished, then hibernates; the next wake is a new slice, with completed steps replayed from storage.

  • A slice's wall clock is continuous, up to 12 hours; a step never starts short. Each slice is minted with a 10-minute credential, and every step begins with at least 5 minutes of it. When a step finishes with less than 5 minutes left, or is about to start with less than 5 minutes left, the platform REFRESHES the credential through the gateway — the same road every ctx.api call travels — and the slice's clock rolls straight on with no pause. So a single step may run up to 10 minutes (a two-minute LLM call, a slow integration, a large export all fit inside one step.do), and a sequence of steps runs as long as it needs, up to a 12-hour ceiling from the slice's start. The guarantee holds at every step.do boundary — including a cleanup step that runs after your handler catches a failed step, which begins on a refreshed budget rather than what the failed step left. Code outside steps spends the clock it runs in, and the next step.do entry refreshes it. A handler that never calls step.do — one long await — gets no refresh and dies at the 10-minute slice deadline.
  • When a refresh is refused, the slice falls back to a pause of a little over 5 minutes. A refresh is refused past the 12-hour ceiling, more than once a minute (one per 60-second window, counted on the platform's epoch-aligned clock, so two refreshes 59 s apart across a window boundary are both honoured), or during a rate-limiter or engine outage. The fallback is the old behaviour: the platform ends the slice with a hibernating sleep — the engine hibernates only for a sleep past its own five-minute grace period, so that is what the fallback sleeps — and the next wake mints a fresh 10-minute slice with a fresh 12-hour ceiling. So a refused refresh costs a little over five minutes, once; an ordinary run past a refusal simply continues on the new slice. Because the fallback is a hibernation, what the run printed before it is not recoverable — it is not on the log record (see Debugging a failing function). Steps running in parallel (Promise.all over two chains) that reach a fallback yield together: it waits for every step in flight to finish before it sleeps, because the engine hibernates only when nothing is running. On a local development server the engine never hibernates: a fallback there is a pause with no fresh credential, and a task run is bounded by one slice.
  • Two budgets do NOT refresh. Subrequests yield instead: they are counted across the slice and do not reset while it refreshes, and when the count passes half the resolved subrequest limit (5 000 of the 10 000 a task slice gets, or half a lower configured value) the next step.do boundary takes the fallback yield rather than refreshing, and the fresh slice resets the count. CPU is the other: a task slice gets 5 minutes of active CPU, and a step that exhausts it is killed by the runtime. The run does not recover from that kill — it fails, with the kill as its error — and the ceiling is per slice, not per step: a refreshed slice keeps counting. So a step's own body must fit in 5 minutes of CPU, though its wall clock may be far longer, and a run whose steps together need more than 5 minutes of CPU must end the slice between them — a step.sleep past the engine's five-minute grace period hibernates it, and the wake starts a fresh CPU budget.
  • The rate ceiling counts starts, not resumes.
  • The output ceiling and outputSchema apply once, to the final return value.

So step.sleep between chunks is for waiting, not for budget: chunk a long job by steps, and sleep when there is something to wait for. A refresh does not count toward the engine's maximum number of steps per run, and neither does a fallback yield — step.sleep and step.sleepUntil are excluded from that allowance, and only your step.do calls count.

A task run's status answer says how its slice is doing: a slice block with when the slice started, its 12-hour ceilingAt, refreshCount and lastRefreshAt, and — once the slice ended — settledAt and settledStatus. primitive functions runs shows the same count in its REFRESHES column, so whether a long run is refreshing or falling back is visible at a glance.

When a run's status settles ​

A task run settles its own row as it ends. The same settlement that writes the invocation record primitive functions logs prints writes the run's status, endedAt and output, so functions runs, functions logs, getStatus and runs wait agree about a run without being polled and without waiting for anything. You never need a staleness timeout of your own: a run that finished a second ago reads finished, and one that finished a day ago and was never polled reads the same — with the value it returned.

running means the run is live, or not yet known to have ended. A run executing, hibernating in step.sleep or waiting out a step retry reads running, and nothing that merely looks at it writes to it. That includes a run the engine is retrying after a platform deploy reset it: the reset is not a failure and the row is never settled for one, so such a run reads running through the five minutes before the engine picks it up again, and slice.resets is what says it happened.

The slice block is how you tell those two apart. settledAt is null while the slice is going and set, with settledStatus, once it ended. Read it one way only: a settled slice is proof the slice ended, and an unsettled slice is not proof that the run is live — the platform writes the invocation record before it settles the slice record, and a settlement it could not write down is logged rather than retried. So treat a running row with an unsettled slice as "not known to have ended", never as "definitely still going".

A running row whose run really did end is repaired on read. If a run's row is ever left behind — a start that died between writing its row and creating its instance, a settlement the platform could not write — the next read of it repairs it: primitive functions runs, getStatus, runs wait and the hard-delete scan each reconcile the run once, asking the engine first and then the platform's own records (the slice record, then the run's invocation record, which lives seven days). A run that ended is settled from what was written down about it, never guessed at.

Two things are fallbacks and nothing more: polling a run, and the platform's own 60-second sweep. Both still work, and neither is what makes a finished run read finished.

There is one deliberate exception, and it is why a healthy start is never mistaken for a lost one: a run whose row was written in the last five minutes and about which nothing has been recorded yet is treated as still starting. It is answered exactly as stored, it is never written, and a hard delete of its function is still refused — the run is about to load that function's code.

When a run fails ​

A run that failed is an answer, not an error: getStatus and waitFor (Starting one and polling it) resolve with status: "failed" rather than throwing, so branch on the status. Why it failed rides on the one structured error:

  • error.message — what went wrong, and where a platform refusal's code is: the message leads with it (OUTPUT_SCHEMA_VIOLATION: …, OUTPUT_TOO_LARGE: …), so match on the prefix. An engine failure is the exception: its message is the engine's own text, and its code is on error.code — see Run error codes.
  • error.code — the platform's classification of this failure, from the closed set in Run error codes. The same value the run row carries as run.errorCode. Branch on this rather than on the message.
  • error.name — the error's own name, not the platform's code. Your thrown error's class name (Error for a plain throw new Error(…)) where the engine preserved it, and absent for a failure read off the persisted row — which is the ordinary case once a run has settled, because the platform keeps only the message there. Treat it as a hint.
  • error.details — anything else the platform sent, when it sent any.

error is absent on a run that has not failed.

Run error codes ​

A failed run carries a code beside its message: run.errorCode on the run row, and error.code on the function run status both clients read. It is a closed set the platform owns, so you can branch on it rather than matching prose.

Two of them are about your handler, five are about the platform's own limits, and six — the ENGINE_* family — are about the engine that ran your code. Nothing in your code caused an ENGINE_* failure. You can branch on the whole family with a prefix check.

CodeWhat happenedWhat to do
FUNCTION_THREWYour handler threw.Read error.message and the run's logs (primitive functions logs <key> --run <runId>).
FUNCTION_NO_HANDLERThe entry file exported no callable default.Export a default function from the file entry names.
FUNCTION_TIMEOUTThe invocation did not finish inside its budget.Shorten the work, or run it as a task (functions start) where the budget is per slice.
FUNCTION_INPUT_INVALIDThe input did not match the declared inputSchema.Fix the caller's input, or the schema.
FUNCTION_OUTPUT_INVALIDA webhook delivery's run returned output that did not match outputSchema. (An HTTP invoke refusing its output answers OUTPUT_SCHEMA_VIOLATION in the envelope instead.)Fix the returned value, or the schema.
OUTPUT_SCHEMA_VIOLATIONA task run's output did not match outputSchema.Fix the returned value, or the schema. The message names the failing paths.
OUTPUT_NOT_SERIALIZABLEThe returned value is not JSON.Return plain data — no functions, no cycles, no undefined where a value is required.
OUTPUT_TOO_LARGEThe output exceeds the 1 MiB platform ceiling.Page the result, or write it to a database and return a handle.
FUNCTION_DELETEDThe function was hard-deleted while the run was starting.Nothing to retry: the stored bundle the run pinned is gone.
FUNCTION_RUNTIME_REFUSEDThe code refused the runtime it was handed (assertRuntime), or a request invocation's step was asked for a wait its budget cannot cover.Call it the other way — invoke for a request, start for a task.
ENGINE_ISOLATE_EVICTEDA platform component the invocation depended on went away mid-call, or the isolate was reset for memory — including an eviction that left the attempt that followed it no retry, which carries the engine's own internal-error message.Nothing in your code caused it. To run the work again, start a new run with a new runKey and the same input — a repeated key replays the failed run and executes nothing — and make any side effect the failed attempt already performed idempotent in your own code. Report a repeat to the platform.
ENGINE_CODE_UPDATEDA platform deploy reset the engine while the attempt ran, and the attempt that followed had no retry left: the run carries the engine's own internal-error message (what resets a running task).Nothing in your code caused it. Start a new run with a new runKey and the same input; a repeated key replays and executes nothing, and side effects the failed attempt performed need your own idempotency.
ENGINE_STORAGE_ERRORThe engine's storage timed out or failed and the object was reset.Nothing in your code caused it. Start a new run with a new runKey and the same input (a repeated key replays), with your own idempotency over side effects already performed. Report a repeat to the platform.
ENGINE_INTERNAL_ERRORThe engine reported an internal error on the attempt, and no reset of this run preceded it — or the platform saw one it could not attribute to a deploy or an eviction.Nothing in your code caused it. Start a new run with a new runKey and the same input (a repeated key replays), with your own idempotency over side effects already performed. Report a repeat to the platform.
ENGINE_SLICE_DEADLINEA platform call was refused because the slice's token deadline had passed.The slice outlived its token. Look for a step that runs longer than the slice ceiling and split it.
ENGINE_INSTANCE_LOSTThe engine no longer reports the instance, and nothing recorded how the run ended.primitive functions logs <key> --run <runId> shows what the platform did record. Start a new run with a new runKey if the work still needs doing.

An engine failure's errorMessage is the engine's own text, not a code-leading sentence: for these the code is on error.code and run.errorCode, never in the message's prefix.

An attempt whose engine died before the platform could write its invocation record leaves no record at all — the writer died with it. For such an attempt the run row's code is the whole statement; there is nothing in functions logs to read.

Reading a task run step by step ​

functions runs answers at the level of the run: status, the runtime it used, start, end, the failure code. For a task run that is rarely enough — the handler's work is in its named steps, and what you want to know is which one it is on, which ones already finished, and where the time went.

bash
primitive functions runs steps <function-id> <run-id>

Each step comes back with its name, its kind (function.step for a step.do, function.sleep for a step.sleep or step.sleepUntil, function.event for a step.waitForEvent), its status, the idle GAP before it, and its duration. A step that is still executing reads running with the time it has been running so far — including a run that is asleep, whose function.sleep step stays running until it wakes. A step that threw reads failed with its message.

The rows are written as the run goes, so they are readable while it is still in flight, and they survive the hibernation of a step.sleep. A resumed run replays its completed steps from the engine's memo rather than re-running them: each step keeps one row, with the timing of the slice that actually executed it. A step that a platform deploy reset keeps that same one row too — it stays running, not failed, until the replay finishes it, and its duration then spans the reset rather than stopping at it. --json emits the same rows as JSON.

Ending a run that will not settle ​

bash
primitive functions runs terminate <function-id> <run-id>

Stops the run where it is and settles its row terminated. Both halves happen, and in that order: stopping the instance without settling the row would leave every reader still reporting running, and settling the row without stopping the instance would leave your code executing — and the next status read reconciles the row against the engine and rolls it forward again.

A run whose instance is already gone — aged out of the engine's retention, or discarded with the document it was working on — still settles, because that is the run you came here to end. A stop that fails for any other reason is an error, and the run is left exactly as it was: the engine is the only thing that knows whether your code is still executing, and a row settled over a live run would both mislead you and remove the retry, since a settled run is reported rather than stopped. Run the command again.

A run that has already finished is reported, not overwritten: the command prints the status it settled at and terminates nothing. Its outcome is the only record of what it actually did — and that is asked of the engine, not only of the run row. A task run settles its own row as it ends, but a settlement that could not be written leaves the row stored running after the run finished; terminating such a run records what it really did instead of burying it under terminated.

Any step still open when the run is terminated is closed with it, however the run was stopped — this command or client.functions.terminate — so the trace of a run stopped mid-step.sleep does not go on reporting that sleep as running. Running the command against a run that is already over closes anything its trace still has open, so a trace left behind by an older stop is repaired by asking again.

Ceilings ​

Every function runs under fixed platform ceilings. [function.limits] may lower any of them and can never raise one — a config asking for more simply gets the platform value.

The two runtimes differ in what they are allowed to spend, which is most of why the runtime matters at all:

LimitRequest runtimeTask runtimeKey
CPU5 000 ms per invocation5 minutes per slicecpuMs (see below)
Outbound subrequests12810 000 per slicesubRequests
Wall clock5 000 ms default, 30 000 ms ceilingcontinuous up to 12 hours; every step begins with at least 5 minutestimeoutMs on the request (request only); see Budgets in a task run
— a step.sleep has to fit the budgetpast it, the invocation fails FUNCTION_RUNTIME_REFUSEDhibernates; hours and days are fine— (start the function as a task)

And the ones that are the same under both:

LimitPlatform valueKey
Invocations per minute, per function1 200ratePerMinute
Response size1 MiB, kept on the run row—
Values per $in / $nin list1 000—

A task run's output is held to the same 1 MiB and is kept on the run row for its life, so it outlives the engine's instance. Over roughly 10 000 characters it is stored by reference and still answered whole; a runs page holding many large outputs may be shorter than the limit you asked for and continue with a cursor.

A burst past the rate ceiling answers 429 with errorCode: "FUNCTION_RATE_LIMITED" and a Retry-After header. Buckets are per app and per function, so one function's traffic never consumes another's. A slot is taken only once a call is going to run: a caller your access rule rejects, or a request your inputSchema refuses, consumes none of the budget. When the limiter itself cannot be reached the call is refused with 503FUNCTION_RATE_UNAVAILABLE rather than run unmetered.

cpuMs is the one ceiling the platform does not hold to the number you write. It is enforced — a function past its budget is stopped, and the invocation answers failed — but the runtime adds an allowance of its own before it stops anything, measured at roughly two seconds. So cpuMs = 50 and cpuMs = 200 buy the same thing in practice, and neither stops your code in fifty or two hundred milliseconds. Lowering cpuMs still lowers the ceiling where the value is large enough to matter; below the allowance it is a declaration of intent rather than a cap you can rely on. If you need a function bounded tightly, bound it with timeoutMs on the request, which is wall clock and is held exactly.

The response-size ceiling is the platform's, not your sandbox's: the platform reads your response with a byte cap and measures the parsed output (and error) with its own encoder. An output over 1 MiB is {"status":"failed", "errorCode":"OUTPUT_TOO_LARGE"} whatever the code inside the sandbox did — page your results or write them to a database and return a handle.

The wall-clock budget starts when the platform accepts the invocation, so loading a function's code the first time comes out of it — and the credential the sandbox carries expires on that same clock, never later than the answer the caller sees.

What a timeout does, and does not, stop

When the budget runs out the caller gets {"status":"timeout"} and the sandbox's request is aborted. The platform cannot kill a running sandbox — the runtime offers no such control — so a function that ignores the abort may keep executing until its cpuMs ceiling stops it.

What it can no longer do is act. The invocation's credential carries an absolute deadline, checked independently on every platform call, so a function that outlives its own invocation is holding something that has already stopped working: its late calls are denied and its side effects never land. Write functions that finish inside their budget; do not rely on work continuing after a timeout, and do not fear it either.

The same deadline bounds what your function starts outside the platform. An ctx.integrations.call is cut off at whichever comes first, the integration's own timeout or what is left of the invocation — including while its response body is still arriving, not merely while it waits for a reply — and answers UPSTREAM_TIMEOUT. A ctx.prompts.run whose provider call passes the deadline answers PROMPT_UPSTREAM_TIMEOUT, which is worth catching separately from an ordinary model failure: it means the budget ran out, not that the model did.

Documentation validated against js-bao-wss-client 3.4.0 · js-bao 0.11.0 · primitive-admin 1.0.62 · primitive-app 3.1.0 — 2026-09-30