Prompts
A prompt is a versioned LLM prompt template kept in your config tree: a prompts/<key>.toml file holding the template text and one or more configurations — provider, model and settings. A server function runs it by key with ctx.prompts.run; the platform renders the template, calls the configured model on the app's own credentials, and hands back the result. Because the prompt lives in config rather than in your function's code, you can iterate on wording, switch models, and check behavior with test cases without touching the code that calls it.
Creating a Prompt
Define a prompt in TOML and push it with config push. The [prompt] table carries its identity; the template text lives on a configuration alongside the provider and model:
# primitive/dev/prompts/summarizer.toml
[prompt]
kind = "chat"
key = "summarizer"
displayName = "Document Summarizer"
[[configs]]
name = "default"
active = true # exactly one entry carries this — it is the config the server runs
provider = "gemini"
model = "models/gemini-3-flash-preview"
[configs.chat]
temperature = 0.5
systemPrompt = "You are a skilled summarizer. Create clear, accurate summaries."
userPromptTemplate = """
Summarize the following text in a {{ input.style || 'brief' }} style:
{{ input.text }}"""primitive config pushBeyond the keys above, a [configs.chat] block accepts topP, maxTokens, outputFormat, reasoningEffort or reasoningBudget (see Bounding the reasoning budget) and a [configs.chat.outputSchema] table. The [[configs]] entry itself carries the keys that are not specific to a kind — name, status, description, provider, model and a [configs.providerConfig] table, which is stored with the config and returned by the API but not applied to the provider request. The prompt itself accepts kind, description and the [prompt.inputSchema] / [prompt.outputSchema] tables. Every one of them round-trips: config pull writes back what the server holds, and a key the CLI does not recognize is reported rather than dropped.
config push applies the complete file. A [prompt] table with no outputSchema key clears any output schema stored on the server, and deleting an optional key of a [[configs]] entry or its [configs.chat] / [configs.decisions] block — systemPrompt, temperature, outputFormat, questions, a reasoning key, and so on — and pushing clears it rather than leaving the stored value live. The entry list itself is the exception: push creates and updates the [[configs]] entries the file lists, but deletes none it omits and clears no activation. Push warns that such a prompt still differs from the server, and config diff keeps reporting it until you config pull or remove the configuration server-side.
Prompt kinds
[prompt].kind says what the prompt runs. It is chat by default — a chat completion rendered from templates, which is every prompt written before the key existed — or decisions, which names OpenRouter's decisions endpoint and its named typed questions. The kind decides the input envelope and the output shape, both of which are asserted per prompt key, so it is fixed at creation: converting a prompt means a new key, and config push declines a changed kind as immutable before it sends anything.
A [[configs]] entry may carry only the block named by its prompt's kind:
kind | block | keys |
|---|---|---|
chat | [configs.chat] | systemPrompt, userPromptTemplate, temperature, topP, maxTokens, outputFormat, outputSchema, reasoningEffort, reasoningBudget, strictOutput |
decisions | [configs.decisions] | questions |
Nothing else about an entry changes with the kind. A [configs.chat] block under a decisions prompt — or a [configs.decisions] block under a chat prompt — is refused at push and at the admin routes alike, naming the block and the kind.
config set does not write a key inside a [[configs]] entry, of either kind — it never has, for any repeated table. It writes prompt.kind and every other [prompt] scalar; edit the blocks in the file.
Running a Prompt From a Function
A function runs the prompt by its key. There is nothing to declare in the function's TOML — no capability line names a prompt:
# primitive/dev/functions/summarize.toml
[function]
key = "summarize"
entry = "functions/summarize/index.ts"
access = "true"// primitive/dev/functions/summarize/index.ts
import { defineFunction } from "primitive-functions";
export default defineFunction(async (input: { text: string }, ctx) => {
const result = await ctx.prompts.run("summarizer", {
variables: { text: input.text, style: "bullet points" },
});
if (!result.success) throw new Error(result.error ?? "prompt failed");
return { summary: result.output };
});The answer is an envelope: success, output (the generated text), error, metrics (durations and token counts) and configId, which names the configuration that ran. The second argument takes variables, a modelOverride for this call only, and a configId to run a specific configuration of this prompt instead of the active one. The model call is bounded by the function's own remaining time. A failure adds upstreamStatus when a provider answered — see when the answer or the call fails.
Who may run it is the function's business. A prompt has no endpoint of its own that a client can call: the function's access gate decides who may make the code run, and that is the whole authorization. What the prompt contributes is its availability — only an active prompt runs, and an inactive or archived one is refused.
Template Variables
You supply values under variables when running a prompt; the template reads them under input:
{{ input.text }} # Caller-provided variable
{{ input.style || 'brief' }} # With fallback
{{ input.items[0].name }} # Nested accessAn unresolved reference fails the render rather than rendering empty, so guard an optional path with a || fallback.
Typed, Structured Output
When a prompt answers JSON, declare its shape once on the prompt and stop parsing it by hand. [prompt.outputSchema] is sent to the provider, the answer is validated against it, and config push renders it into functions/primitive-prompt-types.d.ts so the function's call is typed from it:
# primitive/dev/prompts/categorize.toml
[prompt]
key = "categorize"
displayName = "Categorize"
[prompt.outputSchema]
type = "object"
required = ["category"]
[prompt.outputSchema.properties.category]
type = "string"const answer = await ctx.prompts.run("categorize", {
variables: { text: input.text },
});
if (!answer.success) throw new Error(answer.error ?? "categorize failed");
// `parsed` is the validated JSON value, typed from the schema above.
return { category: answer.parsed.category };parsed appears on the success arm only, which is why the success check is what unlocks it. A field the schema does not declare is a compile error, and so is ctx.prompts.run("no-such-prompt") in a pushed tree.
When the answer or the call fails
Two things can go wrong with the answer itself, and both are success: false with an errorCode — the run happened and was billed, so output, metrics and configId are still there:
errorCode | What happened |
|---|---|
PROMPT_OUTPUT_NOT_JSON | The prompt is declared to answer JSON and the model answered text that does not parse — or that parses to a number JSON cannot represent, such as one that overflows. output has the text. |
PROMPT_OUTPUT_SCHEMA_VIOLATION | It parsed, and the declared outputSchema refuses it. error names the failing paths (value.word must be string). |
A provider failure — the model call itself failing — is success: false with error set and no errorCode, except for the one case below. That is how you tell "the model returned the wrong shape" from "the model failed": an errorCode starting PROMPT_OUTPUT_ is always about the answer.
When the provider answered, the failure also carries upstreamStatus, the provider's own HTTP status, and error names it too — so a 504 after two minutes is distinguishable from a 400 on a refused payload without parsing the message. It does not depend on which provider the configuration names: Gemini and OpenRouter report it the same way. Retry on 408, 429, 502, 503 and 504 — 408 and 504 are the upstream timeouts — and expect any other 4xx to fail the same way next time. The field is absent when there was no provider answer to report: an unset provider key, or a call that never got one.
An upstream timeout is the single provider failure that does set an errorCode, PROMPT_UPSTREAM_TIMEOUT — the same code the platform answers when the invocation's own deadline passes before the call could run, because to a caller the two are the same event and call for the same decision.
One more code is set before the model is ever called: a decisions run whose per-run options are refused answers success: false with errorCode: "PROMPT_CRITERIA_INVALID" and nothing billed — see Options supplied per run.
A validation error names paths and the schema's own constraints and never quotes the model's answer; output is where the answer is, and that is deliberate, so a diagnostic in a log cannot carry generated content.
outputFormat = "json" on the config that runs is enough to get parsed when the text parses, but nothing declared its shape, so it arrives as unknown. Declare [prompt.outputSchema] to get a type. A chat config may carry its own [configs.chat.outputSchema] table too; ctx.prompts.run validates against the prompt's, not the config's.
The generated declaration
functions/primitive-prompt-types.d.ts carries <Key>PromptOutput for every prompts/<key>.toml that declares a [prompt.outputSchema], and a PromptSchemas augmentation keyed by prompt key — that is what types ctx.prompts.run("<key>"). config push regenerates it on every push, beside the other declarations it writes for functions.
A prompt that declares no schema is listed with an empty entry, so its parsed stays optional and unknown — which means removing a schema turns a handler reading parsed.field into a compile error on the next push rather than a runtime undefined.
The prompt key is the one push deploys: the file's [prompt] key when it declares one, its file name otherwise. A file named local.toml carrying key = "categorize" is typed under categorize, because that is the prompt the platform answers on.
Prompt Configurations
A prompt can have multiple configurations for different providers, models, and settings. One configuration is active; a run uses it unless it names a specific configId. This is how you A/B a cheaper model against a higher-quality one without touching the prompt body.
Configs are [[configs]] entries in the prompt's own file — one writer, visible in review:
# prompts/summarizer.toml
[[configs]]
name = "fast"
active = true
provider = "openrouter"
model = "gpt-4o-mini"
[configs.chat]
temperature = 0.3
userPromptTemplate = "Summarize: {{ input.text }}"
[[configs]]
name = "quality"
provider = "gemini"
model = "models/gemini-2.5-pro"
[configs.chat]
temperature = 0.7
userPromptTemplate = "Summarize: {{ input.text }}"primitive config push --only prompt/summarizer # apply the configs
primitive prompts list # each prompt's ID
primitive prompts configs list <prompt-id> # what the server runs
primitive prompts configs get summarizer fast # one config, in full
primitive prompts disable <prompt-id> # stop serving it now
primitive prompts enable <prompt-id>
primitive prompts archive <prompt-id> # archive it (soft delete)active = true marks the live config — move it to switch, and push. Take a config out of service with status = "archived" rather than deleting its block, so the change is a diff someone can read.
Note the two different things called "status". A named config's status (active | archived) says which version is out of service, and it is configuration: you author it here. The PROMPT's availability is not — see Availability is server-owned.
Archiving the whole prompt — as opposed to one of its configs — is primitive prompts archive <prompt-id>, covered in Archive vs prune: ctx.prompts.run and admin execution both refuse an archived prompt, and its configs, executions and analytics keep resolving. Archiving the prompt does not touch the per-config [[configs]] status in the file, which stays yours to author.
Testing Prompts
Define test cases to validate prompt behavior before you activate a change. A case is a TOML file beside the prompt, applied by config push:
# prompts/summarizer.tests/basic-test.toml
[test]
name = "basic-test"
inputVariables = '{"text": "Long article text...", "style": "bullet points"}'
expectedOutputContains = '["•"]'primitive config push --only prompt/summarizer
primitive prompts tests list summarizer
primitive prompts tests run-all summarizerThe prompts tests commands take the prompt's key (the name in prompts/<key>.tests/ and beside the id in prompts list) or its id. A value that names no prompt in the app exits non-zero with No prompt 'summarizr' in app <app-id> rather than an empty list, so "No test cases found." only ever means the prompt exists and has no registered cases.
primitive config fields prompt lists every key of the sidecar. Deleting a case is removing its file and running primitive config push --prune. The file's name is the case's identity — see Test Case Identity. run-all runs the registered cases, not whatever is on disk — the same page covers the push step that closes that gap and config diff's counters.
Verification types include substring contains, regex pattern, JSON subset, and LLM-as-judge.
Trying a Prompt From the CLI
primitive prompts preview renders the template without calling the model, and primitive prompts execute runs it — both are admin diagnostics:
primitive prompts preview <prompt-id> --vars '{"text":"Long article text..."}'
primitive prompts execute <prompt-id> --vars '{"text":"Long article text..."}' --config <config-id>primitive prompts execute deliberately runs an inactive prompt too, so you can trial one before primitive prompts enable puts it in service; ctx.prompts.run refuses an inactive prompt with PROMPT_NOT_EXECUTABLE, in a message naming that command.
Bounding the Reasoning Budget
On a reasoning-by-default model, thinking is most of what a call costs — in latency as well as money. Measured on a merchant-categorisation prompt (google/gemini-3.6-flash through openrouter), reasoning was 49–82% of the output tokens on every call, and the duration tracked the output count almost linearly. A classification task does not want that, and switching models is not the answer: it re-opens accuracy on a prompt whose test cases are tuned to the current one.
A config states the budget in one of two ways, never both:
[configs.chat]
reasoningEffort = "minimal" # none | minimal | low | medium | high
# reasoningBudget = 512 # ...or a whole number of reasoning tokensThe setting belongs to the NAMED config, so a reasoning and a non-reasoning variant of the same prompt are two [[configs]] entries, compared through the ordinary test cases like any other A/B.
Nothing is silently dropped. Each value either maps onto the provider's own spelling or is refused — at push, or at execution when a configId override sends the prompt to a different model:
| Provider / family | reasoningEffort | reasoningBudget |
|---|---|---|
openrouter | OpenRouter's reasoning.effort; none becomes reasoning: { enabled: false } | reasoning.max_tokens |
gemini, 3.x | thinkingConfig.thinkingLevel. Gemini 3 cannot turn thinking off, so none is refused; Pro takes low and high only | Refused — Gemini 3 translates a budget to a level rather than honoring it as a bound. Use reasoningEffort |
gemini, 2.5 | Only none, which means a budget of 0. Any other level is refused | thinkingConfig.thinkingBudget. Gemini 2.5 Pro cannot stop thinking: its floor is 128 tokens |
gemini, 2.0 and earlier | Refused — the model has no reasoning control | Refused |
Two caveats before you rely on a number:
- On
openrouter,reasoningBudgetis an exact bound only on budget-native models; effort-only models translate it to the nearest effort level. The budget is a request, andmetrics.reasoningTokensis the record of what was spent. - Every openrouter request carrying a reasoning setting also carries
provider: { require_parameters: true }, which keeps it away from endpoints that would accept the request and ignore the setting. A model that cannot honor what you asked for therefore fails the execution with the provider's own reason — "Reasoning is mandatory for this endpoint and cannot be disabled", say — instead of quietly running without it.
metrics.reasoningTokens is how you measure the effect. Every path that returns metrics reports it separately from outputTokens: ctx.prompts.run in a server function, the admin execute endpoint, and primitive prompts execute. It is the only honest number — OpenRouter counts reasoning INSIDE its output tokens and Gemini counts it OUTSIDE, so neither headline figure says how much of the decode was deliberation. It is absent when the provider reports none, and 0 when a setting successfully declined it.
Decisions models
OpenRouter serves TypeSafe's System One models (typesafe/jev-1.13) on a different endpoint from chat completions: instead of a rendered template it takes a state value plus named typed questions, and answers each question with a typed answer, its probabilities and a confidence. They are fast and cheap — a transaction categorizer measured 0.24 s median latency at about a seventh of a chat model's cost — and a kind = "decisions" prompt puts that call inside the prompt system, with its config, versioning, test cases and analytics, rather than in an app-owned integration holding its own OpenRouter key.
# config/prompts/transaction-categorizer.toml
[prompt]
kind = "decisions"
key = "transaction-categorizer"
displayName = "Transaction categorizer"
[[configs]]
name = "jev"
active = true
provider = "openrouter" # the only provider a decisions config runs on
model = "typesafe/jev-1.13"
[configs.decisions.questions.category]
type = "choice"
instructions = "Which household category does this bank transaction belong to?"
criteriaSource = "static" # the options are the table below
[configs.decisions.questions.category.criteria]
gas-and-fuel = "Gas & Fuel (Auto & Transport)"
groceries = "Groceries (Food & Restaurants)"A decisions config has no userPromptTemplate and no systemPrompt: the questions ARE the prompt.
state is the input
A decisions run takes one variable, state — the value the questions are asked about. It is passed to the provider verbatim, whatever its JSON shape, and nothing is templated:
const answer = await ctx.prompts.run("transaction-categorizer", {
variables: {
state: {
transaction: {
bank_description: "CHEVRON 0093847",
amount: -61.2,
date: "2026-08-11",
},
},
},
});variables.state is required and is read by presence, not by truthiness: null, 0 and "" are values the endpoint may legitimately be asked about. A run that omits the key altogether is refused before the provider call, so nothing is billed — success: false with an error naming variables.state, identically from ctx.prompts.run, the member execute route and primitive prompts execute. attachments are refused the same way: they are a chat-completion input, and a decisions request has only state.
The answers
output is the answers object serialized, and a server function receives it as parsed with no declaration needed — a decisions run always answers JSON. metrics carries inputTokens, outputTokens, totalTokens and metrics.cost, the price of the call in USD as the provider reported it (absent when the provider reports none).
There are three question types. To get parsed typed, declare the matching [prompt.outputSchema] — typed access comes only from that declaration, never from the questions, because any active config of the prompt can be pinned by configId and only the prompt-level schema is enforced at run time.
choice — pick one of the criteria keys. criteria is a table of at least two option key = description pairs, each description a non-empty string or a non-empty JSON object, and every choice question declares where it comes from with criteriaSource: "static" for a table in the config, as above, or "dynamic" for options each run supplies (see Options supplied per run).
{ "type": "choice", "choice": "gas-and-fuel",
"probabilities": { "gas-and-fuel": 1, "groceries": 0 }, "confidence": 1 }[prompt.outputSchema]
type = "object"
[prompt.outputSchema.properties.category]
type = "object"
[prompt.outputSchema.properties.category.properties.type]
const = "choice"
[prompt.outputSchema.properties.category.properties.choice]
enum = ["gas-and-fuel", "groceries"]
[prompt.outputSchema.properties.category.properties.probabilities]
type = "object"
[prompt.outputSchema.properties.category.properties.confidence]
type = "number"score — a number between the two ends of a scale. Here the criteria is an array of exactly the two endpoint descriptions, low first; the endpoint answers 400 without it. The answer carries the legend it scored against beside the score:
{ "type": "score", "score": 0.01,
"legend": { "0": "not unusual at all", "1": "extremely unusual" },
"probabilities": { "0": 0.99, "1": 0.01 }, "confidence": 0.98 }[configs.decisions.questions.unusual]
type = "score"
instructions = "How unusual is this transaction for a typical household?"
criteria = ["not unusual at all", "extremely unusual"]noul — a bare number, with no confidence:
{ "type": "noul", "noul": 0.81 }Any other key on a question is passed through to the provider, which validates it — except criteriaSource, which is the platform's own and is never sent.
Options supplied per run
Some questions only have options at run time: "which of this household's previous transactions is the same merchant as this one" has a different set of candidates for every transaction. Declare such a question criteriaSource = "dynamic", with no criteria table, and pass the options with each run as variables.criteria.<question>, beside state:
# config/prompts/precedent-matcher.toml
[prompt]
kind = "decisions"
key = "precedent-matcher"
displayName = "Precedent matcher"
[[configs]]
name = "jev"
active = true
provider = "openrouter"
model = "typesafe/jev-1.13"
[configs.decisions.questions.precedent]
type = "choice"
instructions = "Which previous transaction is the same merchant as this one, if any?"
criteriaSource = "dynamic" # each run must pass variables.criteria.precedent
[configs.decisions.questions.category]
type = "choice"
instructions = "Which household category does this bank transaction belong to?"
criteriaSource = "static" # the table below is required; a run may not supply one
[configs.decisions.questions.category.criteria]
gas-and-fuel = "Gas & Fuel (Auto & Transport)"
groceries = "Groceries (Food & Restaurants)"
other = "None of the above"const answer = await ctx.prompts.run("precedent-matcher", {
variables: {
state: { transaction },
criteria: {
precedent: {
p1: { bank_description: "CHEVRON 00938", category: "gas-and-fuel", amount: -42.1, date: "2026-07-30" },
p2: { bank_description: "SHELL 4411", category: "gas-and-fuel", amount: -50.0, date: "2026-07-12" },
none: "No previous transaction is the same merchant",
},
},
},
});The config still owns each question's name, type and instructions, so a run cannot change what is asked — only the options it chooses between. The provider is asked exactly the config's questions, with the run's table as each "dynamic" question's criteria, and the answer comes back typed as for any choice question. The provider's answer is passed through as it came, so a function that acts on choice checks it against the keys it supplied.
A run's options follow the same rule as a config's table — at least two; option keys non-empty with no leading or trailing whitespace; each description a non-empty string or a non-empty JSON object — and are checked before the provider call, so a malformed run is never billed. These all answer success: false with errorCode: "PROMPT_CRITERIA_INVALID", no tokens or cost in metrics, and an error naming the question and the rule:
- a
"dynamic"question with novariables.criteria.<question>, or a table the rule refuses; - options for a
"static"question, for ascoreornoulquestion, or for a name the config does not declare; - a
variables.criteriathat is not an object; - a stored
choicequestion with nocriteriaSource(one written before the key existed) — declare it and push again.
criteriaSource is required on every choice question: push, and the admin create and update routes, refuse one without it, a "static" question without a table, a "dynamic" question with one, and criteriaSource on a score or noul question, naming the question. On a chat prompt variables.criteria is an ordinary variable ({{ input.criteria }}).
Testing a decisions prompt
A test case carries state in its inputVariables and asserts on the answers with expectedJsonSubset — the chosen option is what you pin:
# prompts/transaction-categorizer.tests/chevron.toml
[test]
name = "chevron"
inputVariables = '{"state": {"transaction": {"bank_description": "CHEVRON 0093847", "amount": -61.2, "date": "2026-08-11"}}}'
expectedJsonSubset = '{"category": {"choice": "gas-and-fuel"}}'A case for a "dynamic" question supplies the run's options the same way, beside state — a case's inputVariables is the run's variables:
# prompts/precedent-matcher.tests/chevron-repeat.toml
[test]
name = "chevron-repeat"
inputVariables = '{"state": {"transaction": {"bank_description": "CHEVRON 0093847", "amount": -61.2, "date": "2026-08-11"}}, "criteria": {"precedent": {"p1": {"bank_description": "CHEVRON 00938", "category": "gas-and-fuel", "amount": -42.1, "date": "2026-07-30"}, "none": "No previous transaction is the same merchant"}}}'
expectedJsonSubset = '{"precedent": {"choice": "p1"}}'A case whose options are refused is recorded as a failed run whose one failed check names the question and the rule — on primitive prompts tests run, tests run-all and tests batch start alike — and no provider call is made, so the run history says why the case failed.
A decisions prompt may not be a test case's evaluator: an evaluator prompt is handed the evaluated input and output as template variables, which a decisions run cannot read, so an evaluator must be a chat prompt. Naming one is a 400 at test-case create and update.
Next Steps
- Server Functions — Run a prompt from server-side TypeScript with
ctx.prompts.run - Analytics — Prompt executions are tracked automatically (
prompt.executed, durations, token counts) - Primitive CLI — Full CLI command reference