AIP-15: WORKFLOW.md — agentworkflow/v1 (abstract orchestration manifest)
A markdown + frontmatter format for declaring a multi-step agent workflow's abstract orchestration shape — its steps, branching, parallelism, approval gates, suspend/resume, and compensation. Pairs with the standard `defineWorkflow` / `defineStep` signatures. Implementation lives entirely in the per-step TOOL.md contracts and their AIP-30 DRIVER bindings; workflows themselves are pure orchestration data.
| Field | Value |
|---|---|
| AIP | 15 |
| Title | WORKFLOW.md — agentworkflow/v1 (abstract orchestration manifest) |
| Status | Draft |
| Type | Schema |
| Domain | workflows.sh |
| Requires | AIP-7, AIP-14 (TOOL), AIP-16 (IO), AIP-30 (DRIVER), AIP-58 (RUN — agent step outcome) |
| Resources | ./resources/aip-15 — SKILL.md, ADAPTER.md, WORKFLOW.schema.json, EXAMPLES.md |
Abstract
WORKFLOW.md is a markdown-with-frontmatter file format that packages
a single multi-step agent workflow — its identity, input/output
schemas, ordered + branching + parallel steps, approval gates,
suspend/resume points, retry + compensation policy, and resource
budget. The format is paired with two standard entry-point functions,
defineWorkflow(...) and defineStep(...), whose signatures
any implementation in any language exposes so callers, runtimes, and
adapters share one contract.
It is the workflow analogue of TOOL.md: same posture (human-authored, version-controlled, machine-parseable), same manifest+signature pairing, same governance hookup via AIP-7.
Motivation
A workflow's logical shape — what steps run in what order, what triggers a branch, what causes a wait, who approves a risky step — is independent of any specific runtime's concurrency model, persistence backend, or syntactic flavour. Authoring a workflow once SHOULD give the same logical run anywhere.
Every workflow runtime needs the same small core of primitives:
named steps, sequence, branching, parallelism, suspend/resume,
retry, compensation. The same shapes show up under different verbs
and different syntaxes. WORKFLOW.md extracts the logical shape
into a portable file; defineWorkflow + defineStep are the
standard entry points implementations expose so the manifest's entry
loads the same way under any host.
The manifest is also the source of truth for governance
(AIP-7) audit, work-item tracking
(AIP-13), and skill-based agent authoring (the
companion SKILL.md generates both TOOL.md
and WORKFLOW.md files when an agent is asked to build automation).
Design principles
-
Steps over choreography. The manifest declares a finite set of named steps and a directed acyclic graph (DAG) of relationships between them. No free-form code in the manifest — the DAG is data. Step bodies live in
entryfiles; the manifest references them. -
Approval is a step kind. A workflow that pauses for human approval has a step of
kind: "approval". It's not a runtime concern bolted onto a tool step — it's a first-class node so governance and the planner reason about it the same way. -
Branching is data, not control flow. A
kind: "branch"step declares conditions and target step ids. The runtime evaluates, the manifest doesn't run. This keeps the spec decidable and the adapter trivial. -
Parallelism is explicit. A
kind: "parallel"step lists child step ids that run concurrently. No implicit fan-out from input shapes (which Mastra and others sometimes infer). -
Suspend points are named. A workflow can pause at any
kind: "suspend"step, persist run state, and resume later when a matching event arrives. The manifest names the events; the runtime wires the listeners. -
Compensation is opt-in. Each step MAY declare
compensation: <step-id>whose body undoes the step's effect. Sagas are an alignment of these references, not a separate primitive. -
Unopinionated about runtime. No fields named
mastra*,temporal*, etc. Runtime concerns live in adapter docs.
Specification
File location
Workflows live in a single folder:
.workflows/
invoice-approval/
WORKFLOW.md ← this AIP
workflow.ts ← optional entry (step bodies referenced here)
PROCEDURE.md ← optional vendor-neutral procedure (agencies.sh)
README.md ← optional long-formThe folder name SHOULD match the manifest's id. Workflows MAY
reference tools via TOOL.md ids without an entry
file at all; the manifest then defines the workflow purely
declaratively.
Frontmatter
YAML frontmatter, delimited by --- lines. Case-sensitive.
Required fields
| Field | Type | Description |
|---|---|---|
name | string | Display name (1–80 chars). |
id | string | Machine identifier. Lowercase, digits, dashes. 2–64 chars. Unique within the registry. |
description | string | One-paragraph purpose, written for the LLM caller. ≤2000 chars. |
version | semver | Spec version of THIS file. Bump on incompatible step graph or schema changes. |
inputs | JSON Schema | Workflow-level inputs. Available to any step via $input.*. |
outputs | JSON Schema | Workflow-level output shape. Describes what result produces (or, when result is omitted, the final step's output). |
result | object | string | Optional output value expression: maps step results to the workflow output using the same reference grammar as a step's inputs ($input.*, $steps.<id>.*, literals). Omitted ⇒ the output is the final step's result. Its shape SHOULD satisfy outputs. |
steps | object[] | Ordered array of step declarations (see Steps). |
Optional fields
| Field | Type | Default | Description |
|---|---|---|---|
start | step id | first step | The step the runtime executes first. |
suspendable | boolean | derived | true iff any step has kind: "suspend" or kind: "approval". Adapters MAY use this to pick storage (durable vs in-memory). |
triggers | object[] | [] | Events that start the workflow. See Triggers. Legacy form — for recurring schedules prefer the new routines: field (decoupled, see below). |
routines | (ref | inline)[] | [] | AIP-41 ROUTINE.md refs that fire this workflow. Decoupled from the workflow body — one schedule can drive N workflows; one workflow can be driven by N schedules. Each entry follows the inline | ref | file pattern. Preferred over inline triggers: [{ kind: schedule }]. |
requires | object | {} | Capability requirements per AIP-7. Same shape as TOOL.md. |
approval | string | "per-step" | Default approval class applied to every step lacking its own approval. Same vocabulary as AIP-14: "auto", "always", "on-mutate", "policy:<ref>". |
risk_level | 0–3 | max of step risks | Workflow-level autonomy gate. Used when the whole workflow is dispatched via a parent agent. |
timeout_ms | int | 600000 (10 min) | Hard wall-clock cap on the entire run. |
max_steps | int | 100 | Defence against infinite loops. |
retry | object | none | Workflow-level retry on uncaught errors: { max_attempts, backoff, initial_ms }. Step-level retries take precedence. |
cost_class | string | "metered" | Same as TOOL.md. Baseline; per-step cost ladders up. |
tags | string[] | [] | Discovery tags. |
metadata | object | {} | Free-form, namespaced for adapter hints. |
inputsFiles | object | {} | File-contract IN per AIP-16. Map of <key> → { path, mode?, contentType? }. Host stages each from the workspace at path into the per-run fs root at filename <key> BEFORE the workflow starts. |
outputsFiles | object | {} | File-contract OUT per AIP-16. Map of <key> → { path, mode?, contentType? }. Host syncs each from the per-run fs root at filename <key> to the workspace at path AFTER the workflow completes. path supports <runId> / <workflowId> / <isoDate> interpolation. |
Removed (formerly part of AIP-15, moved to per-step TOOL.md + their AIP-30 DRIVER bindings)
| Field | New home |
|---|---|
code | The TOOL.md a step references is implemented by a DRIVER which carries driver.code (per AIP-26). Workflows have no own bundle. |
run | Same — driver.run. |
runner | Same — driver.runner (per AIP-17). Per-step driver resolution; workflows don't pin runner. |
secrets | Same — driver.auth.ref pointing at SECRETS.md (per AIP-19). |
network | Same — driver.network.egress. The egress allowlist is per-driver, not per-workflow. |
Inline-entry tool steps (tool: { entry: "./workflow.ts#stepFn" }) | Forbidden. Every kind: "tool" step MUST reference an external TOOL.md by id (per AIP-14). Inline bodies break the contract/driver split — the body lives in the driver. Authors who need a one-off body author a sibling TOOL.md + DRIVER.md. |
A WORKFLOW.md carrying any of these fields is invalid. Workflows are pure orchestration — step DAGs, branching, parallelism, approvals, suspend/resume, compensation. Bodies, transport, install, auth, sandbox all live below in the per-step TOOL.md → DRIVER chain.
Inputs / outputs / file contract
The four IO blocks inputs, outputs, inputsFiles, and
outputsFiles follow the contract defined in
AIP-16 — the lifecycle (host stages declared
workspace files into a per-run scratch root, syncs declared outputs
back, injects the reserved _workflowFsRoot input field), the path
interpolation tokens, the concurrency rules, and the error
semantics are all normative there. AIP-15 imports them unchanged.
Workflow-specific bindings:
- The run identifier used to key the scratch root is the workflow
run's
runId. - The path interpolation token
<workflowId>is the workflow manifest'sid. - Step bodies access the scratch root via
inputData._workflowFsRoot(the standard reserved key).
outputs declares the output shape; result declares the output
value. They split the same way a step's outputs (schema) and a
step's inputs (value mapping) do — so the workflow output is authored
declaratively rather than defaulting to whatever the last step happened
to return:
outputs:
type: object
properties:
count: { type: integer }
delivery: { type: object }
result:
count: $steps.shortlist.count
delivery: $steps.sentresult uses the step-inputs reference grammar ($input.*,
$steps.<id>.*, literals; objects resolve recursively). It is OPTIONAL:
omit it and the output is the final step's result, the prior behaviour.
Runner / Secrets / Network
These three blocks compose the workflow's execution context:
runner(AIP-17) — process boundary: engine (subprocess(default) |sandbox|in-process), optional container image, declarativeneeds(language/native/npm/pip), resource limits.secrets(AIP-19) — env-var bindings: vault slugs, OAuth drivers, plain values.network(top-level) — egress allowlist for outbound HTTP.
The downgrade rule (untrusted-source in-process requests forced
to subprocess) and host responsibilities are normative in
AIP-17. AIP-15 imports them unchanged.
Workflow-specific notes:
- Hosts MAY refuse step kinds that require host-side persistence
(
kind: "suspend",kind: "approval", deepkind: "subworkflow"chains) whenrunner.engine: "sandbox"is in use. When refused, the host MUST reject the workflow at registration with a clear error. - A workflow body authored under
.workflows/SHOULD be treated as untrusted by the host and silently downgraded tosubprocessif it requestsin-process.
Steps
Each entry in steps[]:
Common fields (every step kind)
| Field | Type | Description |
|---|---|---|
id | string | Kebab-case, unique within the workflow. |
name | string | Display label. |
kind | enum | "tool", "branch", "parallel", "suspend", "approval", "map", "loop", "subworkflow", "agent", "gate". |
description | string | One-line purpose. |
inputs | object | Mapping from upstream values to this step's input. Keys are the step's input fields; values are JSON-pointer-style references like $input.productUrl, $steps.fetch-page.html, or literals ({ "kind": "literal", "value": 42 }). |
outputs | JSON Schema | Schema of this step's output. Other steps reference it via $steps.<id>.*. |
next | step id | "$end" | Default successor. Branch steps override this with their branches. |
Step-kind specific fields
kind: "tool"
Run a single tool. Step references either:
tool:field — a specific TOOL.md by id (per AIP-14). Implementor is fixed; driver resolution per AIP-30 picks the concrete backend.action:field (preferred) — an AIP-39 ACTION ref. The resolver picks any TOOL withimplements: <action-ref>, considering policy + capability + cost. Decouples the workflow from a specific tool implementation.
Inline-entry bodies are forbidden — the body lives on the AIP-30 DRIVER that implements the TOOL contract, and driver resolution happens at step-execution time per AIP-30's resolver.
Exactly one of tool: or action: MUST be set per step.
A tool step MAY set cacheable: true to cache its output under the run's
cacheKey. Only the step's resolved inputs are hashed — never context
or secrets. Default false.
# Form A — pin a specific TOOL (locks implementation)
- id: fetch-page
kind: tool
tool: pricing-snapshot # TOOL.md id
inputs:
productUrl: $input.productUrl
outputs: { type: object, properties: { html: { type: string } } }
next: parse
retry: { max_attempts: 3, backoff: exponential, initial_ms: 1000 }
timeout_ms: 20000
# Form B — reference an ACTION (resolver picks tool by policy)
- id: commit-draft
kind: tool
action: "@agentik/actions/standard/storage-commit" # ACTION ref
inputs:
message: "${{ ctx.summary }}"
next: notify
# Inheritance: this step inherits the action's mutates,
# risk_level, approval. Step MAY narrow but never widens.
# Per-step pin (escape hatch — overrides resolver, usually omit):
# pinned_provider: apollo-pricing-httpWhy prefer action: : when a new TOOL ships implementing the same
action (e.g. gh.commit joining git.commit for storage:commit),
the workflow benefits automatically. Pinning tool: locks you to one
implementation forever.
When the workflow runtime executes this step, the runtime:
- Loads the TOOL.md by id from the registry.
- Runs the AIP-30 resolver against the call's context to pick a
DRIVER (or honours
pinned_providerwhen present). - Validates
inputsagainst the contract'sinputSchema. - Dispatches to the resolved driver's
execute[<toolId>]. - Validates the result against the contract's
outputSchema. - Maps the result into the workflow's running state per the step's
outputsdeclaration.
The workflow itself is driver-agnostic. The same workflow runs against OpenAI HTTP, Replicate HTTP, or self-hosted SDK depending on workspace policy and resolver decisions per call.
kind: "branch"
Conditionally route to one of N successors.
- id: route
kind: branch
branches:
- when: $steps.parse.tier == "free"
next: handle-free
- when: $steps.parse.tier == "paid"
next: handle-paid
default: $end # taken if no branch matcheswhen is a small expression language (see Expressions).
Adapters MUST support equality, ordering on numbers and strings,
boolean && / || / !, and null checks. Adapters MAY support
more; spec-strict workflows stay within the minimum.
| Field | Type | Description |
|---|---|---|
join | step id | Later sibling where execution resumes after the chosen arm. It MUST come after every arm target. Omitted ⇒ the step right after the last arm's target, so the last arm's body is exactly its target step; declare join to give the last arm several steps. |
fallthrough | boolean | Legacy, discouraged. Default false. true restores the pre-exclusive semantics: the chosen target runs AND every sibling after it in document order, and a no-match without default continues at the next sibling. Incompatible with join. |
Branches are exclusive: only the chosen arm runs. Steps in an untaken
arm never start; the host records each as skipped — a step.skipped
event and StepRecord.status: "skipped", see AIP-58 §5.
kind: "parallel"
Run named child step graphs concurrently; rejoin when all complete.
- id: enrich
kind: parallel
branches:
- id: stripe
next: gather # converges
steps: [...] # nested step list
- id: hubspot
next: gather
steps: [...]
next: gatherEach branch is itself a step list. The parent step's outputs is the
merged shape of its branches. Adapters MAY emit a single concurrent
flow or fan-out across worker queues.
kind: "suspend"
Pause the run until an external event resumes it.
- id: wait-for-payment
kind: suspend
resume:
on: ["stripe.charge.succeeded", "manual.cancel"]
timeout_ms: 86400000 # 24h
on_timeout: cancel # cancel | continue | <step-id>
outputs:
type: object
properties:
eventName: { type: string }
eventPayload: { type: object }
next: charge-succeededStorage: the runtime persists run state at the suspend point and re-dispatches on a matching event. The manifest does not name the storage; that's adapter-specific.
kind: "approval"
Pause for a human decision. A specialised suspend.
- id: legal-review
kind: approval
prompt: "Review the contract draft and approve or reject."
artifacts:
- $steps.draft-contract.fileId
approvers:
- role: legal
- role: founder # any one approver suffices
timeout_ms: 86400000
on_timeout: escalate
on_reject:
next: revise
on_approve:
next: sendGovernance: every approval step writes to the audit log per AIP-7 — actor, decision, timestamp, justification — without the workflow needing to call the audit machinery.
kind: "map"
For-each over a collection. Each item runs the same nested step graph; results aggregate into an array.
- id: process-each-line
kind: map
over: $steps.parse.lines
parallelism: 5 # concurrent items; 0 = unbounded
steps: [...]
outputs:
type: array
items: { type: object, properties: { ... } }kind: "loop"
Repeat the nested step graph while a condition holds. Bounded by
max_iterations.
- id: refine
kind: loop
while: $steps.evaluate.score < 0.8
max_iterations: 5
steps: [...]kind: "subworkflow"
Invoke another WORKFLOW.md. Inputs are mapped, outputs flow into
this step's outputs.
- id: ship
kind: subworkflow
workflow: deploy-to-prod # WORKFLOW.md id
inputs:
branch: main
artifact: $steps.build.artifactPathA subworkflow step MAY instead declare with: — a declarative input
projection for the child workflow. with and inputs are
equivalent: with is an alias of inputs with identical
semantics, using the same reference grammar as every other inputs
mapping (literals, $input.*, $steps.<id>.*; a leading $$
escapes a literal $), resolved against the parent's bindings
when the step runs.
- A
subworkflowstep withoutwith:(and withoutinputs) passes the parent's input object to the child verbatim — unchanged behaviour. - Declaring both
withandinputson the same step MUST be a load-time error naming the step. - A
with:reference to a step id not declared anywhere in the workflow (nested bodies included) MUST be rejected at load time, naming the step and the offending key. - At run time, a reference whose path does not exist in the parent's
bindings MUST be an error naming the subworkflow step and key —
never a silent
undefinedhanded to the child. - Hosts MUST walk nested step lists (map / loop / parallel bodies, branch arms) when applying these rules.
- id: produce-book
kind: subworkflow
workflow: write-book # child WORKFLOW.md id
with:
topic: $input.audience
outline: $steps.plan.outlinekind: "agent"
Spawn or reuse an agent session and send it a prompt, waiting for the
turn to complete. This subsection codifies @agentproto/workflow's
shipped StepAgent interface — packages/workflow/src/types.ts —
and its runtime projection in @agentproto/workflow-runtime
(packages/workflow-runtime/src/types.ts).
| Field | Type | Description |
|---|---|---|
agent.ref | string | App-scoped agent id (e.g. "@my-app/reviewer") the host resolves to a concrete adapter + spawn options at compile time. An unresolvable ref fails compilation, never the run. Omit to spawn a plain adapter, or to reuse via sessionRef. |
prompt | string | Required. The prompt sent to the session. |
adapter | string | Adapter slug for spawning a NEW session. Ignored (and unnecessary) when agent.ref resolves one. Omit both adapter and agent to reuse via sessionRef. |
sessionRef | string | Reuse an earlier agent step's spawned session, by that step's id. |
sandbox | string | object | Run the step's session inside a sandbox — a provider slug or an inline spec ({ provider: ... }). Only meaningful with adapter; a sessionRef reuse inherits the referenced session's placement. Hosts without sandbox support MUST fail the spawn loudly rather than silently running on the host. |
cacheable | boolean | Cache this step's output under the run's cacheKey. Default false. |
maxRetries | integer | Re-prompt-and-retry attempts on outputSchema mismatch. Default 2. |
outputSchema | JSON Schema | Validates the session's final message; re-prompts on mismatch. |
policy | object | awaiting: "auto-allow" (with prompt) | "escalate" (with optional webhookUrl, timeoutMs) | "fail" — what to do when the session asks for something outside its permission. |
options | object | Adapter option id → value, forwarded to a NEW spawn's startSession({ options }). Only meaningful with adapter; ignored on a sessionRef reuse. |
harness | object | Harness pinning for this step's spawn — optional model, effort, role, tools (per-spawn allowlist), skills, cwd, promptFile (+ loader-computed promptSha), and knowledge[] (AIP-10 corpus materialized into the step's cwd under .knowledge/ before the session runs). Precedence when a field is also set elsewhere: step harness > the resolved AGENT.md's frontmatter > caller args > adapter default. |
- id: review-draft
kind: agent
agent:
ref: "@my-app/reviewer"
prompt: "Review the draft at $steps.draft.path and reply with findings."
outputSchema: { type: object, properties: { findings: { type: array, items: { type: string } } } }Outcome (added 2026-09-25, per AIP-58). Whether an agent step's
turn ending counts as succeeded, suspended, or failed is defined
by AIP-58 §3, not by this AIP: a turn ending is never
success by itself — the step's outputSchema MUST validate (or its
declared required artifacts MUST exist, per AIP-16/AIP-58) before it
counts as succeeded. A turn that ends by legibly asking for missing
input resolves suspended { reason: "input-required" }; any other
turn-end that doesn't satisfy the contract resolves
failed { code: "missing-output" }. This AIP's own maxRetries /
outputSchema re-prompt behaviour is unchanged — AIP-58 governs what
happens once retries are exhausted, not the retry loop itself.
kind: "gate"
Run a shell command through the host's subprocess runner (AIP-17) as
a deterministic pass/fail check. Exit code 0 is a pass; any other
exit code is a failure, subject to the step's common retry policy.
This subsection codifies @agentproto/workflow's shipped StepGate
contract — packages/workflow/src/types.ts:370-412.
| Field | Type | Description |
|---|---|---|
command | string | Required. The command to execute. No shell interpolation — the command is invoked as an argv vector (command + args), never through a shell. |
args | string[] | Argv vector (no shell). Each string may carry a LEADING $… run-time reference: a leading $input / $item / $steps.<id> / $index token is resolved per run against the run bindings and any trailing text is appended verbatim ("$input.bookDir/knowledge" → "<resolved bookDir>/knowledge"); $$ escapes a literal $; an unresolvable ref throws naming the step and the arg index. |
cwd | string | Working directory; defaults to the workflow run's own cwd. May carry the same leading-$… ref rule as args. The resolved cwd is made absolute — a relative resolved cwd resolves against the run's cwd, never the daemon process cwd. |
report | string | Path (relative to cwd) of a JSON report file, consulted when stdout doesn't itself parse as JSON. |
on_fail | object | Re-prompt-and-rerun on a failing exit code, bounded by retry.max_attempts. Runs BEFORE each retry attempt after the first: sends the named prior agent step's session (reused via that step's sessionRef) a prompt with this gate's last report injected (reprompt names the step id; with merges extra literal context), waits for its turn, then re-runs the gate command. |
Reporting. The process' stdout, if it parses as JSON, becomes the
gate's report; otherwise the file at report is read and parsed as
JSON. The report — plus ok and exitCode — is bound at
$steps.<id>.report / .ok / .exitCode for later steps to read,
and each attempt's outcome (pass/fail, exit code, attempt index,
duration, and parsed report) is reported as a gate-report event on
the run's event stream as it happens, so per-attempt outcomes are
observable mid-run rather than only at step completion.
- id: check-wordcount
kind: gate
command: node
args: ["scripts/check-wordcount.js", "$input.bookDir/manuscript"]
report: "./report.json"Runtime-internal node types are not authorable
A runtime MAY compile the ten authorable step kinds into internal
node types of its own. Such internal types are NOT manifest-authorable
and are NOT part of the kind enum. This codifies shipped behaviour:
@agentproto/workflow-runtime's node algebra (transform, pipeline,
group, alongside its own tool/branch/map/loop/parallel/
approval/suspend/subworkflow/agent/gate nodes) is a
compilation target for the manifest kinds — evidence:
packages/workflow-runtime/src/types.ts. Manifests listing
kind: "transform" (or pipeline / group) are invalid.
Accepted compile subset (normative)
This AIP describes a step DAG with arbitrary next successors. The
shipped reference compiler (@agentproto/workflow-runtime's
compileWorkflow — packages/workflow-runtime/src/compile-workflow.ts,
module doc) accepts only the linear / structured subset of that
grammar, and conformant implementations SHOULD support at least the
same subset:
- Steps run in document order.
map/loop/parallelnest their child step lists.- Non-linear
nextgotos are rejected with a clear diagnostic. kind: "branch"compiles in its forward-only form: everybranches[].nextanddefaultmust name a later sibling in the same step list — never backward, never into a nestedmap/loop/parallelbody.
Compiling a full goto graph (backward jumps, cross-scope targets) is
a separable follow-up. Retry-style control flow is hand-authored as a
loop step rather than a backward branch. The point of pinning the
subset here is that a manifest that is legal-per-spec but outside the
subset failing to compile must not come as a surprise.
Expressions
The minimum expression grammar adapters MUST support:
expr := value | binary | unary | path
value := string | number | bool | null
path := '$input' ( '.' field )*
| '$steps.' step-id ( '.' field )*
| '$item' ( '.' field )*
| '$index'
binary := expr op expr
op := '==' | '!=' | '<' | '<=' | '>' | '>=' | '&&' | '||'
unary := '!' expr$input is the workflow inputs block; $steps.<id> is a step's
resolved output (no .outputs. segment — the step's output binds
directly under its id); $item / $index are the current element and
position inside a map step. A leading $$ escapes a literal $.
No function calls, no loops in expressions, no string interpolation. Workflows that need richer logic encode it as a tool step with a dedicated body.
Triggers
A workflow MAY declare events that start it.
triggers:
- kind: schedule
cron: "0 9 * * MON"
timezone: Europe/Paris
- kind: webhook
path: /pricing-update
- kind: event
name: stripe.invoice.finalized
- kind: manual
label: "Run pricing snapshot now"Adapters MUST honour kind: "manual" (a UI trigger). Other kinds are
optional and capability-gated; adapters that don't support them MUST
refuse the workflow at registration time with a clear error.
Routines (preferred form for kind: schedule triggers)
triggers: [{ kind: schedule, cron: ... }] is legacy — kept for
back-compat. The decoupled, preferred form for any recurring or
event-driven invocation is the routines: field, which references
AIP-41 ROUTINE.md manifests:
# Preferred — ref to a published or local ROUTINE.md
routines:
- { ref: "@agentik/routines-standard/daily-9am-utc" }
- { file: "./.routines/quarterly-rotation/ROUTINE.md" }
# Inline form — equivalent to a one-shot ROUTINE.md inside this workflow
routines:
- inline:
schedule: { kind: cron, cron: "0 9 * * MON", timezone: "Europe/Paris" }
target: { workflow: { ref: "./" } }
identity: "bot://acme-routines"Decoupling buys you:
- One schedule, many workflows. "Daily 09:00 Paris" defined once.
- Identity attribution at fire time — POLICY-checked.
- Retry, on_failure, history — defined on the routine, not duplicated per workflow.
- Library reuse —
@agentik/routines-standard/*covers most common cadences.
Hosts MAY auto-migrate inline triggers: [{ kind: schedule }] into
anonymous inline routines internally for uniform handling. Authors
SHOULD migrate to routines: for any cadence shared across ≥2 workflows.
Body
Markdown body following the frontmatter. Recommended sections:
## Overview— narrative, when to use vs not.## Diagram— mermaid graph TD / sequenceDiagram / stateDiagram.## Step responsibilities— table of step → owner / SLO.## Errors & recovery— what causes eachnext: error-…route.## Examples— sample run with input + step trace + output.
The body is informational. Workflows MUST function with adapters that read only the frontmatter.
The defineWorkflow and defineStep standard signatures
Every implementation that consumes WORKFLOW.md MUST expose two
functions whose signatures match the contracts below:
defineStep declares one step, defineWorkflow assembles
steps into a runnable workflow.
defineStep (TypeScript notation, normative)
defineStep(definition: StepDefinition): StepHandle
interface StepDefinition {
// Identity — mirrors the manifest fields with the same names.
id: string
name?: string
description?: string
kind: "tool" | "branch" | "parallel" | "suspend"
| "approval" | "map" | "loop" | "subworkflow"
| "agent" | "gate"
inputs?: InputMapping // see manifest spec
outputs?: JSONSchema | unknown
// Common per-step controls — same vocabulary as TOOL.md.
approval?: ApprovalClass
riskLevel?: 0 | 1 | 2 | 3
timeoutMs?: number
retry?: RetryPolicy
compensation?: string // step-id reference
// Kind-specific fields — see manifest spec for the full list.
// (`tool`, `branch`, `parallel`, `suspend`, `approval`, `map`,
// `loop`, `subworkflow`, `agent`, `gate` each carry their own extras.)
[kindSpecific: string]: unknown
}defineWorkflow (TypeScript notation, normative)
defineWorkflow(definition: WorkflowDefinition): WorkflowHandle
interface WorkflowDefinition {
// Identity — mirrors the manifest fields with the same names.
id: string
name?: string
description: string
version?: string
// Schemas — JSON Schema or zod/pydantic-compatible value.
inputSchema: JSONSchema | unknown
outputSchema: JSONSchema | unknown
// The step graph. Either provided up front, or assembled via the
// builder methods below (see `WorkflowHandle`).
steps?: StepHandle[]
start?: string // step-id of the first step
// Workflow-level policies — same vocabulary as the manifest.
approval?: ApprovalClass
riskLevel?: 0 | 1 | 2 | 3
timeoutMs?: number
maxSteps?: number
retry?: RetryPolicy
costClass?: "trivial" | "metered" | "expensive"
triggers?: Trigger[]
tags?: string[]
metadata?: Record<string, unknown>
}
/**
* The handle returned by defineWorkflow. Builder methods append
* steps in order; `commit()` finalises and validates the graph.
*
* Implementations MAY return a fluent builder OR accept the full
* step list up front via `steps:`. Both patterns produce the same
* registered workflow — choose the one that fits the host idiom.
*/
interface WorkflowHandle {
step(step: StepHandle): WorkflowHandle
branch(step: StepHandle): WorkflowHandle // kind: "branch"
parallel(step: StepHandle): WorkflowHandle // kind: "parallel"
approval(step: StepHandle): WorkflowHandle // kind: "approval"
suspend(step: StepHandle): WorkflowHandle // kind: "suspend"
commit(): WorkflowHandle // finalise + validate
}Conformance rules
-
Canonical names. Exports MUST be named
defineWorkflowanddefineStep. Implementations MAY also re-export under host-specific aliases but the canonical names are whatWORKFLOW.mdadapters and SKILL.md authoring guides reference. -
Schemas validated at boundaries. The host MUST validate the workflow's
inputSchemabefore the first step runs and each step'sinputsmapping before its body executes. Steps MUST NOT re-validate; they receive parsed input. -
Step kinds are exhaustive.
defineStepaccepts only the ten kinds enumerated. Unknown kinds MUST be rejected at registration with a clear error — runtimes MAY add proprietary extensions, but they live OUTSIDE the standard signature. This escape hatch does NOT coverkind: "agent"andkind: "gate": those two sit INSIDE the core union (@agentproto/workflow's shippedStepCommon.kind—packages/workflow/src/types.ts), so the eight-kind enumeration this rule previously stated was a real conformance breach of the standard signature, not a proprietary extension. This amendment codifies the shipped ten-kind union. -
approval,riskLevel,timeoutMs,retry,compensationare enforced by the host, not the step body. Steps assume permission was granted and budget is in place. -
Compensation is idempotent. When the host walks back through completed steps to undo on failure, each compensation step MUST tolerate being called repeatedly without observable extra effect. Implementations MUST NOT re-validate compensation completion as a single-shot constraint.
-
No I/O at module load. Same rule as
defineTool. The module containingdefineWorkflow(...)MUST be safely importable without side effects. All I/O happens inside step bodies. -
Suspend points persist. When a step has
kind: "approval"ORkind: "suspend", the implementation MUST persist enough run state to resume from that step after a process restart. This requirement is normative for both step kinds (amended 2026-09-25 — see below). In-memory-only implementations MUST refuse to register suspendable workflows or document the limitation explicitly.This codifies the shipped daemon runner's behaviour: on reload, a run parked awaiting either a human approval or an external suspend event is exempt from being failed — both its
awaitingApprovaland itsawaitingSuspendrecord are durable and re-registered so a subsequent decision or resume still lands, while any other in-flight run (running, orawaiting-inputwith neither record) is markedfailedwith "interrupted by daemon restart" (seepackages/runtime/src/workflow-runner.ts's module doc comment and itsloadRuns/ reload re-registration logic). A decision or resume delivered after a restart still lands for either kind — the event fires, any ledger write happens — but the run's own execution cannot resume mid-workflow, so it settlesfailedwith a message naming that cause, identically for both kinds.2026-09-25 — the
kind: "suspend"carve-out is closed. This rule previously required durable persistence forkind: "approval"only, carryingkind: "suspend"as an explicit, tracked exception (a shipped defect, not a design choice). That exception is removed: the reference runner already persists and re-registersawaitingSuspendthe same way it doesawaitingApproval— the two step kinds get identical restart treatment today, so the earlier carve-out no longer describes shipped behaviour and this rule stops pretending otherwise.
Implementer's guide
For step-by-step guidance on building a defineWorkflow /
defineStep conformant implementation in a specific language or
framework, see ./resources/aip-15/draft/ADAPTER.md.
The AIP only defines the contract; the resource doc walks an
implementer through the projection.
Compensation (sagas)
A workflow that mutates external state SHOULD declare compensation steps so partial failures roll back cleanly. Convention:
- Each step that mutates declares
compensation: <step-id>. - The compensation step takes the original step's outputs as inputs.
- If a downstream step throws, the runtime walks back through already-completed steps in reverse, executing each compensation.
- id: charge-card
kind: tool
tool: stripe-charge
inputs: { amount: $input.amount, customer: $input.customerId }
outputs: { type: object, properties: { chargeId: { type: string } } }
compensation: refund-card
- id: refund-card
kind: tool
tool: stripe-refund
inputs: { chargeId: $steps.charge-card.chargeId }
outputs: { type: object }Compensation MUST itself be idempotent — adapters MAY retry.
Stable identity
id + major version form the workflow's stable identity. Run
records, audit rows, and capability grants key on id@major.
Authoring with SKILL.md
The canonical way to generate a WORKFLOW.md is via a paired
SKILL.md — distributed at
./resources/aip-15/draft/skills/author-workflow/SKILL.md — that an agent
loads when asked to build a workflow. The skill walks the agent
through:
- Decompose the goal into named steps (the planner phase).
- Pick
kindfor each step (tool / branch / parallel / approval / …). - Declare
inputsmappings using the path grammar. - Decide which steps need approval and which run
auto. - Add compensation for every mutating step.
- Validate the manifest against
./resources/aip-15/draft/WORKFLOW.schema.jsonand emit the final pair.
Compatibility
With AIP-26 / AIP-17 / AIP-19 (2026-04-30 revision)
The 2026-04-30 revision adds top-level code:, run:, runtime:,
secrets:, network: blocks defined in their respective sibling
AIPs, and removes the legacy entry: and runtime: shapes from
the manifest. The migration mirrors AIP-14 § Compatibility:
manifest.entry → manifest.run (single-file → string-path form)
manifest.runtime.mode → manifest.runtime.engine
manifest.runtime.env → manifest.secrets (per AIP-19)
manifest.runtime.fs → manifest.inputsFiles / outputsFiles (per AIP-16)
manifest.runtime.network → manifest.network (top-level)
manifest.runtime.cpu_ms_max → manifest.runtime.limits.cpu_ms
manifest.runtime.memory_mb_max → manifest.runtime.limits.memory_mb
manifest.runtime.timeout_ms → manifest.runtime.limits.timeout_msNote: this table previously named the block runner; the canonical
WORKFLOW.schema.json (the "runtime" property) and the shipped
packages/workflow/src/schema.ts both name it runtime — the prose
now aligns with the schema and the implementation.
Hosts SHOULD accept both shapes during the deprecation window.
Output result mapping (2026-06-12 revision)
Adds an optional top-level result: expression that maps step outputs
into the workflow's output value, using the same reference grammar
as a step's inputs ($input.*, $steps.<id>.*, literals;
objects resolve recursively). outputs stays the output shape;
result is the value that fills it. Fully backward-compatible — omit
result and the output remains the final step's result, the prior
behaviour. See EXAMPLES § 12.
Subworkflow with: input projection (2026-09-04 revision)
Adds an optional with: block to kind: "subworkflow" steps — a
declarative input projection for the child workflow, equivalent to
inputs (same reference grammar, resolved against the parent's
bindings). Steps that use neither with: nor inputs are unchanged
and pass the parent's input verbatim. Fully backward-compatible — see
[§ kind: "subworkflow"](#kind-subworkflow).
Suspend persistence extended to kind: "suspend" (2026-09-25 revision)
Conformance rule 7's carve-out — durable restart persistence required
for kind: "approval" only, kind: "suspend" explicitly not-yet-covered
— is removed. Both step kinds now carry the same normative requirement,
matching the reference runner's actual behaviour (it already persists
and re-registers awaitingSuspend the same way as awaitingApproval;
see § rule 7). No field or wire change; this is a
strengthening of an existing MUST, not a new capability. An
implementation that previously relied on the carve-out to leave
kind: "suspend" in-memory-only is non-conforming as of this revision.
Agent step outcome deferred to AIP-58 (2026-09-25 revision)
Adds AIP-58 to requires and defers a kind: "agent"
step's success/suspend/fail determination to AIP-58 §3 rather than
leaving "the turn ended" as an implicit success signal — see
§ kind: "agent"'s Outcome note. No field changes; this
is a clarification of existing behaviour's intended outcome semantics,
not a new step field. Hosts whose current implementation treats a
turn ending as success by default are non-conforming as of this
revision and should adopt AIP-58's outcome rule.
With pre-AIP runtime-specific workflow definitions
For authors arriving from runtime-idiomatic workflow factories
(Mastra createWorkflow, Temporal, etc.):
- Add
WORKFLOW.mdnext to the existing entry, copying step ids and schemas. - Re-express the step bodies behind the standard
defineStepsignature; assemble them viadefineWorkflow().step(...).commit(). - Run the manifest through
WORKFLOW.schema.jsonto verify the shape and the path-expression grammar ininputsmappings.
Hosts MAY accept both legacy and WORKFLOW.md-shaped registrations
during a migration period; the audit-log + work-item shapes from
AIP-7 and AIP-13 are identical either
way as long as mutates / requires / approval are populated on
each step.
Security considerations
WORKFLOW.md is declarative: a malicious manifest can lie about
mutates or grant itself unsafe approval: "auto" on a destructive
step. Hosts MUST treat the manifest as untrusted until verified, and
AIP-7 capability gating MUST run regardless of
manifest claims.
The expression grammar is intentionally minimal so manifests can be statically analysed without sandboxing — a manifest cannot execute code on the host's behalf.
triggers: kind: "webhook" exposes a network surface and MUST be
gated by an explicit host capability. kind: "schedule" exposes
recurring cost and MUST be governed by quota policy.
For runtime isolation considerations (env stripping, sandbox vs in-process privilege, egress enforcement), see AIP-17 § Security considerations.
Open questions
- Variable scoping: should
$steps.<id>be visible inside parallel branches, or only after the join? Current draft: only after the join. - Streaming step outputs: tools that yield progressive results
inside a workflow — surfaced as
outputs.streamon the step? - Distributed runtimes: do we need a
placementhint (local/worker:<queue>/serverless) or punt to runtime? - Cross-workflow signals: a workflow in run A waits for an event raised by run B — same workflow id, different runs. Adapters handle this differently; spec is silent for now.
- Sub-second timing primitives:
kind: "wait"for explicit delays — useful enough to standardise, or punt?
These remain open until enough adapters ship to settle empirically.
Resources
Supporting artifacts for AIP-15. Links open the file on GitHub — markdown and JSON render natively in GitHub's viewer. Browse the full resource tree →
AIP-14: TOOL.md — agenttool/v1 (abstract agent contract)
A markdown + frontmatter format for declaring a single agent tool's abstract contract — its identity, input/output schemas, side-effect profile, approval class, and resource budget. Pairs with the standard `defineTool` signature any implementation exposes. Implementation-specific concerns (transport, code, runner, auth, sandbox) live on the AIP-30 DRIVER layer.
AIP-16: IO.md — shared input/output schema blocks
A composable schema block defining `inputs`, `outputs`, `inputsFiles`, and `outputsFiles` — the data-shape primitives reused by every manifest format that needs to declare what flows in and out of a runnable unit.