step, sleep, and
waitForEvent calls. The engine persists every step result, so a
workflow survives server restarts and deploys. Completed steps never
re-run. A 7-day sleep costs nothing while it waits. An external event
wakes a paused run exactly where it stopped.
Use a workflow when a process spans more time than one function call
should hold:
- onboarding sequences
- agent pipelines with human approval gates
- billing dunning
- report generation with retries
Writing a workflow
Workflows live in aworkflows/ directory next to functions/. One file,
one default-exported workflow(...):
ctx: ctx.runQuery,
ctx.runMutation, ctx.llm, ctx.email, and ctx.scheduler. It runs
under the same idle timeout and cancellation semantics as any action.
Starting and driving
Start a workflow from any mutation or action.ctx.workflows.start
returns the instance id immediately, and the engine’s background driver
takes over from there:
sendEvent resolves { delivered, buffered }. If the run is waiting for
the event, it resumes (delivered: true). If the run is doing something
else, for example running a step, the event is buffered
(buffered: true) and the run’s next waitForEvent with that name
consumes it. Buffered events are consumed oldest first, and survive
restarts. A run holds at most 100 buffered events. An event sent without
data arrives as {}. Sending to a finished run rejects.
Waiting with a timeout
Passtimeout to stop waiting after a duration. The call resolves with
null when the timeout passes first:
sleep: they survive restarts and fire
within about a second of the deadline. Durations use s, m, h, or
d (”60s”, “5m”, “24h”, “7d”). An invalid duration fails the run.
Keys, listing, and cancel
Give a run akey at start, for example the id of the record it works
on. At most one active run exists per workflow name and key. Starting
again with the same key returns the existing run with created: false,
so a retried webhook does not start a second run:
cancel resolves { cancelled: false } when the run had already
finished. Cancel and sendEvent never wait for a step in progress,
including on Postgres replicas.
ctx.workflows.* calls take effect when they are made. In a mutation
that later throws, a started run keeps running and a sent event stays
delivered. Call them from an action, or as the last step of a mutation.
ctx.workflows.list({ name, key, status, limit }) returns runs newest
first, without step history. status is one of pending, running,
sleeping, waiting, completed, failed, cancelled, or active
(any unfinished run). limit defaults to 100, max 1000.
ctx.workflows.get(id) returns one run, or null. Finished runs are
kept for 24 hours.
Steps execute through the background job queue, and sleeps wake on
schedule. SQLite stores state in <app-db>.workflows.db. Postgres
stores state in shared application tables. The admin API drives the same engine for operators:
POST /api/workflows/start, POST /api/workflows/<id>/event (both
with Authorization: Bearer $PYLON_ADMIN_TOKEN).
Inspect runs at GET /api/workflows (query parameters status, name,
key, limit) or GET /api/workflows/<id> (step-by-step results,
timings, retry counts). Cancel a run with
POST /api/workflows/<id>/cancel and an optional { "reason": "..." }
body. POST /api/workflows/start accepts an optional key.
The determinism contract
On every advance, the whole workflow function re-runs from the top. Completed steps replay from their recorded outputs. This gives you plain TypeScript control flow (branches, loops, early returns), with one rule: The sequence ofstep, sleep, or waitForEvent calls must be
identical on every replay, for the same input and step outputs.
- Branch on
wf.inputand on step outputs freely. Both are stable. - Never branch on wall-clock time, randomness, or external state read outside a step. Put those reads inside a step, then branch on its recorded output.
- Step names must be unique within a run. The replay cache is
name-keyed, so a mismatch fails the run loudly rather than reusing
the wrong output. Names starting with
event:ortimeout:are reserved. waitForEventcalls replay in order: the second wait on a name replays the second recorded event for that name.
Retries and failure
A throwing step fails the current advance. The engine retries the same step (default 3 attempts, configurable per workflow):failed with the error and
the step that caused it, inspectable via the API. Code between steps
should be side-effect free.
Workflow step execution is at least once. A step can finish an external
side effect and stop before Pylon records its result. Pylon can then run
that step again. Make step bodies idempotent when they call an external
system.