Skip to main content
A workflow is a TypeScript function that composes step, sleep, and waitForEvent calls. The engine persists every step result, so a workflow survives server restarts and deploys. Completed steps never re-run. A 7-day sleep costs nothing while it waits. An external event wakes a paused run exactly where it stopped. Use a workflow when a process spans more time than one function call should hold:
  • onboarding sequences
  • agent pipelines with human approval gates
  • billing dunning
  • report generation with retries

Writing a workflow

Workflows live in a workflows/ directory next to functions/. One file, one default-exported workflow(...):
Each step’s closure runs with a full action ctx: ctx.runQuery, ctx.runMutation, ctx.llm, ctx.email, and ctx.scheduler. It runs under the same idle timeout and cancellation semantics as any action.

Starting and driving

Start a workflow from any mutation or action. ctx.workflows.start returns the instance id immediately, and the engine’s background driver takes over from there:
Deliver an event to a run the same way, for example from the webhook action that received the confirmation:
sendEvent resolves { delivered, buffered }. If the run is waiting for the event, it resumes (delivered: true). If the run is doing something else, for example running a step, the event is buffered (buffered: true) and the run’s next waitForEvent with that name consumes it. Buffered events are consumed oldest first, and survive restarts. A run holds at most 100 buffered events. An event sent without data arrives as {}. Sending to a finished run rejects.

Waiting with a timeout

Pass timeout to stop waiting after a duration. The call resolves with null when the timeout passes first:
Timeouts are durable, like sleep: they survive restarts and fire within about a second of the deadline. Durations use s, m, h, or d (”60s”, “5m”, “24h”, “7d”). An invalid duration fails the run.

Keys, listing, and cancel

Give a run a key at start, for example the id of the record it works on. At most one active run exists per workflow name and key. Starting again with the same key returns the existing run with created: false, so a retried webhook does not start a second run:
Cancel runs from any mutation or action. No further step runs. A step already in progress can finish, but its result is discarded:
cancel resolves { cancelled: false } when the run had already finished. Cancel and sendEvent never wait for a step in progress, including on Postgres replicas. ctx.workflows.* calls take effect when they are made. In a mutation that later throws, a started run keeps running and a sent event stays delivered. Call them from an action, or as the last step of a mutation. ctx.workflows.list({ name, key, status, limit }) returns runs newest first, without step history. status is one of pending, running, sleeping, waiting, completed, failed, cancelled, or active (any unfinished run). limit defaults to 100, max 1000. ctx.workflows.get(id) returns one run, or null. Finished runs are kept for 24 hours. Steps execute through the background job queue, and sleeps wake on schedule. SQLite stores state in <app-db>.workflows.db. Postgres stores state in shared application tables. The admin API drives the same engine for operators: POST /api/workflows/start, POST /api/workflows/<id>/event (both with Authorization: Bearer $PYLON_ADMIN_TOKEN). Inspect runs at GET /api/workflows (query parameters status, name, key, limit) or GET /api/workflows/<id> (step-by-step results, timings, retry counts). Cancel a run with POST /api/workflows/<id>/cancel and an optional { "reason": "..." } body. POST /api/workflows/start accepts an optional key.

The determinism contract

On every advance, the whole workflow function re-runs from the top. Completed steps replay from their recorded outputs. This gives you plain TypeScript control flow (branches, loops, early returns), with one rule: The sequence of step, sleep, or waitForEvent calls must be identical on every replay, for the same input and step outputs.
  • Branch on wf.input and on step outputs freely. Both are stable.
  • Never branch on wall-clock time, randomness, or external state read outside a step. Put those reads inside a step, then branch on its recorded output.
  • Step names must be unique within a run. The replay cache is name-keyed, so a mismatch fails the run loudly rather than reusing the wrong output. Names starting with event: or timeout: are reserved.
  • waitForEvent calls replay in order: the second wait on a name replays the second recorded event for that name.

Retries and failure

A throwing step fails the current advance. The engine retries the same step (default 3 attempts, configurable per workflow):
Once retries are exhausted, the run lands in failed with the error and the step that caused it, inspectable via the API. Code between steps should be side-effect free. Workflow step execution is at least once. A step can finish an external side effect and stop before Pylon records its result. Pylon can then run that step again. Make step bodies idempotent when they call an external system.

Multiple replicas

Postgres replicas share workflow state. A replica takes a short lease before it advances a run. It renews the lease while the handler runs. A lease token prevents a stale worker from replacing newer state. Another replica resumes the run after an expired lease. SQLite workflow state is local to one machine. Use Postgres for horizontal scaling.