Pre-alpha · weeks 1–10 of 17

Durable execution,
woven to last.

Write ordinary functions in Go, TypeScript or Python. The engine makes them crash-proof, resumable and observable — one binary, no external database, no Kubernetes, no ops team.

Apache-2.0 · Go 1.26+ · zero runtime dependencies

STARTED RESERVE · DONE SETTLE · WAIT
Fig. 1 — the engine at work. The cards are the history: appended, never rewritten. Stop the loom for a second or a year; it resumes on the exact pick it stopped at.
The problem

Every team rebuilds this badly

Payment flows, onboarding sequences, billing cycles, approval chains. They all need to survive crashes — and they're all usually built from Redis queues, cron jobs and a status column.

The DIY version breaks

Queues plus status flags lose state on crashes, can't replay, and breed idempotency bugs. "Did the charge go through before the pod died?" has no answer in your data.

The robust version is enormous

Temporal and Cadence solve it properly — and ask for Cassandra or MySQL, several services, Kubernetes, and someone to run it. Most teams cannot justify that.

The gap this fills: durable execution you can run in under five minutes. One binary, one data directory, a Go SDK. Your workflow crashes, restarts, and picks up exactly where it left off.
One idea, all the way down

The whole engine is one pass of the shuttle

Every workflow is an append-only sequence of events. The engine is a loop over it.

i.

Read the cards

Load the workflow's history — CRC-framed records, fsync'd on every append.

ii.

Replay

Re-run your function against it to find the exact position.

iii.

Schedule the next pick

Dispatch the next activity, arm the timer, or wait on a signal.

iv.

Record the result

Append one durable event. Then do it again, until the cloth is cut.

Append-only history

Each workflow gets a file of length-prefixed, CRC32C-checksummed protobuf records, fsync'd on every append. A step is only "done" once it's durable.

Deterministic replay

To resume, the engine re-runs your function from the top. Completed steps return their recorded results instantly, so side effects never happen twice.

Crash = a no-op

Recovery isn't a special code path — it's the same decision loop, run against the same log. Verified by a test that SIGKILLs the engine mid-workflow.

The pattern book

Proven weaves, ready to run

Four workflows almost every team needs, written against the real SDKs. Each one leans on a different guarantee — pick the thread that matches your problem.

The saga — charge, ship, never half-do it

Reserve stock, charge the card, ship. If the process dies between "charged" and "shipped", a queue-plus-flags build has no answer — the money moved and nothing else did.

On the loom: every step lands on the card chain before it runs. A crash mid-saga replays to the exact pick; exhausted retries hand you an ordinary error, so compensation is just an if-statement.

Saga compensationRetries with backoffDurable history
Goorder_workflow.go
func OrderWorkflow(ctx sdk.Context, input []byte) ([]byte, error) {
	reserved, err := ctx.ExecuteActivity("reserve", input)
	if err != nil {
		return nil, err // failed before any money moved
	}

	// Retries are engine policy (SetRetryPolicy); this error only
	// surfaces once every attempt is exhausted.
	if _, err := ctx.ExecuteActivity("charge", reserved); err != nil {
		ctx.ExecuteActivity("release", reserved) // compensate
		return nil, fmt.Errorf("charge failed, released: %w", err)
	}

	ctx.Sleep("settle", 2*time.Second) // durable across restarts

	return ctx.ExecuteActivity("ship", reserved)
}
On the bench today

What you get today

Everything below is implemented and covered by tests running under -race.

Durable timers

ctx.Sleep("cooldown", 24*time.Hour) — the deadline is on disk, re-armed after a restart.

Signals

Park a workflow until an external event arrives. Signals sent early are buffered, not lost.

Retries with backoff

Per-activity policies with exponential backoff. Retries survive restarts too.

Saga compensation

A failed activity returns an ordinary Go error. Handle it, undo the work, move on.

Leases & heartbeats

A worker that dies mid-task has its work reassigned. Zombie results are fenced out.

Metrics & dashboard

Prometheus /metrics and a built-in web dashboard — both compiled into the binary.

Any language

Write workflows in TypeScript or Python as well as Go. The SDKs are dependency-free — no npm install, no protobuf toolchain.

Where it sits

Not a Temporal replacement — a right-sized option below it

 Queue + DB flagsworkflowdTemporal
Survives a crash mid-flow✗ partial✓✓
Resumes without re-running work✗✓✓
Infrastructure requiredRedis + DBnoneDB cluster + services
Deploy artifactyour app1 binaryseveral services
Time to first workflowdays of glueminutesa sprint
Horizontal HAvariesPhase 4 (Raft)✓
Multi-language SDKsn/aGo, TS, Python (+Java ref)✓
Battle-tested at scale—✗ pre-alpha✓

Choose Temporal when you need multi-language SDKs, huge scale, or a vendor. Choose this when you want durable execution without running a cluster.

Honest status

The pattern, section by section

This is a build-in-public engineering project, currently pre-alpha. Weeks 1–10 of a 17-week plan are done. Don't put payroll on it yet.

✓
Phase 1 — Persistence & the core loop
History store, state machine, gRPC task queue, crash recovery + the Kill Test
✓
Phase 2 — Timers, signals, replay
Durable sleep, signals, deterministic replay of Go code, retries, Saga
✓
Phase 3 — Multi-worker & observability
Leases, heartbeats, Prometheus metrics, visibility API, dashboard
○
Phase 5 — SDK polish & DX
Ergonomic public API, workflow versioning, testing framework, docs
Five minutes, start to finish

Thread the needle

terminal— three ports, one process
git clone https://github.com/olayanju-1234/durable-workflow-engine
cd durable-workflow-engine
make build

# engine + dashboard + metrics, all from one binary
./workflowd --serve

# → gRPC     :7233
# → dashboard :8233
# → metrics   :9090/metrics