Skip to content

Kill it. Resume it.Nothing happens twice.

A durable execution runtime for AI agents, in Rust. Every event is written before the runtime acts on it, so a crashed run comes back and finishes from exactly where it stopped.

agent.tomlagent=d2673e70

Edit it. The next run checkpoints this file and reports the new agent hash. Pin a tool to write and the runtime will never run that call again on its own: if a run stops after the request is recorded but before the result is, it waits for a person to check whether the write happened.

salvor@localfold: salvor-replay (wasm) · model: recordedwrites: 0
salvor runtime=loading log=indexeddb fixture=examples/hero/
the runtime loads on first click, or when this pane scrolls into view.

Hard-refresh this page mid-run. The log is in your browser's IndexedDB. Nothing replays.

The whole agent

agent.toml
model = "claude-opus-4-8"
system_prompt = "Answer in one or two sentences."

That's the whole agent.

Two lines is a durable run. Every model call is written before it is made and again when it lands, so the log is the run. Add [[mcp_servers]] when it needs tools, and pin effect_overrides on any tool that touches something outside the process.

Bring it up

Four ways in. Pick the one your build already uses.

npm install -g @salvor-run/cli

prebuilt binary, no Rust toolchain

These four install the CLI binary. The Python and TypeScript clients for the HTTP API are under Pick a door, below.

Pick a door

Different doors. Same runtime, same log.

Which door is yours depends on where your code already lives.

Rust library

The agent runs inside your own Rust process, with no server and nothing else to keep alive. When that process dies mid-run, recover() reads the log back and carries on from the step it stopped on.

cargo add salvor

use salvor::prelude::*;
use std::sync::Arc;

let store = Arc::new(SqliteStore::open("runs.db")?);
let agent = Agent::builder()
    .model(Config::from_env(), "claude-opus-4-8")
    .system_prompt("Answer in one or two sentences.")
    .build()?;

// start drives the loop; recover() continues a crashed run from the log
match Runtime::new(store).start(&agent, json!({"question": "..."})).await? {
    RunOutcome::Completed { output, .. } => println!("completed: {output}"),
    RunOutcome::Parked { reason, .. } => println!("parked: {reason:?}"),
}

salvor CLI

The same four commands the terminal above runs, against a real SQLite file on disk. Nothing to import and no server to keep running.

salvor run --agent agent.toml --input '"a question"'
salvor list
salvor history 7c1e4a92-3f5b-4d18-9a6c-2e0b8d4f1a37
salvor resume 7c1e4a92-3f5b-4d18-9a6c-2e0b8d4f1a37 --agent agent.toml

This run already completed, so resume refuses it and says so -- "a completed run is not resumable, by design" -- rather than silently restarting a finished job. That refusal is resume's whole point here, not a bug: a run parked at needs_reconciliation refuses too, but asks for resolve first, not resume.

HTTP: salvor runs the loop

Post a run, then read its events as they commit. Salvor calls the model and performs the tools; your code watches and reacts. The stream is plain SSE whose event ids are the log's own sequence numbers, so a dropped connection resumes from the last id it saw with nothing missed and nothing repeated, and curl alone is enough to watch a run. The published Python and TypeScript clients are thin wrappers over that HTTP, and neither holds any recovery logic of its own.

pip install salvor · npm install @salvor-run/client

from salvor import Client

with Client("http://127.0.0.1:8080") as c:
    agent = c.register_agent(open("agent.toml").read())
    run_id = c.start_run(agent, {"question": "..."})
    for event in c.stream_events(run_id):
        print(event.seq, event.kind)

salvor serve reads the model key from ANTHROPIC_API_KEY, or the variable an agent's [llm] api_key_env names instead, from its own environment. A 401 from the Messages API means that variable is unset where serve runs, not in your client.

HTTP: your process runs the loop

Your code decides every step and performs the tool calls itself, in your language, under credentials salvor never holds. Salvor checks each event you append against the log it already has. It makes the model call for you, because the API key stays on the server, and the tool calls stay in your process. The loop that survives being killed is yours. If the provider call throws, or the process dies before the completion is reported, the run parks at needs_reconciliation: salvor will not guess whether the write took effect, so noticing that park and resolving it is your code's job too, and there is no automatic retry for a write.

A tool your process performs is declared by the operator in TOML and has no code behind it on the server. There is no MCP (Model Context Protocol) server here and nothing for salvor to run.

One file per tool, and the flag repeats. A desk that refunds and pays out is two declarations, each with its own effect class and its own answer to trust_completion, and list_client_tools hands your loop the union of them as the definitions the model is given.

This serve needs ANTHROPIC_API_KEY too, in its own environment.

salvor serve --client-tool refund-card.toml --client-tool wire-payout.toml

name = "refund_card"
effect = "write"                  # the operator's word, never the client's
# grants this client the right to close its own call by report alone;
# wire-payout.toml in this example sets it false, so a human closes every payout
trust_completion = true
require_equal = ["amount_cents"]  # reported must equal what the intent said

[input_schema]
type = "object"
required = ["order_id", "amount_cents", "currency"]

# provider_refund_id is the one field a client could not have invented
# without the provider answering, so a completion lacking it is a claim
[output_schema]
type = "object"
required = ["provider_refund_id", "status", "amount_cents"]

# elided here: each field's own [input_schema.properties.*] table
salvor serve --client-tool refund-card.toml

from salvor import Client

# whatever names your loop's definition; salvor records it and hands it
# back on replay, and never resolves it
AGENT = "sha256:client-tools-refund-desk"

with Client("http://127.0.0.1:8080") as c:
    # the operator's declarations are the model's tool definitions
    tools = [{"name": t.name, "input_schema": t.input_schema}
             for t in c.list_client_tools()]

    run = c.open_client_run()
    run.append([run.envelope(0, "RunStarted", agent_def_hash=AGENT, input={})])

    # a model step records its intent AND its completion, so it takes 1 and 2
    step = run.model_step(1, {
        "model": "claude-opus-4-8",
        "max_tokens": 512,
        "tools": tools,
        "messages": [{"role": "user", "content": "ORD-7781 arrived damaged."}],
    })
    call = next(b for b in step.response["content"] if b["type"] == "tool_use")

    opened = run.client_tool_intent(3, call["name"], call["input"])
    # call_provider is YOUR code: the money moves here and salvor never sees
    # it. Perform it under salvor's derived key, so a retry after a crash
    # cannot refund the same order twice.
    charged = call_provider(opened.idempotency_key, call["input"])
    run.client_tool_completion(3, {
        "provider_refund_id": charged["provider_refund_id"],
        "status": charged["status"],
        "amount_cents": charged["amount_cents"],
    })

When salvor may make the call itself, the tool is an MCP server: a subprocess speaking stdio, so it can be written in any language. examples/python-tools and examples/typescript-tools each extend an agent without importing salvor anywhere.

No cure, no pay

Three claims, one gate, and you can run it.

A 109-event control run is recorded, then continued from a fresh store at every one of its 110 prefix boundaries. Every boundary, not a sample.

  • Byte-identical final logEach boundary that completes is asserted equal to the control log, event for event, including every sequence number and recorded timestamp.
  • Zero duplicate writesWrite executions across the whole sweep total exactly the control count, so no boundary re-ran a write the control ran once.
  • A dangling write refusesThe five prefixes that end on a recorded-but-unfinished write refuse with the reconciliation error and execute nothing. A crash mid-write parks for a human rather than guessing.

crates/salvor-runtime/tests/release_gate.rsrelease_gate_kill_at_every_event_boundary_resumes_identically

Reproduce it

The hero fixture is checked in at examples/hero/.

git clone https://github.com/joseym/salvor && cd salvor cargo build ./target/debug/salvor run --fixture examples/hero

the build puts the fixture's tool server on disk; the run then needs no key and no network