Install
openclaw skills install @xizhuomengcontin/orca-replayAnswers questions about a past agent run from its recording rather than from memory, and replays or forks that run. Use when asked why an earlier run did something, or to reproduce a failure.
openclaw skills install @xizhuomengcontin/orca-replayA recording is evidence. Your memory of a session is not, and neither is a transcript you were handed — both are missing the tool results, the exit codes, and the files that changed without anyone mentioning it.
The rule: when a question is about something that already happened, read the trace before you answer. Do not reconstruct it. If a recording exists, guessing is the wrong move even when the guess would have been right.
orca_list_runs — newest first, and it names the run each fork came from. Skip this only when the
user clearly means the most recent one; every other tool defaults to run: "last".
orca_show_run gives the whole timeline: model turns with token counts and stop reasons, tool
calls with arguments and results, shell commands with exit codes, and every file the run changed.
Good for orientation, long for a specific question.
orca_graph is usually the better tool. It returns causal edges — which event produced which. Pass
to: <event seq> to get only the chain that produced one event. That is the shape of an answer
to "why did this happen", where the full timeline is the shape of an answer to "what happened".
recorded and inferred differentlyEvery edge from orca_graph is labelled:
recorded — the recorder watched it happen and wrote it into the trace.inferred — derived just now from a rule the edge names. The trace does not vouch for it.Carry that distinction into your answer. "The trace shows the rm at step 14 removed it" and "this
looks like the rm at step 14, going by timing" are different claims, and flattening them into one
confident sentence is the specific failure this tool exists to prevent. Name the rule when you lean
on an inferred edge.
orca_replay re-runs the recording and reports what could not be reproduced — divergences, and
requests the recording could not serve.
What "offline" covers, and what it does not. Every model response comes from the trace and the
proxy forwards nothing upstream, so no provider is contacted and no tokens are spent. An unmatched
request halts the replay rather than falling through to the network, unless --loose was asked for.
That covers the model traffic. It does not cover the agent's own subprocesses: unless the recording
used --tls-intercept — in which case replay re-establishes interception for the hosts it recorded
— a curl, npm install, git push or database call inside a recorded shell command goes
straight out. Replay is not a sandbox; only a network-isolated container makes it one.
What a matching replay proves, and what it does not. It shows the recorded decisions reproduce against today's environment. It cannot show the failure is deterministic, because the model is not being asked again — the same recorded responses are served back. If the user wants to know whether a fresh run would fail the same way, say that replay cannot answer it; that needs real runs.
Replay re-executes the agent, not just its model traffic. The recorded model responses are
served from the trace, but the agent process runs again for real — so every shell command it issued
runs again too. worktree: true isolates repository files and nothing else. Anything the run
touched outside the tree — /tmp, Docker, a local database, a package manager, another host — is
mutated a second time.
So check before the first replay of a run, not after. Read its shell commands with
orca_show_run and tell the user what will re-execute. If any of it reached outside the working
tree, get approval for that specifically or replay inside a container; do not treat the earlier
worktree answer as covering it. A run that only read files and edited the repository is free and
repeatable, and worth replaying before committing to any explanation.
Pass worktree: true. It replays into a scratch copy and leaves the working tree alone.
Without it, replay is destructive for as long as it runs: it restores the recorded filesystem over the working tree and puts the tree back when the replay ends. Uncommitted work is absent in the meantime, and stays absent if the replay is interrupted before it can restore. Run an in-place replay only when the user has been told that and has agreed to it. "They do not appear to be typing" is not consent.
A replay reporting reused=3/5 on an interactive recording is not a partial failure. Harnesses make
calls for themselves — a quota probe, a session-naming request — and a replay does not repeat them.
orca_compare forks one run onto several models from the same checkpoint: same files, same
conversation prefix, so the model is the only variable. Pick the fork point with orca_checkpoints
and pass it as from.
Grade with verify — a shell command whose exit code is the verdict. Use something the repository
already declares ("npm test", "npm run typecheck") or an explicitly local binary
("./node_modules/.bin/tsc --noEmit"), not npx <tool>: with no local install, npx runs whatever
the registry has under that name, and npx tsc resolves a package deprecated in 2016 that is not
TypeScript.
orca_compare uploads the recording to other people's models, and spends real money doing it.
Each model named receives the same files and conversation prefix the original run had — so whatever
that run touched (source, prompts, configuration, anything a credential was pasted into) is sent to
every provider behind those model ids.
And each fork is a live agent, not a replay. From the fork point onward the model is really
being asked, and whatever it decides to do, it does — its shell commands execute for real, and so
does the verify command you pass. Each fork gets its own worktree, so repository files are
isolated per model; nothing outside the tree is. A fork can also take actions the original run never
took, because it is a different model making fresh decisions.
So the approval has three parts, and they are not the same question:
orca scrub is for when the comparison is worth running but the trace is not safe to
send as-is.orca_show_run, and the same answer if it reached
Docker, a database, a deployment or another host: get approval for that specifically, or run the
comparison in an isolated environment.Never run it to satisfy curiosity the user did not express.
Say so plainly rather than falling back to guessing, and offer to start one.
If orca is already installed:
orca record claude # or codex, opencode, openclaw, grok
If it is not, ask before installing it — a global install changes the user's machine, and that is
their call, not a detail of your task. Install a pinned version rather than whatever latest
resolves to today:
npm i -g orcareplay@0.1.2 # ask first
orca record <agent> runs the agent unmodified behind a local proxy. Nothing about the agent
changes; two environment variables get set. Recording a session now is what makes the next "why did
it do that" answerable.
For a run started with a prompt in argv — orca record claude -- -p "…" — the replay is exact. A
session someone typed into replays approximately, because the prompts were never on the wire and
are recovered from the harness's own transcript; orca replay says which is which rather than
papering over it.
orca export last -o run.html writes one self-contained file. A trace holds whatever the run held,
so run orca scrub before sending one anywhere.
Scrubbing is best-effort, not a guarantee. It matches known key shapes and high-entropy strings; it cannot know that a particular internal hostname, customer name, or unreleased feature is confidential to this user. So scrub, then have the user look at what is actually going out, and get their agreement — do not describe a scrubbed trace as safe on the strength of the scrubber alone.
orca record leave no trace, and
nothing here recovers them. The answer to "why did it do that" in an unrecorded session is
honestly "there is no recording", not a reconstruction.orca record claude -- -p "…") replays byte-for-byte.AskUserQuestion, plan mode) are absent when the same agent runs without one, which can make a
replayed request differ from the recorded one by enough to halt.inferred edges are not evidence. They are derived from a named rule at query time. Treat
them as a reading of the trace, never as something the recorder witnessed.--tls-intercept, and some cannot be reached at all. A recording that came back
empty means the harness was not captured, not that nothing happened.| tool | arguments | notes |
|---|---|---|
orca_list_runs | — | newest first, names the parent of each fork |
orca_show_run | run | the full timeline |
orca_checkpoints | run | where a fork can start |
orca_graph | run, to | causal edges; to narrows to one chain |
orca_replay | run, worktree | offline, free, repeatable |
orca_compare | run, models*, from, verify | spends real tokens |
run accepts a run id or "last", and defaults to "last". Replay traces are skipped when
resolving "last", so it means the newest run you actually recorded.