Execution replay
A screenshot shows the page when a step failed. A replay shows how it got there: every action of the run, in the order of your Given, When and Then, with the page inspectable at each one and the network alongside. It is not a video. It is the DOM at every moment, rendered live, so you can open the inspector on the element the step could not find.
{ "trace": "retain-on-failure" }
npx saffron run # a red scenario now keeps a trace
npx saffron trace # every traced scenario, opening on the failed one
npx saffron trace "Successful login" # or name one
npx saffron trace "features/admin.feature:Successful login" # when two features share the name
saffron trace starts a local server, prints its address and opens it in
your browser. When the last run traced several scenarios (all of them,
with "on"), one replay holds them all: Scenario in the header lists
them by feature, each marked ✓, ✦ or ✗, and the replay opens on a failed
one, else one the agent recorded or healed, else the first. The report links every traced scenario to the same command,
and both IDE panels have an Open Replay action.
What you see
| Left: the scenario | Your steps as written, with the outcome of each. A step opens onto the actions Saffron replayed for it, in the words of the recording ("click button "Log in""), and each action onto the Playwright calls behind it. The failed step is selected when you arrive. |
| Middle: the page | The DOM at the selected action, before it, during it or after it, rendered as a live page: inspect it, select text, scroll it. Above it, a filmstrip of the whole run; click a frame to jump to that moment. |
| Right: Details | The failure, the retries a step needed, the log of each call, and, for a step the run changed, its actions against the committed recording, under Heal review, First recording, New recording or Proposed recording by what kind of run it was. A polled assertion is shown once, with how many times it looked and for how long. |
| Right: Network | Every request of the run; the ones made during the selected action in white. |
| Right: Saffron | What Playwright cannot know: the agent's narrative and adaptations after a heal, what the proposal changes, the {unique:...} values of the run, the screenshots. |
Keys: ↑ ↓ move between steps, ← → between actions, b and a
switch to the page before or after the action. They keep working after a
click on a button or a tab, and leave shortcuts with Cmd, Ctrl or Alt
(select all, say) to the browser. The selected step or action scrolls
into view in the list.
With tracing enabled, Agent · actions show what the agent did during
recording and healing, in words ("click Login button", "read the page"),
alongside the cached attempts; the bookkeeping its browser tools do after
each action is folded into one line. Trace switches between the
original execution and each proof replay, including a failed proof
followed by refinement. Each proof keeps its own step outcomes. In a heal,
a step the agent healed is marked ✦ healed by the agent, and the run
opens on it: the failed attempts of the old recording, then what the agent
did. A first recording heals nothing: a step the agent worked through
reads recorded by the agent. Nor does a proof replay, or a run that
replayed a pending proposal without the agent: there, a step the proposal
changes reads ran the proposed recording. When a run replayed a
pending proposal that failed, and then replayed the committed recording or
recorded again, that first attempt is kept apart under Earlier attempt:
the pending proposal, not mixed into the steps. When a heal fails, the
step it failed on is the one the agent had reached, with the agent's
reason, not the step replay first missed. What the agent looked at
before it named a step (Agent, exploring ·) is filed
under the step it named next; if it never named one, under Setup and
agent exploration. Saffron's own setup before the first step (opening the page) is under
Setup outside the steps; a call outside every step that ran while a
step was running is filed under that step.
A step lists its actions and calls in the order they ran, and ← →
follow that order. A step with no snapshot of its own (one the agent only
inspected) shows the page after the last interaction before it, and says
so; a step that never ran shows no page. An action that took no snapshot of its own (reading
the title, say) shows the page as the last snapshot before it left it. With
several tabs open, only a snapshot of the same tab will do: Playwright 1.63
and later do not record which tab such an action ran in, so the replay says
so rather than show another tab. Only the tabs open at that moment count:
a popup opened later, or one already closed, changes nothing. Under a
replayed step, check the page matches the recording is Saffron making
sure the page has not drifted before the step runs; under the agent's
actions, note the page for the recording is Saffron noting the
page's structure, which that check compares against on later replays.
If the agent tracing bridge cannot start, Saffron warns once, with the reason, and uses the standard browser connection for the rest of the run. Recording and healing continue; those agent calls are absent from the traces.
Highlight target marks elements that Playwright actually matched in the selected snapshot. A missing locator has no matched element to highlight; the viewer says so and shows the attempted locator in Details.
In the Saffron panel, Accept recording promotes the pending recording
and Reject proposal discards it. These act only on the exact proposal
associated with this run. If it has been replaced, accepted or rejected
since, the viewer refuses both. If it became stale against the scenario or
committed cache, or can no longer be read, it cannot be accepted, but
Reject proposal still discards it. Older reports without a proposal
identity remain view-only. An unverified recording, or one this run showed
failing, requires an explicit acknowledgement. The panel reads the
proposal again on every load, so a decision made elsewhere shows on
reload. When the agent also suggests rewriting steps of the feature file,
the panel lists the rewrites: Accept recording drops them, and
saffron accept --with-feature-edit <file> applies them. Propagation to
other recordings still uses the CLI.
What is recorded
| Mode | Keeps |
|---|---|
off (default) |
Nothing. Replay runs at full speed. |
retain-on-failure |
A trace for every scenario that was not green: red, and yellow (recorded or healed by the agent, or a pending proposal replayed). The recommended setting for CI. |
on |
A trace for every scenario, green ones included. For a local investigation, or to see what a passing run does. |
Recording a trace costs some speed (snapshots of the DOM around every
action) and disk (a few hundred kilobytes to a few megabytes per
scenario). Traces live in .saffron/artifacts/<feature>/<scenario>/trace.zip,
with proof sessions in proof-1.zip, proof-2.zip, alongside the
screenshots. They are git-ignored by saffron init, kept for the latest
run only, and removed by saffron prune with the rest of a gone
scenario's evidence.
An open replay keeps the traces and screenshots of the scenario it opened
on, even if another run replaces the files; restart saffron trace to see
the new run. A tab left open from before the restart says to reload the
page rather than load the new run's files. The other scenarios of a replay are read when you first pick
them, and one that a later run has replaced says so instead of showing
the newer run; one that cannot be read (a folder you have no access to,
say) gives the reason. A run stopped before it wrote its report leaves
the previous report pointing at its own, newer files: saffron trace refuses such a trace
and the IDE panels leave out such a screenshot, so a replay never mixes
two runs.
Secrets
{env:VAR} values are typed into the page for real, so a trace would hold
them: in the call that typed them, in the log line that says so, in the
DOM snapshot of the field, and in the network traffic that sends them.
Before the trace is kept, Saffron rewrites the value back to its token
everywhere the trace holds text: the event log, the DOM snapshots and
stylesheets, and every request and response body, URL, header and
WebSocket message, in the spellings the value takes there (JSON escapes,
percent-encoding in URLs and form posts as browsers and servers write it,
HTML entities). That includes a {env:VAR} written as a value in a data
file. The unmasked trace never reaches .saffron/artifacts/, even
if the run is killed while it is being saved.
What it cannot rewrite: a value shorter than three characters (too short
to mask), a value inside base64 or another binary encoding (a Basic auth
header, an uploaded file), and pixels. It also keeps the scheme, host and
port of every URL, so the replay can still load the page's stylesheets
and images: an environment or tenant name in a host name stays visible,
and a variable that holds only an address (BASE_URL=https://staging.acme.test)
is left as it is. A secret anywhere else in a URL is masked, credentials
(user:pass@) included, so a variable holding a whole webhook or
database URL keeps only its scheme and host.
Inside a URL the token is written %7Benv:NAME%7D, the form the browser
resolves it to; in CSS it is written \{env\:NAME\}, so a class or id
named after the value still matches its styles. A value a page displays
in clear text is in the filmstrip; password fields show dots. Look through a
trace before you attach it to a public issue.
How it works
The trace is Playwright's own format, recorded by the browser engine
Saffron replays with, and the page is rendered by the snapshot engine
that ships inside that same Playwright, so capture and rendering use the
same version. A trace you keep for later opens best with the Playwright
version that recorded it. Everything around it is Saffron's: the steps,
the actions, the timeline, the panels and the review data. Playwright's own
viewer can open the same file (npx playwright show-trace .saffron/artifacts/<feature>/<scenario>/trace.zip) and will show the
steps as groups, without the Saffron side.