Files
sanderling/.github/workflows/replay-ui.yml
T
pj 11f72a722a follow-ups from the pr #73 review (#77)
* ci(folio): run gradle on jdk 21 for the metro plugin

the metro gradle plugin folio builds with publishes org.gradle.jvm.version 21
and java 21 class files, so every leg failed at the folio build on a 17
runtime. local builds pass on jdk 25, which is why only ci saw it.

* fix(build): clean pkg/spec/dist, not the dead spec-api path

* chore: point stale spec-api comments at pkg/spec

* fix(spec): publish src so an installed package carries the runtime entries

* fix(testrun): alias the installed spec package so one module graph loads

* fix(spec): export Direction, ScrollAction and LongPressAction from the entry

* docs(spec): cut the package readme to a description and doc links

* docs: say how the cli and spec package versions relate

* fix(verifier): report whether the last action was confirmed applied

Both hosts get applied: true when the runner saw the dispatch succeed and
applied: null when it could not, so an unconfirmed action stops arriving at
the spec as no action at all.

* fix(runner): an apply error leaves the action's fate unknown, not undone

A deadline that fires after the tap was dispatched leaves the effect
committed. Reporting nil made the spec see an effect with no action to cause
it, which is how the counting property convicts a healthy app.

* fix(release): stage the sidecar jar at the renamed embed path

* test(replay-ui): trace fixtures for the vacuity counts

one real green run, one run that rendered nothing, one that judges every property at least once.

* ci(replay-ui): count the steps each property judged

the exit code says no property returned false; it does not say any property was ever evaluated. this reads the trace and reports judged vs declined per property, and fails when the step page never rendered.

* test(replay-ui): cover the summary script from make test

* ci(replay-ui): summarise through the vacuity script

* docs(ci): explain the replay-ui judged/declined counts

* fix(verifier): encode element-valued extractors into the trace

An ax element exports with its find/findAll host functions attached, and
json.Marshal refuses the whole value over them: json: unsupported type:
func(goja.FunctionCall) goja.Value. The encoding failed, curr stayed nil,
and the goja hosts (ios, android) recorded null for every element-valued
extractor in both the per-step diff and the violation witness.

Apply the web host's sanitize rule before marshaling, so one rule encodes
an element on both hosts.

* test(verifier): pin element encoding to one rule on both hosts

* test(runner): assert an element reaches trace.jsonl and its witness

* feat(spec): give state.lastAction an applied field

Three states, not two: no action is a null lastAction, applied: true is an
action the runner confirmed, applied: null is one it dispatched and never
learned the fate of.

* fix(folio): do not attribute an effect to an unconfirmed action

submitChangesBalanceByTypedAmount and createdAccountHasNonZeroBalance both
convict by pinning an effect on the last action, so both decline unless the
runner saw it applied. The fixtures now say which fate they mean.

* test(folio): an unconfirmed submit belongs in the window

The count is an upper bound on the submits a window holds, so the tap that may
have landed counts and committedTransactionsExceedSubmits has nothing to
convict on.

* test(runner): a tap that lands under a failed apply is not a double submit

Drives the real folio counting predicates through the runner against a device
that commits the tap and then times out. The double-submit case is the control:
without it a green proves only that the property never fired.

* test(verifier): pin the three lastAction states on both hosts

The web page is handed the same applied field the goja object exposes, so a
property cannot read one thing on native and another on web.

* docs(spec-language): document the three lastAction states
2026-08-15 15:51:33 +05:30

143 lines
5.0 KiB
YAML

name: replay-ui
# Sanderling fuzzing sanderling's own replay UI. Dispatch-only: it takes minutes
# and it is a demo of the product loop, not a merge gate.
#
# The shape is: produce a real trace, serve it with `sanderling replay`, then run
# a spec against that UI. Any violation fails the job. Six of the seven
# properties in replay-ui/sanderling/spec.ts are cross-panel agreements that hold
# for any trace; the seventh is the stock noUncaughtExceptions. None of them
# needs recalibrating when the fixture changes.
on:
workflow_dispatch:
inputs:
seed:
description: seed for the dogfood run
default: "3"
max-steps:
description: step budget for the dogfood run
default: "80"
permissions:
contents: read
jobs:
dogfood:
timeout-minutes: 45
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- name: Set up Go
uses: actions/setup-go@v5
with:
go-version-file: go.mod
cache: true
- name: Set up bun
uses: oven-sh/setup-bun@v2
with:
bun-version: "1.3.13"
# Pinned stable plus the AppArmor sysctl: the same setup ci.yml's browser
# job needs to get headless Chrome up on ubuntu-latest.
- name: Set up Chrome
uses: browser-actions/setup-chrome@v1
with:
chrome-version: stable
- name: Allow Chrome under unprivileged user namespaces
run: sudo sysctl -w kernel.apparmor_restrict_unprivileged_userns=0
- name: Verify headless Chrome starts
run: |
chrome --version
chrome --headless --no-sandbox --disable-gpu --disable-dev-shm-usage \
--dump-dom 'data:text/html,<title>ok</title>'
# The UI the spec drives is the one embedded in this binary, so the build
# has to come after any change to replay-ui/src.
- name: Build sanderling
run: make sanderling-web
# A trace with a violation and uncaught exceptions in it, so the UI has
# something to render in every panel the spec looks at. No
# --exit-on-violation here: the run is the fixture, and stopping it at the
# first violation would leave a four-step trace to fuzz.
- name: Record a fixture trace
run: |
python3 -m http.server 8792 --bind 127.0.0.1 \
--directory test/browser/testdata/throwing &
ready=""
for _ in $(seq 1 30); do
curl -sf http://127.0.0.1:8792/ >/dev/null && { ready=1; break; }
sleep 1
done
if [ -z "$ready" ]; then
echo "the fixture http server never answered on 127.0.0.1:8792" >&2
exit 1
fi
./bin/sanderling test \
--platform web \
--spec test/browser/testdata/throwing/spec.ts \
--bundle-id http://127.0.0.1:8792/ \
--duration 5m --max-steps 25 --seed 7 \
--output runs/fixture
- name: Serve the trace with sanderling replay
run: |
# Flags before the positional argument: Go's flag package stops
# parsing at the first non-flag word.
./bin/sanderling replay --port 8793 --no-open runs/fixture &
ready=""
for _ in $(seq 1 30); do
curl -sf http://127.0.0.1:8793/api/runs >/dev/null && { ready=1; break; }
sleep 1
done
if [ -z "$ready" ]; then
echo "sanderling replay never served /api/runs on 127.0.0.1:8793" >&2
exit 1
fi
run_id="$(ls runs/fixture | head -1)"
echo "RUN_URL=http://127.0.0.1:8793/runs/$run_id/steps/1" >> "$GITHUB_ENV"
curl -sf "http://127.0.0.1:8793/runs/$run_id/steps/1" >/dev/null
# Inputs go through env rather than into the script text: a `${{ }}` is
# substituted before bash ever sees the line, so a seed of `$(id)` would
# run as a command.
- name: Fuzz the replay UI
run: |
./bin/sanderling test \
--platform web \
--spec replay-ui/sanderling/spec.ts \
--bundle-id "$RUN_URL" \
--duration 10m \
--max-steps "$MAX_STEPS" \
--seed "$SEED" \
--exit-on-violation \
--output runs/dogfood
env:
SEED: ${{ inputs.seed }}
MAX_STEPS: ${{ inputs.max-steps }}
# Exit 0 above means no property returned false. It does not mean any
# property was ever evaluated against real content: they all decline to
# judge when the elements they read are absent, so a run that never
# rendered the step page is green and worthless. This step is what tells
# the two apart, and it fails the job when nothing was judged.
- name: Summarise
if: always()
run: .github/scripts/replay-ui-summary.sh runs/dogfood
env:
SEED: ${{ inputs.seed }}
MAX_STEPS: ${{ inputs.max-steps }}
- name: Upload runs
if: always()
uses: actions/upload-artifact@v4
with:
name: replay-ui-runs
path: runs/
retention-days: 14