ci(examples): one dispatch workflow for every example

folio.yml and replay-ui.yml ran the same operation: build sanderling for
a platform, bring a target up, run a spec against it, classify the trace,
upload the run. They are now one matrix over four examples, each naming
its own runner.

The job is named for what it fuzzes. 'dogfood' named why we run it, not
what runs, the same error as a diagnostic that reports a motivation
instead of an observation.

The matrix is computed by a plan job because jobs.<id>.if cannot read the
matrix context, so a static matrix has no way to leave a leg out. Seeds,
budgets, timeouts, runners and artifact names are unchanged.
This commit is contained in:
pj committed 2026-08-16 01:22:32 +05:30
1 parent 78ea4dd425
commit 52ee0ad424
3 files changed
+193 -419

No files matched your search

+193
View File
@@ -0,0 +1,193 @@
name: examples
# Every example sanderling ships, fuzzed the same way: build sanderling for a
# platform, bring the target up, run a spec against it, classify the trace it
# wrote, upload the run. Only the bring-up differs, and that lives in the
# per-target actions under .github/actions/.
#
# Dispatch-only: these take tens of minutes and they demonstrate the product
# loop, they do not gate a merge.
#
# folio on ios and in the browser expect the bug: folio double-submits a
# transaction on a double tap, so the run is supposed to end with exit 2. Exit 0
# means the fuzzer stopped finding a bug that is still there; exit 1 means the
# harness broke. The two are worth telling apart, which is why
# --exit-on-violation exits 2 and not 1.
#
# folio on android is a health gate. It convicts in four runs out of five, which
# is real evidence but not a gate: the fifth would report a regression it had
# not found. Its budget is set so the conviction it usually gets is a bonus.
#
# the replay ui leg fuzzes sanderling's own replay UI, and any violation fails
# it. Its properties are cross-panel agreements that hold for any trace, so none
# of them needs recalibrating when the fixture changes.
on:
workflow_dispatch:
inputs:
targets:
description: which examples to fuzz
type: choice
options: [all, folio, android, ios, web, replay-ui]
default: all
seed:
description: seed override (0 = each target's calibrated seed)
default: "0"
max-steps:
description: step budget override (0 = each target's calibrated budget)
default: "0"
duration:
description: wall-clock budget override (empty = each target's calibrated budget)
default: ""
permissions:
contents: read
jobs:
plan:
runs-on: ubuntu-latest
permissions: {}
outputs:
examples: ${{ steps.pick.outputs.examples }}
steps:
# The matrix is built here rather than written out under strategy.matrix
# because a job-level `if:` cannot read the matrix context, so a static
# matrix has no way to leave a leg out. jq -c keeps the value on one line,
# which is what makes the $GITHUB_OUTPUT write below safe.
- name: Pick the examples to fuzz
id: pick
run: |
examples='[
{"target":"android","name":"folio on android","app":"folio",
"runs-on":"ubuntu-latest","timeout":90,"sanderling":"android",
"seed":"9","max-steps":"200","duration":"20m","artifact":"folio-android"},
{"target":"ios","name":"folio on ios","app":"folio",
"runs-on":"macos-15","timeout":90,"sanderling":"ios",
"seed":"7","max-steps":"240","duration":"20m","artifact":"folio-ios"},
{"target":"web","name":"folio in the browser","app":"folio",
"runs-on":"ubuntu-latest","timeout":60,"sanderling":"web","chrome":true,
"seed":"3","max-steps":"240","duration":"20m","artifact":"folio-web"},
{"target":"replay-ui","name":"the replay ui","app":"replay-ui",
"runs-on":"ubuntu-latest","timeout":45,"sanderling":"web","chrome":true,
"seed":"3","max-steps":"80","duration":"10m","artifact":"replay-ui-runs"}
]'
picked="$(jq -c --arg want "$TARGETS" \
'map(select($want == "all" or .target == $want or .app == $want))' \
<<<"$examples")"
if [ "$picked" = "[]" ]; then
echo "examples: '$TARGETS' selects no example, so this dispatch would run nothing" >&2
exit 1
fi
echo "examples=$picked" >> "$GITHUB_OUTPUT"
env:
TARGETS: ${{ inputs.targets }}
fuzz:
needs: plan
name: fuzz ${{ matrix.name }}
runs-on: ${{ matrix.runs-on }}
timeout-minutes: ${{ matrix.timeout }}
strategy:
# Each leg is its own evidence. One target failing must not cancel the
# others, which is how these ran as separate jobs.
fail-fast: false
matrix:
include: ${{ fromJSON(needs.plan.outputs.examples) }}
env:
SEED: ${{ inputs.seed != '0' && inputs.seed || matrix.seed }}
MAX_STEPS: ${{ inputs.max-steps != '0' && inputs.max-steps || matrix.max-steps }}
DURATION: ${{ inputs.duration != '' && inputs.duration || matrix.duration }}
IOS_DEVICE: iPhone 16 Pro
steps:
- uses: actions/checkout@v7
- name: Set up Go
uses: actions/setup-go@v7
with:
go-version-file: go.mod
cache: true
- name: Set up bun
uses: oven-sh/setup-bun@0c5077e51419868618aeaa5fe8019c62421857d6 # v2.2.0
with:
bun-version: "1.3.13"
- name: Set up headless Chrome
if: matrix.chrome
uses: ./.github/actions/headless-chrome
- name: Build the folio app
if: matrix.app == 'folio'
uses: ./.github/actions/folio-app
with:
platform: ${{ matrix.target }}
# The UI the replay-ui spec drives is the one embedded in this binary, so
# the build has to come after any change to replay-ui/src.
- name: Build sanderling
run: make "sanderling-$SANDERLING"
env:
SANDERLING: ${{ matrix.sanderling }}
- name: Put folio on the simulator
if: matrix.target == 'ios'
uses: ./.github/actions/folio-simulator
- name: Serve a trace to fuzz
id: fixture
if: matrix.target == 'replay-ui'
uses: ./.github/actions/replay-ui-fixture
- name: Fuzz folio on an emulator
if: matrix.target == 'android'
uses: reactivecircus/android-emulator-runner@a421e43855164a8197daf9d8d40fe71c6996bb0d # v2.38.0
with:
api-level: 34
target: google_apis
arch: x86_64
emulator-options: -no-window -gpu swiftshader_indirect -no-snapshot -noaudio -no-boot-anim
disable-animations: true
script: .github/scripts/folio-run.sh android
- name: Fuzz folio
if: matrix.app == 'folio' && matrix.target != 'android'
run: .github/scripts/folio-run.sh "$TARGET"
env:
TARGET: ${{ matrix.target }}
# Inputs go through env rather than into the script text: a `${{ }}` is
# substituted before bash ever sees the line, so a seed of `$(id)` would
# run as a command.
- name: Fuzz the replay UI
if: matrix.target == 'replay-ui'
run: |
./bin/sanderling test \
--platform web \
--spec replay-ui/sanderling/spec.ts \
--bundle-id "$RUN_URL" \
--duration "$DURATION" \
--max-steps "$MAX_STEPS" \
--seed "$SEED" \
--exit-on-violation \
--output runs/replay-ui
env:
RUN_URL: ${{ steps.fixture.outputs.url }}
# Exit 0 above means no property returned false. It does not mean any
# property was ever evaluated against real content: they all decline to
# judge when the elements they read are absent, so a run that never
# rendered the step page is green and worthless. This step is what tells
# the two apart, and it fails the job when nothing was judged. folio's
# legs make the same call inside folio-run.sh, where the exit code it is
# judging is in scope.
- name: Classify the replay UI run
if: ${{ always() && matrix.target == 'replay-ui' }}
run: .github/scripts/replay-ui-summary.sh runs/replay-ui
- name: Upload the run
if: always()
uses: actions/upload-artifact@v7
with:
name: ${{ matrix.artifact }}
path: runs/
retention-days: 14