The workflows that fuzz the examples are dispatch-only, and GitHub will
not dispatch a workflow that is not on the default branch, so their first
real run is after merge. actionlint and the reference checker are the
only things that can fail before that.
actionlint is pinned by commit, and its tool version is pinned too so a
new release cannot change what CI enforces.
actionlint reads a local action's inputs but never checks its path
exists: uses: ./.github/actions/typo lints clean and fails only when the
job runs. Covers composite action paths, make targets including the ones
the examples matrix builds from $SANDERLING, and the scripts a run: step
invokes plus their executable bit.
Fails when it parses fewer references out of a file than that file
mentions, because a checker that matches nothing reports a safety it
never looked for.
The android health gate grepped the trace for "AddTransactionScreen",
which occurs in exactly one place: the hierarchy dump, as a resource-id.
That is a debug artifact standing in for a fact the spec already reports,
and it was wrong in both directions. Against the real 8.4MB trace with
hierarchy stripped, the old gate failed a healthy 200-step run; against a
trace carrying the marker on a transition frame without the route ever
being reported, it passed and called it healthy.
It reads extractor_changes.route now, whose values come from SCREENS in
the spec, so the gate and the app agree on what being on a screen means.
routeOf answers null on a frame showing two screens, which is exactly the
frame the marker was matching.
The drift check grows to cover both new names: extract("route") and the
SCREENS key. The fixtures are re-derived keeping the route entry and no
hierarchy at all; the full artifacts and the fixtures give byte-identical
verdicts, which is what proves the coupling is gone.
Script and fixtures move together: either alone leaves the suite red.
Four cases on real data: each leg's real verdict, plus the android trace
cut before it reached the transaction screen, which is what proves the
route gate reads a real hierarchy dump.
Also corrects the hand-written fixtures. They set is_error to false on a
plain violation; internal/trace/writer.go tags that field omitempty, so a
real trace omits it entirely. Harmless to the classifier, but a fixture
that does not look like reality is the thing that hides drift.
ios convicted on submitCommitsOneTransactionPerAction, web on both gated
properties, android ran its full 200 steps healthy. Every step is kept;
of each step only step, violations, witnesses and residuals survive.
hierarchy is replaced by the quoted "...Screen" resource ids it held, in
order. It cannot just be dropped: it is 95% of the bytes and also the
only place the android route gate's grep can match, so dropping it flips
that leg from healthy to 'never reached'. 8.4MB to 59KB.
runs/dogfood and the '### replay-ui dogfood' heading carried the same
naming error as the job name: dogfooding is why the run exists, not what
it fuzzes.
buf-setup-action was already pinned with a comment saying why; the other
five rode mutable major tags, so a tag move is an unreviewed change to
what runs. Each major currently resolves to the release named in the
comment, so this freezes today's behaviour rather than changing it.
actions/* stay on major tags: they are first-party to the runner.
folio.yml and replay-ui.yml ran the same operation: build sanderling for
a platform, bring a target up, run a spec against it, classify the trace,
upload the run. They are now one matrix over four examples, each naming
its own runner.
The job is named for what it fuzzes. 'dogfood' named why we run it, not
what runs, the same error as a diagnostic that reports a motivation
instead of an observation.
The matrix is computed by a plan job because jobs.<id>.if cannot read the
matrix context, so a static matrix has no way to leave a leg out. Seeds,
budgets, timeouts, runners and artifact names are unchanged.
Records a trace and serves it with sanderling replay. The step page URL
is a composite output rather than GITHUB_ENV, so it is scoped to the one
step that drives it.
folio-app holds the per-platform toolchain and app build, so a caller
guards one step instead of eight. folio-simulator boots the simulator,
installs folio and leaves the app stopped.
The setup-chrome / apparmor sysctl / launch-check trio is copied across
three jobs. The old comment described setup-chrome v1 semantics: under v2
stable is the default and the alternative is Chrome for Testing latest,
not a dev Chromium, so it is restated for what the pin actually does.
21 cases through a stubbed sanderling: every exit path, the drift check,
a missing trace, a zero-byte trace, an empty glob and a truncated line.
Asserts the flags that reached the binary, not just the exit code.
Invoked as bash -eo pipefail -c, which is what a run: block does. Running
folio-run.sh itself under -e would kill it at the first non-zero
sanderling test, which is the exit code it exists to read.
Nothing tied GATED_PROPERTIES to the spec it gates. Renaming a property
left the classifier matching nothing: ios and web blamed the spec for
finding a different bug, and android silently reclassified a real
conviction as 'judging health only' and stayed green.
replay-ui-summary.sh already makes this check for its own list. The spec
path becomes SPEC-overridable the same way, so the check is testable.
run_dir is empty when the run produced no output directory, and the
fallback made trace ./trace.jsonl. A stray trace in the working directory
was then read as this run's, so a run that wrote nothing reported 'found
the submit bug' and exited 0, defeating the missing-trace check below it.
A refname is attacker-controlled and git permits backtick, $, (, ; and |
in it. Three sites substituted it into a run: block, and NODE_AUTH_TOKEN
sat at job level, so a pushed tag ran arbitrary commands with the publish
credential in reach.
The tag now goes through env:, is validated against an anchored version
pattern before anything consumes it, and reaches the other jobs as a job
output. The token is scoped to the publish step. release-npm declares
contents: read instead of inheriting the repo default.
same hole as the simulator path: devicectl install over an app keeps its data, and the discarded uninstall error hid it. Uninstalling a bundle id that is not installed exits 0 with 'App uninstalled.' on a paired iPhone, so a failure here is always real.
simctl install over an installed app carries its data container across, so discarding the uninstall error reported a clear-state that never happened. Uninstalling an app that is not installed exits 0 on a booted simulator, so every failure here is a real one.
adb uninstall answers Failure [DELETE_FAILED_INTERNAL_ERROR] both when the package was never installed and when it refuses to remove one, so the failure text cannot say which happened and the old code installed over the top either way, keeping the data clear-state was asked to drop. Ask pm path instead, and fall back to pm clear when the app is still there.
it stopped being an equality and became |delta| <= typed, so the old name
demanded more than the property does. renamed with the ci gate's
GATED_PROPERTIES in the same commit so the gate never sees a name it does
not know.
The freshness rule rests on the app popping one entry back to the ledger,
not on anything the frame carries, so the assumption and the measurements
behind it belong next to it.
The interaction that keeps a stale Home card list from ever banking counts
the budget has already forgotten: a submit lands on the ledger, so the
reading that resets the window is a whole action later and the submit is
still in it. Characterization, not a regression: no code changed and it
cannot go red first.
--launcher-activity does not exist in cmd/sanderling/main.go. --device,
--android-app-path and --arm do and were undocumented. runs.md still listed
--max-steps and --exit-on-violation as unshipped, and described --clear-data
as opt-in when the default is already true, contradicting itself ten lines on.
covers hooks, extractors, selectors, properties, actions and the order to write them in, with a complete sample spec that typechecks against the real export surface.
The write finishes before AddTransactionViewModel navigates, but nothing
establishes that Home's total has re-rendered before the frame is read, and
an equality convicts a healthy app for a total one frame behind. A delta of
zero is exactly the shape nine of the eleven measured android false
convictions had. 2x still exceeds x, so all four recorded convictions
survive, checked against the traces.
The trade is real: a balance that moves by LESS than the amount typed is no
longer judged anywhere in this spec.
the leg disables animations and the number was measured with them on. the
first real dispatch carries 4 transitional steps over 200, so the cross-fade
wait does still fire in ci, just far less often.
The check that decides whether the select-all worked looked for an
"editable" attribute maestro's tree does not carry, so it answered
"cannot tell" every time and every erase paid the per-character
fallback. Worse, an open keyboard puts a second focused node in the
tree, one of the IME's own keys, carrying no text: taking the first
focused node would read a field still holding 4096 characters as empty,
which is the one answer that stops the erase early.
Match the text field by class instead. Measured against the real
backend, 4096 characters now clear in 385ms on API 34, 409ms on API 35
and 870ms on API 36, verified empty, where the fallback took ~4s.
maestro's eraseText sends one delete per character through its
instrumentation, measured 29.6 ms/char on the API 34 emulator. The
4096-character string the corpus types cost ~121s to clear, a fifth of a
20 minute run spent on one step, and it recurred every time that field
was typed into again.
Select the content and delete the selection instead: two key events at
any length, measured 0.15s to 1.16s for 4096 characters across API 34,
35 and 36. The result is read back off the tree, and a field that is not
empty, or that the tree cannot report on, is finished off per character
in batches rather than assumed clear.
Fixes#80
New performs the clear-state reset, so the uninstall and reinstall no
longer land underneath a live XCTest session that is already bound to
the app. Launch refuses a clear-state request the driver was not built
for rather than reinstalling under its own session.
An unreadable dumpsys passed a null owner to typeChunks, which switches
the mid-type focus guard off outright and lets the rest of the string
spray into whatever holds the foreground. Fall back to the launched
bundle instead: the guard stays armed, typing still happens, and the
degradation is said out loud rather than assumed away.
The window is an upper bound on the transactions an interval could hold, and
a bound inflated by taps that commit nothing is a bound the app can never
exceed: #78 read a rise of 15 transactions against 37 submits. TxnSubmit is
clickable(enabled = amount.isNotBlank()) and parseCents refuses anything its
regex misses, so a tap whose landing frame shows a refused amount cannot have
committed. Over four recorded android runs that is 19, 11, 25 and 25 of 35,
26, 42 and 42 submit taps.
A relaunch is excepted: a fresh process draws an empty field whatever was
submitted.
Defaulting the count to zero made a dumpsys that said nothing mean
nothing is animating, so a degraded link broke out of the settle early
and handed the runner a frame caught mid-animation. Unknown now waits,
inside the deadline waitForIdle already holds.
adbOutput and readLogcat read to EOF and then waited with no timeout, so
a wedged adb held the step for as long as it liked; one stall over a
remote adb server measured ~100s. The bound has to sit on the read, not
on waitFor: a wedged adb never reaches EOF, so a bounded waitFor after
the read is a line that never runs.
The proof behind the comment: "Travel1" holding 25 transactions and
"Travel12" holding 5 merge to the same string, so no identity key read off
a web card can tell them apart.
Folio rejects a duplicate account name, so the twin the drop rule guards
against is two names the web key cannot tell apart, not two accounts
sharing a name. State what closing the rest would cost and what the tree
would have to carry to close it properly.
The runner now keeps lastAction and marks it relaunched: true where it used
to report nothing at all, so the two properties that demand an effect judge
a step whose process may have died before the write landed.
submitChangesBalanceByTypedAmount and createdAccountHasNonZeroBalance both
decline there. The counting bound does not: a relaunch cannot manufacture a
transaction, and the submit is counted, so declining would throw away the
detection the runner fix restored.
A tap on a text field raises the keyboard and nothing closed it, so the
tree the picker chooses from was missing every app node underneath it,
the submit control included. Close it before the read rather than after
the tap: the picker only ever sees snapshots, and the keyboard is still
on its way up when the tap returns.
Fixes#78
buildDadb hardcoded localhost:5037, so a serial-addressed device always
resolved through this machine's adb server and ADB_SERVER_SOCKET was
ignored. Read the endpoint the way the adb CLI does instead.
Fixes#79
submitCommitsOneTransactionPerAction now states its rule over two windows:
the Home counts it already compared, and the account balance the ledger and
the add-transaction screen redraw on nearly every frame of the transaction
flow. Same rule, and the second window is usually one action wide.