+
+
+ Property-Based Testing for Mobile UI
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
+
Property-Based Testing for Mobile UI
+
Finding harder bugs earlier
+
+
+
+
+
Hey, I'm PJ
+
+
Platform stuff @ OkCredit(YC S18)
+
I'm here to talk about tough bugs
+
+
+
+
+
+
+
Let's talk testing
+
+
+
+
+
+
+
We write unit tests
+
+
+@Test fun `balance of a single $5 debit is -$5`() {
+ val txns = listOf(Transaction(type = debit, amount = 500))
+ assertEquals(-500, balanceOf(txns))
+}
+
+
+
+
+
+
+
We write integration tests
+
+
@Test
+fun `submitting a transaction updates balance`() {
+ /*
+ 1. open the app, log in
+ 2. tap an account
+ 3. tap "Add Transaction"
+ 4. type 50, tap Submit
+ 5. go back home
+ */
+
+ assertEquals(Balance(5000), homeScreen.totalBalance())
+}
+
+
+
+
+
+
+
All tests pass
+
+
+
+
+
+
+
But user reported a bug
+
+
+
+
+
+
There are duplicate transactions,
+ same amount, same time.
+
+
+
+
+
+
Why? We wrote tests for this and they passed????
+
+
+
+
+
+
The problem
+
Example tests cover the paths we thought of.
+
+
Bugs live in the paths we didn't cover.
+
+
+
+
+
Property-Based Testing
+
+
+
Write invariant properties for testing, not just examples
+
The machine generates 100s of random actions. Invariant has to
+ hold for all of them
+
+
+
+
+
+
Property-Based Testing
+
+
+
Example test
+
+ "When I type 50 and tap Submit,
+ balance becomes $50"
+
+
+
+
Property
+
+ "Transaction submit moves the balance exactly by the typed amount"
+
+
+
+
+
+
+
+
+
+
Sanderling
+
+
Autonomous testing based on specifications
+
Explore application on device and finds invalid behaviors
+
Works directly on the device independent of application framework
+
+
+
+
+
+
Sanderling
+
+ Specifying and validating invariant properties
+
Checks all properties against the current state, recording violations
+
Selects the next action based on the current state using a fuzzer, and performs it on device
+
Repeats the process
+
+
+
+
+
+
When we run it for our spec
+
+
+
+
+
+
Found the bug
+
+
+
+
+
+
With example test
+
@Test
+fun `adding $50 credit increases balance by $50`() {
+ // type 50
+ // tap Submit <- just once, obviously
+ // go back home
+
+ assertEquals(5000, home.totalBalanceCents())
+}
+
+
passed
+
+ We never thought to tap twice. The test didn't either.
+
+
+
+
+
+
With property test
+
+Step 14 Tap -> AddTransactionButton
+Step 15 InputText -> TxnAmountField "49"
+Step 16 Tap -> TxnSubmit
+Step 17 Tap -> TxnSubmit
+
+Property violated: submitMovesBalanceByTypedAmount
+ expected balance change: 4900 cents
+ actual balance change: 9800 cents <- 2x the amount
+
+
violation
+
+ Step 17 was random. We never wrote it.
+
+
+
+
+
Recap
+
+
Example-based testing is limited with action space
+
Property-based testing lets us specify invariants that must hold for randomized action space
If you want to use this or contribute, please come talk to me.
+
+
+
+
+
+
+
+
+
+
+
+
+
\ No newline at end of file
diff --git a/talk/pbt-loop.excalidraw b/talk/pbt-loop.excalidraw
new file mode 100644
index 0000000..b8b333c
--- /dev/null
+++ b/talk/pbt-loop.excalidraw
@@ -0,0 +1,29 @@
+{
+ "type": "excalidraw",
+ "version": 2,
+ "source": "excalidraw.com",
+ "elements": [
+ {"id":"b1","type":"rectangle","x":130,"y":20,"width":200,"height":60,"angle":0,"strokeColor":"#c8b89a","backgroundColor":"transparent","fillStyle":"solid","strokeWidth":2,"strokeStyle":"solid","roughness":1,"opacity":100,"seed":101,"version":1,"versionNonce":101,"isDeleted":false,"boundElements":null,"updated":1,"link":null,"locked":false,"groupIds":[]},
+ {"id":"t1","type":"text","x":130,"y":36,"width":200,"height":28,"angle":0,"strokeColor":"#e0d8c8","backgroundColor":"transparent","fillStyle":"solid","strokeWidth":1,"strokeStyle":"solid","roughness":1,"opacity":100,"text":"pick a random action","fontSize":16,"fontFamily":1,"textAlign":"center","verticalAlign":"middle","seed":102,"version":1,"versionNonce":102,"isDeleted":false,"boundElements":null,"updated":1,"link":null,"locked":false,"groupIds":[],"baseline":24,"containerId":null,"originalText":"pick a random action","lineHeight":1.25},
+
+ {"id":"a12","type":"arrow","x":230,"y":80,"width":0,"height":30,"angle":0,"strokeColor":"#c8b89a","backgroundColor":"transparent","fillStyle":"solid","strokeWidth":2,"strokeStyle":"solid","roughness":1,"opacity":100,"points":[[0,0],[0,30]],"lastCommittedPoint":null,"startBinding":null,"endBinding":null,"startArrowhead":null,"endArrowhead":"arrow","seed":103,"version":1,"versionNonce":103,"isDeleted":false,"boundElements":null,"updated":1,"link":null,"locked":false,"groupIds":[]},
+
+ {"id":"b2","type":"rectangle","x":130,"y":110,"width":200,"height":60,"angle":0,"strokeColor":"#c8b89a","backgroundColor":"transparent","fillStyle":"solid","strokeWidth":2,"strokeStyle":"solid","roughness":1,"opacity":100,"seed":104,"version":1,"versionNonce":104,"isDeleted":false,"boundElements":null,"updated":1,"link":null,"locked":false,"groupIds":[]},
+ {"id":"t2","type":"text","x":130,"y":126,"width":200,"height":28,"angle":0,"strokeColor":"#e0d8c8","backgroundColor":"transparent","fillStyle":"solid","strokeWidth":1,"strokeStyle":"solid","roughness":1,"opacity":100,"text":"tap / type on real app","fontSize":16,"fontFamily":1,"textAlign":"center","verticalAlign":"middle","seed":105,"version":1,"versionNonce":105,"isDeleted":false,"boundElements":null,"updated":1,"link":null,"locked":false,"groupIds":[],"baseline":24,"containerId":null,"originalText":"tap / type on real app","lineHeight":1.25},
+
+ {"id":"a23","type":"arrow","x":230,"y":170,"width":0,"height":30,"angle":0,"strokeColor":"#c8b89a","backgroundColor":"transparent","fillStyle":"solid","strokeWidth":2,"strokeStyle":"solid","roughness":1,"opacity":100,"points":[[0,0],[0,30]],"lastCommittedPoint":null,"startBinding":null,"endBinding":null,"startArrowhead":null,"endArrowhead":"arrow","seed":106,"version":1,"versionNonce":106,"isDeleted":false,"boundElements":null,"updated":1,"link":null,"locked":false,"groupIds":[]},
+
+ {"id":"b3","type":"rectangle","x":130,"y":200,"width":200,"height":60,"angle":0,"strokeColor":"#c8b89a","backgroundColor":"transparent","fillStyle":"solid","strokeWidth":2,"strokeStyle":"solid","roughness":1,"opacity":100,"seed":107,"version":1,"versionNonce":107,"isDeleted":false,"boundElements":null,"updated":1,"link":null,"locked":false,"groupIds":[]},
+ {"id":"t3","type":"text","x":130,"y":216,"width":200,"height":28,"angle":0,"strokeColor":"#e0d8c8","backgroundColor":"transparent","fillStyle":"solid","strokeWidth":1,"strokeStyle":"solid","roughness":1,"opacity":100,"text":"check every property","fontSize":16,"fontFamily":1,"textAlign":"center","verticalAlign":"middle","seed":108,"version":1,"versionNonce":108,"isDeleted":false,"boundElements":null,"updated":1,"link":null,"locked":false,"groupIds":[],"baseline":24,"containerId":null,"originalText":"check every property","lineHeight":1.25},
+
+ {"id":"aloop","type":"arrow","x":130,"y":230,"width":0,"height":0,"angle":0,"strokeColor":"#7a7068","backgroundColor":"transparent","fillStyle":"solid","strokeWidth":1,"strokeStyle":"dashed","roughness":1,"opacity":90,"points":[[0,0],[-70,0],[-70,-180],[0,-180]],"lastCommittedPoint":null,"startBinding":null,"endBinding":null,"startArrowhead":null,"endArrowhead":"arrow","seed":109,"version":1,"versionNonce":109,"isDeleted":false,"boundElements":null,"updated":1,"link":null,"locked":false,"groupIds":[]},
+ {"id":"tloop","type":"text","x":20,"y":135,"width":50,"height":20,"angle":0,"strokeColor":"#7a7068","backgroundColor":"transparent","fillStyle":"solid","strokeWidth":1,"strokeStyle":"solid","roughness":1,"opacity":100,"text":"loop","fontSize":13,"fontFamily":1,"textAlign":"center","verticalAlign":"middle","seed":110,"version":1,"versionNonce":110,"isDeleted":false,"boundElements":null,"updated":1,"link":null,"locked":false,"groupIds":[],"baseline":17,"containerId":null,"originalText":"loop","lineHeight":1.25},
+
+ {"id":"aviol","type":"arrow","x":330,"y":230,"width":0,"height":0,"angle":0,"strokeColor":"#c85c5c","backgroundColor":"transparent","fillStyle":"solid","strokeWidth":2,"strokeStyle":"solid","roughness":1,"opacity":100,"points":[[0,0],[70,0],[70,60]],"lastCommittedPoint":null,"startBinding":null,"endBinding":null,"startArrowhead":null,"endArrowhead":"arrow","seed":111,"version":1,"versionNonce":111,"isDeleted":false,"boundElements":null,"updated":1,"link":null,"locked":false,"groupIds":[]},
+
+ {"id":"bviol","type":"rectangle","x":340,"y":290,"width":210,"height":60,"angle":0,"strokeColor":"#c85c5c","backgroundColor":"transparent","fillStyle":"solid","strokeWidth":2,"strokeStyle":"solid","roughness":1,"opacity":100,"seed":112,"version":1,"versionNonce":112,"isDeleted":false,"boundElements":null,"updated":1,"link":null,"locked":false,"groupIds":[]},
+ {"id":"tviol","type":"text","x":340,"y":306,"width":210,"height":28,"angle":0,"strokeColor":"#c85c5c","backgroundColor":"transparent","fillStyle":"solid","strokeWidth":1,"strokeStyle":"solid","roughness":1,"opacity":100,"text":"violation! shrink + report","fontSize":15,"fontFamily":1,"textAlign":"center","verticalAlign":"middle","seed":113,"version":1,"versionNonce":113,"isDeleted":false,"boundElements":null,"updated":1,"link":null,"locked":false,"groupIds":[],"baseline":24,"containerId":null,"originalText":"violation! shrink + report","lineHeight":1.25}
+ ],
+ "appState": {"gridSize": null, "viewBackgroundColor": "transparent"},
+ "files": {}
+}
diff --git a/talk/pbt-loop.png b/talk/pbt-loop.png
new file mode 100644
index 0000000..668c3b1
Binary files /dev/null and b/talk/pbt-loop.png differ
diff --git a/talk/references/quickstorm.txt b/talk/references/quickstorm.txt
new file mode 100644
index 0000000..3e1bf8d
--- /dev/null
+++ b/talk/references/quickstorm.txt
@@ -0,0 +1,98 @@
+Here is a proofread and polished transcript of the presentation. I have corrected the auto-generated transcription errors (such as "quickdrum" to "Quickstrom", "jefferson" to "Jepsen", "peer script" to "PureScript", etc.), removed filler words (like "uh," "um," "you know"), and added proper capitalization, punctuation, and paragraph breaks to make it highly readable while preserving the speaker's original voice and flow.
+
+---
+
+**[Oskar Wickström:]**
+Okay, so I'm going to present Quickstrom today, which is a new tool that I've been working on for specifying and testing web applications. Quickstrom is about autonomous browser testing, and it's based on specifications as opposed to doing browser testing with examples or end-to-end scenarios where you explicitly describe, "take this step, then take this step, do that, and then expect something." Instead, you have these general specifications of what valid behavior of the system is, much like you have properties in property-based testing.
+
+If that's something you've heard of, then Quickstrom, the tool, explores your web application for you and finds behaviors that are invalid. If there are any, it might find those. It does this by basically interacting with the web browser and figuring out what's possible to do. And it can test basically anything that renders to the DOM, which is the Document Object Model of the browser—essentially a tree structure that describes the state of a webpage.
+
+Today I'm going to go through a bit of the background of this project, how it started, and the ideas behind it. I'll present a case study called the *TodoMVC Showdown*, which I've used as a benchmark for the tool itself. Then we'll go into more detail and look at the specification language and how you use Quickstrom to check web applications in practice. Finally, I'll talk a bit about what might happen in the future, some plans and ideas I have, and then we'll take questions at the end.
+
+So, we'll begin with the background here. I have many, many interests in life, and while this is not an exhaustive list, these are some relevant interests that have influenced this project. Since I started programming, I've done web development—started with PHP and all that—and it's been with me ever since. Still, in most of the projects or work that I do, it's somehow related to the web. So that's an interest, and it's also the huge scale of everything using the web. I find that having an impact in this area is very motivating.
+
+One such impact might be to improve browser testing in some way. It seems like most browser testing tools are based on scenarios and examples, so I wanted to see what could be done there. In the last few years, I've been digging more and more into property-based testing. I've used it a lot in my own projects and at work. I've also written a bunch of blog posts, articles, and a short book on property-based testing and how I used it in a screencast video editor.
+
+In that exploration, I've also tried out state machine testing, which can be used for testing user interfaces, APIs, or whatever stateful system. That seems like a relevant thing if you want to test a web application. But if you use state machine property-based tests as they come in some forms, you might end up doing quite a lot of work and boilerplate just to get something up and running. More crucially, I think, is that these usually build on writing a model of your system—like a full, complete functional model of how your system is supposed to work. That means your model inherits all the essential complexity of the system under test, which might be fine, but might also be not so good.
+
+Imagine, for instance, that you're testing a key-value store. The essential complexity of that is pretty low—a key-value store conceptually is pretty simple, depending on its features. But its implementation might be much more complicated if it's a distributed system and you have replication and sharding and whatnot. The test, however, can still be pretty simple. But if you're testing something that is very complicated—like, say, a tax calculation system for Germany or something—then you might have tests or a model that is very complicated if it's supposed to test that. So it depends a bit on your system.
+
+Further, I've been digging into formal methods as a rookie and looking at tools like F*, TLA+, temporal logics, and so on. These all sort of fused together in my mind earlier this year and led me to try to mash them up into something, and this is what I came up with.
+
+I started thinking about the specification language, which would be a combination of linear temporal logic and some sort of functional, expression-oriented language. You can use it to describe or express how these web applications—or systems with web interfaces—change state over time. The state would be something you get from querying the DOM. That was my first idea, and this is where the language of TLA+ started influencing me a bit.
+
+The neat thing about the DOM and browser testing is that you can interact with the DOM and introspect it. So the tool itself could explore your application for you by finding out which actions were possible. This was the idea, at least. I find that this shaves off a large part of the maintenance aspect of these kinds of tests or specifications, because you don't have to couple your tests or specifications to the implementation as much, since you don't have to describe all the actions explicitly. This is sort of piggybacking on the idea that if a user can interact with the system through the browser, then a machine should be able to do the same. In the end, this is meant to run as property-based tests; basically, you could think of this as a special case of property-based testing.
+
+Some high-level goals of this project are that you, as a tester, should focus mostly on specifying and understanding these systems at a high level through these specifications. You shouldn't have to deal with a lot of low-level or complex details. That's a sort of fluffy goal, but something I want to have in mind. One concrete instance of that is that you shouldn't have to deal with timing. You shouldn't have to insert stuff like "sleeping" before you take an action just because you know that "one second after whatever, it's safe to do whatever." So, no sleeps or waits are supported in these specifications, and the tool is meant to deal with that for you.
+
+You're supposed to declaratively express the correctness of the system, and you should also be able to write partial specifications—or, as someone suggested to me yesterday, "smoke specs," which is a nice term. Basically, you should be able to state very loose requirements about your webpages and not have to specify all the details or complexity of them. One nice, reusable specification like that might be regarding modal dialogs—the ones that pop up in front of something else. You might specify that at any point in time, there should be zero or one modal dialog visible, and there can't be multiple. This is more like a UX guideline or a general constraint; it doesn't say anything about the broader functionality of the system, but it is still a specification.
+
+On the non-goal side, one thing I'm not trying to do here is to automatically reinforce or perform deterministic testing. This is rather tricky when interacting with the browser or a system. It might be deterministic, it might not be, and that is very much up to each system under test. This means that shrinking is not always useful, because shrinking sort of breaks down when it's non-deterministic.
+
+Also, I intend not to support any specific web frameworks or technologies. Not only do I not want to maintain that code, but I think the web application should be explicit enough for you to specify, "Okay, we're in a successful state, we're in an error state, we're currently loading data," and so on. If it's not possible to do those sorts of queries and see those states, then it's probably quite hard to use that web application as a user, too. It goes hand in hand: if it's usable and clear what's going on, you should be able to test it. So when people have asked, "Can we hook into the React virtual DOM or this or that framework?" it is explicitly not supported.
+
+Before we go on, I'm just going to quickly show how it works at a high level so you get a sense of what's going on. It begins by navigating to the origin page, which is where it starts. It can navigate to other pages by clicking links and so on, but this is the starting point. Then it begins the main part, which is recording a trace. A trace is a sequence or a list of states and actions—states that have been seen, and actions that have been taken. When it records this trace, it begins by generating some random actions, picks one that is possible to perform at this particular time, performs it, and records it. It then records the state that it sees afterward. It loops until it hits some bound or when it can't make any more progress.
+
+Then it checks the behavior. If you just pick out the states from the trace, you have the behavior. It checks if that behavior satisfies the specification you expressed with linear temporal logic. If not, it might try to shrink it (if you have shrinking enabled), and then it can possibly find a smaller sequence of actions that triggers an error. This is very much like property-based testing.
+
+All right. I'll go into the case study now. I did a case study on TodoMVC. TodoMVC is a—I don't know, ten-year-old or so—project which was basically a showcase for various web frameworks for building single-page apps. They built the same app, a to-do list, multiple times in various combinations of technologies. So I used all these examples as a benchmark for Quickstrom itself.
+
+In the beginning, I wrote a single specification for it; it basically has a new and an old format for TodoMVC. I checked the mainstream implementations—the first bunch of them on the website (todomvc.com)—and found some issues in two implementations: the Angular and Mithril implementations. I submitted an issue to the project, and this was very motivating.
+
+I continued on this track and eventually improved the specification to cover more features of TodoMVC. This was an iterative process because there is no really rigid, complete specification for TodoMVC. There is a plain-text document describing the features you should implement, but it's rather loose on some features. But I checked all the implementations and inferred one specification that worked for most of them. The results were 37 applications passing and 12 failing with various problems. Some you might consider just a missing feature or a technicality that made it hard to test, but some had what you might actually decide is a bug, and four of them were not testable at all (maybe taken down or something). There's a blog post on the details of this case study on my website. There's a link in the slides—I'll share the slides later if you want to follow these links, or you can go to my webpage.
+
+Let's go into a bit more detail, starting with the specification language. I'll show you the current version, which is based on PureScript. PureScript is a statically typed, functional, and pure language that originally started by targeting JavaScript—compiling to JavaScript to run in Node and the browser. Since then, it's gained a few other backends, like Erlang and C++, but this use case is a little different.
+
+First off, I've extended PureScript with new operators for doing the linear temporal logic stuff and querying the DOM. These are special operators that are now available in this language. There is also one thing I haven't noted here, which drastically changes the semantics of the evaluation: PureScript is normally strict, but now it's lazy. This is an experimental thing, but it enables some fun things to be done with control flow. It makes it possible, with some effort, to reuse regular community PureScript packages. This is made possible by an interpreter that I built in Haskell inside Quickstrom, which evaluates PureScript in this special way required for the temporal logic operators.
+
+Those operators essentially change the modality of whatever sub-expression they accept. There are three operators currently: `next`, `always`, and `until`. If you look at `next`, for instance, it takes an expression of some type `A` and returns an `A`. This essentially evaluates the expression in the next state, so you can think of it as changing the point in time in which it's evaluated. Then there are `always` and `until`. `always` takes an expression (like a boolean) and returns true if that is true across all subsequent states. `until` says that the first expression must be true until the second one becomes true.
+
+For querying the DOM, there are two operators: `queryOne` and `queryAll`. These are pretty much the same as `querySelector` and `querySelectorAll` in the browser. They take as arguments a CSS selector, just like in the browser, and they also take a record of element state specifiers. An element in the browser has a lot of attributes and properties, and in this case, you have to explicitly say, "I want to access these particular attributes, CSS styles, or whatever you want to use in your spec."
+
+For some examples (because it's a bit abstract), here is `queryAll` passed the selector `'button'`, so I get all the buttons. I then pass the specifiers `textContent` and `disabled`. Just below, you'll see the type of this expression: I get back an array of records, where each record has a `textContent` string and a `disabled` boolean. In the second example, I use `queryOne` and get back a `Maybe` record instead. This is how you interact with the webpage you're testing, pulling out state from certain parts of the page.
+
+In your spec, you also need to specify actions. Quickstrom currently doesn't just do purely random actions; you can constrain the list of actions it takes. You pass in a list of tuples where the first element is an integer (acting as a relative probability) and the second is the action specifier itself—so, an array of tuples. There are a bunch of predefined ones: you can say "just click everything" or "try to focus anything." You can also tweak these, and in practice, you might have to, because you have a limited amount of time for running these tests and you want them to be effective and have decent coverage. You don't want it to spin around doing useless stuff. You might also want to constrain it and say you just want to click around and interact with a particular part of the web application, not everything.
+
+The specification always looks kind of like this: you have a PureScript module called `Spec`, you import Quickstrom, and then you define `readyWhen`—which says, "when this selector matches an element, we can start testing." You pass it some actions; here I'm just using clicks. Then for the important part—the `proposition`, which is the safety property—you pass in the specification itself. If you're doing a state machine kind of spec, you usually have an initial state predicate, and then you say "always it's going to be either this transition, or that one, and so on." This essentially expresses a state machine.
+
+The transitions usually look sort of like this (simplified): you say that `something` should equal `foo` in this state, and in the next state, `something` should equal `bar`. This expression might build on a DOM query in the end, but the transition basically says we're transitioning from `something` being `foo` to `something` being `bar`.
+
+Because this is PureScript, you can reuse known packages like numbers, strings, arrays, and so on. This is useful because you can do regex matching and various operations. If you're really fancy, you can use monad transformers and go wild. These packages usually have some FFI (Foreign Function Interface) in them, which is normally JavaScript for PureScript. In this case, I've implemented the FFI in Haskell, and they're built into Quickstrom. It's a bit hard to add new packages, but it can be done.
+
+When you want to check a web app with Quickstrom, you use the `check` command. You pass a spec file and an origin URI to start at (which can also just be a local file). If a test fails (whether shrinking happens or not), you'll get output looking something like this: you'll see the trace with the states, the queried elements, what values they had at each state, and the actions in between. It's a bit rough, but this is what we have for now. Here, you might see, "Okay, there's a Not-a-Number (NaN) printed at the end there, so that's probably not good."
+
+You can use Quickstrom on multiple browsers, currently Firefox or Chromium. It uses WebDriver, so it should be possible to support any WebDriver server.
+
+Finally, I want to talk a bit about a concept I call *trailing state changes*, which is a complexity that arose from this kind of testing. You can't really know when state changes will occur. Generically, Quickstrom can't know how many times or when state changes appear. By default, it just records a single state after each action it takes. But you can override this to tell it to record a number of trailing state changes. Say you have an action, and then a state, and you might have more state changes subsequent to that first state. This is often the case when you have asynchronous operations like a fetch or an XHR request.
+
+Let's say you have an application where you click a button—"launch the missiles"—and you see a little spinner saying "missiles on their way." (By the way, it feels weird to talk about launching missiles these days, but anyway.) Then, in the end, you have the last state change which says "something failed; they blew up in the atmosphere" or something. You just had one action—clicking the button—but you had multiple state changes afterward: one immediate and one delayed. You can instruct Quickstrom to await and observe those using command-line options. It uses some fancy tricks to observe the DOM and react immediately so it doesn't have to be super slow.
+
+For the future, I probably want to improve the specification language. It's using PureScript now, but it's not the greatest fit in the end. I think it was a very good start because I could get a capable language early on, but I'll probably switch to some sort of bespoke language. I'm working together with Liam O'Connor right now on seeing what that might be.
+
+I also want to work on better reporting of errors and failures. The list you saw before with the trace is pretty basic and hard to go through. It would be nice to have nicer HTML reports or something similar. I also want to work on coverage, so you can be sure that your specification is exhaustive in some sense, and that you're not just hitting one case and missing another. Related to that, it would be nice to combine this with some sort of targeted search. Right now, it's just randomly picking actions, but if you could have a targeted selection of actions to increase coverage or some other metric, that would be cool.
+
+Stuff like screenshotting could be added. I might want to work on a commercial product for it as well: keep the CLI version open-source and free, and then build a service on top of it where you can do everything in the browser, schedule things, integrate with GitHub and CI, and so on.
+
+If you're interested in learning more about Quickstrom, you can ask me wherever you find me, check out the website, or sign up for the newsletter I post occasionally. With that, I want to say thank you for listening, and I'll take questions.
+
+---
+
+### Q&A Session
+
+**[Moderator:]** Thank you, Oskar. We have three questions posted on Whova. Question number one: Which products are your nearest competitors, and where is Quickstrom better? (Enjoying the talk, by the way, and I like the concept.)
+
+**[Oskar:]** This is where it'll shine through how much I'm failing in my market research, looking at this half academically! There is one AI testing product—I can't remember the name—that does some machine learning for browser testing. That might be a competitor, even if it's not exactly the same concept. Otherwise, the more established scenario-based browser testing tools—where you build the scenarios yourself or record them just by using a web browser—are much more polished and productized. I guess they are very worthy competitors in that sense, but they aren't doing the same thing. I don't know of anything doing exactly the same thing or very close to this.
+
+**[Moderator:]** Okay, one more question from myself. I guess this builds on Selenium or similar. How do you hide waiting for a DOM object to become available, or do you express it in some other way?
+
+**[Oskar:]** Yes, this is built on top of WebDriver, which is a lower-level protocol. I think they extracted it from Selenium at some point, so it's more like an open standard or de facto protocol, but basically Selenium, yes. Quickstrom does a static analysis pass on your specification to see what selectors, attributes, and properties you query for. It then observes all of those in the DOM when it runs. It reacts to state changes in your application instead of doing polling. You don't have to do any waiting; it reacts to the state changes and records them.
+
+**[Moderator:]** Does that answer your question? Well, you will need to visit the app later and post the shorter answer today. There's one more question, and another after this if you have time. The question is: "I wonder in which cases it might be better to turn shrinking off? I thought of it as an essential part of property-based testing. Are there cases where shrinking is not possible, and can shrinking hide bugs one would have found without shrinking?"
+
+**[Oskar:]** The thing with shrinking is that you essentially try to rerun the test with a slightly smaller input, right? The input in this case is a list of actions. But if the system you're testing is non-deterministic, rerunning it even with the exact same input might not produce the same results. You might have one bug exposed in your original list of actions, and then you shrink it, and it might disappear, or you might get another bug. It starts getting confusing. You might not necessarily hide bugs with it, but troubleshooting and rerunning tests will be a lot more confusing and impractical.
+
+I know, for instance, that the testing tool Jepsen doesn't do shrinking for this exact reason. They assume the distributed systems it tests are non-deterministic and dependent on the timing of operations and the network. They've decided not to do shrinking in that context. With Quickstrom, you can have it as an optional thing. Right now, it's enabled by default, but you can set the number of shrinks to zero if you want. It might eventually become an opt-in sort of thing.
+
+**[Moderator:]** All right. I think we have a minute for another question: "Might it make sense to visualize the state change graph?"
+
+**[Oskar:]** Ah, yes! That would be a lovely thing to explore. Right now, there's basically nothing else other than this ugly list that I print in the terminal. It would be nice to build something like a state machine representation, but you would need to identify multiple states as a logical state with different values. That might be tricky to do automatically; I haven't really tried doing that or explored it yet. But it would be very interesting to get a state machine diagram out of these tests and be able to troubleshoot using that representation. Otherwise, it's mostly screenshotting and stuff like that I've been thinking about. But if a state change graph could be inferred from the run, that would be very cool.
diff --git a/talk/repo.svg b/talk/repo.svg
new file mode 100644
index 0000000..50a935b
--- /dev/null
+++ b/talk/repo.svg
@@ -0,0 +1,24 @@
+
diff --git a/talk/testpass.webp b/talk/testpass.webp
new file mode 100644
index 0000000..20458ba
Binary files /dev/null and b/talk/testpass.webp differ