What 30 days of CI in 79 open-source Playwright projects says about red builds
By TurnSignal · September 28, 2026 · 7 min read
In short
- We analysed 30 days (29 August to 28 September 2026) of public GitHub Actions data from 79 open-source projects that run Playwright: 11,618 completed workflow runs and 1,754 job logs.
- 42% of the runs failed, and 49 of the 79 projects had failed runs on their default branch. Red is the normal state of a busy CI, not an exception.
- 903 runs (7.8%, about 1 in 13) were re-run, adding 1,102 extra attempts and roughly 286 hours of CI time whose result was thrown away.
- In 30 of the 59 projects whose logs showed Playwright results, at least one test failed or flaked on three or more different branches: failures that the pull request running them almost certainly did not cause.
What we looked at
We started from public repositories whose GitHub Actions workflows run Playwright, kept those owned by organizations (not personal accounts), written in TypeScript or JavaScript, with recent activity, at least 8 completed runs of the Playwright workflow in the last 30 days and at least 10 end-to-end test files. From the ranked list we took the top 100 and analysed every one; 79 had enough runs in the window to count.
For each project we read the workflow's run list, the attempts of every re-run, and the job logs of up to 20 failed and 5 passing runs. Test names come from Playwright's own end-of-run summary, the block that lists failed and flaky tests after the last test finishes. We did not run any code, open any issue, or contact anyone to produce these numbers, and we name no project below: the point is the pattern, not the teams.
Finding 1: red builds are normal
Across the 79 projects, 42% of completed workflow runs failed (the median project: 43%). That counts the whole workflow, so a lint error or a broken build step counts too, not only Playwright. Still, the number matters because it is what every contributor sees: nearly every other push comes back red.
49 of the 79 projects had at least one failed run on their default branch in the 30 days. While main is red, every open pull request that branches from it inherits the failure, and reviewers have to work out which red is theirs. In 29 of the 59 projects where we could read Playwright results, at least one of the failing tests also failed on the default branch.
Finding 2: re-running is the default fix
903 of 11,618 runs were re-run at least once: 7.8%, about one run in thirteen. Those re-runs added 1,102 extra attempts. Measured as wall-clock time from the start to the end of each earlier attempt (capped at six hours per attempt), that is roughly 286 hours of CI whose result was discarded in favour of the next attempt.
- 65 of the 79 projects re-ran at least one run in the month.
- 27 projects re-ran 10% or more of their runs.
- The median project re-ran 5.8% of its runs, about 0.8 hours of discarded attempts a month; the two busiest lost about 92 and 34 hours.
A re-run is sometimes the right call. The problem is what it erases: on GitHub Actions a re-run keeps the same run and commit, and once it goes green nobody looks at the attempt that failed. The test that forced the re-run stays in the suite, unnamed, and does it again next week.
Finding 3: failing tests travel across branches
A pull request can only break code it touches. So when the same test fails on three or more different branches in a month, the likely causes are outside those pull requests: the test is flaky, it depends on an environment or service that misbehaves, or it was already broken on main. In 30 of the 59 projects with readable Playwright results, at least one test did exactly that.
Across those projects we counted 360 tests that failed or flaked on three or more branches, out of 2,267 distinct failing tests (16%). One project alone contributed 164, most likely an environment problem hitting many tests at once; without it the count is still 196. The median project had 11 distinct failing tests in the logs we read, so a handful of cross-branch tests is a large share of what reviewers deal with.
This is the core cost of a noisy suite. Every contributor whose pull request hits one of those tests has to decide whether it is their fault, usually by re-running (Finding 2) or by asking in chat. The information that would settle it in seconds (this test also failed on four other branches this week) exists in the CI history, but no single report shows it.
Finding 4: retries keep flaky tests green, and alive
With retries enabled, Playwright reruns a failed test inside the same job and reports it as flaky if a later attempt passes. The run stays green. In 29 projects we found tests that only ever appeared as flaky in the logs we read, 298 tests in total (13% of all failing tests). None of them failed a build; all of them cost a retry, and a retry of a slow end-to-end test can take minutes.
Retries are a sensible default for CI, but a flaky test that never turns a build red is a test nobody is asked to fix. The flaky list in each Playwright summary is the cheapest backlog of test fixes a team will ever get, and it scrolls away with the log.
What to do with this in your own suite
- Check main before re-running. If the failing test also fails on the latest default-branch run, re-running your pull request cannot fix it. Find or open the ticket instead.
- Look at a test's history across branches, not only yours. A test that failed on several unrelated branches this week is almost certainly not your change.
- Keep retries, but keep the flaky list. Copy the flaky tests from each Playwright summary somewhere that survives the log, and fix the ones that show up every week first.
- Name the re-run. When a re-run passes, write down which test failed first. After a month the list tells you exactly where the discarded CI hours went.
- Make flaky visible once the suite is stable.
--fail-on-flaky-tests(orfailOnFlakyTests: truein the config) fails a run that had flaky tests, so they cannot hide behind retries.
Limits of this data
- Biased sample. Projects were picked for high failure and re-run rates; the average Playwright project is likely calmer.
- Whole workflow. A failed run may have failed on lint, build or infrastructure before any test ran.
- Lower bounds. We read at most 25 runs per project and up to 6 failed jobs per run, so test counts are lower bounds. 20 of the 79 projects printed no Playwright summary we could read (for example, only the blob reporter runs in CI), so they add to run and re-run numbers but not to test numbers.
- Wall-clock, not billed minutes. Re-run hours measure elapsed time of earlier attempts, not runner minutes (parallel jobs would make billed minutes higher).
Get the same numbers for your repository
The analysis is repeatable for any public repository with a Playwright workflow. If you would like a report for yours (tests ranked by failed and flaky runs, branches hit, and re-run cost over 30 days), email hello@turnsignal.ai with the repository name and we will send it, free.
We did this because it is the problem TurnSignal exists for: on every pull request it labels each failed Playwright test as New failure, Also failing on main, or Known flaky, keeps the flaky history that retries hide, and lists tests that never reported because a shard crashed. It never changes your CI's exit code. Setup is in the docs.
Questions people ask
How often do Playwright CI runs fail in open-source projects?
In our sample of 79 busy open-source projects, 42% of completed workflow runs failed over 30 days, and 49 projects had failed runs on their default branch. The sample was chosen for high failure rates, so typical projects are likely lower.
How much CI time do re-runs waste?
In our sample 7.8% of runs were re-run, adding 1,102 extra attempts and about 286 hours of wall-clock CI time whose result was discarded. The median project lost about 0.8 hours a month; the two busiest lost about 92 and 34 hours.
How can I tell a flaky Playwright failure from a real one?
Check whether the test also fails on main and whether it failed on other branches recently. A test failing on several unrelated branches is almost never caused by the pull request that ran it.
Sources
- GitHub REST API: Workflow runs (including run attempts)
- GitHub REST API: Workflow jobs (job logs)
- GitHub docs: Re-running workflows and jobs
- Playwright docs: Retries (flaky tests)
- Playwright docs: Reporters
- Playwright docs: Command line (--fail-on-flaky-tests)
Playwright details were checked against Playwright 1.63 and its documentation on September 28, 2026.
Try TurnSignal on your next run
Add one reporter line next to your existing reporters. Free to start, no credit card, no repository access.