TurnSignal

Stop re-running failed CI jobs: is this Playwright failure flaky or real?

By TurnSignal · September 28, 2026 · 4 min read

In short

  • Re-running a failed job until it passes is the most common way teams deal with flaky tests, and it teaches everyone to ignore red builds.
  • Before re-running, answer four questions: did it pass on a retry in the same run, does it also fail on main, has this test been flaky before, and does the error touch what you changed?
  • Playwright answers the first question in every run when retries are on. The other three need the test's history across runs and branches.
  • Record what the re-run shows. A test that passes only after a re-run is flaky, even if the final build is green.

The re-run habit

A pull request goes red. The failing test looks unrelated, someone clicks Re-run failed jobs, it passes, and the pull request is merged. It feels efficient. Over a few months it has three costs: real regressions get merged because "it was probably flaky", CI minutes are spent running whole shards again, and nobody fixes the flaky tests because the re-run button makes them bearable.

On GitHub Actions a re-run keeps the same workflow run and commit and increases github.run_attempt. The earlier attempt is still visible in the run, but most people only look at the final result, so the information that a test failed and then passed is effectively lost.

Question 1: did it pass on a retry in the same run?

With retries enabled, Playwright already runs this experiment for you. A test that failed and then passed on a retry is reported as flaky, and one that failed on every attempt is failed. A failed test with two retries has failed three times in a row on the same machine, which is weak evidence for flakiness: re-running the job is unlikely to tell you anything new.

// playwright.config.ts
export default defineConfig({
  retries: process.env.CI ? 2 : 0,
  use: { trace: 'on-first-retry' },
});

Question 2: does it also fail on main?

If the same test fails on the latest run of your default branch, your change did not break it, and re-running your pull request will not fix it either. Check the most recent main run before anything else. If main is red for this test, the right move is to find (or open) the ticket for the broken test, not to re-run.

Question 3: has this test been flaky before?

A test that was flaky twice last week and fails again today is very likely flaky again. A test that has passed 200 times in a row and fails on your branch is very likely broken by your change. That difference is invisible in a single report; you need the test's recent results, ideally across all branches, not only yours.

Question 4: does the error touch what you changed?

Read the error and the trace before deciding. A timeout waiting for a button you renamed is not flaky. A network error from a third-party sandbox in a test for a page you didn't touch probably is. Open the trace (npx playwright show-trace) of the failed attempt: the moment the attempt went wrong usually says which.

If you do re-run, keep what it tells you

  • Name it. If the re-run passes, the test is flaky. Write that down in the pull request or the team channel, with the test name.
  • Count it. Keep a list of tests that needed a re-run. The ones that appear every week are where fixing pays off first.
  • Reproduce it. Locally, npx playwright test path/to/file.spec.ts --repeat-each=20 --retries=0 usually shows an intermittent failure within a few minutes.
  • Stop hiding it. Once the suite is stable, failOnFlakyTests: true (or --fail-on-flaky-tests) makes a flaky test fail the run so it cannot be ignored.

A team rule that works

Agree on one short rule and put it in the pull request template: "Re-run only after checking main and the test's history. If the re-run passes, add the test to the flaky list with a link to the run." It costs a minute per red build, and within a few weeks the list tells you exactly which tests deserve a proper fix, and which failures were real all along.

How TurnSignal answers the four questions

TurnSignal puts the answers next to each failure on the run page and in the pull request comment. Each failed test is labelled New failure (it passed in its latest result on the default branch: most likely your change), Also failing on main (not your change), or Known flaky (flaky in at least one of its previous 20 results). When a CI job is re-run, the new attempt updates the same TurnSignal run, and a test that failed first and passed on the re-run is counted as Flaky, so the information isn't lost when the build turns green.

The Flaky tests page shows which tests were flaky most often over the last 7 to 90 days, with the runs behind each count. The reporter never changes your exit code. The pull request comment works on GitHub only for now; see the docs for setup, or use the TurnSignal GitHub Action.

Questions people ask

Is it bad to re-run failed GitHub Actions jobs?

Not by itself, but re-running until green without recording that the test failed first hides flaky tests and sometimes real regressions. Check whether the test also fails on main and whether it was flaky before.

What is the difference between a Playwright retry and a CI re-run?

A retry happens inside the same Playwright run, and a test that passes on retry is reported as flaky. A CI re-run starts the job again; Playwright sees a brand-new run and does not know the test failed before.

How do I find my flakiest Playwright tests?

Track each test across runs: how often it was flaky or needed a re-run in the last weeks. The tests at the top of that list are the ones to fix first.

Sources

Playwright details were checked against Playwright 1.63 and its documentation on September 28, 2026.

Try TurnSignal on your next run

Add one reporter line next to your existing reporters. Free to start, no credit card, no repository access.

Get started free Read the docs