5 min read

git bisect run: find the regression while you get coffee

A practical guide to letting git binary-search 400 commits for you, with a test script, exit codes that matter, and the traps in front-end repos.

Last month our checkout form stopped showing validation messages in Safari. Nobody knew when it started. The last release we were sure about was four weeks and 412 commits old. Reading through 412 diffs is a day of work. git bisect run found the culprit in nine steps while I made coffee. Here’s how to do that, and the things that make it fail in front-end projects.

Bisect in thirty seconds

git bisect does a binary search over your history. You tell it one commit where the bug exists and one where it doesn’t, and it checks out the commit in the middle. You test, you tell it “good” or “bad”, it halves the range again. For 412 commits, that’s at most ⌈log₂ 412⌉ = 9 rounds.

The manual version:

git bisect start
git bisect bad                 # current HEAD is broken
git bisect good v3.14.0        # this tag was fine
# git checks out a commit in the middle; test it, then:
git bisect good                # or: git bisect bad
# ...repeat until git prints "<sha> is the first bad commit"
git bisect reset               # back to where you started

That’s already much better than reading diffs. But you still have to sit there and run the test nine times. That’s where run comes in.

Letting git drive

git bisect run <cmd> runs a command on each candidate commit and uses its exit code as the verdict:

Exit code Meaning
0 good
1–127, except 125 bad
125 skip: this commit can’t be tested
128 or higher abort the whole bisect

So any test that exits non-zero on failure is already a bisect script. For a unit test:

git bisect start HEAD v3.14.0
git bisect run npx vitest run src/forms/validation.test.ts

git bisect start <bad> <good> saves the separate bad/good steps. Then you walk away.

Writing a test that didn’t exist yet

The interesting case is when the bug has no test, which is always, otherwise CI would have caught it. You need a small script that reproduces the bug and lives outside the repo, because bisect is going to check out old commits where your new test file doesn’t exist.

For the Safari bug, I wrote a Playwright script in /tmp:

// /tmp/bisect-validation.mjs
import { webkit } from 'playwright';

const browser = await webkit.launch();
const page = await browser.newPage();
await page.goto('http://localhost:4173/checkout');
await page.click('button[type=submit]');
const visible = await page.isVisible('#email-error');
await browser.close();
process.exit(visible ? 0 : 1);

And a wrapper that builds the app for each commit:

#!/usr/bin/env bash
# /tmp/bisect.sh
set -u

npm ci --silent || exit 125           # can't install: skip, don't blame
npm run build --silent || exit 125    # can't build: skip

npx vite preview --port 4173 --strictPort >/dev/null 2>&1 &
server=$!
trap 'kill $server 2>/dev/null' EXIT

# wait up to 20s for the server
for i in $(seq 40); do
  curl -sf http://localhost:4173/ >/dev/null && break
  sleep 0.5
done

node /tmp/bisect-validation.mjs

Then:

chmod +x /tmp/bisect.sh
git bisect start HEAD v3.14.0
git bisect run /tmp/bisect.sh

Exit code 125 is the important one

Notice that install and build failures exit with 125, not 1. This is the single most common way to get a wrong answer from bisect run.

Over 400 commits, some won’t build: a half-finished refactor, a dependency bump that was fixed two commits later. If your script reports those as “bad”, bisect narrows in on the broken build instead of your actual bug, and confidently names the wrong commit. With 125, git skips that commit and picks a neighbour.

The opposite mistake is just as bad: set -e at the top of the script. Then any failing command, including your curl readiness loop, exits with 1 and counts as “bad”. I use set -u for typos and handle every exit explicitly.

Front-end traps

Dependencies change between commits. That’s why the script runs npm ci, not npm install. ci installs exactly what the lockfile at that commit says. It’s slower, but the result is deterministic. If your lockfile didn’t change between two candidates, npm’s cache makes the second install fast.

Node versions change too. If the repo has an .nvmrc, and it changed within your range, add nvm use >/dev/null (or fnm use) at the top of the script. Otherwise an old commit builds with today’s Node and fails for reasons that have nothing to do with your bug. Exit 125 protects you from a wrong verdict, but too many skips and bisect gives up with a range instead of a commit.

Ports stay busy. If the preview server from the previous step didn’t die, the next one can’t bind. --strictPort makes it fail loudly instead of silently picking port 4174, and the trap kills it on every exit path.

Flaky tests lie. A test that fails one time in ten will send bisect down the wrong half about one time in ten. Run the check twice and only report “bad” if both fail, or fix the flake first.

Reading the result

When it finishes, git prints the first bad commit and its diff stat. In our case it was a harmless-looking change to a form component: someone replaced <span role="alert"> with a custom element that Safari didn’t upgrade before the validation event fired. Four lines. Nine rounds, about twelve minutes of machine time.

Before you git bisect reset, run git bisect log > /tmp/bisect.log. It records every verdict, which is useful if you want to show a colleague how you got there, or replay the session with git bisect replay after fixing a mistake in your script.

A cheat sheet

git bisect start <bad> <good>
git bisect run ./script.sh       # 0 good, 1-127 bad, 125 skip
git bisect log > bisect.log      # keep the evidence
git bisect reset                 # always, when done

Keep the script outside the repo, skip what can’t build, and be suspicious of flaky tests. Then go get that coffee.