Test impact analysis without build-system lock-in: selecting tests from a code graph
Which tests does a diff reach? Blastline answers from a code graph, not a build system: a safe superset of tests, or the full suite with a stated reason.

Picking which tests a pull request needs to run is a graph-reachability question — changed code, its transitive dependents, and the test files among them — so it can be answered from a code graph instead of a build system or a coverage run. Blastline, our open-source test-impact tool, does exactly that on top of CGraph: it returns a safe superset of tests to run, and when the graph can't vouch for a diff it says so and runs everything.
The problem
Test-impact selection has been stuck in the same three places for years. It is
locked to one build system (Bazel's bazel-diff, the affected commands in Nx
and Turborepo), to one language through coverage instrumentation
(pytest-testmon, Skippy), or to a closed SaaS product. If your repository is a
mix of languages that doesn't live inside Bazel, the practical options are "run
the whole suite on every PR" or "hand-maintain a path map that silently rots."
Both are costs a buyer feels. The first is CI minutes and slow feedback. The second is worse: a selector that skips a test it should have run gives you a green check that means nothing.
What we did
We treated selection as a query against a graph we already build. CGraph
extracts a deterministic graph of symbols, calls, and imports across a dozen-plus
languages; Blastline is the judgment layer on top of its impact query.
blastline tests main..HEAD | xargs vitest run # tests the diff reaches
blastline blast main..HEAD # every transitive dependent, with file:line
The decision that shaped everything was the contract: "run at least these," never "safe to skip." Static graphs miss dynamic dispatch, so the tool is built to fail open. A diff falls back to the full suite, with a machine-readable reason, when a changed file has no node in the graph (configs, lockfiles), when the graph is older than the head commit, when the diff is oversized, or when a sparse-graph guard trips — a graph averaging fewer than 3 edges per file produces subsets that look impressively small and are actually blind, so Blastline refuses to produce one at all.
The alternative we rejected was the one that demos best: return the smallest plausible subset and let coverage gaps sort themselves out. That wins the screenshot and loses the trust of every engineer the first time a regression ships behind a green build.
The result
Before release we replayed 80 historical commits across two repositories, scoring each selection against the tests the commit's author actually changed alongside the code. On a production Next.js app, 19 of 20 commits got a subset averaging about a quarter of the suite, with every genuinely related test selected.
The more useful result came from es-toolkit
(1,508 files): the replay caught our own graph engine silently dropping 650 of
those files. An import written with its source extension (from './chunkBy.ts') produced a stub whose id collided with the real file's node,
and the real file — with its functions and edges — was deleted. After the fix
(CGraph #42), the identical
replay moved like this:
| es-toolkit replay (20 commits) | Commits given a subset | Co-changed tests selected | Mean share of suite run |
|---|---|---|---|
| Before the engine fix | 6 / 20 | 1 / 8 | — |
| After the engine fix | 18 / 20 | 41 / 41 | 4.7% |
Python and Go followed the same loop: the replay exposed blind graphs, a second
guard (disconnected-tests) now refuses to select when tests can't reach the
code, and the extraction gaps were fixed upstream. The full methodology and
tables are in the Blastline repository.
We also run it on this site. It started as an advisory comment on pull
requests that touch the contract between the site and our publishing agent,
and the first two real PRs it commented on both failed open for the same
reason: they edited .env.example, which has no node in the graph, so the
tool said "run the full suite" rather than guess. That is the contract working
as designed, and it is also the honest ergonomic cost — config files are where
a code graph stops knowing things, and a selector should tell you so out loud.
Since then every pull request here builds a graph of src/, lets Blastline
pick the tests, and runs that subset with bun test, or the whole suite when
it fails open.
Takeaway
Test-impact selection doesn't have to come bundled with a build system: if you have a trustworthy code graph, it is a reachability query plus a strict fail-open contract. If slow or untrustworthy CI is taxing your team, that is the kind of developer platform work we take on — and the tools above are MIT-licensed if you'd rather start on your own.