← Back to blog
·5 min read

Test impact analysis without build-system lock-in: selecting tests from a code graph

Which tests does a diff reach? Blastline answers from a code graph, not a build system: a safe superset of tests, or the full suite with a stated reason.

citest-impact-analysiscode-intelligence

A single diff node on the left sends traces through a grid of test tiles; only the seven tiles it actually reaches light up, the rest stay dark.

Picking which tests a pull request needs to run is a graph-reachability question — changed code, its transitive dependents, and the test files among them — so it can be answered from a code graph instead of a build system or a coverage run. Blastline, our open-source test-impact tool, does exactly that on top of CGraph: it returns a safe superset of tests to run, and when the graph can't vouch for a diff it says so and runs everything.

The problem

Test-impact selection has been stuck in the same three places for years. It is locked to one build system (Bazel's bazel-diff, the affected commands in Nx and Turborepo), to one language through coverage instrumentation (pytest-testmon, Skippy), or to a closed SaaS product. If your repository is a mix of languages that doesn't live inside Bazel, the practical options are "run the whole suite on every PR" or "hand-maintain a path map that silently rots."

Both are costs a buyer feels. The first is CI minutes and slow feedback. The second is worse: a selector that skips a test it should have run gives you a green check that means nothing.

What we did

We treated selection as a query against a graph we already build. CGraph extracts a deterministic graph of symbols, calls, and imports across a dozen-plus languages; Blastline is the judgment layer on top of its impact query.

blastline tests main..HEAD | xargs vitest run   # tests the diff reaches
blastline blast main..HEAD                       # every transitive dependent, with file:line

The decision that shaped everything was the contract: "run at least these," never "safe to skip." Static graphs miss dynamic dispatch, so the tool is built to fail open. A diff falls back to the full suite, with a machine-readable reason, when a changed file has no node in the graph (configs, lockfiles), when the graph is older than the head commit, when the diff is oversized, or when a sparse-graph guard trips — a graph averaging fewer than 3 edges per file produces subsets that look impressively small and are actually blind, so Blastline refuses to produce one at all.

Diagram: a diff maps to changed graph nodes, then transitive dependents, then the test files among them, producing a run-at-least-these subset. If any guard trips (a file with no graph node, a graph older than HEAD, an oversized diff, a sparse graph, or tests that cannot reach the changed code), it falls back to the full suite with a machine-readable reason.

The alternative we rejected was the one that demos best: return the smallest plausible subset and let coverage gaps sort themselves out. That wins the screenshot and loses the trust of every engineer the first time a regression ships behind a green build.

The result

Before release we replayed 80 historical commits across two repositories, scoring each selection against the tests the commit's author actually changed alongside the code. On a production Next.js app, 19 of 20 commits got a subset averaging about a quarter of the suite, with every genuinely related test selected.

The more useful result came from es-toolkit (1,508 files): the replay caught our own graph engine silently dropping 650 of those files. An import written with its source extension (from './chunkBy.ts') produced a stub whose id collided with the real file's node, and the real file — with its functions and edges — was deleted. After the fix (CGraph #42), the identical replay moved like this:

Bar chart of the es-toolkit replay over 20 commits: before the engine fix, 6 of 20 commits got a safe test subset and 1 of 8 co-changed tests was selected; after it, 18 of 20 commits got a subset and 41 of 41 co-changed tests were selected.

es-toolkit replay (20 commits) Commits given a subset Co-changed tests selected Mean share of suite run
Before the engine fix 6 / 20 1 / 8 —
After the engine fix 18 / 20 41 / 41 4.7%

Python and Go followed the same loop: the replay exposed blind graphs, a second guard (disconnected-tests) now refuses to select when tests can't reach the code, and the extraction gaps were fixed upstream. The full methodology and tables are in the Blastline repository.

We also run it on this site. It started as an advisory comment on pull requests that touch the contract between the site and our publishing agent, and the first two real PRs it commented on both failed open for the same reason: they edited .env.example, which has no node in the graph, so the tool said "run the full suite" rather than guess. That is the contract working as designed, and it is also the honest ergonomic cost — config files are where a code graph stops knowing things, and a selector should tell you so out loud. Since then every pull request here builds a graph of src/, lets Blastline pick the tests, and runs that subset with bun test, or the whole suite when it fails open.

Takeaway

Test-impact selection doesn't have to come bundled with a build system: if you have a trustworthy code graph, it is a reachability query plus a strict fail-open contract. If slow or untrustworthy CI is taxing your team, that is the kind of developer platform work we take on — and the tools above are MIT-licensed if you'd rather start on your own.

Building something like this?