15 / Research / Aug 2026 - Present
Reducing Wasteful CI/CD Testing with LLM-Based Regression Test Selection
Honours research into whether an LLM can read a commit diff and run a smaller set of tests without missing the ones that reveal regressions.
- Languages evaluated
- 2
- Safety and efficiency measures
- 6
- Week research plan
- 32
Overview
Large software projects often run thousands of tests after every commit. That builds confidence, but it wastes time, machines and money when most of those tests have nothing to do with the change. Regression test selection tries to run only the tests a change could affect, and this project asks whether a large language model can make that call.
The central question is whether an LLM-based selector can reduce wasteful test execution in a CI/CD pipeline without missing the tests that reveal regressions. Safety comes first: a faster pipeline is worthless if it skips the one test that would have caught the bug.
Approach
The selector reads a commit diff, splitting large changes into smaller pieces, and scores every existing test for relevance with a short reason. In two stages, it first ranks change-relevant tests, then a test-count or time budget turns that ranking into the smaller subset that actually runs.
It will be compared at equal budgets against random selection, BM25 text matching and embedding similarity, with an optional hybrid that adds coverage information, so the LLM has to beat cheaper baselines rather than nothing.
Evaluation
The evaluation uses reproducible Java and Python CI failures from BugSwarm, with the full test suite's results as the answer key. It measures failing-test recall, missed failures, commit detection, test reduction, execution-time savings and net pipeline savings once the LLM's own time and cost are counted. Results are reported across several budgets as a safety and efficiency curve, for each language separately and then combined.