
Line coverage certifies that a line executed; it says nothing about whether any test would notice if theline were wrong. Mutation testing answers the second question directly by modifying the programunder test and asking whether the suite complains. Its adoption has been limited by a straightforwardcost — one full suite run per mutant — which for a project of moderate size turns a nightly rebuild intoa multi-hour job. We present dokimos, a mutation-testing tool for Python that combines three costreductions that individually are known but rarely used together in the Python ecosystem: (i) per-testcoverage selection, running each mutant only against the tests that execute its line; (ii) diff scoping,mutating only the lines changed on a pull request; and (iii) verdict caching keyed on the tuple thatdetermines the outcome. Mutations are applied through an import hook rather than by rewriting files ondisk, so an interrupted run leaves the working tree byte-for-byte unchanged. On a representativefixture, dokimos executes 126 mutants in 14.0 s across eight workers and in 0.9 s with a warm cache,versus 397.6 s of a comparable serial baseline. The tool is open source, ships as a GitHub Action, and isapplied to itself: on its own source it reports a mutation score of 64% [55%, 72%] over a 120-mutantsample. We discuss prior work on selective mutation, why the classical cost has kept the technique outof Python continuous integration, and the failure modes that survive the reductions.
