kitty's Parallel Test Runner Said OK While Its Workers Crashed
A missing addSubTest and an ignored exit status let kitty's new parallel test runner report success after failures. Fixing that exposed seven segfaulting workers on my Mac, and Kovid Goyal fixed those his own way.
In August 2026 I was going through kitty, the GPU terminal emulator, looking for things worth a pull request. I had Codex on GPT-5.6 doing a first audit pass over the repo. I picked which findings were worth a maintainer's time and reviewed every line before anything went upstream. Its list had a shell-quoting bug in choose-files at the top, which Kovid Goyal merged on the morning of 12 August. Second on the list was the test runner.
kitty had switched to running its Python tests in parallel two weeks earlier (39ae060d2, "Run python test in parallel", 29 July). This post is about a hole in that runner, what showed up once the hole was closed, and the part Kovid fixed himself. (An earlier exchange with Kovid, about hard links, is in the best way to learn a codebase.)
How the runner collects results
kitty_tests/main.py splits the test list round-robin into up to eight chunks and calls os.fork() once per chunk. Each child runs its tests with a unittest.TestResult subclass called PipeTestResult, which writes one JSON line per event into a pipe: start, ok, fail, error, skip and so on. The parent reads all the pipes, counts records, prints a progress line like Py 211/406, and at the end decides pass or fail from the records it received.
That last decision looked like this:
python_ok = not (py_failures or py_errors or py_unexpected_successes or py_worker_errors)So the parent only knew about failures that arrived as records. Anything that failed without writing a record counted as nothing.
A failing subTest left no record
PipeTestResult implemented addSuccess, addFailure, addError, addSkip and the expected-failure pair, and it had no addSubTest. When a test uses with self.subTest(...) and one of the subtests fails, unittest calls result.addSubTest() for that subtest. Because the test as a whole didn't succeed, it never calls addSuccess for it either. The base class's addSubTest appends to result.failures inside the child, but nothing goes down the pipe.
The parent ended up with a start record for the test and nothing else. A synthetic test with one failing subtest, run through the real collector, printed Ran 0 tests followed by OK, and collect_worker_results() returned success.
The child did know it had failed. run_test_worker exited with status 1 when result.failures was non-empty. The parent then threw that away:
for pid in pids:
os.waitpid(pid, 0)kitty's GLFW, drag-and-drop, font and shell-integration tests all use subTest, so a real assertion failure in any of them would have shown up as one fewer test in the count and a green result.
The first fix
I wanted each channel to mean one thing. The pipe carries test outcomes, including subtests. The exit status says whether the worker process itself survived.
def addSubTest(self, test, subtest, err) -> None:
super().addSubTest(test, subtest, err)
if err is not None:
record_type = 'fail' if issubclass(err[0], test.failureException) else 'error'
self._send({'t': record_type, 'id': str(subtest), 'e': self._elapsed(),
'msg': self._exc_info_to_string(err, test)})The worker now exits 0 after running its tests whether they passed or not, since the failures already went through the pipe, and the parent checks every status:
for pid in pids:
_, status = os.waitpid(pid, 0)
if status != 0:
py_worker_errors.append(
f'Python test worker {pid} exited with status {os.waitstatus_to_exitcode(status)}')Without that split, a failing test would be reported twice (once as a test failure, once as a worker failure). The commit was 63 added lines, most of them two regression tests: one for the subtest record, one where a forked child exits 3 and the collector has to say so.
Then the full suite went red
The focused tests passed. Then I ran the whole suite on my Mac and got this:
Py 211/406 ✓ Go 248 ✓ [37.6s]
======================================================================
WORKER ERROR
----------------------------------------------------------------------
Python test worker 54290 exited with status -11
...
Ran 459 tests in 37.564s
FAILED (worker errors=7)Seven workers had died with signal 11, and 195 Python tests never ran. The code I had changed only reports crashes, so these had been happening before my change too. Going by the python_ok line above, the unpatched runner would have printed Ran 459 tests and OK for this exact run, with the progress line stuck at 211 of 406.
Running the same suite sequentially, with no forking, passed: 406 Python tests and 248 Go tests.
macOS writes a crash report for every one of those children, and each had the same two annotations:
CoreFoundation: *** multi-threaded process forked ***
libsystem_c.dylib: crashed on child side of fork pre-execThe faulting stack was a SIGSEGV (KERN_INVALID_ADDRESS) inside a CoreFoundation preferences lookup, _CFPreferencesCopyValueWithContainer, called from os_log_type_enabled. In other words, a child process was asking the logging system whether a log level was enabled, and that lookup touched framework state that hadn't survived the fork.
Why fork is risky on macOS
POSIX says that when a multi-threaded process calls fork(), the child gets a copy of the whole address space but only the thread that called fork(). Any lock another thread held at that moment stays locked forever in the child, and any half-updated data structure stays half-updated. So until the child calls one of the exec functions, it is only supposed to call async-signal-safe functions. Running a Python interpreter is far outside that list.
Plenty of programs get away with it on Linux. macOS is less forgiving because its system frameworks start their own threads and keep a lot of process-wide state, and CoreFoundation records that a fork happened (that's where the first annotation in the crash reports came from). Python has documented this for years. Since 3.8, multiprocessing defaults to spawn on macOS, and the docs say fork "should be considered unsafe as it can lead to crashes of the subprocess as macOS system libraries may start threads". Since 3.12, os.fork() raises a DeprecationWarning when it can tell the process has multiple threads.
The runner did this right before forking:
# Pre-initialize fonts once before forking so all worker processes inherit
# the warm C-level fontconfig state and their own all_fonts_map() calls are fast.
from kitty.fonts.common import all_fonts_map
all_fonts_map(True)On Linux that warms fontconfig, which is what the comment is about. On macOS all_fonts_map goes through kitty's CoreText backend, which enumerates fonts through CoreText and CoreFoundation APIs. That was my suspect for how the parent picked up framework state and extra threads before the fork. I didn't prove it, and I should have said so more carefully than I did.
My second commit took the blunt route: on macOS, run the Python tests in-process, no forking. I opened kitty#10349, "Report parallel test failures", with both commits (two files, +76/−4).
What Kovid did with it
His first reply:
Why do you think the test launcher is initializing corefoundation? As far as I know it shouldnt be. And the tests are running fine in macOS in CI and on my mac. I have merged parts of your first commit as they made sense.
He had already committed the two halves of the reporting fix to master himself, as 485a12aa12 ("Report python subtest failures when running test suite") and f047fc82d5 ("Report python test worker failures separately from test failures"). Apart from my comment, those diffs are the same as the hunks in my commit. He didn't take the macOS change or my regression-test file.
The first draft of my reply conceded the whole diagnosis, which would have made it look like I'd invented a crash. I threw that out. The next version read like a cross-examination, and I threw that out too. What I posted said I'd based the suspicion on the pre-fork all_fonts_map(True) call going through CoreText, gave the machine (an arm64 M5 Pro MacBook Pro, macOS 26.5.1, kitty's bundled Python 3.14.6) and the crash annotations, and agreed that disabling parallel tests on every Mac was too broad if CI and his machine were fine.
He answered with a commit, 185de897b4, which on macOS replaces fork() with workers started through subprocess.Popen running kitty +runpy, each getting its test IDs as JSON on stdin and its pipe's file descriptor through pass_fds. The commit message gives the reason as avoiding "CoreFoundation libraries not liking fork + threads", and its Fixes #10349 closed my PR about 40 minutes after I opened it. Linux keeps forked workers.
I tested his follow-up cleanup, 00a5ca54ef, twice on the machine that had been crashing. Both runs passed all 395 Python and 247 Go tests, with no worker errors and no new crash reports.
Later that day he pushed dcbffe3fe0, which moves the all_fonts_map(True) call into each exec'd worker on macOS and keeps the pre-fork warm-up only on the Linux fork path. On macOS the parent process no longer touches the font code at all during a test run. The commit message calls it a cleanup, and I read it as one, but it does remove the call I'd pointed at from the macOS parent.
So the reporting fix came from my PR and he applied it, and the crash fix is his, done a better way than mine. Exec'd workers keep the parallelism and make the fork question irrelevant. My version gave up the parallelism on every Mac to avoid a crash I thought only my Mac had.
macOS CI had the same crash
I assumed my machine was the odd one out, and while writing this up I went back to kitty's CI logs to check. It wasn't. On the commit I had audited, both macOS jobs reported 173 of 394 Python tests and passed, while the Linux job in the same run got through all 394. Once the change that reports worker exits landed on master, the next macOS runs went red within minutes: three workers died with SIGSEGV in one job and three with SIGABRT in the other, and Linux stayed green. After the switch to exec'd workers, both macOS jobs ran all 395.
CI had looked healthy for the same reason my first run did, because the runner ignored how its workers exited. The CI runners were on macOS 26.5.2, which also rules out my point release as the explanation. What I still don't know is what made the parent process multi-threaded before the fork. The crashing stack in my reports went through os_log preference loading, with no font code in the frames I captured, and I never tested the all_fonts_map() theory on its own. Exec'd workers made the question moot.
Counting what didn't run
The information was on screen the whole time. The progress line said 211 of 406, and the summary printed Ran 459 tests without comparing it to the 654 the run started with. If I write a runner like this again, the parent will compare the number of results it received against the number of tests it handed out, and treat any gap as a failure, independent of exit codes. An exit status only tells you a worker died. A count also catches a worker that stopped early for some reason nobody has thought of yet.