Skip to content

Cache toolchain probes and keep GC out of parsing - #54

Merged
jserv merged 10 commits into
mainfrom
fast-load
Sep 7, 2026
Merged

jserv merged 10 commits into
mainfrom
fast-load

Conversation

@jserv

@jserv jserv commented Sep 7, 2026 •

Copy link
Copy Markdown
Contributor

Loading a large Kconfig tree spends most of its time somewhere other than parsing. scripts/Kconfig.include defines cc-option, as-instr, ld-option and friends in terms of $(shell,...), so a full load forks a hundred shells that each fork a compiler, and the cyclic garbage collector walks a heap that only grows while the parse builds it. This caches the probe results, always for the lifetime of a Kconfig instance and, when KCONFIG_SHELL_CACHE names a file, across runs as well, and suppresses collection for the duration of the parse.

The cache cannot simply key on the command. $(python,...) runs in this process with os in scope, so a Kconfig file can move os.environ or the working directory partway through a parse, and a result measured before that must never be reused after it. Comparing the raw environment mapping costs about a microsecond, cheap enough to do before every lookup, while the working directory needs a syscall and so is re-read only when in-process Python may have run. Only results measured in the context the cache file is stamped with are written to it.

Collection is suppressed by zeroing the gen0 threshold rather than by calling gc.disable(), so gc.isenabled() keeps reporting what the caller and the Kconfig files believe: a file running $(python,import gc; gc.disable()) still has that survive the parse. The full collection that reclaims a previously dropped tree runs only when the previous parse allocated enough to have left real work behind, since it walks the host application's whole heap and Kconfiglib is usually embedded in a build script.

The remaining commits are the tooling half. menuconfig's jump-to search lived inside a two-hundred-line curses function and could not be tested; extracting it and running the same cases through both tools turned up a real divergence, where guiconfig left stale matches behind after a regex failed to compile. The row renderers and the MENUCONFIG_STYLE parser had no tests either. CI now measures coverage, and .ci/check-format.sh asks git for its file list instead of walking the tree with find, which had it formatting vendored checkouts and __pycache__.

Verified on an Apple M1, macOS 15.6, CPython 3.14.7. python -m pytest tests/ is 308 passed, 8 skipped; ruff check . and .ci/check-format.sh are clean; the suite also passes with tkinter absent (262 passed, 53 skipped) and with a broken Tk install. Parsing a synthetic 3001-file tree goes from 2.15 s to 1.50 s, best of seven interleaved runs in fresh processes. A 200-probe tree goes from 0.77 s cold to 0.002 s with the cache warm. Peak RSS across six repeated loads stays at about 370 MiB against 445 MiB before. .config output is byte-identical with the cache cold, filling, and warm.

Two limitations are deliberate rather than overlooked. Variables that describe the invocation instead of the toolchain, MAKEFLAGS and MAKELEVEL among them, are left out of the cache fingerprint, because including them means a cache filled by make menuconfig never survives a recursive make; the cost is that a command reading one of them directly gets the answer from whichever invocation filled the cache. And os.putenv() writes the process environment without updating os.environ, so it is invisible to the context check. Both are documented in the module docstring.


Summary by cubic

Caches toolchain probe results and keeps the cyclic garbage collector out of parsing, cutting load time on a 3000-file tree from 2.15s to 1.50s. Also fixes UI search bugs and adds tests, coverage, and a benchmark.

Probe cache and GC

  • Probe results for $(cc-option) and friends are memoized per Kconfig instance; setting KCONFIG_SHELL_CACHE persists them across runs, with a cache key covering environment and working directory so only results measured in the stamped context are reused.
  • Probe functions are identified by identity rather than membership, so custom KCONFIG_FUNCTIONS callables with raising __hash__ or __eq__ no longer crash the parse.
  • GC is suppressed during parsing via a zeroed gen0 threshold; a full collection runs only after a large parse, sized by measured allocation, to reclaim dropped trees.

UI and tooling

  • menuconfig's jump-to search is extracted into a testable function; guiconfig now clears stale matches after a bad regex.
  • Added tests for row renderers, style parser, search across both tools, and the probe cache.
  • CI measures coverage with the pass floor set from runner measurements, since local runs on Python 3.14 score differently; the format script uses git's file listing and filters to files on disk.
  • New benchmark script times load and UI phases, outputting plain text or JSON; it skips the GUI phase when tkinter is absent.

Written for commit a91b888. Summary will update on new commits.

Review in cubic

Two thirds of a large Kconfig load is not parsing.
scripts/Kconfig.include defines cc-option, as-instr, ld-option and
friends in terms of $(shell,...), so a full load forks a hundred shells
that each fork a compiler. Those probes are functions of the command and
the toolchain, so remember them: always for the lifetime of a Kconfig
instance, and, when KCONFIG_SHELL_CACHE names a file, across runs as
well.

Keying on the command alone is not enough. $(python,...) runs in this
process with os in scope, so a Kconfig file can move os.environ or the
working directory partway through a parse, and a result measured before
that must never be reused after it. Comparing the raw environment
mapping costs about a microsecond, cheap enough to do before every
lookup, while the working directory needs a syscall and so is re-read
only when in-process Python may have run. Only results measured in the
context the cache file is stamped with are written to it.
Nothing a parse allocates is cyclic garbage: the symbols, menu nodes and
expression tuples all stay live, and the temporaries die by reference
counting. Every collection a parse triggers is therefore a walk over a
heap that only grows and that it cannot free, worth 27 percent of the
parse on a three thousand file tree.

Suppress it by zeroing the gen0 threshold rather than by calling
gc.disable(), so that gc.isenabled() keeps reporting what the caller and
the Kconfig files believe. A file running $(python,import gc;
gc.disable()) still has that survive the parse, which an unconditional
restore would undo.

Suppression promotes everything the parse allocates out of the young
generations, so an instance dropped earlier is reachable only by a full
collection. Run one first, but only when the previous parse moved the
gen0 allocation counter far enough to have left real work behind. A full
collection walks the host application's whole tracked heap, and
Kconfiglib is embedded in build scripts: charging one to every
construction cost twenty small parses inside a large host 0.8 seconds
for nothing.
The search lived inside _jump_to_dialog(), two hundred lines of curses
that need a terminal, so nothing could exercise it. guiconfig has the
same logic in _update_jump_to_matches(), reachable on its own. Lift
menuconfig's copy into _jump_to_matches() and drive both through the
same cases, so the two implementations cannot drift apart unnoticed.

Comparing them once both were reachable turned one up immediately:
guiconfig cleared the visible list when a regex failed to compile but
left the previous matches in _jump_to_matches, contradicting what that
variable is documented to hold. _is_y_mode_choice_sym() in both tools
also returned the tail of an and chain, so a non-choice symbol yielded
None where the callers document a bool.

This replaces tests/test_ui.py, which covered the same ground against a
reimplementation of the logic written in the test file. A copy cannot
catch a change to the original, which is the reason
tests/test_ui_ranges.py already gives for calling the real function.
_node_str() and _value_str() produce every line a user reads and are the
most branch dense functions in either tool, and nothing exercised them.
They are pure functions of a menu node plus a few module globals, so
they run without a terminal or Tk. The same holds for the
MENUCONFIG_STYLE parser, which is the one piece of configuration users
write by hand and so the one most likely to arrive malformed.

The fixture is built so that every branch is reachable: each symbol
type, pinned and unpinned values, select and imply annotations including
the truncation past two sources, choices in y mode, empty against non
empty menus, and promptless symbols in show-all mode.
The selftest job now runs pytest-cov over the library and the three
tools on every job that runs the full suite, prints the table, stores
the report as an artifact, and fails below a floor. The floor is a
backstop against a large drop rather than a ratchet, since total
coverage is about forty per cent and the tooling half of it is thin. It
is set from what these runners report: coverage switched to the
sys.monitoring backend in Python 3.14 and scores the same suite eleven
points higher, so a floor taken from a modern local run fails every job
here.

check-format.sh walked the tree with find and no exclusions, so it
formatted vendored checkouts, build directories and __pycache__, and a
local run failed on files the project does not own. Ask git for the list
instead, which honours gitignore and the local exclude file and needs no
denylist of its own. Without set -e an unchecked cd also left it
gathering files from wherever it was invoked and still exiting zero, so
a format regression would have passed silently.
Times each phase of a load separately, parse, config rendering, the two
menu walks and terminal rendering, so that a regression can be
attributed rather than merely noticed. Plain table by default, JSON for
storing as a CI artifact.

No pass or fail thresholds. Timings taken on shared runners against an
external kernel tree are too noisy to gate on until a stable baseline
exists.
cubic-dev-ai[bot]

This comment was marked as resolved.

A KCONFIG_FUNCTIONS module may register any callable, and testing one
for set membership runs its own __hash__ and __eq__. Those are free to
raise anything, and the exception took the whole parse down: a callable
whose __hash__ raises ValueError aborted before it ever ran. Guarding
TypeError alone covered only the unhashable case.

Both sets hold functions defined in this file, so the question was never
membership, only whether the callable is one of ours. Ask that directly.
guiconfig pulls in tkinter, which is an optional part of a Python
installation, and the unguarded import took the whole benchmark down
with it. The load, config write, menuconfig and terminal timings were
all collectable without it. Return a skipped phase instead, which the
output already renders.
git ls-files --cached still lists a tracked file that has been deleted
in the working tree, so shfmt and black were handed paths that no longer
exist and printed usage errors for them. The find based version this
replaced only ever saw files that were present.
Both gate tests wrote a fixed number of symbols and relied on that
clearing the reclaim threshold. How many tracked objects a symbol
allocates is an interpreter detail, so on a CPython that allocates fewer
the tests would stop exercising the gate rather than fail, which is the
worse outcome. Grow the tree until the flag actually trips.
@jserv
jserv merged commit 7a57bfc into main Sep 7, 2026
20 checks passed
@jserv
jserv deleted the fast-load branch September 7, 2026 07:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant