5.9 KiB
Switcher performance — measured bottlenecks
Benchmark harness: src/project/bench_switcher.cpp (manual target bench_switcher,
not a ctest). Run:
cmake -S . -B build-rel -DCMAKE_BUILD_TYPE=Release
cmake --build build-rel --target bench_switcher
./build-rel/bin/bench_switcher <project-root> /tmp/switcher_metrics.txt
Reference project: ~/prj/r7.20/mobydick — 12,234 files, 17,184 ctags
symbols, 6.3 GB tree.
One-time cost when the switcher opens (UI thread, synchronous)
| Stage | Release | Debug (shipped) |
|---|---|---|
ProjectIndex::listFiles |
~84 ms | ~78 ms |
SymbolIndex::listSymbols |
~974 ms | ~499 ms |
build items + setItems |
~5 ms | ~45-70 ms |
listFilesis agit ls-filessubprocess — acceptable but blocks the UI.listSymbolsis a ctags subprocess (~0.5-1 s) — the dominant open-cost, and it blocks the UI thread. (Cost is in the external binary, not our parse.)
Per-keystroke cost — the interactive bottleneck
PaletteModel::setQuery → re-rank. Originally O(N) every keystroke; the cost
grew with query length even as the result set shrank.
Full-scan (every keystroke re-scores all N), Release:
| Query | files | symbols |
|---|---|---|
server |
~14 ms | ~30 ms |
src server go |
~17 ms | ~36 ms |
Debug (what ships): 270-342 ms (files), 459-598 ms (symbols) per keystroke — unusable.
Fix 1 — incremental narrowing (done)
When a query only appends to the previous one, re-score only the currently
visible subset (the match set is monotonic under lengthening/adding a needle),
turning later keystrokes from O(N) into O(previous matches). Exactness is
guarded by test_palettemodel::incrementalMatchesFullScan (incremental ==
full-scan, incl. the typo tier).
Effect once the set narrows (Release, type-forward):
| Query | full-scan | incremental |
|---|---|---|
files server |
~14 ms | ~6.6 ms |
symbols server |
~30 ms | ~9.0 ms |
Early keystrokes (s, se) remain full-N because almost everything matches —
correct and unavoidable without a prefix index.
Fix 2 — build the installed plugin optimized (done)
The root CMakeLists.txt now defaults CMAKE_BUILD_TYPE to RelWithDebInfo
when the user does not specify one, so an unqualified cmake -B build no longer
ships a -O0 plugin. Measured effect on the symbol switcher (mobydick,
type-forward "serv", 17k symbols): ~340 ms → ~29 ms per keystroke (~12×).
Fix 3 — debounce the filter (done)
PaletteWidget coalesces rapid keystrokes with an 80 ms single-shot timer
before re-ranking (applyQueryNow flushes on Enter so activation always uses
the latest text). A fast typist's burst (s,se,ser…) triggers one re-rank
instead of three, so the broad-query early keystrokes — the only ones that still
scan most of N — are paid at most once per pause, not per character.
Net result after fixes 1-3: the worst single symbol keystroke on a 17k-symbol project is ~29 ms (RelWithDebInfo), and bursts are coalesced; the Alt+G panel is responsive.
Fix 4 — top-K cap, cheap bulk scoring, min query length (done)
Universal Ctags lifted the symbol count on mobydick from ~17k to ~397k
(Kotlin/TypeScript now indexed), which brought the per-keystroke lag back. Three
changes in PaletteModel / FuzzyRanker fix it:
- Score-only bulk pass.
FuzzyRanker::score(..., withRanges=false)skips the highlight-range work (and the oldstd::setunion, replaced by a sort+unique vector) during the full scan; ranges are computed only for the displayed rows. Scores are identical. - Top-K cap (
kMaxVisible = 1000). Scores are computed for all items, but only the best 1000 are kept (std::partial_sort) and given ranges. Nobody scrolls past hundreds of fuzzy hits. - Minimum query length (
kMinQueryChars = 2). A 1-char query matches almost everything, so its full scan is expensive and useless — below the threshold the palette shows the capped full list unranked. The first real scan only runs at 2 chars; from there the incremental path makes every keystroke sub-2 ms.
Measured on mobydick (397k symbols, RelWithDebInfo, type-forward):
| keystroke | before | after |
|---|---|---|
s |
337 ms | ~0 ms (unranked) |
se |
507 ms | 373 ms (the single first scan) |
ser |
586 ms | 1.9 ms |
serv |
704 ms | 1.4 ms |
server |
198 ms | 0.8 ms |
The one remaining O(N) cost is the single first scan at the 2nd char. It is now
parallelised across the thread pool (QtConcurrent::blockingMapped over
per-thread chunks in PaletteModel::rebuild, above kParallelThreshold = 10000
items). The scan is a pure map over read-only data, concatenated in chunk order
and sorted deterministically, so the result is identical to the serial path
regardless of thread timing. Measured first-scan at 397k symbols: ~373 ms →
~61 ms (16 cores). Everything after is incremental (<2 ms).
Indexing off the UI thread (done)
ProjectIndexer (src/project/projectindexer.{h,cpp}, in project_lib, 6
tests) runs listFiles / listSymbols on the thread pool (QtConcurrent +
QFutureWatcher) and delivers results on the UI thread via filesReady /
symbolsReady, caching per project root. The switchers now:
- open the palette immediately (cache hit → instant; miss → a transient "Indexing…" / "Indexing symbols…" placeholder), and
- repopulate via the ready signal when the background job finishes.
A per-root generation counter drops superseded/invalidated results. The cache
is dropped on KateProjectBridge::projectChanged so a project switch re-indexes.
Net effect: the UI thread never blocks on git/ctags; the ~1 s Alt+G stall is
gone (first open shows a placeholder and fills in; subsequent opens are instant
from cache).