kate-deft/docs/PERF.md

5.9 KiB
Raw Blame History

Switcher performance — measured bottlenecks

Benchmark harness: src/project/bench_switcher.cpp (manual target bench_switcher, not a ctest). Run:

cmake -S . -B build-rel -DCMAKE_BUILD_TYPE=Release
cmake --build build-rel --target bench_switcher
./build-rel/bin/bench_switcher <project-root> /tmp/switcher_metrics.txt

Reference project: ~/prj/r7.20/mobydick — 12,234 files, 17,184 ctags symbols, 6.3 GB tree.

One-time cost when the switcher opens (UI thread, synchronous)

Stage Release Debug (shipped)
ProjectIndex::listFiles ~84 ms ~78 ms
SymbolIndex::listSymbols ~974 ms ~499 ms
build items + setItems ~5 ms ~45-70 ms
  • listFiles is a git ls-files subprocess — acceptable but blocks the UI.
  • listSymbols is a ctags subprocess (~0.5-1 s) — the dominant open-cost, and it blocks the UI thread. (Cost is in the external binary, not our parse.)

Per-keystroke cost — the interactive bottleneck

PaletteModel::setQuery → re-rank. Originally O(N) every keystroke; the cost grew with query length even as the result set shrank.

Full-scan (every keystroke re-scores all N), Release:

Query files symbols
server ~14 ms ~30 ms
src server go ~17 ms ~36 ms

Debug (what ships): 270-342 ms (files), 459-598 ms (symbols) per keystroke — unusable.

Fix 1 — incremental narrowing (done)

When a query only appends to the previous one, re-score only the currently visible subset (the match set is monotonic under lengthening/adding a needle), turning later keystrokes from O(N) into O(previous matches). Exactness is guarded by test_palettemodel::incrementalMatchesFullScan (incremental == full-scan, incl. the typo tier).

Effect once the set narrows (Release, type-forward):

Query full-scan incremental
files server ~14 ms ~6.6 ms
symbols server ~30 ms ~9.0 ms

Early keystrokes (s, se) remain full-N because almost everything matches — correct and unavoidable without a prefix index.

Fix 2 — build the installed plugin optimized (done)

The root CMakeLists.txt now defaults CMAKE_BUILD_TYPE to RelWithDebInfo when the user does not specify one, so an unqualified cmake -B build no longer ships a -O0 plugin. Measured effect on the symbol switcher (mobydick, type-forward "serv", 17k symbols): ~340 ms → ~29 ms per keystroke (~12×).

Fix 3 — debounce the filter (done)

PaletteWidget coalesces rapid keystrokes with an 80 ms single-shot timer before re-ranking (applyQueryNow flushes on Enter so activation always uses the latest text). A fast typist's burst (s,se,ser…) triggers one re-rank instead of three, so the broad-query early keystrokes — the only ones that still scan most of N — are paid at most once per pause, not per character.

Net result after fixes 1-3: the worst single symbol keystroke on a 17k-symbol project is ~29 ms (RelWithDebInfo), and bursts are coalesced; the Alt+G panel is responsive.

Fix 4 — top-K cap, cheap bulk scoring, min query length (done)

Universal Ctags lifted the symbol count on mobydick from ~17k to ~397k (Kotlin/TypeScript now indexed), which brought the per-keystroke lag back. Three changes in PaletteModel / FuzzyRanker fix it:

  • Score-only bulk pass. FuzzyRanker::score(..., withRanges=false) skips the highlight-range work (and the old std::set union, replaced by a sort+unique vector) during the full scan; ranges are computed only for the displayed rows. Scores are identical.
  • Top-K cap (kMaxVisible = 1000). Scores are computed for all items, but only the best 1000 are kept (std::partial_sort) and given ranges. Nobody scrolls past hundreds of fuzzy hits.
  • Minimum query length (kMinQueryChars = 2). A 1-char query matches almost everything, so its full scan is expensive and useless — below the threshold the palette shows the capped full list unranked. The first real scan only runs at 2 chars; from there the incremental path makes every keystroke sub-2 ms.

Measured on mobydick (397k symbols, RelWithDebInfo, type-forward):

keystroke before after
s 337 ms ~0 ms (unranked)
se 507 ms 373 ms (the single first scan)
ser 586 ms 1.9 ms
serv 704 ms 1.4 ms
server 198 ms 0.8 ms

The one remaining O(N) cost is the single first scan at the 2nd char. It is now parallelised across the thread pool (QtConcurrent::blockingMapped over per-thread chunks in PaletteModel::rebuild, above kParallelThreshold = 10000 items). The scan is a pure map over read-only data, concatenated in chunk order and sorted deterministically, so the result is identical to the serial path regardless of thread timing. Measured first-scan at 397k symbols: ~373 ms → ~61 ms (16 cores). Everything after is incremental (<2 ms).

Indexing off the UI thread (done)

ProjectIndexer (src/project/projectindexer.{h,cpp}, in project_lib, 6 tests) runs listFiles / listSymbols on the thread pool (QtConcurrent + QFutureWatcher) and delivers results on the UI thread via filesReady / symbolsReady, caching per project root. The switchers now:

  • open the palette immediately (cache hit → instant; miss → a transient "Indexing…" / "Indexing symbols…" placeholder), and
  • repopulate via the ready signal when the background job finishes.

A per-root generation counter drops superseded/invalidated results. The cache is dropped on KateProjectBridge::projectChanged so a project switch re-indexes. Net effect: the UI thread never blocks on git/ctags; the ~1 s Alt+G stall is gone (first open shows a placeholder and fills in; subsequent opens are instant from cache).