docs: remove PERF.md and its references

The performance notes documented one-off benchmarking of the switcher re-rank
and the fixes already applied; it is no longer useful as living documentation.
Drop docs/PERF.md and the references to it in README (layout tree), AGENTS.md
(reference-docs list and the build-type note), and the CMakeLists comment.
This commit is contained in:
Levi Neely 2026-10-08 20:46:31 +02:00
parent 33790043d8
commit 38f5468ce6
4 changed files with 3 additions and 135 deletions

View File

@ -45,7 +45,7 @@ cmake --install build # installs the 6 .so to ~/.local/.../qt6/plu
land where Kate scans.
- An unqualified `cmake -B build` produces an **optimized** build on purpose
(`CMAKE_BUILD_TYPE=RelWithDebInfo`); the palette re-rank and indexing are
several times slower at `-O0`. See `docs/PERF.md`.
several times slower at `-O0`.
## Conventions that matter
@ -106,7 +106,6 @@ plugin's header or add a link dependency between plugins.
## Reference docs
- `docs/PERF.md` — switcher performance analysis and the fixes applied.
- `docs/PLUMBING.md` — configuring the Plan 9 plumber to route into Kate.
- `docs/SAM.md` — the sam structural-regexp dialect and semantics.
- `docs/radials.example.json` — example radial-menu config.

View File

@ -10,8 +10,8 @@ set(QT_MIN_VERSION 6.5.0)
set(KF_MIN_VERSION 6.0.0)
# Default to an optimized build when the user did not pick one. The palette
# re-rank and indexing are measurably ~5-8x slower under -O0 (see docs/PERF.md),
# so an unqualified `cmake -B build` should not ship a Debug plugin.
# re-rank and indexing are measurably ~5-8x slower under -O0, so an unqualified
# `cmake -B build` should not ship a Debug plugin.
if(NOT CMAKE_BUILD_TYPE AND NOT CMAKE_CONFIGURATION_TYPES)
set(CMAKE_BUILD_TYPE RelWithDebInfo CACHE STRING
"Build type (Debug, Release, RelWithDebInfo, MinSizeRel)" FORCE)

View File

@ -288,7 +288,6 @@ src/
plumb/
plumb* + plumbplugin [deft:nineify] plumb — Plan 9 plumber
docs/
PERF.md switcher performance notes
PLUMBING.md configuring the Plan 9 plumber for Kate
SAM.md the sam structural-regexp dialect
radials.example.json example radial-menu config

View File

@ -1,130 +0,0 @@
# Switcher performance — measured bottlenecks
Benchmark harness: `src/project/bench_switcher.cpp` (manual target `bench_switcher`,
not a ctest). Run:
```sh
cmake -S . -B build-rel -DCMAKE_BUILD_TYPE=Release
cmake --build build-rel --target bench_switcher
./build-rel/bin/bench_switcher <project-root> /tmp/switcher_metrics.txt
```
Reference project: `~/prj/r7.20/mobydick` — **12,234 files**, **17,184 ctags
symbols**, 6.3 GB tree.
## One-time cost when the switcher opens (UI thread, synchronous)
| Stage | Release | Debug (shipped) |
|----------------------------|--------:|----------------:|
| `ProjectIndex::listFiles` | ~84 ms | ~78 ms |
| `SymbolIndex::listSymbols` | ~974 ms | ~499 ms |
| build items + `setItems` | ~5 ms | ~45-70 ms |
- `listFiles` is a `git ls-files` subprocess — acceptable but blocks the UI.
- `listSymbols` is a **ctags subprocess (~0.5-1 s)** — the dominant open-cost,
and it blocks the UI thread. (Cost is in the external binary, not our parse.)
## Per-keystroke cost — the interactive bottleneck
`PaletteModel::setQuery` → re-rank. Originally O(N) every keystroke; the cost
*grew* with query length even as the result set shrank.
Full-scan (every keystroke re-scores all N), Release:
| Query | files | symbols |
|------------------|---------:|---------:|
| `server` | ~14 ms | ~30 ms |
| `src server go` | ~17 ms | ~36 ms |
Debug (what ships): **270-342 ms (files)**, **459-598 ms (symbols)** per
keystroke — unusable.
## Fix 1 — incremental narrowing (done)
When a query only **appends** to the previous one, re-score only the currently
visible subset (the match set is monotonic under lengthening/adding a needle),
turning later keystrokes from O(N) into O(previous matches). Exactness is
guarded by `test_palettemodel::incrementalMatchesFullScan` (incremental ==
full-scan, incl. the typo tier).
Effect once the set narrows (Release, type-forward):
| Query | full-scan | incremental |
|----------|----------:|------------:|
| files `server` | ~14 ms | **~6.6 ms** |
| symbols `server` | ~30 ms | **~9.0 ms** |
Early keystrokes (`s`, `se`) remain full-N because almost everything matches —
correct and unavoidable without a prefix index.
## Fix 2 — build the installed plugin optimized (done)
The root `CMakeLists.txt` now defaults `CMAKE_BUILD_TYPE` to `RelWithDebInfo`
when the user does not specify one, so an unqualified `cmake -B build` no longer
ships a `-O0` plugin. Measured effect on the symbol switcher (mobydick,
type-forward "serv", 17k symbols): **~340 ms → ~29 ms per keystroke** (~12×).
## Fix 3 — debounce the filter (done)
`PaletteWidget` coalesces rapid keystrokes with an 80 ms single-shot timer
before re-ranking (`applyQueryNow` flushes on Enter so activation always uses
the latest text). A fast typist's burst (`s`,`se`,`ser`…) triggers one re-rank
instead of three, so the broad-query early keystrokes — the only ones that still
scan most of N — are paid at most once per pause, not per character.
Net result after fixes 1-3: the worst single symbol keystroke on a 17k-symbol
project is ~29 ms (RelWithDebInfo), and bursts are coalesced; the Alt+G panel
is responsive.
## Fix 4 — top-K cap, cheap bulk scoring, min query length (done)
Universal Ctags lifted the symbol count on mobydick from ~17k to **~397k**
(Kotlin/TypeScript now indexed), which brought the per-keystroke lag back. Three
changes in `PaletteModel` / `FuzzyRanker` fix it:
- **Score-only bulk pass.** `FuzzyRanker::score(..., withRanges=false)` skips the
highlight-range work (and the old `std::set` union, replaced by a
sort+unique vector) during the full scan; ranges are computed only for the
displayed rows. Scores are identical.
- **Top-K cap (`kMaxVisible = 1000`).** Scores are computed for all items, but
only the best 1000 are kept (`std::partial_sort`) and given ranges. Nobody
scrolls past hundreds of fuzzy hits.
- **Minimum query length (`kMinQueryChars = 2`).** A 1-char query matches almost
everything, so its full scan is expensive and useless — below the threshold
the palette shows the capped full list unranked. The first real scan only runs
at 2 chars; from there the incremental path makes every keystroke sub-2 ms.
Measured on mobydick (397k symbols, RelWithDebInfo, type-forward):
| keystroke | before | after |
|-----------|-------:|------:|
| `s` | 337 ms | **~0 ms** (unranked) |
| `se` | 507 ms | 373 ms (the single first scan) |
| `ser` | 586 ms | **1.9 ms** |
| `serv` | 704 ms | **1.4 ms** |
| `server` | 198 ms | **0.8 ms** |
The one remaining O(N) cost is the single first scan at the 2nd char. It is now
**parallelised** across the thread pool (`QtConcurrent::blockingMapped` over
per-thread chunks in `PaletteModel::rebuild`, above `kParallelThreshold = 10000`
items). The scan is a pure map over read-only data, concatenated in chunk order
and sorted deterministically, so the result is identical to the serial path
regardless of thread timing. Measured first-scan at 397k symbols: **~373 ms →
~61 ms** (16 cores). Everything after is incremental (<2 ms).
## Indexing off the UI thread (done)
`ProjectIndexer` (`src/project/projectindexer.{h,cpp}`, in `project_lib`, 6
tests) runs `listFiles` / `listSymbols` on the thread pool (`QtConcurrent` +
`QFutureWatcher`) and delivers results on the UI thread via `filesReady` /
`symbolsReady`, caching per project root. The switchers now:
- open the palette **immediately** (cache hit → instant; miss → a transient
"Indexing…" / "Indexing symbols…" placeholder), and
- repopulate via the ready signal when the background job finishes.
A per-root generation counter drops superseded/invalidated results. The cache
is dropped on `KateProjectBridge::projectChanged` so a project switch re-indexes.
Net effect: the UI thread never blocks on git/ctags; the ~1 s Alt+G stall is
gone (first open shows a placeholder and fills in; subsequent opens are instant
from cache).