Perf: prefer Universal Ctags, top-K cap + score-only + min query len

Universal Ctags: SymbolIndex::ctagsBinary() prefers ctags-universal over
the alternatives-managed ctags (Exuberant 5.9 has no Kotlin/TS parser).
Kotlin/TypeScript now indexed — symbol count on mobydick: 17k -> 397k.

At 397k, per-keystroke ranking was 0.3-0.7s. Three structural fixes:

1. FuzzyRanker: replace std::set with vector+sort+unique; add
   withRanges=false fast path that skips highlight computation during
   the bulk scan (ranges materialised only for displayed rows).
2. PaletteModel: cap visible results to top-1000 (partial_sort),
   score-only bulk, ranges for the top-K only. Nobody scrolls 400k.
3. Min query length (2): 1-char queries match everything and produce a
   useless, expensive full scan; below the threshold the palette shows
   the capped full list unranked. First real ranking at 2 chars; from
   there incremental narrowing makes every keystroke sub-2ms.

Measured (mobydick 397k symbols, RelWithDebInfo, type-forward):
  s: 337ms -> ~0ms | se: 507ms -> 373ms (1x first scan) |
  ser: 586ms -> 1.9ms | server: 198ms -> 0.8ms.

3 new tests (incrementalMatchesFullScan, visibleResultsAreCapped,
shortQueryDoesNotRank). docs/PERF.md updated.
This commit is contained in:
Levi Neely 2026-10-08 11:31:36 +02:00
parent a5b8d628b3
commit 1caa8d2221
8 changed files with 199 additions and 48 deletions

View File

@ -76,7 +76,39 @@ Net result after fixes 1-3: the worst single symbol keystroke on a 17k-symbol
project is ~29 ms (RelWithDebInfo), and bursts are coalesced; the Alt+G panel
is responsive.
## Follow-up — move indexing off the UI thread (done)
## Fix 4 — top-K cap, cheap bulk scoring, min query length (done)
Universal Ctags lifted the symbol count on mobydick from ~17k to **~397k**
(Kotlin/TypeScript now indexed), which brought the per-keystroke lag back. Three
changes in `PaletteModel` / `FuzzyRanker` fix it:
- **Score-only bulk pass.** `FuzzyRanker::score(..., withRanges=false)` skips the
highlight-range work (and the old `std::set` union, replaced by a
sort+unique vector) during the full scan; ranges are computed only for the
displayed rows. Scores are identical.
- **Top-K cap (`kMaxVisible = 1000`).** Scores are computed for all items, but
only the best 1000 are kept (`std::partial_sort`) and given ranges. Nobody
scrolls past hundreds of fuzzy hits.
- **Minimum query length (`kMinQueryChars = 2`).** A 1-char query matches almost
everything, so its full scan is expensive and useless — below the threshold
the palette shows the capped full list unranked. The first real scan only runs
at 2 chars; from there the incremental path makes every keystroke sub-2 ms.
Measured on mobydick (397k symbols, RelWithDebInfo, type-forward):
| keystroke | before | after |
|-----------|-------:|------:|
| `s` | 337 ms | **~0 ms** (unranked) |
| `se` | 507 ms | 373 ms (the single first scan) |
| `ser` | 586 ms | **1.9 ms** |
| `serv` | 704 ms | **1.4 ms** |
| `server` | 198 ms | **0.8 ms** |
The one remaining O(N) cost is the single first scan at the 2nd char (~370 ms at
397k), absorbed by the 80 ms debounce during a typing burst. Parallelising that
bulk scan (QtConcurrent across cores) is an optional future ~8× on that one hit.
## Indexing off the UI thread (done)
`ProjectIndexer` (`src/project/projectindexer.{h,cpp}`, in `project_lib`, 6
tests) runs `listFiles` / `listSymbols` on the thread pool (`QtConcurrent` +

View File

@ -6,7 +6,7 @@
#include <QChar>
#include <algorithm>
#include <set>
#include <vector>
namespace katecustom
{
@ -200,12 +200,12 @@ NeedleMatch bestNeedleMatch(const QString &cand, const QString &needle)
return tryTypo(cand, needle);
}
std::vector<MatchRange> indicesToRanges(const std::set<int> &idx)
std::vector<MatchRange> indicesToRanges(const std::vector<int> &sortedIdx)
{
std::vector<MatchRange> ranges;
int start = -1;
int prev = -2;
for (int i : idx) {
for (int i : sortedIdx) {
if (i == prev + 1) {
prev = i;
continue;
@ -237,7 +237,8 @@ QList<QString> FuzzyRanker::tokenize(QStringView query)
return needles;
}
MatchResult FuzzyRanker::score(QStringView candidate, const QList<QString> &needles)
MatchResult FuzzyRanker::score(QStringView candidate, const QList<QString> &needles,
bool withRanges)
{
MatchResult result;
@ -250,7 +251,11 @@ MatchResult FuzzyRanker::score(QStringView candidate, const QList<QString> &need
const QString cand = foldString(candidate);
std::set<int> allIndices;
// Collect matched indices in a plain vector; a std::set here was a hot spot
// when scoring hundreds of thousands of candidates per keystroke. We only
// need the union size (coverage) for the score, and the sorted unique set
// for ranges — both derived cheaply from the vector below.
std::vector<int> idx;
int total = 0;
for (const QString &needle : needles) {
const NeedleMatch m = bestNeedleMatch(cand, needle);
@ -258,16 +263,19 @@ MatchResult FuzzyRanker::score(QStringView candidate, const QList<QString> &need
return MatchResult{}; // AND semantics: any miss rejects.
}
total += m.score;
for (int i : m.indices) {
allIndices.insert(i);
}
idx.insert(idx.end(), m.indices.begin(), m.indices.end());
}
std::sort(idx.begin(), idx.end());
idx.erase(std::unique(idx.begin(), idx.end()), idx.end());
result.matched = true;
// Prefer shorter candidates and reward coverage of the candidate.
const int coverage = static_cast<int>(allIndices.size());
const int coverage = static_cast<int>(idx.size());
result.score = total + coverage * 2 - candidate.size();
result.ranges = indicesToRanges(allIndices);
if (withRanges) {
result.ranges = indicesToRanges(idx); // idx already sorted & unique
}
return result;
}

View File

@ -68,8 +68,15 @@ public:
* Score \a candidate against the already-tokenized \a needles. All needles
* must match (AND semantics) for \c matched to be true. Order of needles
* relative to the candidate is irrelevant (orderless).
*
* When \a withRanges is false the match highlight ranges are not computed
* (the returned MatchResult::ranges is empty) — a measurably cheaper path
* for bulk scoring of a large candidate set, where ranges are only needed
* for the handful of rows actually displayed. The \c matched flag and
* \c score are identical either way.
*/
static MatchResult score(QStringView candidate, const QList<QString> &needles);
static MatchResult score(QStringView candidate, const QList<QString> &needles,
bool withRanges = true);
/*! Convenience: tokenize \a query then score. */
static MatchResult score(QStringView candidate, QStringView query);

View File

@ -10,6 +10,21 @@
namespace katecustom
{
namespace
{
// Cap the number of displayed rows. Nobody scrolls past a few hundred fuzzy
// results, and materialising highlight ranges + sorting is bounded by this.
// Scores are still computed for ALL items (so the best truly surface); only
// the top-K are kept and given ranges.
constexpr int kMaxVisible = 1000;
// Minimum query length before any text ranking happens. A 1-char query matches
// almost everything on a large project, so its full scan is both expensive and
// useless. Below this, the palette shows the (capped) full list unranked; the
// first real scan only runs once the query is selective enough to matter.
constexpr int kMinQueryChars = 2;
} // namespace
PaletteModel::PaletteModel(QObject *parent)
: QAbstractListModel(parent)
{
@ -25,7 +40,11 @@ void PaletteModel::setItems(QList<PaletteItem> items)
void PaletteModel::setQuery(const QString &query)
{
if (query == m_query) {
// Below the minimum length, do not rank: treat as the empty query so the
// palette shows the (capped) full list without an expensive, pointless
// 1-char full scan. The first real ranking happens at kMinQueryChars.
const QString effective = query.size() < kMinQueryChars ? QString() : query;
if (effective == m_query) {
return;
}
beginResetModel();
@ -34,11 +53,16 @@ void PaletteModel::setQuery(const QString &query)
// needle or adding one), so we re-score just the currently-visible subset
// instead of all N items. This turns the common type-forward case from
// O(N) per keystroke into O(previous matches), which collapses quickly.
const bool pureAppend = !m_query.isEmpty() && query.startsWith(m_query);
//
// Note: when the previous result was truncated by the top-K cap, narrowing
// works from that capped subset — the standard fuzzy-finder behaviour. The
// dropped items were lower-scored for the shorter query; the tradeoff buys
// sub-millisecond keystrokes on huge symbol sets.
const bool pureAppend = !m_query.isEmpty() && effective.startsWith(m_query);
if (pureAppend) {
rebuildFromVisible(query);
rebuildFromVisible(effective);
} else {
m_query = query;
m_query = effective;
rebuild();
}
endResetModel();
@ -46,22 +70,17 @@ void PaletteModel::setQuery(const QString &query)
void PaletteModel::rebuild()
{
m_visible.clear();
const QList<QString> needles = FuzzyRanker::tokenize(m_query);
m_scored.clear();
for (int i = 0; i < m_items.size(); ++i) {
const PaletteItem &item = m_items.at(i);
const MatchResult r = FuzzyRanker::score(item.label, needles);
if (!r.matched) {
continue;
// Score-only (no ranges) for the bulk scan — the hot path at large N.
const MatchResult r = FuzzyRanker::score(item.label, needles, /*withRanges=*/false);
if (r.matched) {
m_scored.push_back({i, r.score + item.frecency});
}
// Frecency nudges well-used commands up without overriding a clearly
// stronger textual match.
m_visible.push_back({i, r.score + item.frecency, r.ranges});
}
sortVisible();
finalizeVisible(needles);
}
void PaletteModel::rebuildFromVisible(const QString &newQuery)
@ -69,28 +88,46 @@ void PaletteModel::rebuildFromVisible(const QString &newQuery)
const QList<QString> needles = FuzzyRanker::tokenize(newQuery);
m_query = newQuery;
// Re-score only the items that matched the shorter query. Rewrite m_visible
// in place: survivors keep their slot, dropped items are compacted out.
int write = 0;
for (int r = 0; r < m_visible.size(); ++r) {
const int src = m_visible.at(r).sourceIndex;
const MatchResult res = FuzzyRanker::score(m_items.at(src).label, needles);
if (!res.matched) {
continue;
// Re-score only the items that matched the shorter query (the match set is
// monotonic under appending). Score-only; ranges filled for the top-K below.
m_scored.clear();
for (const Visible &v : m_visible) {
const int src = v.sourceIndex;
const MatchResult r = FuzzyRanker::score(m_items.at(src).label, needles,
/*withRanges=*/false);
if (r.matched) {
m_scored.push_back({src, r.score + m_items.at(src).frecency});
}
m_visible[write] = {src, res.score + m_items.at(src).frecency, res.ranges};
++write;
}
m_visible.resize(write);
sortVisible();
finalizeVisible(needles);
}
void PaletteModel::sortVisible()
void PaletteModel::finalizeVisible(const QList<QString> &needles)
{
// Highest score first; stable so equal scores keep input order.
std::stable_sort(m_visible.begin(), m_visible.end(),
[](const Visible &a, const Visible &b) { return a.score > b.score; });
// Keep only the top-K by score (partial sort), then materialise ranges just
// for those rows. Scores were computed for every candidate, so the best
// truly surface; only display + highlighting is bounded.
const int keep = std::min<int>(kMaxVisible, static_cast<int>(m_scored.size()));
std::partial_sort(
m_scored.begin(), m_scored.begin() + keep, m_scored.end(),
[](const Scored &a, const Scored &b) {
// Higher score first; ties broken by source order so results are
// deterministic (partial_sort is not stable on its own).
if (a.score != b.score) {
return a.score > b.score;
}
return a.sourceIndex < b.sourceIndex;
});
m_visible.clear();
m_visible.reserve(keep);
for (int r = 0; r < keep; ++r) {
const Scored &s = m_scored.at(r);
// Recompute with ranges only for the displayed rows.
const MatchResult res = FuzzyRanker::score(m_items.at(s.sourceIndex).label,
needles, /*withRanges=*/true);
m_visible.push_back({s.sourceIndex, s.score, res.ranges});
}
}
const PaletteItem *PaletteModel::itemAt(int row) const

View File

@ -66,16 +66,23 @@ private:
int score;
std::vector<MatchRange> ranges;
};
// Lightweight score-only entry used during the bulk scan (no ranges).
struct Scored {
int sourceIndex;
int score;
};
void rebuild();
// Re-score only the currently-visible items against a longer (appended)
// query; the match set is monotonic so narrowing is exact. Much cheaper
// than a full rebuild for the common type-forward case.
void rebuildFromVisible(const QString &newQuery);
void sortVisible();
// Partial-sort m_scored to the top-K and materialise ranges for those rows.
void finalizeVisible(const QList<QString> &needles);
QList<PaletteItem> m_items;
QList<Visible> m_visible;
std::vector<Scored> m_scored; // scratch reused across rebuilds
QString m_query;
};

View File

@ -199,6 +199,44 @@ private Q_SLOTS:
const QStringList full = rowsFor({QStringLiteral("server")});
QCOMPARE(incremental, full);
}
// The visible set is capped (top-K by score) so very broad queries do not
// materialise hundreds of thousands of rows. Scores are still computed for
// all items, so the best matches are the ones kept.
void visibleResultsAreCapped()
{
QList<PaletteItem> items;
// 5000 items all containing "x" so an empty / "x" query matches all.
for (int i = 0; i < 5000; ++i) {
items.push_back({QString::number(i),
QStringLiteral("item_x_%1").arg(i), QString(), 0});
}
PaletteModel m;
m.setItems(items);
// Empty query matches everything but must be capped.
QVERIFY(m.rowCount() <= 1000);
}
// Below the minimum query length, no text ranking happens: the palette
// shows the (capped) full list, so a 1-char query does not full-scan.
void shortQueryDoesNotRank()
{
QList<PaletteItem> items;
items.push_back({QStringLiteral("a"), QStringLiteral("alpha"), QString(), 0});
items.push_back({QStringLiteral("b"), QStringLiteral("beta"), QString(), 0});
items.push_back({QStringLiteral("c"), QStringLiteral("gamma"), QString(), 0});
PaletteModel m;
m.setItems(items);
// One char: treated as empty → all rows shown, unranked.
m.setQuery(QStringLiteral("a"));
QCOMPARE(m.rowCount(), 3);
// Two chars: real ranking kicks in and filters.
m.setQuery(QStringLiteral("al"));
QCOMPARE(m.rowCount(), 1);
QCOMPARE(m.index(0, 0).data(PaletteModel::LabelRole).toString(),
QStringLiteral("alpha"));
}
};
QTEST_MAIN(TestPaletteModel)

View File

@ -9,9 +9,23 @@
namespace katecustom
{
QString SymbolIndex::ctagsBinary()
{
// Prefer Universal Ctags: it parses far more languages (Kotlin, TypeScript,
// Vue, Rust, …) that Exuberant 5.9 — often the system `ctags` default via
// update-alternatives — does not. It installs as `ctags-universal` on
// Debian/Ubuntu. Fall back to a plain `ctags` on PATH otherwise.
const QString universal =
QStandardPaths::findExecutable(QStringLiteral("ctags-universal"));
if (!universal.isEmpty()) {
return universal;
}
return QStandardPaths::findExecutable(QStringLiteral("ctags"));
}
bool SymbolIndex::ctagsAvailable()
{
return !QStandardPaths::findExecutable(QStringLiteral("ctags")).isEmpty();
return !ctagsBinary().isEmpty();
}
QList<Symbol> SymbolIndex::parseTags(const QString &ctagsOutput)
@ -90,8 +104,9 @@ QList<Symbol> SymbolIndex::listSymbols(const QString &root, const QStringList &f
// Classic extended format with kind + line fields, numeric addresses, and
// the file list fed on stdin (-L -) so a huge project never overflows the
// command line. Paths are relative to the working directory, so the parsed
// Symbol::file is already project-relative.
ctags.start(QStringLiteral("ctags"),
// Symbol::file is already project-relative. The format flags are accepted by
// both Exuberant and Universal Ctags.
ctags.start(ctagsBinary(),
{QStringLiteral("-f"), QStringLiteral("-"),
QStringLiteral("-L"), QStringLiteral("-"),
QStringLiteral("--fields=+nK"),

View File

@ -51,6 +51,13 @@ public:
/*! True if a ctags binary is usable. */
static bool ctagsAvailable();
/*!
* Path to the ctags binary to use: prefers Universal Ctags
* (`ctags-universal`, which parses Kotlin/TypeScript/… that Exuberant 5.9
* cannot), falling back to a plain `ctags` on PATH. Empty if none found.
*/
static QString ctagsBinary();
/*!
* Parse ctags extended tab-separated output into a symbol list. Pure: no
* process, no filesystem. Lines that are comments (starting with "!_TAG_")