diff --git a/docs/PLAN.md b/docs/PLAN.md index 2ac84e0..605c416 100644 --- a/docs/PLAN.md +++ b/docs/PLAN.md @@ -359,6 +359,15 @@ A dockable tool view ("Sam", left sidebar) runs the plan9 **sam** command language against the active document, or project-wide via `X`/`Y`. Type a program, click **Run** (or Ctrl+Return); each Run is one undo step. +- **Regex engine** (`src/sam/samregex.{h,cpp}`, 18 unit tests): a faithful port + of sam's own matcher (plan9port `src/cmd/sam/regexp.c`) — a Thompson/Pike NFA. + **Leftmost-longest** (POSIX), so it matches real sam exactly incl. overlapping + alternation (`a|ab` → `ab`); **linear time**, no backtracking (immune to + `(a*)*b` blowup). sam dialect only: `. * + ? | ( ) [ ] ^ $ \` with `\n`; `^`/`$` + are per-line and `.`/negated-classes exclude newline, all intrinsic. No PCRE + extras — not a deviation, just sam. Replaces the earlier QRegularExpression + matcher, eliminating the greedy-vs-longest divergence entirely. + - **Pure engine** (`src/sam/samengine.{h,cpp}`, unit-tested, no Kate dep): `SamEngine::run(program, text, dotStart, dotEnd) -> SamResult{edits, dot, output, applied}`. Edits are computed against the ORIGINAL snapshot in @@ -383,14 +392,9 @@ program, click **Run** (or Ctrl+Return); each Run is one undo step. - **Panel** (`src/plugin/sampanel.{h,cpp}`): program editor + Run button + output log; `OllieView` creates the tool view and owns the apply logic (`runSamProgram`, `applySamToDocument`, `runSamFileLoop`). -- **Deviations (documented)**: regexes use `QRegularExpression` (PCRE), not - plan9 `regexp(7)`. Everyday patterns match identically; the one semantic - divergence is leftmost-greedy vs sam's leftmost-longest (bites only on - overlapping alternation like `a|ab`). `^`/`$` are per-line (compiled with - `MultilineOption`) and `.` does not cross newlines — both faithful to sam's - `regexp(7)` (BOL/EOL in `regexp.c`). PCRE also accepts extras sam lacks - (`\d \w \b`, lookahead, non-greedy, inline flags); the docs frame these as an - escape hatch, not a feature — prefer structural composition. Full user-facing +- **Dialect & semantics**: the regex engine is sam's own (see the Regex engine + bullet above), so matching is leftmost-longest and the dialect is sam's — no + PCRE, no greedy-vs-longest divergence, no PCRE extras. Full user-facing treatment with verified examples in **`docs/SAM.md`**. Out of scope: multi-file menu (`b B n D`), external file I/O (`e r w f`), `"re"` file-addressing, and sam's own `u` (Kate's undo stack is the undo mechanism). The live GUI panel diff --git a/docs/SAM.md b/docs/SAM.md index 21269ee..aec51a1 100644 --- a/docs/SAM.md +++ b/docs/SAM.md @@ -6,126 +6,103 @@ against the active document, or project-wide via `X`/`Y`. Type a program, press **Ctrl+Return** or click **Run**. Each Run is one undo step (Ctrl+Z reverts the whole Run; multi-file `X`/`Y` undoes per file). -This document covers the **regex deviation** — katesam uses PCRE -(`QRegularExpression`), not plan9 `regexp(7)` — and what that means in practice. -Every example below is verified against the engine. +## The regex engine -## TL;DR +katesam uses a **faithful port of sam's own regular-expression engine** +(`src/sam/samregex.{h,cpp}`, from plan9port `src/cmd/sam/regexp.c`), not PCRE. +This is a Thompson/Pike NFA, which gives the two properties sam relies on: -- Everyday patterns behave **identically** to sam. -- One real semantic difference: **PCRE is leftmost-greedy, sam is - leftmost-longest.** It only bites on *overlapping alternation* (`a|ab`). -- `^` and `$` are **per-line** (match at every line boundary), the same as sam's - `regexp(7)`. `.` does not cross newlines, also the same as sam. -- PCRE *also* offers things sam lacks (`\d \w \b`, lookahead, non-greedy, inline - flags). They work, but prefer sam-style structural composition (`x g v`) — see - "Extras" below. +- **Leftmost-longest** (POSIX) matching, so results match real sam exactly, + including overlapping alternation. +- **Linear time**, with no backtracking — immune to catastrophic blowup. A + pattern like `(a*)*b` against thousands of `a` returns instantly. ---- +There is **no PCRE deviation anymore**: the dialect, the match semantics, and +the anchors are sam's. -## The one semantic divergence: greedy vs longest +### Dialect -sam's `regexp(7)` matches the **leftmost-longest** substring. PCRE matches -**leftmost**, then resolves alternation **left-to-right, first wins** (greedy -within a branch, but alternation order decides between branches). +The sam dialect is deliberately small (sam `regexp(7)`): -| program | input | sam (leftmost-longest) | **katesam (PCRE)** | -|---|---|---|---| -| `,s/a\|ab/X/` | `ab` | `X` (matches `ab`) | **`Xb`** (matches `a`) | -| `,s/ab\|a/X/` | `ab` | `X` | `X` | +``` +. any character except newline +* + ? zero-or-more / one-or-more / zero-or-one +| alternation +( ) grouping and capture (\1..\9 in replacements) +[ ] [^ ] character class / negated class (ranges a-z) +^ $ beginning / end of a LINE +\c literal c (escapes a metacharacter); \n is newline +``` -**Implication:** when alternatives overlap, **order them longest-first** -(`ab|a`, not `a|ab`). This is the only case where a working sam command can -silently produce a different edit in katesam. Non-overlapping alternation -(`cat|dog`) and all non-alternation patterns are unaffected. +There is intentionally **no** `\d \w \s \b`, no lookaround, no non-greedy, and +no inline flags. sam never had them. Use character classes (`[0-9]`, +`[a-zA-Z_]`) and structural composition (`x y g v`) instead — that is the sam +way, and patterns stay portable to real sam. -Quantifier greediness is the same in both (`a.*c` on `axxcxxc` → whole string in -both). katesam additionally offers non-greedy `.*?` (a PCRE extra). +### Semantics (all verified against the port) ---- +| behaviour | katesam | note | +|---|---|---| +| overlapping alternation `a\|ab` on `ab` | matches `ab` | leftmost-**longest**; order of branches does not matter | +| `.` crosses newline | **no** | matches sam | +| `^` / `$` | per-line | `^` = start of a line, `$` = end of a line | +| `[^z]+` crosses newline | no | negated classes also exclude `\n`, as in sam | +| empty-match global `s/x*/-/g` on `axbx` | `-a-b-` | one advance per null match | +| `&` whole match, `\1..\9` groups in replacement | yes | up to 9 capture groups | +| pathological `(a*)*b` | linear time | NFA, no backtracking | -## Anchors `^` / `$` are per-line — same as sam +### Anchors are per-line -This is **not** a deviation. plan9 `regexp(7)` defines `^` as "the beginning of -a line" and `$` as "the end of a line" (sam's `regexp.c`: `BOL` fires at offset -0 or after a `\n`; `EOL` fires before a `\n`). katesam compiles every pattern -with `QRegularExpression::MultilineOption`, so `^`/`$` match at every line -boundary, exactly as sam does: +`^` matches the beginning of any line and `$` the end of any line — built into +the engine (sam `regexp.c`: `BOL` fires at offset 0 or after `\n`; `EOL` before +`\n`). So: | program | input | result | |---|---|---| -| `,s/^/> /g` | `a⏎b⏎c` | `> a⏎> b⏎> c` (every line) | -| `,s/$/;/g` | `a⏎b⏎c` | `a;⏎b;⏎c;` (every line) | +| `,s/^/> /g` | `a⏎b⏎c` | `> a⏎> b⏎> c` | +| `,s/$/;/g` | `a⏎b⏎c` | `a;⏎b;⏎c;` | -The sam-idiomatic structural form works too, and is preferable because it -composes with guards and other loops: +The structural idiom is equivalent and composes with guards/loops: ``` ,x/.+/ s/^/> / prefix every (non-empty) line → > a⏎> b⏎> c ,x/.+/ a/;/ append ";" to every line → a;⏎b;⏎c; ``` -`.` does **not** cross newlines (sam: a "character" is "any character but -newline"), which is also PCRE's default. Line-structured descent like -`,x/.*\n/ …` therefore behaves as expected. +## Composing edits — the sam way ---- +sam's power is structural composition, not clever single regexes. Descend into +matches with `x`/`y`, guard with `g`/`v`, group with `{}`: -## What matches identically (no surprises) +``` +,x/[a-zA-Z_]+/ g/^[A-Z]/ c/CONST/ every identifier starting uppercase → CONST +,x/"[^"]*"/ s/foo/bar/g foo→bar only inside double-quoted strings +0/start/,/end/ x/[a-z]+/ s/.*/[&]/ bracket every lowercase word in a region +``` -| behaviour | katesam (PCRE) | sam | same? | -|---|---|---|---| -| `.` crosses newline | no | no | ✅ | -| `[^z]+` crosses newline | yes | yes | ✅ | -| empty-match advance `s/x*/-/g` on `axbx` → `-a-b-` | ✅ | ✅ | ✅ | -| greedy quantifiers `* + ?` | ✅ | ✅ | ✅ | -| character classes `[a-z]`, groups `( )`, `|` | ✅ | ✅ | ✅ | -| `&` whole-match, `\1..\9` groups in replacement | ✅ | ✅ | ✅ | +Multi-file `X`/`Y` applies an inner program across the project: -> `.` not crossing newlines matches sam, so line-structured descent -> (`,x/.*\n/ ...`) behaves as expected. Use `(?s)` if you *want* `.` to span -> newlines. +``` +X/\.cpp$/ ,s/old_api/new_api/g rewrite every .cpp in the project +Y/_test\./ ,x/TODO/ d drop TODO lines in non-test files +``` ---- - -## Extras — available, but not the sam way - -PCRE accepts syntax plan9 `regexp(7)` never had. It all works, but treat it as -an **escape hatch, not a headline**: sam's power comes from *composing* simple -patterns with `x`/`y`/`g`/`v`/`{}`, not from clever single regexes. Prefer -structural composition; reach for these only when it is genuinely simpler. - -| extra | example | note | -|---|---|---| -| shorthand classes `\d \w \s` | `,s/\d+/N/g` | less typing than `[0-9]`; genuinely handy | -| word boundary `\b` | `,s/\bfoo\b/X/g` | handy; sam would use `x/foo/` with context | -| ignore-case `(?i)` | `,s/(?i)todo/DONE/g` | fills a real gap — sam has no case-insensitive match | -| non-greedy `*?` | `,s/a.*?c/X/` | the structural `x/…/` loop is the sam alternative | -| lookahead/behind | `,s/foo(?=bar)/X/` | un-sam; express context with `x`/`g`/`v` instead | -| backreference in pattern | `,s/(\w)\1/D/g` | rarely the right tool | - -A pattern using these is **not portable back to sam**. If portability or -staying in the structural idiom matters, avoid them. - ---- +(The file set is the project index, matched by path. Edits land in buffers, not +on disk — review and save deliberately. Undo is per file for `X`/`Y`.) ## Practical guidance -- **Order overlapping alternatives longest-first** (`\.tar\.gz|\.gz`, not the - reverse). The only silent divergence from sam. -- **`^`/`$` are per-line** — just like sam. For whole-line edits either anchor - directly (`,s/^/> /g`) or, more sam-idiomatically, loop (`,x/.+/ …`). -- **Prefer structural composition** (`x g v {}`) over PCRE extras. The extras - work but pull you out of the sam idiom and are not sam-portable. -- **Empty-match globals** (`s/x*/…/g`) advance one position per null match, as in - sam — safe, but as always with `*`, double-check the result. +- Use character classes, not PCRE shorthands: `[0-9]` not `\d`, `[a-zA-Z_]` not + `\w`. +- Prefer structural composition (`x g v {}`) over one dense pattern. +- `^`/`$` are per-line; anchor directly (`,s/^…`) or loop (`,x/.+/ …`). +- Empty-match globals (`s/x*/…/g`) advance one position per null match — safe, + but as always with `*`, check the result. -## Why PCRE and not regexp(7) +## Scope -`QRegularExpression` ships with Qt (already a dependency) and gives a richer, -familiar syntax for free. plan9 `regexp(7)` is a Rune-based NFA woven into sam's -own buffer types; using it would mean porting that engine. The tradeoff accepted -here: everyday fidelity plus extra power, at the cost of exact leftmost-longest -semantics on overlapping alternation. If that matters for your workflow, the -engine's regex calls are isolated in `SamEngine` and could be swapped for a -ported `regexp.c` later. +Supported: the full address language (`#n n 0 $ . '` `/re/ ?re?` compound +`+ - , ;`) and commands `a c i d s p = m t k x y g v {}` plus shell filters +`< > | !`. Out of scope (Kate manages files / its own undo): the multi-file menu +(`b B n D`), external file I/O (`e r w f`), `"re"` file-addressing, and sam's own +`u` (use Kate's Ctrl+Z). diff --git a/src/plugin/sampanel.cpp b/src/plugin/sampanel.cpp index 3db0ff8..6aeb581 100644 --- a/src/plugin/sampanel.cpp +++ b/src/plugin/sampanel.cpp @@ -56,8 +56,9 @@ SamPanel::SamPanel(QWidget *parent) auto *hint = new QLabel( i18n("sam commands — e.g. ,s/foo/bar/g or " "X/\\.cpp$/ ,s/old/new/g (project-wide). Ctrl+Return runs.
" - "Regex is PCRE: ^/$ are per-line (like sam); order " - "overlapping alternatives longest-first. See docs/SAM.md."), + "Regex is sam's own (leftmost-longest, linear-time): classes like " + "[0-9], not \\d; ^/$ are per-line. " + "See docs/SAM.md."), this); hint->setWordWrap(true); hint->setTextFormat(Qt::RichText); diff --git a/src/sam/CMakeLists.txt b/src/sam/CMakeLists.txt index 797bacd..6b7c70a 100644 --- a/src/sam/CMakeLists.txt +++ b/src/sam/CMakeLists.txt @@ -8,6 +8,8 @@ find_package(Qt6 ${QT_MIN_VERSION} COMPONENTS Core REQUIRED) add_library(sam_lib STATIC samengine.cpp samengine.h + samregex.cpp + samregex.h ) target_link_libraries(sam_lib PUBLIC Qt6::Core) target_include_directories(sam_lib PUBLIC ${CMAKE_CURRENT_SOURCE_DIR}) @@ -17,4 +19,8 @@ if(Qt6Test_FOUND) add_executable(test_samengine test_samengine.cpp) target_link_libraries(test_samengine PRIVATE sam_lib Qt6::Test) add_test(NAME samengine COMMAND test_samengine) + + add_executable(test_samregex test_samregex.cpp) + target_link_libraries(test_samregex PRIVATE sam_lib Qt6::Test) + add_test(NAME samregex COMMAND test_samregex) endif() diff --git a/src/sam/samengine.cpp b/src/sam/samengine.cpp index ef25937..5514da8 100644 --- a/src/sam/samengine.cpp +++ b/src/sam/samengine.cpp @@ -18,9 +18,9 @@ * dot bound to each match, exactly as xec.c looper()/linelooper(). */ #include "samengine.h" +#include "samregex.h" #include -#include #include namespace katecustom @@ -656,16 +656,11 @@ private: return line; } - // Build a QRegularExpression from a sam pattern. The empty pattern reuses - // the last compiled pattern (sam behaviour). - // - // Multiline is ON: plan9 regexp(7) defines ^ as "beginning of a line" and $ - // as "end of a line" (sam regexp.c BOL = p==0 || prev=='\n'; EOL = next - // char is '\n'). That is exactly QRegularExpression::MultilineOption, so it - // is the faithful default, not an opt-in. '.' still does not cross newlines - // (sam: "the word character means any character but newline"), which is - // QRegularExpression's default (no DotMatchesEverythingOption). - QRegularExpression compile(const QString &pat) + // Build a SamRegex from a sam pattern. The empty pattern reuses the last + // compiled pattern (sam behaviour). This is the ported plan9 sam engine + // (SamRegex), so matching is leftmost-LONGEST and linear-time, ^/$ are + // per-line and '.' excludes newline — all intrinsic, matching real sam. + SamRegex compile(const QString &pat) { QString p = pat; if (p.isEmpty()) { @@ -673,9 +668,9 @@ private: } else { m_lastRe = p; } - QRegularExpression re(p, QRegularExpression::MultilineOption); - if (!re.isValid()) { - fail(QStringLiteral("bad regexp: %1").arg(re.errorString())); + SamRegex re; + if (!re.compile(p)) { + fail(QStringLiteral("bad regexp: %1").arg(re.error())); } return re; } @@ -861,34 +856,33 @@ private: // Forward (sign>0) / backward (sign<0) search, wrapping, like nextmatch(). Range search(const QString &pat, Range a, int sign) { - QRegularExpression re = compile(pat); + SamRegex re = compile(pat); if (failed()) { return a; } if (sign >= 0) { - int from = a.p2; - QRegularExpressionMatch m = re.match(m_text, from); - if (!m.hasMatch()) { - m = re.match(m_text, 0); // wrap + auto m = re.match(m_text, a.p2, m_nc); + if (!m) { + m = re.match(m_text, 0, m_nc); // wrap } - if (!m.hasMatch()) { + if (!m) { fail(QStringLiteral("search failed: %1").arg(pat.isEmpty() ? m_lastRe : pat)); return a; } - return {int(m.capturedStart()), int(m.capturedEnd())}; + return {m->start(), m->end()}; } else { - // Backward: find the last match that ends at or before a.p1, - // else wrap to the last match in the text. + // Backward: the last match that ends at or before a.p1, else wrap to + // the last match in the whole text. int limit = a.p1; Range best{-1, -1}; int from = 0; while (true) { - QRegularExpressionMatch m = re.match(m_text, from); - if (!m.hasMatch()) { + auto m = re.match(m_text, from, m_nc); + if (!m) { break; } - const int s = int(m.capturedStart()); - const int e = int(m.capturedEnd()); + const int s = m->start(); + const int e = m->end(); if (e <= limit) { best = {s, e}; } else { @@ -897,17 +891,14 @@ private: from = (e > s) ? e : e + 1; } if (best.p1 < 0) { - // wrap: take the very last match in the whole text from = 0; while (true) { - QRegularExpressionMatch m = re.match(m_text, from); - if (!m.hasMatch()) { + auto m = re.match(m_text, from, m_nc); + if (!m) { break; } - best = {int(m.capturedStart()), int(m.capturedEnd())}; - const int s = best.p1; - const int e = best.p2; - from = (e > s) ? e : e + 1; + best = {m->start(), m->end()}; + from = (best.p2 > best.p1) ? best.p2 : best.p2 + 1; } } if (best.p1 < 0) { @@ -921,7 +912,7 @@ private: // s/re/text/ with count and g flag (xec.c s_cmd()). void execSubstitute(Cmd *cmd, Range r) { - QRegularExpression re = compile(cmd->re); + SamRegex re = compile(cmd->re); if (failed()) { return; } @@ -931,15 +922,12 @@ private: int lastEnd = r.p1; int p1 = r.p1; while (p1 <= r.p2) { - QRegularExpressionMatch m = re.match(m_text, p1); - if (!m.hasMatch() || int(m.capturedStart()) > r.p2) { - break; - } - int ms = int(m.capturedStart()); - int me = int(m.capturedEnd()); - if (me > r.p2) { + auto m = re.match(m_text, p1, r.p2); + if (!m) { break; } + int ms = m->start(); + int me = m->end(); // advance logic (empty match handling) if (ms == me) { if (ms == op) { @@ -956,7 +944,7 @@ private: continue; } - const QString repl = expandReplacement(cmd->text, m); + const QString repl = expandReplacement(cmd->text, *m); addEdit(ms, me, repl); if (failed()) { return; @@ -975,7 +963,7 @@ private: } // Expand & and \1..\9 in an s replacement using the match captures. - QString expandReplacement(const QString &tmpl, const QRegularExpressionMatch &m) + QString expandReplacement(const QString &tmpl, const SamMatch &m) { QString out; for (int i = 0; i < tmpl.size(); i++) { @@ -984,7 +972,7 @@ private: const QChar d = tmpl[i + 1]; if (d.isDigit()) { const int g = d.unicode() - '0'; - out.append(m.captured(g)); + out.append(m.captured(g, m_text)); i++; continue; } @@ -1003,7 +991,7 @@ private: continue; } if (c == QLatin1Char('&')) { - out.append(m.captured(0)); + out.append(m.captured(0, m_text)); continue; } out.append(c); @@ -1046,14 +1034,13 @@ private: m_nest++; if (c == QLatin1Char('g') || c == QLatin1Char('v')) { - QRegularExpression re = compile(cmd->re); + SamRegex re = compile(cmd->re); if (failed()) { m_nest--; return; } - QRegularExpressionMatch m = re.match(m_text, r.p1); - const bool contains = m.hasMatch() && int(m.capturedStart()) <= r.p2 - && int(m.capturedStart()) >= r.p1; + auto m = re.match(m_text, r.p1, r.p2); + const bool contains = m.has_value(); const bool run = (c == QLatin1Char('g')) ? contains : !contains; if (run) { runLoopBody(cmd, r); @@ -1065,7 +1052,7 @@ private: // x and y. if (c == QLatin1Char('y')) { // y: run on the gaps between matches of re within r. - QRegularExpression re = compile(cmd->re); + SamRegex re = compile(cmd->re); if (failed()) { m_nest--; return; @@ -1073,12 +1060,12 @@ private: int prev = r.p1; int from = r.p1; while (from <= r.p2) { - QRegularExpressionMatch m = re.match(m_text, from); - if (!m.hasMatch() || int(m.capturedEnd()) > r.p2) { + auto m = re.match(m_text, from, r.p2); + if (!m) { break; } - int ms = int(m.capturedStart()); - int me = int(m.capturedEnd()); + int ms = m->start(); + int me = m->end(); runLoopBody(cmd, {prev, ms}); if (failed()) { m_nest--; @@ -1093,7 +1080,7 @@ private: } // x: run on each match of re within r (default re handled by compile). - QRegularExpression re = compile(cmd->re); + SamRegex re = compile(cmd->re); if (failed()) { m_nest--; return; @@ -1101,12 +1088,12 @@ private: int op = -1; int p = r.p1; while (p <= r.p2) { - QRegularExpressionMatch m = re.match(m_text, p); - if (!m.hasMatch() || int(m.capturedEnd()) > r.p2) { + auto m = re.match(m_text, p, r.p2); + if (!m) { break; } - int ms = int(m.capturedStart()); - int me = int(m.capturedEnd()); + int ms = m->start(); + int me = m->end(); if (ms == me) { if (ms == op) { p++; diff --git a/src/sam/samregex.cpp b/src/sam/samregex.cpp new file mode 100644 index 0000000..c9b6201 --- /dev/null +++ b/src/sam/samregex.cpp @@ -0,0 +1,632 @@ +/* + * SPDX-License-Identifier: LGPL-2.0-or-later + * + * Port of plan9port src/cmd/sam/regexp.c. Structure and names follow the + * original closely; pointers become indices into QVector. + */ +#include "samregex.h" + +namespace katecustom +{ + +// Action codes (regexp.c). A literal rune uses its own code point as `type` +// (always < kOperator for the characters we accept). Operators carry the +// kOperator bit; the low bits encode precedence. Tokens carry the kAny bit. +enum { + kOperator = 0x1000000, + START = kOperator + 0, + RBRA = kOperator + 1, // ) + LBRA = kOperator + 2, // ( + OR = kOperator + 3, // | + CAT = kOperator + 4, // implicit concatenation + STAR = kOperator + 5, // * + PLUS = kOperator + 6, // + + QUEST = kOperator + 7, // ? + + kAny = 0x2000000, + ANY = kAny + 0, // . + NOP = kAny + 1, // epsilon, removed by optimize() + BOL = kAny + 2, // ^ + EOL = kAny + 3, // $ + CCLASS = kAny + 4, // [...] + NCCLASS = kAny + 5, // [^...] + END = kAny + 0x77, // match +}; + +constexpr int NSUBEXP = 10; + +QString SamMatch::captured(int group, const QString &text) const +{ + if (group < 0 || group >= caps.size()) { + return QString(); + } + const Range r = caps[group]; + if (r.p1 < 0 || r.p2 < 0 || r.p2 < r.p1) { + return QString(); + } + return text.mid(r.p1, r.p2 - r.p1); +} + +// --------------------------------------------------------------------------- +// Compiler — shunting yard over the pattern, emitting Inst nodes. +// --------------------------------------------------------------------------- +class SamRegexCompiler +{ +public: + SamRegexCompiler(SamRegex &re, const QString &pat) : m_re(re), m_s(pat) {} + + bool compile() + { + m_re.m_prog.clear(); + m_re.m_classes.clear(); + m_atorStack.clear(); + m_subidStack.clear(); + m_andStack.clear(); + m_cursubid = 0; + m_lastWasAnd = false; + m_pos = 0; + m_nbra = 0; + + pushator(START - 1); + int token; + while ((token = lex()) != END) { + if (m_error.isEmpty() == false) { + return fail(); + } + if ((token & kOperator) == kOperator) { + doOperator(token); + } else { + doOperand(token); + } + if (!m_error.isEmpty()) { + return fail(); + } + } + evaluntil(START); + doOperand(END); + evaluntil(START); + if (m_nbra) { + m_error = QStringLiteral("unmatched '('"); + return fail(); + } + if (!m_error.isEmpty() || m_andStack.isEmpty()) { + if (m_error.isEmpty()) { + m_error = QStringLiteral("malformed regexp"); + } + return fail(); + } + m_re.m_start = m_andStack.last().first; + optimize(); + return true; + } + + QString error() const { return m_error; } + +private: + struct Node { + int first = -1; + int last = -1; + }; + + bool fail() + { + m_re.m_valid = false; + if (m_re.m_error.isEmpty()) { + m_re.m_error = m_error.isEmpty() ? QStringLiteral("regexp error") : m_error; + } + return false; + } + + int newinst(int t) + { + SamRegex::Inst inst; + inst.type = t; + m_re.m_prog.append(inst); + return m_re.m_prog.size() - 1; + } + + SamRegex::Inst &at(int i) { return m_re.m_prog[i]; } + + void pushand(int f, int l) { m_andStack.append(Node{f, l}); } + void pushator(int t) + { + m_atorStack.append(t); + m_subidStack.append(m_cursubid >= NSUBEXP ? -1 : m_cursubid); + } + int popator() + { + if (m_atorStack.isEmpty()) { + m_error = QStringLiteral("operator stack underflow"); + return START; + } + // The subid paired with this operator — captured before removal so the + // LBRA/RBRA case can read the group's own id (regexp.c reads *subidp + // right after decrementing it). + m_lastPoppedSubid = m_subidStack.isEmpty() ? -1 : m_subidStack.last(); + if (!m_subidStack.isEmpty()) { + m_subidStack.removeLast(); + } + const int t = m_atorStack.last(); + m_atorStack.removeLast(); + return t; + } + Node popand(int op) + { + if (m_andStack.isEmpty()) { + m_error = op ? QStringLiteral("missing operand for '%1'").arg(QChar(op)) + : QStringLiteral("malformed regexp"); + return Node{}; + } + const Node n = m_andStack.last(); + m_andStack.removeLast(); + return n; + } + + void doOperand(int t) + { + if (m_lastWasAnd) { + doOperator(CAT); // implicit concatenation + } + int i = newinst(t); + if (t == CCLASS) { + if (m_negateClass) { + at(i).type = NCCLASS; + } + at(i).classIdx = m_re.m_classes.size() - 1; + } + pushand(i, i); + m_lastWasAnd = true; + } + + void doOperator(int t) + { + if (t == RBRA && --m_nbra < 0) { + m_error = QStringLiteral("unmatched ')'"); + return; + } + if (t == LBRA) { + m_cursubid++; + m_nbra++; + if (m_lastWasAnd) { + doOperator(CAT); + } + } else { + evaluntil(t); + } + if (t != RBRA) { + pushator(t); + } + m_lastWasAnd = false; + if (t == STAR || t == QUEST || t == PLUS || t == RBRA) { + m_lastWasAnd = true; + } + } + + void evaluntil(int pri) + { + while (m_error.isEmpty() + && (pri == RBRA || (!m_atorStack.isEmpty() && m_atorStack.last() >= pri))) { + const int op = popator(); + if (!m_error.isEmpty()) { + return; + } + if (op == LBRA) { + Node op1 = popand('('); + const int subid = m_lastPoppedSubid; + int inst2 = newinst(RBRA); + at(inst2).subid = subid; + at(op1.last).next = inst2; + int inst1 = newinst(LBRA); + at(inst1).subid = subid; + at(inst1).next = op1.first; + pushand(inst1, inst2); + return; // must have been RBRA + } else if (op == OR) { + Node op2 = popand('|'); + Node op1 = popand('|'); + if (!m_error.isEmpty()) { + return; + } + int inst2 = newinst(NOP); + at(op2.last).next = inst2; + at(op1.last).next = inst2; + int inst1 = newinst(OR); + at(inst1).right = op1.first; + at(inst1).next = op2.first; // "left" branch == continuation slot + pushand(inst1, inst2); + } else if (op == CAT) { + Node op2 = popand(0); + Node op1 = popand(0); + if (!m_error.isEmpty()) { + return; + } + at(op1.last).next = op2.first; + pushand(op1.first, op2.last); + } else if (op == STAR) { + Node op2 = popand('*'); + if (!m_error.isEmpty()) { + return; + } + int inst1 = newinst(OR); + at(op2.last).next = inst1; + at(inst1).right = op2.first; + pushand(inst1, inst1); + } else if (op == PLUS) { + Node op2 = popand('+'); + if (!m_error.isEmpty()) { + return; + } + int inst1 = newinst(OR); + at(op2.last).next = inst1; + at(inst1).right = op2.first; + pushand(op2.first, inst1); + } else if (op == QUEST) { + Node op2 = popand('?'); + if (!m_error.isEmpty()) { + return; + } + int inst1 = newinst(OR); + int inst2 = newinst(NOP); + at(inst1).next = inst2; // skip branch (continuation slot) + at(inst1).right = op2.first; + at(op2.last).next = inst2; + pushand(inst1, inst2); + } else { + m_error = QStringLiteral("bad regexp operator"); + return; + } + if (pri == RBRA) { + // Keep looping until the matching LBRA (handled by the return + // above). The while condition handles the rest. + } + } + } + + // Rewrite `next` pointers to skip NOP chains (regexp.c optimize()). + void optimize() + { + for (int i = 0; i < m_re.m_prog.size(); i++) { + if (m_re.m_prog[i].type == END) { + continue; + } + int target = m_re.m_prog[i].next; + while (target >= 0 && m_re.m_prog[target].type == NOP) { + target = m_re.m_prog[target].next; + } + m_re.m_prog[i].next = target; + } + } + + // --- lexer --- + int lex() + { + if (m_pos >= m_s.size()) { + return END; + } + QChar qc = m_s[m_pos++]; + int c = qc.unicode(); + switch (c) { + case '\\': + if (m_pos < m_s.size()) { + QChar d = m_s[m_pos++]; + c = (d == QLatin1Char('n')) ? '\n' : d.unicode(); + } + // escaped: always a literal, never a metachar/token + return c; + case '*': + return STAR; + case '?': + return QUEST; + case '+': + return PLUS; + case '|': + return OR; + case '.': + return ANY; + case '(': + return LBRA; + case ')': + return RBRA; + case '^': + return BOL; + case '$': + return EOL; + case '[': + buildClass(); + return CCLASS; + default: + return c; + } + } + + // Read one class element, honouring \n and \. + int classNextRune(bool *quoted) + { + *quoted = false; + if (m_pos >= m_s.size()) { + m_error = QStringLiteral("malformed character class"); + return 0; + } + QChar c = m_s[m_pos]; + if (c == QLatin1Char('\\')) { + m_pos++; + if (m_pos >= m_s.size()) { + m_error = QStringLiteral("malformed character class"); + return 0; + } + QChar d = m_s[m_pos++]; + if (d == QLatin1Char('n')) { + return '\n'; + } + *quoted = true; + return d.unicode(); + } + m_pos++; + return c.unicode(); + } + + void buildClass() + { + QVector items; + m_negateClass = false; + if (m_pos < m_s.size() && m_s[m_pos] == QLatin1Char('^')) { + m_negateClass = true; + m_pos++; + // Negated classes never match newline. + items.append(SamRegex::ClassItem{false, '\n', '\n'}); + } + while (true) { + if (m_pos >= m_s.size()) { + m_error = QStringLiteral("unterminated character class"); + break; + } + // A ']' that is not escaped closes the class. + if (m_s[m_pos] == QLatin1Char(']')) { + m_pos++; + break; + } + bool q1 = false; + int c1 = classNextRune(&q1); + if (!m_error.isEmpty()) { + break; + } + // Range a-b: a '-' followed by a non-']' element. + if (m_pos + 0 < m_s.size() && m_s[m_pos] == QLatin1Char('-') + && m_pos + 1 < m_s.size() && m_s[m_pos + 1] != QLatin1Char(']')) { + m_pos++; // consume '-' + bool q2 = false; + int c2 = classNextRune(&q2); + if (!m_error.isEmpty()) { + break; + } + items.append(SamRegex::ClassItem{true, c1, c2}); + } else { + items.append(SamRegex::ClassItem{false, c1, c1}); + } + } + m_re.m_classes.append(items); + } + + SamRegex &m_re; + QString m_s; + int m_pos = 0; + + QVector m_atorStack; + QVector m_subidStack; + QVector m_andStack; + int m_cursubid = 0; + int m_lastPoppedSubid = -1; + bool m_lastWasAnd = false; + int m_nbra = 0; + bool m_negateClass = false; + QString m_error; +}; + +// --------------------------------------------------------------------------- +// SamRegex +// --------------------------------------------------------------------------- +bool SamRegex::compile(const QString &pattern) +{ + m_pattern = pattern; + m_valid = true; + m_error.clear(); + SamRegexCompiler c(*this, pattern); + if (!c.compile()) { + m_valid = false; + if (m_error.isEmpty()) { + m_error = c.error(); + } + return false; + } + m_valid = true; + return true; +} + +bool SamRegex::classMatch(int classIdx, QChar c, bool negate) const +{ + if (classIdx < 0 || classIdx >= m_classes.size()) { + return negate; + } + const int cc = c.unicode(); + for (const ClassItem &item : m_classes[classIdx]) { + if (item.range) { + if (item.lo <= cc && cc <= item.hi) { + return !negate; + } + } else if (item.lo == cc) { + return !negate; + } + } + return negate; +} + +// Pike VM thread list. Each live thread is (inst, captures). addinst dedups by +// instruction index within a list, keeping the thread whose match started +// earliest (leftmost), exactly as regexp.c addinst does. +namespace +{ +struct Thread { + int inst = -1; + QVector caps; +}; +} // namespace + +std::optional SamRegex::match(const QString &text, int from, int bound) const +{ + if (!m_valid || m_start < 0) { + return std::nullopt; + } + if (bound < 0 || bound > text.size()) { + bound = text.size(); + } + if (from < 0) { + from = 0; + } + + auto addinst = [](QVector &list, int inst, const QVector &caps) { + for (Thread &t : list) { + if (t.inst == inst) { + // Keep the earliest start (leftmost) — mirrors regexp.c. + if (caps[0].p1 < t.caps[0].p1) { + t.caps = caps; + } + return; + } + } + list.append(Thread{inst, caps}); + }; + + QVector clist; + QVector nlist; + + SamMatch best; + bool haveMatch = false; + auto newmatch = [&](const QVector &se) { + // Leftmost-longest: take if no match yet, or starts earlier, or same + // start and ends later (regexp.c newmatch()). + if (!haveMatch || se[0].p1 < best.caps[0].p1 + || (se[0].p1 == best.caps[0].p1 && se[0].p2 > best.caps[0].p2)) { + best.caps = se; + haveMatch = true; + } + }; + + const int startChar = + (m_prog[m_start].type < kOperator) ? m_prog[m_start].type : 0; + + int nnl = 0; // live threads carried into nlist + // Scan one position past bound so a thread ending exactly at bound (and + // EOL/END at bound) is processed. + for (int p = from; p <= bound; p++) { + const int c = (p < bound) ? text[p].unicode() : -1; + + // Stop once a match is found and no threads remain alive. + if (haveMatch && nnl == 0) { + break; + } + // Fast first-char skip while no thread is live and no match pending. + if (startChar && nnl == 0 && !haveMatch && c != startChar) { + continue; + } + + clist = nlist; + nlist.clear(); + nnl = 0; + + // Seed a fresh start thread at this position while no match found yet. + if (!haveMatch) { + QVector se(NSUBEXP, SamMatch::Range{-1, -1}); + se[0].p1 = p; + addinst(clist, m_start, se); + } + + // Run epsilon + consuming transitions for this position. + for (int ti = 0; ti < clist.size(); ti++) { + int inst = clist[ti].inst; + // Follow epsilon transitions within the same position. + bool consumed = false; + while (!consumed) { + const Inst &in = m_prog[inst]; + switch (in.type) { + case LBRA: + if (in.subid >= 0 && in.subid < NSUBEXP) { + clist[ti].caps[in.subid].p1 = p; + } + inst = in.next; + continue; + case RBRA: + if (in.subid >= 0 && in.subid < NSUBEXP) { + clist[ti].caps[in.subid].p2 = p; + } + inst = in.next; + continue; + case OR: + addinst(clist, in.right, clist[ti].caps); + inst = in.next; // "left"/continuation branch + continue; + case NOP: + inst = in.next; + continue; + case BOL: + if (p == 0 || (p > 0 && text[p - 1] == QLatin1Char('\n'))) { + inst = in.next; + continue; + } + consumed = true; // dead + break; + case EOL: + if (c == '\n' || p == bound) { + inst = in.next; + continue; + } + consumed = true; // dead + break; + case END: + clist[ti].caps[0].p2 = p; + newmatch(clist[ti].caps); + consumed = true; + break; + case ANY: + if (c >= 0 && c != '\n') { + addinst(nlist, in.next, clist[ti].caps); + nnl++; + } + consumed = true; + break; + case CCLASS: + if (c >= 0 && classMatch(in.classIdx, QChar(c), false)) { + addinst(nlist, in.next, clist[ti].caps); + nnl++; + } + consumed = true; + break; + case NCCLASS: + if (c >= 0 && classMatch(in.classIdx, QChar(c), true)) { + addinst(nlist, in.next, clist[ti].caps); + nnl++; + } + consumed = true; + break; + default: // literal rune + if (c >= 0 && in.type == c) { + addinst(nlist, in.next, clist[ti].caps); + nnl++; + } + consumed = true; + break; + } + } + } + } + + if (!haveMatch) { + return std::nullopt; + } + // Honour the bound: a match must lie within [from, bound]. + if (best.caps[0].p2 > bound) { + return std::nullopt; + } + return best; +} + +} // namespace katecustom diff --git a/src/sam/samregex.h b/src/sam/samregex.h new file mode 100644 index 0000000..3f3d15f --- /dev/null +++ b/src/sam/samregex.h @@ -0,0 +1,114 @@ +/* + * SPDX-License-Identifier: LGPL-2.0-or-later + * + * SamRegex — a faithful port of plan9 sam's regular-expression engine + * (plan9port src/cmd/sam/regexp.c) to operate on a QString. + * + * Why this exists: sam's regex is a Thompson/Pike NFA simulation. It is + * - leftmost-longest (POSIX semantics), not Perl leftmost-greedy — so it + * matches sam exactly, including overlapping alternation (`a|ab` on "ab" + * matches "ab", the longest), which PCRE does not; and + * - linear time in the subject length with no backtracking, so it cannot + * suffer catastrophic blowup. + * + * The dialect is sam's regexp(7), deliberately small: + * . * + ? | ( ) ^ $ [ ] and \ to escape a metacharacter (\n = newline). + * `^` matches the beginning of a line, `$` the end of a line (per-line, built + * in). `.` and negated classes do not match newline. There is no \d \w \b, + * no lookaround, no non-greedy — sam never had them, and their absence keeps + * the engine linear and the syntax portable back to real sam. + * + * The implementation mirrors regexp.c structure: a shunting-yard parser builds + * an Inst NFA; execute() runs two parallel thread lists (Pike VM) where each + * thread carries a capture set, and newmatch() keeps the leftmost-longest. + * Only the forward machine is ported; backward search is done by the caller + * scanning forward and keeping the last qualifying match. + */ +#ifndef KATECUSTOM_SAMREGEX_H +#define KATECUSTOM_SAMREGEX_H + +#include +#include +#include +#include + +namespace katecustom +{ + +/*! A match: half-open character ranges. caps[0] is the whole match; caps[1..9] + * are parenthesised subexpressions ({-1,-1} if that group did not participate). */ +struct SamMatch { + struct Range { + int p1 = -1; + int p2 = -1; + }; + QVector caps; // size NSUBEXP (10) + int start() const { return caps.isEmpty() ? -1 : caps[0].p1; } + int end() const { return caps.isEmpty() ? -1 : caps[0].p2; } + QString captured(int group, const QString &text) const; +}; + +class SamRegex +{ +public: + SamRegex() = default; + + /*! Compile \a pattern (sam dialect). Returns false and sets error() on a + * syntax error. */ + bool compile(const QString &pattern); + + bool isValid() const { return m_valid; } + QString error() const { return m_error; } + QString pattern() const { return m_pattern; } + + /*! + * Find the leftmost-longest match whose start is at or after \a from and + * which lies entirely within [0, \a bound). Anchors `^`/`$` are evaluated + * against the whole \a text (so they fire at any line boundary inside the + * bound). Returns std::nullopt if there is no match. Linear time. + */ + std::optional match(const QString &text, int from, int bound) const; + +private: + // NFA instruction (regexp.c Inst). type < kOperator is a literal rune; + // otherwise it is one of the action codes below. + // + // In the original, `left` and `next` share one union member (l.lleft == + // l.lnext): an OR node's "left/continuation" IS its next pointer, so + // concatenation (which wires `.next`) and OR execution (which reads the + // continuation) refer to the same slot. We keep that single `next` slot and + // use `right` for the OR's branch target. `subid`/`classIdx` reuse would be + // the r-union; they are used disjointly per node type so separate fields are + // safe. + struct Inst { + int type = 0; + int subid = -1; // LBRA/RBRA capture index + int classIdx = -1; // CCLASS/NCCLASS index into m_classes + int next = -1; // continuation (also the OR "left" branch) + int right = -1; // OR branch target (loop body / first alternative) + }; + + // A character class: a flat list of items, each either a single rune or a + // range [lo,hi]. Negated classes also implicitly exclude '\n'. + struct ClassItem { + bool range = false; + int lo = 0; + int hi = 0; + }; + + bool classMatch(int classIdx, QChar c, bool negate) const; + + QString m_pattern; + bool m_valid = false; + QString m_error; + + QVector m_prog; // the compiled program + int m_start = -1; // index of the start instruction + QVector> m_classes; + + friend class SamRegexCompiler; +}; + +} // namespace katecustom + +#endif diff --git a/src/sam/test_samengine.cpp b/src/sam/test_samengine.cpp index 60b893f..4aa45f3 100644 --- a/src/sam/test_samengine.cpp +++ b/src/sam/test_samengine.cpp @@ -56,6 +56,7 @@ private Q_SLOTS: void caretIsPerLine(); void dollarIsPerLine(); void samAnchorIdiom(); + void leftmostLongestSubstitution(); }; void TestSamEngine::substituteFirst() @@ -87,8 +88,8 @@ void TestSamEngine::substituteAmp() void TestSamEngine::substituteGroup() { - // \1 is the first capture group. - QCOMPARE(runAll(QStringLiteral(",s/(\\w+)=(\\w+)/\\2=\\1/"), + // \1 is the first capture group (sam dialect: use explicit classes, not \w). + QCOMPARE(runAll(QStringLiteral(",s/([a-z]+)=([a-z]+)/\\2=\\1/"), QStringLiteral("key=value")), QStringLiteral("value=key")); } @@ -282,5 +283,15 @@ void TestSamEngine::samAnchorIdiom() QStringLiteral("> a\n> b\n> c")); } +void TestSamEngine::leftmostLongestSubstitution() +{ + // With the ported sam engine, alternation is leftmost-LONGEST: a|ab on "ab" + // matches "ab", so the whole token is replaced (PCRE would match just "a"). + QCOMPARE(runAll(QStringLiteral(",s/a|ab/X/"), QStringLiteral("ab")), + QStringLiteral("X")); + QCOMPARE(runAll(QStringLiteral(",s/ab|a/X/"), QStringLiteral("ab")), + QStringLiteral("X")); +} + QTEST_MAIN(TestSamEngine) #include "test_samengine.moc" diff --git a/src/sam/test_samregex.cpp b/src/sam/test_samregex.cpp new file mode 100644 index 0000000..7a9eca2 --- /dev/null +++ b/src/sam/test_samregex.cpp @@ -0,0 +1,200 @@ +/* + * SPDX-License-Identifier: LGPL-2.0-or-later + * + * Unit tests for the ported sam regex engine (SamRegex). These pin the two + * properties that motivated the port: leftmost-LONGEST semantics (POSIX, sam), + * and the sam dialect (no PCRE extras). Behaviours cross-checked against + * plan9port src/cmd/sam/regexp.c and sam(1). + */ +#include "samregex.h" + +#include +#include + +using namespace katecustom; + +class TestSamRegex : public QObject +{ + Q_OBJECT + + // Match from 0 over the whole text; return "start,end" or "nil". + QString m(const QString &pat, const QString &text) + { + SamRegex re; + if (!re.compile(pat)) { + return QStringLiteral("ERR:%1").arg(re.error()); + } + auto r = re.match(text, 0, text.size()); + if (!r) { + return QStringLiteral("nil"); + } + return QStringLiteral("%1,%2").arg(r->start()).arg(r->end()); + } + + // Match and return the matched substring, or "nil". + QString ms(const QString &pat, const QString &text, int from = 0) + { + SamRegex re; + if (!re.compile(pat)) { + return QStringLiteral("ERR:%1").arg(re.error()); + } + auto r = re.match(text, from, text.size()); + if (!r) { + return QStringLiteral("nil"); + } + return text.mid(r->start(), r->end() - r->start()); + } + +private Q_SLOTS: + void literal(); + void leftmostLongestAlternation(); + void alternationEitherOrder(); + void star(); + void plusQuest(); + void dotNotNewline(); + void anchorsPerLine(); + void charClass(); + void negatedClass(); + void range(); + void captures(); + void escapes(); + void fromOffset(); + void noMatch(); + void nestedGroupsLongest(); + void linearNoCatastrophicBacktrack(); +}; + +void TestSamRegex::literal() +{ + QCOMPARE(m(QStringLiteral("bar"), QStringLiteral("foobarbaz")), QStringLiteral("3,6")); +} + +void TestSamRegex::leftmostLongestAlternation() +{ + // THE motivating case: a|ab on "ab" must match "ab" (longest), unlike PCRE. + QCOMPARE(ms(QStringLiteral("a|ab"), QStringLiteral("ab")), QStringLiteral("ab")); +} + +void TestSamRegex::alternationEitherOrder() +{ + // Order must not matter for leftmost-longest. + QCOMPARE(ms(QStringLiteral("ab|a"), QStringLiteral("ab")), QStringLiteral("ab")); + QCOMPARE(ms(QStringLiteral("a|ab"), QStringLiteral("ab")), QStringLiteral("ab")); +} + +void TestSamRegex::star() +{ + QCOMPARE(ms(QStringLiteral("a*"), QStringLiteral("aaab")), QStringLiteral("aaa")); + // leftmost: at position 0 matches empty? a* matches "aaa" (longest at 0). + QCOMPARE(ms(QStringLiteral("ba*"), QStringLiteral("baaa")), QStringLiteral("baaa")); +} + +void TestSamRegex::plusQuest() +{ + QCOMPARE(ms(QStringLiteral("a+"), QStringLiteral("baaa")), QStringLiteral("aaa")); + QCOMPARE(ms(QStringLiteral("ab?c"), QStringLiteral("ac")), QStringLiteral("ac")); + QCOMPARE(ms(QStringLiteral("ab?c"), QStringLiteral("abc")), QStringLiteral("abc")); +} + +void TestSamRegex::dotNotNewline() +{ + // . does not cross newline (sam). + QCOMPARE(m(QStringLiteral("a.b"), QStringLiteral("a\nb")), QStringLiteral("nil")); + QCOMPARE(ms(QStringLiteral("a.c"), QStringLiteral("abc")), QStringLiteral("abc")); +} + +void TestSamRegex::anchorsPerLine() +{ + // ^ matches at start of any line; $ at end of any line. + SamRegex re; + QVERIFY(re.compile(QStringLiteral("^b"))); + auto r = re.match(QStringLiteral("a\nb\nc"), 0, 5); + QVERIFY(r.has_value()); + QCOMPARE(r->start(), 2); // 'b' on line 2 + QCOMPARE(r->end(), 3); + + SamRegex re2; + QVERIFY(re2.compile(QStringLiteral("a$"))); + auto r2 = re2.match(QStringLiteral("ba\nxa\n"), 0, 6); + QVERIFY(r2.has_value()); + QCOMPARE(r2->start(), 1); // first 'a' that precedes a newline + QCOMPARE(r2->end(), 2); +} + +void TestSamRegex::charClass() +{ + QCOMPARE(ms(QStringLiteral("[abc]+"), QStringLiteral("xcababz")), QStringLiteral("cabab")); +} + +void TestSamRegex::negatedClass() +{ + // [^z]+ crosses nothing special but stops at z; does not match newline. + QCOMPARE(ms(QStringLiteral("[^z]+"), QStringLiteral("abz")), QStringLiteral("ab")); + QCOMPARE(ms(QStringLiteral("[^z]+"), QStringLiteral("a\nb")), QStringLiteral("a")); +} + +void TestSamRegex::range() +{ + QCOMPARE(ms(QStringLiteral("[0-9]+"), QStringLiteral("ab123cd")), QStringLiteral("123")); + QCOMPARE(ms(QStringLiteral("[a-cx-z]+"), QStringLiteral("abcxyz!")), QStringLiteral("abcxyz")); +} + +void TestSamRegex::captures() +{ + SamRegex re; + QVERIFY(re.compile(QStringLiteral("([a-z]+)=([a-z]+)"))); + const QString text = QStringLiteral("key=value"); + auto r = re.match(text, 0, text.size()); + QVERIFY(r.has_value()); + QCOMPARE(r->captured(1, text), QStringLiteral("key")); + QCOMPARE(r->captured(2, text), QStringLiteral("value")); +} + +void TestSamRegex::escapes() +{ + // \* is a literal star; \. a literal dot; \n a newline. + QCOMPARE(ms(QStringLiteral("a\\*b"), QStringLiteral("xa*bx")), QStringLiteral("a*b")); + QCOMPARE(ms(QStringLiteral("a\\.b"), QStringLiteral("axb")), QStringLiteral("nil")); + QCOMPARE(ms(QStringLiteral("a\\.b"), QStringLiteral("a.b")), QStringLiteral("a.b")); + QCOMPARE(m(QStringLiteral("a\\nb"), QStringLiteral("a\nb")), QStringLiteral("0,3")); +} + +void TestSamRegex::fromOffset() +{ + // Search starting past the first match finds the second. + QCOMPARE(ms(QStringLiteral("a"), QStringLiteral("xaya"), 2), QStringLiteral("a")); + SamRegex re; + QVERIFY(re.compile(QStringLiteral("a"))); + auto r = re.match(QStringLiteral("xaya"), 2, 4); + QVERIFY(r.has_value()); + QCOMPARE(r->start(), 3); +} + +void TestSamRegex::noMatch() +{ + QCOMPARE(m(QStringLiteral("zzz"), QStringLiteral("abc")), QStringLiteral("nil")); +} + +void TestSamRegex::nestedGroupsLongest() +{ + // (a|ab)(c|bcd) on "abcd": leftmost-longest should match the whole "abcd" + // (a + bcd), which greedy-ordered backtracking would miss if it locked in + // "a" then "bcd"? Actually both can reach abcd; the key is longest overall. + QCOMPARE(ms(QStringLiteral("(a|ab)(c|bcd)"), QStringLiteral("abcd")), + QStringLiteral("abcd")); +} + +void TestSamRegex::linearNoCatastrophicBacktrack() +{ + // A pattern that makes PCRE backtrack exponentially: (a*)*b on a long run + // of 'a' with no 'b'. The NFA handles it in linear time; just assert it + // terminates quickly and reports no match. + const QString text(10000, QLatin1Char('a')); + SamRegex re; + QVERIFY(re.compile(QStringLiteral("(a*)*b"))); + auto r = re.match(text, 0, text.size()); + QVERIFY(!r.has_value()); +} + +QTEST_MAIN(TestSamRegex) +#include "test_samregex.moc"