diff --git a/docs/PLAN.md b/docs/PLAN.md
index 2ac84e0..605c416 100644
--- a/docs/PLAN.md
+++ b/docs/PLAN.md
@@ -359,6 +359,15 @@ A dockable tool view ("Sam", left sidebar) runs the plan9 **sam** command
language against the active document, or project-wide via `X`/`Y`. Type a
program, click **Run** (or Ctrl+Return); each Run is one undo step.
+- **Regex engine** (`src/sam/samregex.{h,cpp}`, 18 unit tests): a faithful port
+ of sam's own matcher (plan9port `src/cmd/sam/regexp.c`) — a Thompson/Pike NFA.
+ **Leftmost-longest** (POSIX), so it matches real sam exactly incl. overlapping
+ alternation (`a|ab` → `ab`); **linear time**, no backtracking (immune to
+ `(a*)*b` blowup). sam dialect only: `. * + ? | ( ) [ ] ^ $ \` with `\n`; `^`/`$`
+ are per-line and `.`/negated-classes exclude newline, all intrinsic. No PCRE
+ extras — not a deviation, just sam. Replaces the earlier QRegularExpression
+ matcher, eliminating the greedy-vs-longest divergence entirely.
+
- **Pure engine** (`src/sam/samengine.{h,cpp}`, unit-tested, no Kate dep):
`SamEngine::run(program, text, dotStart, dotEnd) -> SamResult{edits, dot,
output, applied}`. Edits are computed against the ORIGINAL snapshot in
@@ -383,14 +392,9 @@ program, click **Run** (or Ctrl+Return); each Run is one undo step.
- **Panel** (`src/plugin/sampanel.{h,cpp}`): program editor + Run button +
output log; `OllieView` creates the tool view and owns the apply logic
(`runSamProgram`, `applySamToDocument`, `runSamFileLoop`).
-- **Deviations (documented)**: regexes use `QRegularExpression` (PCRE), not
- plan9 `regexp(7)`. Everyday patterns match identically; the one semantic
- divergence is leftmost-greedy vs sam's leftmost-longest (bites only on
- overlapping alternation like `a|ab`). `^`/`$` are per-line (compiled with
- `MultilineOption`) and `.` does not cross newlines — both faithful to sam's
- `regexp(7)` (BOL/EOL in `regexp.c`). PCRE also accepts extras sam lacks
- (`\d \w \b`, lookahead, non-greedy, inline flags); the docs frame these as an
- escape hatch, not a feature — prefer structural composition. Full user-facing
+- **Dialect & semantics**: the regex engine is sam's own (see the Regex engine
+ bullet above), so matching is leftmost-longest and the dialect is sam's — no
+ PCRE, no greedy-vs-longest divergence, no PCRE extras. Full user-facing
treatment with verified examples in **`docs/SAM.md`**. Out of scope: multi-file
menu (`b B n D`), external file I/O (`e r w f`), `"re"` file-addressing, and
sam's own `u` (Kate's undo stack is the undo mechanism). The live GUI panel
diff --git a/docs/SAM.md b/docs/SAM.md
index 21269ee..aec51a1 100644
--- a/docs/SAM.md
+++ b/docs/SAM.md
@@ -6,126 +6,103 @@ against the active document, or project-wide via `X`/`Y`. Type a program, press
**Ctrl+Return** or click **Run**. Each Run is one undo step (Ctrl+Z reverts the
whole Run; multi-file `X`/`Y` undoes per file).
-This document covers the **regex deviation** — katesam uses PCRE
-(`QRegularExpression`), not plan9 `regexp(7)` — and what that means in practice.
-Every example below is verified against the engine.
+## The regex engine
-## TL;DR
+katesam uses a **faithful port of sam's own regular-expression engine**
+(`src/sam/samregex.{h,cpp}`, from plan9port `src/cmd/sam/regexp.c`), not PCRE.
+This is a Thompson/Pike NFA, which gives the two properties sam relies on:
-- Everyday patterns behave **identically** to sam.
-- One real semantic difference: **PCRE is leftmost-greedy, sam is
- leftmost-longest.** It only bites on *overlapping alternation* (`a|ab`).
-- `^` and `$` are **per-line** (match at every line boundary), the same as sam's
- `regexp(7)`. `.` does not cross newlines, also the same as sam.
-- PCRE *also* offers things sam lacks (`\d \w \b`, lookahead, non-greedy, inline
- flags). They work, but prefer sam-style structural composition (`x g v`) — see
- "Extras" below.
+- **Leftmost-longest** (POSIX) matching, so results match real sam exactly,
+ including overlapping alternation.
+- **Linear time**, with no backtracking — immune to catastrophic blowup. A
+ pattern like `(a*)*b` against thousands of `a` returns instantly.
----
+There is **no PCRE deviation anymore**: the dialect, the match semantics, and
+the anchors are sam's.
-## The one semantic divergence: greedy vs longest
+### Dialect
-sam's `regexp(7)` matches the **leftmost-longest** substring. PCRE matches
-**leftmost**, then resolves alternation **left-to-right, first wins** (greedy
-within a branch, but alternation order decides between branches).
+The sam dialect is deliberately small (sam `regexp(7)`):
-| program | input | sam (leftmost-longest) | **katesam (PCRE)** |
-|---|---|---|---|
-| `,s/a\|ab/X/` | `ab` | `X` (matches `ab`) | **`Xb`** (matches `a`) |
-| `,s/ab\|a/X/` | `ab` | `X` | `X` |
+```
+. any character except newline
+* + ? zero-or-more / one-or-more / zero-or-one
+| alternation
+( ) grouping and capture (\1..\9 in replacements)
+[ ] [^ ] character class / negated class (ranges a-z)
+^ $ beginning / end of a LINE
+\c literal c (escapes a metacharacter); \n is newline
+```
-**Implication:** when alternatives overlap, **order them longest-first**
-(`ab|a`, not `a|ab`). This is the only case where a working sam command can
-silently produce a different edit in katesam. Non-overlapping alternation
-(`cat|dog`) and all non-alternation patterns are unaffected.
+There is intentionally **no** `\d \w \s \b`, no lookaround, no non-greedy, and
+no inline flags. sam never had them. Use character classes (`[0-9]`,
+`[a-zA-Z_]`) and structural composition (`x y g v`) instead — that is the sam
+way, and patterns stay portable to real sam.
-Quantifier greediness is the same in both (`a.*c` on `axxcxxc` → whole string in
-both). katesam additionally offers non-greedy `.*?` (a PCRE extra).
+### Semantics (all verified against the port)
----
+| behaviour | katesam | note |
+|---|---|---|
+| overlapping alternation `a\|ab` on `ab` | matches `ab` | leftmost-**longest**; order of branches does not matter |
+| `.` crosses newline | **no** | matches sam |
+| `^` / `$` | per-line | `^` = start of a line, `$` = end of a line |
+| `[^z]+` crosses newline | no | negated classes also exclude `\n`, as in sam |
+| empty-match global `s/x*/-/g` on `axbx` | `-a-b-` | one advance per null match |
+| `&` whole match, `\1..\9` groups in replacement | yes | up to 9 capture groups |
+| pathological `(a*)*b` | linear time | NFA, no backtracking |
-## Anchors `^` / `$` are per-line — same as sam
+### Anchors are per-line
-This is **not** a deviation. plan9 `regexp(7)` defines `^` as "the beginning of
-a line" and `$` as "the end of a line" (sam's `regexp.c`: `BOL` fires at offset
-0 or after a `\n`; `EOL` fires before a `\n`). katesam compiles every pattern
-with `QRegularExpression::MultilineOption`, so `^`/`$` match at every line
-boundary, exactly as sam does:
+`^` matches the beginning of any line and `$` the end of any line — built into
+the engine (sam `regexp.c`: `BOL` fires at offset 0 or after `\n`; `EOL` before
+`\n`). So:
| program | input | result |
|---|---|---|
-| `,s/^/> /g` | `a⏎b⏎c` | `> a⏎> b⏎> c` (every line) |
-| `,s/$/;/g` | `a⏎b⏎c` | `a;⏎b;⏎c;` (every line) |
+| `,s/^/> /g` | `a⏎b⏎c` | `> a⏎> b⏎> c` |
+| `,s/$/;/g` | `a⏎b⏎c` | `a;⏎b;⏎c;` |
-The sam-idiomatic structural form works too, and is preferable because it
-composes with guards and other loops:
+The structural idiom is equivalent and composes with guards/loops:
```
,x/.+/ s/^/> / prefix every (non-empty) line → > a⏎> b⏎> c
,x/.+/ a/;/ append ";" to every line → a;⏎b;⏎c;
```
-`.` does **not** cross newlines (sam: a "character" is "any character but
-newline"), which is also PCRE's default. Line-structured descent like
-`,x/.*\n/ …` therefore behaves as expected.
+## Composing edits — the sam way
----
+sam's power is structural composition, not clever single regexes. Descend into
+matches with `x`/`y`, guard with `g`/`v`, group with `{}`:
-## What matches identically (no surprises)
+```
+,x/[a-zA-Z_]+/ g/^[A-Z]/ c/CONST/ every identifier starting uppercase → CONST
+,x/"[^"]*"/ s/foo/bar/g foo→bar only inside double-quoted strings
+0/start/,/end/ x/[a-z]+/ s/.*/[&]/ bracket every lowercase word in a region
+```
-| behaviour | katesam (PCRE) | sam | same? |
-|---|---|---|---|
-| `.` crosses newline | no | no | ✅ |
-| `[^z]+` crosses newline | yes | yes | ✅ |
-| empty-match advance `s/x*/-/g` on `axbx` → `-a-b-` | ✅ | ✅ | ✅ |
-| greedy quantifiers `* + ?` | ✅ | ✅ | ✅ |
-| character classes `[a-z]`, groups `( )`, `|` | ✅ | ✅ | ✅ |
-| `&` whole-match, `\1..\9` groups in replacement | ✅ | ✅ | ✅ |
+Multi-file `X`/`Y` applies an inner program across the project:
-> `.` not crossing newlines matches sam, so line-structured descent
-> (`,x/.*\n/ ...`) behaves as expected. Use `(?s)` if you *want* `.` to span
-> newlines.
+```
+X/\.cpp$/ ,s/old_api/new_api/g rewrite every .cpp in the project
+Y/_test\./ ,x/TODO/ d drop TODO lines in non-test files
+```
----
-
-## Extras — available, but not the sam way
-
-PCRE accepts syntax plan9 `regexp(7)` never had. It all works, but treat it as
-an **escape hatch, not a headline**: sam's power comes from *composing* simple
-patterns with `x`/`y`/`g`/`v`/`{}`, not from clever single regexes. Prefer
-structural composition; reach for these only when it is genuinely simpler.
-
-| extra | example | note |
-|---|---|---|
-| shorthand classes `\d \w \s` | `,s/\d+/N/g` | less typing than `[0-9]`; genuinely handy |
-| word boundary `\b` | `,s/\bfoo\b/X/g` | handy; sam would use `x/foo/` with context |
-| ignore-case `(?i)` | `,s/(?i)todo/DONE/g` | fills a real gap — sam has no case-insensitive match |
-| non-greedy `*?` | `,s/a.*?c/X/` | the structural `x/…/` loop is the sam alternative |
-| lookahead/behind | `,s/foo(?=bar)/X/` | un-sam; express context with `x`/`g`/`v` instead |
-| backreference in pattern | `,s/(\w)\1/D/g` | rarely the right tool |
-
-A pattern using these is **not portable back to sam**. If portability or
-staying in the structural idiom matters, avoid them.
-
----
+(The file set is the project index, matched by path. Edits land in buffers, not
+on disk — review and save deliberately. Undo is per file for `X`/`Y`.)
## Practical guidance
-- **Order overlapping alternatives longest-first** (`\.tar\.gz|\.gz`, not the
- reverse). The only silent divergence from sam.
-- **`^`/`$` are per-line** — just like sam. For whole-line edits either anchor
- directly (`,s/^/> /g`) or, more sam-idiomatically, loop (`,x/.+/ …`).
-- **Prefer structural composition** (`x g v {}`) over PCRE extras. The extras
- work but pull you out of the sam idiom and are not sam-portable.
-- **Empty-match globals** (`s/x*/…/g`) advance one position per null match, as in
- sam — safe, but as always with `*`, double-check the result.
+- Use character classes, not PCRE shorthands: `[0-9]` not `\d`, `[a-zA-Z_]` not
+ `\w`.
+- Prefer structural composition (`x g v {}`) over one dense pattern.
+- `^`/`$` are per-line; anchor directly (`,s/^…`) or loop (`,x/.+/ …`).
+- Empty-match globals (`s/x*/…/g`) advance one position per null match — safe,
+ but as always with `*`, check the result.
-## Why PCRE and not regexp(7)
+## Scope
-`QRegularExpression` ships with Qt (already a dependency) and gives a richer,
-familiar syntax for free. plan9 `regexp(7)` is a Rune-based NFA woven into sam's
-own buffer types; using it would mean porting that engine. The tradeoff accepted
-here: everyday fidelity plus extra power, at the cost of exact leftmost-longest
-semantics on overlapping alternation. If that matters for your workflow, the
-engine's regex calls are isolated in `SamEngine` and could be swapped for a
-ported `regexp.c` later.
+Supported: the full address language (`#n n 0 $ . '` `/re/ ?re?` compound
+`+ - , ;`) and commands `a c i d s p = m t k x y g v {}` plus shell filters
+`< > | !`. Out of scope (Kate manages files / its own undo): the multi-file menu
+(`b B n D`), external file I/O (`e r w f`), `"re"` file-addressing, and sam's own
+`u` (use Kate's Ctrl+Z).
diff --git a/src/plugin/sampanel.cpp b/src/plugin/sampanel.cpp
index 3db0ff8..6aeb581 100644
--- a/src/plugin/sampanel.cpp
+++ b/src/plugin/sampanel.cpp
@@ -56,8 +56,9 @@ SamPanel::SamPanel(QWidget *parent)
auto *hint = new QLabel(
i18n("sam commands — e.g. ,s/foo/bar/g or "
"X/\\.cpp$/ ,s/old/new/g (project-wide). Ctrl+Return runs.
"
- "Regex is PCRE: ^/$ are per-line (like sam); order "
- "overlapping alternatives longest-first. See docs/SAM.md."),
+ "Regex is sam's own (leftmost-longest, linear-time): classes like "
+ "[0-9], not \\d; ^/$ are per-line. "
+ "See docs/SAM.md."),
this);
hint->setWordWrap(true);
hint->setTextFormat(Qt::RichText);
diff --git a/src/sam/CMakeLists.txt b/src/sam/CMakeLists.txt
index 797bacd..6b7c70a 100644
--- a/src/sam/CMakeLists.txt
+++ b/src/sam/CMakeLists.txt
@@ -8,6 +8,8 @@ find_package(Qt6 ${QT_MIN_VERSION} COMPONENTS Core REQUIRED)
add_library(sam_lib STATIC
samengine.cpp
samengine.h
+ samregex.cpp
+ samregex.h
)
target_link_libraries(sam_lib PUBLIC Qt6::Core)
target_include_directories(sam_lib PUBLIC ${CMAKE_CURRENT_SOURCE_DIR})
@@ -17,4 +19,8 @@ if(Qt6Test_FOUND)
add_executable(test_samengine test_samengine.cpp)
target_link_libraries(test_samengine PRIVATE sam_lib Qt6::Test)
add_test(NAME samengine COMMAND test_samengine)
+
+ add_executable(test_samregex test_samregex.cpp)
+ target_link_libraries(test_samregex PRIVATE sam_lib Qt6::Test)
+ add_test(NAME samregex COMMAND test_samregex)
endif()
diff --git a/src/sam/samengine.cpp b/src/sam/samengine.cpp
index ef25937..5514da8 100644
--- a/src/sam/samengine.cpp
+++ b/src/sam/samengine.cpp
@@ -18,9 +18,9 @@
* dot bound to each match, exactly as xec.c looper()/linelooper().
*/
#include "samengine.h"
+#include "samregex.h"
#include
-#include
#include
namespace katecustom
@@ -656,16 +656,11 @@ private:
return line;
}
- // Build a QRegularExpression from a sam pattern. The empty pattern reuses
- // the last compiled pattern (sam behaviour).
- //
- // Multiline is ON: plan9 regexp(7) defines ^ as "beginning of a line" and $
- // as "end of a line" (sam regexp.c BOL = p==0 || prev=='\n'; EOL = next
- // char is '\n'). That is exactly QRegularExpression::MultilineOption, so it
- // is the faithful default, not an opt-in. '.' still does not cross newlines
- // (sam: "the word character means any character but newline"), which is
- // QRegularExpression's default (no DotMatchesEverythingOption).
- QRegularExpression compile(const QString &pat)
+ // Build a SamRegex from a sam pattern. The empty pattern reuses the last
+ // compiled pattern (sam behaviour). This is the ported plan9 sam engine
+ // (SamRegex), so matching is leftmost-LONGEST and linear-time, ^/$ are
+ // per-line and '.' excludes newline — all intrinsic, matching real sam.
+ SamRegex compile(const QString &pat)
{
QString p = pat;
if (p.isEmpty()) {
@@ -673,9 +668,9 @@ private:
} else {
m_lastRe = p;
}
- QRegularExpression re(p, QRegularExpression::MultilineOption);
- if (!re.isValid()) {
- fail(QStringLiteral("bad regexp: %1").arg(re.errorString()));
+ SamRegex re;
+ if (!re.compile(p)) {
+ fail(QStringLiteral("bad regexp: %1").arg(re.error()));
}
return re;
}
@@ -861,34 +856,33 @@ private:
// Forward (sign>0) / backward (sign<0) search, wrapping, like nextmatch().
Range search(const QString &pat, Range a, int sign)
{
- QRegularExpression re = compile(pat);
+ SamRegex re = compile(pat);
if (failed()) {
return a;
}
if (sign >= 0) {
- int from = a.p2;
- QRegularExpressionMatch m = re.match(m_text, from);
- if (!m.hasMatch()) {
- m = re.match(m_text, 0); // wrap
+ auto m = re.match(m_text, a.p2, m_nc);
+ if (!m) {
+ m = re.match(m_text, 0, m_nc); // wrap
}
- if (!m.hasMatch()) {
+ if (!m) {
fail(QStringLiteral("search failed: %1").arg(pat.isEmpty() ? m_lastRe : pat));
return a;
}
- return {int(m.capturedStart()), int(m.capturedEnd())};
+ return {m->start(), m->end()};
} else {
- // Backward: find the last match that ends at or before a.p1,
- // else wrap to the last match in the text.
+ // Backward: the last match that ends at or before a.p1, else wrap to
+ // the last match in the whole text.
int limit = a.p1;
Range best{-1, -1};
int from = 0;
while (true) {
- QRegularExpressionMatch m = re.match(m_text, from);
- if (!m.hasMatch()) {
+ auto m = re.match(m_text, from, m_nc);
+ if (!m) {
break;
}
- const int s = int(m.capturedStart());
- const int e = int(m.capturedEnd());
+ const int s = m->start();
+ const int e = m->end();
if (e <= limit) {
best = {s, e};
} else {
@@ -897,17 +891,14 @@ private:
from = (e > s) ? e : e + 1;
}
if (best.p1 < 0) {
- // wrap: take the very last match in the whole text
from = 0;
while (true) {
- QRegularExpressionMatch m = re.match(m_text, from);
- if (!m.hasMatch()) {
+ auto m = re.match(m_text, from, m_nc);
+ if (!m) {
break;
}
- best = {int(m.capturedStart()), int(m.capturedEnd())};
- const int s = best.p1;
- const int e = best.p2;
- from = (e > s) ? e : e + 1;
+ best = {m->start(), m->end()};
+ from = (best.p2 > best.p1) ? best.p2 : best.p2 + 1;
}
}
if (best.p1 < 0) {
@@ -921,7 +912,7 @@ private:
// s/re/text/ with count and g flag (xec.c s_cmd()).
void execSubstitute(Cmd *cmd, Range r)
{
- QRegularExpression re = compile(cmd->re);
+ SamRegex re = compile(cmd->re);
if (failed()) {
return;
}
@@ -931,15 +922,12 @@ private:
int lastEnd = r.p1;
int p1 = r.p1;
while (p1 <= r.p2) {
- QRegularExpressionMatch m = re.match(m_text, p1);
- if (!m.hasMatch() || int(m.capturedStart()) > r.p2) {
- break;
- }
- int ms = int(m.capturedStart());
- int me = int(m.capturedEnd());
- if (me > r.p2) {
+ auto m = re.match(m_text, p1, r.p2);
+ if (!m) {
break;
}
+ int ms = m->start();
+ int me = m->end();
// advance logic (empty match handling)
if (ms == me) {
if (ms == op) {
@@ -956,7 +944,7 @@ private:
continue;
}
- const QString repl = expandReplacement(cmd->text, m);
+ const QString repl = expandReplacement(cmd->text, *m);
addEdit(ms, me, repl);
if (failed()) {
return;
@@ -975,7 +963,7 @@ private:
}
// Expand & and \1..\9 in an s replacement using the match captures.
- QString expandReplacement(const QString &tmpl, const QRegularExpressionMatch &m)
+ QString expandReplacement(const QString &tmpl, const SamMatch &m)
{
QString out;
for (int i = 0; i < tmpl.size(); i++) {
@@ -984,7 +972,7 @@ private:
const QChar d = tmpl[i + 1];
if (d.isDigit()) {
const int g = d.unicode() - '0';
- out.append(m.captured(g));
+ out.append(m.captured(g, m_text));
i++;
continue;
}
@@ -1003,7 +991,7 @@ private:
continue;
}
if (c == QLatin1Char('&')) {
- out.append(m.captured(0));
+ out.append(m.captured(0, m_text));
continue;
}
out.append(c);
@@ -1046,14 +1034,13 @@ private:
m_nest++;
if (c == QLatin1Char('g') || c == QLatin1Char('v')) {
- QRegularExpression re = compile(cmd->re);
+ SamRegex re = compile(cmd->re);
if (failed()) {
m_nest--;
return;
}
- QRegularExpressionMatch m = re.match(m_text, r.p1);
- const bool contains = m.hasMatch() && int(m.capturedStart()) <= r.p2
- && int(m.capturedStart()) >= r.p1;
+ auto m = re.match(m_text, r.p1, r.p2);
+ const bool contains = m.has_value();
const bool run = (c == QLatin1Char('g')) ? contains : !contains;
if (run) {
runLoopBody(cmd, r);
@@ -1065,7 +1052,7 @@ private:
// x and y.
if (c == QLatin1Char('y')) {
// y: run on the gaps between matches of re within r.
- QRegularExpression re = compile(cmd->re);
+ SamRegex re = compile(cmd->re);
if (failed()) {
m_nest--;
return;
@@ -1073,12 +1060,12 @@ private:
int prev = r.p1;
int from = r.p1;
while (from <= r.p2) {
- QRegularExpressionMatch m = re.match(m_text, from);
- if (!m.hasMatch() || int(m.capturedEnd()) > r.p2) {
+ auto m = re.match(m_text, from, r.p2);
+ if (!m) {
break;
}
- int ms = int(m.capturedStart());
- int me = int(m.capturedEnd());
+ int ms = m->start();
+ int me = m->end();
runLoopBody(cmd, {prev, ms});
if (failed()) {
m_nest--;
@@ -1093,7 +1080,7 @@ private:
}
// x: run on each match of re within r (default re handled by compile).
- QRegularExpression re = compile(cmd->re);
+ SamRegex re = compile(cmd->re);
if (failed()) {
m_nest--;
return;
@@ -1101,12 +1088,12 @@ private:
int op = -1;
int p = r.p1;
while (p <= r.p2) {
- QRegularExpressionMatch m = re.match(m_text, p);
- if (!m.hasMatch() || int(m.capturedEnd()) > r.p2) {
+ auto m = re.match(m_text, p, r.p2);
+ if (!m) {
break;
}
- int ms = int(m.capturedStart());
- int me = int(m.capturedEnd());
+ int ms = m->start();
+ int me = m->end();
if (ms == me) {
if (ms == op) {
p++;
diff --git a/src/sam/samregex.cpp b/src/sam/samregex.cpp
new file mode 100644
index 0000000..c9b6201
--- /dev/null
+++ b/src/sam/samregex.cpp
@@ -0,0 +1,632 @@
+/*
+ * SPDX-License-Identifier: LGPL-2.0-or-later
+ *
+ * Port of plan9port src/cmd/sam/regexp.c. Structure and names follow the
+ * original closely; pointers become indices into QVector.
+ */
+#include "samregex.h"
+
+namespace katecustom
+{
+
+// Action codes (regexp.c). A literal rune uses its own code point as `type`
+// (always < kOperator for the characters we accept). Operators carry the
+// kOperator bit; the low bits encode precedence. Tokens carry the kAny bit.
+enum {
+ kOperator = 0x1000000,
+ START = kOperator + 0,
+ RBRA = kOperator + 1, // )
+ LBRA = kOperator + 2, // (
+ OR = kOperator + 3, // |
+ CAT = kOperator + 4, // implicit concatenation
+ STAR = kOperator + 5, // *
+ PLUS = kOperator + 6, // +
+ QUEST = kOperator + 7, // ?
+
+ kAny = 0x2000000,
+ ANY = kAny + 0, // .
+ NOP = kAny + 1, // epsilon, removed by optimize()
+ BOL = kAny + 2, // ^
+ EOL = kAny + 3, // $
+ CCLASS = kAny + 4, // [...]
+ NCCLASS = kAny + 5, // [^...]
+ END = kAny + 0x77, // match
+};
+
+constexpr int NSUBEXP = 10;
+
+QString SamMatch::captured(int group, const QString &text) const
+{
+ if (group < 0 || group >= caps.size()) {
+ return QString();
+ }
+ const Range r = caps[group];
+ if (r.p1 < 0 || r.p2 < 0 || r.p2 < r.p1) {
+ return QString();
+ }
+ return text.mid(r.p1, r.p2 - r.p1);
+}
+
+// ---------------------------------------------------------------------------
+// Compiler — shunting yard over the pattern, emitting Inst nodes.
+// ---------------------------------------------------------------------------
+class SamRegexCompiler
+{
+public:
+ SamRegexCompiler(SamRegex &re, const QString &pat) : m_re(re), m_s(pat) {}
+
+ bool compile()
+ {
+ m_re.m_prog.clear();
+ m_re.m_classes.clear();
+ m_atorStack.clear();
+ m_subidStack.clear();
+ m_andStack.clear();
+ m_cursubid = 0;
+ m_lastWasAnd = false;
+ m_pos = 0;
+ m_nbra = 0;
+
+ pushator(START - 1);
+ int token;
+ while ((token = lex()) != END) {
+ if (m_error.isEmpty() == false) {
+ return fail();
+ }
+ if ((token & kOperator) == kOperator) {
+ doOperator(token);
+ } else {
+ doOperand(token);
+ }
+ if (!m_error.isEmpty()) {
+ return fail();
+ }
+ }
+ evaluntil(START);
+ doOperand(END);
+ evaluntil(START);
+ if (m_nbra) {
+ m_error = QStringLiteral("unmatched '('");
+ return fail();
+ }
+ if (!m_error.isEmpty() || m_andStack.isEmpty()) {
+ if (m_error.isEmpty()) {
+ m_error = QStringLiteral("malformed regexp");
+ }
+ return fail();
+ }
+ m_re.m_start = m_andStack.last().first;
+ optimize();
+ return true;
+ }
+
+ QString error() const { return m_error; }
+
+private:
+ struct Node {
+ int first = -1;
+ int last = -1;
+ };
+
+ bool fail()
+ {
+ m_re.m_valid = false;
+ if (m_re.m_error.isEmpty()) {
+ m_re.m_error = m_error.isEmpty() ? QStringLiteral("regexp error") : m_error;
+ }
+ return false;
+ }
+
+ int newinst(int t)
+ {
+ SamRegex::Inst inst;
+ inst.type = t;
+ m_re.m_prog.append(inst);
+ return m_re.m_prog.size() - 1;
+ }
+
+ SamRegex::Inst &at(int i) { return m_re.m_prog[i]; }
+
+ void pushand(int f, int l) { m_andStack.append(Node{f, l}); }
+ void pushator(int t)
+ {
+ m_atorStack.append(t);
+ m_subidStack.append(m_cursubid >= NSUBEXP ? -1 : m_cursubid);
+ }
+ int popator()
+ {
+ if (m_atorStack.isEmpty()) {
+ m_error = QStringLiteral("operator stack underflow");
+ return START;
+ }
+ // The subid paired with this operator — captured before removal so the
+ // LBRA/RBRA case can read the group's own id (regexp.c reads *subidp
+ // right after decrementing it).
+ m_lastPoppedSubid = m_subidStack.isEmpty() ? -1 : m_subidStack.last();
+ if (!m_subidStack.isEmpty()) {
+ m_subidStack.removeLast();
+ }
+ const int t = m_atorStack.last();
+ m_atorStack.removeLast();
+ return t;
+ }
+ Node popand(int op)
+ {
+ if (m_andStack.isEmpty()) {
+ m_error = op ? QStringLiteral("missing operand for '%1'").arg(QChar(op))
+ : QStringLiteral("malformed regexp");
+ return Node{};
+ }
+ const Node n = m_andStack.last();
+ m_andStack.removeLast();
+ return n;
+ }
+
+ void doOperand(int t)
+ {
+ if (m_lastWasAnd) {
+ doOperator(CAT); // implicit concatenation
+ }
+ int i = newinst(t);
+ if (t == CCLASS) {
+ if (m_negateClass) {
+ at(i).type = NCCLASS;
+ }
+ at(i).classIdx = m_re.m_classes.size() - 1;
+ }
+ pushand(i, i);
+ m_lastWasAnd = true;
+ }
+
+ void doOperator(int t)
+ {
+ if (t == RBRA && --m_nbra < 0) {
+ m_error = QStringLiteral("unmatched ')'");
+ return;
+ }
+ if (t == LBRA) {
+ m_cursubid++;
+ m_nbra++;
+ if (m_lastWasAnd) {
+ doOperator(CAT);
+ }
+ } else {
+ evaluntil(t);
+ }
+ if (t != RBRA) {
+ pushator(t);
+ }
+ m_lastWasAnd = false;
+ if (t == STAR || t == QUEST || t == PLUS || t == RBRA) {
+ m_lastWasAnd = true;
+ }
+ }
+
+ void evaluntil(int pri)
+ {
+ while (m_error.isEmpty()
+ && (pri == RBRA || (!m_atorStack.isEmpty() && m_atorStack.last() >= pri))) {
+ const int op = popator();
+ if (!m_error.isEmpty()) {
+ return;
+ }
+ if (op == LBRA) {
+ Node op1 = popand('(');
+ const int subid = m_lastPoppedSubid;
+ int inst2 = newinst(RBRA);
+ at(inst2).subid = subid;
+ at(op1.last).next = inst2;
+ int inst1 = newinst(LBRA);
+ at(inst1).subid = subid;
+ at(inst1).next = op1.first;
+ pushand(inst1, inst2);
+ return; // must have been RBRA
+ } else if (op == OR) {
+ Node op2 = popand('|');
+ Node op1 = popand('|');
+ if (!m_error.isEmpty()) {
+ return;
+ }
+ int inst2 = newinst(NOP);
+ at(op2.last).next = inst2;
+ at(op1.last).next = inst2;
+ int inst1 = newinst(OR);
+ at(inst1).right = op1.first;
+ at(inst1).next = op2.first; // "left" branch == continuation slot
+ pushand(inst1, inst2);
+ } else if (op == CAT) {
+ Node op2 = popand(0);
+ Node op1 = popand(0);
+ if (!m_error.isEmpty()) {
+ return;
+ }
+ at(op1.last).next = op2.first;
+ pushand(op1.first, op2.last);
+ } else if (op == STAR) {
+ Node op2 = popand('*');
+ if (!m_error.isEmpty()) {
+ return;
+ }
+ int inst1 = newinst(OR);
+ at(op2.last).next = inst1;
+ at(inst1).right = op2.first;
+ pushand(inst1, inst1);
+ } else if (op == PLUS) {
+ Node op2 = popand('+');
+ if (!m_error.isEmpty()) {
+ return;
+ }
+ int inst1 = newinst(OR);
+ at(op2.last).next = inst1;
+ at(inst1).right = op2.first;
+ pushand(op2.first, inst1);
+ } else if (op == QUEST) {
+ Node op2 = popand('?');
+ if (!m_error.isEmpty()) {
+ return;
+ }
+ int inst1 = newinst(OR);
+ int inst2 = newinst(NOP);
+ at(inst1).next = inst2; // skip branch (continuation slot)
+ at(inst1).right = op2.first;
+ at(op2.last).next = inst2;
+ pushand(inst1, inst2);
+ } else {
+ m_error = QStringLiteral("bad regexp operator");
+ return;
+ }
+ if (pri == RBRA) {
+ // Keep looping until the matching LBRA (handled by the return
+ // above). The while condition handles the rest.
+ }
+ }
+ }
+
+ // Rewrite `next` pointers to skip NOP chains (regexp.c optimize()).
+ void optimize()
+ {
+ for (int i = 0; i < m_re.m_prog.size(); i++) {
+ if (m_re.m_prog[i].type == END) {
+ continue;
+ }
+ int target = m_re.m_prog[i].next;
+ while (target >= 0 && m_re.m_prog[target].type == NOP) {
+ target = m_re.m_prog[target].next;
+ }
+ m_re.m_prog[i].next = target;
+ }
+ }
+
+ // --- lexer ---
+ int lex()
+ {
+ if (m_pos >= m_s.size()) {
+ return END;
+ }
+ QChar qc = m_s[m_pos++];
+ int c = qc.unicode();
+ switch (c) {
+ case '\\':
+ if (m_pos < m_s.size()) {
+ QChar d = m_s[m_pos++];
+ c = (d == QLatin1Char('n')) ? '\n' : d.unicode();
+ }
+ // escaped: always a literal, never a metachar/token
+ return c;
+ case '*':
+ return STAR;
+ case '?':
+ return QUEST;
+ case '+':
+ return PLUS;
+ case '|':
+ return OR;
+ case '.':
+ return ANY;
+ case '(':
+ return LBRA;
+ case ')':
+ return RBRA;
+ case '^':
+ return BOL;
+ case '$':
+ return EOL;
+ case '[':
+ buildClass();
+ return CCLASS;
+ default:
+ return c;
+ }
+ }
+
+ // Read one class element, honouring \n and \.
+ int classNextRune(bool *quoted)
+ {
+ *quoted = false;
+ if (m_pos >= m_s.size()) {
+ m_error = QStringLiteral("malformed character class");
+ return 0;
+ }
+ QChar c = m_s[m_pos];
+ if (c == QLatin1Char('\\')) {
+ m_pos++;
+ if (m_pos >= m_s.size()) {
+ m_error = QStringLiteral("malformed character class");
+ return 0;
+ }
+ QChar d = m_s[m_pos++];
+ if (d == QLatin1Char('n')) {
+ return '\n';
+ }
+ *quoted = true;
+ return d.unicode();
+ }
+ m_pos++;
+ return c.unicode();
+ }
+
+ void buildClass()
+ {
+ QVector items;
+ m_negateClass = false;
+ if (m_pos < m_s.size() && m_s[m_pos] == QLatin1Char('^')) {
+ m_negateClass = true;
+ m_pos++;
+ // Negated classes never match newline.
+ items.append(SamRegex::ClassItem{false, '\n', '\n'});
+ }
+ while (true) {
+ if (m_pos >= m_s.size()) {
+ m_error = QStringLiteral("unterminated character class");
+ break;
+ }
+ // A ']' that is not escaped closes the class.
+ if (m_s[m_pos] == QLatin1Char(']')) {
+ m_pos++;
+ break;
+ }
+ bool q1 = false;
+ int c1 = classNextRune(&q1);
+ if (!m_error.isEmpty()) {
+ break;
+ }
+ // Range a-b: a '-' followed by a non-']' element.
+ if (m_pos + 0 < m_s.size() && m_s[m_pos] == QLatin1Char('-')
+ && m_pos + 1 < m_s.size() && m_s[m_pos + 1] != QLatin1Char(']')) {
+ m_pos++; // consume '-'
+ bool q2 = false;
+ int c2 = classNextRune(&q2);
+ if (!m_error.isEmpty()) {
+ break;
+ }
+ items.append(SamRegex::ClassItem{true, c1, c2});
+ } else {
+ items.append(SamRegex::ClassItem{false, c1, c1});
+ }
+ }
+ m_re.m_classes.append(items);
+ }
+
+ SamRegex &m_re;
+ QString m_s;
+ int m_pos = 0;
+
+ QVector m_atorStack;
+ QVector m_subidStack;
+ QVector m_andStack;
+ int m_cursubid = 0;
+ int m_lastPoppedSubid = -1;
+ bool m_lastWasAnd = false;
+ int m_nbra = 0;
+ bool m_negateClass = false;
+ QString m_error;
+};
+
+// ---------------------------------------------------------------------------
+// SamRegex
+// ---------------------------------------------------------------------------
+bool SamRegex::compile(const QString &pattern)
+{
+ m_pattern = pattern;
+ m_valid = true;
+ m_error.clear();
+ SamRegexCompiler c(*this, pattern);
+ if (!c.compile()) {
+ m_valid = false;
+ if (m_error.isEmpty()) {
+ m_error = c.error();
+ }
+ return false;
+ }
+ m_valid = true;
+ return true;
+}
+
+bool SamRegex::classMatch(int classIdx, QChar c, bool negate) const
+{
+ if (classIdx < 0 || classIdx >= m_classes.size()) {
+ return negate;
+ }
+ const int cc = c.unicode();
+ for (const ClassItem &item : m_classes[classIdx]) {
+ if (item.range) {
+ if (item.lo <= cc && cc <= item.hi) {
+ return !negate;
+ }
+ } else if (item.lo == cc) {
+ return !negate;
+ }
+ }
+ return negate;
+}
+
+// Pike VM thread list. Each live thread is (inst, captures). addinst dedups by
+// instruction index within a list, keeping the thread whose match started
+// earliest (leftmost), exactly as regexp.c addinst does.
+namespace
+{
+struct Thread {
+ int inst = -1;
+ QVector caps;
+};
+} // namespace
+
+std::optional SamRegex::match(const QString &text, int from, int bound) const
+{
+ if (!m_valid || m_start < 0) {
+ return std::nullopt;
+ }
+ if (bound < 0 || bound > text.size()) {
+ bound = text.size();
+ }
+ if (from < 0) {
+ from = 0;
+ }
+
+ auto addinst = [](QVector &list, int inst, const QVector &caps) {
+ for (Thread &t : list) {
+ if (t.inst == inst) {
+ // Keep the earliest start (leftmost) — mirrors regexp.c.
+ if (caps[0].p1 < t.caps[0].p1) {
+ t.caps = caps;
+ }
+ return;
+ }
+ }
+ list.append(Thread{inst, caps});
+ };
+
+ QVector clist;
+ QVector nlist;
+
+ SamMatch best;
+ bool haveMatch = false;
+ auto newmatch = [&](const QVector &se) {
+ // Leftmost-longest: take if no match yet, or starts earlier, or same
+ // start and ends later (regexp.c newmatch()).
+ if (!haveMatch || se[0].p1 < best.caps[0].p1
+ || (se[0].p1 == best.caps[0].p1 && se[0].p2 > best.caps[0].p2)) {
+ best.caps = se;
+ haveMatch = true;
+ }
+ };
+
+ const int startChar =
+ (m_prog[m_start].type < kOperator) ? m_prog[m_start].type : 0;
+
+ int nnl = 0; // live threads carried into nlist
+ // Scan one position past bound so a thread ending exactly at bound (and
+ // EOL/END at bound) is processed.
+ for (int p = from; p <= bound; p++) {
+ const int c = (p < bound) ? text[p].unicode() : -1;
+
+ // Stop once a match is found and no threads remain alive.
+ if (haveMatch && nnl == 0) {
+ break;
+ }
+ // Fast first-char skip while no thread is live and no match pending.
+ if (startChar && nnl == 0 && !haveMatch && c != startChar) {
+ continue;
+ }
+
+ clist = nlist;
+ nlist.clear();
+ nnl = 0;
+
+ // Seed a fresh start thread at this position while no match found yet.
+ if (!haveMatch) {
+ QVector se(NSUBEXP, SamMatch::Range{-1, -1});
+ se[0].p1 = p;
+ addinst(clist, m_start, se);
+ }
+
+ // Run epsilon + consuming transitions for this position.
+ for (int ti = 0; ti < clist.size(); ti++) {
+ int inst = clist[ti].inst;
+ // Follow epsilon transitions within the same position.
+ bool consumed = false;
+ while (!consumed) {
+ const Inst &in = m_prog[inst];
+ switch (in.type) {
+ case LBRA:
+ if (in.subid >= 0 && in.subid < NSUBEXP) {
+ clist[ti].caps[in.subid].p1 = p;
+ }
+ inst = in.next;
+ continue;
+ case RBRA:
+ if (in.subid >= 0 && in.subid < NSUBEXP) {
+ clist[ti].caps[in.subid].p2 = p;
+ }
+ inst = in.next;
+ continue;
+ case OR:
+ addinst(clist, in.right, clist[ti].caps);
+ inst = in.next; // "left"/continuation branch
+ continue;
+ case NOP:
+ inst = in.next;
+ continue;
+ case BOL:
+ if (p == 0 || (p > 0 && text[p - 1] == QLatin1Char('\n'))) {
+ inst = in.next;
+ continue;
+ }
+ consumed = true; // dead
+ break;
+ case EOL:
+ if (c == '\n' || p == bound) {
+ inst = in.next;
+ continue;
+ }
+ consumed = true; // dead
+ break;
+ case END:
+ clist[ti].caps[0].p2 = p;
+ newmatch(clist[ti].caps);
+ consumed = true;
+ break;
+ case ANY:
+ if (c >= 0 && c != '\n') {
+ addinst(nlist, in.next, clist[ti].caps);
+ nnl++;
+ }
+ consumed = true;
+ break;
+ case CCLASS:
+ if (c >= 0 && classMatch(in.classIdx, QChar(c), false)) {
+ addinst(nlist, in.next, clist[ti].caps);
+ nnl++;
+ }
+ consumed = true;
+ break;
+ case NCCLASS:
+ if (c >= 0 && classMatch(in.classIdx, QChar(c), true)) {
+ addinst(nlist, in.next, clist[ti].caps);
+ nnl++;
+ }
+ consumed = true;
+ break;
+ default: // literal rune
+ if (c >= 0 && in.type == c) {
+ addinst(nlist, in.next, clist[ti].caps);
+ nnl++;
+ }
+ consumed = true;
+ break;
+ }
+ }
+ }
+ }
+
+ if (!haveMatch) {
+ return std::nullopt;
+ }
+ // Honour the bound: a match must lie within [from, bound].
+ if (best.caps[0].p2 > bound) {
+ return std::nullopt;
+ }
+ return best;
+}
+
+} // namespace katecustom
diff --git a/src/sam/samregex.h b/src/sam/samregex.h
new file mode 100644
index 0000000..3f3d15f
--- /dev/null
+++ b/src/sam/samregex.h
@@ -0,0 +1,114 @@
+/*
+ * SPDX-License-Identifier: LGPL-2.0-or-later
+ *
+ * SamRegex — a faithful port of plan9 sam's regular-expression engine
+ * (plan9port src/cmd/sam/regexp.c) to operate on a QString.
+ *
+ * Why this exists: sam's regex is a Thompson/Pike NFA simulation. It is
+ * - leftmost-longest (POSIX semantics), not Perl leftmost-greedy — so it
+ * matches sam exactly, including overlapping alternation (`a|ab` on "ab"
+ * matches "ab", the longest), which PCRE does not; and
+ * - linear time in the subject length with no backtracking, so it cannot
+ * suffer catastrophic blowup.
+ *
+ * The dialect is sam's regexp(7), deliberately small:
+ * . * + ? | ( ) ^ $ [ ] and \ to escape a metacharacter (\n = newline).
+ * `^` matches the beginning of a line, `$` the end of a line (per-line, built
+ * in). `.` and negated classes do not match newline. There is no \d \w \b,
+ * no lookaround, no non-greedy — sam never had them, and their absence keeps
+ * the engine linear and the syntax portable back to real sam.
+ *
+ * The implementation mirrors regexp.c structure: a shunting-yard parser builds
+ * an Inst NFA; execute() runs two parallel thread lists (Pike VM) where each
+ * thread carries a capture set, and newmatch() keeps the leftmost-longest.
+ * Only the forward machine is ported; backward search is done by the caller
+ * scanning forward and keeping the last qualifying match.
+ */
+#ifndef KATECUSTOM_SAMREGEX_H
+#define KATECUSTOM_SAMREGEX_H
+
+#include
+#include
+#include
+#include
+
+namespace katecustom
+{
+
+/*! A match: half-open character ranges. caps[0] is the whole match; caps[1..9]
+ * are parenthesised subexpressions ({-1,-1} if that group did not participate). */
+struct SamMatch {
+ struct Range {
+ int p1 = -1;
+ int p2 = -1;
+ };
+ QVector caps; // size NSUBEXP (10)
+ int start() const { return caps.isEmpty() ? -1 : caps[0].p1; }
+ int end() const { return caps.isEmpty() ? -1 : caps[0].p2; }
+ QString captured(int group, const QString &text) const;
+};
+
+class SamRegex
+{
+public:
+ SamRegex() = default;
+
+ /*! Compile \a pattern (sam dialect). Returns false and sets error() on a
+ * syntax error. */
+ bool compile(const QString &pattern);
+
+ bool isValid() const { return m_valid; }
+ QString error() const { return m_error; }
+ QString pattern() const { return m_pattern; }
+
+ /*!
+ * Find the leftmost-longest match whose start is at or after \a from and
+ * which lies entirely within [0, \a bound). Anchors `^`/`$` are evaluated
+ * against the whole \a text (so they fire at any line boundary inside the
+ * bound). Returns std::nullopt if there is no match. Linear time.
+ */
+ std::optional match(const QString &text, int from, int bound) const;
+
+private:
+ // NFA instruction (regexp.c Inst). type < kOperator is a literal rune;
+ // otherwise it is one of the action codes below.
+ //
+ // In the original, `left` and `next` share one union member (l.lleft ==
+ // l.lnext): an OR node's "left/continuation" IS its next pointer, so
+ // concatenation (which wires `.next`) and OR execution (which reads the
+ // continuation) refer to the same slot. We keep that single `next` slot and
+ // use `right` for the OR's branch target. `subid`/`classIdx` reuse would be
+ // the r-union; they are used disjointly per node type so separate fields are
+ // safe.
+ struct Inst {
+ int type = 0;
+ int subid = -1; // LBRA/RBRA capture index
+ int classIdx = -1; // CCLASS/NCCLASS index into m_classes
+ int next = -1; // continuation (also the OR "left" branch)
+ int right = -1; // OR branch target (loop body / first alternative)
+ };
+
+ // A character class: a flat list of items, each either a single rune or a
+ // range [lo,hi]. Negated classes also implicitly exclude '\n'.
+ struct ClassItem {
+ bool range = false;
+ int lo = 0;
+ int hi = 0;
+ };
+
+ bool classMatch(int classIdx, QChar c, bool negate) const;
+
+ QString m_pattern;
+ bool m_valid = false;
+ QString m_error;
+
+ QVector m_prog; // the compiled program
+ int m_start = -1; // index of the start instruction
+ QVector> m_classes;
+
+ friend class SamRegexCompiler;
+};
+
+} // namespace katecustom
+
+#endif
diff --git a/src/sam/test_samengine.cpp b/src/sam/test_samengine.cpp
index 60b893f..4aa45f3 100644
--- a/src/sam/test_samengine.cpp
+++ b/src/sam/test_samengine.cpp
@@ -56,6 +56,7 @@ private Q_SLOTS:
void caretIsPerLine();
void dollarIsPerLine();
void samAnchorIdiom();
+ void leftmostLongestSubstitution();
};
void TestSamEngine::substituteFirst()
@@ -87,8 +88,8 @@ void TestSamEngine::substituteAmp()
void TestSamEngine::substituteGroup()
{
- // \1 is the first capture group.
- QCOMPARE(runAll(QStringLiteral(",s/(\\w+)=(\\w+)/\\2=\\1/"),
+ // \1 is the first capture group (sam dialect: use explicit classes, not \w).
+ QCOMPARE(runAll(QStringLiteral(",s/([a-z]+)=([a-z]+)/\\2=\\1/"),
QStringLiteral("key=value")),
QStringLiteral("value=key"));
}
@@ -282,5 +283,15 @@ void TestSamEngine::samAnchorIdiom()
QStringLiteral("> a\n> b\n> c"));
}
+void TestSamEngine::leftmostLongestSubstitution()
+{
+ // With the ported sam engine, alternation is leftmost-LONGEST: a|ab on "ab"
+ // matches "ab", so the whole token is replaced (PCRE would match just "a").
+ QCOMPARE(runAll(QStringLiteral(",s/a|ab/X/"), QStringLiteral("ab")),
+ QStringLiteral("X"));
+ QCOMPARE(runAll(QStringLiteral(",s/ab|a/X/"), QStringLiteral("ab")),
+ QStringLiteral("X"));
+}
+
QTEST_MAIN(TestSamEngine)
#include "test_samengine.moc"
diff --git a/src/sam/test_samregex.cpp b/src/sam/test_samregex.cpp
new file mode 100644
index 0000000..7a9eca2
--- /dev/null
+++ b/src/sam/test_samregex.cpp
@@ -0,0 +1,200 @@
+/*
+ * SPDX-License-Identifier: LGPL-2.0-or-later
+ *
+ * Unit tests for the ported sam regex engine (SamRegex). These pin the two
+ * properties that motivated the port: leftmost-LONGEST semantics (POSIX, sam),
+ * and the sam dialect (no PCRE extras). Behaviours cross-checked against
+ * plan9port src/cmd/sam/regexp.c and sam(1).
+ */
+#include "samregex.h"
+
+#include
+#include
+
+using namespace katecustom;
+
+class TestSamRegex : public QObject
+{
+ Q_OBJECT
+
+ // Match from 0 over the whole text; return "start,end" or "nil".
+ QString m(const QString &pat, const QString &text)
+ {
+ SamRegex re;
+ if (!re.compile(pat)) {
+ return QStringLiteral("ERR:%1").arg(re.error());
+ }
+ auto r = re.match(text, 0, text.size());
+ if (!r) {
+ return QStringLiteral("nil");
+ }
+ return QStringLiteral("%1,%2").arg(r->start()).arg(r->end());
+ }
+
+ // Match and return the matched substring, or "nil".
+ QString ms(const QString &pat, const QString &text, int from = 0)
+ {
+ SamRegex re;
+ if (!re.compile(pat)) {
+ return QStringLiteral("ERR:%1").arg(re.error());
+ }
+ auto r = re.match(text, from, text.size());
+ if (!r) {
+ return QStringLiteral("nil");
+ }
+ return text.mid(r->start(), r->end() - r->start());
+ }
+
+private Q_SLOTS:
+ void literal();
+ void leftmostLongestAlternation();
+ void alternationEitherOrder();
+ void star();
+ void plusQuest();
+ void dotNotNewline();
+ void anchorsPerLine();
+ void charClass();
+ void negatedClass();
+ void range();
+ void captures();
+ void escapes();
+ void fromOffset();
+ void noMatch();
+ void nestedGroupsLongest();
+ void linearNoCatastrophicBacktrack();
+};
+
+void TestSamRegex::literal()
+{
+ QCOMPARE(m(QStringLiteral("bar"), QStringLiteral("foobarbaz")), QStringLiteral("3,6"));
+}
+
+void TestSamRegex::leftmostLongestAlternation()
+{
+ // THE motivating case: a|ab on "ab" must match "ab" (longest), unlike PCRE.
+ QCOMPARE(ms(QStringLiteral("a|ab"), QStringLiteral("ab")), QStringLiteral("ab"));
+}
+
+void TestSamRegex::alternationEitherOrder()
+{
+ // Order must not matter for leftmost-longest.
+ QCOMPARE(ms(QStringLiteral("ab|a"), QStringLiteral("ab")), QStringLiteral("ab"));
+ QCOMPARE(ms(QStringLiteral("a|ab"), QStringLiteral("ab")), QStringLiteral("ab"));
+}
+
+void TestSamRegex::star()
+{
+ QCOMPARE(ms(QStringLiteral("a*"), QStringLiteral("aaab")), QStringLiteral("aaa"));
+ // leftmost: at position 0 matches empty? a* matches "aaa" (longest at 0).
+ QCOMPARE(ms(QStringLiteral("ba*"), QStringLiteral("baaa")), QStringLiteral("baaa"));
+}
+
+void TestSamRegex::plusQuest()
+{
+ QCOMPARE(ms(QStringLiteral("a+"), QStringLiteral("baaa")), QStringLiteral("aaa"));
+ QCOMPARE(ms(QStringLiteral("ab?c"), QStringLiteral("ac")), QStringLiteral("ac"));
+ QCOMPARE(ms(QStringLiteral("ab?c"), QStringLiteral("abc")), QStringLiteral("abc"));
+}
+
+void TestSamRegex::dotNotNewline()
+{
+ // . does not cross newline (sam).
+ QCOMPARE(m(QStringLiteral("a.b"), QStringLiteral("a\nb")), QStringLiteral("nil"));
+ QCOMPARE(ms(QStringLiteral("a.c"), QStringLiteral("abc")), QStringLiteral("abc"));
+}
+
+void TestSamRegex::anchorsPerLine()
+{
+ // ^ matches at start of any line; $ at end of any line.
+ SamRegex re;
+ QVERIFY(re.compile(QStringLiteral("^b")));
+ auto r = re.match(QStringLiteral("a\nb\nc"), 0, 5);
+ QVERIFY(r.has_value());
+ QCOMPARE(r->start(), 2); // 'b' on line 2
+ QCOMPARE(r->end(), 3);
+
+ SamRegex re2;
+ QVERIFY(re2.compile(QStringLiteral("a$")));
+ auto r2 = re2.match(QStringLiteral("ba\nxa\n"), 0, 6);
+ QVERIFY(r2.has_value());
+ QCOMPARE(r2->start(), 1); // first 'a' that precedes a newline
+ QCOMPARE(r2->end(), 2);
+}
+
+void TestSamRegex::charClass()
+{
+ QCOMPARE(ms(QStringLiteral("[abc]+"), QStringLiteral("xcababz")), QStringLiteral("cabab"));
+}
+
+void TestSamRegex::negatedClass()
+{
+ // [^z]+ crosses nothing special but stops at z; does not match newline.
+ QCOMPARE(ms(QStringLiteral("[^z]+"), QStringLiteral("abz")), QStringLiteral("ab"));
+ QCOMPARE(ms(QStringLiteral("[^z]+"), QStringLiteral("a\nb")), QStringLiteral("a"));
+}
+
+void TestSamRegex::range()
+{
+ QCOMPARE(ms(QStringLiteral("[0-9]+"), QStringLiteral("ab123cd")), QStringLiteral("123"));
+ QCOMPARE(ms(QStringLiteral("[a-cx-z]+"), QStringLiteral("abcxyz!")), QStringLiteral("abcxyz"));
+}
+
+void TestSamRegex::captures()
+{
+ SamRegex re;
+ QVERIFY(re.compile(QStringLiteral("([a-z]+)=([a-z]+)")));
+ const QString text = QStringLiteral("key=value");
+ auto r = re.match(text, 0, text.size());
+ QVERIFY(r.has_value());
+ QCOMPARE(r->captured(1, text), QStringLiteral("key"));
+ QCOMPARE(r->captured(2, text), QStringLiteral("value"));
+}
+
+void TestSamRegex::escapes()
+{
+ // \* is a literal star; \. a literal dot; \n a newline.
+ QCOMPARE(ms(QStringLiteral("a\\*b"), QStringLiteral("xa*bx")), QStringLiteral("a*b"));
+ QCOMPARE(ms(QStringLiteral("a\\.b"), QStringLiteral("axb")), QStringLiteral("nil"));
+ QCOMPARE(ms(QStringLiteral("a\\.b"), QStringLiteral("a.b")), QStringLiteral("a.b"));
+ QCOMPARE(m(QStringLiteral("a\\nb"), QStringLiteral("a\nb")), QStringLiteral("0,3"));
+}
+
+void TestSamRegex::fromOffset()
+{
+ // Search starting past the first match finds the second.
+ QCOMPARE(ms(QStringLiteral("a"), QStringLiteral("xaya"), 2), QStringLiteral("a"));
+ SamRegex re;
+ QVERIFY(re.compile(QStringLiteral("a")));
+ auto r = re.match(QStringLiteral("xaya"), 2, 4);
+ QVERIFY(r.has_value());
+ QCOMPARE(r->start(), 3);
+}
+
+void TestSamRegex::noMatch()
+{
+ QCOMPARE(m(QStringLiteral("zzz"), QStringLiteral("abc")), QStringLiteral("nil"));
+}
+
+void TestSamRegex::nestedGroupsLongest()
+{
+ // (a|ab)(c|bcd) on "abcd": leftmost-longest should match the whole "abcd"
+ // (a + bcd), which greedy-ordered backtracking would miss if it locked in
+ // "a" then "bcd"? Actually both can reach abcd; the key is longest overall.
+ QCOMPARE(ms(QStringLiteral("(a|ab)(c|bcd)"), QStringLiteral("abcd")),
+ QStringLiteral("abcd"));
+}
+
+void TestSamRegex::linearNoCatastrophicBacktrack()
+{
+ // A pattern that makes PCRE backtrack exponentially: (a*)*b on a long run
+ // of 'a' with no 'b'. The NFA handles it in linear time; just assert it
+ // terminates quickly and reports no match.
+ const QString text(10000, QLatin1Char('a'));
+ SamRegex re;
+ QVERIFY(re.compile(QStringLiteral("(a*)*b")));
+ auto r = re.match(text, 0, text.size());
+ QVERIFY(!r.has_value());
+}
+
+QTEST_MAIN(TestSamRegex)
+#include "test_samregex.moc"