sam: port sam's own regex engine (leftmost-longest, linear-time)

Replace QRegularExpression with SamRegex (src/sam/samregex.{h,cpp}), a
faithful port of plan9port src/cmd/sam/regexp.c: a Thompson/Pike NFA.

  - Leftmost-LONGEST (POSIX): matches real sam exactly, including overlapping
    alternation (a|ab on 'ab' -> 'ab'), eliminating the one remaining
    greedy-vs-longest divergence from PCRE.
  - Linear time, no backtracking: immune to catastrophic blowup ((a*)*b over
    10k 'a' returns instantly).
  - sam dialect only: . * + ? | ( ) [ ] ^ $ and \ escaping (\n = newline);
    ^/$ per-line and ./negated-classes exclude newline, all intrinsic. No
    PCRE extras (\d \w \b, lookaround, non-greedy) — sam never had them.

Port notes: shunting-yard compiler + Pike VM with per-thread capture sets and
leftmost-longest newmatch(). Fixed two porting bugs vs the C original: the
l-union aliasing of OR's left/continuation with .next, and reading the popped
subid during popator for correct capture-group ids.

SamEngine now compiles/matches via SamRegex (search, s, x/y/g/v, replacement
captures). Tests: new test_samregex (18) + updated test_samengine (31, incl.
leftmostLongestSubstitution, sam-dialect capture groups). 18/18 ctest suites.

Docs: rewrite docs/SAM.md (no more PCRE deviation; dialect + semantics are
sam's), update PLAN.md and the panel hint.
This commit is contained in:
Levi Neely 2026-10-08 14:44:05 +02:00
parent 4cdd762bb6
commit 6341b1ad01
9 changed files with 1096 additions and 164 deletions

View File

@ -359,6 +359,15 @@ A dockable tool view ("Sam", left sidebar) runs the plan9 **sam** command
language against the active document, or project-wide via `X`/`Y`. Type a
program, click **Run** (or Ctrl+Return); each Run is one undo step.
- **Regex engine** (`src/sam/samregex.{h,cpp}`, 18 unit tests): a faithful port
of sam's own matcher (plan9port `src/cmd/sam/regexp.c`) — a Thompson/Pike NFA.
**Leftmost-longest** (POSIX), so it matches real sam exactly incl. overlapping
alternation (`a|ab` → `ab`); **linear time**, no backtracking (immune to
`(a*)*b` blowup). sam dialect only: `. * + ? | ( ) [ ] ^ $ \` with `\n`; `^`/`$`
are per-line and `.`/negated-classes exclude newline, all intrinsic. No PCRE
extras — not a deviation, just sam. Replaces the earlier QRegularExpression
matcher, eliminating the greedy-vs-longest divergence entirely.
- **Pure engine** (`src/sam/samengine.{h,cpp}`, unit-tested, no Kate dep):
`SamEngine::run(program, text, dotStart, dotEnd) -> SamResult{edits, dot,
output, applied}`. Edits are computed against the ORIGINAL snapshot in
@ -383,14 +392,9 @@ program, click **Run** (or Ctrl+Return); each Run is one undo step.
- **Panel** (`src/plugin/sampanel.{h,cpp}`): program editor + Run button +
output log; `OllieView` creates the tool view and owns the apply logic
(`runSamProgram`, `applySamToDocument`, `runSamFileLoop`).
- **Deviations (documented)**: regexes use `QRegularExpression` (PCRE), not
plan9 `regexp(7)`. Everyday patterns match identically; the one semantic
divergence is leftmost-greedy vs sam's leftmost-longest (bites only on
overlapping alternation like `a|ab`). `^`/`$` are per-line (compiled with
`MultilineOption`) and `.` does not cross newlines — both faithful to sam's
`regexp(7)` (BOL/EOL in `regexp.c`). PCRE also accepts extras sam lacks
(`\d \w \b`, lookahead, non-greedy, inline flags); the docs frame these as an
escape hatch, not a feature — prefer structural composition. Full user-facing
- **Dialect & semantics**: the regex engine is sam's own (see the Regex engine
bullet above), so matching is leftmost-longest and the dialect is sam's — no
PCRE, no greedy-vs-longest divergence, no PCRE extras. Full user-facing
treatment with verified examples in **`docs/SAM.md`**. Out of scope: multi-file
menu (`b B n D`), external file I/O (`e r w f`), `"re"` file-addressing, and
sam's own `u` (Kate's undo stack is the undo mechanism). The live GUI panel

View File

@ -6,126 +6,103 @@ against the active document, or project-wide via `X`/`Y`. Type a program, press
**Ctrl+Return** or click **Run**. Each Run is one undo step (Ctrl+Z reverts the
whole Run; multi-file `X`/`Y` undoes per file).
This document covers the **regex deviation** — katesam uses PCRE
(`QRegularExpression`), not plan9 `regexp(7)` — and what that means in practice.
Every example below is verified against the engine.
## The regex engine
## TL;DR
katesam uses a **faithful port of sam's own regular-expression engine**
(`src/sam/samregex.{h,cpp}`, from plan9port `src/cmd/sam/regexp.c`), not PCRE.
This is a Thompson/Pike NFA, which gives the two properties sam relies on:
- Everyday patterns behave **identically** to sam.
- One real semantic difference: **PCRE is leftmost-greedy, sam is
leftmost-longest.** It only bites on *overlapping alternation* (`a|ab`).
- `^` and `$` are **per-line** (match at every line boundary), the same as sam's
`regexp(7)`. `.` does not cross newlines, also the same as sam.
- PCRE *also* offers things sam lacks (`\d \w \b`, lookahead, non-greedy, inline
flags). They work, but prefer sam-style structural composition (`x g v`) — see
"Extras" below.
- **Leftmost-longest** (POSIX) matching, so results match real sam exactly,
including overlapping alternation.
- **Linear time**, with no backtracking — immune to catastrophic blowup. A
pattern like `(a*)*b` against thousands of `a` returns instantly.
---
There is **no PCRE deviation anymore**: the dialect, the match semantics, and
the anchors are sam's.
## The one semantic divergence: greedy vs longest
### Dialect
sam's `regexp(7)` matches the **leftmost-longest** substring. PCRE matches
**leftmost**, then resolves alternation **left-to-right, first wins** (greedy
within a branch, but alternation order decides between branches).
The sam dialect is deliberately small (sam `regexp(7)`):
| program | input | sam (leftmost-longest) | **katesam (PCRE)** |
|---|---|---|---|
| `,s/a\|ab/X/` | `ab` | `X` (matches `ab`) | **`Xb`** (matches `a`) |
| `,s/ab\|a/X/` | `ab` | `X` | `X` |
```
. any character except newline
* + ? zero-or-more / one-or-more / zero-or-one
| alternation
( ) grouping and capture (\1..\9 in replacements)
[ ] [^ ] character class / negated class (ranges a-z)
^ $ beginning / end of a LINE
\c literal c (escapes a metacharacter); \n is newline
```
**Implication:** when alternatives overlap, **order them longest-first**
(`ab|a`, not `a|ab`). This is the only case where a working sam command can
silently produce a different edit in katesam. Non-overlapping alternation
(`cat|dog`) and all non-alternation patterns are unaffected.
There is intentionally **no** `\d \w \s \b`, no lookaround, no non-greedy, and
no inline flags. sam never had them. Use character classes (`[0-9]`,
`[a-zA-Z_]`) and structural composition (`x y g v`) instead — that is the sam
way, and patterns stay portable to real sam.
Quantifier greediness is the same in both (`a.*c` on `axxcxxc` → whole string in
both). katesam additionally offers non-greedy `.*?` (a PCRE extra).
### Semantics (all verified against the port)
---
| behaviour | katesam | note |
|---|---|---|
| overlapping alternation `a\|ab` on `ab` | matches `ab` | leftmost-**longest**; order of branches does not matter |
| `.` crosses newline | **no** | matches sam |
| `^` / `$` | per-line | `^` = start of a line, `$` = end of a line |
| `[^z]+` crosses newline | no | negated classes also exclude `\n`, as in sam |
| empty-match global `s/x*/-/g` on `axbx` | `-a-b-` | one advance per null match |
| `&` whole match, `\1..\9` groups in replacement | yes | up to 9 capture groups |
| pathological `(a*)*b` | linear time | NFA, no backtracking |
## Anchors `^` / `$` are per-line — same as sam
### Anchors are per-line
This is **not** a deviation. plan9 `regexp(7)` defines `^` as "the beginning of
a line" and `$` as "the end of a line" (sam's `regexp.c`: `BOL` fires at offset
0 or after a `\n`; `EOL` fires before a `\n`). katesam compiles every pattern
with `QRegularExpression::MultilineOption`, so `^`/`$` match at every line
boundary, exactly as sam does:
`^` matches the beginning of any line and `$` the end of any line — built into
the engine (sam `regexp.c`: `BOL` fires at offset 0 or after `\n`; `EOL` before
`\n`). So:
| program | input | result |
|---|---|---|
| `,s/^/> /g` | `a⏎b⏎c` | `> a⏎> b⏎> c` (every line) |
| `,s/$/;/g` | `a⏎b⏎c` | `a;⏎b;⏎c;` (every line) |
| `,s/^/> /g` | `a⏎b⏎c` | `> a⏎> b⏎> c` |
| `,s/$/;/g` | `a⏎b⏎c` | `a;⏎b;⏎c;` |
The sam-idiomatic structural form works too, and is preferable because it
composes with guards and other loops:
The structural idiom is equivalent and composes with guards/loops:
```
,x/.+/ s/^/> / prefix every (non-empty) line → > a⏎> b⏎> c
,x/.+/ a/;/ append ";" to every line → a;⏎b;⏎c;
```
`.` does **not** cross newlines (sam: a "character" is "any character but
newline"), which is also PCRE's default. Line-structured descent like
`,x/.*\n/ …` therefore behaves as expected.
## Composing edits — the sam way
---
sam's power is structural composition, not clever single regexes. Descend into
matches with `x`/`y`, guard with `g`/`v`, group with `{}`:
## What matches identically (no surprises)
```
,x/[a-zA-Z_]+/ g/^[A-Z]/ c/CONST/ every identifier starting uppercase → CONST
,x/"[^"]*"/ s/foo/bar/g foo→bar only inside double-quoted strings
0/start/,/end/ x/[a-z]+/ s/.*/[&]/ bracket every lowercase word in a region
```
| behaviour | katesam (PCRE) | sam | same? |
|---|---|---|---|
| `.` crosses newline | no | no | ✅ |
| `[^z]+` crosses newline | yes | yes | ✅ |
| empty-match advance `s/x*/-/g` on `axbx` → `-a-b-` | ✅ | ✅ | ✅ |
| greedy quantifiers `* + ?` | ✅ | ✅ | ✅ |
| character classes `[a-z]`, groups `( )`, `|` | ✅ | ✅ | ✅ |
| `&` whole-match, `\1..\9` groups in replacement | ✅ | ✅ | ✅ |
Multi-file `X`/`Y` applies an inner program across the project:
> `.` not crossing newlines matches sam, so line-structured descent
> (`,x/.*\n/ ...`) behaves as expected. Use `(?s)` if you *want* `.` to span
> newlines.
```
X/\.cpp$/ ,s/old_api/new_api/g rewrite every .cpp in the project
Y/_test\./ ,x/TODO/ d drop TODO lines in non-test files
```
---
## Extras — available, but not the sam way
PCRE accepts syntax plan9 `regexp(7)` never had. It all works, but treat it as
an **escape hatch, not a headline**: sam's power comes from *composing* simple
patterns with `x`/`y`/`g`/`v`/`{}`, not from clever single regexes. Prefer
structural composition; reach for these only when it is genuinely simpler.
| extra | example | note |
|---|---|---|
| shorthand classes `\d \w \s` | `,s/\d+/N/g` | less typing than `[0-9]`; genuinely handy |
| word boundary `\b` | `,s/\bfoo\b/X/g` | handy; sam would use `x/foo/` with context |
| ignore-case `(?i)` | `,s/(?i)todo/DONE/g` | fills a real gap — sam has no case-insensitive match |
| non-greedy `*?` | `,s/a.*?c/X/` | the structural `x/…/` loop is the sam alternative |
| lookahead/behind | `,s/foo(?=bar)/X/` | un-sam; express context with `x`/`g`/`v` instead |
| backreference in pattern | `,s/(\w)\1/D/g` | rarely the right tool |
A pattern using these is **not portable back to sam**. If portability or
staying in the structural idiom matters, avoid them.
---
(The file set is the project index, matched by path. Edits land in buffers, not
on disk — review and save deliberately. Undo is per file for `X`/`Y`.)
## Practical guidance
- **Order overlapping alternatives longest-first** (`\.tar\.gz|\.gz`, not the
reverse). The only silent divergence from sam.
- **`^`/`$` are per-line** — just like sam. For whole-line edits either anchor
directly (`,s/^/> /g`) or, more sam-idiomatically, loop (`,x/.+/ …`).
- **Prefer structural composition** (`x g v {}`) over PCRE extras. The extras
work but pull you out of the sam idiom and are not sam-portable.
- **Empty-match globals** (`s/x*/…/g`) advance one position per null match, as in
sam — safe, but as always with `*`, double-check the result.
- Use character classes, not PCRE shorthands: `[0-9]` not `\d`, `[a-zA-Z_]` not
`\w`.
- Prefer structural composition (`x g v {}`) over one dense pattern.
- `^`/`$` are per-line; anchor directly (`,s/^…`) or loop (`,x/.+/ …`).
- Empty-match globals (`s/x*/…/g`) advance one position per null match — safe,
but as always with `*`, check the result.
## Why PCRE and not regexp(7)
## Scope
`QRegularExpression` ships with Qt (already a dependency) and gives a richer,
familiar syntax for free. plan9 `regexp(7)` is a Rune-based NFA woven into sam's
own buffer types; using it would mean porting that engine. The tradeoff accepted
here: everyday fidelity plus extra power, at the cost of exact leftmost-longest
semantics on overlapping alternation. If that matters for your workflow, the
engine's regex calls are isolated in `SamEngine` and could be swapped for a
ported `regexp.c` later.
Supported: the full address language (`#n n 0 $ . '` `/re/ ?re?` compound
`+ - , ;`) and commands `a c i d s p = m t k x y g v {}` plus shell filters
`< > | !`. Out of scope (Kate manages files / its own undo): the multi-file menu
(`b B n D`), external file I/O (`e r w f`), `"re"` file-addressing, and sam's own
`u` (use Kate's Ctrl+Z).

View File

@ -56,8 +56,9 @@ SamPanel::SamPanel(QWidget *parent)
auto *hint = new QLabel(
i18n("sam commands — e.g. <tt>,s/foo/bar/g</tt> or "
"<tt>X/\\.cpp$/ ,s/old/new/g</tt> (project-wide). Ctrl+Return runs.<br/>"
"Regex is PCRE: <tt>^</tt>/<tt>$</tt> are per-line (like sam); order "
"overlapping alternatives longest-first. See docs/SAM.md."),
"Regex is sam's own (leftmost-longest, linear-time): classes like "
"<tt>[0-9]</tt>, not <tt>\\d</tt>; <tt>^</tt>/<tt>$</tt> are per-line. "
"See docs/SAM.md."),
this);
hint->setWordWrap(true);
hint->setTextFormat(Qt::RichText);

View File

@ -8,6 +8,8 @@ find_package(Qt6 ${QT_MIN_VERSION} COMPONENTS Core REQUIRED)
add_library(sam_lib STATIC
samengine.cpp
samengine.h
samregex.cpp
samregex.h
)
target_link_libraries(sam_lib PUBLIC Qt6::Core)
target_include_directories(sam_lib PUBLIC ${CMAKE_CURRENT_SOURCE_DIR})
@ -17,4 +19,8 @@ if(Qt6Test_FOUND)
add_executable(test_samengine test_samengine.cpp)
target_link_libraries(test_samengine PRIVATE sam_lib Qt6::Test)
add_test(NAME samengine COMMAND test_samengine)
add_executable(test_samregex test_samregex.cpp)
target_link_libraries(test_samregex PRIVATE sam_lib Qt6::Test)
add_test(NAME samregex COMMAND test_samregex)
endif()

View File

@ -18,9 +18,9 @@
* dot bound to each match, exactly as xec.c looper()/linelooper().
*/
#include "samengine.h"
#include "samregex.h"
#include <QProcess>
#include <QRegularExpression>
#include <algorithm>
namespace katecustom
@ -656,16 +656,11 @@ private:
return line;
}
// Build a QRegularExpression from a sam pattern. The empty pattern reuses
// the last compiled pattern (sam behaviour).
//
// Multiline is ON: plan9 regexp(7) defines ^ as "beginning of a line" and $
// as "end of a line" (sam regexp.c BOL = p==0 || prev=='\n'; EOL = next
// char is '\n'). That is exactly QRegularExpression::MultilineOption, so it
// is the faithful default, not an opt-in. '.' still does not cross newlines
// (sam: "the word character means any character but newline"), which is
// QRegularExpression's default (no DotMatchesEverythingOption).
QRegularExpression compile(const QString &pat)
// Build a SamRegex from a sam pattern. The empty pattern reuses the last
// compiled pattern (sam behaviour). This is the ported plan9 sam engine
// (SamRegex), so matching is leftmost-LONGEST and linear-time, ^/$ are
// per-line and '.' excludes newline — all intrinsic, matching real sam.
SamRegex compile(const QString &pat)
{
QString p = pat;
if (p.isEmpty()) {
@ -673,9 +668,9 @@ private:
} else {
m_lastRe = p;
}
QRegularExpression re(p, QRegularExpression::MultilineOption);
if (!re.isValid()) {
fail(QStringLiteral("bad regexp: %1").arg(re.errorString()));
SamRegex re;
if (!re.compile(p)) {
fail(QStringLiteral("bad regexp: %1").arg(re.error()));
}
return re;
}
@ -861,34 +856,33 @@ private:
// Forward (sign>0) / backward (sign<0) search, wrapping, like nextmatch().
Range search(const QString &pat, Range a, int sign)
{
QRegularExpression re = compile(pat);
SamRegex re = compile(pat);
if (failed()) {
return a;
}
if (sign >= 0) {
int from = a.p2;
QRegularExpressionMatch m = re.match(m_text, from);
if (!m.hasMatch()) {
m = re.match(m_text, 0); // wrap
auto m = re.match(m_text, a.p2, m_nc);
if (!m) {
m = re.match(m_text, 0, m_nc); // wrap
}
if (!m.hasMatch()) {
if (!m) {
fail(QStringLiteral("search failed: %1").arg(pat.isEmpty() ? m_lastRe : pat));
return a;
}
return {int(m.capturedStart()), int(m.capturedEnd())};
return {m->start(), m->end()};
} else {
// Backward: find the last match that ends at or before a.p1,
// else wrap to the last match in the text.
// Backward: the last match that ends at or before a.p1, else wrap to
// the last match in the whole text.
int limit = a.p1;
Range best{-1, -1};
int from = 0;
while (true) {
QRegularExpressionMatch m = re.match(m_text, from);
if (!m.hasMatch()) {
auto m = re.match(m_text, from, m_nc);
if (!m) {
break;
}
const int s = int(m.capturedStart());
const int e = int(m.capturedEnd());
const int s = m->start();
const int e = m->end();
if (e <= limit) {
best = {s, e};
} else {
@ -897,17 +891,14 @@ private:
from = (e > s) ? e : e + 1;
}
if (best.p1 < 0) {
// wrap: take the very last match in the whole text
from = 0;
while (true) {
QRegularExpressionMatch m = re.match(m_text, from);
if (!m.hasMatch()) {
auto m = re.match(m_text, from, m_nc);
if (!m) {
break;
}
best = {int(m.capturedStart()), int(m.capturedEnd())};
const int s = best.p1;
const int e = best.p2;
from = (e > s) ? e : e + 1;
best = {m->start(), m->end()};
from = (best.p2 > best.p1) ? best.p2 : best.p2 + 1;
}
}
if (best.p1 < 0) {
@ -921,7 +912,7 @@ private:
// s/re/text/ with count and g flag (xec.c s_cmd()).
void execSubstitute(Cmd *cmd, Range r)
{
QRegularExpression re = compile(cmd->re);
SamRegex re = compile(cmd->re);
if (failed()) {
return;
}
@ -931,15 +922,12 @@ private:
int lastEnd = r.p1;
int p1 = r.p1;
while (p1 <= r.p2) {
QRegularExpressionMatch m = re.match(m_text, p1);
if (!m.hasMatch() || int(m.capturedStart()) > r.p2) {
break;
}
int ms = int(m.capturedStart());
int me = int(m.capturedEnd());
if (me > r.p2) {
auto m = re.match(m_text, p1, r.p2);
if (!m) {
break;
}
int ms = m->start();
int me = m->end();
// advance logic (empty match handling)
if (ms == me) {
if (ms == op) {
@ -956,7 +944,7 @@ private:
continue;
}
const QString repl = expandReplacement(cmd->text, m);
const QString repl = expandReplacement(cmd->text, *m);
addEdit(ms, me, repl);
if (failed()) {
return;
@ -975,7 +963,7 @@ private:
}
// Expand & and \1..\9 in an s replacement using the match captures.
QString expandReplacement(const QString &tmpl, const QRegularExpressionMatch &m)
QString expandReplacement(const QString &tmpl, const SamMatch &m)
{
QString out;
for (int i = 0; i < tmpl.size(); i++) {
@ -984,7 +972,7 @@ private:
const QChar d = tmpl[i + 1];
if (d.isDigit()) {
const int g = d.unicode() - '0';
out.append(m.captured(g));
out.append(m.captured(g, m_text));
i++;
continue;
}
@ -1003,7 +991,7 @@ private:
continue;
}
if (c == QLatin1Char('&')) {
out.append(m.captured(0));
out.append(m.captured(0, m_text));
continue;
}
out.append(c);
@ -1046,14 +1034,13 @@ private:
m_nest++;
if (c == QLatin1Char('g') || c == QLatin1Char('v')) {
QRegularExpression re = compile(cmd->re);
SamRegex re = compile(cmd->re);
if (failed()) {
m_nest--;
return;
}
QRegularExpressionMatch m = re.match(m_text, r.p1);
const bool contains = m.hasMatch() && int(m.capturedStart()) <= r.p2
&& int(m.capturedStart()) >= r.p1;
auto m = re.match(m_text, r.p1, r.p2);
const bool contains = m.has_value();
const bool run = (c == QLatin1Char('g')) ? contains : !contains;
if (run) {
runLoopBody(cmd, r);
@ -1065,7 +1052,7 @@ private:
// x and y.
if (c == QLatin1Char('y')) {
// y: run on the gaps between matches of re within r.
QRegularExpression re = compile(cmd->re);
SamRegex re = compile(cmd->re);
if (failed()) {
m_nest--;
return;
@ -1073,12 +1060,12 @@ private:
int prev = r.p1;
int from = r.p1;
while (from <= r.p2) {
QRegularExpressionMatch m = re.match(m_text, from);
if (!m.hasMatch() || int(m.capturedEnd()) > r.p2) {
auto m = re.match(m_text, from, r.p2);
if (!m) {
break;
}
int ms = int(m.capturedStart());
int me = int(m.capturedEnd());
int ms = m->start();
int me = m->end();
runLoopBody(cmd, {prev, ms});
if (failed()) {
m_nest--;
@ -1093,7 +1080,7 @@ private:
}
// x: run on each match of re within r (default re handled by compile).
QRegularExpression re = compile(cmd->re);
SamRegex re = compile(cmd->re);
if (failed()) {
m_nest--;
return;
@ -1101,12 +1088,12 @@ private:
int op = -1;
int p = r.p1;
while (p <= r.p2) {
QRegularExpressionMatch m = re.match(m_text, p);
if (!m.hasMatch() || int(m.capturedEnd()) > r.p2) {
auto m = re.match(m_text, p, r.p2);
if (!m) {
break;
}
int ms = int(m.capturedStart());
int me = int(m.capturedEnd());
int ms = m->start();
int me = m->end();
if (ms == me) {
if (ms == op) {
p++;

632
src/sam/samregex.cpp Normal file
View File

@ -0,0 +1,632 @@
/*
* SPDX-License-Identifier: LGPL-2.0-or-later
*
* Port of plan9port src/cmd/sam/regexp.c. Structure and names follow the
* original closely; pointers become indices into QVector<Inst>.
*/
#include "samregex.h"
namespace katecustom
{
// Action codes (regexp.c). A literal rune uses its own code point as `type`
// (always < kOperator for the characters we accept). Operators carry the
// kOperator bit; the low bits encode precedence. Tokens carry the kAny bit.
enum {
kOperator = 0x1000000,
START = kOperator + 0,
RBRA = kOperator + 1, // )
LBRA = kOperator + 2, // (
OR = kOperator + 3, // |
CAT = kOperator + 4, // implicit concatenation
STAR = kOperator + 5, // *
PLUS = kOperator + 6, // +
QUEST = kOperator + 7, // ?
kAny = 0x2000000,
ANY = kAny + 0, // .
NOP = kAny + 1, // epsilon, removed by optimize()
BOL = kAny + 2, // ^
EOL = kAny + 3, // $
CCLASS = kAny + 4, // [...]
NCCLASS = kAny + 5, // [^...]
END = kAny + 0x77, // match
};
constexpr int NSUBEXP = 10;
QString SamMatch::captured(int group, const QString &text) const
{
if (group < 0 || group >= caps.size()) {
return QString();
}
const Range r = caps[group];
if (r.p1 < 0 || r.p2 < 0 || r.p2 < r.p1) {
return QString();
}
return text.mid(r.p1, r.p2 - r.p1);
}
// ---------------------------------------------------------------------------
// Compiler — shunting yard over the pattern, emitting Inst nodes.
// ---------------------------------------------------------------------------
class SamRegexCompiler
{
public:
SamRegexCompiler(SamRegex &re, const QString &pat) : m_re(re), m_s(pat) {}
bool compile()
{
m_re.m_prog.clear();
m_re.m_classes.clear();
m_atorStack.clear();
m_subidStack.clear();
m_andStack.clear();
m_cursubid = 0;
m_lastWasAnd = false;
m_pos = 0;
m_nbra = 0;
pushator(START - 1);
int token;
while ((token = lex()) != END) {
if (m_error.isEmpty() == false) {
return fail();
}
if ((token & kOperator) == kOperator) {
doOperator(token);
} else {
doOperand(token);
}
if (!m_error.isEmpty()) {
return fail();
}
}
evaluntil(START);
doOperand(END);
evaluntil(START);
if (m_nbra) {
m_error = QStringLiteral("unmatched '('");
return fail();
}
if (!m_error.isEmpty() || m_andStack.isEmpty()) {
if (m_error.isEmpty()) {
m_error = QStringLiteral("malformed regexp");
}
return fail();
}
m_re.m_start = m_andStack.last().first;
optimize();
return true;
}
QString error() const { return m_error; }
private:
struct Node {
int first = -1;
int last = -1;
};
bool fail()
{
m_re.m_valid = false;
if (m_re.m_error.isEmpty()) {
m_re.m_error = m_error.isEmpty() ? QStringLiteral("regexp error") : m_error;
}
return false;
}
int newinst(int t)
{
SamRegex::Inst inst;
inst.type = t;
m_re.m_prog.append(inst);
return m_re.m_prog.size() - 1;
}
SamRegex::Inst &at(int i) { return m_re.m_prog[i]; }
void pushand(int f, int l) { m_andStack.append(Node{f, l}); }
void pushator(int t)
{
m_atorStack.append(t);
m_subidStack.append(m_cursubid >= NSUBEXP ? -1 : m_cursubid);
}
int popator()
{
if (m_atorStack.isEmpty()) {
m_error = QStringLiteral("operator stack underflow");
return START;
}
// The subid paired with this operator — captured before removal so the
// LBRA/RBRA case can read the group's own id (regexp.c reads *subidp
// right after decrementing it).
m_lastPoppedSubid = m_subidStack.isEmpty() ? -1 : m_subidStack.last();
if (!m_subidStack.isEmpty()) {
m_subidStack.removeLast();
}
const int t = m_atorStack.last();
m_atorStack.removeLast();
return t;
}
Node popand(int op)
{
if (m_andStack.isEmpty()) {
m_error = op ? QStringLiteral("missing operand for '%1'").arg(QChar(op))
: QStringLiteral("malformed regexp");
return Node{};
}
const Node n = m_andStack.last();
m_andStack.removeLast();
return n;
}
void doOperand(int t)
{
if (m_lastWasAnd) {
doOperator(CAT); // implicit concatenation
}
int i = newinst(t);
if (t == CCLASS) {
if (m_negateClass) {
at(i).type = NCCLASS;
}
at(i).classIdx = m_re.m_classes.size() - 1;
}
pushand(i, i);
m_lastWasAnd = true;
}
void doOperator(int t)
{
if (t == RBRA && --m_nbra < 0) {
m_error = QStringLiteral("unmatched ')'");
return;
}
if (t == LBRA) {
m_cursubid++;
m_nbra++;
if (m_lastWasAnd) {
doOperator(CAT);
}
} else {
evaluntil(t);
}
if (t != RBRA) {
pushator(t);
}
m_lastWasAnd = false;
if (t == STAR || t == QUEST || t == PLUS || t == RBRA) {
m_lastWasAnd = true;
}
}
void evaluntil(int pri)
{
while (m_error.isEmpty()
&& (pri == RBRA || (!m_atorStack.isEmpty() && m_atorStack.last() >= pri))) {
const int op = popator();
if (!m_error.isEmpty()) {
return;
}
if (op == LBRA) {
Node op1 = popand('(');
const int subid = m_lastPoppedSubid;
int inst2 = newinst(RBRA);
at(inst2).subid = subid;
at(op1.last).next = inst2;
int inst1 = newinst(LBRA);
at(inst1).subid = subid;
at(inst1).next = op1.first;
pushand(inst1, inst2);
return; // must have been RBRA
} else if (op == OR) {
Node op2 = popand('|');
Node op1 = popand('|');
if (!m_error.isEmpty()) {
return;
}
int inst2 = newinst(NOP);
at(op2.last).next = inst2;
at(op1.last).next = inst2;
int inst1 = newinst(OR);
at(inst1).right = op1.first;
at(inst1).next = op2.first; // "left" branch == continuation slot
pushand(inst1, inst2);
} else if (op == CAT) {
Node op2 = popand(0);
Node op1 = popand(0);
if (!m_error.isEmpty()) {
return;
}
at(op1.last).next = op2.first;
pushand(op1.first, op2.last);
} else if (op == STAR) {
Node op2 = popand('*');
if (!m_error.isEmpty()) {
return;
}
int inst1 = newinst(OR);
at(op2.last).next = inst1;
at(inst1).right = op2.first;
pushand(inst1, inst1);
} else if (op == PLUS) {
Node op2 = popand('+');
if (!m_error.isEmpty()) {
return;
}
int inst1 = newinst(OR);
at(op2.last).next = inst1;
at(inst1).right = op2.first;
pushand(op2.first, inst1);
} else if (op == QUEST) {
Node op2 = popand('?');
if (!m_error.isEmpty()) {
return;
}
int inst1 = newinst(OR);
int inst2 = newinst(NOP);
at(inst1).next = inst2; // skip branch (continuation slot)
at(inst1).right = op2.first;
at(op2.last).next = inst2;
pushand(inst1, inst2);
} else {
m_error = QStringLiteral("bad regexp operator");
return;
}
if (pri == RBRA) {
// Keep looping until the matching LBRA (handled by the return
// above). The while condition handles the rest.
}
}
}
// Rewrite `next` pointers to skip NOP chains (regexp.c optimize()).
void optimize()
{
for (int i = 0; i < m_re.m_prog.size(); i++) {
if (m_re.m_prog[i].type == END) {
continue;
}
int target = m_re.m_prog[i].next;
while (target >= 0 && m_re.m_prog[target].type == NOP) {
target = m_re.m_prog[target].next;
}
m_re.m_prog[i].next = target;
}
}
// --- lexer ---
int lex()
{
if (m_pos >= m_s.size()) {
return END;
}
QChar qc = m_s[m_pos++];
int c = qc.unicode();
switch (c) {
case '\\':
if (m_pos < m_s.size()) {
QChar d = m_s[m_pos++];
c = (d == QLatin1Char('n')) ? '\n' : d.unicode();
}
// escaped: always a literal, never a metachar/token
return c;
case '*':
return STAR;
case '?':
return QUEST;
case '+':
return PLUS;
case '|':
return OR;
case '.':
return ANY;
case '(':
return LBRA;
case ')':
return RBRA;
case '^':
return BOL;
case '$':
return EOL;
case '[':
buildClass();
return CCLASS;
default:
return c;
}
}
// Read one class element, honouring \n and \<char>.
int classNextRune(bool *quoted)
{
*quoted = false;
if (m_pos >= m_s.size()) {
m_error = QStringLiteral("malformed character class");
return 0;
}
QChar c = m_s[m_pos];
if (c == QLatin1Char('\\')) {
m_pos++;
if (m_pos >= m_s.size()) {
m_error = QStringLiteral("malformed character class");
return 0;
}
QChar d = m_s[m_pos++];
if (d == QLatin1Char('n')) {
return '\n';
}
*quoted = true;
return d.unicode();
}
m_pos++;
return c.unicode();
}
void buildClass()
{
QVector<SamRegex::ClassItem> items;
m_negateClass = false;
if (m_pos < m_s.size() && m_s[m_pos] == QLatin1Char('^')) {
m_negateClass = true;
m_pos++;
// Negated classes never match newline.
items.append(SamRegex::ClassItem{false, '\n', '\n'});
}
while (true) {
if (m_pos >= m_s.size()) {
m_error = QStringLiteral("unterminated character class");
break;
}
// A ']' that is not escaped closes the class.
if (m_s[m_pos] == QLatin1Char(']')) {
m_pos++;
break;
}
bool q1 = false;
int c1 = classNextRune(&q1);
if (!m_error.isEmpty()) {
break;
}
// Range a-b: a '-' followed by a non-']' element.
if (m_pos + 0 < m_s.size() && m_s[m_pos] == QLatin1Char('-')
&& m_pos + 1 < m_s.size() && m_s[m_pos + 1] != QLatin1Char(']')) {
m_pos++; // consume '-'
bool q2 = false;
int c2 = classNextRune(&q2);
if (!m_error.isEmpty()) {
break;
}
items.append(SamRegex::ClassItem{true, c1, c2});
} else {
items.append(SamRegex::ClassItem{false, c1, c1});
}
}
m_re.m_classes.append(items);
}
SamRegex &m_re;
QString m_s;
int m_pos = 0;
QVector<int> m_atorStack;
QVector<int> m_subidStack;
QVector<Node> m_andStack;
int m_cursubid = 0;
int m_lastPoppedSubid = -1;
bool m_lastWasAnd = false;
int m_nbra = 0;
bool m_negateClass = false;
QString m_error;
};
// ---------------------------------------------------------------------------
// SamRegex
// ---------------------------------------------------------------------------
bool SamRegex::compile(const QString &pattern)
{
m_pattern = pattern;
m_valid = true;
m_error.clear();
SamRegexCompiler c(*this, pattern);
if (!c.compile()) {
m_valid = false;
if (m_error.isEmpty()) {
m_error = c.error();
}
return false;
}
m_valid = true;
return true;
}
bool SamRegex::classMatch(int classIdx, QChar c, bool negate) const
{
if (classIdx < 0 || classIdx >= m_classes.size()) {
return negate;
}
const int cc = c.unicode();
for (const ClassItem &item : m_classes[classIdx]) {
if (item.range) {
if (item.lo <= cc && cc <= item.hi) {
return !negate;
}
} else if (item.lo == cc) {
return !negate;
}
}
return negate;
}
// Pike VM thread list. Each live thread is (inst, captures). addinst dedups by
// instruction index within a list, keeping the thread whose match started
// earliest (leftmost), exactly as regexp.c addinst does.
namespace
{
struct Thread {
int inst = -1;
QVector<SamMatch::Range> caps;
};
} // namespace
std::optional<SamMatch> SamRegex::match(const QString &text, int from, int bound) const
{
if (!m_valid || m_start < 0) {
return std::nullopt;
}
if (bound < 0 || bound > text.size()) {
bound = text.size();
}
if (from < 0) {
from = 0;
}
auto addinst = [](QVector<Thread> &list, int inst, const QVector<SamMatch::Range> &caps) {
for (Thread &t : list) {
if (t.inst == inst) {
// Keep the earliest start (leftmost) — mirrors regexp.c.
if (caps[0].p1 < t.caps[0].p1) {
t.caps = caps;
}
return;
}
}
list.append(Thread{inst, caps});
};
QVector<Thread> clist;
QVector<Thread> nlist;
SamMatch best;
bool haveMatch = false;
auto newmatch = [&](const QVector<SamMatch::Range> &se) {
// Leftmost-longest: take if no match yet, or starts earlier, or same
// start and ends later (regexp.c newmatch()).
if (!haveMatch || se[0].p1 < best.caps[0].p1
|| (se[0].p1 == best.caps[0].p1 && se[0].p2 > best.caps[0].p2)) {
best.caps = se;
haveMatch = true;
}
};
const int startChar =
(m_prog[m_start].type < kOperator) ? m_prog[m_start].type : 0;
int nnl = 0; // live threads carried into nlist
// Scan one position past bound so a thread ending exactly at bound (and
// EOL/END at bound) is processed.
for (int p = from; p <= bound; p++) {
const int c = (p < bound) ? text[p].unicode() : -1;
// Stop once a match is found and no threads remain alive.
if (haveMatch && nnl == 0) {
break;
}
// Fast first-char skip while no thread is live and no match pending.
if (startChar && nnl == 0 && !haveMatch && c != startChar) {
continue;
}
clist = nlist;
nlist.clear();
nnl = 0;
// Seed a fresh start thread at this position while no match found yet.
if (!haveMatch) {
QVector<SamMatch::Range> se(NSUBEXP, SamMatch::Range{-1, -1});
se[0].p1 = p;
addinst(clist, m_start, se);
}
// Run epsilon + consuming transitions for this position.
for (int ti = 0; ti < clist.size(); ti++) {
int inst = clist[ti].inst;
// Follow epsilon transitions within the same position.
bool consumed = false;
while (!consumed) {
const Inst &in = m_prog[inst];
switch (in.type) {
case LBRA:
if (in.subid >= 0 && in.subid < NSUBEXP) {
clist[ti].caps[in.subid].p1 = p;
}
inst = in.next;
continue;
case RBRA:
if (in.subid >= 0 && in.subid < NSUBEXP) {
clist[ti].caps[in.subid].p2 = p;
}
inst = in.next;
continue;
case OR:
addinst(clist, in.right, clist[ti].caps);
inst = in.next; // "left"/continuation branch
continue;
case NOP:
inst = in.next;
continue;
case BOL:
if (p == 0 || (p > 0 && text[p - 1] == QLatin1Char('\n'))) {
inst = in.next;
continue;
}
consumed = true; // dead
break;
case EOL:
if (c == '\n' || p == bound) {
inst = in.next;
continue;
}
consumed = true; // dead
break;
case END:
clist[ti].caps[0].p2 = p;
newmatch(clist[ti].caps);
consumed = true;
break;
case ANY:
if (c >= 0 && c != '\n') {
addinst(nlist, in.next, clist[ti].caps);
nnl++;
}
consumed = true;
break;
case CCLASS:
if (c >= 0 && classMatch(in.classIdx, QChar(c), false)) {
addinst(nlist, in.next, clist[ti].caps);
nnl++;
}
consumed = true;
break;
case NCCLASS:
if (c >= 0 && classMatch(in.classIdx, QChar(c), true)) {
addinst(nlist, in.next, clist[ti].caps);
nnl++;
}
consumed = true;
break;
default: // literal rune
if (c >= 0 && in.type == c) {
addinst(nlist, in.next, clist[ti].caps);
nnl++;
}
consumed = true;
break;
}
}
}
}
if (!haveMatch) {
return std::nullopt;
}
// Honour the bound: a match must lie within [from, bound].
if (best.caps[0].p2 > bound) {
return std::nullopt;
}
return best;
}
} // namespace katecustom

114
src/sam/samregex.h Normal file
View File

@ -0,0 +1,114 @@
/*
* SPDX-License-Identifier: LGPL-2.0-or-later
*
* SamRegex — a faithful port of plan9 sam's regular-expression engine
* (plan9port src/cmd/sam/regexp.c) to operate on a QString.
*
* Why this exists: sam's regex is a Thompson/Pike NFA simulation. It is
* - leftmost-longest (POSIX semantics), not Perl leftmost-greedy — so it
* matches sam exactly, including overlapping alternation (`a|ab` on "ab"
* matches "ab", the longest), which PCRE does not; and
* - linear time in the subject length with no backtracking, so it cannot
* suffer catastrophic blowup.
*
* The dialect is sam's regexp(7), deliberately small:
* . * + ? | ( ) ^ $ [ ] and \ to escape a metacharacter (\n = newline).
* `^` matches the beginning of a line, `$` the end of a line (per-line, built
* in). `.` and negated classes do not match newline. There is no \d \w \b,
* no lookaround, no non-greedy — sam never had them, and their absence keeps
* the engine linear and the syntax portable back to real sam.
*
* The implementation mirrors regexp.c structure: a shunting-yard parser builds
* an Inst NFA; execute() runs two parallel thread lists (Pike VM) where each
* thread carries a capture set, and newmatch() keeps the leftmost-longest.
* Only the forward machine is ported; backward search is done by the caller
* scanning forward and keeping the last qualifying match.
*/
#ifndef KATECUSTOM_SAMREGEX_H
#define KATECUSTOM_SAMREGEX_H
#include <QList>
#include <QString>
#include <QVector>
#include <optional>
namespace katecustom
{
/*! A match: half-open character ranges. caps[0] is the whole match; caps[1..9]
* are parenthesised subexpressions ({-1,-1} if that group did not participate). */
struct SamMatch {
struct Range {
int p1 = -1;
int p2 = -1;
};
QVector<Range> caps; // size NSUBEXP (10)
int start() const { return caps.isEmpty() ? -1 : caps[0].p1; }
int end() const { return caps.isEmpty() ? -1 : caps[0].p2; }
QString captured(int group, const QString &text) const;
};
class SamRegex
{
public:
SamRegex() = default;
/*! Compile \a pattern (sam dialect). Returns false and sets error() on a
* syntax error. */
bool compile(const QString &pattern);
bool isValid() const { return m_valid; }
QString error() const { return m_error; }
QString pattern() const { return m_pattern; }
/*!
* Find the leftmost-longest match whose start is at or after \a from and
* which lies entirely within [0, \a bound). Anchors `^`/`$` are evaluated
* against the whole \a text (so they fire at any line boundary inside the
* bound). Returns std::nullopt if there is no match. Linear time.
*/
std::optional<SamMatch> match(const QString &text, int from, int bound) const;
private:
// NFA instruction (regexp.c Inst). type < kOperator is a literal rune;
// otherwise it is one of the action codes below.
//
// In the original, `left` and `next` share one union member (l.lleft ==
// l.lnext): an OR node's "left/continuation" IS its next pointer, so
// concatenation (which wires `.next`) and OR execution (which reads the
// continuation) refer to the same slot. We keep that single `next` slot and
// use `right` for the OR's branch target. `subid`/`classIdx` reuse would be
// the r-union; they are used disjointly per node type so separate fields are
// safe.
struct Inst {
int type = 0;
int subid = -1; // LBRA/RBRA capture index
int classIdx = -1; // CCLASS/NCCLASS index into m_classes
int next = -1; // continuation (also the OR "left" branch)
int right = -1; // OR branch target (loop body / first alternative)
};
// A character class: a flat list of items, each either a single rune or a
// range [lo,hi]. Negated classes also implicitly exclude '\n'.
struct ClassItem {
bool range = false;
int lo = 0;
int hi = 0;
};
bool classMatch(int classIdx, QChar c, bool negate) const;
QString m_pattern;
bool m_valid = false;
QString m_error;
QVector<Inst> m_prog; // the compiled program
int m_start = -1; // index of the start instruction
QVector<QVector<ClassItem>> m_classes;
friend class SamRegexCompiler;
};
} // namespace katecustom
#endif

View File

@ -56,6 +56,7 @@ private Q_SLOTS:
void caretIsPerLine();
void dollarIsPerLine();
void samAnchorIdiom();
void leftmostLongestSubstitution();
};
void TestSamEngine::substituteFirst()
@ -87,8 +88,8 @@ void TestSamEngine::substituteAmp()
void TestSamEngine::substituteGroup()
{
// \1 is the first capture group.
QCOMPARE(runAll(QStringLiteral(",s/(\\w+)=(\\w+)/\\2=\\1/"),
// \1 is the first capture group (sam dialect: use explicit classes, not \w).
QCOMPARE(runAll(QStringLiteral(",s/([a-z]+)=([a-z]+)/\\2=\\1/"),
QStringLiteral("key=value")),
QStringLiteral("value=key"));
}
@ -282,5 +283,15 @@ void TestSamEngine::samAnchorIdiom()
QStringLiteral("> a\n> b\n> c"));
}
void TestSamEngine::leftmostLongestSubstitution()
{
// With the ported sam engine, alternation is leftmost-LONGEST: a|ab on "ab"
// matches "ab", so the whole token is replaced (PCRE would match just "a").
QCOMPARE(runAll(QStringLiteral(",s/a|ab/X/"), QStringLiteral("ab")),
QStringLiteral("X"));
QCOMPARE(runAll(QStringLiteral(",s/ab|a/X/"), QStringLiteral("ab")),
QStringLiteral("X"));
}
QTEST_MAIN(TestSamEngine)
#include "test_samengine.moc"

200
src/sam/test_samregex.cpp Normal file
View File

@ -0,0 +1,200 @@
/*
* SPDX-License-Identifier: LGPL-2.0-or-later
*
* Unit tests for the ported sam regex engine (SamRegex). These pin the two
* properties that motivated the port: leftmost-LONGEST semantics (POSIX, sam),
* and the sam dialect (no PCRE extras). Behaviours cross-checked against
* plan9port src/cmd/sam/regexp.c and sam(1).
*/
#include "samregex.h"
#include <QObject>
#include <QTest>
using namespace katecustom;
class TestSamRegex : public QObject
{
Q_OBJECT
// Match from 0 over the whole text; return "start,end" or "nil".
QString m(const QString &pat, const QString &text)
{
SamRegex re;
if (!re.compile(pat)) {
return QStringLiteral("ERR:%1").arg(re.error());
}
auto r = re.match(text, 0, text.size());
if (!r) {
return QStringLiteral("nil");
}
return QStringLiteral("%1,%2").arg(r->start()).arg(r->end());
}
// Match and return the matched substring, or "nil".
QString ms(const QString &pat, const QString &text, int from = 0)
{
SamRegex re;
if (!re.compile(pat)) {
return QStringLiteral("ERR:%1").arg(re.error());
}
auto r = re.match(text, from, text.size());
if (!r) {
return QStringLiteral("nil");
}
return text.mid(r->start(), r->end() - r->start());
}
private Q_SLOTS:
void literal();
void leftmostLongestAlternation();
void alternationEitherOrder();
void star();
void plusQuest();
void dotNotNewline();
void anchorsPerLine();
void charClass();
void negatedClass();
void range();
void captures();
void escapes();
void fromOffset();
void noMatch();
void nestedGroupsLongest();
void linearNoCatastrophicBacktrack();
};
void TestSamRegex::literal()
{
QCOMPARE(m(QStringLiteral("bar"), QStringLiteral("foobarbaz")), QStringLiteral("3,6"));
}
void TestSamRegex::leftmostLongestAlternation()
{
// THE motivating case: a|ab on "ab" must match "ab" (longest), unlike PCRE.
QCOMPARE(ms(QStringLiteral("a|ab"), QStringLiteral("ab")), QStringLiteral("ab"));
}
void TestSamRegex::alternationEitherOrder()
{
// Order must not matter for leftmost-longest.
QCOMPARE(ms(QStringLiteral("ab|a"), QStringLiteral("ab")), QStringLiteral("ab"));
QCOMPARE(ms(QStringLiteral("a|ab"), QStringLiteral("ab")), QStringLiteral("ab"));
}
void TestSamRegex::star()
{
QCOMPARE(ms(QStringLiteral("a*"), QStringLiteral("aaab")), QStringLiteral("aaa"));
// leftmost: at position 0 matches empty? a* matches "aaa" (longest at 0).
QCOMPARE(ms(QStringLiteral("ba*"), QStringLiteral("baaa")), QStringLiteral("baaa"));
}
void TestSamRegex::plusQuest()
{
QCOMPARE(ms(QStringLiteral("a+"), QStringLiteral("baaa")), QStringLiteral("aaa"));
QCOMPARE(ms(QStringLiteral("ab?c"), QStringLiteral("ac")), QStringLiteral("ac"));
QCOMPARE(ms(QStringLiteral("ab?c"), QStringLiteral("abc")), QStringLiteral("abc"));
}
void TestSamRegex::dotNotNewline()
{
// . does not cross newline (sam).
QCOMPARE(m(QStringLiteral("a.b"), QStringLiteral("a\nb")), QStringLiteral("nil"));
QCOMPARE(ms(QStringLiteral("a.c"), QStringLiteral("abc")), QStringLiteral("abc"));
}
void TestSamRegex::anchorsPerLine()
{
// ^ matches at start of any line; $ at end of any line.
SamRegex re;
QVERIFY(re.compile(QStringLiteral("^b")));
auto r = re.match(QStringLiteral("a\nb\nc"), 0, 5);
QVERIFY(r.has_value());
QCOMPARE(r->start(), 2); // 'b' on line 2
QCOMPARE(r->end(), 3);
SamRegex re2;
QVERIFY(re2.compile(QStringLiteral("a$")));
auto r2 = re2.match(QStringLiteral("ba\nxa\n"), 0, 6);
QVERIFY(r2.has_value());
QCOMPARE(r2->start(), 1); // first 'a' that precedes a newline
QCOMPARE(r2->end(), 2);
}
void TestSamRegex::charClass()
{
QCOMPARE(ms(QStringLiteral("[abc]+"), QStringLiteral("xcababz")), QStringLiteral("cabab"));
}
void TestSamRegex::negatedClass()
{
// [^z]+ crosses nothing special but stops at z; does not match newline.
QCOMPARE(ms(QStringLiteral("[^z]+"), QStringLiteral("abz")), QStringLiteral("ab"));
QCOMPARE(ms(QStringLiteral("[^z]+"), QStringLiteral("a\nb")), QStringLiteral("a"));
}
void TestSamRegex::range()
{
QCOMPARE(ms(QStringLiteral("[0-9]+"), QStringLiteral("ab123cd")), QStringLiteral("123"));
QCOMPARE(ms(QStringLiteral("[a-cx-z]+"), QStringLiteral("abcxyz!")), QStringLiteral("abcxyz"));
}
void TestSamRegex::captures()
{
SamRegex re;
QVERIFY(re.compile(QStringLiteral("([a-z]+)=([a-z]+)")));
const QString text = QStringLiteral("key=value");
auto r = re.match(text, 0, text.size());
QVERIFY(r.has_value());
QCOMPARE(r->captured(1, text), QStringLiteral("key"));
QCOMPARE(r->captured(2, text), QStringLiteral("value"));
}
void TestSamRegex::escapes()
{
// \* is a literal star; \. a literal dot; \n a newline.
QCOMPARE(ms(QStringLiteral("a\\*b"), QStringLiteral("xa*bx")), QStringLiteral("a*b"));
QCOMPARE(ms(QStringLiteral("a\\.b"), QStringLiteral("axb")), QStringLiteral("nil"));
QCOMPARE(ms(QStringLiteral("a\\.b"), QStringLiteral("a.b")), QStringLiteral("a.b"));
QCOMPARE(m(QStringLiteral("a\\nb"), QStringLiteral("a\nb")), QStringLiteral("0,3"));
}
void TestSamRegex::fromOffset()
{
// Search starting past the first match finds the second.
QCOMPARE(ms(QStringLiteral("a"), QStringLiteral("xaya"), 2), QStringLiteral("a"));
SamRegex re;
QVERIFY(re.compile(QStringLiteral("a")));
auto r = re.match(QStringLiteral("xaya"), 2, 4);
QVERIFY(r.has_value());
QCOMPARE(r->start(), 3);
}
void TestSamRegex::noMatch()
{
QCOMPARE(m(QStringLiteral("zzz"), QStringLiteral("abc")), QStringLiteral("nil"));
}
void TestSamRegex::nestedGroupsLongest()
{
// (a|ab)(c|bcd) on "abcd": leftmost-longest should match the whole "abcd"
// (a + bcd), which greedy-ordered backtracking would miss if it locked in
// "a" then "bcd"? Actually both can reach abcd; the key is longest overall.
QCOMPARE(ms(QStringLiteral("(a|ab)(c|bcd)"), QStringLiteral("abcd")),
QStringLiteral("abcd"));
}
void TestSamRegex::linearNoCatastrophicBacktrack()
{
// A pattern that makes PCRE backtrack exponentially: (a*)*b on a long run
// of 'a' with no 'b'. The NFA handles it in linear time; just assert it
// terminates quickly and reports no match.
const QString text(10000, QLatin1Char('a'));
SamRegex re;
QVERIFY(re.compile(QStringLiteral("(a*)*b")));
auto r = re.match(text, 0, text.size());
QVERIFY(!r.has_value());
}
QTEST_MAIN(TestSamRegex)
#include "test_samregex.moc"