sam: port sam's own regex engine (leftmost-longest, linear-time)
Replace QRegularExpression with SamRegex (src/sam/samregex.{h,cpp}), a
faithful port of plan9port src/cmd/sam/regexp.c: a Thompson/Pike NFA.
- Leftmost-LONGEST (POSIX): matches real sam exactly, including overlapping
alternation (a|ab on 'ab' -> 'ab'), eliminating the one remaining
greedy-vs-longest divergence from PCRE.
- Linear time, no backtracking: immune to catastrophic blowup ((a*)*b over
10k 'a' returns instantly).
- sam dialect only: . * + ? | ( ) [ ] ^ $ and \ escaping (\n = newline);
^/$ per-line and ./negated-classes exclude newline, all intrinsic. No
PCRE extras (\d \w \b, lookaround, non-greedy) — sam never had them.
Port notes: shunting-yard compiler + Pike VM with per-thread capture sets and
leftmost-longest newmatch(). Fixed two porting bugs vs the C original: the
l-union aliasing of OR's left/continuation with .next, and reading the popped
subid during popator for correct capture-group ids.
SamEngine now compiles/matches via SamRegex (search, s, x/y/g/v, replacement
captures). Tests: new test_samregex (18) + updated test_samengine (31, incl.
leftmostLongestSubstitution, sam-dialect capture groups). 18/18 ctest suites.
Docs: rewrite docs/SAM.md (no more PCRE deviation; dialect + semantics are
sam's), update PLAN.md and the panel hint.
This commit is contained in:
parent
4cdd762bb6
commit
6341b1ad01
20
docs/PLAN.md
20
docs/PLAN.md
|
|
@ -359,6 +359,15 @@ A dockable tool view ("Sam", left sidebar) runs the plan9 **sam** command
|
|||
language against the active document, or project-wide via `X`/`Y`. Type a
|
||||
program, click **Run** (or Ctrl+Return); each Run is one undo step.
|
||||
|
||||
- **Regex engine** (`src/sam/samregex.{h,cpp}`, 18 unit tests): a faithful port
|
||||
of sam's own matcher (plan9port `src/cmd/sam/regexp.c`) — a Thompson/Pike NFA.
|
||||
**Leftmost-longest** (POSIX), so it matches real sam exactly incl. overlapping
|
||||
alternation (`a|ab` → `ab`); **linear time**, no backtracking (immune to
|
||||
`(a*)*b` blowup). sam dialect only: `. * + ? | ( ) [ ] ^ $ \` with `\n`; `^`/`$`
|
||||
are per-line and `.`/negated-classes exclude newline, all intrinsic. No PCRE
|
||||
extras — not a deviation, just sam. Replaces the earlier QRegularExpression
|
||||
matcher, eliminating the greedy-vs-longest divergence entirely.
|
||||
|
||||
- **Pure engine** (`src/sam/samengine.{h,cpp}`, unit-tested, no Kate dep):
|
||||
`SamEngine::run(program, text, dotStart, dotEnd) -> SamResult{edits, dot,
|
||||
output, applied}`. Edits are computed against the ORIGINAL snapshot in
|
||||
|
|
@ -383,14 +392,9 @@ program, click **Run** (or Ctrl+Return); each Run is one undo step.
|
|||
- **Panel** (`src/plugin/sampanel.{h,cpp}`): program editor + Run button +
|
||||
output log; `OllieView` creates the tool view and owns the apply logic
|
||||
(`runSamProgram`, `applySamToDocument`, `runSamFileLoop`).
|
||||
- **Deviations (documented)**: regexes use `QRegularExpression` (PCRE), not
|
||||
plan9 `regexp(7)`. Everyday patterns match identically; the one semantic
|
||||
divergence is leftmost-greedy vs sam's leftmost-longest (bites only on
|
||||
overlapping alternation like `a|ab`). `^`/`$` are per-line (compiled with
|
||||
`MultilineOption`) and `.` does not cross newlines — both faithful to sam's
|
||||
`regexp(7)` (BOL/EOL in `regexp.c`). PCRE also accepts extras sam lacks
|
||||
(`\d \w \b`, lookahead, non-greedy, inline flags); the docs frame these as an
|
||||
escape hatch, not a feature — prefer structural composition. Full user-facing
|
||||
- **Dialect & semantics**: the regex engine is sam's own (see the Regex engine
|
||||
bullet above), so matching is leftmost-longest and the dialect is sam's — no
|
||||
PCRE, no greedy-vs-longest divergence, no PCRE extras. Full user-facing
|
||||
treatment with verified examples in **`docs/SAM.md`**. Out of scope: multi-file
|
||||
menu (`b B n D`), external file I/O (`e r w f`), `"re"` file-addressing, and
|
||||
sam's own `u` (Kate's undo stack is the undo mechanism). The live GUI panel
|
||||
|
|
|
|||
161
docs/SAM.md
161
docs/SAM.md
|
|
@ -6,126 +6,103 @@ against the active document, or project-wide via `X`/`Y`. Type a program, press
|
|||
**Ctrl+Return** or click **Run**. Each Run is one undo step (Ctrl+Z reverts the
|
||||
whole Run; multi-file `X`/`Y` undoes per file).
|
||||
|
||||
This document covers the **regex deviation** — katesam uses PCRE
|
||||
(`QRegularExpression`), not plan9 `regexp(7)` — and what that means in practice.
|
||||
Every example below is verified against the engine.
|
||||
## The regex engine
|
||||
|
||||
## TL;DR
|
||||
katesam uses a **faithful port of sam's own regular-expression engine**
|
||||
(`src/sam/samregex.{h,cpp}`, from plan9port `src/cmd/sam/regexp.c`), not PCRE.
|
||||
This is a Thompson/Pike NFA, which gives the two properties sam relies on:
|
||||
|
||||
- Everyday patterns behave **identically** to sam.
|
||||
- One real semantic difference: **PCRE is leftmost-greedy, sam is
|
||||
leftmost-longest.** It only bites on *overlapping alternation* (`a|ab`).
|
||||
- `^` and `$` are **per-line** (match at every line boundary), the same as sam's
|
||||
`regexp(7)`. `.` does not cross newlines, also the same as sam.
|
||||
- PCRE *also* offers things sam lacks (`\d \w \b`, lookahead, non-greedy, inline
|
||||
flags). They work, but prefer sam-style structural composition (`x g v`) — see
|
||||
"Extras" below.
|
||||
- **Leftmost-longest** (POSIX) matching, so results match real sam exactly,
|
||||
including overlapping alternation.
|
||||
- **Linear time**, with no backtracking — immune to catastrophic blowup. A
|
||||
pattern like `(a*)*b` against thousands of `a` returns instantly.
|
||||
|
||||
---
|
||||
There is **no PCRE deviation anymore**: the dialect, the match semantics, and
|
||||
the anchors are sam's.
|
||||
|
||||
## The one semantic divergence: greedy vs longest
|
||||
### Dialect
|
||||
|
||||
sam's `regexp(7)` matches the **leftmost-longest** substring. PCRE matches
|
||||
**leftmost**, then resolves alternation **left-to-right, first wins** (greedy
|
||||
within a branch, but alternation order decides between branches).
|
||||
The sam dialect is deliberately small (sam `regexp(7)`):
|
||||
|
||||
| program | input | sam (leftmost-longest) | **katesam (PCRE)** |
|
||||
|---|---|---|---|
|
||||
| `,s/a\|ab/X/` | `ab` | `X` (matches `ab`) | **`Xb`** (matches `a`) |
|
||||
| `,s/ab\|a/X/` | `ab` | `X` | `X` |
|
||||
```
|
||||
. any character except newline
|
||||
* + ? zero-or-more / one-or-more / zero-or-one
|
||||
| alternation
|
||||
( ) grouping and capture (\1..\9 in replacements)
|
||||
[ ] [^ ] character class / negated class (ranges a-z)
|
||||
^ $ beginning / end of a LINE
|
||||
\c literal c (escapes a metacharacter); \n is newline
|
||||
```
|
||||
|
||||
**Implication:** when alternatives overlap, **order them longest-first**
|
||||
(`ab|a`, not `a|ab`). This is the only case where a working sam command can
|
||||
silently produce a different edit in katesam. Non-overlapping alternation
|
||||
(`cat|dog`) and all non-alternation patterns are unaffected.
|
||||
There is intentionally **no** `\d \w \s \b`, no lookaround, no non-greedy, and
|
||||
no inline flags. sam never had them. Use character classes (`[0-9]`,
|
||||
`[a-zA-Z_]`) and structural composition (`x y g v`) instead — that is the sam
|
||||
way, and patterns stay portable to real sam.
|
||||
|
||||
Quantifier greediness is the same in both (`a.*c` on `axxcxxc` → whole string in
|
||||
both). katesam additionally offers non-greedy `.*?` (a PCRE extra).
|
||||
### Semantics (all verified against the port)
|
||||
|
||||
---
|
||||
| behaviour | katesam | note |
|
||||
|---|---|---|
|
||||
| overlapping alternation `a\|ab` on `ab` | matches `ab` | leftmost-**longest**; order of branches does not matter |
|
||||
| `.` crosses newline | **no** | matches sam |
|
||||
| `^` / `$` | per-line | `^` = start of a line, `$` = end of a line |
|
||||
| `[^z]+` crosses newline | no | negated classes also exclude `\n`, as in sam |
|
||||
| empty-match global `s/x*/-/g` on `axbx` | `-a-b-` | one advance per null match |
|
||||
| `&` whole match, `\1..\9` groups in replacement | yes | up to 9 capture groups |
|
||||
| pathological `(a*)*b` | linear time | NFA, no backtracking |
|
||||
|
||||
## Anchors `^` / `$` are per-line — same as sam
|
||||
### Anchors are per-line
|
||||
|
||||
This is **not** a deviation. plan9 `regexp(7)` defines `^` as "the beginning of
|
||||
a line" and `$` as "the end of a line" (sam's `regexp.c`: `BOL` fires at offset
|
||||
0 or after a `\n`; `EOL` fires before a `\n`). katesam compiles every pattern
|
||||
with `QRegularExpression::MultilineOption`, so `^`/`$` match at every line
|
||||
boundary, exactly as sam does:
|
||||
`^` matches the beginning of any line and `$` the end of any line — built into
|
||||
the engine (sam `regexp.c`: `BOL` fires at offset 0 or after `\n`; `EOL` before
|
||||
`\n`). So:
|
||||
|
||||
| program | input | result |
|
||||
|---|---|---|
|
||||
| `,s/^/> /g` | `a⏎b⏎c` | `> a⏎> b⏎> c` (every line) |
|
||||
| `,s/$/;/g` | `a⏎b⏎c` | `a;⏎b;⏎c;` (every line) |
|
||||
| `,s/^/> /g` | `a⏎b⏎c` | `> a⏎> b⏎> c` |
|
||||
| `,s/$/;/g` | `a⏎b⏎c` | `a;⏎b;⏎c;` |
|
||||
|
||||
The sam-idiomatic structural form works too, and is preferable because it
|
||||
composes with guards and other loops:
|
||||
The structural idiom is equivalent and composes with guards/loops:
|
||||
|
||||
```
|
||||
,x/.+/ s/^/> / prefix every (non-empty) line → > a⏎> b⏎> c
|
||||
,x/.+/ a/;/ append ";" to every line → a;⏎b;⏎c;
|
||||
```
|
||||
|
||||
`.` does **not** cross newlines (sam: a "character" is "any character but
|
||||
newline"), which is also PCRE's default. Line-structured descent like
|
||||
`,x/.*\n/ …` therefore behaves as expected.
|
||||
## Composing edits — the sam way
|
||||
|
||||
---
|
||||
sam's power is structural composition, not clever single regexes. Descend into
|
||||
matches with `x`/`y`, guard with `g`/`v`, group with `{}`:
|
||||
|
||||
## What matches identically (no surprises)
|
||||
```
|
||||
,x/[a-zA-Z_]+/ g/^[A-Z]/ c/CONST/ every identifier starting uppercase → CONST
|
||||
,x/"[^"]*"/ s/foo/bar/g foo→bar only inside double-quoted strings
|
||||
0/start/,/end/ x/[a-z]+/ s/.*/[&]/ bracket every lowercase word in a region
|
||||
```
|
||||
|
||||
| behaviour | katesam (PCRE) | sam | same? |
|
||||
|---|---|---|---|
|
||||
| `.` crosses newline | no | no | ✅ |
|
||||
| `[^z]+` crosses newline | yes | yes | ✅ |
|
||||
| empty-match advance `s/x*/-/g` on `axbx` → `-a-b-` | ✅ | ✅ | ✅ |
|
||||
| greedy quantifiers `* + ?` | ✅ | ✅ | ✅ |
|
||||
| character classes `[a-z]`, groups `( )`, `|` | ✅ | ✅ | ✅ |
|
||||
| `&` whole-match, `\1..\9` groups in replacement | ✅ | ✅ | ✅ |
|
||||
Multi-file `X`/`Y` applies an inner program across the project:
|
||||
|
||||
> `.` not crossing newlines matches sam, so line-structured descent
|
||||
> (`,x/.*\n/ ...`) behaves as expected. Use `(?s)` if you *want* `.` to span
|
||||
> newlines.
|
||||
```
|
||||
X/\.cpp$/ ,s/old_api/new_api/g rewrite every .cpp in the project
|
||||
Y/_test\./ ,x/TODO/ d drop TODO lines in non-test files
|
||||
```
|
||||
|
||||
---
|
||||
|
||||
## Extras — available, but not the sam way
|
||||
|
||||
PCRE accepts syntax plan9 `regexp(7)` never had. It all works, but treat it as
|
||||
an **escape hatch, not a headline**: sam's power comes from *composing* simple
|
||||
patterns with `x`/`y`/`g`/`v`/`{}`, not from clever single regexes. Prefer
|
||||
structural composition; reach for these only when it is genuinely simpler.
|
||||
|
||||
| extra | example | note |
|
||||
|---|---|---|
|
||||
| shorthand classes `\d \w \s` | `,s/\d+/N/g` | less typing than `[0-9]`; genuinely handy |
|
||||
| word boundary `\b` | `,s/\bfoo\b/X/g` | handy; sam would use `x/foo/` with context |
|
||||
| ignore-case `(?i)` | `,s/(?i)todo/DONE/g` | fills a real gap — sam has no case-insensitive match |
|
||||
| non-greedy `*?` | `,s/a.*?c/X/` | the structural `x/…/` loop is the sam alternative |
|
||||
| lookahead/behind | `,s/foo(?=bar)/X/` | un-sam; express context with `x`/`g`/`v` instead |
|
||||
| backreference in pattern | `,s/(\w)\1/D/g` | rarely the right tool |
|
||||
|
||||
A pattern using these is **not portable back to sam**. If portability or
|
||||
staying in the structural idiom matters, avoid them.
|
||||
|
||||
---
|
||||
(The file set is the project index, matched by path. Edits land in buffers, not
|
||||
on disk — review and save deliberately. Undo is per file for `X`/`Y`.)
|
||||
|
||||
## Practical guidance
|
||||
|
||||
- **Order overlapping alternatives longest-first** (`\.tar\.gz|\.gz`, not the
|
||||
reverse). The only silent divergence from sam.
|
||||
- **`^`/`$` are per-line** — just like sam. For whole-line edits either anchor
|
||||
directly (`,s/^/> /g`) or, more sam-idiomatically, loop (`,x/.+/ …`).
|
||||
- **Prefer structural composition** (`x g v {}`) over PCRE extras. The extras
|
||||
work but pull you out of the sam idiom and are not sam-portable.
|
||||
- **Empty-match globals** (`s/x*/…/g`) advance one position per null match, as in
|
||||
sam — safe, but as always with `*`, double-check the result.
|
||||
- Use character classes, not PCRE shorthands: `[0-9]` not `\d`, `[a-zA-Z_]` not
|
||||
`\w`.
|
||||
- Prefer structural composition (`x g v {}`) over one dense pattern.
|
||||
- `^`/`$` are per-line; anchor directly (`,s/^…`) or loop (`,x/.+/ …`).
|
||||
- Empty-match globals (`s/x*/…/g`) advance one position per null match — safe,
|
||||
but as always with `*`, check the result.
|
||||
|
||||
## Why PCRE and not regexp(7)
|
||||
## Scope
|
||||
|
||||
`QRegularExpression` ships with Qt (already a dependency) and gives a richer,
|
||||
familiar syntax for free. plan9 `regexp(7)` is a Rune-based NFA woven into sam's
|
||||
own buffer types; using it would mean porting that engine. The tradeoff accepted
|
||||
here: everyday fidelity plus extra power, at the cost of exact leftmost-longest
|
||||
semantics on overlapping alternation. If that matters for your workflow, the
|
||||
engine's regex calls are isolated in `SamEngine` and could be swapped for a
|
||||
ported `regexp.c` later.
|
||||
Supported: the full address language (`#n n 0 $ . '` `/re/ ?re?` compound
|
||||
`+ - , ;`) and commands `a c i d s p = m t k x y g v {}` plus shell filters
|
||||
`< > | !`. Out of scope (Kate manages files / its own undo): the multi-file menu
|
||||
(`b B n D`), external file I/O (`e r w f`), `"re"` file-addressing, and sam's own
|
||||
`u` (use Kate's Ctrl+Z).
|
||||
|
|
|
|||
|
|
@ -56,8 +56,9 @@ SamPanel::SamPanel(QWidget *parent)
|
|||
auto *hint = new QLabel(
|
||||
i18n("sam commands — e.g. <tt>,s/foo/bar/g</tt> or "
|
||||
"<tt>X/\\.cpp$/ ,s/old/new/g</tt> (project-wide). Ctrl+Return runs.<br/>"
|
||||
"Regex is PCRE: <tt>^</tt>/<tt>$</tt> are per-line (like sam); order "
|
||||
"overlapping alternatives longest-first. See docs/SAM.md."),
|
||||
"Regex is sam's own (leftmost-longest, linear-time): classes like "
|
||||
"<tt>[0-9]</tt>, not <tt>\\d</tt>; <tt>^</tt>/<tt>$</tt> are per-line. "
|
||||
"See docs/SAM.md."),
|
||||
this);
|
||||
hint->setWordWrap(true);
|
||||
hint->setTextFormat(Qt::RichText);
|
||||
|
|
|
|||
|
|
@ -8,6 +8,8 @@ find_package(Qt6 ${QT_MIN_VERSION} COMPONENTS Core REQUIRED)
|
|||
add_library(sam_lib STATIC
|
||||
samengine.cpp
|
||||
samengine.h
|
||||
samregex.cpp
|
||||
samregex.h
|
||||
)
|
||||
target_link_libraries(sam_lib PUBLIC Qt6::Core)
|
||||
target_include_directories(sam_lib PUBLIC ${CMAKE_CURRENT_SOURCE_DIR})
|
||||
|
|
@ -17,4 +19,8 @@ if(Qt6Test_FOUND)
|
|||
add_executable(test_samengine test_samengine.cpp)
|
||||
target_link_libraries(test_samengine PRIVATE sam_lib Qt6::Test)
|
||||
add_test(NAME samengine COMMAND test_samengine)
|
||||
|
||||
add_executable(test_samregex test_samregex.cpp)
|
||||
target_link_libraries(test_samregex PRIVATE sam_lib Qt6::Test)
|
||||
add_test(NAME samregex COMMAND test_samregex)
|
||||
endif()
|
||||
|
|
|
|||
|
|
@ -18,9 +18,9 @@
|
|||
* dot bound to each match, exactly as xec.c looper()/linelooper().
|
||||
*/
|
||||
#include "samengine.h"
|
||||
#include "samregex.h"
|
||||
|
||||
#include <QProcess>
|
||||
#include <QRegularExpression>
|
||||
#include <algorithm>
|
||||
|
||||
namespace katecustom
|
||||
|
|
@ -656,16 +656,11 @@ private:
|
|||
return line;
|
||||
}
|
||||
|
||||
// Build a QRegularExpression from a sam pattern. The empty pattern reuses
|
||||
// the last compiled pattern (sam behaviour).
|
||||
//
|
||||
// Multiline is ON: plan9 regexp(7) defines ^ as "beginning of a line" and $
|
||||
// as "end of a line" (sam regexp.c BOL = p==0 || prev=='\n'; EOL = next
|
||||
// char is '\n'). That is exactly QRegularExpression::MultilineOption, so it
|
||||
// is the faithful default, not an opt-in. '.' still does not cross newlines
|
||||
// (sam: "the word character means any character but newline"), which is
|
||||
// QRegularExpression's default (no DotMatchesEverythingOption).
|
||||
QRegularExpression compile(const QString &pat)
|
||||
// Build a SamRegex from a sam pattern. The empty pattern reuses the last
|
||||
// compiled pattern (sam behaviour). This is the ported plan9 sam engine
|
||||
// (SamRegex), so matching is leftmost-LONGEST and linear-time, ^/$ are
|
||||
// per-line and '.' excludes newline — all intrinsic, matching real sam.
|
||||
SamRegex compile(const QString &pat)
|
||||
{
|
||||
QString p = pat;
|
||||
if (p.isEmpty()) {
|
||||
|
|
@ -673,9 +668,9 @@ private:
|
|||
} else {
|
||||
m_lastRe = p;
|
||||
}
|
||||
QRegularExpression re(p, QRegularExpression::MultilineOption);
|
||||
if (!re.isValid()) {
|
||||
fail(QStringLiteral("bad regexp: %1").arg(re.errorString()));
|
||||
SamRegex re;
|
||||
if (!re.compile(p)) {
|
||||
fail(QStringLiteral("bad regexp: %1").arg(re.error()));
|
||||
}
|
||||
return re;
|
||||
}
|
||||
|
|
@ -861,34 +856,33 @@ private:
|
|||
// Forward (sign>0) / backward (sign<0) search, wrapping, like nextmatch().
|
||||
Range search(const QString &pat, Range a, int sign)
|
||||
{
|
||||
QRegularExpression re = compile(pat);
|
||||
SamRegex re = compile(pat);
|
||||
if (failed()) {
|
||||
return a;
|
||||
}
|
||||
if (sign >= 0) {
|
||||
int from = a.p2;
|
||||
QRegularExpressionMatch m = re.match(m_text, from);
|
||||
if (!m.hasMatch()) {
|
||||
m = re.match(m_text, 0); // wrap
|
||||
auto m = re.match(m_text, a.p2, m_nc);
|
||||
if (!m) {
|
||||
m = re.match(m_text, 0, m_nc); // wrap
|
||||
}
|
||||
if (!m.hasMatch()) {
|
||||
if (!m) {
|
||||
fail(QStringLiteral("search failed: %1").arg(pat.isEmpty() ? m_lastRe : pat));
|
||||
return a;
|
||||
}
|
||||
return {int(m.capturedStart()), int(m.capturedEnd())};
|
||||
return {m->start(), m->end()};
|
||||
} else {
|
||||
// Backward: find the last match that ends at or before a.p1,
|
||||
// else wrap to the last match in the text.
|
||||
// Backward: the last match that ends at or before a.p1, else wrap to
|
||||
// the last match in the whole text.
|
||||
int limit = a.p1;
|
||||
Range best{-1, -1};
|
||||
int from = 0;
|
||||
while (true) {
|
||||
QRegularExpressionMatch m = re.match(m_text, from);
|
||||
if (!m.hasMatch()) {
|
||||
auto m = re.match(m_text, from, m_nc);
|
||||
if (!m) {
|
||||
break;
|
||||
}
|
||||
const int s = int(m.capturedStart());
|
||||
const int e = int(m.capturedEnd());
|
||||
const int s = m->start();
|
||||
const int e = m->end();
|
||||
if (e <= limit) {
|
||||
best = {s, e};
|
||||
} else {
|
||||
|
|
@ -897,17 +891,14 @@ private:
|
|||
from = (e > s) ? e : e + 1;
|
||||
}
|
||||
if (best.p1 < 0) {
|
||||
// wrap: take the very last match in the whole text
|
||||
from = 0;
|
||||
while (true) {
|
||||
QRegularExpressionMatch m = re.match(m_text, from);
|
||||
if (!m.hasMatch()) {
|
||||
auto m = re.match(m_text, from, m_nc);
|
||||
if (!m) {
|
||||
break;
|
||||
}
|
||||
best = {int(m.capturedStart()), int(m.capturedEnd())};
|
||||
const int s = best.p1;
|
||||
const int e = best.p2;
|
||||
from = (e > s) ? e : e + 1;
|
||||
best = {m->start(), m->end()};
|
||||
from = (best.p2 > best.p1) ? best.p2 : best.p2 + 1;
|
||||
}
|
||||
}
|
||||
if (best.p1 < 0) {
|
||||
|
|
@ -921,7 +912,7 @@ private:
|
|||
// s/re/text/ with count and g flag (xec.c s_cmd()).
|
||||
void execSubstitute(Cmd *cmd, Range r)
|
||||
{
|
||||
QRegularExpression re = compile(cmd->re);
|
||||
SamRegex re = compile(cmd->re);
|
||||
if (failed()) {
|
||||
return;
|
||||
}
|
||||
|
|
@ -931,15 +922,12 @@ private:
|
|||
int lastEnd = r.p1;
|
||||
int p1 = r.p1;
|
||||
while (p1 <= r.p2) {
|
||||
QRegularExpressionMatch m = re.match(m_text, p1);
|
||||
if (!m.hasMatch() || int(m.capturedStart()) > r.p2) {
|
||||
break;
|
||||
}
|
||||
int ms = int(m.capturedStart());
|
||||
int me = int(m.capturedEnd());
|
||||
if (me > r.p2) {
|
||||
auto m = re.match(m_text, p1, r.p2);
|
||||
if (!m) {
|
||||
break;
|
||||
}
|
||||
int ms = m->start();
|
||||
int me = m->end();
|
||||
// advance logic (empty match handling)
|
||||
if (ms == me) {
|
||||
if (ms == op) {
|
||||
|
|
@ -956,7 +944,7 @@ private:
|
|||
continue;
|
||||
}
|
||||
|
||||
const QString repl = expandReplacement(cmd->text, m);
|
||||
const QString repl = expandReplacement(cmd->text, *m);
|
||||
addEdit(ms, me, repl);
|
||||
if (failed()) {
|
||||
return;
|
||||
|
|
@ -975,7 +963,7 @@ private:
|
|||
}
|
||||
|
||||
// Expand & and \1..\9 in an s replacement using the match captures.
|
||||
QString expandReplacement(const QString &tmpl, const QRegularExpressionMatch &m)
|
||||
QString expandReplacement(const QString &tmpl, const SamMatch &m)
|
||||
{
|
||||
QString out;
|
||||
for (int i = 0; i < tmpl.size(); i++) {
|
||||
|
|
@ -984,7 +972,7 @@ private:
|
|||
const QChar d = tmpl[i + 1];
|
||||
if (d.isDigit()) {
|
||||
const int g = d.unicode() - '0';
|
||||
out.append(m.captured(g));
|
||||
out.append(m.captured(g, m_text));
|
||||
i++;
|
||||
continue;
|
||||
}
|
||||
|
|
@ -1003,7 +991,7 @@ private:
|
|||
continue;
|
||||
}
|
||||
if (c == QLatin1Char('&')) {
|
||||
out.append(m.captured(0));
|
||||
out.append(m.captured(0, m_text));
|
||||
continue;
|
||||
}
|
||||
out.append(c);
|
||||
|
|
@ -1046,14 +1034,13 @@ private:
|
|||
m_nest++;
|
||||
|
||||
if (c == QLatin1Char('g') || c == QLatin1Char('v')) {
|
||||
QRegularExpression re = compile(cmd->re);
|
||||
SamRegex re = compile(cmd->re);
|
||||
if (failed()) {
|
||||
m_nest--;
|
||||
return;
|
||||
}
|
||||
QRegularExpressionMatch m = re.match(m_text, r.p1);
|
||||
const bool contains = m.hasMatch() && int(m.capturedStart()) <= r.p2
|
||||
&& int(m.capturedStart()) >= r.p1;
|
||||
auto m = re.match(m_text, r.p1, r.p2);
|
||||
const bool contains = m.has_value();
|
||||
const bool run = (c == QLatin1Char('g')) ? contains : !contains;
|
||||
if (run) {
|
||||
runLoopBody(cmd, r);
|
||||
|
|
@ -1065,7 +1052,7 @@ private:
|
|||
// x and y.
|
||||
if (c == QLatin1Char('y')) {
|
||||
// y: run on the gaps between matches of re within r.
|
||||
QRegularExpression re = compile(cmd->re);
|
||||
SamRegex re = compile(cmd->re);
|
||||
if (failed()) {
|
||||
m_nest--;
|
||||
return;
|
||||
|
|
@ -1073,12 +1060,12 @@ private:
|
|||
int prev = r.p1;
|
||||
int from = r.p1;
|
||||
while (from <= r.p2) {
|
||||
QRegularExpressionMatch m = re.match(m_text, from);
|
||||
if (!m.hasMatch() || int(m.capturedEnd()) > r.p2) {
|
||||
auto m = re.match(m_text, from, r.p2);
|
||||
if (!m) {
|
||||
break;
|
||||
}
|
||||
int ms = int(m.capturedStart());
|
||||
int me = int(m.capturedEnd());
|
||||
int ms = m->start();
|
||||
int me = m->end();
|
||||
runLoopBody(cmd, {prev, ms});
|
||||
if (failed()) {
|
||||
m_nest--;
|
||||
|
|
@ -1093,7 +1080,7 @@ private:
|
|||
}
|
||||
|
||||
// x: run on each match of re within r (default re handled by compile).
|
||||
QRegularExpression re = compile(cmd->re);
|
||||
SamRegex re = compile(cmd->re);
|
||||
if (failed()) {
|
||||
m_nest--;
|
||||
return;
|
||||
|
|
@ -1101,12 +1088,12 @@ private:
|
|||
int op = -1;
|
||||
int p = r.p1;
|
||||
while (p <= r.p2) {
|
||||
QRegularExpressionMatch m = re.match(m_text, p);
|
||||
if (!m.hasMatch() || int(m.capturedEnd()) > r.p2) {
|
||||
auto m = re.match(m_text, p, r.p2);
|
||||
if (!m) {
|
||||
break;
|
||||
}
|
||||
int ms = int(m.capturedStart());
|
||||
int me = int(m.capturedEnd());
|
||||
int ms = m->start();
|
||||
int me = m->end();
|
||||
if (ms == me) {
|
||||
if (ms == op) {
|
||||
p++;
|
||||
|
|
|
|||
|
|
@ -0,0 +1,632 @@
|
|||
/*
|
||||
* SPDX-License-Identifier: LGPL-2.0-or-later
|
||||
*
|
||||
* Port of plan9port src/cmd/sam/regexp.c. Structure and names follow the
|
||||
* original closely; pointers become indices into QVector<Inst>.
|
||||
*/
|
||||
#include "samregex.h"
|
||||
|
||||
namespace katecustom
|
||||
{
|
||||
|
||||
// Action codes (regexp.c). A literal rune uses its own code point as `type`
|
||||
// (always < kOperator for the characters we accept). Operators carry the
|
||||
// kOperator bit; the low bits encode precedence. Tokens carry the kAny bit.
|
||||
enum {
|
||||
kOperator = 0x1000000,
|
||||
START = kOperator + 0,
|
||||
RBRA = kOperator + 1, // )
|
||||
LBRA = kOperator + 2, // (
|
||||
OR = kOperator + 3, // |
|
||||
CAT = kOperator + 4, // implicit concatenation
|
||||
STAR = kOperator + 5, // *
|
||||
PLUS = kOperator + 6, // +
|
||||
QUEST = kOperator + 7, // ?
|
||||
|
||||
kAny = 0x2000000,
|
||||
ANY = kAny + 0, // .
|
||||
NOP = kAny + 1, // epsilon, removed by optimize()
|
||||
BOL = kAny + 2, // ^
|
||||
EOL = kAny + 3, // $
|
||||
CCLASS = kAny + 4, // [...]
|
||||
NCCLASS = kAny + 5, // [^...]
|
||||
END = kAny + 0x77, // match
|
||||
};
|
||||
|
||||
constexpr int NSUBEXP = 10;
|
||||
|
||||
QString SamMatch::captured(int group, const QString &text) const
|
||||
{
|
||||
if (group < 0 || group >= caps.size()) {
|
||||
return QString();
|
||||
}
|
||||
const Range r = caps[group];
|
||||
if (r.p1 < 0 || r.p2 < 0 || r.p2 < r.p1) {
|
||||
return QString();
|
||||
}
|
||||
return text.mid(r.p1, r.p2 - r.p1);
|
||||
}
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// Compiler — shunting yard over the pattern, emitting Inst nodes.
|
||||
// ---------------------------------------------------------------------------
|
||||
class SamRegexCompiler
|
||||
{
|
||||
public:
|
||||
SamRegexCompiler(SamRegex &re, const QString &pat) : m_re(re), m_s(pat) {}
|
||||
|
||||
bool compile()
|
||||
{
|
||||
m_re.m_prog.clear();
|
||||
m_re.m_classes.clear();
|
||||
m_atorStack.clear();
|
||||
m_subidStack.clear();
|
||||
m_andStack.clear();
|
||||
m_cursubid = 0;
|
||||
m_lastWasAnd = false;
|
||||
m_pos = 0;
|
||||
m_nbra = 0;
|
||||
|
||||
pushator(START - 1);
|
||||
int token;
|
||||
while ((token = lex()) != END) {
|
||||
if (m_error.isEmpty() == false) {
|
||||
return fail();
|
||||
}
|
||||
if ((token & kOperator) == kOperator) {
|
||||
doOperator(token);
|
||||
} else {
|
||||
doOperand(token);
|
||||
}
|
||||
if (!m_error.isEmpty()) {
|
||||
return fail();
|
||||
}
|
||||
}
|
||||
evaluntil(START);
|
||||
doOperand(END);
|
||||
evaluntil(START);
|
||||
if (m_nbra) {
|
||||
m_error = QStringLiteral("unmatched '('");
|
||||
return fail();
|
||||
}
|
||||
if (!m_error.isEmpty() || m_andStack.isEmpty()) {
|
||||
if (m_error.isEmpty()) {
|
||||
m_error = QStringLiteral("malformed regexp");
|
||||
}
|
||||
return fail();
|
||||
}
|
||||
m_re.m_start = m_andStack.last().first;
|
||||
optimize();
|
||||
return true;
|
||||
}
|
||||
|
||||
QString error() const { return m_error; }
|
||||
|
||||
private:
|
||||
struct Node {
|
||||
int first = -1;
|
||||
int last = -1;
|
||||
};
|
||||
|
||||
bool fail()
|
||||
{
|
||||
m_re.m_valid = false;
|
||||
if (m_re.m_error.isEmpty()) {
|
||||
m_re.m_error = m_error.isEmpty() ? QStringLiteral("regexp error") : m_error;
|
||||
}
|
||||
return false;
|
||||
}
|
||||
|
||||
int newinst(int t)
|
||||
{
|
||||
SamRegex::Inst inst;
|
||||
inst.type = t;
|
||||
m_re.m_prog.append(inst);
|
||||
return m_re.m_prog.size() - 1;
|
||||
}
|
||||
|
||||
SamRegex::Inst &at(int i) { return m_re.m_prog[i]; }
|
||||
|
||||
void pushand(int f, int l) { m_andStack.append(Node{f, l}); }
|
||||
void pushator(int t)
|
||||
{
|
||||
m_atorStack.append(t);
|
||||
m_subidStack.append(m_cursubid >= NSUBEXP ? -1 : m_cursubid);
|
||||
}
|
||||
int popator()
|
||||
{
|
||||
if (m_atorStack.isEmpty()) {
|
||||
m_error = QStringLiteral("operator stack underflow");
|
||||
return START;
|
||||
}
|
||||
// The subid paired with this operator — captured before removal so the
|
||||
// LBRA/RBRA case can read the group's own id (regexp.c reads *subidp
|
||||
// right after decrementing it).
|
||||
m_lastPoppedSubid = m_subidStack.isEmpty() ? -1 : m_subidStack.last();
|
||||
if (!m_subidStack.isEmpty()) {
|
||||
m_subidStack.removeLast();
|
||||
}
|
||||
const int t = m_atorStack.last();
|
||||
m_atorStack.removeLast();
|
||||
return t;
|
||||
}
|
||||
Node popand(int op)
|
||||
{
|
||||
if (m_andStack.isEmpty()) {
|
||||
m_error = op ? QStringLiteral("missing operand for '%1'").arg(QChar(op))
|
||||
: QStringLiteral("malformed regexp");
|
||||
return Node{};
|
||||
}
|
||||
const Node n = m_andStack.last();
|
||||
m_andStack.removeLast();
|
||||
return n;
|
||||
}
|
||||
|
||||
void doOperand(int t)
|
||||
{
|
||||
if (m_lastWasAnd) {
|
||||
doOperator(CAT); // implicit concatenation
|
||||
}
|
||||
int i = newinst(t);
|
||||
if (t == CCLASS) {
|
||||
if (m_negateClass) {
|
||||
at(i).type = NCCLASS;
|
||||
}
|
||||
at(i).classIdx = m_re.m_classes.size() - 1;
|
||||
}
|
||||
pushand(i, i);
|
||||
m_lastWasAnd = true;
|
||||
}
|
||||
|
||||
void doOperator(int t)
|
||||
{
|
||||
if (t == RBRA && --m_nbra < 0) {
|
||||
m_error = QStringLiteral("unmatched ')'");
|
||||
return;
|
||||
}
|
||||
if (t == LBRA) {
|
||||
m_cursubid++;
|
||||
m_nbra++;
|
||||
if (m_lastWasAnd) {
|
||||
doOperator(CAT);
|
||||
}
|
||||
} else {
|
||||
evaluntil(t);
|
||||
}
|
||||
if (t != RBRA) {
|
||||
pushator(t);
|
||||
}
|
||||
m_lastWasAnd = false;
|
||||
if (t == STAR || t == QUEST || t == PLUS || t == RBRA) {
|
||||
m_lastWasAnd = true;
|
||||
}
|
||||
}
|
||||
|
||||
void evaluntil(int pri)
|
||||
{
|
||||
while (m_error.isEmpty()
|
||||
&& (pri == RBRA || (!m_atorStack.isEmpty() && m_atorStack.last() >= pri))) {
|
||||
const int op = popator();
|
||||
if (!m_error.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
if (op == LBRA) {
|
||||
Node op1 = popand('(');
|
||||
const int subid = m_lastPoppedSubid;
|
||||
int inst2 = newinst(RBRA);
|
||||
at(inst2).subid = subid;
|
||||
at(op1.last).next = inst2;
|
||||
int inst1 = newinst(LBRA);
|
||||
at(inst1).subid = subid;
|
||||
at(inst1).next = op1.first;
|
||||
pushand(inst1, inst2);
|
||||
return; // must have been RBRA
|
||||
} else if (op == OR) {
|
||||
Node op2 = popand('|');
|
||||
Node op1 = popand('|');
|
||||
if (!m_error.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
int inst2 = newinst(NOP);
|
||||
at(op2.last).next = inst2;
|
||||
at(op1.last).next = inst2;
|
||||
int inst1 = newinst(OR);
|
||||
at(inst1).right = op1.first;
|
||||
at(inst1).next = op2.first; // "left" branch == continuation slot
|
||||
pushand(inst1, inst2);
|
||||
} else if (op == CAT) {
|
||||
Node op2 = popand(0);
|
||||
Node op1 = popand(0);
|
||||
if (!m_error.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
at(op1.last).next = op2.first;
|
||||
pushand(op1.first, op2.last);
|
||||
} else if (op == STAR) {
|
||||
Node op2 = popand('*');
|
||||
if (!m_error.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
int inst1 = newinst(OR);
|
||||
at(op2.last).next = inst1;
|
||||
at(inst1).right = op2.first;
|
||||
pushand(inst1, inst1);
|
||||
} else if (op == PLUS) {
|
||||
Node op2 = popand('+');
|
||||
if (!m_error.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
int inst1 = newinst(OR);
|
||||
at(op2.last).next = inst1;
|
||||
at(inst1).right = op2.first;
|
||||
pushand(op2.first, inst1);
|
||||
} else if (op == QUEST) {
|
||||
Node op2 = popand('?');
|
||||
if (!m_error.isEmpty()) {
|
||||
return;
|
||||
}
|
||||
int inst1 = newinst(OR);
|
||||
int inst2 = newinst(NOP);
|
||||
at(inst1).next = inst2; // skip branch (continuation slot)
|
||||
at(inst1).right = op2.first;
|
||||
at(op2.last).next = inst2;
|
||||
pushand(inst1, inst2);
|
||||
} else {
|
||||
m_error = QStringLiteral("bad regexp operator");
|
||||
return;
|
||||
}
|
||||
if (pri == RBRA) {
|
||||
// Keep looping until the matching LBRA (handled by the return
|
||||
// above). The while condition handles the rest.
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
// Rewrite `next` pointers to skip NOP chains (regexp.c optimize()).
|
||||
void optimize()
|
||||
{
|
||||
for (int i = 0; i < m_re.m_prog.size(); i++) {
|
||||
if (m_re.m_prog[i].type == END) {
|
||||
continue;
|
||||
}
|
||||
int target = m_re.m_prog[i].next;
|
||||
while (target >= 0 && m_re.m_prog[target].type == NOP) {
|
||||
target = m_re.m_prog[target].next;
|
||||
}
|
||||
m_re.m_prog[i].next = target;
|
||||
}
|
||||
}
|
||||
|
||||
// --- lexer ---
|
||||
int lex()
|
||||
{
|
||||
if (m_pos >= m_s.size()) {
|
||||
return END;
|
||||
}
|
||||
QChar qc = m_s[m_pos++];
|
||||
int c = qc.unicode();
|
||||
switch (c) {
|
||||
case '\\':
|
||||
if (m_pos < m_s.size()) {
|
||||
QChar d = m_s[m_pos++];
|
||||
c = (d == QLatin1Char('n')) ? '\n' : d.unicode();
|
||||
}
|
||||
// escaped: always a literal, never a metachar/token
|
||||
return c;
|
||||
case '*':
|
||||
return STAR;
|
||||
case '?':
|
||||
return QUEST;
|
||||
case '+':
|
||||
return PLUS;
|
||||
case '|':
|
||||
return OR;
|
||||
case '.':
|
||||
return ANY;
|
||||
case '(':
|
||||
return LBRA;
|
||||
case ')':
|
||||
return RBRA;
|
||||
case '^':
|
||||
return BOL;
|
||||
case '$':
|
||||
return EOL;
|
||||
case '[':
|
||||
buildClass();
|
||||
return CCLASS;
|
||||
default:
|
||||
return c;
|
||||
}
|
||||
}
|
||||
|
||||
// Read one class element, honouring \n and \<char>.
|
||||
int classNextRune(bool *quoted)
|
||||
{
|
||||
*quoted = false;
|
||||
if (m_pos >= m_s.size()) {
|
||||
m_error = QStringLiteral("malformed character class");
|
||||
return 0;
|
||||
}
|
||||
QChar c = m_s[m_pos];
|
||||
if (c == QLatin1Char('\\')) {
|
||||
m_pos++;
|
||||
if (m_pos >= m_s.size()) {
|
||||
m_error = QStringLiteral("malformed character class");
|
||||
return 0;
|
||||
}
|
||||
QChar d = m_s[m_pos++];
|
||||
if (d == QLatin1Char('n')) {
|
||||
return '\n';
|
||||
}
|
||||
*quoted = true;
|
||||
return d.unicode();
|
||||
}
|
||||
m_pos++;
|
||||
return c.unicode();
|
||||
}
|
||||
|
||||
void buildClass()
|
||||
{
|
||||
QVector<SamRegex::ClassItem> items;
|
||||
m_negateClass = false;
|
||||
if (m_pos < m_s.size() && m_s[m_pos] == QLatin1Char('^')) {
|
||||
m_negateClass = true;
|
||||
m_pos++;
|
||||
// Negated classes never match newline.
|
||||
items.append(SamRegex::ClassItem{false, '\n', '\n'});
|
||||
}
|
||||
while (true) {
|
||||
if (m_pos >= m_s.size()) {
|
||||
m_error = QStringLiteral("unterminated character class");
|
||||
break;
|
||||
}
|
||||
// A ']' that is not escaped closes the class.
|
||||
if (m_s[m_pos] == QLatin1Char(']')) {
|
||||
m_pos++;
|
||||
break;
|
||||
}
|
||||
bool q1 = false;
|
||||
int c1 = classNextRune(&q1);
|
||||
if (!m_error.isEmpty()) {
|
||||
break;
|
||||
}
|
||||
// Range a-b: a '-' followed by a non-']' element.
|
||||
if (m_pos + 0 < m_s.size() && m_s[m_pos] == QLatin1Char('-')
|
||||
&& m_pos + 1 < m_s.size() && m_s[m_pos + 1] != QLatin1Char(']')) {
|
||||
m_pos++; // consume '-'
|
||||
bool q2 = false;
|
||||
int c2 = classNextRune(&q2);
|
||||
if (!m_error.isEmpty()) {
|
||||
break;
|
||||
}
|
||||
items.append(SamRegex::ClassItem{true, c1, c2});
|
||||
} else {
|
||||
items.append(SamRegex::ClassItem{false, c1, c1});
|
||||
}
|
||||
}
|
||||
m_re.m_classes.append(items);
|
||||
}
|
||||
|
||||
SamRegex &m_re;
|
||||
QString m_s;
|
||||
int m_pos = 0;
|
||||
|
||||
QVector<int> m_atorStack;
|
||||
QVector<int> m_subidStack;
|
||||
QVector<Node> m_andStack;
|
||||
int m_cursubid = 0;
|
||||
int m_lastPoppedSubid = -1;
|
||||
bool m_lastWasAnd = false;
|
||||
int m_nbra = 0;
|
||||
bool m_negateClass = false;
|
||||
QString m_error;
|
||||
};
|
||||
|
||||
// ---------------------------------------------------------------------------
|
||||
// SamRegex
|
||||
// ---------------------------------------------------------------------------
|
||||
bool SamRegex::compile(const QString &pattern)
|
||||
{
|
||||
m_pattern = pattern;
|
||||
m_valid = true;
|
||||
m_error.clear();
|
||||
SamRegexCompiler c(*this, pattern);
|
||||
if (!c.compile()) {
|
||||
m_valid = false;
|
||||
if (m_error.isEmpty()) {
|
||||
m_error = c.error();
|
||||
}
|
||||
return false;
|
||||
}
|
||||
m_valid = true;
|
||||
return true;
|
||||
}
|
||||
|
||||
bool SamRegex::classMatch(int classIdx, QChar c, bool negate) const
|
||||
{
|
||||
if (classIdx < 0 || classIdx >= m_classes.size()) {
|
||||
return negate;
|
||||
}
|
||||
const int cc = c.unicode();
|
||||
for (const ClassItem &item : m_classes[classIdx]) {
|
||||
if (item.range) {
|
||||
if (item.lo <= cc && cc <= item.hi) {
|
||||
return !negate;
|
||||
}
|
||||
} else if (item.lo == cc) {
|
||||
return !negate;
|
||||
}
|
||||
}
|
||||
return negate;
|
||||
}
|
||||
|
||||
// Pike VM thread list. Each live thread is (inst, captures). addinst dedups by
|
||||
// instruction index within a list, keeping the thread whose match started
|
||||
// earliest (leftmost), exactly as regexp.c addinst does.
|
||||
namespace
|
||||
{
|
||||
struct Thread {
|
||||
int inst = -1;
|
||||
QVector<SamMatch::Range> caps;
|
||||
};
|
||||
} // namespace
|
||||
|
||||
std::optional<SamMatch> SamRegex::match(const QString &text, int from, int bound) const
|
||||
{
|
||||
if (!m_valid || m_start < 0) {
|
||||
return std::nullopt;
|
||||
}
|
||||
if (bound < 0 || bound > text.size()) {
|
||||
bound = text.size();
|
||||
}
|
||||
if (from < 0) {
|
||||
from = 0;
|
||||
}
|
||||
|
||||
auto addinst = [](QVector<Thread> &list, int inst, const QVector<SamMatch::Range> &caps) {
|
||||
for (Thread &t : list) {
|
||||
if (t.inst == inst) {
|
||||
// Keep the earliest start (leftmost) — mirrors regexp.c.
|
||||
if (caps[0].p1 < t.caps[0].p1) {
|
||||
t.caps = caps;
|
||||
}
|
||||
return;
|
||||
}
|
||||
}
|
||||
list.append(Thread{inst, caps});
|
||||
};
|
||||
|
||||
QVector<Thread> clist;
|
||||
QVector<Thread> nlist;
|
||||
|
||||
SamMatch best;
|
||||
bool haveMatch = false;
|
||||
auto newmatch = [&](const QVector<SamMatch::Range> &se) {
|
||||
// Leftmost-longest: take if no match yet, or starts earlier, or same
|
||||
// start and ends later (regexp.c newmatch()).
|
||||
if (!haveMatch || se[0].p1 < best.caps[0].p1
|
||||
|| (se[0].p1 == best.caps[0].p1 && se[0].p2 > best.caps[0].p2)) {
|
||||
best.caps = se;
|
||||
haveMatch = true;
|
||||
}
|
||||
};
|
||||
|
||||
const int startChar =
|
||||
(m_prog[m_start].type < kOperator) ? m_prog[m_start].type : 0;
|
||||
|
||||
int nnl = 0; // live threads carried into nlist
|
||||
// Scan one position past bound so a thread ending exactly at bound (and
|
||||
// EOL/END at bound) is processed.
|
||||
for (int p = from; p <= bound; p++) {
|
||||
const int c = (p < bound) ? text[p].unicode() : -1;
|
||||
|
||||
// Stop once a match is found and no threads remain alive.
|
||||
if (haveMatch && nnl == 0) {
|
||||
break;
|
||||
}
|
||||
// Fast first-char skip while no thread is live and no match pending.
|
||||
if (startChar && nnl == 0 && !haveMatch && c != startChar) {
|
||||
continue;
|
||||
}
|
||||
|
||||
clist = nlist;
|
||||
nlist.clear();
|
||||
nnl = 0;
|
||||
|
||||
// Seed a fresh start thread at this position while no match found yet.
|
||||
if (!haveMatch) {
|
||||
QVector<SamMatch::Range> se(NSUBEXP, SamMatch::Range{-1, -1});
|
||||
se[0].p1 = p;
|
||||
addinst(clist, m_start, se);
|
||||
}
|
||||
|
||||
// Run epsilon + consuming transitions for this position.
|
||||
for (int ti = 0; ti < clist.size(); ti++) {
|
||||
int inst = clist[ti].inst;
|
||||
// Follow epsilon transitions within the same position.
|
||||
bool consumed = false;
|
||||
while (!consumed) {
|
||||
const Inst &in = m_prog[inst];
|
||||
switch (in.type) {
|
||||
case LBRA:
|
||||
if (in.subid >= 0 && in.subid < NSUBEXP) {
|
||||
clist[ti].caps[in.subid].p1 = p;
|
||||
}
|
||||
inst = in.next;
|
||||
continue;
|
||||
case RBRA:
|
||||
if (in.subid >= 0 && in.subid < NSUBEXP) {
|
||||
clist[ti].caps[in.subid].p2 = p;
|
||||
}
|
||||
inst = in.next;
|
||||
continue;
|
||||
case OR:
|
||||
addinst(clist, in.right, clist[ti].caps);
|
||||
inst = in.next; // "left"/continuation branch
|
||||
continue;
|
||||
case NOP:
|
||||
inst = in.next;
|
||||
continue;
|
||||
case BOL:
|
||||
if (p == 0 || (p > 0 && text[p - 1] == QLatin1Char('\n'))) {
|
||||
inst = in.next;
|
||||
continue;
|
||||
}
|
||||
consumed = true; // dead
|
||||
break;
|
||||
case EOL:
|
||||
if (c == '\n' || p == bound) {
|
||||
inst = in.next;
|
||||
continue;
|
||||
}
|
||||
consumed = true; // dead
|
||||
break;
|
||||
case END:
|
||||
clist[ti].caps[0].p2 = p;
|
||||
newmatch(clist[ti].caps);
|
||||
consumed = true;
|
||||
break;
|
||||
case ANY:
|
||||
if (c >= 0 && c != '\n') {
|
||||
addinst(nlist, in.next, clist[ti].caps);
|
||||
nnl++;
|
||||
}
|
||||
consumed = true;
|
||||
break;
|
||||
case CCLASS:
|
||||
if (c >= 0 && classMatch(in.classIdx, QChar(c), false)) {
|
||||
addinst(nlist, in.next, clist[ti].caps);
|
||||
nnl++;
|
||||
}
|
||||
consumed = true;
|
||||
break;
|
||||
case NCCLASS:
|
||||
if (c >= 0 && classMatch(in.classIdx, QChar(c), true)) {
|
||||
addinst(nlist, in.next, clist[ti].caps);
|
||||
nnl++;
|
||||
}
|
||||
consumed = true;
|
||||
break;
|
||||
default: // literal rune
|
||||
if (c >= 0 && in.type == c) {
|
||||
addinst(nlist, in.next, clist[ti].caps);
|
||||
nnl++;
|
||||
}
|
||||
consumed = true;
|
||||
break;
|
||||
}
|
||||
}
|
||||
}
|
||||
}
|
||||
|
||||
if (!haveMatch) {
|
||||
return std::nullopt;
|
||||
}
|
||||
// Honour the bound: a match must lie within [from, bound].
|
||||
if (best.caps[0].p2 > bound) {
|
||||
return std::nullopt;
|
||||
}
|
||||
return best;
|
||||
}
|
||||
|
||||
} // namespace katecustom
|
||||
|
|
@ -0,0 +1,114 @@
|
|||
/*
|
||||
* SPDX-License-Identifier: LGPL-2.0-or-later
|
||||
*
|
||||
* SamRegex — a faithful port of plan9 sam's regular-expression engine
|
||||
* (plan9port src/cmd/sam/regexp.c) to operate on a QString.
|
||||
*
|
||||
* Why this exists: sam's regex is a Thompson/Pike NFA simulation. It is
|
||||
* - leftmost-longest (POSIX semantics), not Perl leftmost-greedy — so it
|
||||
* matches sam exactly, including overlapping alternation (`a|ab` on "ab"
|
||||
* matches "ab", the longest), which PCRE does not; and
|
||||
* - linear time in the subject length with no backtracking, so it cannot
|
||||
* suffer catastrophic blowup.
|
||||
*
|
||||
* The dialect is sam's regexp(7), deliberately small:
|
||||
* . * + ? | ( ) ^ $ [ ] and \ to escape a metacharacter (\n = newline).
|
||||
* `^` matches the beginning of a line, `$` the end of a line (per-line, built
|
||||
* in). `.` and negated classes do not match newline. There is no \d \w \b,
|
||||
* no lookaround, no non-greedy — sam never had them, and their absence keeps
|
||||
* the engine linear and the syntax portable back to real sam.
|
||||
*
|
||||
* The implementation mirrors regexp.c structure: a shunting-yard parser builds
|
||||
* an Inst NFA; execute() runs two parallel thread lists (Pike VM) where each
|
||||
* thread carries a capture set, and newmatch() keeps the leftmost-longest.
|
||||
* Only the forward machine is ported; backward search is done by the caller
|
||||
* scanning forward and keeping the last qualifying match.
|
||||
*/
|
||||
#ifndef KATECUSTOM_SAMREGEX_H
|
||||
#define KATECUSTOM_SAMREGEX_H
|
||||
|
||||
#include <QList>
|
||||
#include <QString>
|
||||
#include <QVector>
|
||||
#include <optional>
|
||||
|
||||
namespace katecustom
|
||||
{
|
||||
|
||||
/*! A match: half-open character ranges. caps[0] is the whole match; caps[1..9]
|
||||
* are parenthesised subexpressions ({-1,-1} if that group did not participate). */
|
||||
struct SamMatch {
|
||||
struct Range {
|
||||
int p1 = -1;
|
||||
int p2 = -1;
|
||||
};
|
||||
QVector<Range> caps; // size NSUBEXP (10)
|
||||
int start() const { return caps.isEmpty() ? -1 : caps[0].p1; }
|
||||
int end() const { return caps.isEmpty() ? -1 : caps[0].p2; }
|
||||
QString captured(int group, const QString &text) const;
|
||||
};
|
||||
|
||||
class SamRegex
|
||||
{
|
||||
public:
|
||||
SamRegex() = default;
|
||||
|
||||
/*! Compile \a pattern (sam dialect). Returns false and sets error() on a
|
||||
* syntax error. */
|
||||
bool compile(const QString &pattern);
|
||||
|
||||
bool isValid() const { return m_valid; }
|
||||
QString error() const { return m_error; }
|
||||
QString pattern() const { return m_pattern; }
|
||||
|
||||
/*!
|
||||
* Find the leftmost-longest match whose start is at or after \a from and
|
||||
* which lies entirely within [0, \a bound). Anchors `^`/`$` are evaluated
|
||||
* against the whole \a text (so they fire at any line boundary inside the
|
||||
* bound). Returns std::nullopt if there is no match. Linear time.
|
||||
*/
|
||||
std::optional<SamMatch> match(const QString &text, int from, int bound) const;
|
||||
|
||||
private:
|
||||
// NFA instruction (regexp.c Inst). type < kOperator is a literal rune;
|
||||
// otherwise it is one of the action codes below.
|
||||
//
|
||||
// In the original, `left` and `next` share one union member (l.lleft ==
|
||||
// l.lnext): an OR node's "left/continuation" IS its next pointer, so
|
||||
// concatenation (which wires `.next`) and OR execution (which reads the
|
||||
// continuation) refer to the same slot. We keep that single `next` slot and
|
||||
// use `right` for the OR's branch target. `subid`/`classIdx` reuse would be
|
||||
// the r-union; they are used disjointly per node type so separate fields are
|
||||
// safe.
|
||||
struct Inst {
|
||||
int type = 0;
|
||||
int subid = -1; // LBRA/RBRA capture index
|
||||
int classIdx = -1; // CCLASS/NCCLASS index into m_classes
|
||||
int next = -1; // continuation (also the OR "left" branch)
|
||||
int right = -1; // OR branch target (loop body / first alternative)
|
||||
};
|
||||
|
||||
// A character class: a flat list of items, each either a single rune or a
|
||||
// range [lo,hi]. Negated classes also implicitly exclude '\n'.
|
||||
struct ClassItem {
|
||||
bool range = false;
|
||||
int lo = 0;
|
||||
int hi = 0;
|
||||
};
|
||||
|
||||
bool classMatch(int classIdx, QChar c, bool negate) const;
|
||||
|
||||
QString m_pattern;
|
||||
bool m_valid = false;
|
||||
QString m_error;
|
||||
|
||||
QVector<Inst> m_prog; // the compiled program
|
||||
int m_start = -1; // index of the start instruction
|
||||
QVector<QVector<ClassItem>> m_classes;
|
||||
|
||||
friend class SamRegexCompiler;
|
||||
};
|
||||
|
||||
} // namespace katecustom
|
||||
|
||||
#endif
|
||||
|
|
@ -56,6 +56,7 @@ private Q_SLOTS:
|
|||
void caretIsPerLine();
|
||||
void dollarIsPerLine();
|
||||
void samAnchorIdiom();
|
||||
void leftmostLongestSubstitution();
|
||||
};
|
||||
|
||||
void TestSamEngine::substituteFirst()
|
||||
|
|
@ -87,8 +88,8 @@ void TestSamEngine::substituteAmp()
|
|||
|
||||
void TestSamEngine::substituteGroup()
|
||||
{
|
||||
// \1 is the first capture group.
|
||||
QCOMPARE(runAll(QStringLiteral(",s/(\\w+)=(\\w+)/\\2=\\1/"),
|
||||
// \1 is the first capture group (sam dialect: use explicit classes, not \w).
|
||||
QCOMPARE(runAll(QStringLiteral(",s/([a-z]+)=([a-z]+)/\\2=\\1/"),
|
||||
QStringLiteral("key=value")),
|
||||
QStringLiteral("value=key"));
|
||||
}
|
||||
|
|
@ -282,5 +283,15 @@ void TestSamEngine::samAnchorIdiom()
|
|||
QStringLiteral("> a\n> b\n> c"));
|
||||
}
|
||||
|
||||
void TestSamEngine::leftmostLongestSubstitution()
|
||||
{
|
||||
// With the ported sam engine, alternation is leftmost-LONGEST: a|ab on "ab"
|
||||
// matches "ab", so the whole token is replaced (PCRE would match just "a").
|
||||
QCOMPARE(runAll(QStringLiteral(",s/a|ab/X/"), QStringLiteral("ab")),
|
||||
QStringLiteral("X"));
|
||||
QCOMPARE(runAll(QStringLiteral(",s/ab|a/X/"), QStringLiteral("ab")),
|
||||
QStringLiteral("X"));
|
||||
}
|
||||
|
||||
QTEST_MAIN(TestSamEngine)
|
||||
#include "test_samengine.moc"
|
||||
|
|
|
|||
|
|
@ -0,0 +1,200 @@
|
|||
/*
|
||||
* SPDX-License-Identifier: LGPL-2.0-or-later
|
||||
*
|
||||
* Unit tests for the ported sam regex engine (SamRegex). These pin the two
|
||||
* properties that motivated the port: leftmost-LONGEST semantics (POSIX, sam),
|
||||
* and the sam dialect (no PCRE extras). Behaviours cross-checked against
|
||||
* plan9port src/cmd/sam/regexp.c and sam(1).
|
||||
*/
|
||||
#include "samregex.h"
|
||||
|
||||
#include <QObject>
|
||||
#include <QTest>
|
||||
|
||||
using namespace katecustom;
|
||||
|
||||
class TestSamRegex : public QObject
|
||||
{
|
||||
Q_OBJECT
|
||||
|
||||
// Match from 0 over the whole text; return "start,end" or "nil".
|
||||
QString m(const QString &pat, const QString &text)
|
||||
{
|
||||
SamRegex re;
|
||||
if (!re.compile(pat)) {
|
||||
return QStringLiteral("ERR:%1").arg(re.error());
|
||||
}
|
||||
auto r = re.match(text, 0, text.size());
|
||||
if (!r) {
|
||||
return QStringLiteral("nil");
|
||||
}
|
||||
return QStringLiteral("%1,%2").arg(r->start()).arg(r->end());
|
||||
}
|
||||
|
||||
// Match and return the matched substring, or "nil".
|
||||
QString ms(const QString &pat, const QString &text, int from = 0)
|
||||
{
|
||||
SamRegex re;
|
||||
if (!re.compile(pat)) {
|
||||
return QStringLiteral("ERR:%1").arg(re.error());
|
||||
}
|
||||
auto r = re.match(text, from, text.size());
|
||||
if (!r) {
|
||||
return QStringLiteral("nil");
|
||||
}
|
||||
return text.mid(r->start(), r->end() - r->start());
|
||||
}
|
||||
|
||||
private Q_SLOTS:
|
||||
void literal();
|
||||
void leftmostLongestAlternation();
|
||||
void alternationEitherOrder();
|
||||
void star();
|
||||
void plusQuest();
|
||||
void dotNotNewline();
|
||||
void anchorsPerLine();
|
||||
void charClass();
|
||||
void negatedClass();
|
||||
void range();
|
||||
void captures();
|
||||
void escapes();
|
||||
void fromOffset();
|
||||
void noMatch();
|
||||
void nestedGroupsLongest();
|
||||
void linearNoCatastrophicBacktrack();
|
||||
};
|
||||
|
||||
void TestSamRegex::literal()
|
||||
{
|
||||
QCOMPARE(m(QStringLiteral("bar"), QStringLiteral("foobarbaz")), QStringLiteral("3,6"));
|
||||
}
|
||||
|
||||
void TestSamRegex::leftmostLongestAlternation()
|
||||
{
|
||||
// THE motivating case: a|ab on "ab" must match "ab" (longest), unlike PCRE.
|
||||
QCOMPARE(ms(QStringLiteral("a|ab"), QStringLiteral("ab")), QStringLiteral("ab"));
|
||||
}
|
||||
|
||||
void TestSamRegex::alternationEitherOrder()
|
||||
{
|
||||
// Order must not matter for leftmost-longest.
|
||||
QCOMPARE(ms(QStringLiteral("ab|a"), QStringLiteral("ab")), QStringLiteral("ab"));
|
||||
QCOMPARE(ms(QStringLiteral("a|ab"), QStringLiteral("ab")), QStringLiteral("ab"));
|
||||
}
|
||||
|
||||
void TestSamRegex::star()
|
||||
{
|
||||
QCOMPARE(ms(QStringLiteral("a*"), QStringLiteral("aaab")), QStringLiteral("aaa"));
|
||||
// leftmost: at position 0 matches empty? a* matches "aaa" (longest at 0).
|
||||
QCOMPARE(ms(QStringLiteral("ba*"), QStringLiteral("baaa")), QStringLiteral("baaa"));
|
||||
}
|
||||
|
||||
void TestSamRegex::plusQuest()
|
||||
{
|
||||
QCOMPARE(ms(QStringLiteral("a+"), QStringLiteral("baaa")), QStringLiteral("aaa"));
|
||||
QCOMPARE(ms(QStringLiteral("ab?c"), QStringLiteral("ac")), QStringLiteral("ac"));
|
||||
QCOMPARE(ms(QStringLiteral("ab?c"), QStringLiteral("abc")), QStringLiteral("abc"));
|
||||
}
|
||||
|
||||
void TestSamRegex::dotNotNewline()
|
||||
{
|
||||
// . does not cross newline (sam).
|
||||
QCOMPARE(m(QStringLiteral("a.b"), QStringLiteral("a\nb")), QStringLiteral("nil"));
|
||||
QCOMPARE(ms(QStringLiteral("a.c"), QStringLiteral("abc")), QStringLiteral("abc"));
|
||||
}
|
||||
|
||||
void TestSamRegex::anchorsPerLine()
|
||||
{
|
||||
// ^ matches at start of any line; $ at end of any line.
|
||||
SamRegex re;
|
||||
QVERIFY(re.compile(QStringLiteral("^b")));
|
||||
auto r = re.match(QStringLiteral("a\nb\nc"), 0, 5);
|
||||
QVERIFY(r.has_value());
|
||||
QCOMPARE(r->start(), 2); // 'b' on line 2
|
||||
QCOMPARE(r->end(), 3);
|
||||
|
||||
SamRegex re2;
|
||||
QVERIFY(re2.compile(QStringLiteral("a$")));
|
||||
auto r2 = re2.match(QStringLiteral("ba\nxa\n"), 0, 6);
|
||||
QVERIFY(r2.has_value());
|
||||
QCOMPARE(r2->start(), 1); // first 'a' that precedes a newline
|
||||
QCOMPARE(r2->end(), 2);
|
||||
}
|
||||
|
||||
void TestSamRegex::charClass()
|
||||
{
|
||||
QCOMPARE(ms(QStringLiteral("[abc]+"), QStringLiteral("xcababz")), QStringLiteral("cabab"));
|
||||
}
|
||||
|
||||
void TestSamRegex::negatedClass()
|
||||
{
|
||||
// [^z]+ crosses nothing special but stops at z; does not match newline.
|
||||
QCOMPARE(ms(QStringLiteral("[^z]+"), QStringLiteral("abz")), QStringLiteral("ab"));
|
||||
QCOMPARE(ms(QStringLiteral("[^z]+"), QStringLiteral("a\nb")), QStringLiteral("a"));
|
||||
}
|
||||
|
||||
void TestSamRegex::range()
|
||||
{
|
||||
QCOMPARE(ms(QStringLiteral("[0-9]+"), QStringLiteral("ab123cd")), QStringLiteral("123"));
|
||||
QCOMPARE(ms(QStringLiteral("[a-cx-z]+"), QStringLiteral("abcxyz!")), QStringLiteral("abcxyz"));
|
||||
}
|
||||
|
||||
void TestSamRegex::captures()
|
||||
{
|
||||
SamRegex re;
|
||||
QVERIFY(re.compile(QStringLiteral("([a-z]+)=([a-z]+)")));
|
||||
const QString text = QStringLiteral("key=value");
|
||||
auto r = re.match(text, 0, text.size());
|
||||
QVERIFY(r.has_value());
|
||||
QCOMPARE(r->captured(1, text), QStringLiteral("key"));
|
||||
QCOMPARE(r->captured(2, text), QStringLiteral("value"));
|
||||
}
|
||||
|
||||
void TestSamRegex::escapes()
|
||||
{
|
||||
// \* is a literal star; \. a literal dot; \n a newline.
|
||||
QCOMPARE(ms(QStringLiteral("a\\*b"), QStringLiteral("xa*bx")), QStringLiteral("a*b"));
|
||||
QCOMPARE(ms(QStringLiteral("a\\.b"), QStringLiteral("axb")), QStringLiteral("nil"));
|
||||
QCOMPARE(ms(QStringLiteral("a\\.b"), QStringLiteral("a.b")), QStringLiteral("a.b"));
|
||||
QCOMPARE(m(QStringLiteral("a\\nb"), QStringLiteral("a\nb")), QStringLiteral("0,3"));
|
||||
}
|
||||
|
||||
void TestSamRegex::fromOffset()
|
||||
{
|
||||
// Search starting past the first match finds the second.
|
||||
QCOMPARE(ms(QStringLiteral("a"), QStringLiteral("xaya"), 2), QStringLiteral("a"));
|
||||
SamRegex re;
|
||||
QVERIFY(re.compile(QStringLiteral("a")));
|
||||
auto r = re.match(QStringLiteral("xaya"), 2, 4);
|
||||
QVERIFY(r.has_value());
|
||||
QCOMPARE(r->start(), 3);
|
||||
}
|
||||
|
||||
void TestSamRegex::noMatch()
|
||||
{
|
||||
QCOMPARE(m(QStringLiteral("zzz"), QStringLiteral("abc")), QStringLiteral("nil"));
|
||||
}
|
||||
|
||||
void TestSamRegex::nestedGroupsLongest()
|
||||
{
|
||||
// (a|ab)(c|bcd) on "abcd": leftmost-longest should match the whole "abcd"
|
||||
// (a + bcd), which greedy-ordered backtracking would miss if it locked in
|
||||
// "a" then "bcd"? Actually both can reach abcd; the key is longest overall.
|
||||
QCOMPARE(ms(QStringLiteral("(a|ab)(c|bcd)"), QStringLiteral("abcd")),
|
||||
QStringLiteral("abcd"));
|
||||
}
|
||||
|
||||
void TestSamRegex::linearNoCatastrophicBacktrack()
|
||||
{
|
||||
// A pattern that makes PCRE backtrack exponentially: (a*)*b on a long run
|
||||
// of 'a' with no 'b'. The NFA handles it in linear time; just assert it
|
||||
// terminates quickly and reports no match.
|
||||
const QString text(10000, QLatin1Char('a'));
|
||||
SamRegex re;
|
||||
QVERIFY(re.compile(QStringLiteral("(a*)*b")));
|
||||
auto r = re.match(text, 0, text.size());
|
||||
QVERIFY(!r.has_value());
|
||||
}
|
||||
|
||||
QTEST_MAIN(TestSamRegex)
|
||||
#include "test_samregex.moc"
|
||||
Loading…
Reference in New Issue