| #
1c00c021
|
| 09-Jul-2026 |
Andreas Gohr <gohr@cosmocode.de> |
fix(parser): validate inline formatting closers with a single memoized scan
The inline formatting modes only open a span when a valid closer exists ahead. That check was a lookahead built on CONTENT
fix(parser): validate inline formatting closers with a single memoized scan
The inline formatting modes only open a span when a valid closer exists ahead. That check was a lookahead built on CONTENT_UNTIL_PARA, tested character by character up to the next paragraph break and re-evaluated from scratch for every opener candidate — openers times paragraph length. With pcre.jit=0 a crafted 32KB page took 16s and an ordinary 34KB page with long paragraphs 37s; with the JIT on (the PHP default) the per-character lookahead exhausted the JIT stack, the match silently failed, and the formatting — or everything after it — rendered as plain text.
The check also decided the wrong thing. It scanned raw text, so a closer lookalike inside content the lexer consumes atomically — a nowiki or %% span, a backtick code span, a link, a URL — counted as a real closer even though the mode's exit pattern can never fire there. And it ignored the enclosing span: an inner delimiter whose only closer lay past the closer of the mode it sits in was entered anyway, so a stray delimiter paired with one in a following sibling span and dragged the boundary along — the `*` in ''glob/*.conf'' joined the `*` of the next ''...'' span and corrupted the paragraph; the same held for //, ** and __ inside monospace, and an emphasis opened inside ((...)) ran past the footnote's )) and the enclosing bold's **.
Each formatting mode now declares its closer through Lexer::addCloserPattern(), mirroring addExitPattern(), and the lexer answers "does a valid closer exist ahead" with one anchored possessive scan per range instead of a lookahead per opener:
- The scan runs left to right from the opener, hopping over opaque spans derived from already-registered patterns — a plain or special match is consumed in one step, an entry into a verbatim mode (nowiki, the backtick code spans) extends to that mode's first exit — so a closer lookalike inside consumed content is never mistaken for a closer. Each hop finds the earliest of boundary, closer, or opaque span in a single leftmost search, keeping the check linear. - An opener is rejected when the nearest enclosing mode that has a closer of its own would close before the opener's own closer, so a delimiter that can never close within its span stays literal. That ancestor is found by walking the mode stack past modes that declare no closer (plugins, footnotes); the nearest guarded ancestor suffices, as it was itself validated against its own when it opened. - Both verdicts are memoized and reset per parse() run: a proven closer validates every earlier candidate, and a proven closer-free range rejects every later candidate before the next boundary. With the lexer consuming each opened span, the whole parse is linear in document size.
Closer patterns match the closing delimiter itself with flanking context in lookarounds — the convention exit patterns already follow — so closer positions compare exactly across modes and a closer directly after an inner opener is seen. AbstractFormatting derives the closer from the exit pattern and registers it with the paragraph break as the boundary, preserving the rule that formatting never spans paragraphs; a mode with other needs can pass a different boundary or none.
Footnote declares its )) as a closer rather than guarding its (( entry with a (?=.*)) lookahead, so the footnote becomes a boundary the scan sees and formatting inside it no longer pairs across the )); its closer takes no paragraph boundary, as footnotes are block-level. GfmEmphasis gains a closer pattern so single * emphasis is validated the same way, while its entry lookahead still enforces CommonMark nearest-delimiter pairing. GfmEmphasis and GfmStrong span bodies cannot contain their delimiter, so their in-pattern lookaheads stay linear on their own; the GFM backtick span bodies get deterministic alternatives with possessive quantifiers, removing their per-character backtracking.
CONTENT_UNTIL_PARA is removed: any entry pattern built with it recreates the quadratic scan. ParallelRegex gains escapePattern() so embedded closer fragments follow the lexer's bare-parenthesis convention, reports PREG_JIT_STACKLIMIT_ERROR so a future JIT exhaustion surfaces instead of silently truncating, and no longer rewrites its registered patterns in place while compiling the compound regex.
The adversarial 32KB page drops to 0.1s, the 37-second benign page to 0.1s, and a 128KB variant stays under 0.6s.
show more ...
|
| #
864d6c6d
|
| 21-Apr-2026 |
Andreas Gohr <gohr@cosmocode.de> |
fix Lexer so exit-pattern lookbehinds see chars consumed by prior tokens
Lexer::reduce used to hand PCRE a shrinking tail of the subject — each matched token was chopped off the front of $raw and th
fix Lexer so exit-pattern lookbehinds see chars consumed by prior tokens
Lexer::reduce used to hand PCRE a shrinking tail of the subject — each matched token was chopped off the front of $raw and the next preg_match ran on what remained. Once a token was consumed, the bytes before the cursor were gone, and any lookbehind assertion in a subsequent pattern silently failed.
The bug was latent for DokuWiki's entire history because literal exit patterns like `\*\*`, `</file>`, or `%%` don't care what's behind them. It surfaced with c3755410a ("require non-whitespace adjacency for inline formatting delimiters"), which added `(?<=[^\s])` to Strong, Emphasis, Underline, Monospace, Subscript, Superscript and Deleted at once. After that commit, `**[[link]]**` stopped closing — the `]` that would satisfy the lookbehind had just been consumed by the link match, so Strong stayed open until end-of-section and swallowed everything after it (list items, headings, the lot).
Fix:
* Lexer::parse and Lexer::reduce track a byte offset into $raw instead of mutating $raw. $initialLength and the shrinking-length arithmetic for absolute match positions are replaced by straight offset increments; the no-progress guard and the trailing-unmatched dispatch both shift to the same cursor.
* ParallelRegex::split takes an optional $offset and passes it to preg_match together with PREG_OFFSET_CAPTURE. PCRE scans from the offset forward but still sees the whole subject, so lookbehinds work across already-consumed tokens. The secondary preg_split call used to carve out pre/post is no longer needed — PREG_OFFSET_CAPTURE gives the match start for free, saving one regex operation per reduce() step.
Regression tests at all three layers:
* ParallelRegexTest — offset plumbing and pre/match accounting. * LexerTest::testIndexLookbehindAcrossConsumedToken — exit-pattern lookbehind targeting the `/>` of a self-closing `<a/>` that was consumed as a SPECIAL token on the previous step. Fails under the old Lexer. * FormattingTest — `**[[link]]**` and `**foo//bar//**` round-trip with correct open/close instructions through the full pipeline.
Also updates ListsTest::testUnorderedListStrong, whose expectations documented the pre-fix buggy behaviour ("formatting able to spread across list items"). With the fix, bold correctly stays within a single list item; the expected call sequence and the comment are updated to match.
show more ...
|
| #
c3755410
|
| 20-Apr-2026 |
Andreas Gohr <gohr@cosmocode.de> |
require non-whitespace adjacency for inline formatting delimiters
An opening delimiter must now be followed by a non-whitespace character, and a closing delimiter must be preceded by one. Empty deli
require non-whitespace adjacency for inline formatting delimiters
An opening delimiter must now be followed by a non-whitespace character, and a closing delimiter must be preceded by one. Empty delimiter pairs (****, ____, '''', <sub></sub>, <sup></sup>, <del></del>) no longer match and stay literal.
Rationale: this matches Markdown's flanking-delimiter rules and eliminates accidental bolding of sequences like `** note**` at the start of a sentence. Well-formed uses (**bold**, //italic//, __underline__) are unchanged.
Affected modes: Strong, Emphasis, Underline, Monospace, Subscript, Superscript, Deleted.
BREAKING: content that was already malformed but previously rendered as formatted (e.g. `**foo bar **`) now stays literal.
show more ...
|
| #
10fb3d65
|
| 20-Apr-2026 |
Andreas Gohr <gohr@cosmocode.de> |
prevent inline formatting from matching across paragraph boundaries
The Lexer compiles all patterns with the `s` (DOTALL) flag via ParallelRegex::getPerlMatchingFlags(), which makes `.` match newlin
prevent inline formatting from matching across paragraph boundaries
The Lexer compiles all patterns with the `s` (DOTALL) flag via ParallelRegex::getPerlMatchingFlags(), which makes `.` match newlines. Inline formatting modes use lookaheads like `\*\*(?=.*\*\*)` to verify a closing delimiter exists, so with DOTALL a lone `**` happily matched its "closer" many paragraphs later, swallowing blank lines into a single <strong> run.
Add CONTENT_UNTIL_PARA on AbstractMode — a regex snippet matching any character unless it would start a paragraph break (blank line, possibly with horizontal whitespace). Update all inline formatting entry patterns (Strong, Emphasis, Underline, Monospace, Subscript, Superscript, Deleted) to use it in their closing-delimiter lookaheads.
Emphasis also gets a real closing-`//` check; its previous lookahead just verified "content exists with a non-colon char" without requiring the closing delimiter at all.
Single newlines inside a delimiter pair still match (multi-line formatting); only blank lines end it.
BREAKING: This means you no longer can mark multiple paragraphs as bold or strike them out. On the other hand it prevents accidentally breaking the page layout by missing a closing delimiter (as reported many many times over the years) eg. #1025 #3588 #1056
show more ...
|
| #
504c13e8
|
| 18-Apr-2026 |
Andreas Gohr <andi@splitbrain.org> |
migrate parser tests to modern namespaced classes
Move old-style parser tests from _test/tests/inc/parser/ to namespaced test classes under _test/tests/Parsing/ParserMode/ and _test/tests/Parsing/Le
migrate parser tests to modern namespaced classes
Move old-style parser tests from _test/tests/inc/parser/ to namespaced test classes under _test/tests/Parsing/ParserMode/ and _test/tests/Parsing/Lexer/.
- Add ParserTestBase with assertCalls() helper for comparing handler call sequences with byte index stripping - Split lexer.test.php into ParallelRegexTest, StateStackTest, and LexerTest with a RecordingHandler that extends Handler - Merge handler_parse_highlight_options tests into CodeTest - Add new tests for previously untested modes: Nocache, Notoc, Rss, and all individual formatting modes (Strong, Emphasis, etc.) - Modernize test code: [] syntax, lowercase null, correct assertEquals argument order, replace deprecated withConsecutive and string callables - Renderer tests remain in old location (renderers not yet migrated)
show more ...
|