| #
b22130d2
|
| 08-Jul-2026 |
Andreas Gohr <gohr@cosmocode.de> |
fix(parser): stop group-quantifier body patterns from exhausting memory
Several patterns repeat a group — (?:…)* / (?:…)+ — over a body that can span a large region. On the non-JIT PCRE engine each
fix(parser): stop group-quantifier body patterns from exhausting memory
Several patterns repeat a group — (?:…)* / (?:…)+ — over a body that can span a large region. On the non-JIT PCRE engine each iteration of a group repetition keeps its own backtracking frame, so a big body retains one frame per byte: a linear, unbounded memory spike that can fatal a request (and, with JIT on, trips pcre.jit_stacklimit so the lexer dumps the rest of the document as literal text). A repetition over a single item (.* , [^x]+) is optimized by PCRE and unaffected.
Make each body possessive: the repeated group cannot match the delimiter that follows it, so it always stops at the first closer and never needs to backtrack.
- media {{…}}, windowssharelink, gfm_quote, gfm_link: possessive body. - email validation (used by emaillink and mail_isvalid): possessive local-part and domain groups; a long dotted `a.a.a…` local part or domain was a ReDoS vector. Match results are unchanged.
show more ...
|
| #
e248eb6f
|
| 05-Jul-2026 |
Andreas Gohr <andi@splitbrain.org> |
fix(parser): don't drop document content after a blank indented line
Preformatted's exit pattern (?=\n[^ \t\n]) is a lookahead-only exit: it consumes no byte so the boundary \n stays in the stream f
fix(parser): don't drop document content after a blank indented line
Preformatted's exit pattern (?=\n[^ \t\n]) is a lookahead-only exit: it consumes no byte so the boundary \n stays in the stream for a following block mode (an <hr> or header after an indented code block) to anchor on.
When that exit fired with nothing else consumed - a "blank" line that is empty after its indent, followed by a column-0 line - the match was zero-width at the current offset and the lexer's no-advance guard aborted the whole parse, silently discarding the rest of the document. Reachable on ordinary trailing whitespace ("abc\n \nmore") and inside footnotes.
Exempt zero-width MODE_EXIT matches from the guard: popping the mode stack is real progress and leaves the boundary byte for the parent mode to consume on the next iteration. The stack strictly shrinks on each such exit, so this cannot loop; the infinite-loop protection is unchanged for every other zero-width match. Preformatted's pattern is left as-is, so <hr>/header after an indented code block still works, including after a blank indented line.
Add regressions for the blank-line, blank-line-after-code, and boundary-preservation (hr after a blank indented line) cases.
show more ...
|
| #
47a02a10
|
| 04-Jun-2026 |
Andreas Gohr <gohr@cosmocode.de> |
Parsing: make parse syntax a per-parse value, drop ModeInterface
The active parse's syntax flavour is a per-parse question, not process- global state: within a single request a plugin can render bun
Parsing: make parse syntax a per-parse value, drop ModeInterface
The active parse's syntax flavour is a per-parse question, not process- global state: within a single request a plugin can render bundled DokuWiki-syntax text inside an otherwise-Markdown page. Yet ModeRegistry was a singleton that read $conf['syntax'] and the $PARSER_MODES global, and every mode reached it through ModeRegistry::getInstance() — so the flavour lived in shared mutable state that two parses in one request would fight over.
Make the registry a short-lived value instead:
- ModeRegistry is constructed once per parse with an explicit $syntax and injected into Parser, Handler and every mode. getSyntax() / isDwPreferred() / isMdPreferred() consult $this->syntax; the DOKU_UNITTEST-gated mode-list cache hack is gone (each registry is fresh, nothing to invalidate). - p_get_instructions() is now the single place in the pipeline where $conf['syntax'] is read; from there the flavour travels as a parameter. No code under inc/Parsing/ reads $conf['syntax'] directly anymore — the five syntax-reading modes (Preformatted, GfmHr, GfmEscape, Externallink, GfmQuote) route through $this->registry.
Keep the two concepts apart, as documented in the ModeRegistry and AbstractMode docblocks: the user's configured *preference* stays in $conf['syntax'] for UI code (toolbar, settings), while the active parse's syntax is a parameter carried by the registry.
$PARSER_MODES is demoted to a deprecated, read-only mirror, published during loadPluginModes() — third-party syntax plugins (columnlist, alphalist2, phpwikify, skipentity) and the bundled info plugin read the global directly, often from their constructors, so the taxonomy must stay visible there. No core code reads the mirror.
Fold ModeInterface into AbstractMode while here: getSort()/handle() are abstract, the connect callbacks carry defaults, and the public $Lexer "FIXME should be done by setter" becomes setLexer()/getLexer() injected by Parser::addMode() alongside the registry. Nested-content resolution moves to the allowedCategories()/filterAllowedModes() hooks, resolved once when the registry is attached.
Tests build their own parser/registry through ParserTestBase::setSyntax() instead of mutating $conf and calling the removed ModeRegistry::reset().
show more ...
|
| #
dccbd514
|
| 05-May-2026 |
Andreas Gohr <andi@splitbrain.org> |
GfmQuote: accept ^> line starts so quotes can follow tables and lists
GfmTable, DW Table, and DW Listblock all consume the boundary \n on their way out. A pure-lookahead exit pattern at that boundar
GfmQuote: accept ^> line starts so quotes can follow tables and lists
GfmTable, DW Table, and DW Listblock all consume the boundary \n on their way out. A pure-lookahead exit pattern at that boundary would trip the lexer's no-advance safety check, because tables and lists exit right after consuming a marker token and have no leading unmatched content for the lookahead to attach to (unlike Preformatted, whose body leaves code lines as UNMATCHED right before the boundary).
Fix this on the consumer side: change the first-line anchor from \n> to (?:^|\n)>. With the lexer's m flag, ^ matches at offset 0 and at any position immediately following a \n in the subject, including the position right after a \n that a preceding mode just consumed. Subsequent quote lines keep the \n> anchor.
Adds three handoff tests in GfmQuoteTest covering GfmTable, DW Table, and DW Listblock. Resolves GFM spec example 201.
show more ...
|
| #
13a62f81
|
| 04-May-2026 |
Andreas Gohr <andi@splitbrain.org> |
rename syntax flavors 'dokuwiki' / 'markdown' to 'dw' / 'md'
Symmetry with the existing 'dw+md' / 'md+dw' setting values.
|
| #
309a0852
|
| 30-Apr-2026 |
Andreas Gohr <andi@splitbrain.org> |
replace DW Quote with unified GfmQuote
GfmQuote covers blockquote parsing for both DokuWiki and GFM dialects in a single mode. Same quote_open/quote_close handler instructions; a DW-preferred post-p
replace DW Quote with unified GfmQuote
GfmQuote covers blockquote parsing for both DokuWiki and GFM dialects in a single mode. Same quote_open/quote_close handler instructions; a DW-preferred post-pass flattens sub-parsed paragraph wrapping into linebreak calls so existing pages keep their <br/>-between-lines rendering. MD-preferred keeps the <p>-wrapped spec shape.
Block content (lists, fenced code, tables) inside `>` quotes now renders, since the body is sub-parsed. Headers stay excluded (BASEONLY) — TOC and section-edit anchors don't compose with <blockquote>, same rationale as GfmListblock.
Convert ModeRegistry's sub-parser cache into an acquire/release pool to support same-key re-entrancy: a list inside a quote re-enters gfm_quote during the list-item sub-parse, and the inner call needs its own parser instance even though the exclusion key matches. GfmListblock is updated to use the new acquire/release primitives.
show more ...
|