| 846f53c1 | 08-Jul-2026 |
Andreas Gohr <gohr@cosmocode.de> |
fix(search): lock the metadata registry read-modify-write
updateMetadataRegistry() read metadata.idx, merged in new keys and wrote the file back without holding a lock across the sequence. Two index
fix(search): lock the metadata registry read-modify-write
updateMetadataRegistry() read metadata.idx, merged in new keys and wrote the file back without holding a lock across the sequence. Two indexer processes each registering a different new key could both read the same registry and the later writer would clobber the other's key, dropping it. A dropped key is not cleared from its collection on deletePage(), leaving orphaned index entries.
Guard the whole read-merge-write with a dedicated metadata lock so concurrent registrations can no longer lose a key.
show more ...
|
| d9043e78 | 08-Jul-2026 |
Andreas Gohr <gohr@cosmocode.de> |
fix(search): make DirectCollection::getEntitiesWithData() linear
The method resolved each entity name with a separate retrieveRow() call, each of which rescans the entity index from the start, makin
fix(search): make DirectCollection::getEntitiesWithData() linear
The method resolved each entity name with a separate retrieveRow() call, each of which rescans the entity index from the start, making the whole operation quadratic in the number of entities. This runs on MetadataSearch::getPages('title') and hurts large wikis. Collect the entity IDs in one pass and resolve their names with a single batched retrieveRows() read instead.
show more ...
|
| 45fa8bb2 | 08-Jul-2026 |
Andreas Gohr <gohr@cosmocode.de> |
fix(search): stop FileIndex::retrieveRows() scanning past the last row
The early-exit check compared the array_shift() result against false, but array_shift() returns null on an empty array, so the
fix(search): stop FileIndex::retrieveRows() scanning past the last row
The early-exit check compared the array_shift() result against false, but array_shift() returns null on an empty array, so the break never fired and every call read the index file to EOF after collecting the last requested row. Compare against null and skip the file entirely when nothing is requested.
show more ...
|
| 61dea710 | 08-Jul-2026 |
Andreas Gohr <gohr@cosmocode.de> |
fix(search): wait for a contended index lock and apply dperm
The Lock rewrite (c66b5ec65) made Lock::acquire() fail immediately when the lock directory already existed, where the old indexer lock re
fix(search): wait for a contended index lock and apply dperm
The Lock rewrite (c66b5ec65) made Lock::acquire() fail immediately when the lock directory already existed, where the old indexer lock retried until the holder released it. Concurrent saves or indexer runs then aborted indexing instead of serializing (recoverable only via the .indexed tag on the next edit), and the lock directory was created without the configured directory permissions, unlike every other directory DokuWiki creates.
Mirror the io_lock() convention: wait for a contended lock, bounded by a tunable wait timeout, before throwing; keep clearing locks older than five minutes as stale; and chmod the new lock directory to $conf['dperm']. The give-up path still throws IndexLockException, so the self-healing retry via the .indexed tag is unchanged.
show more ...
|
| 7595d8da | 06-Jul-2026 |
splitbrain <86426+splitbrain@users.noreply.github.com> |
Rector and PHPCS fixes |
| ebaa8481 | 05-Jul-2026 |
Andreas Gohr <andi@splitbrain.org> |
fix(search): restore histogram() as a Collection method with a BC shim
The indexer rework dropped Doku_Indexer::histogram(), which plugins reach through idx_get_indexer()->histogram() to build tag c
fix(search): restore histogram() as a Collection method with a BC shim
The indexer rework dropped Doku_Indexer::histogram(), which plugins reach through idx_get_indexer()->histogram() to build tag clouds (e.g. tagfilter via the 'subject' metadata index). Calling it fatalled through LegacyIndexer::__call as an undefined method.
Add histogram() to AbstractCollection: it sums each token's frequency across all entities and returns them ordered by frequency, filtered by min/max/minlen, handling both length-split (fulltext) and single-group (metadata) collections. DirectCollection overrides it for the 1:1 title case (count shared values). LegacyIndexer::histogram() is a deprecated shim dispatching by $key to the matching collection.
show more ...
|
| 9e4de2f7 | 05-Jul-2026 |
Andreas Gohr <andi@splitbrain.org> |
fix(search): don't let collections release a caller-owned index lock
Indexer::addPage() locks the page index once and hands the same instance to several collections, each of which lock() and unlock(
fix(search): don't let collections release a caller-owned index lock
Indexer::addPage() locks the page index once and hands the same instance to several collections, each of which lock() and unlock() it around their work.
AbstractCollection::lock() locked and tracked every index for later release. When a shared AbstractIndex was already locked by the caller, locking it again was a no-op but the collection still tracked it, so the collection's unlock() released the caller's lock and flipped the shared index read-only. Each later collection then re-acquired and re-released it, leaving windows in which another process could grab the page lock mid-operation; contention could abort addPage() with the title index updated but the fulltext index stale.
Treat a shared index that is already writable as owned by its creator: leave its lock untouched and only manage the locks the collection acquires itself.
Add a regression test asserting a caller-held index lock survives repeated collection lock()/unlock() cycles.
show more ...
|
| 4e61aa71 | 05-Jul-2026 |
Andreas Gohr <andi@splitbrain.org> |
fix(search): create index lock directories inside data/locks
Index locks were built by concatenating the configured lock directory and the index name without a path separator.
That produced paths l
fix(search): create index lock directories inside data/locks
Index locks were built by concatenating the configured lock directory and the index name without a path separator.
That produced paths like data/lockspage.index instead of data/locks/page.index, so index lock directories were created in the data root rather than inside the configured lock directory.
Fix this by joining the lock directory and index name with an explicit slash, and update the lock test to assert the correct location and guard against the old broken path.
show more ...
|
| bc12b8fe | 05-Jul-2026 |
Andreas Gohr <andi@splitbrain.org> |
fix(search): read split index suffixes correctly
AbstractIndex::max() used /(\d)+\.idx$/ to determine the highest split index suffix. For filenames like w10.idx or w12.idx that regex captured only t
fix(search): read split index suffixes correctly
AbstractIndex::max() used /(\d)+\.idx$/ to determine the highest split index suffix. For filenames like w10.idx or w12.idx that regex captured only the last digit, so the maximum token-length group was truncated to 9. Wildcard searches therefore skipped all fulltext index shards for words with 10 or more characters.
Tighten the match to the basename and require the current index name followed immediately by digits: ^<idx>(\d+)\.idx$. This captures the full numeric suffix and also avoids counting unrelated index families that share the same prefix, such as treating wiki2.idx as a numbered shard of the w index.
Add regressions for both behaviors: multi-digit suffixes are parsed correctly, and same-prefix indexes are ignored when determining the maximum shard number.
show more ...
|
| b1f7ba64 | 05-Jul-2026 |
Andreas Gohr <andi@splitbrain.org> |
fix(search): fatals on invalid short wildcard terms
Searches such as "wiki a*" could crash in QueryEvaluator::opAnd() with a TypeError. The query parser still emitted the wildcard term, but the full
fix(search): fatals on invalid short wildcard terms
Searches such as "wiki a*" could crash in QueryEvaluator::opAnd() with a TypeError. The query parser still emitted the wildcard term, but the fulltext search later dropped it because its non-wildcard base was below the minimum search term length.
That left the evaluator with an AND operator whose right-hand operand was missing, causing a stack underflow during RPN evaluation.
Fix this in two places: - filter invalid short wildcard terms already in Tokenizer::getWords() when wildcard parsing is enabled, so they never enter the parsed query - harden QueryEvaluator against missing operands for AND/OR/NOT to avoid fatals if tokens disappear during later processing
Add regression tests covering the parser output for "wiki a*" and the evaluator behavior when an operand is missing.
show more ...
|
| e282ae55 | 05-Jul-2026 |
Andreas Gohr <andi@splitbrain.org> |
fix(search): search index corruption in updateTuple
Require an end-of-tuple boundary when removing an existing tuple in TupleOps::updateTuple(), so updating row ID 17 no longer corrupts tuples for I
fix(search): search index corruption in updateTuple
Require an end-of-tuple boundary when removing an existing tuple in TupleOps::updateTuple(), so updating row ID 17 no longer corrupts tuples for IDs like 170 or 171.
Add regression tests covering both replacement and deletion when one numeric row ID is a decimal prefix of another.
show more ...
|
| 2cda0166 | 17-Jun-2026 |
Andreas Gohr <gohr@cosmocode.de> |
Indexer: signal nothing-to-do via boolean return instead of void
The TaskRunner runs indexing, sitemap, digest and changelog-trim tasks in sequence and relies on each task returning false when it di
Indexer: signal nothing-to-do via boolean return instead of void
The TaskRunner runs indexing, sitemap, digest and changelog-trim tasks in sequence and relies on each task returning false when it did no work so the next one is tried. The indexer rewrite changed addPage(), deletePage() and renamePage() to return void and only abort via exceptions, breaking that contract: indexing always looked like work was done and the following tasks never ran.
Restore the boolean return on these three methods (true when work was done, false when there was nothing to do) while still using exceptions to signal errors, and propagate it through TaskRunner::runIndexer(). runIndexer() also no longer forces reindexing on every call.
The legacy compatibility layer is adjusted to match: LegacyIndexer and idx_addPage() forward the boolean, mapping SearchExceptions back to the historic error-message/false returns. LegacyIndexer::renamePage() restores the 'page is not in index' message that the move plugin expects.
Closes #4661
show more ...
|
| 79dae64d | 17-Jun-2026 |
Andreas Gohr <gohr@cosmocode.de> |
Indexer: treat same-second save and index as up to date
needsIndexing() compared the .indexed tag mtime against the page mtime with <=, so a page that was saved and indexed within the same second wa
Indexer: treat same-second save and index as up to date
needsIndexing() compared the .indexed tag mtime against the page mtime with <=, so a page that was saved and indexed within the same second was always reported as still needing indexing. Require the page to be strictly newer than the index tag instead, so an equal mtime correctly counts as up to date.
show more ...
|
| 2ff7e61c | 10-Jun-2026 |
Andreas Gohr <gohr@cosmocode.de> |
fix(indexer): explicitly handle renames
In an attempt to simplify the index handling, the newly refactored indexer implemented a rename as delete+add sequence.
This had unintended consequences for
fix(indexer): explicitly handle renames
In an attempt to simplify the index handling, the newly refactored indexer implemented a rename as delete+add sequence.
This had unintended consequences for the move plugin which may move several pages at once, requiring a working index even while some pages have already been moved while others still remain at their old location.
Related to #4646
show more ...
|
| 6e39b4e3 | 28-May-2026 |
Andreas Gohr <andi@splitbrain.org> |
refactor(search): extract LegacyIndexer wrapper for BC contract
Move the deprecated helpers (lookupKey, addMetaKeys, renameMetaValue, getPID, lookup) off Indexer and into a new LegacyIndexer wrapper
refactor(search): extract LegacyIndexer wrapper for BC contract
Move the deprecated helpers (lookupKey, addMetaKeys, renameMetaValue, getPID, lookup) off Indexer and into a new LegacyIndexer wrapper. The wrapper also restores the Doku_Indexer return contract (true|string) around addPage/deletePage/renamePage/clear so plugins using the legacy API keep working without try/catch.
idx_get_indexer() now returns the LegacyIndexer; getPages stays on Indexer because plugins call it directly on Indexer instances.
fixes #4645
show more ...
|
| 53307a6b | 09-May-2026 |
Andreas Gohr <andi@splitbrain.org> |
Delete inc/Search/concept.txt
The contents have been added to the wiki |
| 8788dbbd | 06-May-2026 |
splitbrain <86426+splitbrain@users.noreply.github.com> |
Rector and PHPCS fixes |
| 4f29a5b9 | 06-May-2026 |
Andreas Gohr <andi@splitbrain.org> |
SearchIndex: fix comment position
single line comment moved to the wrong line on reformatting |
| 06053dca | 10-Apr-2026 |
Andreas Gohr <andi@splitbrain.org> |
SearchIndex: remove write side effect from retrieveRow()
retrieveRow() padded the index file when the requested RID was beyond the current length. This was an optimization for subsequent changeRow()
SearchIndex: remove write side effect from retrieveRow()
retrieveRow() padded the index file when the requested RID was beyond the current length. This was an optimization for subsequent changeRow() calls, but changeRow() already handles padding on its own. The side effect was also inconsistent with retrieveRows() which is a pure read.
Added a cross-index integration test verifying RID consistency across entity, token, frequency and reverse indexes when multiple entities share tokens.
show more ...
|
| 5d034a75 | 08-Apr-2026 |
Andreas Gohr <andi@splitbrain.org> |
SearchIndex: increase index version |
| 9369b4a9 | 08-Apr-2026 |
Andreas Gohr <andi@splitbrain.org> |
SearchIndex: rector, phpcs, type hint fixes |
| db8be586 | 08-Apr-2026 |
Andreas Gohr <andi@splitbrain.org> |
SearchIndex: review fixes — auto-save MemoryIndex, cast TupleOps counts, style cleanups
- MemoryIndex: auto-save dirty data on unlock/destruction to prevent silent index corruption when indexes ar
SearchIndex: review fixes — auto-save MemoryIndex, cast TupleOps counts, style cleanups
- MemoryIndex: auto-save dirty data on unlock/destruction to prevent silent index corruption when indexes are used in tandem - TupleOps::parseTuples(): cast exploded count strings to int - FileIndex::retrieveRow(): document the write-on-read padding behavior - Fix whitespace issues in ApiCore, common.php, Sitemap/Mapper - Update concept.txt to reflect MemoryIndex auto-save behavior
show more ...
|
| 2a22d4b9 | 08-Apr-2026 |
Andreas Gohr <andi@splitbrain.org> |
SearchIndex: document Tokenizer::isValidSearchTerm() in concept.txt |
| 1148921d | 08-Apr-2026 |
Andreas Gohr <andi@splitbrain.org> |
SearchIndex: unify CollectionSearch API and optimize search pipeline
- Remove separate lookup() API from CollectionSearch. All searches now use addTerm()/execute() with a single unified pipeline.
SearchIndex: unify CollectionSearch API and optimize search pipeline
- Remove separate lookup() API from CollectionSearch. All searches now use addTerm()/execute() with a single unified pipeline. - Add matches() predicate to Term using efficient string functions (===, str_starts_with, str_ends_with, str_contains) instead of regex. - Add caseInsensitive() support on CollectionSearch and Term for metadata/title searches where indexed values preserve case. - Remove callback support from MetadataSearch::lookupKey() — the only real usage (case-insensitive substring) is replaced by caseInsensitive() + wildcards. - Remove min-length validation from Term. Add Tokenizer::isValidSearchTerm() for callers that need it (FulltextSearch, Indexer::lookup). - Optimize execute() from 4 group passes to 2: scan tokens + resolve frequencies in one pass per group, batch entity name resolution, then populate Terms. - Store full match detail in Term: entity → token → frequency. New accessors getMatches(), getEntityTokens(), getEntityFrequencies() derive different views from this single data structure. - Term no longer used as scratch pad by CollectionSearch. Index-internal data (token IDs, entity IDs) stays local to execute(). Terms receive only final resolved results. - Use title from search results in MetadataSearch::pageLookupCallBack() instead of re-fetching via p_get_first_heading(). - Update concept.txt documentation.
show more ...
|
| b9d7a615 | 07-Apr-2026 |
Andreas Gohr <andi@splitbrain.org> |
SearchIndex: updated documentation
to be moved into the wiki later |