History log of /plugin/totext/plugin.info.txt (Results 1 – 4 of 4)
Revision Date Author Comments
# 00f006a0 15-Jul-2026 Andreas Gohr <andi@splitbrain.org>

Version upped


# 9cdb356d 14-Jul-2026 Andreas Gohr <andi@splitbrain.org>

Version upped


# 0fbee189 22-Jun-2026 Andreas Gohr <gohr@cosmocode.de>

Return body text and metadata from a single extraction

Every extractor now returns an ExtractionResult (body text plus a
normalised metadata map) from one parse of the file, instead of just a
text s

Return body text and metadata from a single extraction

Every extractor now returns an ExtractionResult (body text plus a
normalised metadata map) from one parse of the file, instead of just a
text string. Metadata keys are a single canonical vocabulary shared
across all formats - Title, Author, Subject, Keywords, Description,
Created, Modified, Language, Producer, plus Copyright (image-only) - so
consumers never have to special-case the source format.

- ExtractionResult: final readonly value object carrying $text,
$metadata and a per-half ?Throwable $textError / $metadataError (with
isComplete()). Text and metadata are extracted independently: an
extractor throws only at its total-failure gate (container won't unzip,
PDF won't parse, file unreadable) and otherwise records a failed half's
error while still returning the salvaged half.

- OOXML/ODF metadata: new AbstractOoxmlExtractor and AbstractOdfExtractor
intermediaries declare each family's source map once (docProps/core.xml
+ app.xml; meta.xml); the concrete extractors only switch their extends
clause. AbstractZipXmlExtractor gains a generic mapMetadataFromXml()
walker.

- PDF: reads the Info dictionary (Title/Author/Subject/Keywords/Created/
Modified/Producer-else-Creator) alongside the body text from the same
parse. normalizePdfString() is a temporary shim for UTF-16BE literal
strings that prinsfrank <= v3.1.0 returns mojibake-expanded.

- Images: now metadata-only (text is ''); field-map keys renamed to the
canonical vocabulary (Caption->Description, Date->Created,
Camera->Producer).

- helper: extract() and extractMetadata() added; extractText() keeps its
string return and its throw-on-failure contract (rethrows textError),
so the docsearch plugin keeps working unchanged.

- cli: prints body text followed by metadata by default; --text/-t and
--meta/-m restrict to one. A salvaged-but-partial extraction warns on
STDERR and exits 0; nothing usable exits non-zero.

- ExtractionException::wrap() centralises wrapping a caught error as an
ExtractionException (passing one through if it already is).

Tests assert the real Tika samples' embedded metadata and the
partial-failure paths. The pre-existing PDF body-text Form-XObject
failure is unchanged.

show more ...


# ca76a2e6 17-Jun-2026 Andreas Gohr <gohr@cosmocode.de>

Initial implementation of the totext plugin

Extract plain text from various file formats using PHP only (no shell-outs,
no external binaries). Provides a CLI component (cli_plugin_totext) and a
help

Initial implementation of the totext plugin

Extract plain text from various file formats using PHP only (no shell-outs,
no external binaries). Provides a CLI component (cli_plugin_totext) and a
helper component (helper_plugin_totext).

Supported formats:
- OOXML: docx, xlsx, pptx
- OpenDocument: odt, ods, odp
- PDF (via bundled smalot/pdfparser)
- Plain-text family: txt, csv, md, markdown, log, text
- Image metadata (EXIF/IPTC): jpg, jpeg, tif, tiff

Includes a fixture-based test suite covering every extractor, the factory
routing, and the helper end to end.

show more ...