| #
0fbee189
|
| 22-Jun-2026 |
Andreas Gohr <gohr@cosmocode.de> |
Return body text and metadata from a single extraction
Every extractor now returns an ExtractionResult (body text plus a normalised metadata map) from one parse of the file, instead of just a text s
Return body text and metadata from a single extraction
Every extractor now returns an ExtractionResult (body text plus a normalised metadata map) from one parse of the file, instead of just a text string. Metadata keys are a single canonical vocabulary shared across all formats - Title, Author, Subject, Keywords, Description, Created, Modified, Language, Producer, plus Copyright (image-only) - so consumers never have to special-case the source format.
- ExtractionResult: final readonly value object carrying $text, $metadata and a per-half ?Throwable $textError / $metadataError (with isComplete()). Text and metadata are extracted independently: an extractor throws only at its total-failure gate (container won't unzip, PDF won't parse, file unreadable) and otherwise records a failed half's error while still returning the salvaged half.
- OOXML/ODF metadata: new AbstractOoxmlExtractor and AbstractOdfExtractor intermediaries declare each family's source map once (docProps/core.xml + app.xml; meta.xml); the concrete extractors only switch their extends clause. AbstractZipXmlExtractor gains a generic mapMetadataFromXml() walker.
- PDF: reads the Info dictionary (Title/Author/Subject/Keywords/Created/ Modified/Producer-else-Creator) alongside the body text from the same parse. normalizePdfString() is a temporary shim for UTF-16BE literal strings that prinsfrank <= v3.1.0 returns mojibake-expanded.
- Images: now metadata-only (text is ''); field-map keys renamed to the canonical vocabulary (Caption->Description, Date->Created, Camera->Producer).
- helper: extract() and extractMetadata() added; extractText() keeps its string return and its throw-on-failure contract (rethrows textError), so the docsearch plugin keeps working unchanged.
- cli: prints body text followed by metadata by default; --text/-t and --meta/-m restrict to one. A salvaged-but-partial extraction warns on STDERR and exits 0; nothing usable exits non-zero.
- ExtractionException::wrap() centralises wrapping a caught error as an ExtractionException (passing one through if it already is).
Tests assert the real Tika samples' embedded metadata and the partial-failure paths. The pre-existing PDF body-text Form-XObject failure is unchanged.
show more ...
|
| #
ca76a2e6
|
| 17-Jun-2026 |
Andreas Gohr <gohr@cosmocode.de> |
Initial implementation of the totext plugin
Extract plain text from various file formats using PHP only (no shell-outs, no external binaries). Provides a CLI component (cli_plugin_totext) and a help
Initial implementation of the totext plugin
Extract plain text from various file formats using PHP only (no shell-outs, no external binaries). Provides a CLI component (cli_plugin_totext) and a helper component (helper_plugin_totext).
Supported formats: - OOXML: docx, xlsx, pptx - OpenDocument: odt, ods, odp - PDF (via bundled smalot/pdfparser) - Plain-text family: txt, csv, md, markdown, log, text - Image metadata (EXIF/IPTC): jpg, jpeg, tif, tiff
Includes a fixture-based test suite covering every extractor, the factory routing, and the helper end to end.
show more ...
|