Return body text and metadata from a single extractionEvery extractor now returns an ExtractionResult (body text plus anormalised metadata map) from one parse of the file, instead of just atext s
Return body text and metadata from a single extractionEvery extractor now returns an ExtractionResult (body text plus anormalised metadata map) from one parse of the file, instead of just atext string. Metadata keys are a single canonical vocabulary sharedacross all formats - Title, Author, Subject, Keywords, Description,Created, Modified, Language, Producer, plus Copyright (image-only) - soconsumers never have to special-case the source format.- ExtractionResult: final readonly value object carrying $text, $metadata and a per-half ?Throwable $textError / $metadataError (with isComplete()). Text and metadata are extracted independently: an extractor throws only at its total-failure gate (container won't unzip, PDF won't parse, file unreadable) and otherwise records a failed half's error while still returning the salvaged half.- OOXML/ODF metadata: new AbstractOoxmlExtractor and AbstractOdfExtractor intermediaries declare each family's source map once (docProps/core.xml + app.xml; meta.xml); the concrete extractors only switch their extends clause. AbstractZipXmlExtractor gains a generic mapMetadataFromXml() walker.- PDF: reads the Info dictionary (Title/Author/Subject/Keywords/Created/ Modified/Producer-else-Creator) alongside the body text from the same parse. normalizePdfString() is a temporary shim for UTF-16BE literal strings that prinsfrank <= v3.1.0 returns mojibake-expanded.- Images: now metadata-only (text is ''); field-map keys renamed to the canonical vocabulary (Caption->Description, Date->Created, Camera->Producer).- helper: extract() and extractMetadata() added; extractText() keeps its string return and its throw-on-failure contract (rethrows textError), so the docsearch plugin keeps working unchanged.- cli: prints body text followed by metadata by default; --text/-t and --meta/-m restrict to one. A salvaged-but-partial extraction warns on STDERR and exits 0; nothing usable exits non-zero.- ExtractionException::wrap() centralises wrapping a caught error as an ExtractionException (passing one through if it already is).Tests assert the real Tika samples' embedded metadata and thepartial-failure paths. The pre-existing PDF body-text Form-XObjectfailure is unchanged.
show more ...
Initial implementation of the totext pluginExtract plain text from various file formats using PHP only (no shell-outs,no external binaries). Provides a CLI component (cli_plugin_totext) and ahelp
Initial implementation of the totext pluginExtract plain text from various file formats using PHP only (no shell-outs,no external binaries). Provides a CLI component (cli_plugin_totext) and ahelper component (helper_plugin_totext).Supported formats:- OOXML: docx, xlsx, pptx- OpenDocument: odt, ods, odp- PDF (via bundled smalot/pdfparser)- Plain-text family: txt, csv, md, markdown, log, text- Image metadata (EXIF/IPTC): jpg, jpeg, tif, tiffIncludes a fixture-based test suite covering every extractor, the factoryrouting, and the helper end to end.