Release 4.0.0 - 8/18/2026 This section is the complete delta from 3.x. It includes everything first released in 4.0.0-alpha-1 and 4.0.0-beta-1; those sections below are stubs. Upgrading from 3.x? Start with the migration guides at https://tika.apache.org/docs -- "Migrating to Tika 4.x", "Migrating Tika Server to 4.x" and "Metadata Changes in Tika 4.x". They carry the detail and the code examples behind the summaries here. Important architectural change: parsing now runs in a forked process where possible. tika-server's endpoints, tika-app's -a/--async and -f/--fork, and tika-grpc all parse in forked, crash-isolated tika-pipes workers. Applications embedding Tika should consider getting the same isolation from PipesForkParser (tika-pipes-fork-parser) rather than parsing in-process with AutoDetectParser. Note that the project does not treat denial of service -- memory exhaustion, CPU exhaustion, a crashed process -- as a security issue when files are parsed outside these isolated paths; see https://tika.apache.org/security-model.html. BREAKING CHANGES --- Platform, packaging, configuration and output format (everyone) --- * Tika 4.x requires Java 17 or later; 3.x built and ran on Java 11. All published artifacts are compiled with --release 17 (TIKA-4685). * tika-app and tika-server-standard now ship as zip distributions with an adjacent lib/ directory; the published jars are thin launchers and fail with NoClassDefFoundError if run on their own. This catches tika-server-standard hardest, because its jar is still on Maven Central: unzip the distribution and run from inside it (TIKA-4733). * tika-parsers-standard-package is now a pom, not a jar. Users must add pom in Maven or @pom in Gradle (TIKA-4712). * The default content handler is now Markdown. tika-app, tika-server (the /tika and /rmeta endpoints) and the async/pipes CLI emit Markdown content by default instead of XHTML/XML (plain text for the async CLI). Request the previous format explicitly: tika-app -x/--xml, the server's /tika/xml and /rmeta/xml paths, the async CLI --handler x (TIKA-4663). * Configuration moved from XML to JSON. TikaConfig and the org.apache.tika.config XML-configuration API are removed: TikaConfig, ConfigBase, Field, Param, ParamField, LoadErrorHandler, InitializableProblemHandler, TikaConfigSerializer and TikaTaskTimeout are gone. Use TikaLoader from tika-serialization. tika-app --convert-config-xml-to-json converts a 3.x parsers section as a starting point; every other section needs manual migration (TIKA-4544, TIKA-4545, TIKA-4553, TIKA-4565). * An unregistered component name in a default-parser, default-detector or default-encoding-detector "exclude" list now throws a TikaConfigException at config load instead of logging a WARN, so a 3.x config that named the component by class name or misspelled it now refuses to start. Use the registered name (e.g. "pdf-parser"); tika-app --list-parser-names prints them (TIKA-3268, TIKA-4808). * Metadata keys were renamed for consistency and provenance. Every Tika-asserted key now lives under a single tk: prefix, replacing 3.x's scattered X-TIKA:, tika:, tika_pg:, rendering:, signature: and imagereader: prefixes and bare names such as resourceName; names Tika coined inside format namespaces are kebab-cased (pdf:hasMarkedContent -> pdf:has-marked-content) while names from a file or an external standard keep their spelling; and open key families gained prefixes (audio:, ner:, envi:, ogg:streams-, grobid:, iso19115:, gdal:, geotopic:, mif:, idml:). Code using the TikaCoreProperties / TikaPagedText / Rendering constants is unaffected. Code that references keys by String has two paths: update the strings with the key-for-key tables in metadata-changes-4x.adoc, or turn on the compatibility filter below and migrate on your own schedule (TIKA-4816). * The opt-in legacy-key-migration-filter restores 3.x key spellings at the emit edge (default direction V4_TO_V3), so an unmigrated consumer keeps working against 4.x output; V3_TO_V4 maps 3.x names forward instead. tika-core bundles metadata-migration-3x-4x.json, the machine-readable rename/drop table (TIKA-4797). * The reserved tk: (and legacy X-TIKA:) namespace is now a trust boundary for String-keyed writes. Metadata#set/add(String, String) throw IllegalArgumentException on a reserved key instead of 3.x's silent success, where a document-controlled property named X-TIKA:Parsed-By could overwrite Tika's own value, and Property's public factories reject reserved names outright. Document- and tool-derived names now go through Metadata#add(KeyPrefix, String, String) -- append-only, skip-and-WARN on hostile names -- or its Instant overload for source-typed dates (TIKA-4816). * Metadata no longer implements CreativeCommons, Geographic, HttpHeaders, Message, ClimateForcast, TIFF or TikaMimeKeys: inherited constants move to their home interface, e.g. Metadata.CONTENT_TYPE becomes HttpHeaders.CONTENT_TYPE (now a Property, though the key string is unchanged). TikaMimeKeys and ClimateForcast are deleted outright; ClimateForecast (corrected spelling) replaces the latter, with its keys under cf: (TIKA-4816). * Other Metadata API changes: setAll(Properties) removed with no replacement -- it bypassed both the limiter and the reserved-key guard, so use putAll(Metadata) or individual set/add calls; PassthroughPrefix renamed KeyPrefix; the Property factories internalClosedChoise / internalOpenChoise / externalClosedChoise / externalOpenChoise renamed to ...Choice with no forwarders; the dead enum constants PropertyType.STRUCTURE and ValueType.{LOCALE, MIME_TYPE, PROPER_NAME, URL, XPATH} removed; package org.apache.tika.metadata.writefilter renamed to ...metadata.writelimiter. Metadata's serialVersionUID also changed, so a 3.x-serialized instance now fails with InvalidClassException instead of deserializing into an object that throws on first write (TIKA-4816). --- Java API (library integrators) --- * The core SPI signatures changed. Parser.parse takes a TikaInputStream instead of an InputStream (there is no InputStream overload), Detector.detect takes (TikaInputStream, Metadata, ParseContext), and EmbeddedDocumentExtractor's shouldParseEmbedded/parseEmbedded gained a ParseContext and take a TikaInputStream. Every third-party implementation must be updated; callers can wrap with TikaInputStream.get(...). The Tika facade still accepts an InputStream, but Tika.detect(InputStream, ...) no longer returns the caller's stream at its original position. The detector still resets the TikaInputStream it reads -- but that read-ahead is buffered inside an internal wrapper that detect() discards, so the caller's own stream comes back advanced. Pass a TikaInputStream you own (and rewind it), or re-open the source (TIKA-4399, TIKA-4541, TIKA-4569). * TikaInputStream no longer caches by default. A stream is consumed in passthrough mode unless enableRewind() is called at position 0; rewind()/getFile()/getPath() after reading without enableRewind() throw instead of silently spooling. A parser that read part of a stream and then asked for a file worked in 3.x and now fails. Digesters call enableRewind() themselves (TIKA-4618, TIKA-4623). * Parsing with a concrete parser (not AutoDetectParser) and an empty ParseContext no longer auto-generates an AutoDetectParser to handle embedded files: they are silently skipped, with no content and no exception. Nor does it auto-generate a Detector to identify them; they are reported as application/octet-stream instead. Set Parser.class and Detector.class in the ParseContext, or go through AutoDetectParser, which does this for you (TIKA-4819). * EmbeddedDocumentExtractorFactory and friends are removed; ParsingEmbeddedDocumentExtractor and UnpackExtractor are now stateless singletons (use INSTANCE) that take the enclosing ParseContext as a method parameter rather than capturing one at construction. Code that supplied a custom factory should bind an EmbeddedDocumentExtractor instance directly. EmbeddedDocumentUtil's instance API is likewise removed in favor of statics that take a ParseContext explicitly (TIKA-4819). * ParseContext configuration is now resolved per component instance rather than per config class, because a class-keyed write leaked one component's config to every other component binding the same config class. Two consequences: parseContext.get(SomeConfig.class) no longer returns a JSON-resolved config, so a third-party component following the PDFBoxRenderer pattern must be handed its config explicitly; and precedence is inverted -- a JSON config now beats a programmatic context.set(XConfig.class, ...), which used to win (TIKA-4808). * ForkParser and the entire org.apache.tika.fork package are removed from tika-core. Out-of-process parsing is now provided by PipesForkParser in the new tika-pipes-fork-parser module -- the recommended parser for untrusted documents. tika-app's -f/--fork routes through it, and --fork-timeout is rejected rather than silently ignored (TIKA-4554, TIKA-4571, TIKA-4651). * Unified timeout model across the library, pipes and server: a total-task budget plus a progress/stall timeout, composed recursively over embedded documents. TikaTimeoutException is now a checked exception, and several parser/pipes config fields were renamed (*TimeoutSeconds / *TimeoutMs -> *TimeoutMillis, including a unit change for Tess4J) (TIKA-4813). * Parsers and detectors no longer expose bean setters/getters for their settings. Configuration moves to per-component *Config objects supplied through the ParseContext (e.g. GeoParserConfig, DWGParserConfig, AmazonTranscribeConfig, MagikaDetector/SiegfriedDetector configs) (TIKA-4758). * The encoding detectors moved out of parser packages into org.apache.tika.detect.* and into new tika-encoding-detector-* modules: org.apache.tika.parser.txt.{CharsetDetector,CharsetMatch, Icu4jEncodingDetector,UniversalEncodingDetector,BOMDetector,...} are now org.apache.tika.detect.icu4j.*, org.apache.tika.detect.universal.* and org.apache.tika.detect.BOMDetector, and org.apache.tika.parser.html.HtmlEncodingDetector is now org.apache.tika.detect.html.HtmlEncodingDetector. NonDetectingEncodingDetector is removed (TIKA-4685, TIKA-4720). * MetadataListFilter has been renamed MetadataFilter, and the 3.x MetadataFilter has been removed (TIKA-4546). * API changes in the EmbeddedStreamTranslator (TIKA-4518), and DigestingParser is removed (TIKA-4607). * BasicContentHandlerFactory.parseHandlerType now throws IllegalArgumentException for an unrecognized handler name instead of silently returning the supplied default (TIKA-4809). --- tika-server --- * All parsing now runs out-of-process through tika-pipes. /tika, /rmeta, /meta, /unpack, /detect, /pipes and /async share a fixed pool of numClients forked worker JVMs (default derived from host cores), so a parser crash, OOM or timeout no longer takes down the server. The cost is a sizing decision 3.x never asked of you: numClients is both the server's concurrency ceiling and its CPU/memory footprint, and each fork's heap is set with pipes.forkedJvmArgs (e.g. -Xmx1g), not the server JVM's. Size both deliberately; see the cpu-sizing docs (TIKA-4809). * Capability flags are default-deny and split in two. enableUnsecureFeatures no longer exists -- a config still carrying it fails to start with an "Unrecognized field" error -- and is replaced by allowPipes (gates /pipes and /async) and allowPerRequestConfig (gates the /config endpoints and the multipart config part). /status is no longer gated and is enabled simply by listing it under endpoints. tika-grpc gains the same allowPerRequestConfig flag plus allowComponentModifications, which gates runtime Save/Delete of fetchers and pipes iterators (TIKA-4764). * Endpoints removed: /translate/* (unusable as shipped), /tika/main and /tika/form/main (Boilerpipe; use /tika/text), and the /tika/form family. The 3.x /tika/config and /tika/form/config forms are replaced by the /tika/config* multipart POSTs, which require allowPerRequestConfig (TIKA-4809). * Endpoints collapsed: /detect/stream is now /detect, and /language/stream and /language/string are both /language. Behavior changed with the rename: /detect now runs in the fork pool, so it can return 429, 503 or 413, and a failure reading the body is a 500 where 3.x returned 200 with application/octet-stream as if detection had succeeded; /language caps input at the first 100,000 characters and uses the default LanguageDetector on the classpath, where 3.x pinned Optimaize (TIKA-4809). * Output-format routing on /tika changed. The bare /tika endpoint returns Markdown (was XHTML); use /tika/xml for XHTML. /tika/text is body-only again, as in 3.x. /tika/json and /tika/config/json default to the server default (markdown) rather than hardcoded plain text. The Accept header no longer selects the output format -- 3.x routed bare /tika among plain text, HTML and XHTML by Accept (nondeterministically for */*); now the path names the format. An unrecognized handler name in the path is a 400 listing the valid types, instead of silently falling back to the default (TIKA-4663, TIKA-4809). * Per-request configuration headers are removed, and are now silently ignored if sent: writeLimit, throwOnWriteLimitReached, maxEmbeddedResources/maxEmbeddedCount, X-Tika-Handler and the meta_* metadata-injection family. The limits move to parse-context (output-limits.writeLimit, output-limits.throwOnWriteLimit, embedded-limits.maxCount); X-Tika-Handler becomes an explicit handler path; meta_* has no replacement, and with per-request config off by default a caller can no longer bound the output of a single request. The X-Tika-OCR* and X-Tika-PDF* families were removed earlier in the 4.x line (TIKA-4809). * Caller errors now map to accurate HTTP status codes instead of always returning 200 or 500. A saturated worker pool returns 429, a crashed/timed-out/OOM worker returns 503, an unknown or reserved fetcher/emitter or bad handler returns 400, and an over-limit body returns 413; the 429 and 503 responses carry a Retry-After header. Error bodies are now JSON ({"status":"TIMEOUT"}, with a message field when one is available) where 3.x returned plain text such as "Parse failed: TIMEOUT" (TIKA-4809). * The raw /tika family's 422 responses carry the extracted content only; the exception is no longer appended to the body -- use /rmeta for the structured exception (TIKA-4809). * /meta now runs through the same pipes-backed parser as the other extraction endpoints, so it gains their crash isolation and their error handling: a container exception comes back as 200 with tk:exception:container-exception instead of 500, and /meta/{field} returns 422 instead of 500 or 400. A request with no Accept header now returns JSON; 3.x returned CSV, still available via Accept: text/csv. /meta also no longer returns a language field -- it parses with the ignore handler, so there is no text to detect from; configure a language-detection metadata filter and use /rmeta or /tika/json instead (TIKA-4809). * /async requires an object body {"tuples":[...]} instead of a bare JSON array, validates fetcher/emitter ids at POST time (400), rejects a batch larger than the queue's total capacity with 400 instead of throttling it, and one bad tuple no longer stops the async workers. /pipes returns the same JSON body as /tika/rmeta/unpack -- {"status":, "message":...} -- instead of a /pipes-only {"status":"ok"|"process_crash"} shape, returns 400 with the reason for a malformed request body, and rejects emit strategies other than EMIT_ALL, whose passed-back data the /pipes response cannot carry (TIKA-4809). * Many server config keys were removed or renamed (logLevel, idBase, digest, returnStackTrace, port ranges, the spawn-child options, ...), and an unrecognized key now fails startup with an error naming it; see migrating-tika-server-4x.adoc for the key-by-key migration. One change no startup error will flag: taskTimeoutMillis is now parse-context.timeout-limits.totalTaskTimeoutMillis, and its default grew from 5 minutes to 1 hour (TIKA-4809, TIKA-4813). * Request bodies are now capped by maxRequestSizeBytes, defaulting to 1 GiB; larger requests are rejected with 413, including over-limit chunked uploads, which previously surfaced as an empty 500 (TIKA-4809). * The 'endpoints' allowlist now also gates SPI-provided resources; a discovered resource binds only when its root endpoint is enabled (TIKA-4809). * Fetcher-based streaming is removed: the InputStreamFactory pattern for fetching documents via the fetcherName/fetchKey headers is gone, and all documents now go through the pipes infrastructure. The no-op -a/--pluginsConfig flag is removed and now fails option parsing, --help exits 0, and the tika-server-client module is removed (TIKA-4809). --- tika-pipes and tika-grpc --- * tika-pipes implementation modules are now pf4j plugins, reorganized by resource (tika-pipes-solr) vs task (tika-pipes-fetcher-solr). Core classes moved to tika-pipes-core, and the file-system components moved out of it into their own tika-pipes-file-system plugin (TIKA-4334, TIKA-4519, TIKA-4543). * FetchEmitTuple JSON now names the per-tuple parse context "parse-context" (was "parseContext") and rejects unknown tuple fields with an error naming the field (TIKA-4809). * The pipes config keys staleFetcherTimeoutSeconds and staleFetcherDelaySeconds have been removed; a config still carrying them fails startup (TIKA-4809). * TimeoutLimits: progressTimeoutMillis of 0 combined with a positive totalTaskTimeoutMillis is now rejected at config load; it would kill every task immediately (TIKA-4809). * The http-fetcher now verifies TLS certificates and hostnames by default; set verifySsl:false to opt out (TIKA-4809). * SolrJ moves from 8.11.4 to 10.0.0; the Solr fetcher, emitter and pipes iterator no longer support Solr 8 (TIKA-4789). * tika-grpc: the generated Java classes moved from package org.apache.tika to org.apache.tika.pipes.grpc.proto, so every generated type moves and Java gRPC clients must update their imports. This is a source break only: the proto package ("tika") and the service name ("Tika") are unchanged, so the wire protocol is identical and clients in other languages are unaffected (TIKA-4808). --- tika-app and tika-eval-app --- * tika-app's batch mode is gone. The -bc/batch directory-to-directory command line (backed by the removed tika-batch module) has no successor flag; use -a/--async, which runs the same work through tika-pipes (TIKA-4333, TIKA-4340). * tika-core/tika-app: NetworkParser and tika-app's -c/--client= network-client mode were removed -- they dispatched raw sockets to an arbitrary user-supplied host with no auth or TLS. Use tika-server instead (TIKA-4808). * tika-eval-app's command line changed: the FileProfile sub-command is removed, the -bc batch-config option is gone, and extract directories are now named with -e/--extracts (Profile) and -a/--extractsA + -b/--extractsB (Compare); -i/--inputDir, -d/--db, -c/--config, -n/--numWorkers and -m/--maxExtractLength replace the 3.x spellings (TIKA-4342, TIKA-4450, TIKA-4452, TIKA-4507). * tika-app: the inline short forms -eX (output encoding) and -pX (document password) were removed from standard mode; use --encoding=X and --password=X. Prefix-matching them silently swallowed single-dash long names, so every single-dash long name is now rejected with a message naming the two-dash form (TIKA-4808). --- Parser, detector and output behavior --- * The tika-langdetect-tika module is removed (TikaLanguageDetector, LanguageIdentifier, LanguageProfile, LanguageProfilerBuilder, ProfilingWriter). tika-app, tika-server and tika-eval now bundle the new CharSoup detector (tika-langdetect-charsoup) instead of tika-langdetect-optimaize, so the language reported by default changes (TIKA-4662). * PDF: extractIncrementalUpdateInfo now defaults to true (was false), so every PDF parse emits pdf:incremental-update-count and related keys without configuration. parseIncrementalUpdates remains false (TIKA-4354, TIKA-4358). * Audio cover art is now extracted as embedded documents from MP3 (ID3v2 APIC/PIC), MP4 (covr), Vorbis and FLAC. Embedded-document counts and /rmeta list lengths for audio files change (TIKA-4801). * The DOM-based OOXML extractors are removed (XWPFWordExtractorDecorator, XSLFPowerPointExtractorDecorator, POIXMLTextExtractorDecorator, XPSTextExtractor) and with them the OfficeParserConfig keys useSAXDocxExtractor and useSAXPptxExtractor. The SAX extractors are the only implementation (TIKA-4692, TIKA-4708). * The MuPDF renderer is removed (org.apache.tika.renderer.pdf.mutool); PDF page rendering for OCR now uses the PDFBox renderer or the new PopplerRenderer (TIKA-4664). * The legacy ExternalParser is removed; external parsers now require explicit JSON configuration. CompositeExternalParser and ExternalParsersFactory, which loaded tika-external-parsers.xml definitions from the classpath automatically, are gone (TIKA-4707). * Headers are no longer injected into the body/content of MSG files (TIKA-4345). Please open a ticket if you need this behavior across email formats. --- Removed modules and classes --- * Removed modules with no direct replacement: tika-batch (TIKA-4333), tika-dl (TIKA-4499), the advanced media module (TIKA-4500), tika-fuzzing (TIKA-4506), tika-age-recogniser (TIKA-4343), tika-server-eval (TIKA-4555), the dotnet bindings (TIKA-4332) and snaps deployment (TIKA-4502). * Removed parsers and classes with no direct replacement: SentimentAnalysisParser (TIKA-4574); ObjectRecognitionParser and the Tensorflow recognisers/captioners; PooledTimeSeriesParser; org.apache.tika.parser.pdf.AccessChecker (replaced by PDFParserConfig.AccessCheckMode); and the jempbox-based JempboxExtractor / XMPMetadataExtractor / pdf.xmpschemas.* classes superseded by the unified XMP extractor (TIKA-4775). * Smaller removals: ExceptionUtils.trimMessage moved from tika-core to tika-eval-core; HttpClientFactory's inert redirect-host allowlist accessors (get/setAllowedHostsForRedirect) are gone (TIKA-4809); and org.apache.tika.utils.{RereadableInputStream,AnnotationUtils}, org.apache.tika.io.{IOUtils,InputStreamFactory}, org.apache.tika.sax.DIFContentHandler, org.apache.tika.parser.{AutoDetectParserFactory,ParserFactory} and org.apache.tika.parser.internal.Activator are removed. NEW FEATURES * tika-pipes gains three parse modes: NO_PARSE (detect only, no parse), CONTENT_ONLY (emitters write raw content, no metadata envelope) and UNPACK (write embedded bytes out). All three are wired through tika-app and tika-server (TIKA-4631, TIKA-4637, TIKA-4656). * New tika-pipes plugins: an Elasticsearch emitter (TIKA-4672), an Atlassian JWT fetcher (TIKA-4604), Google Drive, Microsoft Graph, Azure Blob, JSON and HTTP plugins, and an Apache Ignite ConfigStore for runtime fetcher/emitter configuration (TIKA-4583, TIKA-4587, TIKA-4598). * New inference and OCR modules: tika-inference and tika-vlm add vision-language-model parsers (Claude, Gemini, OpenAI) that emit vlm:prompt-tokens / vlm:completion-tokens, and tika-parser-tess4j-module adds in-process Tesseract OCR (TIKA-4665, TIKA-4666, TIKA-4667, TIKA-4690). * Unified XMP extraction across containers; adds HEIF/HEIC and WebP XMP, including Samsung/Google Motion Photo (TIKA-4775). * tika-app and tika-server can load extra jars (additional EncodingDetectors, Parsers, etc.) from the directory named by the -Dtika.extras.dir system property, without repackaging the application. Off by default; the directory is a trusted code location whose contents run with full process privileges. The jars are forwarded onto forked pipes/server workers too, so they are available where parsing happens (TIKA-4755). * A Markdown parser with structured, lossless XHTML output, complementing the Markdown content handler (TIKA-4770). * Content-based detection of ASN.1/DER crypto containers, enabled by default: the PKCS#7/CMS, PKCS#12 and RFC 5544 timestamped-data families gained magic where 3.x had globs only, and Pkcs7Parser refines the smime-type on the output Content-Type at parse time. Files that detected as application/octet-stream in 3.x may now detect as a crypto type; see configuration/detectors.adoc for the opt-in detect-time refinement (TIKA-1997, TIKA-2856). * New detection: Android binary XML (application/vnd.android.axml, TIKA-4747) and Frictionless Data packages (TIKA-4643); improved mp3/aac (TIKA-4612) and grib (TIKA-4655) detection. OTHER CHANGES * Release artifacts are now channel-specific. Maven Central gets slim per-module jars (plus pom, sources and javadoc); the Apache dist area gets runnable zip distributions (tika-app, tika-server-standard, tika-eval-app) and drop-in pf4j plugin zips; Docker Hub gets ready-to-run images. Fat/shaded artifacts no longer go to Maven Central (TIKA-4733). * The charset, junk-text and language detection stack was rewritten: language-aware charset detection, a universal junk detector, wider Unicode handling, the new CharSoup language detector, and more efficient common-token lookups via bloom filters (TIKA-4662, TIKA-4671, TIKA-4675, TIKA-4691, TIKA-4719, TIKA-4731, TIKA-4745, TIKA-4754, TIKA-4810). * Pipes now carries small documents to the forked worker inside the request instead of writing them to disk first. Content at or below the new pipes.maxInlineBytes (default 10 MB) rides in the request and is served in the worker by the reserved __bytes fetcher, touching no disk; larger content is written out once as before, and a stream already backed by a file keeps its file. Set maxInlineBytes to 0 to spool every non-empty body (TIKA-4808). * PipesClient/PipesServer IPC now enforces a configurable payload limit (pipes.maxIpcPayloadBytes, default 100 MB) in both directions. Results that exceed the limit return PAYLOAD_LIMIT_EXCEEDED instead of causing heap exhaustion; crash messages are also size-capped (TIKA-4793). * tika-server requests now carry only their own parse-context entries to the forked worker, which supplies the config defaults itself. Previously the server sent its config's parse-context with every request, so the worker treated the operator's timeout-limits as caller input and clamped them at pipes.maxTotalTaskTimeoutMillis (TIKA-4808). * tika-grpc now routes fetchAndParse through the PipesParser client pool instead of one shared single-threaded PipesClient: concurrent calls no longer crash the worker, pipes.numClients and (for the first time) pipes.useSharedServer take effect, and pool saturation surfaces in-band as CLIENT_UNAVAILABLE_WITHIN_MS. An interrupted call recycles its worker, so a pooled client cannot go back to the queue dirty (TIKA-4815). fetchAndParseServerSideStreaming now completes the call after delivering its reply, instead of leaving the client waiting forever (TIKA-4804). * New audio/video metadata: audio:bitrate, audio:is-variable-bitrate, audio:has-drm, audio:channels, video:frame-rate, video:bitrate and MP4 sample size; ID3 TCOP and Vorbis COPYRIGHT map to xmpDM:copyright, EXIF GPS altitude maps to geo:alt, and the presentation start of delayed QuickTime timed-metadata tracks is exposed (TIKA-4777, TIKA-4779, TIKA-4780, TIKA-4781, TIKA-4800, TIKA-4802). * Parsing and robustness fixes across formats: CHM (TIKA-4783), ID3 UTF-16 (TIKA-4784), MPEG2/2.5 Layer III frame sizing (TIKA-4791), .doc empty comments (TIKA-4718), OOXML hyphenation and field-code hyperlinks (TIKA-4646, TIKA-4683), RTF attachments in HTML decapsulation (TIKA-4710), image extraction (TIKA-4736), embedded-file extension calculation (TIKA-4808), and general media-file robustness (TIKA-4812). MAPI properties no longer overwrite better-fitting Dublin Core terms (TIKA-4806). Embedded-file naming was streamlined (TIKA-4689). * PDFs whose %PDF- header is preceded by a print-composition job ticket are no longer detected as text/x-matlab: up to 50 %% comment or blank lines may now precede it, extending the TIKA-3328 rule past its 512-byte reach (TIKA-4782). * MagicDetector now compiles its regular expression once, in the constructor, instead of recompiling it on every match (TIKA-4796). * tika-eval-core is no longer published as a fat jar (TIKA-4414) and tika-grpc no longer shades gRPC (TIKA-4709). * Fix concurrency bug in TikaToXMP (TIKA-4393). * Dependency upgrades throughout the 4.0.0 line, including Jetty 11 -> 12.1.12, CXF 4.0 -> 4.2.3 and SolrJ 8.11.4 -> 10.0.0, plus routine library updates (TIKA-4327). Release 4.0.0-beta-1 - 6/29/2026 Prerelease. Its changes are folded into the 4.0.0 section above. Release 4.0.0-alpha-1 - 5/4/2026 Prerelease. Its changes are folded into the 4.0.0 section above. Release 3.3.0 - 3/18/2026 * Switch to poi-ooxml-full (TIKA-4563). * Users need to add "allowAbsolutePaths=true" for the FileSystemFetcher to fetch an absolute path (TIKA-4649). * Add a markdown option for content handlers (TIKA-4563). * Improve zip parsing (TIKA-4650). * Add detection of compressed bmp (TIKA-4511). * Allow per file timeouts in tika-pipes (TIKA-4497). * Add matroska detector (TIKA-1180). * Allow multiple values for many Dublin Core keys (TIKA-4466). * Extract macros by default in tika-app's commandline and gui (TIKA-4472). * Improve extraction of Javascript from PDFs (TIKA-4465). Release 3.2.3 - 9/11/2025 * Allow backwards compatibility with versions of commons-compress before 1.28.0 (TIKA-4469). * Fix XFA parsing within PDFs when woodstox is on the classpath as in tika-server (TIKA-4482). * Dependency updates. Release 3.2.2 - 8/6/2025 * Fix for CVE-2025-54988. * Improve detection of encrypted ODT files (TIKA-4459). * Dependency updates (TIKA-4455). Release 3.2.1 - 6/26/2025 * Fix POIFSContainerDetector regression when wrapping an InputStream in a TikaInputStream (TIKA-4441). * Important bug fix for zip-based detection on a non-TikaInputStream (TIKA-4424). * Improve text extraction from EMF (TIKA-4432). * Dependency updates (TIKA-4421). Release 3.2.0 - 05/21/2025 * Detect inline images in MSG files (TIKA-4391). * Improve extraction of metadata in MSG files (TIKA-4381). * Fix concurrency bug in TikaToXMP (TIKA-4393). * Fix potential GDAL deadlock (TIKA-4385). * Improve extraction of properties from msg files (TIKA-4381). * Include internal attachment path in tika-eval reports (TIKA-4374). * Upgrade jsoup to 1.20.1 with workaround for change in self-closing tag behavior (TIKA-4419). * Upgrade dependencies (TIKA-4379). Release 3.1.0 - 01/28/25 * Allow users to turn off the injection of some headers into the content stream of MSG files (TIKA-4345). * Add a wrapper for Google's magika detector (TIKA-4344). * Add support for MachO via Alexey Pelykh (TIKA-4309). * Add logic to inject spaces in XPS files based on font widths via Ruairidh Williamson (TIKA-4315). * Mime type "application/json" is now a sub class of "text/javascript" not "application/javascript" (TIKA-4336) * Remove tagsoup from the project entirely. Note that some of the tags produced by the SourceCodeParser are slightly different (TIKA-4338) Release 3.0.0 - 10/15/2024 * Fix regression in TextAndCSVParser (TIKA-4278). Release 3.0.0-BETA2 - 07/09/2024 BREAKING CHANGES * Updated PST parser to use standard Message metadata keys and improved handling of embedded files (TIKA-4248). * Convenience methods for XML readers were moved from ParseContext to XMLReaderUtils (TIKA-4259). Other Changes * Add GRPC server (TIKA-4181). * Improved configurability in tika-pipes (TIKA-4243). * Add optional PST parser based on libpst/readpst (TIKA-4250). Release 3.0.0-BETA - 12/01/2023 BREAKING CHANGES * Require Java 11 (TIKA-4128). * The boilerpipe handler has been moved to the tika-handler-boiler-pipe package (TIKA-4138). * We've migrated HTML parsing to the JSoup parser instead of TagSoup. If you have a custom configuration on the HTMLParser, you'll need to change that to o.a.t.p.html.JSoupParser (TIKA-1599). * Removed xerces2 as a dependency (TIKA-4135). * tika-core now has a scope of "provided" in most non-app modules (TIKA-4191). * Tika will look for "custom-mimetypes.xml" directly on the classpath, NOT under "/org/apache/tika/mime/". (TIKA-4147). * Return media type "text/javascript" instead of "application/javascript" to follow RFC-9239. (TIKA-4119). Other Changes/Updates * Improve detection of sqlite3-based file formats (TIKA-4187). * Upgrade PDFBox to 3.0.1 (TIKA-3347) * Deprecated AbstractParser for removal in 4.x (TIKA-4132). * Fix bug in DateUtils that stripped timezone information from incoming Calendar objects (TIKA-4126). * The InputStreamDigester now calculates stream length (TIKA-4016). Release 2.9.0 - 8/23/2023 * With user configuration, the PDFParser can now throw an EncryptedDocumentException for Microsoft IRM PDF containers with encrypted payloads. Separately, the PDFParser now throws an EncryptedDocumentException instead of an IOException if the security handler cannot be found (TIKA-4082). * Fix bug that led to duplicate extraction of macros from some OLE2 containers (TIKA-4116). * Parse iframe's srcdoc as an embedded file (TIKA-3109). * Add detection of warc.gz as a specialization of gz and parse as if a standard WARC (TIKA-4048). * Allow users to modify the attachment limit size in the /unpack resource (TIKA-4039) * Fixed write limit bug in RecursiveParserWrapper (TIKA-4055). * Add mime detection for many files with thanks to Gregory Lepore (TIKA-3992). * Fixed iWork 13 keynote detection on files with wrong extension (TIKA-4111). Release 2.8.0 - 5/11/2023 * Enable counting and/or parsing of incremental updates in PDFs. This is an experimental feature and may change in later releases (TIKA-4017). * Fixed bug that prevented the the loading of CompositeExternalParser in tika-app and tika-server-standard. This parser will call exiftool and ffmpeg if those are installed, as was the behavior in Tika 1.x. Exclude org.apache.tika.parser.external.CompositeExternalParser if you do not want this behavior (TIKA-4022). * Removed the shading of tika-parsers-standard-module (TIKA-4038). * Enable optional extraction of file system metadata in FileSystemFetcher (TIKA-4035). * Allow pretty printing in FileSystemEmitter (TIKA-4034). * Add detection for and a new mime type for older postscript-based Adobe Illustrator "application/illustrator+ps" files (TIKA-3971). * Add magic detection for canon raw file types: crw, cr2 and cr3 (TIKA-3991). * Add detection for ONIX message files (TIKA-4011). * Add detection and a parser for ActiveMime files (TIKA-3987). * Add extraction of rendition layout value and version from Epub (TIKA-4013). * Improve embedded file extraction from PDFs (TIKA-4012). * Improve metadata extraction from WARCs (TIKA-4018). * Update to PDFBox 2.0.28 (TIKA-4016). * Users may now avoid the ZeroByteFileException via a setting on the AutoDetectParserConfig (TIKA-3976). * Fix bug in closing elements in the presence of elements in RTF files (TIKA-3972). * Improve extraction of embedded file names in .docx (TIKA-3968). * Normalize author, title, subject and description to their Dublin Core properties in the HTMLParser (TIKA-3963). Release 2.7.0 - 1/31/2023 * Add SVG detection for svg files that lack the xml header (TIKA-3308). * Migrate to a live fork of Universal Charset Detector (TIKA-3213). * Improve handling of text-based attachments inside .eml files (TIKA-3959). * Add tika-parser-nlp-package to release artifacts (TIKA-3958). * Remove need for element in classes that extend ConfigBase (TIKA-3946). * Add X-TIKA:embedded_id_path to ensure unique embedded file paths (TIKA-3942). * Fix bug that prevented digests when the fallback/EmptyParser was called (TIKA-3939). * Remove log4j 1.2.x (and slf4j-log4j12 which now redirects to slf4j-reload4j) from all modules (TIKA-3935). * Upgrade mime4j to 0.8.9 (TIKA-3950). * Refactor date parsing for emails (TIKA-3957) * Upgrade to Bouncy Castle 1.71 and jdk18on jars (TIKA-3933). * Add a JDBCPipesReporter (TIKA-3931). * Add multivalued field strategy option in jdbc-emitter (TIKA-3930). Default is now 'concatenate' with ', ' as the delimiter. * Downgrade logging in PipesClient for each parse from info to debug. Release 2.6.0 - 11/3/2022 * Add optional Siegfried detector (TIKA-3901). * Move OverrideDetector's functionality to the CompositeDetector (TIKA-3904). * The FileCommandDetector has been refactored to have the same behavior as the Siegfried detector; see setUseMime in the javadoc (TIKA-3902). * Fix bug in OpenSearch emitter that prevented upserts on documents with embedded files (TIKA-3882). * Extract PDF actions and triggers into the file's metadata (TIKA-3887). * Add a tika-async-cli module (TIKA-3885). * Fetch keys sent via headers to tika server are now URL decoded (TIKA-3864). Release 2.5.0 - 09/30/2022 * Improved extraction of PDF subset info for PDF/UA, PDF/VT, and PDF/X. NOTE: we no longer append PDF/A information, e.g. 'version="A-1b"' to the 'dc:format'. Users must now get that information from the 'pdfa:PDFVersion' key or from 'pdfaid:conformance' and 'pdfaid:part' (TIKA-3844). * Avoid infinite loop in bookmark extraction from PDFs (TIKA-3832). * Upgraded to slf4j 2.0.1 (TIKA-3842). * Added upsert option for the OpenSearch emitter (TIKA-3855). * Extract PDF signature information at the document level into the metadata (TIKA-3852). * Enable configuration of digests via AutoDetectParserConfig (TIKA-3853). * Use commons-io byte array streams via PJ Fanning (TIKA-3843). * Upgrade to PDFBox 2.0.27 (TIKA-3866). * Upgrade to JempBox 1.8.17 (TIKA-3856). * Add extraction of ODF version from ODF files (TIKA-3840). * tika-parser-html-commons (BoilerPipeHandler) is no longer a a dependency of tika-parser-html-module. tika-app and tika-server-standard have added a dependency on tika-parser-html-commons. However, users who are managing custom dependencies and who want the BoilerPipeHandler will have to now include the tika-parser-html-commons dependency (TIKA-1484). * Add unrar as an optional parser (TIKA-3800). * Refactor FuzzingCLI to use PipesParser (TIKA-3799). * ServiceLoader's loadServiceProviders() now guarantees unique classes (TIKA-3797). * Fix bug that prevented setting of includeHeadersAndFooters for xls, xlsx, doc and docx via tika-config (TIKA-3796). * Fix bug that prevented specification of rendered image type via http header in the PDFParser (TIKA-3794). * Fix bug causing some Exif dates to be decoded wrongly on timezones different than UTC (TIKA-3815). * Numerous dependency upgrades (TIKA-3795). * Add ALPHA-level initial releases of JDBCEmitter, FileSystemStatusReporter and OpenSearchPipesReporter. These may have breaking changes in subsequent releases. Release 2.4.1 - 06/14/2022 * Implement bulk upload in the OpenSearch emitter (TIKA-3791). * Implement tika-server client via pipes mode (TIKA-3790). * Custom embedded parsers and EmbeddedDocumentHandlers can now add metadata to the container file's metadata (TIKA-3789). * Record embedded file exceptions in the container file's metadata (TIKA-3788). * Allow continuation of parsing after write limit has been reached (TIKA-3787). * Allow pass-through of 'Content-Length' header to metadata in TikaResource (TIKA-3786). * Add embedded depth to profiles tables in tika-eval (TIKA-3775). * Add stop() method to TikaServerCli so that it can be run with Apache Commons Daemon (TIKA-1570). * Fixed bug in ordering of Parsers during service loading (TIKA-3750). * Users can expand system properties from the forking process into forked tika-server processes (TIKA-3748). * Fix a few files being wrongly detected as EML (TIKA-3771). * Fix ignoreCharsets param of Icu4jEncodingDetector (TIKA-3774). Release 2.4.0 - 04/23/2022 * NOTE: To save on resources, we no longer include the deeplearning4j dependencies in the tika-dl jar. The dependencies for the tika-dl package must be provided by users. See: https://github.com/apache/tika/blob/main/tika-parsers/tika-parsers-ml/tika-dl/pom.xml for the dependencies that must be provided at run-time (TIKA-3676). * NOTE: Added prefix "dwg-custom:" to DWG custom metadata properties (TIKA-3731). * Add initial, BETA-grade TLS encryption option for tika-server; configuration may change in future releases (TIKA-3719). * Allow specification of fetcherName and fetchKey via query parameters in request URI in tika-server (TIKA-3714). * Add basic parsers for WARC and WACZ in tika-parsers-standard (TIKA-3697). * Add MetadataWriteFilter capability to improve memory profile in Metadata objects (TIKA-3695). * Allow configurability of the ContentHandlerDecorator used by the AutoDetectParser (TIKA-3723). * Allow configurability of the EmbeddedDocumentExtractor used by the AutoDetectParser (TIKA-3711). * Add detection for Frictionless Data packages and WACZ (TIKA-3696). * Add detection for DGN files with gratitude and credit to Steven Frew's tika-dgn-detector (TIKA-3721). * Add parser for metadata from DGN 8 files via Dan Coldrick (TIKA-3721). * Add a fetcher and emitter for Azure blob storage (TIKA-3707). * Add detection for files encrypted by Microsoft's Rights Management Service (TIKA-3666). * Fixed regression in 2.3.0 that led to more embedded filenames than appropriate being written to the content (TIKA-3711). * tika-server now clones forking process' environment variables into forked process (TIKA-3715). * Add an optional /eval endpoint for tika-eval profile or compare capabilities in tika-server (TIKA-3689). * Add a Parsed-By-Full-Set metadata item to record all parsers that processed a file (TIKA-3716). * Add metadata filters for Optimaize and OpenNLP language detectors (TIKA-3717). * Upgrade to PDFBox 2.0.26 (TIKA-3726). * Upgrade deeplearning4j to 1.0.0-M2 (TIKA-3458 and PR#527). * Various dependency upgrades, including POI, dl4j, gson, jackson, twelvemonkeys, log4j2 and others (TIKA-3675 and many PRs from dependabot). * Switch cipher from ECB to GCM in HttpClientFactory (TIKA-3724). Release 2.3.0 - 02/02/2022 * Upgrade to Apache POI 5.2.0. This is the first upgrade to POI 5.x and represents a major refactoring. Users may experience significantly more logging (TIKA-3164). * Upgrade to log4j2 2.17.1 (TIKA-3638). * Improve consistency in reporting package-entry divs across all parsers for embedded files (TIKA-3644). This leads to some more text (embedded file names) in files with many embedded attachments. * Improve configuration of maps as params for parsers in TikaConfig (TIKA-3645). * Improve identification of iWorks 13 files and add parsing for thumbnails, some metadata and attachments (TIKA-3634). Skip handling of .iwa files, which are not yet supported. * Limit the default in-memory processing (maxMainMemoryBytes) in the PDFParser to 512MB as in the 1.x branch (TIKA-3642). * Added IDML Parser from 1.x series to 2.x series (TIKA-3188). * Extract annotation types and subtypes for PDFs into metadata (TIKA-3653). * Add metadata value for PDFs that contain 3D annotations (TIKA-3653). * Add parser for Translation Memory eXchange (TMX) files (TIKA-3660). * Add Bill of Materials (Maven BOM) for centralized module version management (TIKA-3367). Release 2.2.1 - 12/19/2021 * Fix multithreading bug for ooxml files (TIKA-3627). * Upgrade log4j to 2.17.0 (TIKA-3625). * Upgrade to PDFBox 2.0.25 (TIKA-3622) * Fix bug that prevented metadata keys in the UnpackerResource in tika-server (TIKA-3624). * Upgrade log4j to 2.16.0 (TIKA-3623) Release 2.2.0 - 12/13/2021 * Add support for OneNote files downloaded from O365 (TIKA-3446). * Fix logic bug in PipesServer that prevented concatenation of content from attachments (TIKA-3609). * Improve extraction of embedded files from MSOffice files created by non-Microsoft tools (TIKA-3526). * Added back ability to ignore load errors in TikaConfig (TIKA-3575). * Make SecureContentHandler and other parameters configurable in AutoDetectParser programmatically and via tika-config.xml (TIKA-3594). * Fix default logging in tika-app in batch mode (TIKA-3589). * Fix bug that prevented specifying a config with the long --config= option in tika-app in batch mode (TIKA-3589). * Fix thread starvation after numerous restarts in PipesClient (TIKA-3588). * Fix race condition when starting multiple forked servers on multiple ports (TIKA-3586). * Add timeout per task to be configured via headers for tika-server's legacy endpoints /tika and /rmeta. Note that this timeout greater than taskTimeoutMillis (TIKA-3582). * Add metadata item for whether or not a PDF has a collection/ is a Portfolio PDF (TIKA-3579). * Add detection of ESRI Layer files (TIKA-3570). * Add detection of JPEG XL, MARC, ICC profiles, NES-ROM file types (TIKA-3562 and TIKA-3563) * Remove duplicate "subject" metadata keys that were intended for backwards compatibility within 1.x only (TIKA-3564). * Fix Open Office mime types to be subclasses of application/zip and no longer require OPCPackageDetector-last ordering of zip detectors (TIKA-3556). * Improve robustness and features of the httpfetcher (TIKA-3543) * Add optional fetch ranges to FetchEmitTuple to allow range fetching from, e.g. http or s3 (TIKA-3542). * Exclude dependencies on jsoup and ehcache in ucar grib/cdm (TIKA-3003). Release 2.1.0 - 08/18/2021 MAJOR CHANGES in 2.1.0: * Improved packaging for tika-parsers-extended. Use the tika-parser-scientific-package and tika-parser-sqlite3-package artifacts if you want fat jars with dependencies. (TIKA-3510) * Tika app writes UTF-8 when an encoding is not specified; the legacy behavior was UTF-8 on Mac OS, but System default on other OSs (TIKA-3515). * Change the default rendering strategy for PDFs from NO_TEXT to ALL (TIKA-3520). Other changes: * Fixed bug that pointed to the wrong tessdata directory if the user specified a tesseract path but not also a tessdata path (TIKA-3518). * Fixed bug in Icu4j's encoding detector where it would return non-standard names for charsets, e.g. IBM424_rtl is now returned as IBM424 (TIKA-3516). * Add a simple UrlFetcher in tika-core as a basic alternative to tika-fetcher-http (TIKA-3527). * Add tika-pipes support for Google Cloud Storage (TIKA-3524). * Fix markup ordering errors in xhtml output for ODT files (TIKA-2242). * Fix serialization of embedded docs in OpenSearch emitter and fix embedded documents not being indexed in some use cases in the Solr emitter (TIKA-3490). * Add pipesClientId system property to PipesServer so that each forked process can log to its own logger (TIKA-3480). * Add DateNormalizingMetadataFilter let users ensure that all dates emitted to Solr/OpenSearch are in UTC. Users can configure which timezone they'd like to use in cases where the file format does not store a timezone (TIKA-3496). * Breaking change in the Solr and OpenSearch emitters. To achieve the SKIP or CONCATENATE attachment strategy, modify the parseMode in the pipesiterators or in the FetchEmitTuple (TIKA-3494). Release 2.0.0 - 07/07/2021 * Cleanup of fetcher integration with tika-server. * Update dependencies. Release 2.0.0-BETA - 05/19/2021 * Refactor pipes module for resilience * Add transcribe capability (TIKA-94). Release 2.0.0-ALPHA - 01/13/2021 BREAKING CHANGES in 2.0.0 * General * OCR is now triggered automatically for PDFs if tesseract is on the user's path see (https://cwiki.apache.org/confluence/display/TIKA/TikaOCR#TikaOCR-disable-ocr) for how to disable OCR. * We upgraded from log4j to log4j2 in tika-app, tika-server and anywhere else we used to use log4j. * By default, when rendering a page for OCR, the PDFParser does not render glyphs/text. * Removed deprecated Metadata keys/properties (TIKA-1974). * Removed deprecated PDFPreflightParser (TIKA-3437). * Removed dangerous calls to read an inputstream or convert to bytes without specifying a charset * Parsers can be configured via tika-config.xml on instantiation. We have moved away from configuration via .properties files because of confusion among users. This affects the PDFParser, TesseractOCRParser and the StringsParser. * Changed namespaces of translator implementations (o.a.t.language.translate.impl) to avoid split-package with tika-core * tika-parsers * The parser modules have been broken into three main modules: tika-parsers-standard, tika-parsers-extended and tika-parsers-ml. Users may now need to add tika-parsers-extended's tika-parser-scientific-module or tika-parser-sqlite3-module to tika-app and tika-server to include parsers that used to be included by default (for example: envi, gdal, grib, isatab, netcdf, sqlite3). * PDFParser -- a) see above on OCR. b) This parser no longer warns if the jpeg2000 dependency is not included. Tika now relies on PDFBox to log an error if a jpeg2000 image should be processed but can't because the required external dependency is not available. See https://pdfbox.apache.org/2.0/dependencies.html#jai-image-io for the non-ASF-2.0-compatible jpeg2000 library. * CompressorParser -- users must add the com.github.luben:zstd-jni dependency to the classpath to process zstd files. This is an optional library that is no longer bundled in tika-parsers-standard-package because it contains native libs. * ChmParser was moved to org.apache.tika.parser.microsoft.chm * RTFParser was moved to org.apache.tika.parser.microsoft.rtf * We are now using non-shaded versions of xmpcore with namespaces com.adobe.internal.* vs com.adobe.*. * tika-app * See above on default inclusion of only tika-parsers-standard. * tika-server * tika-server now by default forks a process to isolate the parsing in the forked process (this was called the -spawnChild option in tika-1.x). Clients must now expect that tika-server will restart on OOM, timeouts, crashes or after parsing a large number of files. When this happens tika-server will restand and not receive connections for brief periods. The less robust, legacy behavior of not forking a process is available with "-noFork"= * Most of tika-server's legacy configuration via the commandline has been moved into configuration via a tika-config.xml file. * tika-server's "enableFileUrl" has been removed in favor of a FileSystemFetcher. * tika-server's /metadata endpoint requires tika-server-standard to write XMP/rdf output. This output is not available in tika-server-core. * In tika-server, for those parsers that can be configured per parse via a config object passed in through the ParseContext, the config object will only update those fields that the user has modified. The config object will no longer fully reset all settings to the default settings per parse. This has a more intuitive "update the base/configured settings" with what has been changed in the config object. * tika-eval * tika-eval's default profile and comparison reports no longer include tag reports. Users can get the report configs that include tags (*-tags.xml): https://github.com/apache/tika/tree/main/tika-eval/tika-eval-app/src/main/resources Release 1.27 - 06/30/2021 * Migrate MP4 parsing to Drew Noakes' metadata-extractor (TIKA-3459). To revert to legacy parser turn off NoakesMP4Parser and turn on MP4Parser via tika-config.xml. * Prevent rare infinite loop in tika-server's -spawnChild mode when restart fails because of failure to bind to the port (TIKA-3441). * Improve likelihood that tesseract will not be orphaned on jvm restart in tika-server (TIKA-3441). * Deprecate experimental PDFPreflightParser (TIKA-3437). * Apply encoding detection to zip entry names via Ryan421 (TIKA-3374). * Add json output for /tika endpoint in tika-server (TIKA-3352). * Tika's PDFParser should use the underlying file if one is passed in via a TikaInputStream (TIKA-3350) Release 1.26 - 03/24/2021 * Fix thread safety bug in OpenOffice parser (TIKA-3334). * The "writeLimit" header now pertains to the combined characters written per container document (and embedded documents) in the /rmeta endpoint in tika-server (TIKA-3325); it no longer functions only per container or embedded document. * Extract more embedded files in PDFs by recursively processing the embedded file tree (TIKA-3332). * Allow for case insensitive headers for configuration of the PDFParser and the TesseractOCRParser in tika-server via Subhajit Das (TIKA-3320). * Improve detection and parsing of XPS files (TIKA-3316). * General dependency upgrades (TIKA-3244). * Great optimization in ForkParser (TIKA-3237). * Fix parsing of emails attached to other emails in PST files (TIKA-3004). * MP3 parser should output the xmpDM:duration metadata as seconds not milliseconds, consistent with the other Audio and Video parsers (TIKA-3318). * MP4 parser check if any of the Compatible Brands match when identifying the subtype (TIKA-3310). Release 1.25 - 11/25/2020 * Fix inconsistent license in xmpcore (TIKA-3204). * General upgrades including some dependencies with recently found security vulnerabilities (TIKA-3119). * Add detection and a parser for flat ODF files (TIKA-3159). * Add extraction of macros from ODF files (TIKA-3161). * Add mime detection for hprof and hprof text files (TIKA-3144). * Add TextSignature and TextProfileSignature to tika-eval (TIKA-3145 and TIKA-3146) * Create a metadata filter to trigger tika-eval stats post parsing (TIKA-3140) * Add a configurable metadata-filter for the RecursiveParserWrapper (TIKA-3137) * Parameterize writeLimit and maxEmbeddedResources for RecursiveParserWrapper in tika-server (TIKA-3133) * Add status endpoint to tika-server (TIKA-3129). * Remove whitelist/blacklist terminology (TIKA-3120) * Add detection for parquet files (TIKA-3115). * Add detection and parsing for bplist (TIKA-3104). * Enable metadata value filtering for RecursiveParserWrapper (TIKA-3137) * Add a basic parser for plist files based on com.googlecode.plist:dd-plist (TIKA-3104). * Read hyperlinked images from ODT files (TIKA-3156). * Updated GrobidRESTParser to use new API location (TIKA-3191). * Add FileProfiler to tika-eval (TIKA-3216). * Add status endpoint to tika-server (TIKA-3129). * Improved handling of zip files with STORED entries with data descriptor (TIKA-3196). * Add parsers for XLZ, IDML and MIF (TIKA-2976, TIKA-3188 and TIKA-3189). * Add the beginnings of a format-aware fuzzing module (TIKA-3083). * Add wrapper for Linux 'file' command for mime detection (TIKA-3215). * Added ability to skip parsing of embedded files in Tika Server (TIKA-3227). Release 1.24.1 - 4/17/2020 * Allow gzip compression of input and output streams for tika-server (TIKA-3073). Release 1.24 - 3/11/2019 * Add scripts to run tika-server as a service via Eric Pugh, and add these scripts and jar as a new artifact in the release (TIKA-3010). * Upgrade Drew Noakes' metadata-extractor (TIKA-2952). * Enable optional extraction of structural tags in PDFs (alpha-grade) (TIKA-3026). * Tika app's --extract mode now outputs to STDOUT (TIKA-3035). * Add an optional Preflight parser for PDFs (TIKA-3055). * Improve detection of some zip-based formats (TIKA-3057). * Upgrade metadata-extractor to 2.13.0 (TIKA-2952). * Upgrade to POI 4.1.2 (TIKA-3047). * Extract XMP from PSD files (TIKA-3050). * Added XMLProfiler as an optional parser to profile XFA and XMP in PDFs (TIKA-3045). * Extract inline images that rely on the DCT filter from PDFs (TIKA-3041). * Upgrade to PDFBox 2.0.19 (TIKA-3033). * Fix bug in ASM parser configuration (TIKA-2992). * Upgrade to java-libpst 0.9.3 (TIKA-2546). * Fixed XLIFF12Parser failures with ToXMLHandler (TIKA-3014). Release 1.23 - 12/02/2019 * NOTE: The PDFParser now relies on OCRDPI to render page images when users configure OCR on rendered page images. This will have the effect of increasing rendered image size (TIKA-2624). * NOTE: tika-server no longer returns 415 for file types for which there is no parser. * Fix bug in AUTO OCR strategy in the PDFParser (TIKA-3002). * Fix incorrect height and width metadata extraction from JPEG images (TIKA-2630). * Upgrade to POI 4.1.1 (TIKA-2851). * Upgrade to PDFBox 2.0.17 (TIKA-2951). * Ensure that the PDFParser respects custom configuration of Tesseract from tika-config.xml via Eric Pugh (TIKA-2970). * Add parser for XLIFF v1.2 files (TIKA-2975). * Add mime type detection support for WebAssembly (TIKA-2894), HEIF / HEIC images (TIKA-2942), Digilite FDF (TIKA-2988); and xml-root detection for XFDF (TIKA-2990) and XDP (TIKA-2989). * Add an XLZ Parser (TIKA-2976). * Fix deadlock with ForkParser when InputStream throws IOException (TIKA-2892). Release 1.22 - 07/29/2019 * NOTE: tika-server no longer hard-codes the HtmlParser to handle XML files (TIKA-2910). Users must now configure that behavior via a tika-config.xml file. * NOTE: Known regression: PDFBOX-4587 -- PDF passwords with codepoints between 0xF000 and 0XF0000 will cause an exception. * Add parser for HWP v5 files via SooMyung Lee (soomyung) and JinSup Kim (ddoleye) (TIKA-2909). * Fix order of closing streams to avoid "Failed to close temporary resource" exception in TesseractOCRParser (TIKA-2908). * Improve AutoDetectReader performance by caching encoding detector (TIKA-1568). * Prevent RTFParser from outputting illegal tag combinations (TIKA-2889). * Fix RereadableInputStream to release all resources (TIKA-2903). * Implement custom language identifier in the tika-eval module based on OpenNLP's language detector; add 18 languages and add common words lists for all 121 languages (TIKA-2790). * Fix NPE in MimeTypesReader.releaseParser() via Eamonn Saunders (TIKA-2896). * Fix RTFParser to extract more content (TIKA-2883). * Add clientSubmitTime to the metadata extracted from PST files (TIKA-2898). * Improve StreamingZipContainerDetector for xltx, xltm and several other file formats (TIKA-2886). Release 1.21 - 05/14/2019 * Add optional AUTO mode to OCR'ing of PDFs. If tesseract is installed and on the path, and this option is selected programmatically or via TikaConfig(), the PDFParser will use heuristics to decide whether or not to run OCR per page on PDFs. (TIKA-2749) * The ZipContainerDetector's default behavior was changed to run streaming detection up to its markLimit. Users can get the legacy behavior (spool-to-file/rely-on-underlying-file-in-TikaInputStream) by setting markLimit=-1. The POIFSContainerDetector requires an underlying file; it will try to spool the file to disk; if the file's length is > markLimit, it will not attempt detection; set markLimit to -1 for legacy behavior (TIKA-2849). * Upgrade PDFBox to 2.0.14 (TIKA-2834). * Add CSV detection and replace TXTParser with TextAndCSVParser; users can turn off CSV detection by excluding the TextAndCSVParser and adding back the TXTParser via tika-config (TIKA-2833). * Add a CSVParser. CSV detection is currently based solely on filename and/or information conveyed via Metadata (TIKA-2826). * General upgrades: asm, bouncycastle, commons-codec, commons-lang3, cxf, guava, h2, httpcomponents, jackcess, junrar, Lucene, mime4j, opennlp, parso, sqlite-jdbc (provided), zstd-jni (provided) (TIKA-2824) * Bundle xerces2 with tika-parsers (TIKA-2802). * Upgrade jaxb to 2.3.2 (TIKA-2819). * Upgrade jackson to 2.9.8 (TIKA-2717). * Update tika-eval's common tokens lists (TIKA-2822). * Handle bad tags in tika-eval more robustly (TIKA-2810). * Add reports for tags in tika-eval (TIKA-2809). * Extract text from SDT element within textboxes in .docx files (TIKA-2807). * Try to handle truncated OOXML files more robustly (TIKA-2765). Release 1.20 - 12/17/2018 * Upgrade to POI 4.0.1 (TIKA-2751). * Integrate/parameterize new angles handling in PDFBox (TIKA-2779). * Upgrade to PDFBox 2.0.13 (TIKA-2788). * Prevent content within