Release 4.0.0 - 8/18/2026
This section is the complete delta from 3.x. It includes everything first
released in 4.0.0-alpha-1 and 4.0.0-beta-1; those sections below are stubs.
Upgrading from 3.x? Start with the migration guides at
https://tika.apache.org/docs -- "Migrating to Tika 4.x", "Migrating Tika
Server to 4.x" and "Metadata Changes in Tika 4.x". They carry the detail
and the code examples behind the summaries here.
Important architectural change: parsing now runs in a forked process
where possible. tika-server's endpoints, tika-app's -a/--async and -f/--fork,
and tika-grpc all parse in forked, crash-isolated tika-pipes workers.
Applications embedding Tika should consider getting the same isolation from
PipesForkParser (tika-pipes-fork-parser) rather than parsing in-process with
AutoDetectParser. Note that the project does not treat denial of service --
memory exhaustion, CPU exhaustion, a crashed process -- as a security issue
when files are parsed outside these isolated paths; see
https://tika.apache.org/security-model.html.
BREAKING CHANGES
--- Platform, packaging, configuration and output format (everyone) ---
* Tika 4.x requires Java 17 or later; 3.x built and ran on Java 11. All
published artifacts are compiled with --release 17 (TIKA-4685).
* tika-app and tika-server-standard now ship as zip distributions with an
adjacent lib/ directory; the published jars are thin launchers and fail
with NoClassDefFoundError if run on their own. This catches
tika-server-standard hardest, because its jar is still on Maven Central:
unzip the distribution and run from inside it (TIKA-4733).
* tika-parsers-standard-package is now a pom, not a jar. Users must add
pom in Maven or @pom in Gradle (TIKA-4712).
* The default content handler is now Markdown. tika-app, tika-server (the
/tika and /rmeta endpoints) and the async/pipes CLI emit Markdown content
by default instead of XHTML/XML (plain text for the async CLI). Request
the previous format explicitly: tika-app -x/--xml, the server's /tika/xml
and /rmeta/xml paths, the async CLI --handler x (TIKA-4663).
* Configuration moved from XML to JSON. TikaConfig and the
org.apache.tika.config XML-configuration API are removed: TikaConfig,
ConfigBase, Field, Param, ParamField, LoadErrorHandler,
InitializableProblemHandler, TikaConfigSerializer and TikaTaskTimeout are
gone. Use TikaLoader from tika-serialization. tika-app
--convert-config-xml-to-json converts a 3.x parsers section as a starting
point; every other section needs manual migration
(TIKA-4544, TIKA-4545, TIKA-4553, TIKA-4565).
* An unregistered component name in a default-parser, default-detector or
default-encoding-detector "exclude" list now throws a TikaConfigException
at config load instead of logging a WARN, so a 3.x config that named the
component by class name or misspelled it now refuses to start. Use the
registered name (e.g. "pdf-parser"); tika-app --list-parser-names prints
them (TIKA-3268, TIKA-4808).
* Metadata keys were renamed for consistency and provenance. Every
Tika-asserted key now lives under a single tk: prefix, replacing 3.x's
scattered X-TIKA:, tika:, tika_pg:, rendering:, signature: and
imagereader: prefixes and bare names such as resourceName; names Tika
coined inside format namespaces are kebab-cased (pdf:hasMarkedContent ->
pdf:has-marked-content) while names from a file or an external standard
keep their spelling; and open key families gained prefixes (audio:, ner:,
envi:, ogg:streams-, grobid:, iso19115:, gdal:, geotopic:, mif:, idml:).
Code using the TikaCoreProperties / TikaPagedText / Rendering constants is
unaffected. Code that references keys by String has two paths: update the
strings with the key-for-key tables in metadata-changes-4x.adoc, or turn
on the compatibility filter below and migrate on your own schedule
(TIKA-4816).
* The opt-in legacy-key-migration-filter restores 3.x key spellings at the
emit edge (default direction V4_TO_V3), so an unmigrated consumer keeps
working against 4.x output; V3_TO_V4 maps 3.x names forward instead.
tika-core bundles metadata-migration-3x-4x.json, the machine-readable
rename/drop table (TIKA-4797).
* The reserved tk: (and legacy X-TIKA:) namespace is now a trust boundary
for String-keyed writes. Metadata#set/add(String, String) throw
IllegalArgumentException on a reserved key instead of 3.x's silent
success, where a document-controlled property named X-TIKA:Parsed-By could
overwrite Tika's own value, and Property's public factories reject reserved
names outright. Document- and tool-derived names now go through
Metadata#add(KeyPrefix, String, String) -- append-only, skip-and-WARN on
hostile names -- or its Instant overload for source-typed dates
(TIKA-4816).
* Metadata no longer implements CreativeCommons, Geographic, HttpHeaders,
Message, ClimateForcast, TIFF or TikaMimeKeys: inherited constants move to
their home interface, e.g. Metadata.CONTENT_TYPE becomes
HttpHeaders.CONTENT_TYPE (now a Property, though the key string is
unchanged). TikaMimeKeys and
ClimateForcast are deleted outright; ClimateForecast (corrected spelling)
replaces the latter, with its keys under cf: (TIKA-4816).
* Other Metadata API changes: setAll(Properties) removed with no replacement
-- it bypassed both the limiter and the reserved-key guard, so use
putAll(Metadata) or individual set/add calls; PassthroughPrefix renamed
KeyPrefix; the Property factories internalClosedChoise / internalOpenChoise
/ externalClosedChoise / externalOpenChoise renamed to ...Choice with no
forwarders; the dead enum constants PropertyType.STRUCTURE and
ValueType.{LOCALE, MIME_TYPE, PROPER_NAME, URL, XPATH} removed; package
org.apache.tika.metadata.writefilter renamed to ...metadata.writelimiter.
Metadata's serialVersionUID also changed, so a 3.x-serialized instance now
fails with InvalidClassException instead of deserializing into an object
that throws on first write (TIKA-4816).
--- Java API (library integrators) ---
* The core SPI signatures changed. Parser.parse takes a TikaInputStream
instead of an InputStream (there is no InputStream overload),
Detector.detect takes (TikaInputStream, Metadata, ParseContext), and
EmbeddedDocumentExtractor's shouldParseEmbedded/parseEmbedded gained a
ParseContext and take a TikaInputStream. Every third-party implementation
must be updated; callers can wrap with TikaInputStream.get(...). The Tika
facade still accepts an InputStream, but Tika.detect(InputStream, ...) no
longer returns the caller's stream at its original position. The detector
still resets the TikaInputStream it reads -- but that read-ahead is
buffered inside an internal wrapper that detect() discards, so the
caller's own stream comes back advanced. Pass a TikaInputStream you own
(and rewind it), or re-open the source
(TIKA-4399, TIKA-4541, TIKA-4569).
* TikaInputStream no longer caches by default. A stream is consumed in
passthrough mode unless enableRewind() is called at position 0;
rewind()/getFile()/getPath() after reading without enableRewind() throw
instead of silently spooling. A parser that read part of a stream and then
asked for a file worked in 3.x and now fails. Digesters call
enableRewind() themselves (TIKA-4618, TIKA-4623).
* Parsing with a concrete parser (not AutoDetectParser) and an empty
ParseContext no longer auto-generates an AutoDetectParser to handle
embedded files: they are silently skipped, with no content and no
exception. Nor does it auto-generate a Detector to identify them; they are
reported as application/octet-stream instead. Set Parser.class and
Detector.class in the ParseContext, or go through AutoDetectParser, which
does this for you (TIKA-4819).
* EmbeddedDocumentExtractorFactory and friends are removed;
ParsingEmbeddedDocumentExtractor and UnpackExtractor are now stateless
singletons (use INSTANCE) that take the enclosing ParseContext as a method
parameter rather than capturing one at construction. Code that supplied a
custom factory should bind an EmbeddedDocumentExtractor instance directly.
EmbeddedDocumentUtil's instance API is likewise removed in favor of statics
that take a ParseContext explicitly (TIKA-4819).
* ParseContext configuration is now resolved per component instance rather
than per config class, because a class-keyed write leaked one component's
config to every other component binding the same config class. Two
consequences: parseContext.get(SomeConfig.class) no longer returns a
JSON-resolved config, so a third-party component following the
PDFBoxRenderer pattern must be handed its config explicitly; and precedence
is inverted -- a JSON config now beats a programmatic
context.set(XConfig.class, ...), which used to win (TIKA-4808).
* ForkParser and the entire org.apache.tika.fork package are removed from
tika-core. Out-of-process parsing is now provided by PipesForkParser in
the new tika-pipes-fork-parser module -- the recommended parser for
untrusted documents. tika-app's -f/--fork routes through it, and
--fork-timeout is rejected rather than silently ignored
(TIKA-4554, TIKA-4571, TIKA-4651).
* Unified timeout model across the library, pipes and server: a total-task
budget plus a progress/stall timeout, composed recursively over embedded
documents. TikaTimeoutException is now a checked exception, and several
parser/pipes config fields were renamed (*TimeoutSeconds / *TimeoutMs ->
*TimeoutMillis, including a unit change for Tess4J) (TIKA-4813).
* Parsers and detectors no longer expose bean setters/getters for their
settings. Configuration moves to per-component *Config objects supplied
through the ParseContext (e.g. GeoParserConfig, DWGParserConfig,
AmazonTranscribeConfig, MagikaDetector/SiegfriedDetector configs)
(TIKA-4758).
* The encoding detectors moved out of parser packages into
org.apache.tika.detect.* and into new tika-encoding-detector-* modules:
org.apache.tika.parser.txt.{CharsetDetector,CharsetMatch,
Icu4jEncodingDetector,UniversalEncodingDetector,BOMDetector,...} are now
org.apache.tika.detect.icu4j.*, org.apache.tika.detect.universal.* and
org.apache.tika.detect.BOMDetector, and
org.apache.tika.parser.html.HtmlEncodingDetector is now
org.apache.tika.detect.html.HtmlEncodingDetector.
NonDetectingEncodingDetector is removed (TIKA-4685, TIKA-4720).
* MetadataListFilter has been renamed MetadataFilter, and the 3.x
MetadataFilter has been removed (TIKA-4546).
* API changes in the EmbeddedStreamTranslator (TIKA-4518), and
DigestingParser is removed (TIKA-4607).
* BasicContentHandlerFactory.parseHandlerType now throws
IllegalArgumentException for an unrecognized handler name instead of
silently returning the supplied default (TIKA-4809).
--- tika-server ---
* All parsing now runs out-of-process through tika-pipes. /tika, /rmeta,
/meta, /unpack, /detect, /pipes and /async share a fixed pool of
numClients forked worker JVMs (default derived from host cores), so a
parser crash, OOM or timeout no longer takes down the server. The cost is
a sizing decision 3.x never asked of you: numClients is both the server's
concurrency ceiling and its CPU/memory footprint, and each fork's heap is
set with pipes.forkedJvmArgs (e.g. -Xmx1g), not the server JVM's. Size
both deliberately; see the cpu-sizing docs (TIKA-4809).
* Capability flags are default-deny and split in two. enableUnsecureFeatures
no longer exists -- a config still carrying it fails to start with an
"Unrecognized field" error -- and is replaced by allowPipes (gates /pipes
and /async) and allowPerRequestConfig (gates the /config endpoints and the
multipart config part). /status is no longer gated and is enabled simply
by listing it under endpoints. tika-grpc gains
the same allowPerRequestConfig flag plus allowComponentModifications, which
gates runtime Save/Delete of fetchers and pipes iterators (TIKA-4764).
* Endpoints removed: /translate/* (unusable as shipped), /tika/main and
/tika/form/main (Boilerpipe; use /tika/text), and the /tika/form family.
The 3.x /tika/config and /tika/form/config forms are replaced by the
/tika/config* multipart POSTs, which require allowPerRequestConfig
(TIKA-4809).
* Endpoints collapsed: /detect/stream is now /detect, and /language/stream
and /language/string are both /language. Behavior changed with the rename:
/detect now runs in the fork pool, so it can return 429, 503 or 413, and a
failure reading the body is a 500 where 3.x returned 200 with
application/octet-stream as if detection had succeeded; /language caps
input at the first 100,000 characters and uses the default LanguageDetector
on the classpath, where 3.x pinned Optimaize (TIKA-4809).
* Output-format routing on /tika changed. The bare /tika endpoint returns
Markdown (was XHTML); use /tika/xml for XHTML. /tika/text is body-only
again, as in 3.x. /tika/json and /tika/config/json default to the server
default (markdown) rather than hardcoded plain text. The Accept header no
longer selects the output format -- 3.x routed bare /tika among plain
text, HTML and XHTML by Accept (nondeterministically for */*); now the
path names the format. An unrecognized handler name in the path is a 400
listing the valid types, instead of silently falling back to the default
(TIKA-4663, TIKA-4809).
* Per-request configuration headers are removed, and are now silently
ignored if sent: writeLimit, throwOnWriteLimitReached,
maxEmbeddedResources/maxEmbeddedCount, X-Tika-Handler and the meta_*
metadata-injection family. The limits move to parse-context
(output-limits.writeLimit, output-limits.throwOnWriteLimit,
embedded-limits.maxCount); X-Tika-Handler becomes an explicit handler path;
meta_* has no replacement, and with per-request config off by default a
caller can no longer bound the output of a single request. The
X-Tika-OCR* and X-Tika-PDF* families were removed earlier in the 4.x line
(TIKA-4809).
* Caller errors now map to accurate HTTP status codes instead of always
returning 200 or 500. A saturated worker pool returns 429, a
crashed/timed-out/OOM worker returns 503, an unknown or reserved
fetcher/emitter or bad handler returns 400, and an over-limit body returns
413; the 429 and 503 responses carry a Retry-After header. Error bodies are
now JSON ({"status":"TIMEOUT"}, with a message field when one is
available) where 3.x returned plain text such as "Parse failed: TIMEOUT"
(TIKA-4809).
* The raw /tika family's 422 responses carry the extracted content only; the
exception is no longer appended to the body -- use /rmeta for the
structured exception (TIKA-4809).
* /meta now runs through the same pipes-backed parser as the other
extraction endpoints, so it gains their crash isolation and their error
handling: a container exception comes back as 200 with
tk:exception:container-exception instead of 500, and /meta/{field} returns
422 instead of 500 or 400. A request with no Accept header now returns
JSON; 3.x returned CSV, still available via Accept: text/csv. /meta also no
longer returns a language field -- it parses with the ignore handler, so
there is no text to detect from; configure a language-detection metadata
filter and use /rmeta or /tika/json instead (TIKA-4809).
* /async requires an object body {"tuples":[...]} instead of a bare JSON
array, validates fetcher/emitter ids at POST time (400), rejects a batch
larger than the queue's total capacity with 400 instead of throttling it,
and one bad tuple no longer stops the async workers. /pipes returns the
same JSON body as /tika/rmeta/unpack -- {"status":,
"message":...} -- instead of a /pipes-only {"status":"ok"|"process_crash"}
shape, returns 400 with the reason for a malformed request body, and
rejects emit strategies other than EMIT_ALL, whose passed-back data the
/pipes response cannot carry (TIKA-4809).
* Many server config keys were removed or renamed (logLevel, idBase,
digest, returnStackTrace, port ranges, the spawn-child options, ...), and
an unrecognized key now fails startup with an error naming it; see
migrating-tika-server-4x.adoc for the key-by-key migration. One change no
startup error will flag: taskTimeoutMillis is now
parse-context.timeout-limits.totalTaskTimeoutMillis, and its default grew
from 5 minutes to 1 hour (TIKA-4809, TIKA-4813).
* Request bodies are now capped by maxRequestSizeBytes, defaulting to 1 GiB;
larger requests are rejected with 413, including over-limit chunked
uploads, which previously surfaced as an empty 500 (TIKA-4809).
* The 'endpoints' allowlist now also gates SPI-provided resources; a
discovered resource binds only when its root endpoint is enabled
(TIKA-4809).
* Fetcher-based streaming is removed: the InputStreamFactory pattern for
fetching documents via the fetcherName/fetchKey headers is gone, and all
documents now go through the pipes infrastructure. The no-op
-a/--pluginsConfig flag is removed and now fails option parsing, --help
exits 0, and the tika-server-client module is removed (TIKA-4809).
--- tika-pipes and tika-grpc ---
* tika-pipes implementation modules are now pf4j plugins, reorganized by
resource (tika-pipes-solr) vs task (tika-pipes-fetcher-solr). Core classes
moved to tika-pipes-core, and the file-system components moved out of it
into their own tika-pipes-file-system plugin
(TIKA-4334, TIKA-4519, TIKA-4543).
* FetchEmitTuple JSON now names the per-tuple parse context "parse-context"
(was "parseContext") and rejects unknown tuple fields with an error naming
the field (TIKA-4809).
* The pipes config keys staleFetcherTimeoutSeconds and
staleFetcherDelaySeconds have been removed; a config still carrying them
fails startup (TIKA-4809).
* TimeoutLimits: progressTimeoutMillis of 0 combined with a positive
totalTaskTimeoutMillis is now rejected at config load; it would kill
every task immediately (TIKA-4809).
* The http-fetcher now verifies TLS certificates and hostnames by default;
set verifySsl:false to opt out (TIKA-4809).
* SolrJ moves from 8.11.4 to 10.0.0; the Solr fetcher, emitter and pipes
iterator no longer support Solr 8 (TIKA-4789).
* tika-grpc: the generated Java classes moved from package org.apache.tika
to org.apache.tika.pipes.grpc.proto, so every generated type moves and Java
gRPC clients must update their imports. This is a source break only: the
proto package ("tika") and the service name ("Tika") are unchanged, so the
wire protocol is identical and clients in other languages are unaffected
(TIKA-4808).
--- tika-app and tika-eval-app ---
* tika-app's batch mode is gone. The -bc/batch directory-to-directory command
line (backed by the removed tika-batch module) has no successor flag; use
-a/--async, which runs the same work through tika-pipes
(TIKA-4333, TIKA-4340).
* tika-core/tika-app: NetworkParser and tika-app's -c/--client=
network-client mode were removed -- they dispatched raw sockets to an
arbitrary user-supplied host with no auth or TLS. Use tika-server
instead (TIKA-4808).
* tika-eval-app's command line changed: the FileProfile sub-command is
removed, the -bc batch-config option is gone, and extract directories are
now named with -e/--extracts (Profile) and -a/--extractsA + -b/--extractsB
(Compare); -i/--inputDir, -d/--db, -c/--config, -n/--numWorkers and
-m/--maxExtractLength replace the 3.x spellings
(TIKA-4342, TIKA-4450, TIKA-4452, TIKA-4507).
* tika-app: the inline short forms -eX (output encoding) and -pX (document
password) were removed from standard mode; use --encoding=X and
--password=X. Prefix-matching them silently swallowed single-dash long
names, so every single-dash long name is now rejected with a message
naming the two-dash form (TIKA-4808).
--- Parser, detector and output behavior ---
* The tika-langdetect-tika module is removed (TikaLanguageDetector,
LanguageIdentifier, LanguageProfile, LanguageProfilerBuilder,
ProfilingWriter). tika-app, tika-server and tika-eval now bundle the new
CharSoup detector (tika-langdetect-charsoup) instead of
tika-langdetect-optimaize, so the language reported by default changes
(TIKA-4662).
* PDF: extractIncrementalUpdateInfo now defaults to true (was false), so
every PDF parse emits pdf:incremental-update-count and related keys
without configuration. parseIncrementalUpdates remains false
(TIKA-4354, TIKA-4358).
* Audio cover art is now extracted as embedded documents from MP3 (ID3v2
APIC/PIC), MP4 (covr), Vorbis and FLAC. Embedded-document counts and
/rmeta list lengths for audio files change (TIKA-4801).
* The DOM-based OOXML extractors are removed (XWPFWordExtractorDecorator,
XSLFPowerPointExtractorDecorator, POIXMLTextExtractorDecorator,
XPSTextExtractor) and with them the OfficeParserConfig keys
useSAXDocxExtractor and useSAXPptxExtractor. The SAX extractors are the
only implementation (TIKA-4692, TIKA-4708).
* The MuPDF renderer is removed (org.apache.tika.renderer.pdf.mutool); PDF
page rendering for OCR now uses the PDFBox renderer or the new
PopplerRenderer (TIKA-4664).
* The legacy ExternalParser is removed; external parsers now require explicit
JSON configuration. CompositeExternalParser and ExternalParsersFactory,
which loaded tika-external-parsers.xml definitions from the classpath
automatically, are gone (TIKA-4707).
* Headers are no longer injected into the body/content of MSG files
(TIKA-4345). Please open a ticket if you need this behavior across email
formats.
--- Removed modules and classes ---
* Removed modules with no direct replacement: tika-batch (TIKA-4333),
tika-dl (TIKA-4499), the advanced media module (TIKA-4500), tika-fuzzing
(TIKA-4506), tika-age-recogniser (TIKA-4343), tika-server-eval
(TIKA-4555), the dotnet bindings (TIKA-4332) and snaps deployment
(TIKA-4502).
* Removed parsers and classes with no direct replacement:
SentimentAnalysisParser (TIKA-4574); ObjectRecognitionParser and the
Tensorflow recognisers/captioners; PooledTimeSeriesParser;
org.apache.tika.parser.pdf.AccessChecker (replaced by
PDFParserConfig.AccessCheckMode); and the jempbox-based JempboxExtractor /
XMPMetadataExtractor / pdf.xmpschemas.* classes superseded by the unified
XMP extractor (TIKA-4775).
* Smaller removals: ExceptionUtils.trimMessage moved from tika-core to
tika-eval-core; HttpClientFactory's inert redirect-host allowlist
accessors (get/setAllowedHostsForRedirect) are gone (TIKA-4809); and
org.apache.tika.utils.{RereadableInputStream,AnnotationUtils},
org.apache.tika.io.{IOUtils,InputStreamFactory},
org.apache.tika.sax.DIFContentHandler,
org.apache.tika.parser.{AutoDetectParserFactory,ParserFactory} and
org.apache.tika.parser.internal.Activator are removed.
NEW FEATURES
* tika-pipes gains three parse modes: NO_PARSE (detect only, no parse),
CONTENT_ONLY (emitters write raw content, no metadata envelope) and
UNPACK (write embedded bytes out). All three are wired through tika-app
and tika-server (TIKA-4631, TIKA-4637, TIKA-4656).
* New tika-pipes plugins: an Elasticsearch emitter (TIKA-4672), an
Atlassian JWT fetcher (TIKA-4604), Google Drive, Microsoft Graph, Azure
Blob, JSON and HTTP plugins, and an Apache Ignite ConfigStore for
runtime fetcher/emitter configuration (TIKA-4583, TIKA-4587, TIKA-4598).
* New inference and OCR modules: tika-inference and tika-vlm add
vision-language-model parsers (Claude, Gemini, OpenAI) that emit
vlm:prompt-tokens / vlm:completion-tokens, and
tika-parser-tess4j-module adds in-process Tesseract OCR
(TIKA-4665, TIKA-4666, TIKA-4667, TIKA-4690).
* Unified XMP extraction across containers; adds HEIF/HEIC and WebP XMP,
including Samsung/Google Motion Photo (TIKA-4775).
* tika-app and tika-server can load extra jars (additional EncodingDetectors,
Parsers, etc.) from the directory named by the -Dtika.extras.dir system
property, without repackaging the application. Off by default; the
directory is a trusted code location whose contents run with full process
privileges. The jars are forwarded onto forked pipes/server workers too,
so they are available where parsing happens (TIKA-4755).
* A Markdown parser with structured, lossless XHTML output, complementing
the Markdown content handler (TIKA-4770).
* Content-based detection of ASN.1/DER crypto containers, enabled by
default: the PKCS#7/CMS, PKCS#12 and RFC 5544 timestamped-data families
gained magic where 3.x had globs only, and Pkcs7Parser refines the
smime-type on the output Content-Type at parse time. Files that detected
as application/octet-stream in 3.x may now detect as a crypto type; see
configuration/detectors.adoc for the opt-in detect-time refinement
(TIKA-1997, TIKA-2856).
* New detection: Android binary XML (application/vnd.android.axml,
TIKA-4747) and Frictionless Data packages (TIKA-4643); improved mp3/aac
(TIKA-4612) and grib (TIKA-4655) detection.
OTHER CHANGES
* Release artifacts are now channel-specific. Maven Central gets slim
per-module jars (plus pom, sources and javadoc); the Apache dist area
gets runnable zip distributions (tika-app, tika-server-standard,
tika-eval-app) and drop-in pf4j plugin zips; Docker Hub gets ready-to-run
images. Fat/shaded artifacts no longer go to Maven Central (TIKA-4733).
* The charset, junk-text and language detection stack was rewritten:
language-aware charset detection, a universal junk detector, wider
Unicode handling, the new CharSoup language detector, and more efficient
common-token lookups via bloom filters (TIKA-4662, TIKA-4671, TIKA-4675,
TIKA-4691, TIKA-4719, TIKA-4731, TIKA-4745, TIKA-4754, TIKA-4810).
* Pipes now carries small documents to the forked worker inside the request
instead of writing them to disk first. Content at or below the new
pipes.maxInlineBytes (default 10 MB) rides in the request and is served in
the worker by the reserved __bytes fetcher, touching no disk; larger
content is written out once as before, and a stream already backed by a
file keeps its file. Set maxInlineBytes to 0 to spool every non-empty body
(TIKA-4808).
* PipesClient/PipesServer IPC now enforces a configurable payload limit
(pipes.maxIpcPayloadBytes, default 100 MB) in both directions. Results
that exceed the limit return PAYLOAD_LIMIT_EXCEEDED instead of causing
heap exhaustion; crash messages are also size-capped (TIKA-4793).
* tika-server requests now carry only their own parse-context entries to the
forked worker, which supplies the config defaults itself. Previously the
server sent its config's parse-context with every request, so the worker
treated the operator's timeout-limits as caller input and clamped them at
pipes.maxTotalTaskTimeoutMillis (TIKA-4808).
* tika-grpc now routes fetchAndParse through the PipesParser client pool
instead of one shared single-threaded PipesClient: concurrent calls no
longer crash the worker, pipes.numClients and (for the first time)
pipes.useSharedServer take effect, and pool saturation surfaces in-band as
CLIENT_UNAVAILABLE_WITHIN_MS. An interrupted call recycles its worker, so a
pooled client cannot go back to the queue dirty (TIKA-4815).
fetchAndParseServerSideStreaming now completes the call after delivering
its reply, instead of leaving the client waiting forever (TIKA-4804).
* New audio/video metadata: audio:bitrate, audio:is-variable-bitrate,
audio:has-drm, audio:channels, video:frame-rate, video:bitrate and MP4
sample size; ID3 TCOP and Vorbis COPYRIGHT map to xmpDM:copyright, EXIF
GPS altitude maps to geo:alt, and the presentation start of delayed
QuickTime timed-metadata tracks is exposed (TIKA-4777, TIKA-4779,
TIKA-4780, TIKA-4781, TIKA-4800, TIKA-4802).
* Parsing and robustness fixes across formats: CHM (TIKA-4783), ID3 UTF-16
(TIKA-4784), MPEG2/2.5 Layer III frame sizing (TIKA-4791), .doc empty
comments (TIKA-4718), OOXML hyphenation and field-code hyperlinks
(TIKA-4646, TIKA-4683), RTF attachments in HTML decapsulation
(TIKA-4710), image extraction (TIKA-4736), embedded-file extension
calculation (TIKA-4808), and general media-file robustness (TIKA-4812).
MAPI properties no longer overwrite better-fitting Dublin Core terms
(TIKA-4806). Embedded-file naming was streamlined (TIKA-4689).
* PDFs whose %PDF- header is preceded by a print-composition job ticket are
no longer detected as text/x-matlab: up to 50 %% comment or blank lines may
now precede it, extending the TIKA-3328 rule past its 512-byte reach
(TIKA-4782).
* MagicDetector now compiles its regular expression once, in the
constructor, instead of recompiling it on every match (TIKA-4796).
* tika-eval-core is no longer published as a fat jar (TIKA-4414) and
tika-grpc no longer shades gRPC (TIKA-4709).
* Fix concurrency bug in TikaToXMP (TIKA-4393).
* Dependency upgrades throughout the 4.0.0 line, including Jetty 11 ->
12.1.12, CXF 4.0 -> 4.2.3 and SolrJ 8.11.4 -> 10.0.0, plus routine
library updates (TIKA-4327).
Release 4.0.0-beta-1 - 6/29/2026
Prerelease. Its changes are folded into the 4.0.0 section above.
Release 4.0.0-alpha-1 - 5/4/2026
Prerelease. Its changes are folded into the 4.0.0 section above.
Release 3.3.0 - 3/18/2026
* Switch to poi-ooxml-full (TIKA-4563).
* Users need to add "allowAbsolutePaths=true" for the FileSystemFetcher to fetch
an absolute path (TIKA-4649).
* Add a markdown option for content handlers (TIKA-4563).
* Improve zip parsing (TIKA-4650).
* Add detection of compressed bmp (TIKA-4511).
* Allow per file timeouts in tika-pipes (TIKA-4497).
* Add matroska detector (TIKA-1180).
* Allow multiple values for many Dublin Core keys (TIKA-4466).
* Extract macros by default in tika-app's commandline and gui (TIKA-4472).
* Improve extraction of Javascript from PDFs (TIKA-4465).
Release 3.2.3 - 9/11/2025
* Allow backwards compatibility with versions of commons-compress before 1.28.0 (TIKA-4469).
* Fix XFA parsing within PDFs when woodstox is on the classpath as in tika-server (TIKA-4482).
* Dependency updates.
Release 3.2.2 - 8/6/2025
* Fix for CVE-2025-54988.
* Improve detection of encrypted ODT files (TIKA-4459).
* Dependency updates (TIKA-4455).
Release 3.2.1 - 6/26/2025
* Fix POIFSContainerDetector regression when wrapping an InputStream in
a TikaInputStream (TIKA-4441).
* Important bug fix for zip-based detection on a non-TikaInputStream (TIKA-4424).
* Improve text extraction from EMF (TIKA-4432).
* Dependency updates (TIKA-4421).
Release 3.2.0 - 05/21/2025
* Detect inline images in MSG files (TIKA-4391).
* Improve extraction of metadata in MSG files (TIKA-4381).
* Fix concurrency bug in TikaToXMP (TIKA-4393).
* Fix potential GDAL deadlock (TIKA-4385).
* Improve extraction of properties from msg files (TIKA-4381).
* Include internal attachment path in tika-eval reports (TIKA-4374).
* Upgrade jsoup to 1.20.1 with workaround for change in self-closing tag behavior (TIKA-4419).
* Upgrade dependencies (TIKA-4379).
Release 3.1.0 - 01/28/25
* Allow users to turn off the injection of some headers into the content stream of MSG
files (TIKA-4345).
* Add a wrapper for Google's magika detector (TIKA-4344).
* Add support for MachO via Alexey Pelykh (TIKA-4309).
* Add logic to inject spaces in XPS files based on font widths via Ruairidh Williamson (TIKA-4315).
* Mime type "application/json" is now a sub class of "text/javascript" not "application/javascript" (TIKA-4336)
* Remove tagsoup from the project entirely. Note that
some of the tags produced by the SourceCodeParser are slightly different (TIKA-4338)
Release 3.0.0 - 10/15/2024
* Fix regression in TextAndCSVParser (TIKA-4278).
Release 3.0.0-BETA2 - 07/09/2024
BREAKING CHANGES
* Updated PST parser to use standard Message metadata keys and improved
handling of embedded files (TIKA-4248).
* Convenience methods for XML readers were moved from ParseContext to
XMLReaderUtils (TIKA-4259).
Other Changes
* Add GRPC server (TIKA-4181).
* Improved configurability in tika-pipes (TIKA-4243).
* Add optional PST parser based on libpst/readpst (TIKA-4250).
Release 3.0.0-BETA - 12/01/2023
BREAKING CHANGES
* Require Java 11 (TIKA-4128).
* The boilerpipe handler has been moved to the tika-handler-boiler-pipe
package (TIKA-4138).
* We've migrated HTML parsing to the JSoup parser instead of TagSoup. If
you have a custom configuration on the HTMLParser, you'll need to change
that to o.a.t.p.html.JSoupParser (TIKA-1599).
* Removed xerces2 as a dependency (TIKA-4135).
* tika-core now has a scope of "provided" in most non-app modules (TIKA-4191).
* Tika will look for "custom-mimetypes.xml" directly on the classpath, NOT
under "/org/apache/tika/mime/". (TIKA-4147).
* Return media type "text/javascript" instead of "application/javascript"
to follow RFC-9239. (TIKA-4119).
Other Changes/Updates
* Improve detection of sqlite3-based file formats (TIKA-4187).
* Upgrade PDFBox to 3.0.1 (TIKA-3347)
* Deprecated AbstractParser for removal in 4.x (TIKA-4132).
* Fix bug in DateUtils that stripped timezone information from
incoming Calendar objects (TIKA-4126).
* The InputStreamDigester now calculates stream length (TIKA-4016).
Release 2.9.0 - 8/23/2023
* With user configuration, the PDFParser can now throw an EncryptedDocumentException
for Microsoft IRM PDF containers with encrypted payloads. Separately,
the PDFParser now throws an EncryptedDocumentException instead of an IOException
if the security handler cannot be found (TIKA-4082).
* Fix bug that led to duplicate extraction of macros from some OLE2 containers (TIKA-4116).
* Parse iframe's srcdoc as an embedded file (TIKA-3109).
* Add detection of warc.gz as a specialization of gz and parse as if a standard WARC (TIKA-4048).
* Allow users to modify the attachment limit size in the /unpack resource (TIKA-4039)
* Fixed write limit bug in RecursiveParserWrapper (TIKA-4055).
* Add mime detection for many files with thanks to Gregory Lepore (TIKA-3992).
* Fixed iWork 13 keynote detection on files with wrong extension (TIKA-4111).
Release 2.8.0 - 5/11/2023
* Enable counting and/or parsing of incremental updates in PDFs. This
is an experimental feature and may change in later releases (TIKA-4017).
* Fixed bug that prevented the the loading of CompositeExternalParser in tika-app and
tika-server-standard. This parser will call exiftool and ffmpeg if those are installed, as was
the behavior in Tika 1.x. Exclude org.apache.tika.parser.external.CompositeExternalParser
if you do not want this behavior (TIKA-4022).
* Removed the shading of tika-parsers-standard-module (TIKA-4038).
* Enable optional extraction of file system metadata in FileSystemFetcher (TIKA-4035).
* Allow pretty printing in FileSystemEmitter (TIKA-4034).
* Add detection for and a new mime type for older postscript-based
Adobe Illustrator "application/illustrator+ps" files (TIKA-3971).
* Add magic detection for canon raw file types: crw, cr2 and cr3 (TIKA-3991).
* Add detection for ONIX message files (TIKA-4011).
* Add detection and a parser for ActiveMime files (TIKA-3987).
* Add extraction of rendition layout value and version from Epub (TIKA-4013).
* Improve embedded file extraction from PDFs (TIKA-4012).
* Improve metadata extraction from WARCs (TIKA-4018).
* Update to PDFBox 2.0.28 (TIKA-4016).
* Users may now avoid the ZeroByteFileException via a
setting on the AutoDetectParserConfig (TIKA-3976).
* Fix bug in closing elements in the presence of elements
in RTF files (TIKA-3972).
* Improve extraction of embedded file names in .docx (TIKA-3968).
* Normalize author, title, subject and description to their Dublin Core
properties in the HTMLParser (TIKA-3963).
Release 2.7.0 - 1/31/2023
* Add SVG detection for svg files that lack the xml header (TIKA-3308).
* Migrate to a live fork of Universal Charset Detector (TIKA-3213).
* Improve handling of text-based attachments inside .eml files (TIKA-3959).
* Add tika-parser-nlp-package to release artifacts (TIKA-3958).
* Remove need for element in classes that extend ConfigBase (TIKA-3946).
* Add X-TIKA:embedded_id_path to ensure unique embedded file paths (TIKA-3942).
* Fix bug that prevented digests when the fallback/EmptyParser
was called (TIKA-3939).
* Remove log4j 1.2.x (and slf4j-log4j12 which now redirects to slf4j-reload4j) from
all modules (TIKA-3935).
* Upgrade mime4j to 0.8.9 (TIKA-3950).
* Refactor date parsing for emails (TIKA-3957)
* Upgrade to Bouncy Castle 1.71 and jdk18on jars (TIKA-3933).
* Add a JDBCPipesReporter (TIKA-3931).
* Add multivalued field strategy option in jdbc-emitter (TIKA-3930).
Default is now 'concatenate' with ', ' as the delimiter.
* Downgrade logging in PipesClient for each parse from info to debug.
Release 2.6.0 - 11/3/2022
* Add optional Siegfried detector (TIKA-3901).
* Move OverrideDetector's functionality to the CompositeDetector (TIKA-3904).
* The FileCommandDetector has been refactored to have the same
behavior as the Siegfried detector; see setUseMime in the javadoc (TIKA-3902).
* Fix bug in OpenSearch emitter that prevented upserts on
documents with embedded files (TIKA-3882).
* Extract PDF actions and triggers into the file's metadata (TIKA-3887).
* Add a tika-async-cli module (TIKA-3885).
* Fetch keys sent via headers to tika server are now URL decoded (TIKA-3864).
Release 2.5.0 - 09/30/2022
* Improved extraction of PDF subset info for PDF/UA, PDF/VT, and PDF/X.
NOTE: we no longer append PDF/A information, e.g. 'version="A-1b"'
to the 'dc:format'. Users must now get that information from the
'pdfa:PDFVersion' key or from 'pdfaid:conformance'
and 'pdfaid:part' (TIKA-3844).
* Avoid infinite loop in bookmark extraction from PDFs (TIKA-3832).
* Upgraded to slf4j 2.0.1 (TIKA-3842).
* Added upsert option for the OpenSearch emitter (TIKA-3855).
* Extract PDF signature information at the document level
into the metadata (TIKA-3852).
* Enable configuration of digests via AutoDetectParserConfig (TIKA-3853).
* Use commons-io byte array streams via PJ Fanning (TIKA-3843).
* Upgrade to PDFBox 2.0.27 (TIKA-3866).
* Upgrade to JempBox 1.8.17 (TIKA-3856).
* Add extraction of ODF version from ODF files (TIKA-3840).
* tika-parser-html-commons (BoilerPipeHandler) is no longer a
a dependency of tika-parser-html-module. tika-app and tika-server-standard
have added a dependency on tika-parser-html-commons. However,
users who are managing custom dependencies and who want the BoilerPipeHandler
will have to now include the tika-parser-html-commons dependency
(TIKA-1484).
* Add unrar as an optional parser (TIKA-3800).
* Refactor FuzzingCLI to use PipesParser (TIKA-3799).
* ServiceLoader's loadServiceProviders() now guarantees
unique classes (TIKA-3797).
* Fix bug that prevented setting of includeHeadersAndFooters
for xls, xlsx, doc and docx via tika-config (TIKA-3796).
* Fix bug that prevented specification of rendered image type
via http header in the PDFParser (TIKA-3794).
* Fix bug causing some Exif dates to be decoded wrongly on
timezones different than UTC (TIKA-3815).
* Numerous dependency upgrades (TIKA-3795).
* Add ALPHA-level initial releases of JDBCEmitter,
FileSystemStatusReporter and OpenSearchPipesReporter.
These may have breaking changes in subsequent releases.
Release 2.4.1 - 06/14/2022
* Implement bulk upload in the OpenSearch emitter (TIKA-3791).
* Implement tika-server client via pipes mode (TIKA-3790).
* Custom embedded parsers and EmbeddedDocumentHandlers
can now add metadata to the container file's
metadata (TIKA-3789).
* Record embedded file exceptions in the container
file's metadata (TIKA-3788).
* Allow continuation of parsing after write limit has
been reached (TIKA-3787).
* Allow pass-through of 'Content-Length' header to metadata
in TikaResource (TIKA-3786).
* Add embedded depth to profiles tables in tika-eval (TIKA-3775).
* Add stop() method to TikaServerCli so that it can be run
with Apache Commons Daemon (TIKA-1570).
* Fixed bug in ordering of Parsers during service loading (TIKA-3750).
* Users can expand system properties from the forking
process into forked tika-server processes (TIKA-3748).
* Fix a few files being wrongly detected as EML (TIKA-3771).
* Fix ignoreCharsets param of Icu4jEncodingDetector (TIKA-3774).
Release 2.4.0 - 04/23/2022
* NOTE: To save on resources, we no longer include the
deeplearning4j dependencies in the tika-dl jar. The dependencies for the
tika-dl package must be provided by users. See:
https://github.com/apache/tika/blob/main/tika-parsers/tika-parsers-ml/tika-dl/pom.xml
for the dependencies that must be provided at run-time (TIKA-3676).
* NOTE: Added prefix "dwg-custom:" to DWG custom metadata properties (TIKA-3731).
* Add initial, BETA-grade TLS encryption option for tika-server;
configuration may change in future releases (TIKA-3719).
* Allow specification of fetcherName and fetchKey via query parameters
in request URI in tika-server (TIKA-3714).
* Add basic parsers for WARC and WACZ in tika-parsers-standard (TIKA-3697).
* Add MetadataWriteFilter capability to improve memory profile in
Metadata objects (TIKA-3695).
* Allow configurability of the ContentHandlerDecorator used
by the AutoDetectParser (TIKA-3723).
* Allow configurability of the EmbeddedDocumentExtractor used
by the AutoDetectParser (TIKA-3711).
* Add detection for Frictionless Data packages and WACZ (TIKA-3696).
* Add detection for DGN files with gratitude and credit
to Steven Frew's tika-dgn-detector (TIKA-3721).
* Add parser for metadata from DGN 8 files via Dan Coldrick (TIKA-3721).
* Add a fetcher and emitter for Azure blob storage (TIKA-3707).
* Add detection for files encrypted by Microsoft's Rights Management Service
(TIKA-3666).
* Fixed regression in 2.3.0 that led to more embedded filenames
than appropriate being written to the content (TIKA-3711).
* tika-server now clones forking process' environment variables
into forked process (TIKA-3715).
* Add an optional /eval endpoint for tika-eval profile or compare
capabilities in tika-server (TIKA-3689).
* Add a Parsed-By-Full-Set metadata item to record all parsers that processed
a file (TIKA-3716).
* Add metadata filters for Optimaize and OpenNLP language detectors (TIKA-3717).
* Upgrade to PDFBox 2.0.26 (TIKA-3726).
* Upgrade deeplearning4j to 1.0.0-M2 (TIKA-3458 and PR#527).
* Various dependency upgrades, including POI, dl4j, gson, jackson,
twelvemonkeys, log4j2 and others (TIKA-3675 and many PRs from dependabot).
* Switch cipher from ECB to GCM in HttpClientFactory (TIKA-3724).
Release 2.3.0 - 02/02/2022
* Upgrade to Apache POI 5.2.0. This is the first upgrade to POI
5.x and represents a major refactoring. Users may experience
significantly more logging (TIKA-3164).
* Upgrade to log4j2 2.17.1 (TIKA-3638).
* Improve consistency in reporting package-entry divs across
all parsers for embedded files (TIKA-3644). This leads
to some more text (embedded file names) in files with
many embedded attachments.
* Improve configuration of maps as params for parsers in
TikaConfig (TIKA-3645).
* Improve identification of iWorks 13 files and add parsing
for thumbnails, some metadata and attachments (TIKA-3634).
Skip handling of .iwa files, which are not yet supported.
* Limit the default in-memory processing (maxMainMemoryBytes) in
the PDFParser to 512MB as in the 1.x branch (TIKA-3642).
* Added IDML Parser from 1.x series to 2.x series (TIKA-3188).
* Extract annotation types and subtypes for PDFs into metadata (TIKA-3653).
* Add metadata value for PDFs that contain 3D annotations (TIKA-3653).
* Add parser for Translation Memory eXchange (TMX) files (TIKA-3660).
* Add Bill of Materials (Maven BOM) for centralized module version management (TIKA-3367).
Release 2.2.1 - 12/19/2021
* Fix multithreading bug for ooxml files (TIKA-3627).
* Upgrade log4j to 2.17.0 (TIKA-3625).
* Upgrade to PDFBox 2.0.25 (TIKA-3622)
* Fix bug that prevented metadata keys in the UnpackerResource
in tika-server (TIKA-3624).
* Upgrade log4j to 2.16.0 (TIKA-3623)
Release 2.2.0 - 12/13/2021
* Add support for OneNote files downloaded from O365 (TIKA-3446).
* Fix logic bug in PipesServer that prevented concatenation of
content from attachments (TIKA-3609).
* Improve extraction of embedded files from MSOffice files created
by non-Microsoft tools (TIKA-3526).
* Added back ability to ignore load errors in TikaConfig (TIKA-3575).
* Make SecureContentHandler and other parameters configurable in
AutoDetectParser programmatically and via tika-config.xml (TIKA-3594).
* Fix default logging in tika-app in batch mode (TIKA-3589).
* Fix bug that prevented specifying a config with the long
--config= option in tika-app in batch mode (TIKA-3589).
* Fix thread starvation after numerous restarts in
PipesClient (TIKA-3588).
* Fix race condition when starting multiple forked
servers on multiple ports (TIKA-3586).
* Add timeout per task to be configured via headers
for tika-server's legacy endpoints /tika and /rmeta.
Note that this timeout greater than taskTimeoutMillis (TIKA-3582).
* Add metadata item for whether or not a PDF has a collection/
is a Portfolio PDF (TIKA-3579).
* Add detection of ESRI Layer files (TIKA-3570).
* Add detection of JPEG XL, MARC, ICC profiles, NES-ROM file types
(TIKA-3562 and TIKA-3563)
* Remove duplicate "subject" metadata keys that were intended
for backwards compatibility within 1.x only (TIKA-3564).
* Fix Open Office mime types to be subclasses of application/zip
and no longer require OPCPackageDetector-last ordering of zip
detectors (TIKA-3556).
* Improve robustness and features of the httpfetcher (TIKA-3543)
* Add optional fetch ranges to FetchEmitTuple to allow range fetching from,
e.g. http or s3 (TIKA-3542).
* Exclude dependencies on jsoup and ehcache in ucar grib/cdm (TIKA-3003).
Release 2.1.0 - 08/18/2021
MAJOR CHANGES in 2.1.0:
* Improved packaging for tika-parsers-extended. Use the tika-parser-scientific-package and
tika-parser-sqlite3-package artifacts if you want fat jars with dependencies. (TIKA-3510)
* Tika app writes UTF-8 when an encoding is not specified; the legacy behavior
was UTF-8 on Mac OS, but System default on other OSs (TIKA-3515).
* Change the default rendering strategy for PDFs from NO_TEXT to ALL (TIKA-3520).
Other changes:
* Fixed bug that pointed to the wrong tessdata directory if the user specified
a tesseract path but not also a tessdata path (TIKA-3518).
* Fixed bug in Icu4j's encoding detector where it would return non-standard
names for charsets, e.g. IBM424_rtl is now returned as IBM424 (TIKA-3516).
* Add a simple UrlFetcher in tika-core as a basic alternative
to tika-fetcher-http (TIKA-3527).
* Add tika-pipes support for Google Cloud Storage (TIKA-3524).
* Fix markup ordering errors in xhtml output for ODT files (TIKA-2242).
* Fix serialization of embedded docs in OpenSearch emitter
and fix embedded documents not being indexed in some use
cases in the Solr emitter (TIKA-3490).
* Add pipesClientId system property to PipesServer so that each
forked process can log to its own logger (TIKA-3480).
* Add DateNormalizingMetadataFilter let users ensure that all dates
emitted to Solr/OpenSearch are in UTC. Users can configure which
timezone they'd like to use in cases where the file format does
not store a timezone (TIKA-3496).
* Breaking change in the Solr and OpenSearch emitters. To achieve
the SKIP or CONCATENATE attachment strategy, modify the
parseMode in the pipesiterators or in the FetchEmitTuple (TIKA-3494).
Release 2.0.0 - 07/07/2021
* Cleanup of fetcher integration with tika-server.
* Update dependencies.
Release 2.0.0-BETA - 05/19/2021
* Refactor pipes module for resilience
* Add transcribe capability (TIKA-94).
Release 2.0.0-ALPHA - 01/13/2021
BREAKING CHANGES in 2.0.0
* General
* OCR is now triggered automatically for PDFs if tesseract
is on the user's path see (https://cwiki.apache.org/confluence/display/TIKA/TikaOCR#TikaOCR-disable-ocr)
for how to disable OCR.
* We upgraded from log4j to log4j2 in tika-app, tika-server and anywhere else
we used to use log4j.
* By default, when rendering a page for OCR, the PDFParser does not render glyphs/text.
* Removed deprecated Metadata keys/properties (TIKA-1974).
* Removed deprecated PDFPreflightParser (TIKA-3437).
* Removed dangerous calls to read an inputstream or convert to bytes
without specifying a charset
* Parsers can be configured via tika-config.xml on instantiation.
We have moved away from configuration via .properties files because
of confusion among users. This affects the PDFParser, TesseractOCRParser
and the StringsParser.
* Changed namespaces of translator implementations (o.a.t.language.translate.impl) to avoid
split-package with tika-core
* tika-parsers
* The parser modules have been broken into three main modules:
tika-parsers-standard, tika-parsers-extended and tika-parsers-ml.
Users may now need to add tika-parsers-extended's
tika-parser-scientific-module or tika-parser-sqlite3-module to tika-app and
tika-server to include parsers that used to be included by default
(for example: envi, gdal, grib, isatab, netcdf, sqlite3).
* PDFParser -- a) see above on OCR. b) This parser no longer warns if the jpeg2000
dependency is not included. Tika now relies on PDFBox to log an error if a jpeg2000
image should be processed but can't because the required external dependency is
not available. See https://pdfbox.apache.org/2.0/dependencies.html#jai-image-io
for the non-ASF-2.0-compatible jpeg2000 library.
* CompressorParser -- users must add the com.github.luben:zstd-jni dependency to
the classpath to process zstd files. This is an optional library that is no longer bundled
in tika-parsers-standard-package because it contains native libs.
* ChmParser was moved to org.apache.tika.parser.microsoft.chm
* RTFParser was moved to org.apache.tika.parser.microsoft.rtf
* We are now using non-shaded versions of xmpcore with namespaces com.adobe.internal.*
vs com.adobe.*.
* tika-app
* See above on default inclusion of only tika-parsers-standard.
* tika-server
* tika-server now by default forks a process to isolate the parsing
in the forked process (this was called the -spawnChild option
in tika-1.x). Clients must now expect that tika-server
will restart on OOM, timeouts, crashes or after parsing a
large number of files. When this happens tika-server will restand and not
receive connections for brief periods. The less robust, legacy behavior
of not forking a process is available with "-noFork"=
* Most of tika-server's legacy configuration via the commandline has been moved
into configuration via a tika-config.xml file.
* tika-server's "enableFileUrl" has been removed in favor of a FileSystemFetcher.
* tika-server's /metadata endpoint requires tika-server-standard to write XMP/rdf output.
This output is not available in tika-server-core.
* In tika-server, for those parsers that can be configured per parse via a config object
passed in through the ParseContext, the config object will only update those fields
that the user has modified. The config object will no longer
fully reset all settings to the default settings per parse.
This has a more intuitive "update the base/configured settings" with
what has been changed in the config object.
* tika-eval
* tika-eval's default profile and comparison reports no longer include tag reports.
Users can get the report configs that include tags (*-tags.xml):
https://github.com/apache/tika/tree/main/tika-eval/tika-eval-app/src/main/resources
Release 1.27 - 06/30/2021
* Migrate MP4 parsing to Drew Noakes' metadata-extractor (TIKA-3459).
To revert to legacy parser turn off NoakesMP4Parser and turn on MP4Parser
via tika-config.xml.
* Prevent rare infinite loop in tika-server's -spawnChild mode
when restart fails because of failure to bind to the port (TIKA-3441).
* Improve likelihood that tesseract will not be orphaned on
jvm restart in tika-server (TIKA-3441).
* Deprecate experimental PDFPreflightParser (TIKA-3437).
* Apply encoding detection to zip entry names via Ryan421 (TIKA-3374).
* Add json output for /tika endpoint in tika-server (TIKA-3352).
* Tika's PDFParser should use the underlying file if one is passed in
via a TikaInputStream (TIKA-3350)
Release 1.26 - 03/24/2021
* Fix thread safety bug in OpenOffice parser (TIKA-3334).
* The "writeLimit" header now pertains to the combined characters
written per container document (and embedded documents) in the /rmeta
endpoint in tika-server (TIKA-3325); it no longer functions only
per container or embedded document.
* Extract more embedded files in PDFs by recursively processing the
embedded file tree (TIKA-3332).
* Allow for case insensitive headers for configuration of the PDFParser
and the TesseractOCRParser in tika-server via Subhajit Das (TIKA-3320).
* Improve detection and parsing of XPS files (TIKA-3316).
* General dependency upgrades (TIKA-3244).
* Great optimization in ForkParser (TIKA-3237).
* Fix parsing of emails attached to other emails in PST files (TIKA-3004).
* MP3 parser should output the xmpDM:duration metadata as seconds not
milliseconds, consistent with the other Audio and Video parsers (TIKA-3318).
* MP4 parser check if any of the Compatible Brands match when identifying
the subtype (TIKA-3310).
Release 1.25 - 11/25/2020
* Fix inconsistent license in xmpcore (TIKA-3204).
* General upgrades including some dependencies with
recently found security vulnerabilities (TIKA-3119).
* Add detection and a parser for flat ODF files (TIKA-3159).
* Add extraction of macros from ODF files (TIKA-3161).
* Add mime detection for hprof and hprof text files (TIKA-3144).
* Add TextSignature and TextProfileSignature to tika-eval (TIKA-3145 and TIKA-3146)
* Create a metadata filter to trigger tika-eval stats post parsing (TIKA-3140)
* Add a configurable metadata-filter for the RecursiveParserWrapper (TIKA-3137)
* Parameterize writeLimit and maxEmbeddedResources for RecursiveParserWrapper
in tika-server (TIKA-3133)
* Add status endpoint to tika-server (TIKA-3129).
* Remove whitelist/blacklist terminology (TIKA-3120)
* Add detection for parquet files (TIKA-3115).
* Add detection and parsing for bplist (TIKA-3104).
* Enable metadata value filtering for RecursiveParserWrapper (TIKA-3137)
* Add a basic parser for plist files based on com.googlecode.plist:dd-plist (TIKA-3104).
* Read hyperlinked images from ODT files (TIKA-3156).
* Updated GrobidRESTParser to use new API location (TIKA-3191).
* Add FileProfiler to tika-eval (TIKA-3216).
* Add status endpoint to tika-server (TIKA-3129).
* Improved handling of zip files with STORED entries with
data descriptor (TIKA-3196).
* Add parsers for XLZ, IDML and MIF (TIKA-2976, TIKA-3188 and TIKA-3189).
* Add the beginnings of a format-aware fuzzing module (TIKA-3083).
* Add wrapper for Linux 'file' command for mime detection (TIKA-3215).
* Added ability to skip parsing of embedded files in Tika Server (TIKA-3227).
Release 1.24.1 - 4/17/2020
* Allow gzip compression of input and output streams for tika-server (TIKA-3073).
Release 1.24 - 3/11/2019
* Add scripts to run tika-server as a service via Eric Pugh,
and add these scripts and jar as a new artifact in the release (TIKA-3010).
* Upgrade Drew Noakes' metadata-extractor (TIKA-2952).
* Enable optional extraction of structural tags in PDFs (alpha-grade) (TIKA-3026).
* Tika app's --extract mode now outputs to STDOUT (TIKA-3035).
* Add an optional Preflight parser for PDFs (TIKA-3055).
* Improve detection of some zip-based formats (TIKA-3057).
* Upgrade metadata-extractor to 2.13.0 (TIKA-2952).
* Upgrade to POI 4.1.2 (TIKA-3047).
* Extract XMP from PSD files (TIKA-3050).
* Added XMLProfiler as an optional parser to profile XFA and XMP
in PDFs (TIKA-3045).
* Extract inline images that rely on the DCT filter from PDFs (TIKA-3041).
* Upgrade to PDFBox 2.0.19 (TIKA-3033).
* Fix bug in ASM parser configuration (TIKA-2992).
* Upgrade to java-libpst 0.9.3 (TIKA-2546).
* Fixed XLIFF12Parser failures with ToXMLHandler (TIKA-3014).
Release 1.23 - 12/02/2019
* NOTE: The PDFParser now relies on OCRDPI to render page images when
users configure OCR on rendered page images. This will have the effect
of increasing rendered image size (TIKA-2624).
* NOTE: tika-server no longer returns 415 for file types for which there
is no parser.
* Fix bug in AUTO OCR strategy in the PDFParser (TIKA-3002).
* Fix incorrect height and width metadata extraction from JPEG images (TIKA-2630).
* Upgrade to POI 4.1.1 (TIKA-2851).
* Upgrade to PDFBox 2.0.17 (TIKA-2951).
* Ensure that the PDFParser respects custom configuration of Tesseract
from tika-config.xml via Eric Pugh (TIKA-2970).
* Add parser for XLIFF v1.2 files (TIKA-2975).
* Add mime type detection support for WebAssembly (TIKA-2894),
HEIF / HEIC images (TIKA-2942), Digilite FDF (TIKA-2988);
and xml-root detection for XFDF (TIKA-2990) and XDP (TIKA-2989).
* Add an XLZ Parser (TIKA-2976).
* Fix deadlock with ForkParser when InputStream throws IOException (TIKA-2892).
Release 1.22 - 07/29/2019
* NOTE: tika-server no longer hard-codes the HtmlParser to handle
XML files (TIKA-2910). Users must now configure that behavior
via a tika-config.xml file.
* NOTE: Known regression: PDFBOX-4587 -- PDF passwords with codepoints
between 0xF000 and 0XF0000 will cause an exception.
* Add parser for HWP v5 files via SooMyung Lee (soomyung) and
JinSup Kim (ddoleye) (TIKA-2909).
* Fix order of closing streams to avoid "Failed to close temporary resource"
exception in TesseractOCRParser (TIKA-2908).
* Improve AutoDetectReader performance by caching encoding
detector (TIKA-1568).
* Prevent RTFParser from outputting illegal tag combinations (TIKA-2889).
* Fix RereadableInputStream to release all resources (TIKA-2903).
* Implement custom language identifier in the tika-eval module based on
OpenNLP's language detector; add 18 languages and add common words
lists for all 121 languages (TIKA-2790).
* Fix NPE in MimeTypesReader.releaseParser() via Eamonn Saunders (TIKA-2896).
* Fix RTFParser to extract more content (TIKA-2883).
* Add clientSubmitTime to the metadata extracted from PST files (TIKA-2898).
* Improve StreamingZipContainerDetector for xltx, xltm and
several other file formats (TIKA-2886).
Release 1.21 - 05/14/2019
* Add optional AUTO mode to OCR'ing of PDFs. If tesseract is installed
and on the path, and this option is selected programmatically
or via TikaConfig(), the PDFParser will use heuristics to decide
whether or not to run OCR per page on PDFs. (TIKA-2749)
* The ZipContainerDetector's default behavior was changed to run
streaming detection up to its markLimit. Users can get the
legacy behavior (spool-to-file/rely-on-underlying-file-in-TikaInputStream)
by setting markLimit=-1. The POIFSContainerDetector requires an underlying file;
it will try to spool the file to disk; if the file's length is > markLimit,
it will not attempt detection; set markLimit to -1 for legacy behavior (TIKA-2849).
* Upgrade PDFBox to 2.0.14 (TIKA-2834).
* Add CSV detection and replace TXTParser with TextAndCSVParser;
users can turn off CSV detection by excluding the TextAndCSVParser
and adding back the TXTParser via tika-config (TIKA-2833).
* Add a CSVParser. CSV detection is currently based solely on filename
and/or information conveyed via Metadata (TIKA-2826).
* General upgrades: asm, bouncycastle, commons-codec, commons-lang3, cxf,
guava, h2, httpcomponents, jackcess, junrar, Lucene, mime4j, opennlp, parso,
sqlite-jdbc (provided), zstd-jni (provided) (TIKA-2824)
* Bundle xerces2 with tika-parsers (TIKA-2802).
* Upgrade jaxb to 2.3.2 (TIKA-2819).
* Upgrade jackson to 2.9.8 (TIKA-2717).
* Update tika-eval's common tokens lists (TIKA-2822).
* Handle bad tags in tika-eval more robustly (TIKA-2810).
* Add reports for tags in tika-eval (TIKA-2809).
* Extract text from SDT element within textboxes in .docx files (TIKA-2807).
* Try to handle truncated OOXML files more robustly (TIKA-2765).
Release 1.20 - 12/17/2018
* Upgrade to POI 4.0.1 (TIKA-2751).
* Integrate/parameterize new angles handling in
PDFBox (TIKA-2779).
* Upgrade to PDFBox 2.0.13 (TIKA-2788).
* Prevent content within and elements
to be written in the ToTextContentHandler (TIKA-2550).
* Switch child to parent communication to a shared memory-mapped
file in tika-server's -spawnChild mode.
* Fix bug in tika-server when run in legacy mode (not -spawnChild)
that caused it to return 503 on documents submitted after
it hit an OutOfMemoryError (TIKA-2776).
* Upgrade jaxb-runtime and javax.activation (TIKA-2778).
* tika-app in batch mode now requires an interrupt or
kill signal to the parent process to stop the parent
and the child processes (TIKA-2780).
* Bulk upgrade of dependencies (TIKA-2775).
* Improve language id efficiency in tika-eval (TIKA-2777).
* Upgrade sqlite "provided" dependency to 3.25.2 (TIKA-2773).
* Remove duplication of notes in PPT slides (TIKA-2735)
* Use -javaHome or $JAVA_HOME (if they exist) when
spawning child in tika-server's -spawnChild mode.
* Fixed closing of styles around Hyperlinks in Word Parser
Contributed by Ronan O'Sullivan (TIKA-2599).
Release 1.19.1 - 10/4/2018
* Update PDFBox to 2.0.12, jempbox to 1.8.16
and jbig2 to 3.0.2 (TIKA-2745).
* Fix regression in parser for MP3 files (TIKA-2730).
* Updated Python Dependency Check for TesseractOCR (TIKA-2740).
* Improve SAXParser robustness (TIKA-2727).
* Remove dependency on slf4j-log4j12 by upgrading jmatio (TIKA-2742).
* Replace com.sun.xml.bind:jaxb-impl and jaxb-core with
org.glassfish.jaxb:jaxb-runtime and jaxb-core (TIKA-2743)
Release 1.19 - 9/14/2018
* Require Java 8 (TIKA-2679).
* Enable building with Java 11 (TIKA-2668)
* Add an option to make tika-server robust against infinite loops,
OOMs, and memory leaks (TIKA-2725).
* Allow configuration of the Tesseract parser via the standard
tika-config.xml options (TIKA-2705).
* Improve handling of empty cells across table-based
formats (TIKA-2479).
* Add a Standards compliant HTML encoding detector
via Gerard Bouchar (TIKA-2673).
* Improved XML parsing -- limited default entity expansions to 20.
To raise this limit, add -Djdk.xml.entityExpansionLimit=XXX to
your commandline.
* Mime magic improvements for Olympus RAW (TIKA-2658), interpreted
server-side languages via HTTP (TIKA-2648), MHTML (TIKA-2723)
* Add absolute timeout to ForkParser rather than testing
for active (TIKA-2656).
* Make the RecursiveParserWrapper work with the ForkParser (TIKA-2655).
* Allow the ForkParser to specify a directory containing tika-app.jar
for use by the ForkServer. This allows users to keep most of the
parser dependencies out of their code; and it allows for an easy
addition of optional jars for Parser dependencies,
such as the xerial sqlite jar (TIKA-2653).
* Use a pool for SAXParsers and DOMBuilders rather than creating
a new parser/builder for every parse.
For better performance, set XMLReaderUtils.setPoolSize() to the
number of threads you're using with Tika (TIKA-2645).
* Add the RecursiveParserWrapperHandler to improve the RecursiveParserWrapper
API slightly (TIKA-2644).
* Upgraded to Commons-Compress 1.18 (TIKA-2707).
* Upgraded to Apache POI 4.0.0 (TIKA-2552).
* Upgraded to Apache PDFBox 2.0.11 (TIKA-2681).
* Upgraded to deeplearning4j 1.0.0-beta2 (TIKA-2672).
* Upgraded jmatio to 1.4 (TIKA-2667)
* Upgraded Apache Lucene to 7.4.0 in tika-eval and tika-examples (TIKA-2695).
* Upgraded junrar to 1.0.1 (TIKA-2664).
* Numerous other upgrades (TIKA-2692).
* Excluded Spring as a transitive dependency (TIKA-2721).
Release 1.18 - 4/20/2018
* Upgrade jackson to 2.9.5 (TIKA-2634).
* Add support for brotli (TIKA-2621).
* Upgrade PDFBox to 2.0.9 and include new jbig2-imageio
from org.apache.pdfbox (TIKA-2579 and TIKA-2607).
* Support for TIFF images in PDF files (TIKA-2338)
* Detection of full encrypted 7z files (TIKA-2568)
* Various new mimes and typo fixes in tika-mimetypes.xml
via Andreas Meier (TIKA-2527).
* Revert to listenForAllRecords=false in ExcelExtractor
via Grigoriy Alekseev (TIKA-2590)
* Add workaround to identify TIFFs that might confuse
commons-compress's tar detection via Daniel Schmidt
(TIKA-2591)
* Ignore non-IANA supported charsets in HTML meta-headers
during charset detection in HTMLEncodingDetector
via Andreas Meier (TIKA-2592)
* Add detection and parsing of zstd (if user provides
com.github.luben:zstd-jni) via Andreas Meier (TIKA-2576)
* Allow for RFC822 detection for files starting with "dkim-"
and/or "x-" via Andreas Meier (TIKA-2578 and TIKA-2587)
* Extract xlsx files embedded in OLE objects within PPT and PPTX
via Brian McColgan (TIKA-2588).
* Extract files embedded in HTML and javascript inside HTML
that are stored in the Data URI scheme (TIKA-2563).
* Extract text from grouped text boxes in PPT (TIKA-2569).
* Extract language metadata item from PDF files via Matt Sheppard (TIKA-2559)
* RFC822 with multipart/mixed, first text element should be treated
as the main body of the email, not an attachment (TIKA-2547).
* Swap out com.tdunning:json for com.github.openjson:openjson to avoid
jar conflicts (TIKA-2556).
* No longer hardcode HtmlParser for XML files in tika-server (TIKA-2551).
* Require Java 8 (TIKA-2553).
* Add a parser for XPS (TIKA-2524).
* Mime magic for Dolby Digital AC3 and EAC3 files
* Fixed bug where TesseractOCRParser ignores configured ImageMagickPath,
and set rotation script to ignore Python warnings (TIKA-2509)
* Upgrade geo-apis to 3.0.1 (TIKA-2535)
* Mime definition and magic improvements for text-based programming
and config formats (TIKA-2554, TIKA-2567, TIKA-1141)
* Added local Docker image build using dockerfile-maven-plugin to allow
images to be built from source (TIKA-1518).
* Support for SAS7BDAT data files (TIKA-2462)
* Handle .epub files using .htm rather than .html extensions for the
embedded contents (TIKA-1288)
* Mime magic for ACES Images (TIKA-2628) and DPX Images (TIKA-2629)
* For sparse XLSX and XLSB files, always output missing cells to
the left of filled ones (matching XLS), and optionally output
missing rows on all 3 formats if requested via the
OfficeParserContext (TIKA-2479)
Release 1.17 - 12/8/2017
***NOTE: THIS IS THE LAST VERSION OF TIKA THAT WILL RUN
ON Java 7. The next versions will require Java 8***
* Fix thread-safety in ChmExtractor (TIKA-2519).
* Upgrade cxf to 3.0.16 (TIKA-2516).
* Allow users to configure maxMainMemoryBytes for PDFs via shrike (PR-213).
* Extract underline and strikethrough in docx (TIKA-2347 and TIKA-2512).
* Cache TikaConfig in EmbeddedDocumentUtil for better performance
in documents with large number of attachments (TIKA-2511).
* Extract media files from ooxml (TIKA-2510).
* Standardize the way the Image and Video captioning
dockers and extraction work (TIKA-2400, GitHub-208)
* Upgrade to xmpcore 5.1.3 (TIKA-2034).
* Upgrade to metadata-extractor 2.10.1 (TIKA-2486).
* Upgrade to OpenNLP 1.8.3 (TIKA-2502).
* Upgrade to Jackson 2.9.2 (TIKA-2501).
* Catch potential NPE in getting InputStream for attachments
in PST file (TIKA-2488).
* Upgrade to PDFBox 2.0.8 (TIKA-2489).
* Allow configuration of markLimit in EncodingDetectors
via tika-config.xml (TIKA-2485).
* RFC822Parser now selects the best alternative for
multipart/alternative body components. This aligns with the
behavior of the OutlookParser (TIKA-2478). Users can select
legacy behavior via the "extractAllAlternatives" parameter
in the RFC822 parser definition in tika-config.xml.
* Narrow mime detection for ms-owner files and add detection
for .nls files (TIKA-2469).
* Fix bug in CharsetDetector that led to different detected charsets
depending on whether user setText with a byte[] or an InputStream
via Sean Story (TIKA-2475).
* Remove JAXB for easier use with Java 9 via Robert Munteanu (TIKA-2466).
* Upgrade to POI 3.17 (TIKA-2429).
* Enabling extraction of standard references from text (TIKA-2449).
* Load external custom mimetypes XML from system property
tika.custom-mimetypes (TIKA-2460).
* Extract number of tiffs in a multi-page tiff (TIKA-2451).
* Fix detection of emails extracted from mbox (TIKA-2456).
* Add OverrideDetector and allow PSTParser to specify body content type
as text or html -- to avoid incorrect auto-detection of
rfc/mbox, etc. (TIKA-2454)
* AutoDetectParser throws ZeroByteFileException for zero-byte files after
detection on the file extension (TIKA-2450).
* Extract phonetic runs in docx with experimental SAX parser (TIKA-2448).
* Extract phonetic runs from xls and allow users to turn off extraction
of phonetic runs in both xls and xlsx (TIKA-2440).
* OOXML locale should be set by POI's LocaleUtil not Locale.getDefault().
Fix unit tests to be robust against different locales in OOXML
and ExcelParser (TIKA-2438).
* Upgrade to PDFBox 2.0.7 (TIKA-2431).
* Tika now has support for automatic image captioning, that
combines Computer Vision and Natural Language Processing to
automatically generate a readable caption for an image
(TIKA-2262, TIKA-2355, TIKA-2402, Gh-198, Gh-196, Gh-189).
* Add TestCorruptedFiles to allow devs to test parsers against
corrupted input files (TIKA-2430).
* Correct Mimetype definition for Windows batch files (CMD and BAT)
which are the same (TIKA-2445)
* PSDParser memory use improvements (TIKA-2447)
* Add underline extraction from Word documents (doc/docx) via Stuart Hendren
as well as strikethrough extraction in docx (TIKA-2347, GitHub-173)
* Corrected Tesseract OCR rotation.py script and made it a configurable
option via Peter Weiss (TIKA-2385)
Release 1.16 - 7/7/2017
* Exclude jj2000 from edu.ucar grip to avoid potential
license conflicts with ASL 2.0
* Add Age recognition using Ensemble model for Linear regression
and Apache OpenNLP Maximum Entropy. Tika can now detect age from
text (TIKA-1988).
* Add Tika Deep Learning support for the VGG16 model for
Very Deep Convolutional Networks for Large-Scale Image Recognition.
Now Tika supports both Inception v3/v4 and VGG16 based image
recognition (TIKA-2298).
* Extract macros from PPT (TIKA-2089).
* Extract absolute path for last saved location when available
in .xlsx and .xlsb (TIKA-2335).
* Rename SentimentParser to SentimentAnalysisParser to
prevent conflict with dependency (TIKA-2368).
* tika-app now extracts inline images in PDFs by
default, and it includes a warning to users that this is not the
default behavior elsewhere in Tika (TIKA-2374).
* Allow configurability of warnings for problems during
parser initialization (TIKA-2389).
* Upgrade to Jackcess 2.1.8 (TIKA-2380).
* Upgrade to POI 3.17-beta1 (TIKA-2336).
* Remove non-ASL-2.0-compatible org.json (TIKA-1804).
* Allow extraction of