The Apache Tika project is pleased to announce the release of Apache Tika 4.1.0. The release contents have been pushed out to the main Apache release site and to the Maven Central sync.
Apache Tika is a toolkit for detecting and extracting metadata and structured text content from various documents using existing parser libraries. Apache Tika 4.1.0 adds significant new capabilities in multimodal inference. This version also includes bug fixes and other improvements, including: OCR and enrichment engines now selected by name through a "text-recognizers" list, substantial performance improvements via spooling to disk less often, tika-server gaining named configuration presets and Micrometer metrics. Note: the project plans to change the default unpack format to frictionless in 4.2.0; please chime in on the dev list if this will be a problem for you. Details can be found in the changes file: https://www.apache.org/dist/tika/4.1.0/CHANGES-4.1.0.txt and in our 4.x docs site: https://tika.apache.org/docs/4.1.x/ Apache Tika is available on the download page: https://tika.apache.org/download.html Apache Tika will be available shortly in binary form or for use with Maven from the Central Repository: https://repo1.maven.org/maven2/org/apache/tika/ When downloading, please remember to verify the downloads using signatures found: https://www.apache.org/dist/tika/KEYS For more information on Apache Tika, visit the project home page: https://tika.apache.org/ Many, many thanks to fellow devs and our larger community! Extra special shout out to Tilman and Oleg for voting through RC1, RC2 and finally RC3! -- Tim Allison, on behalf of the Apache Tika community
