Here are
113 public repositories
matching this topic...
The Apache Tika toolkit detects and extracts metadata and text from over a thousand different file types (such as PPT, XLS, and PDF).
Updated
Jul 28, 2021
Java
Elasticsearch File System Crawler (FS Crawler)
Updated
Jul 28, 2021
Java
Spark-Crawler: Apache Nutch-like crawler that runs on Apache Spark.
Updated
Jul 15, 2021
Java
A cross-platform command line tool for parallelised content extraction and analysis.
Updated
Jul 26, 2021
Java
Use the Java Tika text extraction library on the .NET platform
Updated
Jan 28, 2020
Rich Text Format
Viewers for statistics and dashboarding of Domain Search Engine data
Updated
Jan 19, 2016
Python
Interactive Image similarity and Visual Search and Retrieval application
Updated
Apr 25, 2021
JavaScript
ImageCat is an Apache OODT RADIX application that uses Apache Solr, Apache Tika and Apache OODT to ingest 10s of millions of files (images,but could be extended to other files) in place, and to extract metadata and OCR information from those files/images using Tika and Tesseract OCR.
Updated
Aug 26, 2018
Java
Tika-Similarity uses the Tika-Python package (Python port of Apache Tika) to compute file similarity based on Metadata features.
Updated
Jun 2, 2021
Python
Apache Tika bindings for PHP: extract text and metadata from documents, images and other formats
Code for Machine Learning with TensorFlow: 2nd Edition Published by Manning Publications
Updated
Jun 8, 2021
Jupyter Notebook
R Interface to Apache Tika
Extract and Visualize location from any file
Updated
Jun 10, 2021
JavaScript
Convenience Docker images for Apache Tika Server
Updated
Jul 20, 2021
Shell
Geographic Place, Date/time, and Pattern entity extraction toolkit along with text extraction from unstructured data and GIS outputters.
Updated
Jun 21, 2021
Java
pdf2html is a module which helps to convert PDF file to HTML pages using Apache Tika. This module also helps to generate thumbnail image for PDF file using Apache PDFBox.
Updated
Jul 20, 2021
JavaScript
Distributed, fault tolerant batch processing for Natural Language Applications and Search, using remote partitioning
Updated
Jun 23, 2021
Java
Apache NiFi Custom Processor Extracting Text From Files with Apache Tika
Updated
Jun 15, 2021
Java
Java web application taking IPFS hashes, extracting (textual) content and metadata through Apache's Tika.
Updated
Nov 19, 2020
Java
Small box of pandora to prototype your app with ready for use backend. This is just my compilation of different solutions occasionally applied in hackathons and challenges
A ruby wrapper for the Tika jar (tika-app.jar) that extracts text in a lot of formats from PDF, xls, doc, etc files
Updated
Oct 22, 2020
DIGITAL Command Language
Python bindings for Apache Tika
Updated
Aug 20, 2020
Python
Distill information about amendments to the Oregon Revised Statutes.
Updated
Jan 13, 2020
Haskell
A suite of Machine Learning / Deep Learning Dockerfiles to allow Apache Tika to extract objects and to produce textual captions for images and video
qDesktopSearch - a Qt5 Desktop App for indexing & searching the files of the local machine
Extract text from a document by Apache Tika
Updated
Jun 22, 2021
TypeScript
An Elasticsearch engine plugin for Moodle's Global Search
TYPO3 Extension: solr_file_indexer
Apache Tika Server as a Background Service in Node.js
Updated
Jun 27, 2021
JavaScript
Improve this page
Add a description, image, and links to the
tika
topic page so that developers can more easily learn about it.
Curate this topic
Add this topic to your repo
To associate your repository with the
tika
topic, visit your repo's landing page and select "manage topics."
Learn more
You can’t perform that action at this time.
You signed in with another tab or window. Reload to refresh your session.
You signed out in another tab or window. Reload to refresh your session.
[Updated after reading sotera.github.io/newman/features].
In the "Top Addresses" screenshot below, jeb@jeb.org shows 79% in the donut plot and 0.988 in the bar plot.

I thought the 0.988 was a proportion -- w