Over a million developers have joined DZone.
{{announcement.body}}
{{announcement.title}}

Introducing Hydra: An Open Source Document Processing Framework

DZone's Guide to

Introducing Hydra: An Open Source Document Processing Framework

· Java Zone ·
Free Resource

Get the Edge with a Professional Java IDE. 30-day free trial.

 The above presentation details the document-processing framework named Hydra that was developed by Findwise.

This presentation will detail the document-processing framework called Hydra that has been developed by Findwise. It is intended as a description of the framework and the problem it aims to solve. We will first discuss the need for scalable document processing, outlining that there is a missing link between the open source chain to bridge the gap between source system and the search engine, then will move on to describe the design goals of Hydra, as well as how it has been implemented to meet those demands on flexibility, robustness and ease of use. This session will end by discussing some of the possibilities that this new pipeline framework can offer, such as freely seamlessly scaling up the solution during peak loads, metadata enrichment as well as proposed integration with Hadoop for Map/Reduce tasks such as page rank calculations.

Get the Java IDE that understands code & makes developing enjoyable. Level up your code with IntelliJ IDEA. Download the free trial.

Topics:

Published at DZone with permission of

Opinions expressed by DZone contributors are their own.

{{ parent.title || parent.header.title}}

{{ parent.tldr }}

{{ parent.urlSource.name }}