Skip to main navigation Skip to search Skip to main content

Ontology-based information extraction and integration from heterogeneous data sources

  • Paul Buitelaar
  • , Philipp Cimiano
  • , Anette Frank
  • , Matthias Hartung
  • , Stefania Racioppa
  • German Research Centre for Artificial Intelligence (DFKI)
  • Institut für Technik der Informationsverarbeitung(ITIV)
  • Heidelberg University

Research output: Contribution to a Journal (Peer & Non Peer)Articlepeer-review

90 Citations (Scopus)

Abstract

In this paper we present the design, implementation and evaluation of SOBA, a system for ontology-based information extraction from heterogeneous data resources, including plain text, tables and image captions. SOBA is capable of processing structured information, text and image captions to extract information and integrate it into a coherent knowledge base. To establish coherence, SOBA interlinks the information extracted from different sources and detects duplicate information. The knowledge base produced by SOBA can then be used to query for information contained in the different sources in an integrated and seamless manner. Overall, this allows for advanced retrieval functionality by which questions can be answered precisely. A further distinguishing feature of the SOBA system is that it straightforwardly integrates deep and shallow natural language processing to increase robustness and accuracy. We discuss the implementation and application of the SOBA system within the SmartWeb multimodal dialog system. In addition, we present a thorough evaluation of the different components of the system. However, an end-to-end evaluation of the whole SmartWeb system is out of the scope of this paper and has been presented elsewhere by the SmartWeb consortium.

Original languageEnglish
Pages (from-to)759-788
Number of pages30
JournalInternational Journal of Human Computer Studies
Volume66
Issue number11
DOIs
Publication statusPublished - Nov 2008
Externally publishedYes

Keywords

  • Information extraction
  • Knowledge integration
  • Ontology-based natural language processing
  • Question answering

Fingerprint

Dive into the research topics of 'Ontology-based information extraction and integration from heterogeneous data sources'. Together they form a unique fingerprint.

Cite this