Automatic Document Analyzer and Classifier

Total Page:16

File Type:pdf, Size:1020Kb

Automatic Document Analyzer and Classifier ADAC Automatic Document Analyzer and Classifier A. Guitouni A.-C. Boury-Brisset DRDC Valcartier L. Belfares K. Tiliki Université Laval C. Poirier Intellaxiom Inc. Defence R&D Canada – Valcartier Technical Report DRDC Valcartier TR 2004-265 October 2006 ADAC Automatic Document Analyzer and Classifier A. Guitouni A.-C. Boury-Brisset DRDC Valcartier L. Belfares K. Tiliki Université Laval C. Poirier Intellaxiom Inc. Defence R&D Canada - Valcartier Technical Report DRDC Valcartier TR 2004-265 October 2006 Author A. Guitouni, A.-C. Boury-Brisset, L. Belfares, K. Tiliki and C. Poirier Approved by Dr. E. Bossé Section Head / Decision Support System Section Approved for release by G. Bérubé Chief Scientist © Her Majesty the Queen as represented by the Minister of National Defence, 2006 © Sa majesté la reine, représentée par le ministre de la Défense nationale, 2006 Abstract Military organizations have to deal with an increasing number of documents coming from different sources and in various formats (paper, fax, e-mails, electronic documents, etc.) The documents have to be screened, analyzed and categorized in order to interpret their contents and gain situation awareness. These documents should be categorized according to their contents to enable efficient storage and retrieval. In this context, intelligent techniques and tools should be provided to support this information management process that is currently partially manual. Integrating the recently acquired knowledge in different fields in a system for analyzing, diagnosing, filtering, classifying and clustering documents with a limited human intervention would improve efficiently the quality of information management with reduced human resources. A better categorization and management of information would facilitate correlation of information from different sources, avoid information redundancy, improve access to relevant information, and thus better support decision-making processes. DRDC Valcartier’s ADAC system (Automatic Document Analyzer and Classifier) incorporates several techniques and tools for document summarization and semantic analysis based on ontology of a certain domain (e.g. terrorism), and algorithms of diagnosis, classification and clustering. In this document, we describe the architecture of the system and the techniques and tools used at each step of the document processing. For the first prototype implementation, we focused on the terrorism domain to develop the document corpus and related ontology. Résumé Les organisations militaires font face à une augmentation notable du nombre de documents provenant de différentes sources en formats divers (papier, télécopie, courriels, documents électroniques, etc.) Ces documents doivent être scrutés, analysés et catégorisés afin d’en interpréter le contenu pour comprendre la situation. Ils doivent donc être catégorisés selon leur contenu pour un meilleur archivage et une recherche ultérieure plus efficace. Dans ce contexte, des techniques et des outils évolués devront donc être développés pour appuyer et mener ce processus de gestion de l’information qui, actuellement, est essentiellement effectué de façon manuelle. L’intégration de connaissances nouvelles provenant de différents domaines au sein d’un même système pour la gestion documentaire traitant notamment l’analyse, le diagnostic, le filtrage, la classification et l’organisation de documents devrait permettre d’en améliorer notablement l’efficacité, et ce, avec un minimum d’intervention humaine. Une meilleure gestion devrait faciliter l’intégration d’informations provenant de diverses sources, éliminer toute redondance, améliorer l’accès à l’information pertinente et fournir ainsi, en bout de ligne, un meilleur soutien au processus de prise de décision. Le système ADAC (Automatic Document Analyzer and Classifier) conçu à RDDC Valcartier incorpore différentes techniques et outils pour le résumé et l’analyse sémantique basée sur l’ontologie d’un domaine particulier (p. ex. celui du terrorisme), et des algorithmes de diagnostic, de classification et l’organisation de documents. Dans ce rapport, nous décrivons l’architecture du système, ainsi que les techniques et outils utilisés à chaque étape du traitement d’un document. Pour l’implantation du prototype, l’accent a été mis sur le domaine du terrorisme pour développer une ontologie, ainsi qu’une collection de documents adaptée. DRDC Valcartier TR 2004-265 i This page intentionally left blank. ii DRDC Valcartier TR 2004-265 Executive summary Military organizations, in particular intelligence or command centers have to deal with an increasing number of documents coming from different sources and in various formats (paper, fax, e-mail messages, electronic documents, etc.). These documents must be analyzed in order to interpret their contents and gain situation awareness. These documents should be categorized according to their content to enable efficient storage and retrieval. In this context, intelligent techniques and tools should be provided to support this information management process that is currently partly manual. Automatic, intelligent processing of documents is at the intersection of many fields of research, especially Linguistics and Artificial Intelligence, including natural language processing, pattern recognition, semantic analysis and ontology. Integrating the recently acquired knowledge in these fields in a system for analyzing, diagnosing, filtering, classifying and clustering documents with limited human intervention would improve efficiently the quality of information management with reduced human resources. A better categorization and management of information would facilitate the correlation of information from different sources, avoid information redundancy, improve access to relevant information, and thus better support decision-making processes. This is the purpose of the work we have undertaken at DRDC Valcartier as part of the Common Operational Picture for 21st Century Technology Demonstration project. The ADAC system (Automatic Document Analyzer and Classifier) incorporates several techniques and tools for document summarization and semantic analysis based on the ontology of a certain domain (e.g. terrorism), and algorithms of diagnostic, classification and clustering. A document is processed through the following steps: i) Summarization: large documents are summarized to provide a synthesized view of their content; ii) Statistical and semantic analysis: the document is indexed by identifying the attributes that best characterize it. Both statistical analysis and semantic processing exploiting domain ontology are carried out at this stage; iii) Diagnosis: intercept relevant document matching criteria provided by the user (e.g. document on a particular subject) in order to execute an appropriate action (e.g. alert); iv) Filtering/classification: classify/categorize the document in predefined hierarchical classes; and v) Clustering: assign the document to the most similar group of previously processed documents. External actions can then be triggered on specific classes of documents (e.g. alerts, visualization and data mining). Using a launching agent, ADAC checks periodically the presence of new documents and processes them. The diagnostic and filtering/classification tests may be processed on previously analyzed documents if new directives require it. In this report, we describe the architecture of the system and the techniques and tools used at each step of the document processing. For the first prototype implementation, we have chosen to focus our document corpus and related ontology on the terrorism domain. Guitouni, A, Boury-Brisset, A.-C., Belfares, L., Tiliki, K., Poirier, C., 2006. ADAC: Automatic Document Analyzer and Classifier, DRDC Valcartier, TR 2004-265, Defence R&D Canada. DRDC Valcartier TR 2004-265 iii Sommaire Les organisations militaires, en particulier les cellules de renseignement et les centres de commandement, doivent traiter un nombre sans cesse croissant d’informations provenant de différentes sources sous divers formats (papier, fax, courriels, documents électroniques, etc.). Ces documents doivent être scrutés et analysés afin d’en interpréter le contenu pour une meilleure gestion de situation. Ils doivent donc être catégorisés selon leur sujet pour permettre, d’une part, un archivage efficace et, d’autre part, pour faciliter une recherche ultérieure. Dans ce contexte, des techniques et des outils avancés devront donc être développés pour soutenir et mener ce processus de gestion de l’information, qui à l’heure actuelle est essentiellement effectué manuellement. La compréhension automatique de documents est un domaine de recherche multi-disciplinaire touchant en particulier la linguistique computationnelle et l’intelligence artificielle, notamment le traitement de la langue naturelle, la reconnaissance de formes, l’analyse sémantique et ontologique. L’intégration dans un même système des résultats de recherches récentes dans différents champs de connaissances reliés à la gestion documentaire traitant notamment de l’analyse, du diagnostic, du filtrage, et de la classification de documents devrait permettre d’en améliorer considérablement l’efficacité avec un minimum d’intervention humaine. Une meilleure catégorisation et une gestion adéquate de l’information devraient faciliter l’aggrégation d’informations provenant de diverses sources, éliminer toute redondance, améliorer l’accès à l’information pertinente et ainsi fournir un meilleur soutien au processus de
Recommended publications
  • Evaluation of an Ontology-Based Knowledge-Management-System
    Information Services & Use 25 (2005) 181–195 181 IOS Press Evaluation of an Ontology-based Knowledge-Management-System. A Case Study of Convera RetrievalWare 8.0 1 Oliver Bayer, Stefanie Höhfeld ∗, Frauke Josbächer, Nico Kimm, Ina Kradepohl, Melanie Kwiatkowski, Cornelius Puschmann, Mathias Sabbagh, Nils Werner and Ulrike Vollmer Heinrich-Heine-University Duesseldorf, Institute of Language and Information, Department of Information Science, Universitaetsstraße 1, D-40225 Duesseldorf, Germany Abstract. With RetrievalWare 8.0TM the American company Convera offers an elaborated software in the range of Information Retrieval, Information Indexing and Knowledge Management. Convera promises the possibility of handling different file for- mats in many different languages. Regarding comparable products one innovation is to be stressed particularly: the possibility of the preparation as well as integration of an ontology. One tool of the software package is useful in order to produce ontologies manually, to process existing ontologies and to import the very. The processing of search results is also to be mentioned. By means of categorization strategies search results can be classified dynamically and presented in personalized representations. This study presents an evaluation of the functions and components of the system. Technological aspects and modes of opera- tion under the surface of Convera RetrievalWare will be analysed, with a focus on the creation of libraries and thesauri, and the problems posed by the integration of an existing thesaurus. Broader aspects such as usability and system ergonomics are integrated in the examination as well. Keywords: Categorization, Convera RetrievalWare, disambiguation, dynamic classification, information indexing, information retrieval, ISO 5964, knowledge management, ontology, user interface, taxonomy, thesaurus 1.
    [Show full text]
  • Evaluation of an Ontology-Based Knowledge- Management-System. a Case Study of Convera Retrievalware 8.0
    See discussions, stats, and author profiles for this publication at: http://www.researchgate.net/publication/228969390 Evaluation of an ontology-based knowledge- management-system. A case study of Convera RetrievalWare 8.0 ARTICLE in INFORMATION SERVICES & USE · JANUARY 2005 CITATIONS DOWNLOADS VIEWS 2 11 210 10 AUTHORS, INCLUDING: Cornelius Puschmann Mathias Sabbagh Zeppelin University Heinrich-Heine-Universität Düsseldorf 30 PUBLICATIONS 62 CITATIONS 1 PUBLICATION 2 CITATIONS SEE PROFILE SEE PROFILE Available from: Mathias Sabbagh Retrieved on: 22 June 2015 Information Services & Use 25 (2005) 181–195 181 IOS Press Evaluation of an Ontology-based Knowledge-Management-System. A Case Study of Convera RetrievalWare 8.0 1 Oliver Bayer, Stefanie Höhfeld ∗, Frauke Josbächer, Nico Kimm, Ina Kradepohl, Melanie Kwiatkowski, Cornelius Puschmann, Mathias Sabbagh, Nils Werner and Ulrike Vollmer Heinrich-Heine-University Duesseldorf, Institute of Language and Information, Department of Information Science, Universitaetsstraße 1, D-40225 Duesseldorf, Germany Abstract. With RetrievalWare 8.0TM the American company Convera offers an elaborated software in the range of Information Retrieval, Information Indexing and Knowledge Management. Convera promises the possibility of handling different file for- mats in many different languages. Regarding comparable products one innovation is to be stressed particularly: the possibility of the preparation as well as integration of an ontology. One tool of the software package is useful in order to produce ontologies manually, to process existing ontologies and to import the very. The processing of search results is also to be mentioned. By means of categorization strategies search results can be classified dynamically and presented in personalized representations. This study presents an evaluation of the functions and components of the system.
    [Show full text]
  • Geethanjali College of Engineering and Technology Cheeryal(V), Keesara(M), R.R.Dist-501 301
    Geethanjali College of Engineering and Technology Cheeryal(V), Keesara(M), R.R.Dist-501 301 DEPARTMENT OF CSE QC Format No.: Name of the Subject/Lab Course: Information Retrieval Systems Programme: UG ------------------------------------------------------------------------------------------------- Branch: CSE Version No: 1.0 Year: 4 Updated On: Semester: 1 No.of Pages: 100 ----------------------------------------------------------------------------------------------------------- Classification Status: (Unrestricted/Restricted) Distribution List: ----------------------------------------------------------------------------------------------------------- Prepared By: 1.Name: M . Raja Krishna Kumar 2.Sign: 3.Design: Associate Professor 4.Date: ----------------------------------------------------------------------------------------------------------- Verified By: * For Q.C Only: 1.Name: 1.Name: 2.Sign: 2.Sign: 3.Design: 3.Design: 4.Date 4.Date ----------------------------------------------------------------------------------------------------------- Approved By:(HOD) 1.Name: Prof. Dr. P. V. S. Srinivas 2.Sign: 3.Date: Department of CSE Information Retrieval Systems Subject File By M. Raja Krishna Kumar AMIETE(IT), M.TECH(IP), (Ph.D.(CSE)) Associate Professor Page 2 of 100 Contents S.No. Name of the Topic Page No 1. Detailed Lecture Notes on all Units 2. 3. 4. 5. Page 3 of 100 JNTU Syllabus UNIT-I Introduction: Definition, Objectives, Functional Overview, Relationship to DBMS, Digital libraries and Data Warehouses. Information Retrieval System
    [Show full text]
  • Introduction to FAST Search Server 2010 for Sharepoint
    www.it-ebooks.info www.it-ebooks.info Working with Microsoft ® FAST™ Search Server 2010 for SharePoint ® Mikael Svenson Marcus Johansson Robert Piddocke www.it-ebooks.info Published with the authorization of Microsoft Corporation by: O’Reilly Media, Inc. 1005 Gravenstein Highway North Sebastopol, California 95472 Copyright © 2012 by Mikael Svenson, Marcus Johansson, Robert Piddocke All rights reserved. No part of the contents of this book may be reproduced or transmitted in any form or by any means without the written permission of the publisher. ISBN: 978-0-7356-6222-3 1 2 3 4 5 6 7 8 9 LSI 7 6 5 4 3 2 Printed and bound in the United States of America. Microsoft Press books are available through booksellers and distributors worldwide. If you need support related to this book, email Microsoft Press Book Support at [email protected]. Please tell us what you think of this book at http://www.microsoft.com/learning/booksurvey. Microsoft and the trademarks listed at http://www.microsoft.com/about/legal/en/us/IntellectualProperty/ Trademarks/EN-US.aspx are trademarks of the Microsoft group of companies. All other marks are property of their respective owners. The example companies, organizations, products, domain names, email addresses, logos, people, places, and events depicted herein are fictitious. No association with any real company, organization, product, domain name, email address, logo, person, place, or event is intended or should be inferred. This book expresses the author’s views and opinions. The information contained in this book is provided without any express, statutory, or implied warranties.
    [Show full text]
  • User Requirements and Functional Specification of The
    User Requirements and Functional Specification of the EuroWordNet project Version 5, Final October, 1996 Laura Bloksma£ Pedro Luis Díez-Orzas$ Piek Vossen£ Deliverable D001, WP1, EuroWordNet, LE2-4003 £ Computer Centrum Letteren, University of Amsterdam $ Novell Linguistic Development, Antwerp Identification number LE-4003-D-001 Type Document Title User requirements and functional specification of EuroWordNet Status Final Deliverable D001 Work Package WP1 Task T1 Period covered March - June 1996 Date October, 1996 Version 5 Number of pages 66 Authors Laura Bloksma, Pedro Díez-Orzas, Piek Vossen, WP/Task responsible Novell Project contact point Piek Vossen Computer Centrum Letteren University of Amsterdam Spuistraat 134 1012 VB Amsterdam The Netherlands tel. +31 20 525 4624 fax. +31 20 525 4429 e-mail: [email protected] http://www.let.uva.nl/CCL/EuroWordNet.html EC project officer Jose Soler Status Public Actual distribution Project Consortium The EuroWordNet User-Group The EuroWordNet WWW page Suplementary notes Key words Lexical semantic databases, Information Retrieval, Language Engineering Abstract In this document the general design of the EuroWordNet database is described based on the user-requirements and the technical state of the art for building multilingual semantic resources. The user- requirements are discussed from two different perspectives: the actual use of the resource in a multilingual information retrieval system developed by Novell Linguistic Development and the potential use of the resource by a diverse group of institutes and companies in Europe, constituted by the EuroWordNet user-group. The purpose of the latter group is to create a wider awareness of the use of this type of resources and to establish cooperation with other groups that build such resources to develop standards and make resources compatible.
    [Show full text]
  • Document and Records Management
    Technology Evaluation and Comparison Report www.butlergroup.com DocumentDocument andand RecordsRecords ManagementManagement Managing Information for Compliance, Efficiency, and Value February 2005 For more information on Butler Group’s Products and Services or to register FREE to receive TECHwatch, written by Martin Butler and Tim Jennings, together with a monthly guide to our Research and Events, visit www.butlergroup.com Founder and President Important Notice Martin Butler We have relied on data and information which we reasonably believe to be up-to-date and correct when preparing this Report, but because it comes Research from a variety of sources outside of our direct control, we cannot guarantee Susan Clarke that all of it is entirely accurate or up-to-date. Mike Davis This Report is of a general nature and not intended to be specific, Richard Edwards customised, or relevant to the requirements of any particular set of circumstances. The interpretations contained in the Report are non-unique and you are responsible for carrying out your own interpretation of the data and information upon which this Report was based. Accordingly, Butler Direct Limited is not responsible for your use of this Report in any specific circumstances, or for your interpretation of this Report. Published by Butler Direct Limited The interpretation of the data and information in this Report is based on Published February 2005 generalised assumptions and by its very nature is not intended to produce © Butler Direct Limited accurate or specific results. Accordingly, it is your responsibility to use your own relevant professional skill and judgement to interpret the data and All rights reserved.
    [Show full text]
  • Search Insights 2020
    Search Insights 2020 The Search Network February 2020 Contents Introduction 1 The Cambrian explosion (of search), Paul Cleverley 4 Benchmarking enterprise search – a perspective from Denmark, 8 Kurt Kragh Sørensen The advent of natural language information retrieval, Max Irwin 12 Microsoft Search in Office 365, Agnes Molnar 15 Content integration, Valentin Richter 20 Skills for effective relevance engineering, Charlie Hull 23 The importance of informed query log analysis, Martin White 26 Good practice in taxonomy project management, Helen Lippell 31 Changes in open source search, Elizabeth Haubert 35 Searching for expertise and experts, Martin White 39 Search resources: books and blogs 45 Enterprise search chronology 47 Search vendors 50 Search integrators 52 Glossary 54 This work is licensed under the Creative Commons Attribution 2.0 UK: England & Wales License. To view a copy of this license, visit https://creativecommons.org/licenses/by/2.0/uk/ or send a letter to Creative Commons, PO Box 1866, Mountain View, CA 94042, USA. Editorial services provided by Val Skelton ([email protected]) Design & Production by Simon Flegg - Hub Graphics Ltd (www.hubgraphics.co.uk) Search Insights 2020 Introduction The Search Network is a community of expertise. It was set up in October 2017 by a group of eight search implementation specialists working in Europe and North Amer- ica. We have known each other for at least a decade and share a common passion for search that delivers business value through providing employees with access to infor- mation and knowledge that enables them to make decisions that benefit the organisa- tion and their personal career objectives.
    [Show full text]
  • Digital Notes on Information Retrieval Systems (R17a1209) B.Tech Iv Year
    DIGITAL NOTES ON INFORMATION RETRIEVAL SYSTEMS (R17A1209) B.TECH IV YEAR - I SEM (2020-2021) DEPARTMENT OF INFORMATION TECHNOLOGY MALLA REDDY COLLEGE OF ENGINEERING & TECHNOLOGY (Autonomous Institution – UGC, Govt. of India) (Affiliated to JNTUH, Hyderabad, Approved by AICTE - Accredited by NBA & NAAC – ‘A’ Grade - ISO 9001:2015 Certified) Maisammaguda, Dhulapally (Post Via. Hakimpet), Secunderabad – 500100, Telangana State, INDIA. MRCET-IT Page 1 MALLA REDDY COLLEGE OF ENGINEERING & TECHNOLOGY DEPARTMENT OF INFORMATION TECHNOLOGY IV Year B.Tech IT –I Sem L T /P/D C 3 -/-/- 3 (R17A1209)INFORMATION RETRIEVAL SYSTEMS (Core Elective IV) OBJECTIVES Study fundamentals of DBMS, Data warehouse and Digital libraries Learn various preprocessing techniques and indexing approaches in text mining Know various clustering approaches and study different similarity measures Study various search techniques in information retrieval systems Know different cognitive approaches used in text retrieval systems and evaluation approaches Study retrieval in multimedia systems and know various evaluation measures Know about query languages and online IRsystem UNIT-I Introduction: Definition, Objectives, Functional Overview, Relationship to DBMS, Digital libraries and Data Warehouses. Information Retrieval System Capabilities: Search, Browse, Miscellaneous UNIT-II Cataloging and Indexing: Objectives, Indexing Process, Automatic Indexing, Information Extraction. Data Structures: Introduction, Stemming Algorithms, Inverted file structures, N-gram data structure,
    [Show full text]
  • Information Storage and Retrieval Systems: Theory and Implementation
    INFORMATION STORAGE AND RETRIEVAL SYSTEMS Theory and Implementation Second Edition THE KLUWER INTERNATIONAL SERIES ON INFORMATION RETRIEVAL Series Editor W. Brace Croft University of Massachusetts, Amherst Also in the Series: MULTIMEDIA INFORMATION RETRIEVAL: Content-Based Information Retrieval from Large Text and Audio Databases, by Peter Schäuble; ISBN: 0-7923-9899-8 INFORMATION RETRIEVAL SYSTEMS: Theory and Implementation, by Gerald Kowalski; ISBN: 0-7923-9926-9 CROSS-LANGUAGE INFORMATION RETRIEVAL, edited by Gregory Grefenstette; ISBN: 0-7923-8122-X TEXT RETRIEVAL AND FILTERING: Analytic Models of Performance, by Robert M. Losee; ISBN: 0-7923-8177-7 INFORMATION RETRIEVAL: UNCERTAINTY AND LOGICS: Advanced Models for the Representation and Retrieval of Information, by Fabio Crestani, Mounia Lalmas, and Cornelis Joost van Rijsbergen; ISBN: 0- 7923-8302-8 DOCUMENT COMPUTING: Technologies for Managing Electronic Document Collections, by Ross Wilkinson, Timothy Arnold-Moore, Michael Fuller, Ron Sacks-Davis, James Thom, and Justin Zobel; ISBN: 0-7923-8357-5 AUTOMATIC INDEXING AND ABSTRACTING OF DOCUMENT TEXTS, by Marie-Francine Moens; ISBN 0-7923-7793-1 ADVANCES IN INFORMATIONAL RETRIEVAL: Recent Research from the Center for Intelligent Information Retrieval, by W. Bruce Croft; ISBN 0- 7923-7812-1 INFORMATION STORAGE AND RETRIEVAL SYSTEMS Theory and Implementation Second Edition by Gerald J. Kowalski Central Intelligence Agency Mark T. Maybury The MITRE Corporation KLUWER ACADEMIC PUBLISHERS NEW YORK, BOSTON, DORDRECHT, LONDON, MOSCOW eBook
    [Show full text]
  • Convera 10 8 Final.Fm
    Convera (Offline) © 2013 by Stephen E. Arnold, www.arnoldit.com Convera positions itself as a knowledge discovery platform, a marketing angle that vendors have fol- lowed. Convera faces challenges delivering its vision to licensees. Sizzle is not the steak in search. Author’s note: This is an unpublished, preliminary draft of a description originally destined for a client report. The information is provided as part of ArnoldIT’s archiving project. The information in this draft may not be used without prior writ- ten permission. The information in this document was written before Convera went out of business with the sale of its remaining assets to Vertical Search Works. Convera ushered in the era of selling “everything plus the kitchen sink” search. The firm was among the first to package search as “concept searching,” “knowl- edge management” and “text analytics”, thus kicking off an era of calling search something to capture more revenue. The company’s contribution to search was to lay out a road map of where information retrieval would go in the next decade. Convera narrowed it focus to vertical search or eCommerce search. Upon its dis- solution, Convera professionals moved to consulting, engineering services, or other search vendors. This information is a rough draft and is frozen. 1 Introduction Excalibur Technologies, backed by the low-profile investment firm Allen & Company, was the precursor of Convera. Based in the Washington, DC area, Convera was formed by Excalibur Technologies combined with Intel’s Inter- active Media Services division. It is a leading provider of content manage- ment solutions that unlock the value of digital content.
    [Show full text]
  • Going Beyond Simple Keyword Search in the Next Generation of Information Search Tools
    Going beyond simple keyword search in the next generation of Information Search Tools Anastasio Molano Denodo Technologies Inc. Almirante Francisco Moreno, 5 28040 Madrid - Spain [email protected] Index NLP technologies go beyond traditional Information 1. Introduction Retrieval techniques enabling a system to accomplish a 2. Language Engineering Techniques and human-like understanding of text, and thus, permitting to Resources extract useful meaning from unstructured text. Lexical Resources NLP Techniques Search companies such as Ask Jeeves, Convera, Northern 3. Market situation and Prospects Light, Verity, SmartLogik, Q-Go, and Cognit among European initiatives and market prospects others, have incorporated NLP techniques in their search Research in Spain and market prospects solutions. Iberoamerican initiatives Expectations are high, as these tools are having a great 4. Conclusions impact on the industry, especially on large companies corporate Intranet searchers, and generally, in those Introduction applications in which searching efficiently over large document repositories is crucial (e.g. Digital Libraries, Medicine databases, Legal databases, Competitive Wouldn’t be nice if you could receive an exact answer Intelligence tools, etc.). The current relevance of when you query a search engine, instead of a list of multilingual, cross language and interactive retrieval will URL’s? Questions such as “What is an iceberg” or “What further increase demand on this kind of technologies. is the distance between Rome and Paris?” would receive a precise answer, rather than a list of related documents. This Given the size of digital information universally available will be possible in the short future, thanks to the evolution today, along searching itself, other complementary of Natural Language Processing techniques (NLP in short).
    [Show full text]
  • A Survey of Concept-Based Information Retrieval Tools on the Web
    A Survey of Concept-based Information Retrieval Tools on the Web Hele-Mai HAAV Tanel-Lauri LUBI Institute of Cybernetics at Tallinn Technical University Akadeemia tee 21, 12618 Tallinn [email protected] [email protected] Abstract: In order to solve the problem of information overkill on the web current information retrieval tools need to be improved. Much more "intelligence" should be embedded to search tools to manage effectively search, retrieval, filtering and presenting relevant information. This can be done by concept-based (or ontology driven) information retrieval, which is considered as one of the high-impact technologies for the next ten years. Nevertheless, most of commercial products of search and retrieval category do not report about concept-based search features. The paper provides an overview of concept-based information retrieval techniques and software tools currently available as prototypes or commercial products. Tools are evaluated using feature classification, which incorporates general characteristics of tools and their information retrieval features. 1. Introduction and Motivation Current information retrieval tools mostly use keyword search, which is unsatisfactory option because of its low precision and recall. In this paper, we consider concept-based information retrieval model as a new and promising way of improving search on the web. Informally, concept-based information retrieval is search for information objects based on their meaning rather than on the presence of the keywords in the object. In the last 5 years, concept-based information retrieval tools have been created and used mostly in academic and industrial research environments [Guarino et al 1999, Woods 1998]. For example, in the survey of information retrieval vendors by R.
    [Show full text]