A Comparative Evaluation of Off-The-Shelf Distributed Semantic

Total Page:16

File Type:pdf, Size:1020Kb

A Comparative Evaluation of Off-The-Shelf Distributed Semantic COGNITIVE NEUROPSYCHOLOGY, 2016 VOL. 33, NOS. 3–4, 175–190 http://dx.doi.org/10.1080/02643294.2016.1176907 A comparative evaluation of off-the-shelf distributed semantic representations for modelling behavioural data Francisco Pereiraa, Samuel Gershmanb , Samuel Ritterc and Matthew Botvinickd aMedical Imaging Technologies, Siemens Healthcare, USA; bDepartment of Psychology and Center for Brain Science, Harvard University, Cambridge, USA; cPrinceton Neuroscience Institute, Princeton University, Princeton, USA; dGoogle DeepMind, London, UK ABSTRACT ARTICLE HISTORY In this paper we carry out an extensive comparison of many off-the-shelf distributed semantic Received 15 May 2015 vectors representations of words, for the purpose of making predictions about behavioural Revised 31 March 2016 results or human annotations of data. In doing this comparison we also provide a guide for how Accepted 5 April 2016 vector similarity computations can be used to make such predictions, and introduce many KEYWORDS resources available both in terms of datasets and of vector representations. Finally, we discuss Distributed semantic the shortcomings of this approach and future research directions that might address them. representation; evaluation; semantic space; semantic vector 1. Introduction computational tasks we are trying to solve (and the more general problem of concept representation in 1.1. Distributed semantic representations the brain) require models that are general enough to We are interested in one particular aspect of concep- encompass the entire English vocabulary as well as tual representation – the meaning of a word – arbitrary linguistic combinations, our focus will be on insofar as it is used in the performance of semantic distributional semantic models. Existing hand-engin- tasks. The study of concepts in general has a long eered systems cannot yet be used to address all the and complex history, and we will not attempt to do tasks that we consider. it justice here (see Margolis & Laurence, 1999, and Common to many distributional semantic models is Murphy, 2002). Researchers have approached the the idea that semantic representations can be con- problem of modelling meaning in diverse ways. One ceived as vectors in a metric space, such that proximity approach is to build representations of a concept – in vector space captures a geometric notion of seman- aword used in one specific sense – by hand, using tic similarity (Turney & Pantel, 2010). This idea has some combination of linguistic, ontological, and fea- been important both for psychological theorizing tural knowledge. Examples of this approach include (Howard, Shankar, & Jagadisan, 2011; Landauer & WordNet (Miller, Beckwith, Fellbaum, Gross, & Miller, Dumais, 1997; Lund & Burgess, 1996; McNamara, 1990), Cyc (Lenat, 1995), and semantic feature norms 2011; Steyvers, Shiffrin, & Nelson, 2004) as well as for collected by various research groups (e.g., McRae, building practical natural language processing Cree, Seidenberg, and McNorgan, 2005, and Vinson systems (Collobert & Weston, 2008; Mnih & Hinton, & Vigliocco, 2008). An alternative approach, known 2007; Turian, Ratinov, & Bengio, 2010). However, as distributional semantics, starts from the idea that vector space models are known to have a number of words occurring in similar linguistic contexts – sen- weaknesses. The psychological structure of similarity tences, paragraphs, documents – are semantically appears to disagree with some aspects of the geome- similar (see Sahlgren, 2008, for a review). A major prac- try implied by vector space models, as evidenced by tical advantage of distributional semantics is that it asymmetry of similarity judgments and violations of enables automatic extraction of semantic represen- the triangle inequality (Griffiths, Steyvers, & Tenen- tations by analysing large corpora of text. Since the baum, 2007; Tversky, 1977). Furthermore, many CONTACT Francisco Pereira [email protected] © 2016 Informa UK Limited, trading as Taylor & Francis Group 176 F. PEREIRA ET AL. vector space models do not deal gracefully with polys- brain activation, as measured with functional mag- emy or word ambiguity (but see Jones, Gruenenfelder, netic resonance imaging (fMRI), in response to seman- & Recchia, 2011, and Turney & Pantel, 2010). Recently, tic stimuli (e.g., a picture of an object together with the a number of different researchers have started focus- word naming it, or the word alone). Such models learn ing on producing vector representations for specific a mapping between the degree to which a dimension meanings of words (Huang, Socher, Manning, & Ng, in a distributed semantic representation vector is 2012; Neelakantan, Shankar, Passos, & McCallum, present and its effect on the overall spatial pattern 2015; Reisinger & Mooney, 2010; Yao & Van Durme, of brain activation. These models can be inverted to 2011), but these are still of limited use without some decode semantic vectors from patterns of brain acti- degree of manual intervention to pick which mean- vation, which allow validation of the mappings by clas- ings to use in generating predictions. We discuss sifying the mental state in new data; this can be done these together with other available approaches in by comparing the decoded vectors with “true” vectors Section 2.1. In the work reported here, we do not extracted from a text corpus. attempt to address these issues directly; our goal is Reviewing this literature is beyond the scope of this to compare the effectiveness of different vector rep- paper, but we will highlight particularly relevant work. resentations of words, rather than comparing them The seminal publication in this area is Mitchell et al. with other kinds of models. (2008), which showed that it was possible to build such forward models, and use them to make predic- tions about new imaging data. The authors rep- 1.2. Modelling human data resented concepts by semantic vectors where Ever since Landauer and Dumais (1997) demonstrated dimensions corresponded to different verbs; the that distributed semantic representations could be vector for a particular concept was derived from co- used to make predictions about human performance occurrence counts of the word naming the concept in semantic tasks, numerous researchers have used and each of those verbs (e.g., the verb “run” co- measures of (dis)similarity between word vectors – occurs more often with animate beings than inani- cosine similarity, euclidean distance, correlation – for mate objects). Subsequently, Just, Cherkassky, Aryal, that purpose. There are now much larger test datasets and Mitchell (2010) produced more elaborate than the TOEFL synonym test used in Landauer and vectors from human judgments, with each dimension Dumais (1997), containing hundreds to thousands of corresponding to one of tens of possible semantic fea- judgments on tasks such as word association, tures. In both cases, this allowed retrospective analogy, and semantic relatedness and similarity, as interpretation of patterns of activation corresponding described in Section 2.3. The availability of LSA as a to each semantic dimension (e.g., ability to manipulate web service1 for calculating similarity between words corresponded to activation in the motor cortex). Other or documents has also allowed researchers to use it groups re-analysing the data from Mitchell et al. (2008) as a means of obtaining a kind of “ground truth” for showed that superior decoding performance could be purposes such as generating stimuli (e.g., Green, obtained by using distributed semantic represen- Kraemer, Fugelsang, Gray, & Dunbar, 2010). In parallel tations rather than human postulated features (e.g., with all this work, researchers within the machine Pereira, Detre, & Botvinick, 2011, and Liu, Palatucci, & learning community have developed many other dis- Zhang, 2009). In particular, Pereira et al. (2011) used tributed semantic representations, mostly used as a topic model of a small corpus of Wikipedia articles components of systems carrying out a variety of to learn a semantic representation where each dimen- natural language processing tasks, ranging from infor- sion corresponded to an interpretable dimension mation retrieval to sentiment classification (Wang & shared by a number of related semantic categories. Manning, 2012). Furthermore, the semantic vectors from brain Beyond behavioural data, distributed semantic rep- images for related concepts exhibited similarity struc- resentations have been used in cognitive neuro- ture that echoed the similarity structure present in science, in the study of how semantic information is word association data, and could also be used to gen- represented in the brain. More specifically, they have erate words pertaining to the mental contents at the been used as components of forward models of time the images were acquired. A systematic COGNITIVE NEUROPSYCHOLOGY 177 comparison of the effectiveness of various kinds of dis- – for all the most commonly used off-the-shelf tributed semantic representations in decoding can be representations. found in Murphy, Talukdar, and Mitchell (2012a). This To do this, we had to assume that the information work has led researchers to consider distributed derived from text corpora suffices to make behav- semantic representations as a core component of ioural predictions; existing literature, and our own forward models of brain activation in semantic tasks, experience, tell us that this is the case. But is this the or even to try to incorporate
Recommended publications
  • Semantic Computing
    SEMANTIC COMPUTING Lecture 7: Introduction to Distributional and Distributed Semantics Dagmar Gromann International Center For Computational Logic TU Dresden, 30 November 2018 Overview • Distributional Semantics • Distributed Semantics – Word Embeddings Dagmar Gromann, 30 November 2018 Semantic Computing 2 Distributional Semantics Dagmar Gromann, 30 November 2018 Semantic Computing 3 Distributional Semantics Definition • meaning of a word is the set of contexts in which it occurs • no other information is used than the corpus-derived information about word distribution in contexts (co-occurrence information of words) • semantic similarity can be inferred from proximity in contexts • At the very core: Distributional Hypothesis Distributional Hypothesis “similarity of meaning correlates with similarity of distribution” - Harris Z. S. (1954) “Distributional structure". Word, Vol. 10, No. 2-3, pp. 146-162 meaning = use = distribution in context => semantic distance Dagmar Gromann, 30 November 2018 Semantic Computing 4 Remember Lecture 1? Types of Word Meaning • Encyclopaedic meaning: words provide access to a large inventory of structured knowledge (world knowledge) • Denotational meaning: reference of a word to object/concept or its “dictionary definition” (signifier <-> signified) • Connotative meaning: word meaning is understood by its cultural or emotional association (positive, negative, neutral conntation; e.g. “She’s a dragon” in Chinese and English) • Conceptual meaning: word meaning is associated with the mental concepts it gives access to (e.g. prototype theory) • Distributional meaning: “You shall know a word by the company it keeps” (J.R. Firth 1957: 11) John Rupert Firth (1957). "A synopsis of linguistic theory 1930-1955." In Special Volume of the Philological Society. Oxford: Oxford University Press. Dagmar Gromann, 30 November 2018 Semantic Computing 5 Distributional Hypothesis in Practice Study by McDonald and Ramscar (2001): • The man poured from a balack into a handleless cup.
    [Show full text]
  • A Comparison of Word Embeddings and N-Gram Models for Dbpedia Type and Invalid Entity Detection †
    information Article A Comparison of Word Embeddings and N-gram Models for DBpedia Type and Invalid Entity Detection † Hanqing Zhou *, Amal Zouaq and Diana Inkpen School of Electrical Engineering and Computer Science, University of Ottawa, Ottawa ON K1N 6N5, Canada; [email protected] (A.Z.); [email protected] (D.I.) * Correspondence: [email protected]; Tel.: +1-613-562-5800 † This paper is an extended version of our conference paper: Hanqing Zhou, Amal Zouaq, and Diana Inkpen. DBpedia Entity Type Detection using Entity Embeddings and N-Gram Models. In Proceedings of the International Conference on Knowledge Engineering and Semantic Web (KESW 2017), Szczecin, Poland, 8–10 November 2017, pp. 309–322. Received: 6 November 2018; Accepted: 20 December 2018; Published: 25 December 2018 Abstract: This article presents and evaluates a method for the detection of DBpedia types and entities that can be used for knowledge base completion and maintenance. This method compares entity embeddings with traditional N-gram models coupled with clustering and classification. We tackle two challenges: (a) the detection of entity types, which can be used to detect invalid DBpedia types and assign DBpedia types for type-less entities; and (b) the detection of invalid entities in the resource description of a DBpedia entity. Our results show that entity embeddings outperform n-gram models for type and entity detection and can contribute to the improvement of DBpedia’s quality, maintenance, and evolution. Keywords: semantic web; DBpedia; entity embedding; n-grams; type identification; entity identification; data mining; machine learning 1. Introduction The Semantic Web is defined by Berners-Lee et al.
    [Show full text]
  • Distributional Semantics
    Distributional Semantics Computational Linguistics: Jordan Boyd-Graber University of Maryland SLIDES ADAPTED FROM YOAV GOLDBERG AND OMER LEVY Computational Linguistics: Jordan Boyd-Graber UMD Distributional Semantics 1 / 19 j j From Distributional to Distributed Semantics The new kid on the block Deep learning / neural networks “Distributed” word representations Feed text into neural-net. Get back “word embeddings”. Each word is represented as a low-dimensional vector. Vectors capture “semantics” word2vec (Mikolov et al) Computational Linguistics: Jordan Boyd-Graber UMD Distributional Semantics 2 / 19 j j From Distributional to Distributed Semantics This part of the talk word2vec as a black box a peek inside the black box relation between word-embeddings and the distributional representation tailoring word embeddings to your needs using word2vec Computational Linguistics: Jordan Boyd-Graber UMD Distributional Semantics 3 / 19 j j word2vec Computational Linguistics: Jordan Boyd-Graber UMD Distributional Semantics 4 / 19 j j word2vec Computational Linguistics: Jordan Boyd-Graber UMD Distributional Semantics 5 / 19 j j word2vec dog cat, dogs, dachshund, rabbit, puppy, poodle, rottweiler, mixed-breed, doberman, pig sheep cattle, goats, cows, chickens, sheeps, hogs, donkeys, herds, shorthorn, livestock november october, december, april, june, february, july, september, january, august, march jerusalem tiberias, jaffa, haifa, israel, palestine, nablus, damascus katamon, ramla, safed teva pfizer, schering-plough, novartis, astrazeneca, glaxosmithkline, sanofi-aventis, mylan, sanofi, genzyme, pharmacia Computational Linguistics: Jordan Boyd-Graber UMD Distributional Semantics 6 / 19 j j Working with Dense Vectors Word Similarity Similarity is calculated using cosine similarity: dog~ cat~ sim(dog~ ,cat~ ) = dog~ · cat~ jj jj jj jj For normalized vectors ( x = 1), this is equivalent to a dot product: jj jj sim(dog~ ,cat~ ) = dog~ cat~ · Normalize the vectors when loading them.
    [Show full text]
  • Facing the Facts of Fake: a Distributional Semantics and Corpus Annotation Approach Bert Cappelle, Pascal Denis, Mikaela Keller
    Facing the facts of fake: a distributional semantics and corpus annotation approach Bert Cappelle, Pascal Denis, Mikaela Keller To cite this version: Bert Cappelle, Pascal Denis, Mikaela Keller. Facing the facts of fake: a distributional semantics and corpus annotation approach. Yearbook of the German Cognitive Linguistics Association, De Gruyter, 2018, 6 (9-42). hal-01959609 HAL Id: hal-01959609 https://hal.archives-ouvertes.fr/hal-01959609 Submitted on 18 Dec 2018 HAL is a multi-disciplinary open access L’archive ouverte pluridisciplinaire HAL, est archive for the deposit and dissemination of sci- destinée au dépôt et à la diffusion de documents entific research documents, whether they are pub- scientifiques de niveau recherche, publiés ou non, lished or not. The documents may come from émanant des établissements d’enseignement et de teaching and research institutions in France or recherche français ou étrangers, des laboratoires abroad, or from public or private research centers. publics ou privés. Facing the facts of fake: a distributional semantics and corpus annotation approach Bert Cappelle, Pascal Denis and Mikaela Keller Université de Lille, Inria Lille Nord Europe Fake is often considered the textbook example of a so-called ‘privative’ adjective, one which, in other words, allows the proposition that ‘(a) fake x is not (an) x’. This study tests the hypothesis that the contexts of an adjective-noun combination are more different from the contexts of the noun when the adjective is such a ‘privative’ one than when it is an ordinary (subsective) one. We here use ‘embeddings’, that is, dense vector representations based on word co-occurrences in a large corpus, which in our study is the entire English Wikipedia as it was in 2013.
    [Show full text]
  • Sentiment, Stance and Applications of Distributional Semantics
    Sentiment, stance and applications of distributional semantics Maria Skeppstedt Sentiment, stance and applications of distributional semantics Sentiment analysis (opinion mining) • Aims at determining the attitude of the speaker/ writer * In a chunk of text (e.g., a document or sentence) * Towards a certain topic • Categories: * Positive, negative, (neutral) * More detailed Movie reviews • Excruciatingly unfunny and pitifully unromantic. • The picture emerges as a surprisingly anemic disappointment. • A sensitive, insightful and beautifully rendered film. • The spark of special anime magic here is unmistak- able and hard to resist. • Maybe you‘ll be lucky , and there‘ll be a power outage during your screening so you can get your money back. Methods for sentiment analysis • Train a machine learning classifier • Use lexicon and rules • (or combine the methods) Train a machine learning classifier Annotating text and training the model Excruciatingly Excruciatingly unfunny unfunny Applying the model pitifully unromantic. The model pitifully unromantic. Examples of machine learning Frameworks • Scikit-learn • NLTK (natural language toolkit) • CRF++ Methods for sentiment analysis • Train a machine learning classifier * See for instance Socher et al. (2013) that classified English movie review sentences into positive and negative sentiment. * Trained their classifier on 11,855 manually labelled sentences. (R. Socher, A. Perelygin, J. Wu, J. Chuang, C. D. Manning, A. Y. Ng, C. Potts, Recursive deep models for semantic composition- ality over a sentiment treebank) Methods based on lexicons and rules • Lexicons for different polarities • Sometimes with different strength on the words * Does not need the large amount of training data that is typically required for machine learning methods * More flexible (e.g, to add new sentiment poles and new languages) Commercial use of sentiment analysis I want to retrieve everything positive and negative that is said about the products I sell.
    [Show full text]
  • Building Web-Interfaces for Vector Semantic Models with the Webvectors Toolkit
    Building Web-Interfaces for Vector Semantic Models with the WebVectors Toolkit Andrey Kutuzov Elizaveta Kuzmenko University of Oslo Higher School of Economics Oslo, Norway Moscow, Russia [email protected] eakuzmenko [email protected] Abstract approaches, which allowed to train distributional models with large amounts of raw linguistic data We present WebVectors, a toolkit that very fast. The most established word embedding facilitates using distributional semantic algorithms in the field are highly efficient Continu- models in everyday research. Our toolkit ous Skip-Gram and Continuous Bag-of-Words, im- has two main features: it allows to build plemented in the famous word2vec tool (Mikolov web interfaces to query models using a et al., 2013b; Baroni et al., 2014), and GloVe in- web browser, and it provides the API to troduced in (Pennington et al., 2014). query models automatically. Our system Word embeddings represent the meaning of is easy to use and can be tuned according words with dense real-valued vectors derived from to individual demands. This software can word co-occurrences in large text corpora. They be of use to those who need to work with can be of use in almost any linguistic task: named vector semantic models but do not want to entity recognition (Siencnik, 2015), sentiment develop their own interfaces, or to those analysis (Maas et al., 2011), machine translation who need to deliver their trained models to (Zou et al., 2013; Mikolov et al., 2013a), cor- a large audience. WebVectors features vi- pora comparison (Kutuzov and Kuzmenko, 2015), sualizations for various kinds of semantic word sense frequency estimation for lexicogra- queries.
    [Show full text]
  • Distributional Semantics for Understanding Spoken Meal Descriptions
    DISTRIBUTIONAL SEMANTICS FOR UNDERSTANDING SPOKEN MEAL DESCRIPTIONS Mandy Korpusik, Calvin Huang, Michael Price, and James Glass∗ MIT Computer Science and Artificial Intelligence Laboratory, Cambridge, MA 02139, USA fkorpusik, calvinh, pricem, [email protected] ABSTRACT This paper presents ongoing language understanding experiments conducted as part of a larger effort to create a nutrition dialogue system that automatically extracts food concepts from a user’s spo- ken meal description. We first discuss the technical approaches to understanding, including three methods for incorporating word vec- tor features into conditional random field (CRF) models for seman- tic tagging, as well as classifiers for directly associating foods with properties. We report experiments on both text and spoken data from an in-domain speech recognizer. On text data, we show that the ad- dition of word vector features significantly improves performance, achieving an F1 test score of 90.8 for semantic tagging and 86.3 for food-property association. On speech, the best model achieves an F1 test score of 87.5 for semantic tagging and 86.0 for association. Fig. 1. A diagram of the flow of the nutrition system [4]. Finally, we conduct an end-to-end system evaluation through a user study with human ratings of 83% semantic tagging accuracy. Index Terms— CRF, Word vectors, Semantic tagging • We demonstrate a significant improvement in semantic tag- ging performance by adding word embedding features to a classifier. We explore features for both the raw vector values 1. INTRODUCTION and similarity to prototype words from each category. In ad- dition, improving semantic tagging performance benefits the For many patients with obesity, diet tracking is often time- subsequent task of associating foods with properties.
    [Show full text]
  • Distributional Semantics Today Introduction to the Special Issue Cécile Fabre, Alessandro Lenci
    Distributional Semantics Today Introduction to the special issue Cécile Fabre, Alessandro Lenci To cite this version: Cécile Fabre, Alessandro Lenci. Distributional Semantics Today Introduction to the special issue. Revue TAL, ATALA (Association pour le Traitement Automatique des Langues), 2015, Sémantique distributionnelle, 56 (2), pp.7-20. hal-01259695 HAL Id: hal-01259695 https://hal.archives-ouvertes.fr/hal-01259695 Submitted on 26 Jan 2016 HAL is a multi-disciplinary open access L’archive ouverte pluridisciplinaire HAL, est archive for the deposit and dissemination of sci- destinée au dépôt et à la diffusion de documents entific research documents, whether they are pub- scientifiques de niveau recherche, publiés ou non, lished or not. The documents may come from émanant des établissements d’enseignement et de teaching and research institutions in France or recherche français ou étrangers, des laboratoires abroad, or from public or private research centers. publics ou privés. Distributional Semantics Today Introduction to the special issue Cécile Fabre* — Alessandro Lenci** * University of Toulouse, CLLE-ERSS ** University of Pisa, Computational Linguistics Laboratory ABSTRACT. This introduction to the special issue of the TAL journal on distributional seman- tics provides an overview of the current topics of this field and gives a brief summary of the contributions. RÉSUMÉ. Cette introduction au numéro spécial de la revue TAL consacré à la sémantique dis- tributionnelle propose un panorama des thèmes de recherche actuels dans ce champ et fournit un résumé succinct des contributions acceptées. KEYWORDS: Distributional Semantic Models, vector-space models, corpus-based semantics, se- mantic proximity, semantic relations, evaluation of distributional ressources. MOTS-CLÉS : Sémantique distributionnelle, modèles vectoriels, proximité sémantique, relations sémantiques, évaluation de ressources distributionnelles.
    [Show full text]
  • Distributional Semantics
    Distributional semantics Distributional semantics is a research area that devel- by populating the vectors with information on which text ops and studies theories and methods for quantifying regions the linguistic items occur in; paradigmatic sim- and categorizing semantic similarities between linguis- ilarities can be extracted by populating the vectors with tic items based on their distributional properties in large information on which other linguistic items the items co- samples of language data. The basic idea of distributional occur with. Note that the latter type of vectors can also semantics can be summed up in the so-called Distribu- be used to extract syntagmatic similarities by looking at tional hypothesis: linguistic items with similar distributions the individual vector components. have similar meanings. The basic idea of a correlation between distributional and semantic similarity can be operationalized in many dif- ferent ways. There is a rich variety of computational 1 Distributional hypothesis models implementing distributional semantics, includ- ing latent semantic analysis (LSA),[8] Hyperspace Ana- The distributional hypothesis in linguistics is derived logue to Language (HAL), syntax- or dependency-based from the semantic theory of language usage, i.e. words models,[9] random indexing, semantic folding[10] and var- that are used and occur in the same contexts tend to ious variants of the topic model. [1] purport similar meanings. The underlying idea that “a Distributional semantic models differ primarily with re- word is characterized by the company it keeps” was pop- spect to the following parameters: ularized by Firth.[2] The Distributional Hypothesis is the basis for statistical semantics. Although the Distribu- • tional Hypothesis originated in linguistics,[3] it is now re- Context type (text regions vs.
    [Show full text]
  • Arguments For: Semantic Folding and Hierarchical Temporal Memory
    [email protected] Mariahilferstrasse 4 1070 Vienna, Austria http://www.cortical.io ARGUMENTS FOR: SEMANTIC FOLDING AND HIERARCHICAL TEMPORAL MEMORY Replies to the most common Semantic Folding criticisms Francisco Webber, September 2016 INTRODUCTION 2 MOTIVATION FOR SEMANTIC FOLDING 3 SOLVING ‘HARD’ PROBLEMS 3 ON STATISTICAL MODELING 4 FREQUENT CRITICISMS OF SEMANTIC FOLDING 4 “SEMANTIC FOLDING IS JUST WORD EMBEDDING” 4 “WHY NOT USE THE OPEN SOURCED WORD2VEC?” 6 “SEMANTIC FOLDING HAS NO COMPARATIVE EVALUATION” 7 “SEMANTIC FOLDING HAS ONLY LOGICAL/PHILOSOPHICAL ARGUMENTS” 7 FREQUENT CRITICISMS OF HTM THEORY 8 “JEFF HAWKINS’ WORK HAS BEEN MOSTLY PHILOSOPHICAL AND LESS TECHNICAL” 8 “THERE ARE NO MATHEMATICALLY SOUND PROOFS OR VALIDATIONS OF THE HTM ASSERTIONS” 8 “THERE IS NO REAL-WORLD SUCCESS” 8 “THERE IS NO COMPARISON WITH MORE POPULAR ALGORITHMS OF DEEP LEARNING” 8 “GOOGLE/DEEP MIND PRESENT THEIR PERFORMANCE METRICS, WHY NOT NUMENTA” 9 FMRI SUPPORT FOR SEMANTIC FOLDING THEORY 9 REALITY CHECK: BUSINESS APPLICABILITY 9 GLOBAL INDUSTRY UNDER PRESSURE 10 “WE HAVE TRIED EVERYTHING …” 10 REFERENCES 10 Arguments for Semantic Folding and Hierarchical Temporal Memory Theory Introduction During the last few years, the big promise of computer science, to solve any computable problem given enough data and a gold standard, has triggered a race for machine learning (ML) among tech communities. Sophisticated computational techniques have enabled the tackling of problems that have been considered unsolvable for decades. Face recognition, speech recognition, self- driving cars; it seems like a computational model could be created for any human task. Anthropology has taught us that a sufficiently complex technology is indistinguishable from magic by the non-expert, indigenous mind.
    [Show full text]
  • A Vector Space for Distributional Semantics for Entailment ∗
    A Vector Space for Distributional Semantics for Entailment ∗ James Henderson and Diana Nicoleta Popa Xerox Research Centre Europe [email protected] and [email protected] unk f g f Abstract ⇒ ¬ unk 1 0 0 0 Distributional semantics creates vector- f 1 1 0 0 space representations that capture many g 1 0 1 0 forms of semantic similarity, but their re- f 1 0 0 1 ¬ lation to semantic entailment has been less clear. We propose a vector-space model Table 1: Pattern of logical entailment between which provides a formal foundation for a nothing known (unk), two different features f and g known, and the complement of f ( f) known. distributional semantics of entailment. Us- ¬ ing a mean-field approximation, we de- velop approximate inference procedures ness with a distributional-semantic model of hy- and entailment operators over vectors of ponymy detection. probabilities of features being known (ver- Unlike previous vector-space models of entail- sus unknown). We use this framework ment, the proposed framework explicitly models to reinterpret an existing distributional- what information is unknown. This is a crucial semantic model (Word2Vec) as approxi- property, because entailment reflects what infor- mating an entailment-based model of the mation is and is not known; a representation y en- distributions of words in contexts, thereby tails a representation x if and only if everything predicting lexical entailment relations. In that is known given x is also known given y. Thus, both unsupervised and semi-supervised we model entailment in a vector space where each experiments on hyponymy detection, we dimension represents something we might know.
    [Show full text]
  • Embeddings in Natural Language Processing
    Embeddings in Natural Language Processing Theory and Advances in Vector Representation of Meaning Mohammad Taher Pilehvar Tehran Institute for Advanced Studies Jose Camacho-Collados Cardiff University SYNTHESIS LECTURESDRAFT ON HUMAN LANGUAGE TECHNOLOGIES M &C Morgan& cLaypool publishers ABSTRACT Embeddings have been one of the dominating buzzwords since the early 2010s for Natural Language Processing (NLP). Encoding information into a low-dimensional vector representation, which is easily integrable in modern machine learning algo- rithms, has played a central role in the development in NLP. Embedding techniques initially focused on words but the attention soon started to shift to other forms: from graph structures, such as knowledge bases, to other types of textual content, such as sentences and documents. This book provides a high level synthesis of the main embedding techniques in NLP, in the broad sense. The book starts by explaining conventional word vector space models and word embeddings (e.g., Word2Vec and GloVe) and then moves to other types of embeddings, such as word sense, sentence and document, and graph embeddings. We also provide an overview on the status of the recent development in contextualized representations (e.g., ELMo, BERT) and explain their potential in NLP. Throughout the book the reader can find both essential information for un- derstanding a certain topic from scratch, and an in-breadth overview of the most successful techniques developed in the literature. KEYWORDS Natural Language Processing, Embeddings, Semantics DRAFT iii Contents 1 Introduction ................................................. 1 1.1 Semantic representation . 3 1.2 One-hot representation . 4 1.3 Vector Space Models . 5 1.4 The Evolution Path of representations .
    [Show full text]