Phone Merger Specification for Multilingual ASR: The Motorola Polyphone Network Lynette Melnar and Jim Talley Motorola Human Interface Labs, Voice Dialog Systems Lab, USA E-mail:
[email protected],
[email protected] ABSTRACT 2. MOTPOLY: OVERVIEW This paper describes the Motorola Polyphone Network MotPoly is a knowledge-based, hierarchically arranged (MotPoly), a hierarchical, universal phone phone merger specification network for shared correspondence network that defines allowable phone multilingual and multi-dialect acoustic modeling. The mergers for shared acoustic modeling in multilingual and organization of MotPoly is language-independent and multi-dialect automatic speech recognition (ML-ASR). phone merger is defined by an internally ranked system of MotPoly’s organization is defined by phonetic similarity relative phonetic similarity, phonological contrastiveness, and other language-independent phonological factors. and phone frequency. MotPoly functions in ML-ASR as a Unlike other approaches to shared acoustic modeling, phone merger framework that constrains data-driven MotPoly can be effectively used in systems where acoustic modeling strategies. Because MotPoly is not computational resources are limited, such as portable biased toward any particular language, language family, devices. Furthermore, it is less constrained by language or language type, it is a universal, static definition of data availability than other approaches. With MotPoly as phone merger and can be used unmodified to specify part of an overall strategy, Motorola’s Voice Dialog likely mergers for an unlimited number of phones from Systems Lab’s ML-ASR team was able to define a set of any collection of languages or dialects. multilingual acoustic models whose size was only 23% of the largest monolingual model set but whose overall MotPoly is of particular value in applications that are performance was higher than the monolingual models by restricted in computational and language resources.