Augmented Javanese Speech Levels Machine Translation 45

Total Page:16

File Type:pdf, Size:1020Kb

Augmented Javanese Speech Levels Machine Translation 45 Augmented Javanese Speech Levels Machine Translation 45 Augmented Javanese Speech Levels Machine Translation Aji P. Wibawa1, Andrew Nafalski2, and Wayan F. Mahmudy3, Non-members ABSTRACT language [1, 7, 8]. Furthermore, the selection of incor- rect vocabularies [4] indicates that they lack mastery This paper presents the development of the hy- of speech levels and do not know how to use them brid corpus-based machine translation for Javanese appropriately in verbal communication. In fact, the language. The system is designed to deal with the acquisition of speech levels among teenagers can be complexity of politeness expression and speech levels classified as very poor: 36.45 out of 100. This finding of Javanese that is considered as a local language with was revealed in a research on the use of speech levels the biggest number of users in Indonesia. Statistical by youngsters in Solo [8] and the result was based on features are embedded to increase the performance written vocabulary translation tests. of the system. The edit shifting distance is applied Realizing that they cannot handle this polite form, due to increase the alignment efficiency. However, im- younger speakers usually switch into Indonesian lan- proper alignment contributed by recorded impossible guage(bahasa Indonesia), which they can handle more pair and insufficient data training is still detected. easily and they believe to be more reliable to use in This paper proposes a new improvement of the de- the global era [7-9]. If this continues, the krama form- veloped alignment algorithm based on the impossible a unique characteristic Javanese-is in the danger of pair restriction. Based on experimental results, the diminishing. In addition, unskilled educators and the new developed algorithm is more accurate (A=93.8%) lack of speech levels' guides may be exacerbating the even though the number of training data is less than problem [10, 11]. While teachers are expected to serve the old one (A=87.9%). as language models at school [1], some of them use inappropriate words and levels [4] - a fact which fur- Keywords: Javanese Speech Levels, Corpus, Ma- ther suggests that Javanese speech levels dying out. chine Translation, Impossible Pair Limitation Hence, a machine translation should be developed to protect the Javanese speech levels from being extinct. 1. INTRODUCTION A machine translator has been developed in order Respecting others properly by using paralinguis- to protect the existence of the speech level [12, 13]. tic forms and proper speech is part of the politeness The translator provides bilingual pragmatic transla- in Javanese culture. The appropriate degree of po- tion between speech levels [12] and also Indonesian liteness is often articulated in the form of deference [13]. The translation knowledge is based on a bi-text in communication. Speech levels [1, 2], speech styles alignment that reinforced by edit shifting distance al- [3] or Javanese language politeness [4] are frequently gorithm [14]. The results are impressive; however, used to name the degree of deference. word repetition [13] and pragmatic translation mis- The levels of speech and associated politeness takes [12] are detected during the translation. forms are being neglected with fewer speakers being This paper is an extended version of [12], focused conversant in them even though Javanese is consid- on modifying the bilingual text alignment due to ered as the most widely used regional language in In- avoid the unwanted mistakes as well as increasing donesia [5, 6].Negative tendency is detected concern- the translation accuracy. The novel algorithm will be ing the use of Javanese speech levels among teenagers. compared to the previous model with extra training They may select an incorrect speech level to address a data. high-status person since they are unable to transform the local politeness value into its equivalent refined 2. THEJAVANESE SPEECH LEVELS Manuscript received on March 31, 2013 ; revised on June 26, Javanese linguists [2, 15, 16] divide the speech 2013. levels into three classes: krama, madya and ngoko. Final manuscript received August 10, 2013. The classes can be further classified into nine 1 The author is with Department of Electrical Engineer- sub-levels: mudha-krama (MK), kramantara (KA), ing, State University of Malang, Indonesia., E-mail: ajipw@ um.ac.id wredha-krama (WK), madya-krama (MdK), madyan- 2 The author is with School of Engineering, Uni- tara (Md A), madya-ngoko (Md Ng), basa-antya versity of South Australia, Australia., E-mail: An- (BA), antya-basa (AB), ngoko-lugu (Ng L). The sub- [email protected] 3 The author is with Department of Computer Science, Braw- levels has been simplified into four categories; ngoko ijaya University, Indonesia ., E-mail: [email protected] (Ng), ngoko alus (NgA), krama (Kr) and krama alus 46 ECTI TRANSACTIONS ON COMPUTER AND INFORMATION TECHNOLOGY VOL.8, NO.1 May 2014 (KrA) in the first Javanese Congress in 1991[4]. The- ample of alignment when translating bapakku man- example of simplified of Javanese speech levels and gan ana omahku into bapak kula dhahar wonten its lexical characteristics is shown in Table 1. griya kula is shown in Figure 1. The top part shows direct alignmentthat produces inaccurate translation Table 1: Example of Simplified Speech Levels and whereas the bottom part shows the correct alignment. Lexical Features. Fig.1: Example of Direct and Correct Javanese Speech Level Alignment. The Javanese sentence has a basic structure of SVO (subject, verb, and object). The meaning of sen- tences is pragmatically diverse and related to subject and verb agreement (SVA). The SVA derives from non-linguistic factors (i.e. social status, ages and re- lationships) [17]. For instance, sentences (1) and (2) show the differences of SVA based onthe social status of the subjects where both are operating at krama level. (1) Murid-muridsawegnedha (students are eat- ing). (2) Guru-guru sawegdhahar (teachers are eating). The underlined words in both sentences have a same meaning (eating). However, the first sentence must use the word nedha because the subject of the first sentence is students who obviously have lower social status than their teachers in sentence (2). In consequence, the word nedha is used instead of dha- har, which is only suitable for honoured people. 3. REVIEW OF MACHINE TRANSLATION Machine translation (MT), a branch of computa- One-to-one word translation in Javanese may pro- tional linguistics, is simply defined as the automatic duce poor, inaccurate and inappropriate results [14] computer-assisted translation of bilingual or multilin- since one word may be literally translated into two gual natural languages[18]. Based on its knowledge words at another speech level, and vice versa.As a base, MT is classified into two categories: Rule-based result, the source and target sentences may consist Machine Translation (RBMT) and Corpus-based Ma- of an unequal number of words, known as asym- chine Translation (CBMT). metric formation phenomenon. For example, the RBMT uses linguistic information such as seman- sentence awake dhewe ditimbali ibumu (Ng) is tic, morphological, and syntactic information as its translated as kita dipuntimbali ibu panjenen- knowledge base. The involvement of linguists in gan (Kr). Both sentences have four words; however, knowledge base development is unavoidable. They translating the words one-to-one based only their or- also build transfer rules between languages that are der in the sentence, produces an inaccurate transla- relatively complex and time consuming, especially for tion. In fact, the sentences are composed of three those with very different structures. As a result, the asymmetric pairs (i.e. awake dhewe$kita; ditim- development costs increases in an effort to achieve the bali?dipuntimbali; ibumu$ibu panjenengan). An ex- required quality threshold. The participation of lan- Augmented Javanese Speech Levels Machine Translation 47 guage experts is desperately needed when more adap- tive RBMT is desired. The linguists have to retrain the developed RBMT by adding new rules, vocabu- lary and other linguistic information, and this con- tributes further to extra development time and costs. On the other hand, the knowledge of CBMT is based on large sets of bilingual text that are recog- nised as corpora[19, 20]. The two kinds of CBMT are example-based machine translation (EBMT) and statistical machine translation (SMT). The availabil- ity of corpora is a key factor in reducing development costs. Once available, the development time is re- duced in parallel with the translation development cost. In contrast with RBMT, the role of linguists in the development of CBMT is purely optional. When Fig.2: The Design of Javanese Machine Transla- a new instance of translation arises and unknown tion. words occur, the corpus may be updated automat- ically. While RBMT is highly efficient for more gen- translation process uses the texts in both source lan- eral translation, CBMT may work better for a specific guage (S) and target language (T ). Each text is di- domain. Although some inconsistency may be found, vided into sentences (s and t ) and can be formu- for example, in pure SMT technique [21, 22] RBMT i j lated as in (1) and (2). Here, n and n represent generates more natural translation than CBMT. i j the number of sentences in the source and target text Both SMT and EBMT derive mainly from large respectively. bilingual texts; that is known as a corpus (i.e. corpora in the plural form) [19, 20]. While SMT focuses more on word combinations and their occurrence, EBMT S = fs : Sj0 ≤ i < n g (1) focuses on text segmentation [20], phrase memorisa- i i f j ≤ g tion [22]and analogical sentence recombination [19]. T = tj : T 0 i < nj (2) Moreover, SMT has a more obvious definition than Afterwards, all correspondence sentences are EBMT in terms of selecting the most appropriate parsed into a set of words (w and w ) from the first translation.
Recommended publications
  • From Arabic Style Toward Javanese Style: Comparison Between Accents of Javanese Recitation and Arabic Recitation
    From Arabic Style toward Javanese Style: Comparison between Accents of Javanese Recitation and Arabic Recitation Nur Faizin1 Abstract Moslem scholars have acceptedmaqamat in reciting the Quran otherwise they have not accepted macapat as Javanese style in reciting the Quran such as recitationin the State Palace in commemoration of Isra` Miraj 2015. The paper uses a phonological approach to accents in Arabic and Javanese style in recitingthe first verse of Surah Al-Isra`. Themethod used here is analysis of suprasegmental sound (accent) by usingSpeech Analyzer programand the comparison of these accents is analyzed by descriptive method. By doing so, the author found that:first, there is not any ideological reason to reject Javanese style because both of Arabic and Javanese style have some aspects suitable and unsuitable with Ilm Tajweed; second, the suitability of Arabic style was muchthan Javanese style; third, it is not right to reject recitingthe Quran with Javanese style only based on assumption that it evokedmistakes and errors; fourth, the acceptance of Arabic style as the art in reciting the Quran should risedacceptanceof the Javanese stylealso. So, rejection of reciting the Quranwith Javanese style wasnot due to any reason and it couldnot be proofed by any logical argument. Keywords: Recitation, Arabic Style, Javanese Style, Quran. Introduction There was a controversial event in commemoration of Isra‘ Mi‘raj at the State Palacein Jakarta May 15, 2015 ago. The recitation of the Quran in the commemoration was recitedwithJavanese style (langgam).That was not common performance in relation to such as official event. Muhammad 58 Nur Faizin, From Arabic Style toward Javanese Style Yasser Arafat, a lecture of Sunan Kalijaga State Islamic University Yogyakarta has been reciting first verse of Al-Isra` by Javanese style in the front of state officials and delegationsof many countries.
    [Show full text]
  • Ka И @И Ka M Л @Л Ga Н @Н Ga M М @М Nga О @О Ca П
    ISO/IEC JTC1/SC2/WG2 N3319R L2/07-295R 2007-09-11 Universal Multiple-Octet Coded Character Set International Organization for Standardization Organisation Internationale de Normalisation Международная организация по стандартизации Doc Type: Working Group Document Title: Proposal for encoding the Javanese script in the UCS Source: Michael Everson, SEI (Universal Scripts Project) Status: Individual Contribution Action: For consideration by JTC1/SC2/WG2 and UTC Replaces: N3292 Date: 2007-09-11 1. Introduction. The Javanese script, or aksara Jawa, is used for writing the Javanese language, the native language of one of the peoples of Java, known locally as basa Jawa. It is a descendent of the ancient Brahmi script of India, and so has many similarities with modern scripts of South Asia and Southeast Asia which are also members of that family. The Javanese script is also used for writing Sanskrit, Jawa Kuna (a kind of Sanskritized Javanese), and Kawi, as well as the Sundanese language, also spoken on the island of Java, and the Sasak language, spoken on the island of Lombok. Javanese script was in current use in Java until about 1945; in 1928 Bahasa Indonesia was made the national language of Indonesia and its influence eclipsed that of other languages and their scripts. Traditional Javanese texts are written on palm leaves; books of these bound together are called lontar, a word which derives from ron ‘leaf’ and tal ‘palm’. 2.1. Consonant letters. Consonants have an inherent -a vowel sound. Consonants combine with following consonants in the usual Brahmic fashion: the inherent vowel is “killed” by the PANGKON, and the follow- ing consonant is subjoined or postfixed, often with a change in shape: §£ ndha = § NA + @¿ PANGKON + £ DA-MAHAPRANA; üù n.
    [Show full text]
  • Proposal for a Gurmukhi Script Root Zone Label Generation Ruleset (LGR)
    Proposal for a Gurmukhi Script Root Zone Label Generation Ruleset (LGR) LGR Version: 3.0 Date: 2019-04-22 Document version: 2.7 Authors: Neo-Brahmi Generation Panel [NBGP] 1. General Information/ Overview/ Abstract This document lays down the Label Generation Ruleset for Gurmukhi script. Three main components of the Gurmukhi Script LGR i.e. Code point repertoire, Variants and Whole Label Evaluation Rules have been described in detail here. All these components have been incorporated in a machine-readable format in the accompanying XML file named "proposal-gurmukhi-lgr-22apr19-en.xml". In addition, a document named “gurmukhi-test-labels-22apr19-en.txt” has been provided. It provides a list of labels which can produce variants as laid down in Section 6 of this document and it also provides valid and invalid labels as per the Whole Label Evaluation laid down in Section 7. 2. Script for which the LGR is proposed ISO 15924 Code: Guru ISO 15924 Key N°: 310 ISO 15924 English Name: Gurmukhi Latin transliteration of native script name: gurmukhī Native name of the script: ਗੁਰਮੁਖੀ Maximal Starting Repertoire [MSR] version: 4 1 3. Background on Script and Principal Languages Using It 3.1. The Evolution of the Script Like most of the North Indian writing systems, the Gurmukhi script is a descendant of the Brahmi script. The Proto-Gurmukhi letters evolved through the Gupta script from 4th to 8th century, followed by the Sharda script from 8th century onwards and finally adapted their archaic form in the Devasesha stage of the later Sharda script, dated between the 10th and 14th centuries.
    [Show full text]
  • M. Ricklefs an Inventory of the Javanese Manuscript Collection in the British Museum
    M. Ricklefs An inventory of the Javanese manuscript collection in the British Museum In: Bijdragen tot de Taal-, Land- en Volkenkunde 125 (1969), no: 2, Leiden, 241-262 This PDF-file was downloaded from http://www.kitlv-journals.nl Downloaded from Brill.com09/29/2021 11:29:04AM via free access AN INVENTORY OF THE JAVANESE MANUSCRIPT COLLECTION IN THE BRITISH MUSEUM* he collection of Javanese manuscripts in the British Museum, London, although small by comparison with collections in THolland and Indonesia, is nevertheless of considerable importance. The Crawfurd collection, forming the bulk of the manuscripts, provides a picture of the types of literature being written in Central Java in the late eighteenth and early nineteenth centuries, a period which Dr. Pigeaud has described as a Literary Renaissance.1 Because they were acquired by John Crawfurd during his residence as an official of the British administration on Java, 1811-1815, these manuscripts have a convenient terminus ad quem with regard to composition. A large number of the items are dated, a further convenience to the research worker, and the dates are seen to cluster in the four decades between AD 1775 and AD 1815. A number of the texts were originally obtained from Pakualam I, who was installed as an independent Prince by the British admini- stration. Some of the manuscripts are specifically said to have come from him (e.g. Add. 12281 and 12337), and a statement in a Leiden University Bah ad from the Pakualaman suggests many other volumes in Crawfurd's collection also derive from this source: Tuwan Mister [Crawfurd] asked to be instructed in adat law, with examples of the Javanese usage.
    [Show full text]
  • The Local Wisdom in Javanese Thinking Culture Within Hanacaraka Philosophy
    THE LOCAL WISDOM IN JAVANESE THINKING CULTURE WITHIN HANACARAKA PHILOSOPHY Fitriana Kartika Sari STKIP PGRI Ponorogo email: [email protected] Abstract (Title: The Local Wisdom In Javanese Thinking Culture Within Hanacaraka Philosophy). Hanacaraka refers to the written language of Javanese, Javanese script. The name is based on the order of the twenty-Javanese-script, which is started by the series of hanacaraka script. This study aims to describe the local wisdom content of Javanese Thinking culture within the philosophical meaning of hanacaraka. This study used a semiotic approach applied to the main data source, namely Sêrat Kérata Basa. The results showed that (1) the philosophical meaning of the hanacaraka script was fully equipped with the wisdom of Javanese thinking culture in the form of symbolism and othak-athik mathuk which is supported by deep contemplation which involves all the common senses, including the inner sensitivity; (2) hanacaraka philosophy depicts God creature’s (human) circumstances, which is equipped with creation, feeling, and effort; human cannot avoid all God’s fate upon them until the end of their life; human conditions that achieve life harmony due to the ability to unify the manifestation of God and human characteristics, which heats up within themselves; human condition which put God’s order highly by doing all the orders and avoiding all of His warnings. Keywords: local wisdom, Javanese thinking, Hanacaraka INTRODUCTION because human is always eager to know and Javanese local wisdom is the principal find the authentic truth for everything. Human factor of thoughtful thinking which covers tries to seek for the revelating indicator of all then aspects of knowledge, faith, mysteries deeply through the use of all the comprehension or understanding, customs, common senses to gain satisfying conclusive and ethics.
    [Show full text]
  • Design of Javanese Text to Speech Application
    Design of Javanese Text to Speech Application Yulia, Liliana, Rudy Adipranata, Gregorius Satia Budhi Informatics Department, Industrial Technology Faculty, Petra Christian University Surabaya, Indonesia [email protected] Abstract—Javanese is one of the many regional languages used in Indonesia. Javanese language is used by most of the population in Java. But now along with the development of the era, the use of regional languages including Javanese language is to be re- duced especially among the younger generation. One way to help conserve the use of Javanese language is to utilize information technologies, one of them is by developing a text to speech appli- cation that can be used to find out how the pronunciation of Ja- vanese language. In this paper, we discussed the design for Java- nese text to speech applications uses finite state automata. The design result will be used as rules to separate syllables when im- plementing text to speech application. Index Terms—Javanese language; Finite state automata; Text to speech. Figure 1: Basic Javanese characters I. INTRODUCTION In addition to the basic characters, the Javanese character Javanese language is a language widely spoken by the peo- has supplementary characters, consist of symbols for express- ple of Java. It is one of the regional languages of many region- ing vowels as well as a combination of two specific conso- al languages spoken in Indonesia. As one of the assets of na- nants. This supplementary characters is called sandhangan tional culture, Javanese language needs to be preserved. The and can be seen in Figure 2 [5]. younger generation is now more interested in learning a for- Symbol Example Read eign language, rather than the native Indonesian local lan- guage.
    [Show full text]
  • Ahom Range: 11700–1174F
    Ahom Range: 11700–1174F This file contains an excerpt from the character code tables and list of character names for The Unicode Standard, Version 14.0 This file may be changed at any time without notice to reflect errata or other updates to the Unicode Standard. See https://www.unicode.org/errata/ for an up-to-date list of errata. See https://www.unicode.org/charts/ for access to a complete list of the latest character code charts. See https://www.unicode.org/charts/PDF/Unicode-14.0/ for charts showing only the characters added in Unicode 14.0. See https://www.unicode.org/Public/14.0.0/charts/ for a complete archived file of character code charts for Unicode 14.0. Disclaimer These charts are provided as the online reference to the character contents of the Unicode Standard, Version 14.0 but do not provide all the information needed to fully support individual scripts using the Unicode Standard. For a complete understanding of the use of the characters contained in this file, please consult the appropriate sections of The Unicode Standard, Version 14.0, online at https://www.unicode.org/versions/Unicode14.0.0/, as well as Unicode Standard Annexes #9, #11, #14, #15, #24, #29, #31, #34, #38, #41, #42, #44, #45, and #50, the other Unicode Technical Reports and Standards, and the Unicode Character Database, which are available online. See https://www.unicode.org/ucd/ and https://www.unicode.org/reports/ A thorough understanding of the information contained in these additional sources is required for a successful implementation.
    [Show full text]
  • 2016 Semi Finalists Medals
    2016 US Physics Olympiad Semi Finalists Medal Rankings StudentMedal School City State Abbott, Ryan WHopkinsBronze Medal SchoolNew Haven CT Alton, James SLakesideHonorable Mention High SchoolEvans GA ALUMOOTIL, VARKEY TCanyonHonorable Mention Crest AcademySan Diego CA An, Seung HwanGold Medal Taft SchoolWatertown CT Ashary, Rafay AWilliamHonorable Mention P Clements High SchoolSugar Land TX Balaji, ShreyasSilver Medal John Foster Dulles High SchoolSugar Land TX Bao, MikeGold Medal Cambridge Educational InstituteChino Hills CA Beasley, NicholasGold Medal Stuyvesant High SchoolNew York NY BENABOU, JOSHUA N Gold Medal Plandome NY Bhattacharyya, MoinakSilver Medal Lynbrook High SchoolSan Jose CA Bhattaram, Krishnakumar SLynbrookBronze Medal High SchoolSan Jose CA Bhimnathwala, Tarung SBronze Medal Manalapan High SchoolManalapan NJ Boopathy, AkhilanGold Medal Lakeside Upper SchoolSeattle WA Cao, AntonSilver Medal Evergreen Valley High SchoolSan Jose CA Cen, Edward DBellaireHonorable Mention High SchoolBellaire TX Chadraa, Dalai BRedmondHonorable Mention High SchoolRedmond WA Chakrabarti, DarshanBronze Medal Northside College Preparatory HSChicago IL Chan, Clive ALexingtonSilver Medal High SchoolLexington MA Chang, Kevin YBellarmineSilver Medal Coll PrepSan Jose CA Cheerla, NikhilBronze Medal Monta Vista High SchoolSan Jose CA Chen, AlexanderSilver Medal Princeton High SchoolPrinceton NJ Chen, Andrew LMissionSilver Medal San Jose High SchoolFremont CA Chen, Benjamin YArdentSilver Medal Academy for Gifted YouthIrvine CA Chen, Bryan XMontaHonorable
    [Show full text]
  • 1 PASANGAN DAN SANDHANGAN DALAM AKSARA JAWA Oleh
    PASANGAN DAN SANDHANGAN DALAM AKSARA JAWA1 oleh: Sri Hertanti Wulan [email protected] Jurusan Pendidikan Bahasa Daerah FBS UNY Aksara nglegena yang digunakan dalam ejaan bahasa Jawa pada dasarnya terdiri atas dua puluh aksara yang bersifat silabik Darusuprapta (2002: 12-13). Masing-masing aksara mempunyai aksara pasangan, yakni aksara yang berfungsi untuk menghubungkan suku kata mati/tertutup dengan suku kata berikutnya, kecuali suku kata yang tertutup dengan wignyan (.. ), layar (....), dan cecak (....). Pasangan – pasangan tersebut antara lain: a) Aksara pasangan wutuh, ditulis di bawah aksara yang diberi pasangan dan tidak disambung, yaitu antara lain: Tabel 1 Pasangan Wutuh pasangan Wujud Contoh Ra … dalan ramé= Ya … tumbas yuyu = Ga … dalan gĕdhé = .… dolan ngomah = nga b) Aksara pasangan tugelan ditulis di belakang aksara yang diberi pasangan dan tidak disambungkan dengan aksara yang diberi pasangan, yaitu antara lain: 1Disampaikan dalam PPM PelatihanAksaraJawadan PendirianHanacaraka Centre sebagaiRevitalisasiFungsiAksaraJawa kerjasama FBS UNY dan Dinas Dikpora DIY. Dilaksanakan di Dikpora DIY, Senin 28 Oktober 2013. 1 Tabel 2.1 Pasangan Tugelan pasangan Wujud Contoh ha … adhĕm hawané = pa … bakul pĕlĕm = sa … dalan sĕpi = c) Aksara pasangan tugelan ditulis di bawah aksara yang diberi pasangan dan tidak disambungkan dengan aksara yang diberi pasangan, yaitu antara lain: Tabel 2.2 Pasangan Tugelan pasangan Wujud Contoh kilèn kalen= ka … wis takon = ta … tas larang = la … Pasangan – pasangan tersebut, bila mendapatkan sandhangan
    [Show full text]
  • Source Readings in Javanese Gamelan and Vocal Music, Volume 2
    THE UNIVERSITY OF MICHIGAN CENTER FOR SOUTH AND SOUTHEAST ASIAN STUDIES MICHIGAN PAPERS ON SOUTH AND SOUTHEAST ASIA Editorial Board Alton L. Becker Karl L. Hutterer John K. Musgrave Peter E. Hook, Chairman Ann Arbor, Michigan USA KARAWITAN SOURCE READINGS IN JAVANESE GAMELAN AND VOCAL MUSIC Judith Becker editor Alan H. Feinstein assistant editor Hardja Susilo Sumarsam A. L. Becker consultants Volume 2 MICHIGAN PAPERS ON SOUTH AND SOUTHEAST ASIA; Center for South and Southeast Asian Studies The University of Michigan Number 30 Open access edition funded by the National Endowment for the Humanities/ Andrew W. Mellon Foundation Humanities Open Book Program. Library of Congress Catalog Card Number: 82-72445 ISBN 0-89148-034-X Copyright ©' 1987 by Center for South and Southeast Asian Studies The University of Michigan Publication of this book was assisted in part by a grant from the Publications Program of the National Endowment for the Humanities. Additional funding or assistance was provided by the National Endowment for the Humanities (Transla- tions); the Southeast Asia Regional Council, Association for Asian Studies; The Rackham School of Graduate Studies, The University of Michigan; and the School of Music, The University of Michigan. Printed in the United States of America ISBN 978-0-89-148034-1 (hardcover) ISBN 978-0-47-203819-0 (paper) ISBN 978-0-47-212769-6 (ebook) ISBN 978-0-47-290165-4 (open access) The text of this book is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License: https://creativecommons.org/licenses/by-nc-nd/4.0/ CONTENTS PREFACE: TRANSLATING THE ART OF MUSIC A.
    [Show full text]
  • ISO/IEC JTC1/SC2/WG2 N 2005 Date: 1999-05-29
    ISO INTERNATIONAL ORGANIZATION FOR STANDARDIZATION ORGANISATION INTERNATIONALE DE NORMALISATION --------------------------------------------------------------------------------------- ISO/IEC JTC1/SC2/WG2 Universal Multiple-Octet Coded Character Set (UCS) -------------------------------------------------------------------------------- ISO/IEC JTC1/SC2/WG2 N 2005 Date: 1999-05-29 TITLE: ISO/IEC 10646-1 Second Edition text, Draft 2 SOURCE: Bruce Paterson, project editor STATUS: Working paper of JTC1/SC2/WG2 ACTION: For review and comment by WG2 DISTRIBUTION: Members of JTC1/SC2/WG2 1. Scope This paper provides a second draft of the text sections of the Second Edition of ISO/IEC 10646-1. It replaces the previous paper WG2 N 1796 (1998-06-01). This draft text includes: - Clauses 1 to 27 (replacing the previous clauses 1 to 26), - Annexes A to R (replacing the previous Annexes A to T), and is attached here as “Draft 2 for ISO/IEC 10646-1 : 1999” (pages ii & 1 to 77). Published and Draft Amendments up to Amd.31 (Tibetan extended), Technical Corrigenda nos. 1, 2, and 3, and editorial corrigenda approved by WG2 up to 1999-03-15, have been applied to the text. The draft does not include: - character glyph tables and name tables (these will be provided in a separate WG2 document from AFII), - the alphabetically sorted list of character names in Annex E (now Annex G), - markings to show the differences from the previous draft. A separate WG2 paper will give the editorial corrigenda applied to this text since N 1796. The editorial corrigenda are as agreed at WG2 meetings #34 to #36. Editorial corrigenda applicable to the character glyph tables and name tables, as listed in N1796 pages 2 to 5, have already been applied to the draft character tables prepared by AFII.
    [Show full text]
  • An Introduction to Indic Scripts
    An Introduction to Indic Scripts Richard Ishida W3C [email protected] HTML version: http://www.w3.org/2002/Talks/09-ri-indic/indic-paper.html PDF version: http://www.w3.org/2002/Talks/09-ri-indic/indic-paper.pdf Introduction This paper provides an introduction to the major Indic scripts used on the Indian mainland. Those addressed in this paper include specifically Bengali, Devanagari, Gujarati, Gurmukhi, Kannada, Malayalam, Oriya, Tamil, and Telugu. I have used XHTML encoded in UTF-8 for the base version of this paper. Most of the XHTML file can be viewed if you are running Windows XP with all associated Indic font and rendering support, and the Arial Unicode MS font. For examples that require complex rendering in scripts not yet supported by this configuration, such as Bengali, Oriya, and Malayalam, I have used non- Unicode fonts supplied with Gamma's Unitype. To view all fonts as intended without the above you can view the PDF file whose URL is given above. Although the Indic scripts are often described as similar, there is a large amount of variation at the detailed implementation level. To provide a detailed account of how each Indic script implements particular features on a letter by letter basis would require too much time and space for the task at hand. Nevertheless, despite the detail variations, the basic mechanisms are to a large extent the same, and at the general level there is a great deal of similarity between these scripts. It is certainly possible to structure a discussion of the relevant features along the same lines for each of the scripts in the set.
    [Show full text]