Sentence Complexity Estimation for Chinese-Speaking Learners of Japanese
Total Page:16
File Type:pdf, Size:1020Kb
Sentence Complexity Estimation for Chinese-speaking Learners of Japanese Jun Liu Yuji Matsumoto Graduate School of Information Science Graduate School of Information Science Nara Institute of Science and Technology Nara Institute of Science and Technology 8916-5 Takayama, Ikoma, Nara 630-0192, Japan 8916-5 Takayama, Ikoma, Nara 630-0192, Japan [email protected] [email protected] Song, 2011; Horiba, 2012; Gilakjani and Sabouri, Abstract 2016). Reading comprehension is also influenced by the complexity of the reading material. Texts It is fairly challenging for a foreign containing highly demanding vocabularies and language learner to read and understand highly complex sentence structures are likely to Japanese texts containing words of high disturb the learners’ reading comprehension. difficulty level or low frequency and Learners of Japanese language from kanji complicated linguistic structures. Because a background countries benefit substantially from large number of Chinese characters (kanji kanji characters commonly used in both Japanese in Japanese) are commonly used both in and Chinese when they read Japanese sentences or Chinese and Japanese, the more confusing documents. However, it is more challenging for problem for Japanese language learners them to read and learn various Japanese functional from kanji background countries is the expressions with varied meanings and usages. acquisition of various complex Japanese The selection of appropriate reading material functional expressions. In this study, we matching the learners’ individual capabilities is propose a method utilizing Japanese kanji highly likely to enable language learners to read in characters, particularly Japanese–Chinese a more focused and selective manner. To support homographs with identical or similar learners in gathering useful information from texts meanings, as a critical feature of sentence- more effectively, certain online public Japanese complexity estimation for Chinese- reading-assistance systems such as Reading Tutor1, speaking learners of Japanese language. Asunaro 2 , Rikai 3 , and WWWJDIC 4 are highly Experimental results have partially effective. These systems are adequately demonstrated the effectiveness of our constructed for providing an internet learning method in enhancing the accuracy of environment where learners can make complete sentence-complexity estimation. use of information from the internet for their Japanese language study, and a few of them are 1 Introduction specifically designed to enable language learners to Enhancing reading capability is one of the understand Japanese texts by offering words with important purposes in second language teaching their corresponding difficulty level information or and learning. There are various factors that impact translation (Toyoda 2016). However, these systems learners’ reading comprehension. A few of these 1 factors involve the learners’ vocabulary knowledge, http://language.tiu.ac.jp/ 2 https://hinoki-project.org/asunaro/ grammar knowledge, reading strategies, interest, 3 http://www.rikai.com/perl/Home.pl attitude, and motivation (Koda, 2007; Han and 4 http://nihongo.monash.edu/cgi-bin/wwwjdic?9T 296 31st Pacific Asia Conference on Language, Information and Computation (PACLIC 31), pages 296–302 Cebu City, Philippines, November 16-18, 2017 Copyright c 2017 Jun Liu and Yuji Matsumoto do not take the learners’ native language that focuses on grammar and utilizes grammar background into account. Moreover, these systems templates. provide learners with limited information on the In recent years, a few Japanese text difficulty grammatical difficulty of all the various types of evaluation systems have been developed to support Japanese functional expressions, which learners Japanese language learners (Hasebe and Lee, 2015; actually intend to learn as a part of the procedure Lee and Hasebe, 2016). For example, JReadability5 for learning Japanese. can analyze input text and estimate its readability In Section 2 of this paper, we introduce some to categorize it as belonging to one of six difficulty previous works. In Section 3, we describe our levels, on the basis of five characteristics: average method for ranking example sentences of Japanese length of sentence; percentage of kango (words of functional expressions by utilizing Japanese– Chinese origin), percentage of wago (words of Chinese homographs with identical or similar Japanese origin), percentage of verbs, and meanings, as a critical feature. Section 4 describes percentage of particles. the several experiments conducted to examine the However, JReadability too does not sufficiently effectiveness of our method. Finally, in Section 5, consider the various types of Japanese functional we conclude and describe future work. expressions with varying difficulty levels. The prediction value calculated by this system is more 2 Previous Research reliable for long texts (approximately 1000 characters) and not for single sentences. Text difficulty or text readability evaluation is one of the challenges in natural language 3 General Method processing (NLP) owing to the linguistic complexity generated from both vocabulary and Japanese and Chinese share a large quantity of grammar. Researchers have been actively homographs that use identical kanji characters exploring methods to evaluate text difficulty (both in simplified Chinese and traditional (Gonzalez-Dios et al., 2014; Hancke, Vajjala, and Chinese). Table 1 presents a few examples of Meurers, 2012; Vajjala and Meurers, 2012; Xia, Japanese–Chinese homographs. These words play Kochmar and Briscoe, 2016). a significant role while reading Japanese or For English texts, there are numerous popular Chinese texts. According to a report by Wang formulas such as Flesch Reading Ease (Flesch (2001), approximately 80–95% Japanese–Chinese 1948) and Flesch-Kincaid Grade Level, all of homographs are used to express identical or similar which are used for several applications such as meanings in both the languages. Foreign language compilation of reading materials for language learners from kanji background countries can learners. Collins–Thompson and Callan (2004) straightforwardly understand the meaning of these proposed a language modeling method to estimate words according to kanji characters. This is the readability of English and French texts. occasionally more convenient than grammar for For Japanese texts, Tateishi, Ono, and Yamada foreign language learners from kanji background (1988a; 1988b) introduced a formula based on six countries to learn Japanese. surface characteristics: average number of For Japanese language learners, a vital challenge characters per sentence, average number of Roman is to master a large number of complex functional letters and symbols, average number of hiragana expressions. Hence, providing appropriate example characters, average number of kanji characters, sentences for learners based on their individual average number of katakana characters, and ratio Japanese language capabilities are highly likely to of touten (comma) to kuten (period). Formula- aid the enhancement of the efficiency of learning based approaches have also been used or teaching various Japanese functional expressions. Japanese to young native speakers (Shibasaki and In order to achieve this goal, we utilize Sawai, 2007; Sato, Matsuyoshi, and Kondoh, 2008; Japanese–Chinese homographs as a new feature, Shibasaki and Tamaoka, 2010). To evaluate text which is more or less dissimilar from previous difficulty level for foreign language learners of research, to estimate sentence difficulty and select Japanese, Wang and Andersen (2016) introduced an approach for evaluating Japanese text difficulty 5 http://jreadability.net 297 the most appropriate example sentences as learning content for Japanese functional expressions. Japanese vocabulary Difficulty level 山岳(mountains) Japanese Chinese Meaning 養う(to cultivate) N1 社会(society) 社会(society) Identical 忙しない(busy) 技術(technology) 技术(technology) Identical 前提(Presupposition) 東西(east and west) 东西(east and Similar 迫る(to press) N2 west; thing) 勇ましい(brave) 培養(culture) 培养(culture; Similar 愛情(love) train) 含める(to include) N3 手紙 手纸 (letters) (toilet paper) Dissimilar 巨大(huge) 勉強(study) 勉强(reluctantly) Dissimilar 複雑(complex) 捨てる(to throw away) N4 Table 1: Examples of Japanese–Chinese 挨拶(greeting) homographs. 学校(school) 3.1 Difficulty Level Evaluation Standard 明るい(bright) N5 To estimate the difficulties of example sentences, 始まる(begin) we follow the standard of the Japanese Language Japanese grammar Difficulty level Proficiency Test (JLPT). The JLPT consists of five べからざる(must not) levels: N1, N2, N3, N4, and N5. The least difficult がてら(while doing something) SP1 6 level is N5, and the most difficult level is N1 . を顧みず(regardless of) Since 2010, the JLPT official lists of vocabulary からといって(just because) and grammar have not been published in books, we に加えて(in addition to) SP2 referenced a few books (Xu and Reika, 2013a; Xu and Reika, 2013b) and online learning websites7,8, に違いない(without a doubt) all of which provide lists of the JLPT vocabulary にとって(to) and grammar with difficulty levels ranging from に比べて(compare) SP3 N1–N5. Here, we consider levels N3/SP3 and わけがない(it is impossible that) lower as “easy” level, levels N2/SP2 and above as かもしれない(maybe) difficult level. A few examples of vocabulary and ことができる(can) SP4 grammar in JLPT are presented in Table 2. みたいだ(similar to) 3.2 List of Japanese–Chinese Homographs てから(after) 前に Japanese language learners from kanji