Journal of Global Tourism Research, Volume 5, Number 2, 2020

Original Article

A study of English tourist guidebooks at airports in in

Hiromi Ban (Graduate School of Engineering, Nagaoka University of Technology, [email protected]) Takashi Oyabu (Nihonkai International Exchange Center, [email protected])

Abstract is located in the Hokuriku region in Japan. One of the main targets of the tourism industry in Ishikawa is to increase the number of tourists from foreign countries. In order to achieve this goal, it is necessary to provide foreign tourists with “language service.” In this study, in order to understand the state of language service provided to foreign tourists, what linguistic characteristics can be found in English guidebooks at and Airport, which are local airports in Japan, were investigated and compared with guidebooks available at international airports in Japan and the U.S. In short, frequency characteris- tics of character- and word-appearance were investigated using a program written in C++. These characteristics were approximated by an exponential function. Furthermore, the percentage of Japanese junior high school required vocabulary and American basic vocabulary was calculated to obtain the difficulty-level as well as the K-characteristic of each material. As a result, it was clearly shown that English guidebooks available at airports in the Hokuriku region have a similar tendency to literary writings in the char- acteristics of character-appearance. Besides, the values of the K-characteristic for the guidebooks are high, and the difficulty level is low in terms of the American basic vocabulary.

Keywords guidebooks at local airports in Japan have some interesting data mining, metrical linguistics, statistical analysis, text min- characteristics regarding character- and word-appearance. ing, tourist guidebook 2. Method of analysis and materials 1. Introduction The materials analyzed here are English guidebooks availa- Ishikawa Prefecture, located in the Hokuriku region in Ja- ble at Komatsu, Toyama, Narita, Kansai and Chubu. Moreover, pan, has a population of about 1.1 million, and its capital is San Francisco was taken as an example city. Ishikawa is blessed with natural beauty and of an overseas international airport because San Francisco is a traditional cultures, which attract a lot of tourists. These days, popular tourist destination in the United States. The following one of the main targets of the tourism industry in Ishikawa is guidebooks were selected with paying attention to unify the to increase the number of tourists from foreign countries. In or- topics as much as possible. der to achieve this goal, it is necessary to provide foreign tour- ists with a “language service,” which motivates foreigners to • Material 1: HOKURIKU JAPAN, , Ishikawa & Toyama, go sightseeing more easily. This “language service” means to RESORT OF WONDERS AND FASCINATION, Hot spring serve benefits and convenience to foreign tourists by enhancing route blessed with four seasons, Mar. 2000, Komatsu Air- signs, pamphlets and homepages in several languages. It will port become a key word for the increase of foreign tourists [Ban and • Material 2: TOYAMA – Japan, Oct. 2007, and TOYAMA City Oyabu, 2019; National Association of Language, Business and Guide, Nov. 2006, Tourism Education, 2020]. • Material 3: Tourist Guide, Around Narita International Air- While some foreigners who visit Kyoto often extend their port, May 2008, Narita International Airport trip to Kanazawa which is located about two hours away by • Material 4: Have a nice day in KANSAI, Visitor’s guide, vol. limited express train, other tourists also come to use regular 5, Feb. 2008, Kansai International Airport flights from and Shanghai or charter flights from Tai- • Material 5: Aichi, Gifu, Mie, Shizuoka, Fukui, , AC- wan to Komatsu Airport, located one hour or less away from CESS MAP, June 2007, Chubu Centrair International Airport Kanazawa city by car. Moreover, there are regular flights from • Material 6: san francisco guide®, where to go & what to do, Dalian to Toyama Airport which is located in the vicinity of Aug. 2010, San Francisco International Airport (SFO) Kanazawa, and it is likely that tourists who visit Toyama will also visit Ishikawa Prefecture. Due to the circulation, the publication of Material 1 is older In this study, in order to understand the state of “language than other materials. service” provided to foreign tourists, English guidebooks at In addition, English textbooks for Japanese junior high Komatsu Airport and Toyama Airport, which are local air- school students (NEW HORIZON English Course 1, 2 and 3 ports in Japan, were examined, and compared with guidebooks (2010, Shoseki Co., Ltd.) (hereinafter referred to as “JHS available at Narita, Kansai, Chubu, and San Francisco interna- 1, 2 and 3”)) and those for Japanese high school students (UNI- tional airports. As a result, it was clearly shown that English CORN ENGLISH COURSE I, II and READING (2010, Bun-ei-

Union Press 141 H. Ban and T. Oyabu: A study of English tourist guidebooks at airports in Hokuriku region in Japan

do Publishing Co., Ltd.) (“HS 1, 2 and 3”)) were also analyzed. 0.125 HS 3 The computer program for this analysis is composed of C++. 0.120 Characters 3.Narita HS 2 Besides the characteristics of character- and word-appearance 0.115 1.Komatsu 2.Toyama for each piece of material, various information such as the 0.110 HS 1 5.Chubu “number of sentences,” the “number of paragraphs,” the “mean 0.105 JHS 3 e n t b 4.Kansai c i JHS 2 i 0.100 length,” the “number of words per sentence,” etc. can be ex- f

e f 0.095 tracted by this program [Ban et al., 2016a]. o C 0.090 JHS 1 0.085 3. Results 0.080 (y = 0.0096x + 0.0114) 3.1 Characteristics of character-appearance 0.075 Zipf’s law being referred to, frequencies of character- and 6.SFO 0.070 word-appearance were examined. First, the frequently used 6.0 6.5 7.0 7.5 8.0 8.5 9.0 9.5 10.0 10.5 11.0 11.5 12.0 characters in each material and their frequency were derived. Coefficient c The frequencies of the 50 most frequently used characters in- Figure 2: Dispersions of coefficients c and b for character- cluding blanks, capitals, small letters, and punctuations were appearance plotted on a descending scale. The vertical shaft shows the degree of the frequency and the horizontal shaft shows the or- der of character-appearance. The vertical shaft is scaled with a approximated by [y = 0.0096x + 0.0114]. The values of coef- logarithm. Figure 1 shows the results for Material 1. ficients c and b for Materials 1 and 2 are high: the values of c are 10.540 and 10.811, and those of b are 0.1130 and 0.1129.

100 Previously, various English writings were analyzed and it was reported that, as for the 50 most frequently used characters, Characters Komatsu Airport blank there is a positive correlation between the coefficients c and b, e 10 a and that the more journalistic the material is, the lower the val- rh t i s o ues of c and b are, and that the more literary the material is, the 1 higher the values of c and b are [Ban et al., 2015a]. Thus, while y = 10.54e–0.113x the material at San Francisco International Airport is rather

Frequency (%) journalistic, the tourist guidebooks available at local airports in 0.1 Japan have a similar tendency to English literary writings.

0.01 3.2 Characteristics of word-appearance 1 5 9 13 17 21 25 29 33 37 41 45 49 Next, frequently used words in each material and their fre- Order of appearance quency were derived. Table 1 shows the top 20 words most Figure 1: Frequency characteristics of character-appearance in frequently used in each material. The article THE is the most Material 1 frequently used word in every material except JHS 1. While OF is the second most frequently used word in the five guide- Between the 26th and 27th places, there is an inflection point books in Japan, AND is the second most frequently used word caused by the difference in declines, and a relatively larger de- for Material 6. In the cases of Materials 1 and 2, as well as JHS cline is observed at the 27th place and thereafter. This charac- 1 and JHS 2, the frequency of CAN is high, which ranks at 15 teristic curve was approximated by the following exponential and 12 respectively. On the other hand, in the cases of Materi- function: als 3, 4 and 5, the frequencies of JAPAN and JAPANESE are high. Besides, nouns related to tourism, such as FESTIVAL in y = c * exp (–bx) (1) Material 1, VISITORS, TEMPLES, STREET and TRANSIT can be seen at the 8th to 20th in guidebooks. From this function, coefficients c and b can be derived [Ban Just as in the case of characters, the frequencies of the 50 and Oyabu, 2013]. In the case of Material 1, as shown in Figure most frequently used words in each material were plotted. Each 1, values, c = 10.54 and b = 0.113 were obtained. characteristic curve was approximated by the same exponential The distribution of coefficients c and b extracted from each function. The distribution of c and b is shown in Figure 3. As material is shown in Figure 2. for the coefficient c, the values for Materials 1 and 2 are high: There is a linear relationship between c and b for all materi- they are 1.9973 (Material 1) and 2.0042 (Material 2), compared als. While the values of coefficients c and b for Material 6 are with the value for Material 6 (1.5527). Moreover, the value of lowest, those for HS 3, HS 2 and Material 3 are high. With coefficient c gradually increases in the order of Material 1, Ma- regard to the English textbooks, values of c and b are larger for terial 2 and Material 3. This order corresponds with the coef- higher grades. The values for all the six tourist guidebooks are ficients c for character-appearance, and the intervals of the val-

142 Journal of Global Tourism Research, Volume 5, Number 2, 2020 ) I a it is to in of as on we for the are but she and that was with have HS 3 Unicorn R ( ) I a it is to in of as on we for the my but she and had that was they HS 2 Unicorn 2 ( ) I a it is to in of he on for his the are my one and that was they HS 1 people Unicorn 1 ( ) I a it is to in of for the are my but she and you this was very have JHS 3 people Horizon 3 Horizon ( ) I a it is to in of he on we for the are but can and you was have about JHS 2 Horizon 2 Horizon ( ) I a is at to in do it's we the are my I'm yes can you this like your have JHS 1 Horizon 1 Horizon ( - a is at to in of or by on for the SFO and San map with from cisco Fran street public transit 6. a is to in of as an on for the are hot and this city was with from Chubu Japan famous 5. Table 1: High-frequency 1: Table words each for material a is at to in of as its by on for the are can and you with from Japan Kansai temple 4. a is to in of as an on for the are and this that from Narita pride many Narita visitors 3. Japanese a it is at to in of as by on for the are can and you this with from Toyama Toyama 2. a it is to in of as by on for the are has can and this with from which festival Komatsu 1. 2 3 4 5 6 7 8 9 1 11 17 16 14 18 10 15 19 13 12 20

143 H. Ban and T. Oyabu: A study of English tourist guidebooks at airports in Hokuriku region in Japan

0.055 value for Material 6 (64.349), which is the lowest of all the 12 materials. The values for Materials 1 and 2 are high: they are 0.053 Words 1.Komatsu 118.882 (Material 1) and 107.047 (Material 2), which are 54.533 2.Toyama 0.051 4.Kansai and 42.678 higher than that for Material 6. As for textbooks,

b 0.049 6.SFO 5.Chubu HS 3 3.Narita the values for JHS and those for HS are 70.358 to 78.935 and e n t

c i 79.643 to 85.488, which are similar respectively, and the former i

f 0.047

e f are lower than the latter. o

C 0.045 HS 1 HS 2 The results showing higher K-characteristics for Materials 1 0.043 JHS 2 and 2 than for Material 6 coincide with the aforementioned ten- 0.041 dency regarding coefficients c and b for character- and word- JHS 1 0.039 JHS 3 appearance. In addition, higher K-characteristic values for 1.5 1.6 1.7 1.8 1.9 2.0 2.1 textbooks for HS than those for JHS coincide with the tendency Coefficient c regarding coefficients c and b for character-appearance and Figure 3: Dispersions of coefficients c and b for word-appear- coefficient b for word-appearance. This correlation between ance K-characteristic and the coefficients for character- and word- appearance needs to be studied in the future. ues in both cases are very similar as well. On the hand, as for 3.3 Degree of difficulty the coefficient b, the value for Material 1 is the highest and that In order to show how difficult the materials for readers are, for Material 2 is the second highest of all. All the six guide- the degree of difficulty for each material through the variety books have higher values than all the six textbooks. Besides, of words and their frequency was derived [Ban et al., 2015b; the values of coefficients c and b for word-appearance for Ma- Ban et al., 2016b]. That is, two parameters to measure difficulty terials 1, 2 and 3, and those for three textbooks for high school were used; one is for word-type or word-sort (Dws), and the students are similar respectively, and they might be regarded as other is for the frequency or the number of words (Dwn). The two clusters. equation for each parameter is as follows: As a method of featuring words used in a writing, a statisti- cian named Udny Yule suggested an index called the “K-char- Dws = (1 – nrs / ns) (3) acteristic” in 1944 [Yule, 1944]. This can express the richness of vocabulary in writings by measuring the probability of any Dwn = {1 – ( 1 / nt * Σn(i))} (4) randomly selected pair of words being identical. He tried to identify the author of The Imitation of Christ using this index. where nt means the total number of words, ns means the total

This K-characteristic is defined as follows: number of word-sort, nrs means the required English vocabu- lary in Japanese junior high schools or American basic vocabu- 4 2 K = 10 (S2 / S1 – 1 / S1) (2) lary by The American Heritage Picture Dictionary (American Heritage Dictionaries, Houghton Mifflin, 2003), and n(i) means where if there are fi words used xi times in a writing, S1 = Σ xifi , the respective number of each required or basic word. Thus, 2 S2 = Σxi fi. it can be calculated how many required or basic words are not The K-characteristic for each material was examined. The contained in each piece of material in terms of word-sort and results are shown in Figure 4. According to the figure, the val- frequency. ues for the five guidebooks in Japan are high: they range form Thus, the values of both Dws and Dwn were calculated to show 97.682 (Material 3) to 124.897 (Material 5), compared with the how difficult the materials are for readers, and to show at which

5. Chubu 1. Komatsu (124.897) 2. Toyama (118.882) 4. Kansai (107.027) 6. SFO 3. Narita (99.684) (64.349) (97.682)

JHS 2 HS 1 HS 3 HS 2 (85.488) JHS 3 JHS 1 (78.935) (79.642) (82.301) (70.358) (71.259)

60 70 80 90 100 110 120 130

Figure 4: K-characteristic for each material

144 Journal of Global Tourism Research, Volume 5, Number 2, 2020

level of English the materials are, compared with other materi- other parts of speech because the meaning of each word was als. Then, in order to make the judgments of difficulty easier not checked. for the general public, one difficulty parameter was derived from Dws and Dwn using the following principal component 3.4.1 Mean word length analysis: As for the “mean word length,” it is 5.861 letters for Mate- rial 1, which is the shortest of all the six guidebook materials. z = a1 * Dws + a2 * Dwn (5) In the case of Material 2, it is 5.937 letters, which being equal to Material 4, is the third longest of all. The mean word length where a1 and a2 are the weights used to combine Dws and Dwn. of Material 6 (6.004 letters) is longer than any other material. Using the variance-covariance matrix, the first principal com- It seems that this is because Material 6 contains many long- ponent z was extracted: [z = 0.7071 * Dws + 0.7071 * Dwn] both length terms such as COLLECTION (10 times), ENTERTAIN- for required and basic vocabularies, from which the principal MANT (13), FISHERMAN’S (45), NEIGHBORHOOD(S) (15), component scores were calculated. The results are shown in RESTAURANT(S) (32) and WATERFRONT (9). Figure 5. According to Figure 5, all the 6 guidebooks are more difficult 3.4.2 Number of words per sentence than English textbooks, and Material 6 is by far the most dif- The “number of words per sentence” for Material 1 is 17.836 ficult of all. In the case of the required vocabulary, Material 4 words and that for Material 2 is 17.099 words. They are the is the most difficult, and Material 1 is the second most difficult second and the third longest of all the materials. All the five of the five guidebooks in Japan. The difficulty of Material 1 is guidebooks in Japan have more number of words per sentence similar to that of Material 4. The difficulty level decreases in than Material 6 (14.806 words). Thus, it can be said that Eng- the order of Materials 4, 1, 5, 2 and 3. On the other hand, in the lish tourist guidebooks at Japanese airports are characterized case of the basic vocabulary, Material 3 is the most difficult, by a large number of words per sentence. Material 3 (18.145 and Material 2 is the easiest of the five guidebook materials in words) has the highest number of all. From this point of view, Japan. Material 5 is the second most difficult after Material 3, as well as the result of the difficulty derived through the variety and its difficulty is almost equal to Materials 1 and 4. of words and their frequency in terms of the basic vocabulary, Thus, although Material 1 is difficult in the case of the Japa- Material 3 seems to be rather difficult to read. nese required vocabulary, English guidebooks available at local airports in Hokuriku region are easy in terms of the American 3.4.3 Frequency of relatives basic vocabulary. Therefore, the materials seem to be easier for The “frequency of relatives” for Material 2 is 1.414 %, which Americans to read. is the second highest, and the one for Material 1 is 1.033 %, which is the fourth highest of all the guidebooks. The fre- 3.4 Other characteristics quency for Material 2 is as high as that for Material 3 (1.540 Other metrical characteristics of each material were com- %). The one for Material 5, whose percentage is only 0.472 %, pared. The results of the “mean word length,” the “number of is the lowest of all. Therefore, it can be assumed that the Eng- words per sentence,” etc. are shown in Table 2. Although the lish guidebooks at Toyama and Narita Airports tend to contain “frequency of prepositions,” the “frequency of relatives,” etc. more complex sentences, the material seems to be difficult to were counted, some of the words counted might be used as read from this point of view, as well as in terms of the variety

3

Difficulty → ) y t 6.SFO u l

c 2

f i 3.Narita d i 5.Chubu 4.Kansai o f 1 1.Komatsu

d e 2.Toyama r a HS 3

( G 0 HS 2 –3.5 –2.5 –1.5 –0.5 HS 1 0.5 1.5 2.5 3.5 Required (Japan) (Grade of difficulty →) –1

–2 JHS 3 ) S .

JHS 2 . –3 ( U c i JHS 1 s

B a –4

Figure 5: Principal component scores of difficulty

145 H. Ban and T. Oyabu: A study of English tourist guidebooks at airports in Hokuriku region in Japan ) 76 260 1.217 4.412 3.865 5.568 0.977 8.393 3,594 1,005 2.383 HS3 15.778 15.052 15,857 88,289 Unicorn R ( ) 75 261 890 5.517 3.410 4.616 1.215 8.707 2.421 0.801 2,657 HS 2 14.810 13.780 67,662 12,264 Unicorn 2 ( ) 73 163 633 1.745 5.478 9.324 3.883 3.883 3.926 0.694 2,059 0.802 8,083 HS 1 14.769 12.769 44,279 Unicorn 1 ( ) 71 317 177 764 1.119 5.161 1.791 8.183 0.331 1.927 3.395 2,594 JHS 3 12.188 13,387 10.684 Horizon 3 Horizon ( ) 69 394 227 799 1.736 1.736 1.530 7.299 1.392 3.599 2,876 4.994 0.223 11.788 15.511 JHS 2 14,362 Horizon 2 Horizon ( ) 69 251 233 497 9.110 1.792 5.335 1,339 1.077 5.096 0.897 0.263 2.694 6,824 17.476 JHS 1 Horizon 1 Horizon ( 79 199 968 SFO 1.130 3.919 0.475 3,657 1.040 0.266 4.864 6.004 11.647 14,332 14.806 86,046 6. 43 69 101 787 2.159 0.472 0.472 1.649 1,699 0.530 0.950 5.906 2.349 Chubu 13.954 16.822 10,034 5. 77 132 287 2.174 2.917 0.746 2.610 1,671 5.937 4,874 0.699 0.842 16.983 15.292 28,936 Table 2: Metrical 2: dataTable each for material 4. Kansai 71 54 179 1,169 3.315 0.810 2.778 1.324 0.833 0.833 1.540 5.964 3,248 Narita 18.145 19,372 15.306 3. 74 252 120 1.414 2.157 5.937 0.861 0.974 1,423 2.100 3.028 4,309 17.099 14.202 25,583 Toyama 2. 75 147 385 1.033 2.619 5.861 0.728 0.728 0.797 1.545 3.567 1,925 6,867 17.836 15.367 40,245 Komatsu 1. Total num.Total characters of num.Total character-type of num.Total words of num.Total word-type of num.Total sentences of num.Total paragraphs of Mean word length Words/sentence Sentences/paragraph Commas/sentence Repetition a word of Freq. prepositions of (%) Freq. relatives of (%) Freq. auxiliaries of (%) Freq. pers. of pronouns (%)

146 Journal of Global Tourism Research, Volume 5, Number 2, 2020

of words and their frequency. rials after 9-letter words.

3.4.4 Frequency of auxiliaries 3.6 Positioning of each material There are two kinds of auxiliaries in a broad sense. One After the aforementioned results being standardized, a po- expresses the tense and voice, such as BE which makes up the sitioning of all the materials was made with a principal com- progressive form and the passive form, the perfect tense HAVE, ponent analysis by correlation procession. The following 22 and DO in interrogative sentences or negative sentences. The items were considered: the values of coefficient c for character- other is a modal auxiliary, such as WILL or CAN which ex- appearance, coefficient b for character-appearance, coefficient presses the mood or attitude of the speaker [Ban et al., 2017]. c for word-appearance, coefficient b for word-appearance, and In this study, only modal auxiliaries were targeted. As a result, K-characteristic, the principal component scores of difficulty while the “frequency of auxiliaries” for Material 2 (0.974 %) is using the required vocabulary, and scores of difficulty us- the highest and Material 1 (0.728 %) is the third highest of all ing the basic vocabulary, and the total numbers of characters, the guidebook materials, Material 6 contains 0.266 % auxilia- character-type, words, word-type, sentences, and paragraphs, ries, which is the lowest of all. Therefore, it might be said that the mean word length, the numbers of words per sentence, sen- while the writers of English guidebooks available at Japanese tences per paragraph, commas per sentence, and repetition of a airports tend to communicate their subtle thoughts and feelings word, and the frequencies of prepositions, relatives, auxiliaries, by using auxiliary verbs, the style of Material 6 can be called and personal pronouns. Figure 7 shows the results. more assertive. In Figure 7, both Material 1 and Material 2 are located next to Material 4. Therefore, it can be said that the literary style as 3.5 Word-length distribution a whole of the English guidebooks available at the airports in In addition, word-length distribution for each material was the Hokuriku region in Japan is similar to the style of the Kan- examined. The results are shown in Figure 6. The vertical sai International Airport. shaft shows the degree of frequency with the word length as a As for the Hokuriku region, the number of limited express variable. As for all the guidebook materials, the frequency of trains whose departure and arrival is in the Osaka district is 3-letter words is the highest. The frequency of 3-letter words much larger than that for the Kanto and Chubu areas. Then, the ranges from 17.334 % (Material 3) to 21.307 % (Material 5). Hokuriku region seems to have received more influence of the The frequency of 5-letter words such as ENJOY, WATER and Kansai area. Moreover, the characteristics of spoken language WHICH for Materials 1 and 2 is higher than in other 10 mate- in the Hokuriku region seem to be comparatively similar to rials. While in the case of Material 1, the frequency decreases those in the Kansai area. Thus, it is very interesting that also after 4-letter words, in the case of Material 2, although the fre- the English guidebooks analyzed in this study have more influ- quency decreases until 7-letter words, the frequency of 8-letter ence of the Kansai area. words such as FESTIVAL, GOKAYAMA and VISITORS is 0.604 % higher than that of 7-letter words. 4. Conclusion Besides, although Materials 1 and 2 have almost equal fre- Some characteristics of character- and word-appearance quencies to other guidebooks regarding 8-letter words, the for English tourist guidebooks at local airports in Hokuriku degree of decrease for them gets a little higher than other mate- region in Japan were investigated, compared with those for

25 1. Komatsu Word length 2. Toyama 3. Narita 20 4. Kansai 5. Chubu 6. SFO 15 JHS 1 JHS 2 JHS 3 HS 1

Percentage 10 HS 2 HS 3

5

0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 Number of letters

Figure 6: Word-length distribution for each material

147 H. Ban and T. Oyabu: A study of English tourist guidebooks at airports in Hokuriku region in Japan

15

HS 3 10 HS 2

5 HS 1 6. SFO JHS 2 0 JHS 3 1. Komatsu JHS 1 –5 2. Toyama 4. Kansai 3. Narita

–10 5. Chubu

–15 –20 –15 –10 –5 0 5 10 15

Figure 7: Positioning of each material guidebooks available at Narita, Kansai, Chubu, and San Fran- terials for environmentology. International Journal of Busi- cisco international airports. In this analysis, an approximate ness and Economics, Vol. 5, No. 1, 21-32. equation of an exponential function was used to extract the Ban, H. and Oyabu, T. (2019). Feature extraction of the “Tour- characteristics of each material using coefficients c and b of ism English Proficiency Test” using data mining. Journal of the equation. Moreover, the percentage of Japanese junior high Global Tourism Research, Vol. 4, No. 1, 27-34. school required vocabulary and American basic vocabulary Ministry of Land, Infrastructure, Transport and Tourism (ed.) was calculated to obtain the difficulty-level as well as the K- (2020). White paper on tourism, 2018 ed., National Printing characteristic. As a result, it was clearly shown that English Bureau. guidebooks available at local airports in Hokuriku have a simi- Yule, G. U. (1944). The statistical study of literary vocabulary. lar tendency to literary writings in the characteristics of char- Cambridge University Press. acter-appearance. Besides, the values of the K-characteristic for the guidebooks are high, and the difficulty level is low in terms (Received October 5, 2020; accepted October 31, 2020) of the American basic vocabulary. In the future, to examine new guidebooks published after the opening of the and compare with the results educed in this study is being planned.

References Ban, H., Kimura, H., and Oyabu, T. (2015a). Text mining of English materials for business management. International Journal of Engineering & Technical Research, Vol. 3, No. 8, 238-243. Ban, H., Kimura, H., and Oyabu, T. (2016a). Feature extraction of English guidebooks for Hokuriku region in Japan. Jour- nal of Global Tourism Research, Vol. 1, No. 1, 71-76. Ban, H., Kimura, H., and Oyabu, T. (2016b). Text mining of English articles on the Noto Hanto Earthquake in 2007. Journal of Global Tourism Research, Vol. 1, No. 2, 115-120. Ban, H., Kimura, H., and Oyabu, T. (2017). Metrical feature extraction of English books on Tourism. Journal of Global Tourism Research, Vol. 2, No. 1, 67-72. Ban, H., Oguri, R., and Kimura, H. (2015b). Difficulty-level classification for English writings. Transactions on Machine Learning and Artificial Intelligence, Vol. 3, No. 3, 24-32. Ban, H. and Oyabu, T. (2013). Text data mining of English ma-

148