Tampilkan postingan dengan label TEACHING. Tampilkan semua postingan
Tampilkan postingan dengan label TEACHING. Tampilkan semua postingan

Kamis, 23 Desember 2010

Exploring ESL Learners' Use of Hypermedia Reading Glosses


GULCAN ERCETIN
Bogazici University
Abstract:
This study explores the types of annotations intermediate and advanced ESL learners preferred to use while they were engaged in reading a hypermedia text. The study also investigated learners' attitudes towards reading in a hypermedia environment. The participants were 84 ESL adult learners enrolled at the Center for English as a Second Language at the University of Arizona. Data were collected through a tracking tool, a reading comprehension test, a questionnaire, and interviews. Quantitative data indicated that the intermediate ESL learners accessed annotations more frequently than those in the advanced group. However, they did not differ in the amount of time they spent on annotations. On the other hand, the advanced learners performed better on the reading comprehension test. Both groups of learners preferred word definitions to pronunciations of words and graphics in order to get information about the text at the word level. However, they preferred videos and graphics to get extra information about the topic. Analysis of the qualitative data revealed that hypermedia reading had a positive impact on the participants' attitudes towards reading on the computer. The participants indicated that the hypermedia environment made reading more enjoyable and comprehensible.
INTRODUCTION
New digital technologies such as hypermedia, hypertext, or multimedia have great potential for teaching and learning because of the innovative ways they present information and the control and freedom they give learners over their learning (Means, Blando, Olson, Morocco, Remz, & Zorfass., 1993; Heller, 1990; Marsh & Kumar, 1992; Roblyer, Edwards, & Havriluk, 1997; Popkewitz & Shutkin, 1995). Hypermedia combines hypertext and multimedia within one system (Jonassen, 1996). In other words, hypermedia environments such as those found on the web or in electronic books utilize text, sound, graphics, video, and animation in the same document and present information in a nonlinear fashion through nodes and links. Thus, hypermedia is a type of multimedia distinguished by the richness and depth of information it provides through links and multiple presentation modes (Preece, 1993). Although hypermedia provides "flexible information environments" (Goldman, 1996), reading hypermedia documents poses certain challenges for readers. Successful reading comprehension in a hypermedia environment goes beyond effective uses of top-down and bottom-up processes; it requires additional reading skills to cope with the demands of the new environment. For instance, readers need to be able to interpret visual images, video, film, charts, and tables (Lemke, 1998), navigate through complex and continually changing systems of information, (Leu, 1999), make decisions concerning when to read a definition or an explanation (Venezky, 1994), distinguish relevant and reliable information and make connections between discrete bodies of information and their relative importance (Synder, 1998; Landow, 1992), and monitor their reading in order not to become distracted from their reading purpose.
Hypermedia texts are mostly organized in a nonlinear manner because information is provided through nodes or links which may take readers to different types of glosses or annotations to help them understand the text better. Hypermedia glosses may be presented before or during reading; they may function to highlight or clarify important points or simply to provide lexical or syntactic information; their focus may be textual or extratextual; they may be provided within the body of text or outside the text; they may come in the form of text, images, sound recordings, or videos (Roby, 1999). Glossing is particularly useful in second language reading (L2 reading); words or phrases that are judged to be outside learners' current competence may be explained through glosses or annotations (Widdowson, 1984).1 Thus, a given text may be made comprehensible for L2 readers without reducing its authenticity.
REVIEW OF RELATED LITERATURE
Few research studies have investigated ESL learners' use of computerized reading glosses. For instance, questions such as "Do ESL learners find computerized glosses useful for reading comprehension?" or "What types of glosses do they prefer and find useful?" remain largely unanswered. One tool that goes a long way to facilitate research is the use of a tracking tool which allows researchers to follow what learners actually do when they are engaged in a particular hypermedia learning task. Tracking learners' interaction with the text may help researchers trace and explore learners' reading strategies (Blake, 1992) and provide valid means for process-oriented research (Hulstijn, 1993). Thus, such technology may provide insights into both the product and the process of learning (Collentine, 2000).
Aust, Kelley, and Roby (1993) investigate fifth-semester Spanish learners' preferences regarding computerized and conventional dictionaries, which were referred to as hyper-references and paper references, respectively. The researchers found that participants who had access to bilingual hyperreferences made more consultations than those who had access to monolingual hyperreferences, bilingual paper references, or monolingual paper references. Monolingual hyperreferences were also used more frequently than the bilingual and monolingual paper references. However, whether references were bilingual or monolingual did not make a difference when they had access to conventional references. Thus, the study suggests that learners prefer computerized bilingual dictionaries for second language reading. However, no differences among the groups were found with respect to reading comprehension.
Roby (1999) reports on an experimental study which investigated Spanish learners' use of paper and computer dictionary and glosses. While the dictionaries provided lexical information, the glosses contained "the meaning of an item in context." Roby found that learners who had access to computer dictionaries looked up substantially more words than those who had access to paper dictionaries. Moreover, the participants in the dictionary + gloss condition read the text in less time than those in the dictionary-alone condition. This study did not find any differences among the groups on reading comprehension. Unfortunately, insufficient numbers of such studies do not allow us to make any generalizations about L2 learners' use of computerized glosses and their attitudes about computerized reading.
Lomicka (1998) describes a pilot study which investigated the use of various types of glosses (i.e., images, references, questions, pronunciation, and translations in English) by 12 intermediate-level French learners. The tracker data revealed that the participants preferred definitional glosses to other types of glosses. Despite the small sample size and lack of statistical evidence, this study is useful because of its attempt to explore learners' preferences for different types of glosses. The purpose of the study presented here is to systematically explore intermediate and advanced ESL learners' interaction with a hypermedia text. The study investigates the following two questions: (a) What types of annotations do learners prefer to use when they are engaged in reading a hypermedia document? and (b) Are there any differences between intermediate and advanced learners with regard to their use of hypermedia annotations?
METHODOLOGY
Participants
Eighty-four participants learning English for academic purposes at the University of Arizona took part in this study. They came from a variety of language and cultural backgrounds: Arabic (29), Spanish (16), Japanese (13), Korean (10), Chinese (5), Indonesian (2), Portuguese (2), Thai (2), Bulgarian (1), Turkish (1), German (1), Pular (1), and Vietnamese (1). The average age of the participants was 24.06 years, ranging from 17-40 years of age. Students were placed at proficiency levels based on their performance on the Listening and Structure sections of a standardized placement test: A Comprehensive English Language Test for Learners of English. There were 34 intermediate- and 50 advanced-level learners.
Materials
An electronic reading text entitled "Stephen Hawking's Universe" was adapted from the Public Broadcasting Service web site (www.pbs.org/wnet/hawking/html/home.html). In order to aid reading comprehension, the text was annotated using multiple forms of digital media created by using Macromedia Director, version 7.1.2. Based on the interactive theory of reading (Rumelhart 1977; Kintsch & Van Dijk, 1978; Bernhardt, 1991), the annotations were aimed at facilitating both top-down and bottom-up processes. Thus, they were categorized as (a) textual annotations—those providing information about the text such as definition of a word and its pronunciation, or (b) contextual annotations—those providing background information about the topic.
Textual annotations were provided within the body of text, whereas contextual annotations were provided outside the text. When participants clicked on a highlighted word, phrase, or a background information button, they could see in what form or forms of media the information was available (i.e., text, graphics, sound, or video). They were then able to choose and view as many annotations as were provided in the hyperactive link (see Figure 1).
Textual annotations involved three types of media: text, sound, and graphics. Text provided the definition of a word or phrase and its grammatical form—noun, adjective, verb, and so forth. Sound annotations provided the pronunciation of a word or phrase via audio recordings. Graphics annotations provided photos or drawings to illustrate the word, notion, or idea. Contextual annotations were of four types: text, sound, graphics, and movie. Text provided extra information on the topic in textual format. Sound involved audio recordings related to the topic. Graphics involved pictures or drawings related to the topic. Finally, movie involved digital videos about the topic.
Sample Hyperactive Reading Passage
0x01 graphic
In summary, the annotations are categorized as follows:
1. textual text annotations: dictionary definitions of words,
2. textual audio annotations: pronunciations of words,
3. textual graphics annotations: pictures or drawings describing the meanings of words,
4. contextual text annotations: extra information about the topic in the form of text,
5. contextual audio annotations: sound recordings providing extra information about the topic,
6. contextual graphics annotations: pictures, drawings, or photos providing extra information about the topic, and
7. contextual video annotations: digitized movies providing extra information about the topic.
The software tracked every interaction of the readers with the text, including which annotations readers had chosen to view, how much time (in seconds) they spent on a particular annotation, and the order in which they selected the annotations. This information was saved in a log file for each reader.
Once the participants finished reading the text, they were given a 14-item reading comprehension test on the computer. This test consisted of short-answer, multiple-choice, and open-ended questions. Participants did not have access to the reading text during the test, but they were allowed to take notes while reading the text and use their notes during the test. Upon quitting the test, each participant’s responses to the questions were also saved in a log file.
Procedures
The participants took part in the study during their regular class periods. They were told that this study was aimed at exploring their reading strategies in a computer environment. It was also emphasized that participation in the study was voluntary and that it would not affect their performance in the class.
The data collection consisted of two major phases. The first phase lasted two class periods. During this part of the study, participants were given a demonstration on how the software worked. Then, they were asked to read the text for its content and use the annotations to help them understand the text better. They were told that they would be given a comprehension test when they finished reading, that they would not have access to the text during the test, and that they should take notes on the paper provided. Although the comprehension test did not include any information from the annotations, this was not indicated to the participants so that they would not opt to read the text only instead of using the annotations. When they quit the software, two log files were saved on the hard disk: one for the participants' interaction with the text and the other for their answers to the comprehension test questions. Finally, some of the participants filled out a background questionnaire the same day, while others did so the next day. The background questionnaire involved questions regarding the demographic information about the participants as well as their perceptions about the usefulness of annotations.
The second phase of the study involved interviews with 20 volunteering participants within three days after they took part in the study. The goal of the interviews was to obtain in-depth data about the participants' attitudes about reading a text in a hypermedia environment and the usefulness of the annotations.
Data Analysis and Results
Before discussing the results of the analyses, a brief discussion about the difficulty of the text is in order. The data regarding the difficulty of the text were obtained from the questionnaire the participants completed after they participated in the study. The participants were asked to rate Vocabulary, Grammar, and Content from 1 to 5, 1 being very easy and 5 being very difficult.
Independent-samples chi-square tests were conducted in order to ascertain whether the differences between the two groups were statistically significant. Results indicated that the groups were statistically different in their ratings of Vocabulary (χ2 = 9.97, p =.006) and Grammar (χ 2 = 8.25, p = .01), but not Content (χ 2 = 4.16, p = .12). In other words, the intermediate group considered Vocabulary and Grammar significantly more difficult than the advanced group. However, the difficulty level of Content was similar for both groups.
Since the difficulty level of certain aspects of the text was different for the groups, it was hypothesized that the two groups would interact with the text differently. In other words, they would differ in their use of the annotations incorporated into the text. The frequency with which the annotations were accessed, the amount of time spent on annotations, as well as the total amount of time spent reading the text were variables investigated with regard to participants' interaction with the text.
Frequency of Access to Hypermedia Annotations
The number of clicks made by the participants to view the annotations determined the frequency of access to annotations. Participants could click on a given annotation as many times as they wished. Whether there were any differences in each group with regard to access to each type of annotation and whether the two groups differed in their access to annotations were analyzed by a two-way mixed-design ANOVA with one grouping variable (proficiency level) and one within-subjects variable (annotation type). Hence, language proficiency was the independent variable with two levels (intermediate and advanced), whereas the annotation type was the dependent variable with seven levels measured on an interval/ratio scale.
In summary, contextual annotations were accessed significantly more frequently than the textual annotations. Among the textual annotations definitions of words were accessed the most, and among the contextual annotations video and text annotations were accessed the most.
The Amount of Time Spent on Annotations
The determination of the amount of time participants spent on hypermedia annotations was based on the amount of time they viewed the annotations. Since the annotations appeared only while the mouse button was down, the amount of time the participants kept their hands on the mouse was considered to be the time they viewed a given annotation. Annotations viewed for less than one second were not tabulated
Total Time Spent on Reading the Text
The results of the independent samples t test indicated that the difference in the total amount of time spent in reading the text between the two groups was significant (t(82) = 2.78, p = .007). Thus, the intermediate group spent significantly more time on reading or studying the text than the advanced group.
Performance on the Reading Comprehension Test
The performance of the groups on the reading comprehension test was compared through an independent samples t test. Results of the descriptive statistics showed that the mean score for the intermediate group was lower (M = 9.16) than that of the advanced group (M = 11.72). The highest possible score on this test was 22. Both distributions were normally distributed, and variances between the groups were found to be homogenous.
 Insights from the Questionnaires and Interviews
A qualitative investigation through the questionnaire and interviews provided valuable information regarding the participants' perceptions of the usefulness of annotations and their experience of reading in a hypermedia environment.
All the participants who were interviewed indicated that they enjoyed reading on the computer; it was different from the readings they did in class. They defined reading on the computer as "more interesting," "easier," and "comprehensible." For all of the participants, definitions of words were "necessary" or "essential." Some considered pronunciation not very important to understanding the text, but, for some, it was still useful because, as stated by Participant 4, "Pronunciation is one of the most difficult things in English. Any opportunity to improve pronunciation is very important." Finally, for some participants contextual annotations offered too much information, and they thought it was not necessary to use the contextual annotations to understand the text. For instance, Participant 7 viewed most of the contextual annotations, but he also stated: "My new understanding of the universe is from the reading, not from the movies or picture."
Summary
The participants did not agree on the usefulness of contextual annotations in aiding reading comprehension, but they did agree that word definitions were essential. Interviews indicated that the participants with prior knowledge utilized contextual annotations differently from those without prior knowledge. For those with prior knowledge, extra information about the topic was utilized because of "curiosity" or "interest," while those without prior knowledge utilized them to understand the concepts better.
The amount of time it took the intermediate group to read the text was much longer than the advanced group. However, the advanced group performed better on the reading comprehension test. Thus, despite the efforts of the intermediate group to understand the text by utilizing more annotations and spending more time on studying or reading the text, they still were not able to reach the level of performance displayed by the advanced group.
DISCUSSION
This study provided evidence for the positive impact of hypermedia on the learners' perceptions of their reading experience. The provision of authentic input and interaction between the reader and the text were found to be important features of hypermedia, leading to positive attitudes. The authentic input aroused the readers' interest because the linguistic input was contextualized and, therefore, easier to understand. As indicated by one of the participants, they were able to fulfill their curiosity about "what Stephen Hawking looked like and how he sounded."
Interviews revealed that the participants enjoyed interacting with the text at their own pace and selecting information based on their own needs and interests. The individualized reading gave them control over the reading process. Hence, the interactivity between the reader and the text motivated the participants in the reading process.
Finally, an examination of the specific types of annotations used by the learners revealed that both intermediate and advanced learners preferred word definitions in order to decode the text. They preferred visual information (i.e., graphics and videos) to get schematic information about the topic. The usefulness of these annotations was rated highly by the participants.
Pedagogical Implications
Interactivity, authenticity, and multimodal learning are important characteristics of a hypermedia environment for L2 reading. Interaction between readers and text provides individualized learning (Soo, 1999) and promotes learner autonomy (Healey, 1999). Learners get a chance to read at their own pace, select information based on their interests and needs, and take responsibility for their learning. As one of the participants put it, students may not be able to ask questions in a classroom environment because they feel shy or they cannot keep up with the pace of the class. These problems disappear in a hypermedia environment. Learners take control over their reading and make decisions about the amount of time they spend on reading, the pace of their reading, and the path they choose to construct a meaning. Moreover, students get actively involved in the reading process. They take the initiative to select the relevant information in the process of meaning construction. Hence, reading becomes a more personal and meaningful process for learners because they create their own meaning based on the way they interact with the text.
Provision of authentic input poses less challenge for L2 readers in a hypermedia environment because they are not only presented with authentic language but also with the means to deal with it. Annotations were considered to be highly useful for reading comprehension in this study. Easy access to annotations through multiple forms of media makes the reading of authentic texts more manageable and motivating for L2 readers. The participants of this study complained to their teachers why they could not read "like this all the time."
Limitations of the Study
This study has several limitations, which may caution us about the results obtained. First, the target population was ESL students learning English for academic purposes. The sample consisted of intermediate and advanced level ESL learners who were enrolled at a specific US university in a given semester to learn English for academic purposes. Unless the study is replicated in other learning contexts with different samples, the findings cannot be generalized to the general target population. Second, the reading comprehension test used in this study caused several problems. Some participants perceived that the whole purpose of the study was to test them on their reading ability. These participants did not utilize the annotations because they wanted to spend more time on the test. They also indicated that had there not been a test, they would have interacted with the text differently. Third, other factors might have influenced learners' interaction with a hypermedia text such as reading goals, learner styles, reading strategies, experience with computers, and reader's interest in the topic. These factors were not investigated in the study presented here, and they may be more related to performance and annotation use. Finally, the study did not investigate whether the difference on reading comprehension between the two groups was due to proficiency level, annotation use, or other factors. A more controlled study would be necessary to investigate the effectiveness of hypermedia annotations for reading comprehension.
Recommendations for Future Research
This study was exploratory in nature. Participants' interaction with text was not controlled for the sake of simulating a real-life task. However, an experimental study, which controls access to annotation types and investigates its relationship with reading comprehension may provide us with the true effect of annotations on performance. Second, in addition to experimental data, more qualitative data would also provide deeper insights into the reading process. For instance, using think-aloud protocols while learners are interacting with the text would provide information about the students' thinking processes and reading strategies in a hypermedia environment. Third, a longitudinal study in a context, where hypermedia is integrated into the curriculum and where learners are experienced users, would provide more valid findings than a study conducted at a given time with learners of varying computer experience. Finally, considering other variables such as learner styles, motivation, cultural background, gender, reading strategies, and how these variables may affect annotation use would add substantially to our understanding of the phenomenon.

Using Electronic Mail as a Medium for Foreign Language Study and Instruction


Ken R. Lunde
University of Wisconsin-Madison
Abstract:
This article describes how electronic mail can be used to send and receive foreign character sets, using the Japanese character set as an example. Electronic mail is fast, inexpensive, and can be stored, modified, and printed. This modern communication tool can be used to accelerate the traditional penpal process, and can act as a medium for instruction through correspondence courses. In addition, other computerized information, such as computer software and digitized speech can be sent using electronic mail. This opens many doors for future trends in computer-aided instruction (CAI).
Introduction
Electronic mail is an efficient means of communicating at the local and international scale. While it is easy to send text which uses only the 94 printable American Standard Code for Information Interchange (ASCII) characters, character sets which use more than these 94 characters pose problems. This article illustrates how these problems have been solved for the Japanese character set which contains nearly 7,000 characters, and how this ability to send foreign character sets can be applied to foreign language study and instruction.
This article has been divided into three parts to allow readers to skip over parts which are of less interest to them. Part I describes what electronic mail is and its advantages over other forms of communication; all readers are encouraged to read through this. Part 2 contains technical information for those who are interested in how Japanese is transmitted using electronic mail. Part 3 contains information on how electronic mail can be used for foreign language study and instruction, and examples of electronic mail in practical usage.
1. Electronic Mail
Electronic mail is a modern way to communicate across campus or across the globe. It makes use of computer networks to relay messages to their respective destinations. This adds a fourth method of communication to the three listed below:
Telephone
Facsimile
Conventional Mail
There are three factors to consider when communicating: speed, cost, and ease of storage for future reference, modification, or printing. For the sake of this article, I will suppose that we are communicating with Japan.
The first factor, speed, is found in facsimile and the telephone which communicate in real time. The next fastest method of communication is electronic mail. Electronic mail messages take anywhere from 30 minutes to six hours to travel from the United States to Japan. When sending electronic mail within the United States, the travel time is reduced to just a few minutes. The slowest method of communication is conventional mail which may take up to one week.
The second factor, cost, is an important consideration. Electronic mail may be the cheapest method of communication depending on where you are and who you are. Telephone and facsimile are more expensive since a direct long-distance connection is required. Electronic mail, for example, is entirely subsidized by the university for faculty and staff at the University of Wisconsin-Madison, but students are allowed to use only a pay-as-you-go electronic mail service.
The third factor, ease of storage, is found only in electronic mail. Each character is stored electronically with its own unique electronic value, and thus can be inserted into word processing applications for storage, modification, and printing. A facsimile is a remote photocopy, so the unique electronic values for each individual character are not recoverable.
There are other advantages as well. Just as text files can be sent using electronic mail, computer software (in encoded form) also can be sent. In fact, public domain software may be acquired by logging in to remote hosts using telnet (remote login) or file transfer protocol (ftp). In addition, there are also mailing lists which regularly broadcast information on electronic mail networks.
2. Technical Aspects
2.1. The Japanese Character Set
The ASCII character set is a seven-bit character set. In other words, each character is represented by seven binary bits, each bit having two possible values, on or off. This gives us a maximum of 128 representable characters. However, only 94 of these 128 ASCII characters are printable (numerals, symbols, and the Alphabet). The rest are unprintable, and include control characters such as <CR> (carriage return), <LF> (line feed), etc.
2.2. Electronic Transmission of the Japanese Character Set
The Japanese have developed what are known as Kanji-In and Kanji-Out escape sequences; the reason why Kanji-In and Kanji-Out are called escape sequences will become apparent in the following paragraph. The Kanji-In escape sequence commands Japanese terminals to begin to treat two ASCII characters as one Japanese character. This means that two bytes represent a single character. The Kanji-Out escape sequence, on the other hand, commands Japanese terminals to return to the normal ASCII mode, namely that one byte represents one character. Using two seven-bit bytes to represent one character creates a matrix in which the first byte is the row, and the second byte is the column. As mentioned above, JIS only uses the 94 printable ASCII characters in this coding scheme; this allows a maximum of 8,836 characters to be represented which can clearly handle the 6,877 standard Japanese characters. The remaining 1,959 spaces are used for user-defined or corporation specific characters.
3. Examples of Applications and Practical Usage
3.1. Applications in Foreign Language Study and Instruction
Readers of the CALICO Journal are aware that CAI software and video tapes are currently being used to improve foreign language study and instruction. Now I will describe how electronic mail communication technology can be applied to foreign language study and instruction.
One traditional method used by foreign language students to improve their abilities in their target language is that of obtaining penpals. All correspondence with these penpals is usually done by conventional mail, and consequently much time elapses between letters. Perhaps only three letters are exchanged per semester. Electronic mail accelerates this process considerably. For example, in a time span of only two months I have received eight electronic mail letters from one of my Japanese penpals. Electronic mail correspondence has many benefits: the chances that mail will cross are reduced, and students get more practice reading and composing text in their target language. The first issue is a logistical matter. Foreign language departments or individuals can purchase terminals which display foreign languages. Dedicated terminals are, however, not always required since there is a wide variety of software which allows computers to emulate terminals, even foreign language terminals. This means that computers which foreign language departments or individuals already possess may be used as electronic mail terminals. The second issue has a variety of solutions. For starters, many universities throughout the United States have set up exchange programs with foreign universities, and an electronic mail penpal exchange program could enhance this relationship. Penpals and their electronic mail addresses can also be obtained by replying to articles posted in electronic mail news broadcasts. The most difficult foreign electronic mail address to obtain is the first one. After that they pile up fast! The third issue is educational in nature. Students and teachers must learn the fundamentals of word processing in their target languages, and learn how to use electronic mail. In the case of Japanese, the principles of Kana-to-Kanji Conversion must be learned. This can be accomplished through frequent tutorial sessions.
Electronic mail also has the potential to be used for administering correspondence courses. Correspondence courses are traditionally sent to students using conventional mail, and are tailored for people who do not have the time to attend formal classes due to job conflicts. As with correspondence with penpals, electronic mail also accelerates this process. Correspondence courses which are administered by electronic mail are not limited to people living in one's own country, but could be administered on a global scale due to the speed of electronic mail.
 3.2. Examples of Electronic Mail in Practical Use
 I use electronic mail on a daily basis for sending and receiving Japanese text. I am currently corresponding with three Japanese acquaintances, all of whom work for corporations in or around Tokyo. Although the reason I correspond with them is for the sake of obtaining useful information, my Japanese language skills have improved significantly as a side-effect of frequent correspondence. Imagine how much I could benefit if I concentrated my efforts on improving my Japanese! I simply like the idea that I can communicate with Japan hundreds of times faster and at a much lower cost than conventional mail.
These topics are called newsgroups, and a small sample of them is listed below:
Newsgroup Name
Description of Newsgroup
fj.ai
Artificial intelligence discussions
fj.books
Books of all genres, shapes, and sizes
fi.comp.text
Text processing issues and methods
fj.followup
Follow-ups to articles in net.general
fj.general
*Important* and timely announcements of interest to all
fj.junet
General discussion about JUNET itself
fj.misc
Various discussions when there are no groups to match
fj.rec.animation
Discussions about animated movies
fj.rec.ham
Discussions about ham radio
fj.rec.idol
General topics about idols (i.e., popular singers)
fj.rec.misc
Recreational/participant topics not covered elsewhere
fj.soc.misc
Socially-oriented topics not covered elsewhere
fj.sys.mac
Discussions about the Apple Macintosh & Lisa
fj.sys.pc98
Discussions about NEC's PC-9800 series & other computers
Each JUNET News broadcast contains many articles. The poster's address is always included in the header of each article so that replies can be sent either directly to the poster or as an article to be posted in a future JUNET News broadcast.
Although I am not using electronic mail to its fullest educational potential, I do notice a significant improvement in my abilities in reading and composing text in Japanese. Those who wish to use electronic mail strictly as a learning tool will, I am sure, experience marked improvement in their target language skills.
I recently discovered that there is already a course which makes use of electronic mail to enhance foreign language study. This course is offered at the University of Toronto, Canada.
This course, called Computer-Assisted Composition in Japanese and Chinese, while not devoted solely to correspondence using electronic mail, does give students much practice in composing text in their target language. The students studying Japanese are given the option to use electronic mail to correspond with their peers in Japan, namely students studying English at the University of Tokyo, to exchange ideas and information.
The goals of this course are to motivate the students to use their target language creatively, to promote interaction in their target language, to enhance the cultural and intellectual component of foreign language study, and to improve the students' ability to read Chinese characters. The results were that the students displayed marked improvement in their character production, reading comprehension, and word processing skills in their target language.
For further information regarding this course, please contact Professor Kazuko Nakajima, Department of East Asian Studies, University of Toronto, 130 St. George Street, Toronto, Ontario, M5S IA5, Canada (BITNET: nakajima@utorepas).
Conclusion
The use of electronic mail for foreign language instruction does not, of course, replace formal classroom instruction, but instead complements it. Electronic mail is simply a modern communication tool which has the potential for use as an instructional or learning aid. Electronic mail communication can significantly improve students' reading and composition skills in their target language; spoken and listening comprehension skills can be improved in a classroom environment under the direct supervision of a native speaker who is qualified to teach, or in the foreign country itself.
Teachers should carefully consider how to use electronic mail as an instructional tool. If penpal correspondence using electronic mail is to be used as a motivational tool, it may be best to leave it out of the classroom environment since any evaluation or grading by an instructor may deter students from freely communicating in their target language.

Using a Computer in Foreign Language Pronunciation Training: What Advantages?

 


Maxine Eskenazi
Carnegie Mellon University
Abstract:
This paper looks at how speech-interactive CALL can help the classroom teacher carry out recommendations from immersion-based approaches to language instruction. Emerging methods for pronunciation tutoring are demonstrated from Carnegie Mellon University's FLUENCY project, addressing not only phone articulation but also speech prosody, responsible for the intonation and rhythm of utterances. New techniques are suggested for eliciting freely constructed yet specifically targeted utterances in speech-interactive CALL. In addition, pilot experiments are reported that demonstrate new methods for detecting and correcting errors by mining the speech signal for information about learners' deviations from native speakers' pronunciation.
INTRODUCTION
The ever growing speed and memory of commercially available computers, coupled with decreasing price, is making feasible the idea of creating computer-assisted language learning (CALL) that is speech-interactive. Even though the hardware conditions for an ideal automatic training system exist, can the same be said of state-of-the-art automatic speech recognition (ASR) and of our knowledge of the variability of the speech signal--the main stumbling block to higher quality speech recognition? Has the technology come far enough for systems to be able to teach pronunciation effectively? To answer these questions, we will first specify what is believed to contribute to successful language learning under a direct approach, drawing largely from principles described by Celce Murcia and Goodwin (1991). We will then list pedagogical recommendations following from this approach, such as providing language samples from many different speakers. Next, we will look at what speech-interactive CALL can do to help the classroom teacher carry out these recommendations. We illustrate with emerging methods for pronunciation tutoring from the FLUENCY project at Carnegie Mellon University (CMU) (Eskenazi, 1996), methods that support both articulation of phonemes and use of prosody—the intonation and rhythm of speech. The emphasis here is on pronunciation in the context of overall language learning. Proficient pronunciation is essential to language learning because below a certain level of pronunciation, even if grammar and vocabulary have been mastered, communication obviously cannot take place.
WHAT CONTRIBUTES TO SUCCESS IN TARGET LANGUAGE PRONUNCIATION?
Conditions for Success and Pedagogical Recommendations Based on Immersion
Many foreign language instructors agree (Celce Murcia & Goodwin, 1991) that living in a country where the target language is spoken is the best way to become fluent—a total immersion situation. They also generally agree (Kenworthy, 1987; Laroy, 1995; Richards & Rodgers, 1986) on which conditions of living abroad are critical to effective language learning:
• Learners hear large quantities of speech.
• Learners hear many different native speakers.
• Learners produce large quantities of utterances on their own.
• Learners receive pertinent feedback.
• The context in which the language is practiced has significance.
These conditions cover the external environment of language learning. From each we can extract recommendations for how to learn language under less than total immersion conditions.1 These recommendations cannot always be carried out in classroom contexts, thus presenting opportunity and motivation for ASR technology to complement teaching.
Recommendation 1
Learners hear large quantities of speech. For language learners who are not living in the country of the target language, immersion
courses consisting of six to eight hours daily are often the best alternative for exposing learners to the language. An ideal ratio of one student-one teacher would provide maximum speaking and feedback time. This situation is not always feasible. On the one hand, most students have other daily activities and, on the other, employing human teachers for eight hours a day is expensive (Bernstein, 1994). Moreover, immersion classes usually have five to ten students, and attending to individual needs reduces the amount of time the teacher speaks to the class.
Recommendation 2
Learners hear many different native speakers. This recommendation implies employing many native teachers with a diversity of voice types and dialects. However, the variety of native speakers available locally is limited, as is the number of people that a school can afford to hire. Traditional educational materials that promote wider exposure, such as audio and video cassettes, tend to be non-interactive, and their audio quality can degrade over time.
Recommendation 3
Learners produce large quantities of utterances on their own. Ideally, the student is in a one-on-one setting where the teacher encourages short conversations, constantly eliciting the student's speech. In reality, students in the classroom share the teacher's attention. The amount of time they spend individually producing speech and participating in conversation is thus reduced.
Recommendation 4

Learners receive pertinent feedback. In immersion contexts feedback that leads to correction of form or content may occur in two ways. Implicit feedback comes when speaker and listener realize that the message did not get across. A clarification dialogue usually takes place ("I beg your pardon?" "What did you say?"), ending with a corrected message that is understood. Less often, when culture and interpersonal context permit, the listener offers explicit correction, such as pointing out the error or repeating what the speaker said but with correction. In the ideal classroom, teachers offer implicit and explicit feedback at just the right times, keeping a balance between not intervening too often, to avoid discouraging the student, and intervening often enough to keep an error from becoming a hard-to-break habit. Expert teachers adapt the pace of correction—how often they intervene—to fit the student's personality. In reality, however, not all teachers use the same techniques and, in the classroom, are not always able to adapt these techniques to individuals. When class size increases, the amount of feedback to the individual student decreases.
Recommendation 5
The context in which the language is practiced has significance. Living in the country where the target language is spoken gives learners the practical need to speak. Their utterances have immediate significance. To accommodate this recommendation, the ideal language classroom includes fast-paced games and everyday conversations that create meaningful contexts (Bowen, 1975; Brumfit, 1984; Crookall & Carpenter, 1990). The student has to respond rapidly and utter new terms in these contexts. In reality, classroom size again reduces the individual learner's time for participating in such activities.
Conditions for Success and Pedagogical Recommendations Based on Structured Intervention
There are two additional conditions that appear critical for learning pronunciation but that do not follow from immersion—indeed, they follow from an assumption of structured intervention that departs from pure immersion: 1) Learners feel at ease in the language learning situation. Whereas the very young language learner perceives and tries out new sounds easily, older learners lose this ability. Embarrassment or fear may inhibit the learner from trying new sounds or even from speaking, whether in a total immersion or a classroom environment (Laroy, 1995). 2) There is ongoing assessment of learners' progress. Language learning appears most efficient when the teacher constantly monitors progress to guide appropriate remediation or advancement.
These conditions lead to pedagogical recommendations that may be particularly hard to carry out in the classroom.
Recommendation 6
Learners feel at ease. A key dimension of the learner's "internal" environment is self-confidence and motivation. Although there are techniques to boost student confidence in the classroom (Laroy, 1995; Krashen, 1982)—such as correcting only when necessary, reinforcing good pronunciation, and avoiding negative feedback—these may not overcome learners' inhibitions. Laroy (1995) finds that when students are asked in front of peers to make sounds that do not exist in their native language, these students tend to feel ill at ease. As a result, they may stop trying completely or may only make sounds from their native language. One-on-one teaching is important at this point, allowing students to "perform" in front of the teacher alone, not in front of a whole class, until they are comfortable with the newly learned sounds.
Recommendation 7
There is ongoing assessment. To adapt training to individual needs, the teacher ideally monitors each student's moment-by-moment progress, assessing strong and weak points, and judges where to focus effort next. The effective teacher takes into account what the student feels is useful, thus keeping students involved in their own progress (Celce Murcia & Goodwin, 1991; Laroy, 1995). In reality, classroom teachers cannot maintain steady monitoring of each student at this level of detail.
WHERE CAN SPEECH-INTERACTIVE CALL MAKE A CONTRIBUTION?
It is not feasible to carry out these seven recommendations fully in the traditional language classroom, given constraints on teaching time and materials. The ideal CALL system could help toward realizing these recommendations by providing individualized practice and feedback in a safe environment and sending back regular progress reports to the teacher (Wyatt, 1988). The human teacher must still do the high-level, subtle work of creating a positive atmosphere for the production of new sounds and stress patterns, explaining fine conceptual differences between a student's native language and the target language, and exploring cultural differences (Bernstein, 1994).
For each of our recommendations we will consider where automatic functions, in the form of both ASR and CD-ROM, can support the classroom. We draw examples from the FLUENCY project and from other systems featured in this volume.
CALL Can Help Learners Hear Large Quantities of Speech
With the decreasing cost and increasing capacity of computer memory and storage, CALL can offer users a choice of many prerecorded utterances. CD-ROMs afford high-quality sound and video clips of speakers, giving learners a chance to see articulatory movements used in producing new sounds (e.g., LaRocca, 1994). The teacher no longer has to find or record native speakers, although tools can be provided for teachers to add new speakers to the data set. The highly available digitized speech supplements the teacher's speech without incurring additional cost at each use. It also allows individualized access to particular samples of speech.
 ASR-Based CALL Can Help Learners Produce Large Quantities of Utterances on Their Own
Limitations of Traditional ASR-based CALL
A major problem in speech-interactive CALL, in commercial products especially, is that learners remain relatively passive (Wachowicz & Scott, this issue). Although learners may be asked to voice an answer to a question, this by design involves either parroting an utterance just presented or reading one of a small set of written choices (Bernstein, 1994; Bernstein & Franco, 1995). Learners get no practice in constructing their own utterances (i.e., choosing vocabulary and assembling syntax). The commercially available AuraLang package (Auralog, 1995), for example, is an appealing language teaching system that feeds to ASR the user's pronunciation of one of three written sentences. Each choice leads the dialogue along a different path. A certain degree of realism is attained, but students do not actively construct utterances.
Techniques for Extending the Limitations: Sentence Elicitation
The FLUENCY project has developed a technique that enables users of speech-interactive CALL to participate more actively in constructing utterances (Eskenazi, 1996). In traditional speech-interactive CALL, ASR works well because the system "knows" what a speaker will say and matches exemplars of the phones; it expects (pre-stored in memory) against the incoming signal (what was actually said). The technique developed in FLUENCY, by contrast, makes it possible to predict enough of what the speaker will say to satisfy the needs of the recognizer while giving speakers apparent freedom to construct utterances on their own. The technique is based on sentence elicitation, modeled on the drills used in the once prevalent Audio-Lingual Method (Modern Language Materials, 1964) and the British Broadcasting Company tutorials (Allen, 1968).
Several studies have addressed whether specifically targeted speech data can be collected using sentence elicitation (Hansen, Novick, & Sutton, 1996; Isard & Eskenazi, 1991; Pean, Williams, & Eskenazi, 1993). Results confirm that a given prompt sentence in a carefully constructed exercise elicits at most one to three distinct response sentences from normal speakers. Students can practice constructing answers to the same elicitation sentences as often as they wish, at no additional cost in teacher time and materials. Availability and patience are other qualities that enable the system to support our recommendation of having learners produce large quantities of utterances on their own.
ASR-Based CALL Can Provide Learners With Pertinent Corrective Feedback
Teachers often ask what type of corrective feedback speech recognition can furnish. This section will address two aspects of the question: whether and what types of errors can be detected successfully, and what methods are effective in telling students about errors and showing them how to make corrections.
Can Errors Be Detected? Phone Errors Versus Prosody Errors
Error detection procedures differ as follows. Phone-based errors are identified in forced alignment mode. Given an expected utterance, the recognizer takes the actual utterance and returns the placement in time of phones and words on the speech signal. By this method the learner's recognition scores can be compared to the mean recognition scores for native speakers—all uttering the same sentence in the same speaking style—and the learner's errors can thereby be identified and located (Bernstein & Franco, 1995). For prosodic errors, however, only duration can be obtained from the output of the recognizer. That is, when the recognizer returns the phones and their scores, it can also return the duration of the phones. Frequency and intensity, on the other hand, are measured on the speech signal before it is sent to the recognizer but after it is preprocessed. Intensity is usually obtained by using a technique known as cepstral analysis. Fundamental frequency is obtained from an algorithm that detects peaks in the signal and measures the distance between them. Speakers as individuals vary greatly on the three components of prosody. For example, some people speak louder or faster in general than do others. Thus, it is important that measures of the three be expressed in relative terms, such as the duration of one syllable compared to the next.
Phone Error Detection: A Pilot Study of ASR-based Comparisons of Native and Nonnative Speakers
Although researchers have been cautious about using ASR to pinpoint phone errors, recent work in the FLUENCY project shows that the recognizer can be used in this task if the context is well chosen (Eskenazi, 1996). Demonstrating this is a pilot study of native and nonnative speakers uttering responses in elicitation exercises.
Method
Ten native speakers of American English were recorded (5 male and 5 female) and 20 speakers of other languages (one male and one female from each of the following L1s: French, German, Hebrew, Hindi, Italian, Korean, Mandarin, Portuguese, Russian, Spanish).2 Expert language teachers were asked to listen to the sentences recorded by each speaker and to judge where there was an error, what it was, and how (and when) they would intervene to correct it. Teachers marked these judgments on phonemically labeled copies of the target sentences. The agreement between human teachers and ASR detection was used as a preliminary indication of the validity of automatic error detection. n speech-interactive CALL. The SPELL foreign language teaching system (Rooney, Hiller, Laver, & Jack, 1992) addresses both fundamental frequency, or pitch, and duration. Pitch detection, like speech recognition, is by no means a perfected technique. But Bagshaw, Hiller, & Jack's (1993) work on better pitch detectors for SPELL shows that algorithms can be made more precise within a specific application. This work compared the student's pitch contours to those of native speakers to demonstrate the informativeness of pitch detection. Pitch detection was incorporated into SPELL and the output interpreted in visual and auditory feedback for the student. SPELL assumes that suprasegmental (prosodic) aspects of speech should be tied to segmental (phonemic) information—for example, by showing pitch trajectories (contours over segments) and pitch anchor points (centers of stressed vowels). SPELL also addresses speech rhythm, showing segmental duration and acoustic features of vowel quality (predicting strong vs. weak vowels).
Tajima, Port, and Dalby (1994) and Tajima, Dalby, and Port (1996) have addressed duration. They studied how timing changes in speech affect the intelligibility of nonnative speakers and created remedial training supported by ASR. By using speech that is practically devoid of segmental content (ma ma ma Ö), they separate the segmental and suprasegmental aspects of the speech signal to focus on one aspect—temporal pattern training.
The FLUENCY project has looked at how to detect changes in duration, pitch, and intensity to find where a nonnative speaker deviates from acceptable native values. Prosody training in FLUENCY is linked to segmental aspects, with students producing meaningful phones. We aim to detect deviations independently of L1 and L2 so that if a learning system is ported to a new target language, its prosody detection does not have to be changed fundamentally. We have promising results from a pilot study, reported below, using hand-labeled features of the spectrogram.
Prosody Error Detection: A Pilot Study of ASR-based Comparisons of Native and Nonnative Speakers
Method
For the English sentence data recorded in the pilot study on phones, we additionally asked human teachers to mark the location and type of prosodic errors of each speaker on transcriptions of the sentences. We first examined the speech signal to determine whether the information used by teachers to detect errors could be characterized in the spectrogram. After examining phone-, syllable-, and word-sized segments, we developed three measures, one for each component of prosody. We compared these with human teachers' judgments of places where prosody needed improvement in each sentence and refined the measures until they showed close agreement with human judgments. These measures then define the features we want to extract automatically from the speech signal to diagnose where students need improvement.
Duration Results
The first measure was duration of the speech signal, measured on the waveform. The results of the duration comparisons are given in Figure 2. The duration of one voiced segment was compared to the duration of the preceding one ("ratio of seg1/seg2" on the vertical axis) to make the observations independent of individual variations in speaking rate. Pitch Results
The second measure we developed was the total number of pitch peaks present in the speech signal, calculated for each segment.4 Again, results were compared between neighboring segments. We were able to detect pitch deviations related to duration as well as independent of it. For example, mfrc raised pitch much higher on /EHK/ in "extra" than on the following vocalic segment /STRA/, probably because /EHK/ is also longer (see Figure 2). However, the speaker mpeg varied pitch independently of duration.
Intensity Results
The third measure developed, for intensity, was the average of all the cepstral values over a given vocalic segment. To address relative rather than absolute intensity, we compared these values segment-to-segment with those of neighboring vocalic segments, as with duration and pitch. The resulting curves and spread of speaker space, shown in Figure 3, differ in general aspect from the results in Figures 1 and 2. Outliers were indicated that matched teachers' judgments about relative stress centers in utterances. For example, msjh shows stress displaced within the "I/did/want" region, mbob displaced stress within "did/want/to," and msjh, among others, within the "ex/tra/in" region. The speakers' changes in amplitude appeared to be independent of duration and pitch.
Two-by-Two Comparison of Average Intensity (Amplitude) on Voiced Segments
0x01 graphic
Implications
Our pilot study suggests that the spectrogram can be mined for measures of speech prosody that have diagnostic value and are consistent with what expert teachers say they would detect and correct. We are now rendering these measures automatically detectable. Being separate from one another, the three measures of prosody, once analyzed in an utterance, could be expressed in visual displays for the learner that show pitch, duration, or amplitude. A learner's utterance could then be compared with a native speaker's utterance on each dimension to illustrate differences. Our results suggest that the components of prosody are not totally independent of each other. We saw this particularly in the dependency of pitch on duration. We suggest that correction first address the three components separately, then address their combined effect. Instruction could begin by exercising pitch and duration changes independently, then give practice on changing pitch and duration together.
An Argument for Early Prosody Instruction
Early prosody instruction, starting the first year of language study, could be a boon to learning both syntax and phone articulation. Because speakers prepare the syntax of a sentence they want to say at about the same point as they prepare prosody, incorrect word order will not fit the "song" that it is to be sung to. Self-correction then comes into play as students rearrange syntax to give a better fit to prosody. (Because the "song" is considered as a whole and the syntax as a concatenation of elements, the student should tend to rearrange syntax and not prosody.) Phones may benefit from early prosody training, for example, in the case of stressed and unstressed vowels in English. If a target vowel is unstressed and the Spanish speaker uses a tense (stressed) vowel that is close to the target in articulatory space, self-correction should follow because the speaker's longer tense vowel will not "fit the song" well. For example, the unstressed "this" in the sentence "I want this present" is shorter and softer than the surrounding vowels. Practice of correct prosody in this sentence should aid pronunciation of "this" by lessening emphasis on and shortening the / IH/ sound. Follow-up exercises could put "this" into new contexts, such as "This is yours," where the word is not so short and the speaker must make more effort to retain the shortened form just learned.
Effective Correction in Speech-interactive CALL
Learners' difficulties with phones and prosody, which our pilot studies suggest can be readily detected in the speech waveform, become targets for focused correction in CALL. The system that only detects pronunciation errors (e.g., parts of TriplePlayPlus by Syracuse Language Systems, 1994) is of limited aid. Learners will make random, trial-and-error attempts to correct the reported error. There may be little true amelioration and even negative effects if learners make a series of poor attempts at a sound. Such unsupervised repetitions could reinforce poor pronunciation to the point of becoming a hard-to-correct habit (Morley, 1994).
Effective correction requires that recognizer results be interpreted, as by putting them into a visually comprehensible form and comparing them to native speech. Our work in FLUENCY suggests that how recognizer results are best interpreted for instruction differs between phone correction and prosody correction. This suggestion stems from the fact that phones are different from one language to another while prosody is produced in the same way across languages. Whereas students must be guided as to tongue and teeth placement for a new phone, they don't need instruction on how to increase pitch if they have normal hearing: They only need to be shown when to increase and decrease it, and by how much.
Correcting Phone Errors
There has been some success in using minimal pairs—contrasting sounds in context in the target language, such as "I want a beet"/"I want a bit" (see Dalby & Kewley-Port, this issue; Wachowicz & Scott, this issue). Effective teachers often go further, with instructions on how to change articulator position and duration. This kind of instruction is important because if a sound does not already belong to a learner's phonetic repertory, the learner will associate it with a close speech sound that is in the repertory. For example, anglophones beginning to speak French typically hear and pronounce the French sound /y/ (in tu) as the English sound /u/ (in "too"); but they can be taught to use liprounding to approximate French /y/.
Automatic systems can teach articulator placement for new sounds, adding graphical views, for example, of the inside of the mouth (LaRocca, 1994). This instruction can be likened to gymnastics; the learner "feels" when the articulators are correctly in place and practices with the recognizer to confirm this. Learners can train their ears to recognize the new sounds and relate them to what they feel their muscles doing. Akhane-Yamada et al. (1996) suggest that learning to perceive sound distinctions helps in their production.
Phone articulation training can be L1-independent. A target vowel, for example, can be taught by starting with a close cardinal vowel (e.g., /a/, / i/, and /u/ have a high probability of existing in most L1s). A better solution, requiring more computer memory and linguistic knowledge, is to start with a close vowel in the learner's particular L1. Taking into account the learner's L1 can help anticipate errors and point to pertinent articulatory hints (Kenworthy, 1987). Thus, knowing that French has no lax vowels lets teachers of English to French speakers focus on how to go from a tense vowel to a close lax vowel ("peat" to "pit").
Correcting Prosody Errors
Based on work in FLUENCY, we propose that the visual display more than oral instructions will be critical to prosody correction. The key is for learners to see where the curve representing their production differs from the native speaker's curve. Prosody displays can benefit from the wealth of work on automated systems that teach the deaf to speak. For example, Video Voice (Micro Video, 1989) uses histograms to represent intensity (over time) and xy curves for pitch (over time). Duration is implicit in the time axis of the intensity histogram. Video Voice compares what the student says to a native speaker's prerecorded exemplar. For pitch the student sees the two frequency curves and, guided by hints, tries to increase
or decrease pitch at relevant points to come closer to the exemplar. Trials within the FLUENCY project confirm the importance of visual details to help learners understand the display, for example, using a continuous line as opposed to a divided contour for pitch.
ASR-based CALL Can Provide Significant Contexts for Language Practice
CALL can simulate authentic contexts using multimedia and multimodal displays in ways discussed elsewhere in this volume (e.g., Rypa & Price; Wachowicz & Scott). Learners can participate in one-to-one conversations with one or more simulated or videotaped interlocutors. The cue for the student to speak can be realistic, such as having a character on the screen turn head and eyes toward the user (or the camera).
ASR-Based CALL Can Put Learners at Ease
The computer can prove the ideal partner for putting a language learner at ease in speaking. Whereas the human teacher judges the student's production, the computer can be viewed as neutral. It can support continual practice of unusual sounds until students have enough confidence to go before others. The system becomes what Wyatt (1988) calls a collaborative tool rather than a facilitative one, with students assuming the role of judges of their own productions. This role not only has pedagogical backing (Celce Murcia & Goodwin, 1991) but can also benefit system performance. For example, if an exercise requires making a fine phonetic distinction that the recognizer detects poorly, the system can mislead and frustrate the student by giving errant pronunciation scores and, on that basis, deciding what to present next. However, if the system simply displays recognition results without pronunciation scores and allows students to decide whether they did well or need further practice, then ASR-derived error is less problematic. The student gains a sense of control over the chain of events but the teacher can still intervene to insist on more practice.
ASR-Based CALL Can Provide Ongoing Assessment
CALL today can enable rapid, constant assessment of the learner. The system can provide more details more rapidly than a teacher grading tests (Bernstein & Franco, 1995). The feedback given to the teacher can go beyond pronunciation scoring. In traditional computer-aided instruction, learners are scored right or wrong on a given question and the scores tallied at the end of the session. But for a system that gives visual data to help learners decide where to correct themselves, feedback to the teacher can include learners' own decisions as to their strong and weak points. For example, in a lesson on how to emphasize content words in utterances, if the learner decides to work on duration rather than pitch or amplitude, we can assume either that duration presented more of a problem or that the learner did not have time for the other two aspects. In any case, the teacher who receives the system's report can immediately test progress in the aspect the learner worked on and recommend what to work on in the next session.
Latency of response can also be measured (Bernstein & Franco, 1995) to obtain an even clearer view of where learners are having difficulties. Responses that took more time to formulate can be noted, as can progress in decreasing latencies over a session.
CONCLUSION
Speech-interactive CALL brings to pronunciation instruction a wealth of new, sometimes unforeseen, techniques. Increases in computer memory and storage for expanded exposure to many speakers and for multimedia corrective feedback can reproduce some of the advantages of total immersion learning. There is still much to be done. Teachers and computer scientists need to collaborate more closely to refine ASR-based tools and to invent and validate new teaching methods to build on the advantages of the new medium.