By Ana P. Biazon Rocha 

We teachers usually have to make several decisions at once when giving feedback on pronunciation: Should I interrupt the learner? Should I correct the error directly? Will the learner understand what I am correcting? And, perhaps most importantly, what will help them actually improve? Based on that, and drawing from Mroz (2018), and from the latest IATEFL PronSIG webinar with Carol Johnson and Walcyr Cardoso, back on 22 August 2026, ‘From Feedback to Learning: Integrating Speech Recognition in Pronunciation Teaching and Assessment’, in this post, we will rethink how to provide feedback on pronunciation with the possible integration of speech recognition.

1. Focus on intelligibility

Feedback on pronunciation is paramount in language learning. As we frequently emphasise here in our blog, pronunciation is not simply about sounding ‘right’, but mainly about communication. According to Derwing and Munro (2015), intelligibility, the extent to which a listener understands what a speaker intends to say, should be distinguished from accentedness. Therefore, the goal of pronunciation teaching should not be eliminating a learner’s accent, but helping them become clearly understood (For an overview on terms such as intelligibility and accentedness, please check this previous post). 

This shift in perspective has important consequences for feedback. Instead of treating pronunciation feedback as a process of identifying and correcting mistakes, it should be a way of helping learners notice how their pronunciation affects communication, decide what they want to change and practise making that change (Rocha, 2026). 

2. Systematic and useful feedback on pronunciation

In many classrooms, pronunciation is usually dealt with when a problem happens to arise: a learner says something unclearly, and the teacher responds. This kind of incidental feedback can be useful since it is immediate, connected to a real communicative need and does not necessarily interrupt the lesson for long. However, if pronunciation is addressed only when errors occur, it risks becoming an ‘add-on’ rather than an integral part of language teaching. Hence, a more systematic approach to pronunciation teaching, in which pronunciation is deliberately incorporated into lesson planning, is needed (Rocha, 2026).

In addition, while corrective feedback can help learners notice features of language that they might otherwise overlook, pronunciation feedback is not automatically effective simply because the teacher has corrected something. In other words, correcting pronunciation does not necessarily mean improving it. As a result, pronunciation feedback should be:

  • clear: students need to understand what the feedback means
  • relevant: feedback should be provided according to the lesson outcomes and/or learners’ goals/needs
  • consistent and regular: feedback should be an integral part of everyday lessons 

Another key factor is practice. Feedback without an opportunity to try again is easily forgotten. After receiving feedback, learners need opportunities to use the target pronunciation feature again through activities such as dialogues, recordings, drills, games or other purposeful tasks (Rocha, 2026).

3. What can speech recognition add?

Automatic speech recognition (ASR) converts spoken language into text. A simple ASR example is when we use the microphone function in our phones to transcribe a message rather than typing it. In language learning, this creates an intriguing feedback mechanism: learners speak, the system produces a transcription and the transcription provides evidence about what the system appears to have understood. 

Working with 16 university learners of French, Mroz (2018) investigated how learners used ASR on their smartphones to develop awareness of the intelligibility of their own speech. The study found that ASR could credibly simulate, to some extent, how speech might be understood by another speaker, helping learners identify successes and gaps in their intelligibility.

Based on that, we can observe that the transcription is not simply a score, but something for learners to look at, interpret and respond to.

Imagine a student is asked to say: ‘She lives on a farm’.

They know what they intended to say, but the speech-recognition system produces a different word. The mismatch creates a small but potentially powerful moment: ‘I said one thing, but the system heard something else’, making the ‘intelligibility gap’ visible.

Accordingly, Mroz (2018) describes learners engaging in cycles of speaking, tracking the ASR output, noticing possible problems, trying again and making adjustments, which can support individualised formative self-evaluation. For this reason, Mroz (2018) and Johnson and Cardoso agree that ASR can be used as a feedback tool rather than a mere pronunciation judge.

4. Confirmation is feedback too

One particularly useful insight from Johnson and Cardoso’s webinar about the integration of ASR in pronunciation teaching is that feedback does not always have to tell learners that they are wrong. Many times the general concept of feedback and its most famous type, corrective feedback, overlap: the learner makes a mistake, teachers help them correct it. In contrast, if ASR transcribes what the learner intended to say, that can provide confirmative feedback, which means that what they are doing is right. This is essential to boost their confidence and provide reassurance on their learning process, instead of only focusing on mistakes or problems. As one of the participants in Mroz’s (2018) study notes: ‘It’s a confidence booster, just knowing that it does actually understand a lot of what I say. It does make me feel good about myself a little bit! (p. 626)’.

And here’s some food for thought: how often do we provide more chances for confirmative feedback rather than corrective feedback in our pronunciation lessons? 

5. Integrating ARS as controlled practice

Johnson and Cardoso emphasised that ASR may be particularly useful during controlled practice, when learners are working intensively on a particular pronunciation feature. For example, imagine we are focusing on the /ɪ/ vs // distinction (ship/sheep). We can give learners two possible sentences:

A. The sheep is on the farm.
B. The ship is in the harbour.

The learner chooses one secretly and says it into the ASR system. If they say: The sheep is on the farm, and ASR produces: The sheep is on the farm, that is confirmative feedback. On the other hand, if the learner intended: The sheep is on the farm, but ASR produces: The ship is on the farm, the learner now has something to investigate. They can try again, perhaps exaggerating the vowel contrast, and see whether the transcription changes.

In this case, the learner is not just trying to produce an isolated sound correctly but communicate a particular meaning: ‘Did my pronunciation allow the listener/system to identify the word I intended?’ This is much closer to the communicative goal of pronunciation teaching.

Besides, the fact that the learner can try again is a great learning opportunity, as they can figure out what they were doing differently and experiment with their pronunciation until they achieve the intended result. Learners then become more autonomous: they can practise using ASR independently, while the teacher can remain available to help them interpret the feedback and address difficulties that the technology cannot explain. As mentioned previously, Mroz (2018) found that learners could use ASR as a diagnostic tool to identify both successes and gaps in their intelligibility, suggesting that it could support more autonomous pronunciation practice. Thus, the aim is not simply to get the ‘right’ transcription, but the learning that happens between one attempt and the next.

Based on Mroz (2018) and Johnson and Cardoso’s webinar, a five-stage cycle can be a useful starting point if you want to experiment with ASR in your lessons, as described in the image below.

A visual representatrion of a simple classroom cycle for ASR-supported pronunciation practice

Image 1: A simple classroom cycle for ASR-supported pronunciation practice

6. Is ASR the solution to our problems?

Although ASR can show learners what was transcribed, unlike teachers, it cannot necessarily explain why the transcription occurred, whether the pronunciation is appropriate in a particular communicative context, or what the learner should do next. This is especially important because ASR is not perfectly accurate, and recognition problems may lead to anxiety and frustration. As a participant in Mroz’s (2018) study mentions: ‘I have to tell myself not to feel discouraged’ (p. 626). Many times the learner is pronouncing something accurately but other factors such as background noise, quality of the equipment used, how close or distant the learner is from the computer microphone, etc., might interfere in the transcription. ASR also looks at the semantics of the sentence being produced. So sometimes it can change or even correct what the learner uttered, not providing a true transcription of what the learner really pronounced. 

Similarly, Mroz (2018) mentions that ASR may sometimes identify unintelligibility where a human listener would actually understand the learner. Therefore, ASR is a technological tool, not a human teacher who can understand and interpret what is happening. 

As mentioned by Johnson and Cardoso, another important thing to consider is that ASR works well with segmental features such as vowel and consonant sounds. However, these are not the only pronunciation features that should be focused on. Suprasegmentals such as stress, intonation and linking sounds are critical for intelligibility and can deeply impact communication (For an overview on how to teach suprasegmentals, please check this previous post). 

In conclusion, pronunciation teaching should be about helping learners communicate effectively, so when teachers prioritise intelligibility, pronunciation feedback can become a tool for learner empowerment rather than simply correction (Rocha, 2026). Speech recognition can be one more way of making this learning process visible, helping learners engage with the feedback given: ‘You get to see what another person would hear you as’ (author’s italics) (Mroz, 2018, p. 628).

Finally, it is important to mention that this post presents a brief overview on speech recognition in pronunciation teaching and learning. For further information, please check McCrocklin and Levis (2026) and Mroz (2018). 

For more discussions, ideas and tips on how to teach pronunciation, please use our PronSIG blog as your guide! Don’t forget to leave your comments below and follow PronSIG on social media.

References

Derwing, T., & Munro, M. (2015). Pronunciation fundamentals: Evidence-based perspectives for L2 teaching and research. John Benjamins.

McCrocklin, S. & Levis, J. (2026). Automatic speech recognition and pronunciation learning. Language Teaching, 59(3), 293-309. doi.org/10.1017/S0261444825100852 

Mroz, A. (2018). Seeing how people hear you: French learners experiencing intelligibility through automatic speech recognition. Foreign Language Annals, 51(3), 617–637. https://doi.org/10.1111/flan.12348  

Rocha, A. P. B. (2026). Balancing systematic pronunciation teaching and feedback. Speak Out! (74), 46-54.