AI-Generated Audio Summary
When people learn that our speaking apps run on advanced AI, they usually ask which model we use. It is a fair question, and yes, we use strong, current models, the best available for real-time voice. But it is also the least interesting part of the story.
Here is the thing. Choosing a capable model is the easy part. Any competent developer can connect to one. What decides whether a learner actually improves is something else entirely: whether the system understands how people learn to speak a second language. That knowledge does not come from the model. It comes from us.
What “Harnessing AI” Really Means
There is a phrase that has become common lately, and it captures how we work: harnessing AI. A large language model (an LLM) is raw capability on its own, fluent but unguided. To turn it into a tutor a student can trust, we build around it: guardrails that keep conversations safe and on task, tools that handle live audio and transcription, a curated knowledge base the tutor draws on so her guidance stays accurate and aligned to the exam, what engineers call retrieval, or RAG, memory so she recognizes a learner across sessions, and careful prompt engineering that shapes how she behaves in a real conversation.
We will spare you the full recipe. The databases, the retrieval, the engineering underneath are real work, but most learners do not need to know how the pipes are laid. Here is what matters: the model is the engine, and the teaching is everything we build around it. Every choice, from how long the tutor speaks to when she corrects an error, is a design decision grounded in how people learn.
The Teacher Inside the Machine
Here is a concrete example. Recently we refined how Verónica, our AI tutor, adapts to each learner in the speaking practice app. While testing the app, we ran straight into a familiar problem: she would set up an activity by explaining several steps at once, and by the time she finished, the first step was hard to recall.
Anyone who has taught a second language will recognize this instantly. It is not a flaw in the model. It is cognitive load. A learner working in their second language, their L2, is already spending mental effort on comprehension, vocabulary, and pronunciation. Stack a multi-step instruction on top of that, and the instruction itself becomes the obstacle.
So we redesigned how Verónica paces a conversation. She now gives one step at a time, keeps her turns short, checks that each point landed before moving on, and calibrates her language to the learner, simpler and slower when someone is struggling, more demanding as they grow. If you have studied second language acquisition, you will recognize the principles at work: comprehensible input, Krashen’s idea that we acquire language best when it sits just beyond our current level; scaffolding, which grows out of Vygotsky’s work on what a learner can do with support; and giving the learner room to produce language, not only receive it. We did not bolt those ideas on afterward. They are how we decided the tutor should behave from the start, even if getting the execution right took some refining.
Lowering the Affective Filter
There is another reason a patient AI partner helps, and it is one of the most important ideas in language learning. Anxiety, embarrassment, and the fear of getting it wrong raise what the linguist Stephen Krashen called the affective filter, an emotional barrier that blocks new language from taking hold. A learner who is afraid to make a mistake in front of classmates, a professor, or an examiner is a learner whose mind is only half available for acquisition.
Verónica lowers that filter. She never sighs, never rushes, never makes a learner feel judged. Someone can stumble through the same hard scenario five times at eleven at night, try phrasings they would never risk in a real room, and simply keep going. Mistakes become practice instead of embarrassment. That safety is not a pleasant extra; it is a condition for learning, and we designed for it on purpose.
Honest Feedback, Not a Flattering Number
Here is another design decision, and it goes to the heart of trust. Every speaking exam judges more than words. It judges delivery: pronunciation, pace, the small hesitations that tell an examiner whether a candidate is at ease. It is tempting to hand the learner a tidy scorecard with a number for each of those traits, so we built exactly that. Then we changed it, on purpose: the delivery feedback now comes from the app actually listening to the student’s recording, not from guessing at a transcript.
Why did we change it? The old scorecard tried to grade pronunciation from the transcript, and a transcript is only text. You cannot hear pronunciation in text. A number like “pronunciation: two out of three” pulled from a transcript is not measuring anything; it only looks precise, and a learner who trusts it practices the wrong thing. That is the number we removed.
So we drew a line. The written feedback now scores only what the text can honestly show: how well the response met the task, the development of ideas, grammar, vocabulary, academic register, and whether the learner stayed in the target language. Pronunciation and delivery are assessed separately, from the actual recording, where the app can truly hear how the learner sounded, which is exactly how a real examiner works. After a spoken task the learner gets an honest read on their pronunciation, their rhythm, and their audible hesitations, drawn from the audio, not guessed from a page.
The rest of the feedback follows the same rule: say only what you can defend. Every observation quotes something the learner actually said. We cap the growth areas at three, because a page of twenty corrections is cognitive load again, not help. And when a learner slips into their first language or drops into a casual register, we do not scold; we show the natural target-language phrasing and the more academic version side by side, as a skill to build. Feedback a learner cannot trust is worse than no feedback at all, and no model, however advanced, decides that for you. Knowing it is the educator’s job.
Two Vocabularies in the Same Room
Building this well means moving between two vocabularies in the same breath. One is the language of engineers: models, guardrails, latency, tokens. The other is the language of teachers: comprehensible input, the affective filter, cognitive load, register, the L2 learner. At most technology companies, those two vocabularies live on different floors. At Enabling Learning, they live in the same room, and often in the same person.
That is the real advantage. We are not a software company that hired an education advisor, and we are not educators who handed the technology to someone else. We are classroom teachers, instructional designers, and researchers in bilingual education and second language acquisition who also build the software. We come at the same problem from several angles at once, and the product carries all of that expertise inside it, whether or not the learner ever sees it.
Why We Do Not Chase Every New Model
The field moves quickly. A newer, more capable voice model seems to arrive every few weeks, and it is tempting to adopt each one the day it ships. We do not work that way. Keeping up with the technology and chasing it are not the same thing. We stay current, and we use models capable enough to truly listen to a learner’s voice, but we adopt each one because it serves the learner, not because it is new.
When a new model appears, we evaluate it through a pedagogical lens first. The question is not whether it is newer or scores higher on some general benchmark. The question is whether it understands a second-language speaker with an accent, responds naturally in the target language, and helps the learner more than what we already have. We test it carefully, and we adopt it only when it genuinely serves students. Sometimes the newest model wins. Sometimes it does not, and we keep the one that teaches better. The learner, not the release date, decides.
A Powerful Tutor, Not a Replacement
We want to be clear about what these apps are and are not. Verónica is more than a conversation partner. She is an AI tutor, a whole system we have built across our platform, and in these apps we have given her the power to hold a real-time voice conversation. She gives learners something they can rarely get enough of: patient, judgment-free practice, available at any hour, in the exact situations their exam or their classroom will demand. For many learners, that is the missing piece between studying and performing.
But she is not a replacement for a teacher, and we do not design her to be. Good teaching still depends on human educators who know their students, their communities, and their goals. Our aim is to amplify that work, to give teachers and learners a tool that multiplies practice and feedback, while the human relationship stays at the center. Technology in service of learning, not the other way around.
We Keep Getting Better
These apps will keep improving. As stronger models arrive and pass our pedagogical tests, the conversations will grow more natural and more capable. As we learn from how students actually use them, we will keep refining the teaching. What will not change is the order of our priorities: education first, technology second. We use the best tools available, and we shape them with what we know about how people learn.
Get Started
If you want to see this in practice, our two live-voice speaking apps are ready now. BTLPT Speaking helps Texas bilingual teacher candidates rehearse academic Spanish in realistic scenarios, and Academic English Speaking helps TOEFL and graduate-school applicants build fluency and confidence in English. Both let you practice real conversations with Verónica, at your own pace, from your browser.
If you missed it, our earlier post, Studying Is Not the Same as Speaking, explains why live practice matters so much for speaking exams. Consider this the companion piece: how we build that practice to actually teach. Universities and programs interested in institutional access can contact us for a walkthrough.
References
Krashen, S. D. (1982). Principles and practice in second language acquisition. Pergamon Press.
Swain, M. (1985). Communicative competence: Some roles of comprehensible input and comprehensible output in its development. In S. M. Gass & C. G. Madden (Eds.), Input in second language acquisition (pp. 235–253). Newbury House.
Sweller, J. (1988). Cognitive load during problem solving: Effects on learning. Cognitive Science, 12(2), 257–285. https://doi.org/10.1207/s15516709cog1202_4
Vygotsky, L. S. (1978). Mind in society: The development of higher psychological processes (M. Cole, V. John-Steiner, S. Scribner, & E. Souberman, Eds.). Harvard University Press.
Wood, D., Bruner, J. S., & Ross, G. (1976). The role of tutoring in problem solving. Journal of Child Psychology and Psychiatry, 17(2), 89–100. https://doi.org/10.1111/j.1469-7610.1976.tb00381.x
Verónica is an AI tutor built by Enabling Learning to support, not replace, the work of human educators. Our speaking apps are designed for adult and professional learners preparing for certification and academic exams.





Leave a Comment