LingoLeap Research Team · September 2026
Listen and Repeat looks almost too simple to study. You hear one sentence and say it back. There is no argument to build and no template to remember. Yet when we looked at real practice, the same failures kept appearing: a sentence would start cleanly, then lose words, endings, or rhythm as it became longer.
We wanted to know whether those impressions were real or just memorable examples. So we aggregated 152,218 scored sentence attempts from 23,390 LingoLeap practice sessions collected between January and September 2026.
The strongest pattern was sentence length. Learners reproduced 78.8% of prompts with six words or fewer, but only 50.0% once prompts reached sixteen words. We also found a small set of consistently difficult sounds, frequent loss of word-final -s, and a steady score decline as detected pauses accumulated.
Below we walk through the setup, the five patterns we found, what might explain them, and what the data still cannot tell us.
Prompt load
≤6 → 16+ words
More material to retain
Reproduction
78.8% → 50.0%
Less of the prompt returns
Delivery
2.22/5 at 5+ pauses
More breaks, lower score
Fine detail
11.4% final -s omitted
Small endings disappear
We kept the questions deliberately narrow. The data can show recurring patterns in LingoLeap practice; it cannot, by itself, prove why they happen.
The dataset contains 23,390 real practice sessions and 152,218scored sentence attempts. Each recording was processed by automatic speech pronunciation assessment, which returned word timings and word- and phoneme-level accuracy scores from 0 to 100. Each sentence also received a LingoLeap platform score from 0 to 5.
We define reproduction rate as the share of the target sentence a speaker repeats, recovered by aligning the target sentence to the recognised words. A pause is a silence greater than 250ms between two words.
One detail matters here: an attempt is not a person. A session can contain several sentences, and the same learner can contribute more than one session. The aggregate available to this page does not include a deduplicated learner count or an independent-prompt count. We therefore treat these as descriptive practice records, not 152,218 independent test takers.
We excluded 11,512 items with no recognised speech from content-level calculations. We report denominators, percentages, group contrasts, and score differences, but not significance tests or confidence intervals—the repeated learners and prompts are not independent observations in this aggregate extract.
This was the clearest pattern in the dataset. Reproduction falls from 78.8% on the shortest sentences to 50.0% on the longest, and the score falls with it.
What surprised us is what did not change. Average pronunciation accuracy stays in a narrow range, from 84.8 to 84.5, while reproduction changes by 28.8 percentage points. In other words, learners do not appear to pronounce every surviving word dramatically worse; they simply bring back less of the sentence.
Greater memory and processing load is one plausible explanation [1]. But we cannot isolate it here: longer prompts may also contain harder vocabulary, syntax, or sound sequences. This is a recurring pattern in the records, not a controlled experiment on sentence length.
| Sentence length | Attempts | Avg score | Δ score | Reproduced | Δ reproduced |
|---|---|---|---|---|---|
| ≤6 words | 18,892 | 3.74 | 0.00 | 78.8% | 0.0 pp |
| 7–9 words | 47,733 | 3.51 | -0.23 | 73.7% | -5.1 pp |
| 10–12 words | 49,501 | 3.15 | -0.59 | 65.8% | -13.0 pp |
| 13–15 words | 23,034 | 2.79 | -0.95 | 57.2% | -21.6 pp |
| 16+ words | 13,058 | 2.61 | -1.13 | 50.0% | -28.8 pp |
/θ/ — the “th” in thin — is the lowest-scoring sound at 61.6/100; 37.6% of its 17,575 measured occurrences fall below the provider’s 60-point threshold.
| Phoneme | Occurrences | Mean accuracy | Below 60 |
|---|---|---|---|
| /θ/ | 17,575 | 61.6 | 37.6% |
| /ŋ/ | 61,203 | 69.8 | 27.2% |
| /z/ | 188,178 | 74.3 | 23.0% |
| /oʊ/ | 83,991 | 75.2 | 25.0% |
| /s/ | 286,021 | 79.3 | 18.8% |
| /d/ | 212,537 | 80.7 | 16.5% |
| /w/ | 94,081 | 81 | 14.9% |
| /t/ | 401,948 | 81.5 | 16.3% |
| /h/ | 60,167 | 81.6 | 17.3% |
| /n/ | 409,462 | 81.8 | 16.4% |
The gap is large enough to be useful, but it is easy to over-interpret. These are automatic-assessment outputs, not phonetician judgements. Microsoft defines AccuracyScore as acoustic similarity to a native pronunciation model and treats word scores below 60 as mispronunciations [3].
It may be part of the story—cross-language speech research predicts phoneme-specific variation [2]—but this page does not test that hypothesis. First-language status is unavailable for most sessions, so we keep that comparison in a separate analysis.
First-language variation
Which sounds are hardest depends on the speaker's native language. The contrast between Chinese-first-language and other learners is examined in a companion study on Chinese learners’ pronunciation.
Word-final -s is dropped on 11.4% of the 237,757 words that carry it.
This is the kind of error that is easy to miss by ear. The main word is still there, so the repetition feels correct, but a small grammatical signal has vanished. Even when the ending survives, the closing sibilant averages only 74.3/100 — audible but weakly formed. Endings are acoustically brief and can be difficult to preserve under time pressure. The present data do not identify whether memory load, articulation, or recognition error caused an individual omission. The loss is meaning-bearing: dropping the -s turns closes into close and weekends into weekend; the same happens to past-tense -ed and to numbers and times.
Target
The reading room closes at nine on weekends.
Common loss
The reading room closes at nine on weekends.
-s analysis.Score falls monotonically with each added pause — from 3.68/5 for the cleanest deliveries to 2.22/5 at five or more breaks.
| Pauses used | Attempts | Avg score | Δ vs none |
|---|---|---|---|
| None (straight through) | 66,778 | 3.68 | 0.00 |
| 1 | 42,342 | 3.11 | -0.57 |
| 2 | 25,452 | 2.78 | -0.90 |
| 3 | 11,930 | 2.55 | -1.13 |
| 4 | 4,270 | 2.41 | -1.27 |
| 5+ | 1,446 | 2.22 | -1.46 |
The nearly step-by-step decline looks compelling, but it is not a pause penalty estimate. Lower-proficiency or less-complete attempts may both pause more and score lower.
Much less than we expected. The average score is 2.82when a break lands at an annotated split point and 2.73when it lands mid-phrase. That small descriptive gap is not enough to support a strong placement claim.
Content words carrying the hardest sounds and clusters score lowest — for example turned (41/100) and loaner (44/100).
| Word | Times seen | Avg accuracy |
|---|---|---|
| turned | 61 | 41 |
| loaner | 104 | 44 |
| portraits | 189 | 54.2 |
| closed | 391 | 61.3 |
| sightlines | 314 | 62.7 |
| known | 149 | 63.3 |
| rollers | 263 | 65.5 |
| turnstiles | 315 | 67.2 |
| brochures | 74 | 68 |
| counters | 91 | 68.6 |
The findings describe two observable dimensions: reproduction and pause behaviour on one side, and automatic pronunciation assessment on the other. Working-memory research makes processing load a plausible interpretation of the length pattern [1], while language-assessment research cautions that automated speech scores require construct-level validation [4]. These mechanisms remain hypotheses here; the dataset establishes patterns, not their causes.
For learners, the descriptive results still identify concrete diagnostic targets: compare short and long prompts, inspect whether the sentence ending survives, and review low-scoring phonemes. The companion method guide turns those observations into a practice protocol, but its recommendations should not be read as an experimentally estimated treatment effect.
Every figure above is an average over thousands of attempts. Each individual Listen and Repeat sentence also has its own page — the real target sentence, how it splits into chunks, which specific words are dropped, and where speakers pause. The concrete case behind the aggregate:
Listen and Repeat is not failing learners in one mysterious way. The records point to several smaller problems that arrive together: longer prompts are reproduced less completely, a few sounds are consistently fragile, grammatical endings disappear, and extra pauses accompany lower scores.
Our current best interpretation is that memory pressure and articulation interact, but that is still an interpretation. The next useful step is to test these patterns within the same learner and the same sentence over time. That would tell us whether targeted practice changes the outcome, rather than merely describing who currently performs better.
For now, the most practical use of the data is diagnostic: find the sentence length where your reproduction begins to break, check whether the final words and endings survive, and isolate the sounds that repeatedly score low. Our Listen and Repeat method guide turns those checks into a practice routine.
LingoLeap Research Team. (2026). TOEFL Listen and Repeat: Evidence from 152,218 Scored Sentence Attempts. LingoLeap Research, version 1.0.
https://lingoleap.ai/jp/research/listen-and-repeat
This study covers 23,390 real Listen and Repeat sessions — 152,218 individually scored sentence attempts — from TOEFL 2026 practice on LingoLeap since January 2026.
/θ/, the "th" in "thin", is the hardest by a clear margin: it averaged 61.6 out of 100 and was mispronounced on 37.6% of attempts. Many first languages have no equivalent sound, so it is often replaced with /s/ or /t/.
In this descriptive dataset, the share reproduced falls from 78.8% at six words or fewer to 50.0% past sixteen, while the average platform score falls from 3.74 to 2.61 out of 5. The observational design does not establish sentence length as the sole cause of that difference.
No. Scores are LingoLeap AI scores on a 0–5 per-sentence scale, and pronunciation is measured by automatic assessment. They are a large, consistent signal for how the task behaves, not official ETS results.