LingoLeap

What 152,218 Listen and Repeat attempts taught us

LingoLeap Research Team · September 2026

Listen and Repeat looks almost too simple to study. You hear one sentence and say it back. There is no argument to build and no template to remember. Yet when we looked at real practice, the same failures kept appearing: a sentence would start cleanly, then lose words, endings, or rhythm as it became longer.

We wanted to know whether those impressions were real or just memorable examples. So we aggregated 152,218 scored sentence attempts from 23,390 LingoLeap practice sessions collected between January and September 2026.

The strongest pattern was sentence length. Learners reproduced 78.8% of prompts with six words or fewer, but only 50.0% once prompts reached sixteen words. We also found a small set of consistently difficult sounds, frequent loss of word-final -s, and a steady score decline as detected pauses accumulated.

Below we walk through the setup, the five patterns we found, what might explain them, and what the data still cannot tell us.

01

Prompt load

≤6 → 16+ words

More material to retain

02

Reproduction

78.8% → 50.0%

Less of the prompt returns

03

Delivery

2.22/5 at 5+ pauses

More breaks, lower score

04

Fine detail

11.4% final -s omitted

Small endings disappear

A visual map of the four signals examined in this analysis; arrows show reading order, not proven causality.
Section 1

What we wanted to understand

We kept the questions deliberately narrow. The data can show recurring patterns in LingoLeap practice; it cannot, by itself, prove why they happen.

  • What changes when a sentence gets longer?
  • Which sounds and word endings repeatedly receive low pronunciation-assessment results?
  • What happens to the platform score as a learner pauses more often?
Section 2

The setup

The dataset contains 23,390 real practice sessions and 152,218scored sentence attempts. Each recording was processed by automatic speech pronunciation assessment, which returned word timings and word- and phoneme-level accuracy scores from 0 to 100. Each sentence also received a LingoLeap platform score from 0 to 5.

We define reproduction rate as the share of the target sentence a speaker repeats, recovered by aligning the target sentence to the recognised words. A pause is a silence greater than 250ms between two words.

One detail matters here: an attempt is not a person. A session can contain several sentences, and the same learner can contribute more than one session. The aggregate available to this page does not include a deduplicated learner count or an independent-prompt count. We therefore treat these as descriptive practice records, not 152,218 independent test takers.

We excluded 11,512 items with no recognised speech from content-level calculations. We report denominators, percentages, group contrasts, and score differences, but not significance tests or confidence intervals—the repeated learners and prompts are not independent observations in this aggregate extract.

Section 3

Pattern 1: Longer sentences lose more of the prompt

This was the clearest pattern in the dataset. Reproduction falls from 78.8% on the shortest sentences to 50.0% on the longest, and the score falls with it.

0%25%50%75%100%79%≤6 w3.74/574%7–9 w3.51/566%10–12 w3.15/557%13–15 w2.79/550%16+ w2.61/5bars: % of sentence reproduced · below: avg scoreLingoLeap Research
Figure 1. Share of the target sentence reproduced, by sentence-length band (n = 152,218 attempts). Average score for each band is shown beneath the axis. LingoLeap Research · 2026

What surprised us is what did not change. Average pronunciation accuracy stays in a narrow range, from 84.8 to 84.5, while reproduction changes by 28.8 percentage points. In other words, learners do not appear to pronounce every surviving word dramatically worse; they simply bring back less of the sentence.

Why might this happen?

Greater memory and processing load is one plausible explanation [1]. But we cannot isolate it here: longer prompts may also contain harder vocabulary, syntax, or sound sequences. This is a recurring pattern in the records, not a controlled experiment on sentence length.

Sentence lengthAttemptsAvg scoreΔ scoreReproducedΔ reproduced
≤6 words18,8923.740.0078.8%0.0 pp
7–9 words47,7333.51-0.2373.7%-5.1 pp
10–12 words49,5013.15-0.5965.8%-13.0 pp
13–15 words23,0342.79-0.9557.2%-21.6 pp
16+ words13,0582.61-1.1350.0%-28.8 pp
Table 1. Attempts, average score, and reproduction rate by sentence-length band. LingoLeap Research · 2026
Section 4

Pattern 2: A few sounds are much harder than the rest

/θ/ — the “th” in thin — is the lowest-scoring sound at 61.6/100; 37.6% of its 17,575 measured occurrences fall below the provider’s 60-point threshold.

020406080100/θ/ thin61.6/ŋ/ sing69.8/z/ zoo74.3// goat75.2/s/ see79.3/d/ day80.7/w/ way81/t/ tea81.5/h/ hat81.6/n/ no81.8LingoLeap Research
Figure 2. Average pronunciation accuracy (0–100) for the ten hardest sounds. Dashed line marks the 60-point threshold below which a word is counted mispronounced; red bars fall below it. LingoLeap Research · 2026
PhonemeOccurrencesMean accuracyBelow 60
/θ/17,57561.637.6%
/ŋ/61,20369.827.2%
/z/188,17874.323.0%
/oʊ/83,99175.225.0%
/s/286,02179.318.8%
/d/212,53780.716.5%
/w/94,0818114.9%
/t/401,94881.516.3%
/h/60,16781.617.3%
/n/409,46281.816.4%
Table 2. The ten lowest-scoring phonemes, with occurrence counts and the share below the 60-point accuracy threshold. LingoLeap Research · 2026

The gap is large enough to be useful, but it is easy to over-interpret. These are automatic-assessment outputs, not phonetician judgements. Microsoft defines AccuracyScore as acoustic similarity to a native pronunciation model and treats word scores below 60 as mispronunciations [3].

Is first language the explanation?

It may be part of the story—cross-language speech research predicts phoneme-specific variation [2]—but this page does not test that hypothesis. First-language status is unavailable for most sessions, so we keep that comparison in a separate analysis.

First-language variation

Which sounds are hardest depends on the speaker's native language. The contrast between Chinese-first-language and other learners is examined in a companion study on Chinese learners’ pronunciation.

Section 5

Pattern 3: The word survives, but the ending disappears

Word-final -s is dropped on 11.4% of the 237,757 words that carry it.

This is the kind of error that is easy to miss by ear. The main word is still there, so the repetition feels correct, but a small grammatical signal has vanished. Even when the ending survives, the closing sibilant averages only 74.3/100 — audible but weakly formed. Endings are acoustically brief and can be difficult to preserve under time pressure. The present data do not identify whether memory load, articulation, or recognition error caused an individual omission. The loss is meaning-bearing: dropping the -s turns closes into close and weekends into weekend; the same happens to past-tense -ed and to numbers and times.

Target

The reading room closes at nine on weekends.

Common loss

The reading room closes at nine on weekends.

A visual example of the feature counted in the word-final -s analysis.
Section 6

Pattern 4: Every additional pause comes with a lower score

Score falls monotonically with each added pause — from 3.68/5 for the cleanest deliveries to 2.22/5 at five or more breaks.

2.02.53.03.54.03.6803.1112.7822.5532.4142.225+detected pauses
Figure 3. Average LingoLeap platform score by detected pause count (n = 152,218 scored attempts). LingoLeap Research · 2026
Pauses usedAttemptsAvg scoreΔ vs none
None (straight through)66,7783.680.00
142,3423.11-0.57
225,4522.78-0.90
311,9302.55-1.13
44,2702.41-1.27
5+1,4462.22-1.46
Table 3. Average score by the number of pauses used (a pause is a silence over 250ms between two words). LingoLeap Research · 2026

The nearly step-by-step decline looks compelling, but it is not a pause penalty estimate. Lower-proficiency or less-complete attempts may both pause more and score lower.

Does pause placement matter?

Much less than we expected. The average score is 2.82when a break lands at an annotated split point and 2.73when it lands mid-phrase. That small descriptive gap is not enough to support a strong placement claim.

Section 7

Pattern 5: A small set of words repeatedly scores low

Content words carrying the hardest sounds and clusters score lowest — for example turned (41/100) and loaner (44/100).

turned41loaner44portraits54.2closed61.3sightlines62.7known63.3
Figure 4. Average automatic pronunciation accuracy for six of the lowest-scoring words in the filtered dataset. LingoLeap Research · 2026
WordTimes seenAvg accuracy
turned6141
loaner10444
portraits18954.2
closed39161.3
sightlines31462.7
known14963.3
rollers26365.5
turnstiles31567.2
brochures7468
counters9168.6
Table 4. The ten lowest-accuracy words, among words attempted at least fifty times. LingoLeap Research · 2026
Section 8

What might connect these patterns

The findings describe two observable dimensions: reproduction and pause behaviour on one side, and automatic pronunciation assessment on the other. Working-memory research makes processing load a plausible interpretation of the length pattern [1], while language-assessment research cautions that automated speech scores require construct-level validation [4]. These mechanisms remain hypotheses here; the dataset establishes patterns, not their causes.

For learners, the descriptive results still identify concrete diagnostic targets: compare short and long prompts, inspect whether the sentence ending survives, and review low-scoring phonemes. The companion method guide turns those observations into a practice protocol, but its recommendations should not be read as an experimentally estimated treatment effect.

Section 9

From aggregate patterns to real sentences

Every figure above is an average over thousands of attempts. Each individual Listen and Repeat sentence also has its own page — the real target sentence, how it splits into chunks, which specific words are dropped, and where speakers pause. The concrete case behind the aggregate:

Section 10

Closing thoughts

Listen and Repeat is not failing learners in one mysterious way. The records point to several smaller problems that arrive together: longer prompts are reproduced less completely, a few sounds are consistently fragile, grammatical endings disappear, and extra pauses accompany lower scores.

Our current best interpretation is that memory pressure and articulation interact, but that is still an interpretation. The next useful step is to test these patterns within the same learner and the same sentence over time. That would tell us whether targeted practice changes the outcome, rather than merely describing who currently performs better.

For now, the most practical use of the data is diagnostic: find the sentence length where your reproduction begins to break, check whether the final words and endings survive, and isolate the sounds that repeatedly score low. Our Listen and Repeat method guide turns those checks into a practice routine.

Section 11

Dataset note

LingoLeap Research Team. (2026). TOEFL Listen and Repeat: Evidence from 152,218 Scored Sentence Attempts. LingoLeap Research, version 1.0.

https://lingoleap.ai/hi/research/listen-and-repeat

Section 12

Notes and sources

  1. [1] Baddeley, A., Gathercole, S., & Papagno, C. (1998). The phonological loop as a language learning device. Psychological Review, 105(1), 158–173.
  2. [2] Saito, K. (2019). Effects of second language pronunciation teaching revisited. Language Learning, 69(3), 652–708.
  3. [3] Microsoft. (2026). Use pronunciation assessment. Microsoft Learn.
  4. [4] Xi, X. (2023). Examining the validity of automated speech scoring systems. Journal of the Acoustical Society of Japan, 79(3), 170–176.
Section 13

Questions we hear from learners

How many TOEFL Listen and Repeat attempts were analyzed?

This study covers 23,390 real Listen and Repeat sessions — 152,218 individually scored sentence attempts — from TOEFL 2026 practice on LingoLeap since January 2026.

What is the hardest English sound for TOEFL speakers?

/θ/, the "th" in "thin", is the hardest by a clear margin: it averaged 61.6 out of 100 and was mispronounced on 37.6% of attempts. Many first languages have no equivalent sound, so it is often replaced with /s/ or /t/.

Does sentence length affect the Listen and Repeat score?

In this descriptive dataset, the share reproduced falls from 78.8% at six words or fewer to 50.0% past sixteen, while the average platform score falls from 3.74 to 2.61 out of 5. The observational design does not establish sentence length as the sole cause of that difference.

Are these official ETS scores?

No. Scores are LingoLeap AI scores on a 0–5 per-sentence scale, and pronunciation is measured by automatic assessment. They are a large, consistent signal for how the task behaves, not official ETS results.