This page documents the criteria LingoLeap scores a spoken answer against and the exact fields the report hands back, task by task. It is a specification, not a sales page: every criterion named below is the one the grader is instructed to apply, and every field named below is one the result screen can display.
By the LingoLeap Research Team · Published:
You get four things. First, a transcript of what the speech recogniser heard. Second, machine measurements of the audio on a 0-100 scale — accuracy, fluency, prosody and completeness, plus a combined pronunciation figure. Third, a task score from 0 to 5 in half-point steps, broken into the criteria that belong to that task: accuracy, completeness and intelligibility for Listen and Repeat; topic relevance, elaboration, fluency, pronunciation and grammar & vocabulary for Take an Interview. Each criterion carries its own written feedback, not just a number. Fourth, task-specific repair material — for Listen and Repeat, the words you missed or changed and per-word pronunciation guidance with IPA and syllable stress; for Take an Interview, a model response, grammar corrections, vocabulary upgrades and the key points a strong answer would have covered.
The 2026 TOEFL Speaking section has two task types and eleven scored questions in roughly eight minutes. There is no Read Aloud task and no independent or integrated speaking task of the pre-2026 kind, so anything written for those older tasks does not describe what you will be scored on.
Listen and Repeat
7 items. You hear one sentence and repeat it inside a response window of 8, 10 or 12 seconds depending on sentence length.
Take an Interview
4 questions. You answer each spoken question within 45 seconds. Early questions are personal and factual; later ones ask for an opinion with reasons.
The two tasks are scored against different criteria because they elicit different speech. A rubric written for one of them tells you nothing useful about the other.
A submitted answer goes through four stages, and the result screen fills in as each one lands. Knowing the order explains why the pronunciation numbers appear before the written feedback.
1. Audio assessment
The recording is transcribed and assessed phoneme by phoneme. This stage produces the transcript and the 0-100 accuracy, fluency, prosody and completeness figures, plus a per-word breakdown that marks each word as correct, mispronounced, omitted or inserted.
2. Rubric scoring
The transcript and the task prompt are scored against the rubric for that task type on a 0-5 scale in half-point steps, producing the overall task score, each criterion score and the written feedback attached to each criterion.
3. Grammar and language check
For the Interview task, the answer is checked for grammar errors and weak word choices, each returned as the original wording, a correction and an explanation of why the correction is better.
4. Expert feedback
For Listen and Repeat, every word scored below the pronunciation threshold gets its own coaching entry: what went wrong, the IPA, the syllable breakdown with stress, the sounds to focus on, how to practise it and similar words to practise with.
Listen and Repeat is scored on how faithfully you reproduced the sentence you heard. The scale runs 0-5 with half points allowed.
Accuracy
How closely the words you produced match the words in the prompt. Substituting a synonym, changing a tense marker or transposing two words all count against accuracy — the task asks for repetition, not paraphrase.
Completeness
How much of the sentence you produced at all. Repeating the opening and trailing off, or dropping a content word from a longer sentence, is a completeness problem rather than an accuracy one.
Intelligibility
Whether a listener could understand you without effort. Imprecise pronunciation that makes a content word ambiguous, running words together, or struggling over a phrase all reduce intelligibility.
scoreThe overall task score, 0-5 in half-point steps.accuracy, completeness, intelligibilityOne score each on the same 0-5 scale, so you can see which of the three pulled the total down.accuracy_feedback, completeness_feedback, intelligibility_feedbackWritten feedback for each criterion separately, rather than one undifferentiated comment.missing_wordsThe list of words from the prompt that were missing or changed in what you said.incorrect_wordsFor each substitution, three things: the word expected, what was heard instead, and the type of error.pronunciation_suggestionsFor each word that scored below the threshold: its per-word accuracy score, the error type (mispronunciation, omission or insertion), what went wrong, the IPA, the syllables with stress marked, the sounds to focus on, step-by-step practice instructions, and similar words to drill.specific_mistakesEach mistake as a type, exactly what happened, how it affected the score, and the fix.overall_feedbackA summary naming your strongest area and the one thing to work on next.Illustrative structure — not a real graded response
To make the shape concrete: if the prompt were "The committee approved the revised budget on Thursday" and the recogniser heard "The committee approved the revised budget Thursday", the report would carry an entry in missing_words for the dropped function word, and a pronunciation entry for any word whose per-word accuracy fell below the threshold, with its IPA and syllable stress. The scores below are deliberately left blank — no score on this page is taken from a real graded answer.
| Field | What it would contain |
|---|---|
| missing_words | the function word that was dropped |
| incorrect_words | expected / heard / error type, for each substitution |
| pronunciation_suggestions | IPA, syllables with stress, sounds to focus on, practice steps, similar words |
| accuracy, completeness, intelligibility | one 0-5 score and one written comment each |
The Interview task is scored on how well you addressed the question and how clearly you did it. The scale again runs 0-5 with half points allowed.
Topic relevance
Whether the answer is on topic and actually addresses the question that was asked, rather than a neighbouring one or a memorised script.
Elaboration
Whether the answer is developed. A response that is on topic but consists mostly of language recycled from the question scores low here even if it is perfectly pronounced.
Fluency
Conversational pace, and whether pauses are natural or frequent and lengthy enough to make the delivery choppy. Frequent filler words count here.
Pronunciation
Intelligibility, rhythm and intonation — whether they convey meaning or make the listener work.
Grammar & vocabulary
Range and accuracy, judged by whether they let you express precise meanings or visibly restrict what you can say.
scoreThe overall task score, 0-5 in half-point steps.Five criterion scorestopic_relevance, elaboration, fluency, pronunciation and grammar_vocabulary, each on the 0-5 scale.Five criterion commentsA separate written comment for each of the five, so a low total is traceable to the criterion that caused it.model_responseA well-elaborated, fluent answer to the same question at the top of the scale, so you can compare structure rather than guess at it.grammar_issuesEach error as the original wording, the correction, and why it is wrong.vocabulary_suggestionsEach weak choice as the original word or phrase, a better alternative, and why the alternative is better.fluency_tipsSpecific tips for improving speaking fluency on this kind of answer.key_points_to_includeThe points a strong answer to this question would have covered — the fastest way to see what your answer left out.overall_feedbackA summary with specific suggestions for improvement.The 0-100 figures on the report are not the task score and do not convert to a band. They come from a phoneme-level assessment of the audio itself and exist to tell you which part of your delivery is weakest.
Accuracy
Averaged over the words in the recording: how closely each word's pronunciation matched the expected pronunciation.
Fluency
The proportion of the recording spent actually speaking rather than pausing, expressed as a percentage of elapsed time.
Prosody
Intonation, stress and rhythm, assessed separately from whether the individual sounds were right.
Completeness
The share of words produced with no detected error, counting omissions against you and ignoring insertions.
The single pronunciation number is not an average. The four figures are sorted and the lowest is weighted at 0.4, the other three at 0.2 each. Your weakest dimension therefore moves it roughly twice as much as any other — which is the point: it is designed to stop a strong score on three dimensions from hiding a bad one.
The same stage marks each word with an error type — none, mispronunciation, omission or insertion — and its own accuracy score. This is what the per-word pronunciation coaching on a Listen and Repeat report is built from.
A task is scored 0-5. The TOEFL 2026 score you care about is a 1-6 band. LingoLeap converts between them in two steps: the 0-5 task score maps to a comparable 0-30 section score, and the 0-30 score maps to the 1-6 band using the published speaking conversion. The table below is generated from that conversion at build time, so it cannot drift out of step with the app.
| Task score (0-5) | Comparable 0-30 | Band (1-6) | CEFR |
|---|---|---|---|
| 5.0 | 30 | 6.0 | C2 |
| 4.5 | 27 | 5.5 | C1 |
| 4.0 | 25 | 5.0 | C1 |
| 3.5 | 22 | 4.0 | B2 |
| 3.0 | 20 | 4.0 | B2 |
| 2.5 | 17 | 3.0 | B1 |
| 2.0 | 14 | 2.5 | A2 |
| 1.5 | 11 | 2.0 | A2 |
| 1.0 | 8 | 1.5 | A1 |
| 0.5 | 4 | 1.0 | A1 |
| 0.0 | 0 | 1.0 | A1 |
Two limits. First, the table converts a single task score; a section band on the real test reflects all eleven scored questions, not one. Second, ETS has not published how raw speaking performance becomes a band score, so the 0-5 to 0-30 step is LingoLeap's own mapping and should be read as an estimate, not as an official equivalence. See our page on what ETS has not published about the 2026 TOEFL. /toefl/2026-open-questions
In the order that saves the most time.
Three things are deliberately absent.
No accuracy or calibration claim
We do not publish an agreement figure between these scores and official ETS scores, and nothing on this page should be read as one. Until that study exists and is published, treat every score here as an estimate.
No real graded sample
Every example on this page is a description of the report's structure. No score, transcript or piece of feedback here is taken from a real graded answer, because publishing one requires a test taker's consent.
No claim about official scoring internals
The score-level descriptions above are the criteria our grader applies. Where ETS has not published a detail — most importantly how a raw speaking performance becomes a band — we say so rather than filling the gap.
Record a Listen and Repeat item or an Interview question in a realistic 2026 mock test and read the criteria on your own speech instead of on an example.
Take a free mock testTOEFL Listen and Repeat
The task itself: response windows, what raters listen for, and how to train for it.
TOEFL Speaking 2026
Both 2026 speaking tasks, timing, and the section format end to end.
TOEFL score conversion
How the 0-30 and 1-6 scales line up across all four sections.
What ETS has not published
The open questions about 2026 scoring, including how raw speaking performance becomes a band.
What TOEFL 2026 writing feedback contains
The same specification for written answers: the criteria for Write an Email and Academic Discussion, the returned fields and the band conversion.