TOEFLIELTS
Türkçe

TOEFL 2026 Writing Feedback: What You Actually Get

This page documents the criteria LingoLeap scores a written answer against and the exact fields the report hands back, task by task. It is a specification, not a sales page: every criterion named below is the one the grader is instructed to apply, and every field named below is one the result screen can display.

By the LingoLeap Research Team · Published:

What feedback do you get on a TOEFL 2026 writing answer?

You get five things. First, a task score from 0 to 5 in half-point steps, broken into the four criteria that belong to that task: content, syntactic and lexical variety, social conventions and accuracy for Write an Email; content and elaboration, response to the discussion, syntactic and lexical variety and language accuracy for Write for an Academic Discussion. Each criterion carries its own written comment and, where there is something to fix, quoted extracts from your own answer with a suggested rewrite and an explanation. Second, a minimal-edit grammar and spelling pass that lists each error as a one-to-three-word original, the correction, a reason and an error type. Third, a rewritten version of your answer at the level of a 5, together with the specific changes that got it there. Fourth, a vocabulary analysis: a CEFR distribution across A1 to C2, your total word count, your text annotated level by level, and a list of lower-level words with higher-level replacements. Fifth, a sentence-variety analysis that classifies every sentence as simple, compound or complex, records its opener type and length, and suggests rewrites; plus a mindmap and keyword list built from the rewritten answer. Alongside all of it the report carries your word, sentence and paragraph counts and the score converted to the 0-30 scale, the 1-6 band, CEFR and an IELTS equivalent.

The 2026 writing tasks

The 2026 TOEFL Writing section has three task types and twelve questions in roughly 23 minutes, presented in a fixed linear order. Only two of the three are graded against a rubric, and those two are the ones this page documents. Nothing written for the older independent or integrated essay describes what you will be scored on now.

Build a Sentence (~6 minutes)

You arrange given words or word groups into one grammatically correct sentence. It is scored correct or incorrect, not on a 0-5 rubric, so there is no per-criterion written feedback to describe.

Write an Email (~7 minutes)

You read a campus scenario and write a short email for a specific purpose and audience. Scored 0-5 in half-point steps against four criteria.

Write for an Academic Discussion (~10 minutes)

You read a professor's prompt and two student replies, then write your own contribution. Official 2026 materials say an effective response contains at least 100 words. Scored 0-5 in half-point steps against four criteria.

Timings and the task inventory here are the ones published on our own TOEFL Writing 2026 guide, which sources them from official 2026 practice materials. /toefl/writing

What happens to your answer

Submitting an answer runs a fixed sequence of six steps. The report fills in as each one finishes, which is why the score can appear before the vocabulary and sentence panels do.

1. Scoring

The answer, the task prompt and a set of response statistics go to the grader with the full 0-5 rubric for that task type. It returns an overall score, the four criterion scores, a written comment on each and quoted examples.

2. Grammar and spelling pass

A separate, stricter pass lists errors as minimal edits. The instruction is explicit that each correction changes only one to three words and does not rewrite or rephrase the sentence.

3. Revision

A rewritten version of your answer at the level of a 5, built from your own content rather than a generic model. For the email task the instruction forbids placeholders such as [Last Name], so the rewrite uses the names the scenario actually gives.

4. Vocabulary analysis

Word-level CEFR classification. Where the CEFR dictionary is available it does the classification and the model is used only for the upgrade suggestions; otherwise the model does both.

5. Sentence-variety analysis

Every sentence is classified by type, opener and length, connectors are sorted into basic and advanced, and rewrites are proposed for the weakest ones.

6. Mindmap and keywords

An outline of the rewritten answer plus a keyword list, each entry carrying an IPA pronunciation, part of speech, definition and the exact sentence from the rewrite where the word appears.

Write an Email: criteria and fields

The email grader is instructed to apply the four official 2026 scoring dimensions. The name on the left is the rubric dimension and the field name in the result; the name in brackets is the label the report prints above it.

The four criteria

content — shown as "Communicative Purpose"

Whether your elaboration supports the communicative purpose of the email. At 5 it effectively supports it; at 3 it only partially supports it; at 2 it is limited or irrelevant.

syntactic_lexical_variety — shown as "Language Facility"

Syntactic variety and idiomatic word choice. At 5 the variety is effective and the word choice precise and idiomatic; at 3 the range of syntax and vocabulary is moderate; at 1 the language is telegraphic with very limited vocabulary.

social_conventions — shown as "Social Conventions"

Politeness, register, organisation of information, and how you formulate actions such as requests, refusals and criticisms. At 5 these are consistently appropriate; at 3 there are some noticeable errors; at 2 they are inconsistent or inappropriate.

accuracy — shown as "Grammar & Mechanics"

Lexical and grammatical accuracy. At 5 there are almost no errors, with only the common typos expected under timed conditions; at 4 there are few; at 2 errors accumulate across sentence structure and language use.

The 0-5 levels

  • 5 — Fully successful. The response is effective, clearly expressed, and shows consistent facility in the use of language.
  • 4 — Generally successful. The response is mostly effective and easily understood; language facility is adequate to the task.
  • 3 — Partially successful. The response generally accomplishes the task, but limitations in language facility may keep parts of the message from being fully clear and effective.
  • 2 — Mostly unsuccessful. The response attempts the task but is mostly ineffective; the message may be limited or difficult to interpret.
  • 1 — Unsuccessful. The attempt is ineffective and the message may be limited to the point of being unintelligible, with any coherent language mostly borrowed from the prompt.
  • 0 — Non-scorable. Blank, rejects the topic, is not in English, is entirely copied from the prompt, is entirely unconnected to it, or consists of arbitrary keystrokes.

The overall score is not an average

For the email task the grader is told explicitly that the overall score is not a simple average of the four criterion scores, and that it is limited by weaknesses in vocabulary and social conventions: the overall score should not exceed the lower of syntactic_lexical_variety and social_conventions by more than half a point. Correct grammar on its own does not produce a high score. That is why an answer can come back with a strong accuracy figure and a middling overall.

Fields returned for Write an Email

  • scoreThe overall task score, 0-5 in half-point steps.
  • content, syntactic_lexical_variety, social_conventions, accuracyThe four criterion scores, each 0-5 in half-point steps.
  • [criterion]_feedbackOne written comment per criterion — four in all.
  • [criterion]_examplesPer criterion, a list of objects carrying original, suggestion and explanation, quoting your own text. The grader is instructed to return an empty list rather than invent an example when there is nothing genuine to fix, so an empty panel means no finding, not a missing one.
  • overall_feedbackA summary comment, written in your interface language.
  • revised_emailYour email rewritten to the level of a 5, using the names and details the scenario supplies rather than placeholders.
  • grammar_errorsA list of objects carrying original_text, corrected_text and explanation, from the revision step.
  • word_choice_improvementsA list of objects carrying original_text, improved_text and explanation.
  • social_convention_issuesA list of objects carrying issue and suggestion — the register and politeness problems specifically, kept separate from grammar.

Write for an Academic Discussion: criteria and fields

The discussion grader applies a different set of four dimensions. One of them has no counterpart in the email task: whether you actually engage with what the other students wrote.

The four criteria

content_and_elaboration — shown as "Content & Elaboration"

Whether you give relevant, well-elaborated explanations, examples or details to support your position. At 5 the elaboration is well developed with relevant examples; at 3 it may be missing, unclear or irrelevant in parts.

response_to_discussion — shown as "Response to Discussion"

Whether you contribute to the ongoing discussion in your own words, typically by responding to the arguments in the short texts you were given. At 4 there is real engagement with others' arguments; at 2 that engagement is limited and the other students may be largely ignored.

syntactic_lexical_variety — shown as "Syntactic & Lexical Variety"

Variety of sentence structures and precise, idiomatic word choice. At 5 complex syntax is used effectively; at 2 the range of syntax and vocabulary is limited.

language_accuracy — shown as "Language Accuracy"

Lexical and grammatical errors. The rubric states that a top score requires a response with almost no errors; at 3 errors in sentence structure, word forms or idiomatic language are noticeable.

The 0-5 levels

  • 5 — Fully successful. The response is relevant and very clearly expressed, with consistent facility in the use of language.
  • 4 — Generally successful. A relevant contribution to the discussion; the ideas are easily understood.
  • 3 — Partially successful. Mostly relevant and mostly understandable, with some facility in the use of language.
  • 2 — Mostly unsuccessful. The attempt is ineffective and the ideas may be hard to follow because of language limitations.
  • 1 — Unsuccessful. Few or no coherent ideas and severely limited language control; coherent language is mostly borrowed from the prompt.
  • 0 — Non-scorable. Blank, not in English, entirely copied from the prompt, or entirely unconnected to the topic.

Scored holistically, unlike the email

Where the email grader is given an arithmetic ceiling, the discussion grader is told the overall score reflects a holistic assessment of all four dimensions: strong content with adequate language can reach the 3.5-4 range, and excellent content with sophisticated language the 4.5-5 range. The two tasks therefore combine their criterion scores in genuinely different ways, and the same four subscores would not necessarily produce the same overall on both.

Fields returned for Academic Discussion

  • scoreThe overall task score, 0-5 in half-point steps.
  • content_and_elaboration, response_to_discussion, syntactic_lexical_variety, language_accuracyThe four criterion scores, each 0-5 in half-point steps.
  • [criterion]_feedbackOne written comment per criterion — four in all.
  • [criterion]_examplesPer criterion, a list of objects carrying original, suggestion and explanation, quoting your own text, or an empty list where there is nothing genuine to fix.
  • overall_feedbackA summary comment, written in your interface language.
  • revised_responseYour contribution rewritten to the level of a 5, at roughly the target length for the task.
  • grammar_errorsA list of objects carrying original_text, corrected_text and explanation, from the revision step.
  • content_improvementsA list of objects carrying issue and improvement, saying what was lacking in the original and how the rewrite addressed it.

The rest of the report

Four more panels are produced for both rubric-scored tasks. They are the same for the email and the discussion, because the same modules run on both.

Minimal grammar corrections

Each item carries original, corrected, reason and type, where original and corrected are one to three words. This pass is deliberately separate from the criterion feedback: it is a proofreading list, not an assessment.

Vocabulary analysis

vocabulary_distribution gives the share of your words at each CEFR level from A1 to C2; total_words gives the count; annotated_text returns your own text with levels marked inline; vocabulary_upgrades lists lower-level words with a higher-level replacement, both levels and an explanation; vocabulary_level_summary is a short written assessment.

Sentence-variety analysis

sentence_analysis classifies every sentence as simple, compound or complex, records its opener type and word count, and attaches a rewrite where one is needed. Alongside it come sentence_type_distribution, opener_distribution, sentence_lengths, average_sentence_length, a verdict on length variety, a connector_analysis splitting the connectors you used into basic and advanced, a variety_score and a ranked top_improvements list.

Mindmap and keywords

An outline of the rewritten answer — for an email, greeting, self-introduction, purpose, background, supporting details, request, closing and sign-off — plus a keyword list where each entry carries the word, its IPA pronunciation, part of speech, definition and the exact sentence from the rewrite that uses it.

Counts and conversions

The report also carries n_words, n_sentences and n_paragraphs for your answer, and the score expressed four ways: raw_score out of 5, score out of 30, band_score out of 6, and cefr, plus an ielts equivalent.

From a 0-5 task score to a 1-6 band

A task score is converted in three steps: 0-5 to the familiar 0-30 scale, 0-30 to the 2026 band scale of 1-6 using the writing-specific table, and the band to a CEFR level. The table below is computed at build time from the same constants the application uses, so it cannot drift from the app.

Task score (0-5)0-30 scale2026 band (1-6)CEFR
5.0306.0C2
4.5275.5C1
4.0255.0C1
3.5224.5B2
3.0204.0B2
2.5174.0B2
2.0143.0B1
1.5112.5A2
1.082.0A2
0.541.5A1
0.001.0A1

Two task scores share a band

Because the 0-30 band table treats 17 to 20 as a single band, a 2.5 and a 3.0 both land on band 4.0. The gap between them is real on the 0-30 scale and invisible on the band scale. If you are tracking progress, track the raw 0-5 criterion scores, not the band.

One step in this chain is ours, not ETS's

ETS publishes the 0-5 writing rubric and publishes the 1-6 band scale, but has not published how raw writing performance is converted into a band. The 0-5 to 0-30 step above is our own mapping, chosen so the familiar 0-30 figure stays available; it is not an official conversion and should not be read as one. We track this and the other unpublished details at /toefl/2026-open-questions

How to read your feedback

The report is long. This is the order that gets the most out of it in the least time.

  1. 1Read the four criterion scores before the overall. The overall is a single number that hides which of the four is holding you back, and on the email task it is deliberately capped by the weakest of vocabulary and social conventions.
  2. 2Find your lowest criterion and read only its written comment and its examples first. Those examples quote your own sentences, so they tell you what to change rather than what to improve in general.
  3. 3Treat an empty examples list as a finding, not an omission. The grader is instructed to return nothing rather than fabricate a correction, so an empty panel means that criterion had no genuine fix to offer.
  4. 4Read the minimal grammar corrections separately from the criterion feedback. They are a proofreading pass, and a long list of one-word typos does not by itself mean a low accuracy score.
  5. 5Compare the revised version against your own sentence by sentence rather than reading it straight through. The content_improvements or word_choice_improvements list tells you which changes were the load-bearing ones.
  6. 6Use the vocabulary distribution as a trend, not a target. A high share of A1-A2 words is the specific weakness the email rubric caps your score for; the upgrades list is the shortest route to moving it.
  7. 7On the discussion task, check response_to_discussion first if you write well but score mid-range. It is the criterion with no counterpart anywhere else, and a fluent answer that ignores the other students is exactly what it penalises.

What this page does not claim

Three things are deliberately absent, and their absence is the point.

No accuracy or calibration figure

We have not published a study comparing these scores against official ETS writing scores, so this page states no agreement rate, margin of error or accuracy claim. Treat the score as practice feedback, not a prediction.

No sample score or sample feedback

There is no worked example on this page. We hold no graded 2026 writing response we are able to publish, and inventing one to illustrate the format would defeat the purpose of a specification.

No invented ETS internals

Where ETS has not published a detail — most importantly how raw writing performance becomes a band — this page says so rather than filling the gap.

Frequently asked questions

What criteria is a TOEFL 2026 email scored on?
Four: content, which asks whether your elaboration supports the communicative purpose; syntactic and lexical variety; social conventions, covering politeness, register, organisation and how you formulate requests, refusals and criticisms; and accuracy, meaning lexical and grammatical errors. Each is scored 0-5 in half-point steps and carries its own written comment. In the LingoLeap report these appear under the labels Communicative Purpose, Language Facility, Social Conventions and Grammar & Mechanics.
What criteria is the Academic Discussion task scored on?
Four, and they are not the same four as the email: content and elaboration; response to the discussion, meaning whether you contribute in your own words to what the other students argued; syntactic and lexical variety; and language accuracy. Each is scored 0-5 in half-point steps with its own written comment.
Is the overall writing score the average of the four criterion scores?
No, and the two tasks differ. For Write an Email the grader is instructed that the overall is not a simple average and is limited by weaknesses in vocabulary and social conventions, so it should not exceed the lower of those two by more than half a point — correct grammar alone does not lift it. For Write for an Academic Discussion the overall is a holistic judgement across all four dimensions rather than an arithmetic result.
How accurate is LingoLeap's TOEFL writing score?
We do not know, and we will not guess. We have published no study comparing these scores against official ETS writing scores, so we make no accuracy or calibration claim anywhere on this page. What we can state precisely is what the score is made of, which is what this page documents. Use it as practice feedback, not as a prediction of your official result.
Do you correct my grammar, or just comment on it?
Both, in two separate passes. The accuracy or language accuracy criterion gives a score and a written comment with quoted examples from your answer. A separate proofreading pass returns a list of minimal corrections, each changing only one to three words, with the original, the correction, a reason and an error type. It is instructed not to rewrite or rephrase sentences.
Do I get a model answer?
You get a revision of your own answer at the level of a 5, not a generic model. For the email task the rewrite is required to use the names and details from the scenario rather than placeholders such as [Last Name]. It comes with the grammar errors it fixed and, depending on the task, the word-choice improvements and social-convention issues, or a list of what was lacking in content and how the rewrite addressed it.
Why do some feedback sections come back empty?
Because the grader is instructed to return an empty list rather than fabricate an example when a criterion has no genuine improvement to suggest, and not to show identical original and suggestion pairs. An empty examples panel means that criterion found nothing to fix.
Is Build a Sentence graded the same way?
No. Build a Sentence is scored correct or incorrect rather than on the 0-5 rubric, so there is no per-criterion breakdown or written feedback of the kind described here. This page covers the two rubric-scored tasks, Write an Email and Write for an Academic Discussion.

See the report on your own writing

Write a 2026 email or academic discussion response in a full mock test and get the criterion scores, quoted corrections, revision, vocabulary and sentence analysis described above.

Take a free mock test

Related guides