TOEFLIELTS
English
LingoLeap Research

Do templates hurt your TOEFL score? We measured 3,692 answers.

Every prep site asserts that AI scoring flags template phrases. None of them publishes a number. We scored 3,692 Academic Discussion responses across 1,467 tasks and looked.

LingoLeap Research, 17 September 2026. Method, data handling and every limitation are on this page.

The finding

A template opener is associated with +0.04 points on the 0-30 scale, 95% CI -0.12 to +0.20.

That is a null, and it is a null with power: the study had 80% power to detect a difference of 0.23 points, and the confidence interval rules out any penalty larger than 0.12. This does not mean templates help — every positive number here is smaller than the grader's own resolution. It means the penalty everyone asserts did not appear.

Read this before you quote us

These are our AI grader's scores, not ETS's e-rater.

The scores in this study are produced by LingoLeap's own rubric-trained AI grader. We have shown what happens when an AI scoring system trained on the Academic Discussion rubric scores 3,692 responses. We have not tested the engine ETS uses, and nobody outside ETS can. The sites claiming e-rater flags templates have no data either; the difference is that we say so.

How common are template openers?

36.2% of responses (1,337 of 3,692) open with a stock phrase, and 55.8% (2,060 of 3,692) contain at least one anywhere. Template use is normal, not a fringe tactic. Any page telling learners that templates are a red flag is telling better than one in three of the cohort that their normal writing is a red flag.

The second finding in this table is that the specific phrases prep sites warn about are not the phrases learners use. "There are several reasons" appears as an opener twice in 3,692 responses. "It is widely acknowledged" twice. "I could not agree more" zero times, and that is a genuine zero from a matcher verified against a known-positive case first. The real template vocabulary is opinion-framing: the top two phrases alone are 24% of openers.

PhraseAs opener%Anywhere%
In my opinion48513.1%66217.9%
From / In my perspective38810.5%55315.0%
I firmly / strongly believe1413.8%44812.1%
When it comes to752.0%1123.0%
From my point of view611.7%792.1%
Personally speaking / Personally, I541.5%671.8%
As far as I am concerned / know521.4%912.5%
I wholeheartedly concur / agree511.4%722.0%
I strongly agree511.4%812.2%
In my view270.7%1153.1%
With the development / advent of100.3%1082.9%
In today's society / world50.1%340.9%
There are several / many reasons20.1%130.4%
It is widely acknowledged / believed20.1%310.8%
I could not agree more00.0%00.0%
Any stock phrase1,33736.2%2,06055.8%

Table 1. Prevalence of fifteen stock phrases, n = 3,692. "As opener" means the phrase matches in the first 90 characters.

The score difference is +0.04 points

Responses that open with a template average 25.198 on the 0-30 scale. Responses that do not average 25.158. The difference is +0.040, 95% CI -0.119 to +0.199, t = +0.49. There is no detectable effect in either direction.

With these group sizes, the minimum detectable effect at 80% power and two-sided alpha = 0.05 is 0.232 points. This is not an underpowered null hiding a large effect: a penalty of a quarter of a point or more would have shown up, and did not.

GroupnMean score (0-30)
Template opener1,33725.198
No template opener2,35525.158

The same result from the other end

If the grader punished templates, the top band would be depleted of them. It is not. Top scorers use template openers at essentially the corpus rate.

CohortnShare using a template opener
Score 28 or above44735.8%
Score 27 or above1,29035.6%
Whole corpus3,69236.2%
Score 22 or below75334.8%

Dose-response

By the number of distinct stock phrases in the response: 0 phrases 25.18 (n = 1,632), 1 phrase 25.19 (n = 1,705), 2 phrases 25.20 (n = 306), 3 or more 24.45 (n = 49). Flat through two phrases; the 3-or-more cell is 49 responses and too small to call.

The sub-scores point the opposite way from the claim

Each response carries three 0-5 sub-scores. This is the test of the specific mechanism the search results assert, which is a language penalty on formulaic phrasing. If templated language were being flagged as formulaic, Language Use would fall while Relevance stayed flat.

DimensionTemplateNo templateDifference95% CIt
Relevance & Contribution4.3984.372+0.025±0.027+1.81
Clarity & Elaboration3.9003.874+0.026±0.027+1.84
Language Use & Grammar3.7753.759+0.016±0.030+1.03

Table 2. Sub-score means, n = 1,335 template / 2,348 no-template for each dimension.

The observed pattern is the reverse ordering. Language Use is the flattest of the three, and no dimension is negative. Read this carefully, because it is easy to over-read: all three differences are positive and all three are practically zero. A difference of 0.025 on a 5-point scale quantised in 0.5 steps is one twentieth of the smallest step the grader can award. Nothing here says templates help. What it says is that the asserted mechanism does not appear.

After controlling for length the picture is the same: Relevance beta = +0.029 (t = 2.09), Clarity beta = +0.030 (t = 2.17), Language Use beta = +0.018 (t = 1.16). Three dimensions were tested, so t values near 2 should be read as marginal, and the effect sizes make the question academic.

Confounds and robustness

A number is only defensible if the obvious alternative explanations were ruled out. Five were.

Length

Length does predict score: r(length, score) = +0.134 across n = 3,692, and in an OLS a log-point of length is worth +1.99 score points (t = +12.8). Template users write slightly shorter responses, 158.2 versus 161.2 words, a difference of -3.0 ± 2.7, so the confound biases against templates if anything. Controlling for it changes nothing. Within length-quintile pooled difference: +0.024, 95% CI ±0.154, and quintile by quintile the sign alternates (+0.25, +0.22, -0.14, +0.13, -0.24), which is what noise around zero looks like. OLS of score on template and log words: template beta = +0.059, SE = 0.081, t = +0.73.

Prompt difficulty

Scores could vary by task. Restricting to the 260 tasks that contain both a template and a non-template response (2,165 responses) and pooling the within-task differences gives +0.235, 95% CI +0.004 to +0.468, bootstrapped over tasks with 4,000 reps. This is the only estimate in the study that is marginally distinguishable from zero, and it is positive, which is the opposite of a penalty. It uses 59% of the corpus and one sixth of the tasks, so it is the weakest of the three estimates. We report it because leaving it out would be selective.

Treatment definition

Defining the treatment as a stock phrase anywhere in the response rather than only in the opener (2,060 versus 1,632) gives an overall difference of -0.005 ± 0.159; Relevance +0.006, Clarity +0.007, Language Use -0.012. Still null on every dimension.

Interface locale

88% of records carry the cn interface locale. That is the UI language, not the response language: 0 of 3,247 cn records contain more than 5% CJK characters in the response text, confirming the responses are English in both groups. Split anyway, the cn difference is -0.012 ± 0.168 (n = 1,183 / 2,057) and the en difference is +0.428 ± 0.495 (n = 151 / 293). Both null; the en cell is small.

Duplicate responses

3,629 of the 3,692 response texts are distinct; 63 (1.7%) are exact duplicates of another response. Deduplicating changes nothing: prevalence 36.4%, difference +0.037, 95% CI -0.123 to +0.197 (n = 1,320 / 2,309).

The one exception: “When it comes to…”

Of the fifteen phrases, nine occur as openers at least 30 times and could be tested individually. Eight are indistinguishable from the baseline, with differences from -0.08 to +0.57 and confidence intervals spanning zero or nearly so. One is not.

Responses opening with “When it comes to…” (n = 75) average 23.573, which is 1.585 ± 0.512 below the no-template baseline of 25.158. The effect survives every control. These responses are longer than baseline, 209.8 versus 161.2 words, and still score lower, so length-adjusted the gap widens to -2.20 (SE 0.286, t = -7.7). It is spread over 56 distinct tasks with no single task contributing more than 7 responses, so it is not one bad prompt. Nine phrases were tested; at a Bonferroni threshold this one still clears comfortably. Removing it from the treatment set moves the main contrast to +0.136 ± 0.160, still null.

OpenernMeanvs baseline 25.158
When it comes to…7523.573-1.585 ± 0.512
In my opinion48525.388+0.229 ± 0.222
From / In my perspective38825.134-0.024 ± 0.254
I firmly / strongly believe14125.248+0.090 ± 0.378
From my point of view6125.557+0.399 ± 0.470
I wholeheartedly concur / agree5125.725+0.567 ± 0.574
I strongly agree5125.686+0.528 ± 0.514
Personally speaking / Personally, I5425.352+0.193 ± 0.611
As far as I am concerned / know5225.077-0.081 ± 0.686

Table 3. Per-opener means for the nine phrases with n >= 30, against the no-template baseline of 25.158.

What we can and cannot say

What we can say: one opener is a marker of lower-scoring responses. What we cannot say: that the phrase causes it. The drop falls on all three dimensions equally (Relevance -0.246, Clarity -0.234, Language Use -0.246), which is the signature of generally weaker writing rather than a phrase-level language flag. The evidence is correlational and n = 75. Treat it as a flag, not a rule.

Method

Population
Every non-legacy Academic Discussion task in LingoLeap's task registry (1,480 tasks) and all of its representative sample slugs, 3,729 distinct slugs. Academic Discussion is the only 2026-format task with a real corpus, which is why it is the population.
Retrieval
Each sample was fetched from the evaluation-sample endpoint with 8 concurrent workers and backoff-and-retry on HTTP 503. The backend rate-limits above roughly 8 concurrent reads; an unthrottled first pass at 16 workers lost 784 slugs to 503s, which is a retrieval artefact and not missing data, because every one of them came back on retry. Final: 3,723 records retrieved, 6 hard 404s, 0 unrecovered.
Two fields, and they are different things
The record nests a content field and a markdown field. The content field is the learner's response, and every measurement of template language is taken on that field only. The markdown field is the evaluator's commentary, which for cn-locale records (88% of the corpus) is largely written in Chinese and contains a rewritten model essay. A regex run over the markdown field would measure the grader's prose, not the learner's. It is used only to read scores.
Scoring fields
A numeric score field is present on 3,258 records. For the remaining 465 the score is parsed out of the markdown. On the 2,459 records where both are present they agree on 2,459 of 2,459 (100%), which is the validation for using the parsed value as a fallback. Per-dimension sub-scores parse from the markdown on 3,703 of 3,723 records (99.5%).
Exclusions and analysis set
13 records with no response text and 8 records scoring 0 with 0.0 on all three sub-dimensions (grader failures, not essays) were dropped. Analysis set: n = 3,692 responses across 1,467 distinct tasks. Mean score 25.17 (SD 2.42), median 26, IQR 24-27. Mean length 160 words, median 154.
Treatment definition
Fifteen stock phrases, derived from the data rather than assumed. Every response's first 120 characters were tokenised and the leading 3-, 4- and 5-word sequences counted across the corpus; the recurring formulae were then written as case-insensitive regexes. A response is a template-opener if any stock phrase matches in its first 90 characters.
Known-positive verification
Before any counting, the matcher was run against a response opening “As far as I am concerned, I firmly believe that…”, which matched, and against a non-formulaic opener, which did not. Both behaved as expected, so every zero in Table 1 is a zero from a matcher proven to fire on a positive case.
Sampling
These are representative sample slugs from the task registry, up to about 3 per task, not a random draw from all 16,479 Academic Discussion responses. They are selected for being published sample pages, which may over-represent well-formed responses. Nothing in the selection is correlated with template use, but the population is published samples, not all submissions.

Reproducing this

Every figure on this page comes from a single analysis script that re-fetches the corpus and reprints each number. The population, the retrieval method, the field distinction, the exclusions and the treatment definition are stated above in enough detail for a journalist or a competitor to check the work.

What this data does not show

The strongest version of a claim is the one that gets you caught. Six things this study does not establish:

  1. 1

    It is not evidence about ETS's e-rater or SpeechRater

    The scores here are produced by LingoLeap's own rubric-trained AI grader. We have measured how an AI scoring system trained on the Academic Discussion rubric treats templates. We have not tested the engine ETS uses, and nobody outside ETS can.

  2. 2

    It is not evidence that templates help

    Every positive number in this study is smaller than the grader's own resolution. The correct verdict is no detectable effect, not that templates are safe and certainly not that templates work.

  3. 3

    It does not cover speaking

    Academic Discussion is writing. Claims about spoken templates are untested here.

  4. 4

    It does not generalise to all TOEFL responses

    The population is LingoLeap's published sample responses to non-legacy Academic Discussion tasks, up to about 3 per task.

  5. 5

    It says nothing about content templates

    We measured opening formulae. Memorised example paragraphs, invented statistics and pre-written body arguments are something this study did not look at. A learner who memorises an entire essay is doing something else.

  6. 6

    It cannot separate the template from the template user

    This is observational. Nobody was randomly assigned a template. The within-task and within-length-quintile estimates narrow the confound but do not remove it.

Frequently asked questions

Do TOEFL templates hurt your score?
In our measurement of 3,692 AI-scored Academic Discussion responses, a template opener is associated with +0.04 points on the 0-30 scale, 95% CI -0.12 to +0.20. That is no detectable effect in either direction, and the study had 80% power to detect a difference of 0.23 points. These are LingoLeap's rubric-trained AI grader's scores, not ETS's e-rater, which nobody outside ETS can test.
Does AI scoring penalize template phrases?
Not in our data, and the sub-scores point the opposite way from the claim. If templated language were flagged as formulaic, the Language Use and Grammar sub-score would fall. It is the flattest of the three dimensions at +0.016 of 5 (t = 1.03), while Relevance is +0.025 and Clarity +0.026. No dimension is penalised. All three differences are practically zero.
How many TOEFL test takers actually use templates?
36.2% of the 3,692 responses open with a stock phrase and 55.8% contain at least one anywhere. The most common openers are opinion-framing phrases: In my opinion (13.1% of responses as an opener) and From or In my perspective (10.5%). The phrases prep sites usually warn about are rare: There are several reasons appears as an opener twice in 3,692 responses.
Is there any opening phrase that does correlate with lower scores?
One. Responses opening with When it comes to (n = 75) average 23.573 against a no-template baseline of 25.158, a gap of 1.585 ± 0.512 that widens to -2.20 after adjusting for length. The drop falls on all three sub-dimensions equally, which is the signature of generally weaker writing rather than a penalty on the phrase. The evidence is correlational: treat it as a flag, not a rule.
Does a longer response score higher, and does that explain the result?
Length does predict score, r = +0.134, and a log-point of length is worth +1.99 points in an OLS. It does not explain the result. Template users write slightly shorter responses (158.2 versus 161.2 words), so the confound biases against templates. Within length quintiles the pooled difference is +0.024 ± 0.154, and an OLS controlling for log words gives a template coefficient of +0.059 (SE 0.081, t = +0.73).
Does this mean I should use a template?
It means the penalty everyone asserts did not appear in 3,692 scored responses. It does not mean templates help: every positive number here is smaller than the grader's resolution. It also says nothing about memorised body paragraphs, invented statistics or pre-written arguments, because we measured opening formulae only. A template that frames your own reasoning is not the same thing as a memorised essay.

Get your own responses scored

The grader in this study is the one that scores your practice. Write an Academic Discussion response and see the three sub-scores for yourself.

Start practising free