Do templates hurt your TOEFL score? We measured 3,692 answers.
Every prep site asserts that AI scoring flags template phrases. None of them publishes a number. We scored 3,692 Academic Discussion responses across 1,467 tasks and looked.
LingoLeap Research, 17 September 2026. Method, data handling and every limitation are on this page.
The finding
A template opener is associated with +0.04 points on the 0-30 scale, 95% CI -0.12 to +0.20.
That is a null, and it is a null with power: the study had 80% power to detect a difference of 0.23 points, and the confidence interval rules out any penalty larger than 0.12. This does not mean templates help — every positive number here is smaller than the grader's own resolution. It means the penalty everyone asserts did not appear.
Read this before you quote us
These are our AI grader's scores, not ETS's e-rater.
The scores in this study are produced by LingoLeap's own rubric-trained AI grader. We have shown what happens when an AI scoring system trained on the Academic Discussion rubric scores 3,692 responses. We have not tested the engine ETS uses, and nobody outside ETS can. The sites claiming e-rater flags templates have no data either; the difference is that we say so.
How common are template openers?
36.2% of responses (1,337 of 3,692) open with a stock phrase, and 55.8% (2,060 of 3,692) contain at least one anywhere. Template use is normal, not a fringe tactic. Any page telling learners that templates are a red flag is telling better than one in three of the cohort that their normal writing is a red flag.
The second finding in this table is that the specific phrases prep sites warn about are not the phrases learners use. "There are several reasons" appears as an opener twice in 3,692 responses. "It is widely acknowledged" twice. "I could not agree more" zero times, and that is a genuine zero from a matcher verified against a known-positive case first. The real template vocabulary is opinion-framing: the top two phrases alone are 24% of openers.
| Phrase | As opener | % | Anywhere | % |
|---|---|---|---|---|
| In my opinion | 485 | 13.1% | 662 | 17.9% |
| From / In my perspective | 388 | 10.5% | 553 | 15.0% |
| I firmly / strongly believe | 141 | 3.8% | 448 | 12.1% |
| When it comes to | 75 | 2.0% | 112 | 3.0% |
| From my point of view | 61 | 1.7% | 79 | 2.1% |
| Personally speaking / Personally, I | 54 | 1.5% | 67 | 1.8% |
| As far as I am concerned / know | 52 | 1.4% | 91 | 2.5% |
| I wholeheartedly concur / agree | 51 | 1.4% | 72 | 2.0% |
| I strongly agree | 51 | 1.4% | 81 | 2.2% |
| In my view | 27 | 0.7% | 115 | 3.1% |
| With the development / advent of | 10 | 0.3% | 108 | 2.9% |
| In today's society / world | 5 | 0.1% | 34 | 0.9% |
| There are several / many reasons | 2 | 0.1% | 13 | 0.4% |
| It is widely acknowledged / believed | 2 | 0.1% | 31 | 0.8% |
| I could not agree more | 0 | 0.0% | 0 | 0.0% |
| Any stock phrase | 1,337 | 36.2% | 2,060 | 55.8% |
Table 1. Prevalence of fifteen stock phrases, n = 3,692. "As opener" means the phrase matches in the first 90 characters.
The score difference is +0.04 points
Responses that open with a template average 25.198 on the 0-30 scale. Responses that do not average 25.158. The difference is +0.040, 95% CI -0.119 to +0.199, t = +0.49. There is no detectable effect in either direction.
With these group sizes, the minimum detectable effect at 80% power and two-sided alpha = 0.05 is 0.232 points. This is not an underpowered null hiding a large effect: a penalty of a quarter of a point or more would have shown up, and did not.
| Group | n | Mean score (0-30) |
|---|---|---|
| Template opener | 1,337 | 25.198 |
| No template opener | 2,355 | 25.158 |
The same result from the other end
If the grader punished templates, the top band would be depleted of them. It is not. Top scorers use template openers at essentially the corpus rate.
| Cohort | n | Share using a template opener |
|---|---|---|
| Score 28 or above | 447 | 35.8% |
| Score 27 or above | 1,290 | 35.6% |
| Whole corpus | 3,692 | 36.2% |
| Score 22 or below | 753 | 34.8% |
Dose-response
By the number of distinct stock phrases in the response: 0 phrases 25.18 (n = 1,632), 1 phrase 25.19 (n = 1,705), 2 phrases 25.20 (n = 306), 3 or more 24.45 (n = 49). Flat through two phrases; the 3-or-more cell is 49 responses and too small to call.
The sub-scores point the opposite way from the claim
Each response carries three 0-5 sub-scores. This is the test of the specific mechanism the search results assert, which is a language penalty on formulaic phrasing. If templated language were being flagged as formulaic, Language Use would fall while Relevance stayed flat.
| Dimension | Template | No template | Difference | 95% CI | t |
|---|---|---|---|---|---|
| Relevance & Contribution | 4.398 | 4.372 | +0.025 | ±0.027 | +1.81 |
| Clarity & Elaboration | 3.900 | 3.874 | +0.026 | ±0.027 | +1.84 |
| Language Use & Grammar | 3.775 | 3.759 | +0.016 | ±0.030 | +1.03 |
Table 2. Sub-score means, n = 1,335 template / 2,348 no-template for each dimension.
The observed pattern is the reverse ordering. Language Use is the flattest of the three, and no dimension is negative. Read this carefully, because it is easy to over-read: all three differences are positive and all three are practically zero. A difference of 0.025 on a 5-point scale quantised in 0.5 steps is one twentieth of the smallest step the grader can award. Nothing here says templates help. What it says is that the asserted mechanism does not appear.
After controlling for length the picture is the same: Relevance beta = +0.029 (t = 2.09), Clarity beta = +0.030 (t = 2.17), Language Use beta = +0.018 (t = 1.16). Three dimensions were tested, so t values near 2 should be read as marginal, and the effect sizes make the question academic.
Confounds and robustness
A number is only defensible if the obvious alternative explanations were ruled out. Five were.
Length
Length does predict score: r(length, score) = +0.134 across n = 3,692, and in an OLS a log-point of length is worth +1.99 score points (t = +12.8). Template users write slightly shorter responses, 158.2 versus 161.2 words, a difference of -3.0 ± 2.7, so the confound biases against templates if anything. Controlling for it changes nothing. Within length-quintile pooled difference: +0.024, 95% CI ±0.154, and quintile by quintile the sign alternates (+0.25, +0.22, -0.14, +0.13, -0.24), which is what noise around zero looks like. OLS of score on template and log words: template beta = +0.059, SE = 0.081, t = +0.73.
Prompt difficulty
Scores could vary by task. Restricting to the 260 tasks that contain both a template and a non-template response (2,165 responses) and pooling the within-task differences gives +0.235, 95% CI +0.004 to +0.468, bootstrapped over tasks with 4,000 reps. This is the only estimate in the study that is marginally distinguishable from zero, and it is positive, which is the opposite of a penalty. It uses 59% of the corpus and one sixth of the tasks, so it is the weakest of the three estimates. We report it because leaving it out would be selective.
Treatment definition
Defining the treatment as a stock phrase anywhere in the response rather than only in the opener (2,060 versus 1,632) gives an overall difference of -0.005 ± 0.159; Relevance +0.006, Clarity +0.007, Language Use -0.012. Still null on every dimension.
Interface locale
88% of records carry the cn interface locale. That is the UI language, not the response language: 0 of 3,247 cn records contain more than 5% CJK characters in the response text, confirming the responses are English in both groups. Split anyway, the cn difference is -0.012 ± 0.168 (n = 1,183 / 2,057) and the en difference is +0.428 ± 0.495 (n = 151 / 293). Both null; the en cell is small.
Duplicate responses
3,629 of the 3,692 response texts are distinct; 63 (1.7%) are exact duplicates of another response. Deduplicating changes nothing: prevalence 36.4%, difference +0.037, 95% CI -0.123 to +0.197 (n = 1,320 / 2,309).
The one exception: “When it comes to…”
Of the fifteen phrases, nine occur as openers at least 30 times and could be tested individually. Eight are indistinguishable from the baseline, with differences from -0.08 to +0.57 and confidence intervals spanning zero or nearly so. One is not.
Responses opening with “When it comes to…” (n = 75) average 23.573, which is 1.585 ± 0.512 below the no-template baseline of 25.158. The effect survives every control. These responses are longer than baseline, 209.8 versus 161.2 words, and still score lower, so length-adjusted the gap widens to -2.20 (SE 0.286, t = -7.7). It is spread over 56 distinct tasks with no single task contributing more than 7 responses, so it is not one bad prompt. Nine phrases were tested; at a Bonferroni threshold this one still clears comfortably. Removing it from the treatment set moves the main contrast to +0.136 ± 0.160, still null.
| Opener | n | Mean | vs baseline 25.158 |
|---|---|---|---|
| When it comes to… | 75 | 23.573 | -1.585 ± 0.512 |
| In my opinion | 485 | 25.388 | +0.229 ± 0.222 |
| From / In my perspective | 388 | 25.134 | -0.024 ± 0.254 |
| I firmly / strongly believe | 141 | 25.248 | +0.090 ± 0.378 |
| From my point of view | 61 | 25.557 | +0.399 ± 0.470 |
| I wholeheartedly concur / agree | 51 | 25.725 | +0.567 ± 0.574 |
| I strongly agree | 51 | 25.686 | +0.528 ± 0.514 |
| Personally speaking / Personally, I | 54 | 25.352 | +0.193 ± 0.611 |
| As far as I am concerned / know | 52 | 25.077 | -0.081 ± 0.686 |
Table 3. Per-opener means for the nine phrases with n >= 30, against the no-template baseline of 25.158.
What we can and cannot say
What we can say: one opener is a marker of lower-scoring responses. What we cannot say: that the phrase causes it. The drop falls on all three dimensions equally (Relevance -0.246, Clarity -0.234, Language Use -0.246), which is the signature of generally weaker writing rather than a phrase-level language flag. The evidence is correlational and n = 75. Treat it as a flag, not a rule.
Method
- Population
- Every non-legacy Academic Discussion task in LingoLeap's task registry (1,480 tasks) and all of its representative sample slugs, 3,729 distinct slugs. Academic Discussion is the only 2026-format task with a real corpus, which is why it is the population.
- Retrieval
- Each sample was fetched from the evaluation-sample endpoint with 8 concurrent workers and backoff-and-retry on HTTP 503. The backend rate-limits above roughly 8 concurrent reads; an unthrottled first pass at 16 workers lost 784 slugs to 503s, which is a retrieval artefact and not missing data, because every one of them came back on retry. Final: 3,723 records retrieved, 6 hard 404s, 0 unrecovered.
- Two fields, and they are different things
- The record nests a content field and a markdown field. The content field is the learner's response, and every measurement of template language is taken on that field only. The markdown field is the evaluator's commentary, which for cn-locale records (88% of the corpus) is largely written in Chinese and contains a rewritten model essay. A regex run over the markdown field would measure the grader's prose, not the learner's. It is used only to read scores.
- Scoring fields
- A numeric score field is present on 3,258 records. For the remaining 465 the score is parsed out of the markdown. On the 2,459 records where both are present they agree on 2,459 of 2,459 (100%), which is the validation for using the parsed value as a fallback. Per-dimension sub-scores parse from the markdown on 3,703 of 3,723 records (99.5%).
- Exclusions and analysis set
- 13 records with no response text and 8 records scoring 0 with 0.0 on all three sub-dimensions (grader failures, not essays) were dropped. Analysis set: n = 3,692 responses across 1,467 distinct tasks. Mean score 25.17 (SD 2.42), median 26, IQR 24-27. Mean length 160 words, median 154.
- Treatment definition
- Fifteen stock phrases, derived from the data rather than assumed. Every response's first 120 characters were tokenised and the leading 3-, 4- and 5-word sequences counted across the corpus; the recurring formulae were then written as case-insensitive regexes. A response is a template-opener if any stock phrase matches in its first 90 characters.
- Known-positive verification
- Before any counting, the matcher was run against a response opening “As far as I am concerned, I firmly believe that…”, which matched, and against a non-formulaic opener, which did not. Both behaved as expected, so every zero in Table 1 is a zero from a matcher proven to fire on a positive case.
- Sampling
- These are representative sample slugs from the task registry, up to about 3 per task, not a random draw from all 16,479 Academic Discussion responses. They are selected for being published sample pages, which may over-represent well-formed responses. Nothing in the selection is correlated with template use, but the population is published samples, not all submissions.
Reproducing this
Every figure on this page comes from a single analysis script that re-fetches the corpus and reprints each number. The population, the retrieval method, the field distinction, the exclusions and the treatment definition are stated above in enough detail for a journalist or a competitor to check the work.
What this data does not show
The strongest version of a claim is the one that gets you caught. Six things this study does not establish:
- 1
It is not evidence about ETS's e-rater or SpeechRater
The scores here are produced by LingoLeap's own rubric-trained AI grader. We have measured how an AI scoring system trained on the Academic Discussion rubric treats templates. We have not tested the engine ETS uses, and nobody outside ETS can.
- 2
It is not evidence that templates help
Every positive number in this study is smaller than the grader's own resolution. The correct verdict is no detectable effect, not that templates are safe and certainly not that templates work.
- 3
It does not cover speaking
Academic Discussion is writing. Claims about spoken templates are untested here.
- 4
It does not generalise to all TOEFL responses
The population is LingoLeap's published sample responses to non-legacy Academic Discussion tasks, up to about 3 per task.
- 5
It says nothing about content templates
We measured opening formulae. Memorised example paragraphs, invented statistics and pre-written body arguments are something this study did not look at. A learner who memorises an entire essay is doing something else.
- 6
It cannot separate the template from the template user
This is observational. Nobody was randomly assigned a template. The within-task and within-length-quintile estimates narrow the confound but do not remove it.
Frequently asked questions
Do TOEFL templates hurt your score?
Does AI scoring penalize template phrases?
How many TOEFL test takers actually use templates?
Is there any opening phrase that does correlate with lower scores?
Does a longer response score higher, and does that explain the result?
Does this mean I should use a template?
Get your own responses scored
The grader in this study is the one that scores your practice. Write an Academic Discussion response and see the three sub-scores for yourself.
Start practising freeRelated guides
TOEFL templates
The full template library for every 2026 writing and speaking task.
Academic Discussion template
The task this study measured, with a structure you fill with your own reasoning.
Writing rubrics
The three dimensions the sub-scores in this study decompose into.
LingoLeap Research
Every original-data study we have published.