Four criteria, equally weighted: Fluency and Coherence, Lexical Resource, Grammatical Range and Accuracy, and Pronunciation. Each is scored on the nine-band scale, and the overall is the average, rounded.
Which means a candidate at 7, 7, 6, 6 and a candidate at 8, 6, 6, 6 both land near the same overall while needing completely different work. That is the case for looking past the headline number.
Fluency and Coherence
Not speed. Fluency here means the ability to keep going at a natural pace without the effort showing, and coherence means the answer holds together.
What gets noticed, in rough order of cost:
- Hesitation for language rather than for content. Pausing to think about what to say is normal and largely forgiven. Pausing to hunt for a word is the thing being measured.
- Self-repairs. Starting a sentence, abandoning it, starting again. One or two are human; a pattern of them says the sentence was not planned before it began.
- Length of run. How many words you get out between pauses. Short runs read as effortful even when the vocabulary is good.
- Discourse markers. Whether ideas are linked at all, and whether the linking is varied or the same three connectors on repeat.
Fillers are worth a specific note. "Um" and "you know" are not automatically penalised - native speakers use them constantly. They cost you when they cluster in place of content, because then they are hesitation wearing a disguise.
Lexical Resource
The most commonly misunderstood of the four. Candidates read "vocabulary" and reach for long words, which is close to the opposite of what the descriptors reward.
What is being assessed is range and precision: whether you have enough words to talk about unfamiliar topics, whether you pick the right one, and whether you can paraphrase when you do not have it.
- Range means not recycling. Using "important" six times scores below using "significant", "crucial" and "central" where each fits.
- Precision beats difficulty. A common word used exactly outperforms a rare one used approximately, and a misused rare word is worse than no rare word.
- Collocation is where advanced candidates gain most. "Heavy traffic" and "make a decision" are unremarkable to a native ear and obviously wrong when swapped.
- Paraphrase under pressure is explicitly rewarded. Not knowing a word is fine; stalling because you do not know it is not.
Idioms are the classic trap. The descriptors mention idiomatic language at the upper bands, so candidates memorise a handful and deploy them regardless of fit. An examiner notices a forced idiom immediately, and it reads as rehearsed rather than resourceful.
Grammatical Range and Accuracy
Two things, and candidates usually optimise for the wrong one.
Accuracy is measured as the proportion of error-free clauses. Range is whether complex structures appear at all: subordination, conditionals, relative clauses, a spread of tenses, the passive where it belongs.
The trap is that these pull against each other. Speak only in short simple sentences and your accuracy will look excellent while your range caps the score around band 5 or 6. The descriptors are explicit that a candidate who attempts complex structures with some errors outscores one who avoids them entirely.
So the practical advice is uncomfortable: take grammatical risks. A conditional that comes out slightly wrong is worth more than a safe sentence that comes out perfectly.
Pronunciation
The criterion candidates worry about for the wrong reason. No accent scores higher than another. A Nigerian, Indian, Scottish or Brazilian accent can all reach band 9. What is assessed is:
- Intelligibility - how much effort the listener has to make, and whether they ever lose you entirely.
- Word and sentence stress - whether emphasis lands where meaning requires it.
- Intonation - whether the pitch pattern carries meaning, or runs flat.
- Sense grouping - whether words are chunked into meaningful units or delivered one at a time.
Rhythm is the feature that separates bands most reliably. English compresses unstressed syllables hard; speakers of syllable-timed languages tend to give every syllable equal weight, which stays intelligible but sounds effortful and caps the score.
What to do with all this
Three things, in order.
First, get your four criteria separately rather than as an overall. The gap between your best and worst is usually a band and a half, and that gap is where the entire opportunity sits.
Second, work on one criterion at a time. Attention is the constraint - trying to fix grammar, vocabulary and pausing inside the same two-minute answer reliably fixes none of them.
Third, get the measurements rather than the impression. "You hesitate too much" is not actionable. "Eleven hesitation pauses, average run of six words, four self-repairs" tells you what to count next time.
An app that scores on the device can hand you those numbers for every answer you record, which is the part a weekly lesson with a tutor cannot do - not because the tutor is worse, but because nobody counts pauses in real time.