PTE Academic's Speaking & Writing section opens the test and sets the tone for everything after it: you talk into a microphone while an automated speech-recognition engine scores you in real time, with no examiner in the room and no chance to restart a recording. That format rewards a very specific skill set — steady pacing, clear articulation, and answers that hit their content targets quickly — rather than charisma or vocabulary range. This guide explains how the scoring actually works, walks through each of the five speaking task types with concrete technique, and lays out a practice routine you can run in the weeks before your test.
How PTE Speaking is scored
Every scored speaking task is marked against three broad criteria: content (did you say the right things, or reproduce the source text accurately), oral fluency (rhythm, phrasing, and continuity of speech, judged by whether it sounds like natural connected speech rather than a string of isolated words), and pronunciation (how clearly individual sounds and word stress are produced, judged against a standard model rather than any one accent).
Speaking tasks contribute not just to the Speaking communicative skill score but also feed into your overall score profile — Read Aloud content, for instance, also counts toward reading, and Repeat Sentence content counts toward listening. This is why examiners of the PTE system treat speaking accuracy as connected to your other section scores, not siloed.
Most speaking items are scored automatically by Pearson's AI engine the moment you finish recording. Describe Image and Re-tell Lecture are among the task types where human raters also periodically review a sample of responses alongside the automated score, so consistency between what you say and how you say it matters on every attempt, not just the ones you think are being checked.
How the AI actually listens
The scoring engine is a speech-recognition and prosody model, not a person having a conversation with you. It cannot infer meaning from tone of voice, humour, or hesitant self-correction the way a human listener might. Practically, that has a few consequences worth building your technique around.
First, timing is unforgiving: the microphone starts recording after a beep and stops automatically, so a slow start or a long pause mid-answer eats into your real speaking time and can be read as broken fluency. Second, the engine expects continuous, evenly-paced speech — audibly restarting a sentence, trailing off, or leaving long silent gaps lowers the fluency score even if the words you do get out are accurate. Third, because it is a recognition model, mumbling, extreme rushing, or speaking away from the microphone will reduce your pronunciation score even when your grammar and vocabulary are fine.
None of this means you need a 'neutral' or invented accent. The engine is trained on many accents and rewards clear articulation and correct stress patterns, not the sound of any single nationality's English.
A practical practice routine
Because the format is so mechanical, PTE speaking responds well to targeted, repeatable drilling rather than general 'speaking practice'. A routine that works for most candidates spends dedicated time on each of the weaknesses the scoring model actually penalises.
Record everything. You cannot judge your own pacing or spot filler words by feeling alone — play the recording back and check it against the time limit and against a transcript of what you meant to say.
Daily, 10–15 minutes: Read Aloud on unfamiliar text, focusing purely on steady pace and not stopping to self-correct.
3–4 times a week: Repeat Sentence drills using short audio clips, building up sentence length gradually.
3–4 times a week: Describe Image and Re-tell Lecture using a fixed template so you never run out of things to say before time is up.
Weekly: a full timed mock of all five speaking tasks back-to-back, to build stamina and get used to the beep-and-record rhythm.
Ongoing: read your Read Aloud and Answer Short Question scripts aloud to a phone recorder and listen for words you consistently mispronounce or mumble.
Task by task
Read Aloud
30–40s to prepare, then read; text is up to 60 words Content, Pronunciation, Oral fluency (also feeds Reading)
How to do it
1Use the preparation time to silently scan for hard words, numbers, and where the natural sentence breaks are — mark them mentally rather than trying to memorise the passage.
2Read in phrases, not word by word: group two to four words together and let your voice fall slightly at commas and full stops, the way you would in normal conversation.
3Keep a steady pace even if you stumble on a word — a brief recovery is far less costly than a long pause or restarting the sentence.
4Make sure your voice reaches a normal, confident volume for the whole sentence; trailing off at the end of long sentences is a common fluency deduction.
Common mistakes
Racing through the text to 'finish early', which flattens intonation and hurts fluency scoring.
Stopping completely and restarting after a misread word instead of continuing smoothly.
Reading in a flat monotone with no phrase grouping, which reads as unnatural to the fluency model.
Repeat Sentence
Sentence plays once (3–9 seconds), then you have 15 seconds to speak Content, Pronunciation, Oral fluency (also feeds Listening)
How to do it
1Start speaking almost immediately after the beep — waiting to 'get it perfect in your head' wastes response time and adds silence at the start.
2Prioritise reproducing as many content words as possible over getting every function word exact; content is scored heavily on word-for-word match but partial answers still earn partial credit.
3If you miss a word in the middle, keep going and fill the gap with your best guess rather than stopping — an unbroken attempt scores better than silence.
4Practise with sentences of increasing length so your working-memory span for holding a sentence stretches naturally over time.
Common mistakes
Staying silent because you didn't catch the whole sentence — even a partial, confident repetition beats no answer.
Adding words that weren't in the original, which is penalised as inaccurate content.
Speaking too quietly or trailing off, which reduces the pronunciation score even when the words are correct.
Describe Image
25s to prepare, then 40s to speak Content, Pronunciation, Oral fluency; reviewed by both AI and human raters
How to do it
1Use a fixed template so you never freeze: an opening line naming the image type, one or two sentences on the overall trend or main feature, then two to three sentences on specific data points or details, and a closing sentence on the overall takeaway.
2In the 25 seconds of prep, identify the single biggest feature of the chart or image (the highest bar, biggest change, most unusual element) — that becomes your main sentence.
3Keep talking for close to the full 40 seconds; stopping early because you've 'said everything' costs content score, since raters expect a reasonably full description.
4Use approximate figures rather than trying to read out exact numbers from a small chart — 'just under half' is safer and faster than misreading a precise value.
Common mistakes
Listing every single data point mechanically instead of grouping them around a trend or comparison.
Running out of things to say after 15–20 seconds and going silent for the remaining time.
Ignoring axis labels or units and describing numbers without context.
Re-tell Lecture
Lecture audio/video plays (up to ~90s), 10s to prepare, then 40s to speak Content, Pronunciation, Oral fluency; reviewed by both AI and human raters
How to do it
1While the lecture plays, jot down keywords only — the topic, any numbers, and the two or three points the speaker returns to or emphasises.
2Open with a one-sentence summary of the lecture's overall topic, then move through your noted points roughly in the order they were mentioned.
3Paraphrase rather than trying to quote the lecture verbatim; the scoring rewards accurate content and coherent structure, not exact wording.
4If the lecture had a clear structure (problem then solution, cause then effect), mirror that structure in your re-telling so it sounds organised.
Common mistakes
Trying to write full sentences of notes while listening, which means missing later content.
Speaking only about the first third of the lecture because notes for the rest ran out.
Going silent for several seconds while trying to recall a detail instead of moving on to the next point.
Answer Short Question
Audio plays a short question, then 10s to answer Content only (single word or short phrase)
How to do it
1Answer with the shortest correct word or phrase possible — this task rewards accuracy, not elaboration or full sentences.
2Answer immediately; the question is usually general-knowledge or vocabulary-based (e.g. antonyms, categories, common facts), so there is little to gain from long thinking time.
3If unsure of the exact term, give your best single-word guess rather than staying silent, since an attempt can still be marked correct or partially credited.
Common mistakes
Answering in a full sentence when a single word was expected, which doesn't help the score and wastes the short response window.
Overthinking and running out of the 10-second window without answering at all.
Common questions
Is PTE Speaking scored by a person or by AI?
It is scored automatically by Pearson's AI scoring engine immediately after you respond. For some task types — including Describe Image and Re-tell Lecture — human raters also review a portion of responses alongside the automated score, so both accurate content and clear, natural delivery matter.
Does my accent affect my PTE Speaking score?
The scoring model is trained to recognise a wide range of native and non-native English accents, so having an accent is not itself penalised. What is scored is clear pronunciation of sounds and correct word stress — mumbling, very quiet speech, or heavily distorted sounds will lower the pronunciation score regardless of accent.
What happens if I pause for too long while speaking?
Long silences during a speaking response are read by the scoring model as broken fluency and can also mean the microphone stops recording before you finish, since response time is fixed and automatic. It's better to keep speaking, even with minor errors, than to stop and think mid-answer.
Should I aim to use the full response time on every speaking task?
For Describe Image and Re-tell Lecture, yes — stopping well before time is up usually means you haven't given enough content and can cost marks. For Repeat Sentence and Answer Short Question, the goal is accuracy and appropriate brevity, not filling time, since these responses are naturally short.
Can I redo a speaking response if I make a mistake?
No. Each speaking item is recorded once within its fixed time window and cannot be replayed or re-recorded during the test. If you misspeak, the better strategy is to correct yourself briefly and keep going smoothly rather than stopping, since an unbroken attempt scores better than a restarted one.
Now record yourself
Reading strategy only gets you so far. Linguently gives you timed PTE speaking tasks with AI feedback on fluency and pronunciation, plus real conversation partners to practise with.