Speaking · Task 3 — Describing a Scene

CELPIP Speaking Task 3 Template — Describing a Scene

Master CELPIP Speaking Task 3 with a proven describe-a-scene template, spatial-vocabulary banks, a 60-second model answer, and a clear CLB 7-to-11 roadmap.

9 min read

CELPIP Speaking Task 3 asks you to describe a single picture to a listener who cannot see it. You get 30 seconds to prepare and 60 seconds to speak, and the goal is simple: paint the scene so clearly that someone listening could draw it. This guide gives you a systematic structure, two reusable phrase banks, a fill-in template, a scored model answer, and a CLB 7-to-11 roadmap so you stop rambling and start describing with control.

The trick is that there is no "right answer" to find — the picture is fully visible to you. What the examiner rewards is organised, detailed, fluent description. Most test-takers lose marks not because they lack words, but because they jump randomly around the image, repeat "there is a man... there is a woman... there is a dog," and run out of things to say at the 40-second mark. A repeatable route through the scene fixes all of that.

What the examiner is looking for

Task 3 is scored on the four standard CELPIP Speaking criteria. Here is what each means for this task and how to hit it.

CriterionWhat it means hereHow to hit it
Content & CoherenceA logical, complete tour of the picture — overview, then a systematic sweepMove foreground to background (or left to right) and never backtrack randomly
VocabularyPrecise nouns, action verbs, position words and adjectivesName objects specifically ("a striped umbrella," not "a thing"); use verbs of posture and motion
ListenabilitySmooth pace, clear linking, natural stress and pausingUse transitions ("In the foreground...," "To the right...") and avoid long silent gaps
Task FulfilmentYou genuinely describe a scene in detail for the full 60 secondsKeep speaking until the timer ends; add detail rather than stopping early

The single biggest differentiator is Content & Coherence: a describer who works in a deliberate order always sounds more competent than one who lists objects at random, even if the random speaker knows more words.

The 30-second plan: the foreground-to-background sweep

You cannot plan every word in 30 seconds, but you can lock in a route. The strongest structure for Task 3 is a four-zone sweep that gives the listener a mental map before you fill in detail.

Step 1 — Give a one-line overview (about 8 seconds)

Open by telling the listener where we are and the general situation, so everything after has context. Resist diving straight into one tiny object. Something like: "This picture shows a busy outdoor farmers' market on what looks like a sunny weekend morning." In one breath the listener knows the setting, the mood, and roughly how many people to expect. This is your topic sentence — it earns Content & Coherence marks immediately.

Step 2 — Describe the foreground (about 18 seconds)

Move to what is closest and most prominent. Name the main people or objects, say where they are using position language, and — crucially — say what they are doing using the present continuous. "In the foreground, on the left, a young woman is paying for a basket of apples while the vendor is handing her some change." Notice the combination: position word + subject + present-continuous verb. That formula is the engine of the whole answer.

Step 3 — Move to the centre and the people (about 18 seconds)

Now sweep to the middle of the scene and the secondary characters. Keep using present continuous and add small human details that make the scene feel alive. "In the centre, two children are pointing at a fruit stand, and behind them a man wearing a green apron is arranging vegetables." Mentioning clothing, gestures and reactions shows range of vocabulary and keeps you talking past the 40-second wall where most people freeze.

Step 4 — Finish with the background and overall impression (about 16 seconds)

Close by zooming out to what is far away, then add a one-line interpretation of the mood. "In the background, I can see striped tents and a few trees, and the whole place looks lively and cheerful." Ending with atmosphere signals to the listener (and the rater) that you have given a complete picture, not a half-finished list. If you still have a few seconds, add one more small detail rather than going silent.

The fill-in template

Memorise this skeleton and slot the picture's details into the brackets. The bracketed verbs should almost always be present continuous (is/are + -ing).

OVERVIEW (~8s):
This picture shows [setting / place] on what looks like a [time of day / weather] [day].
It seems to be [general situation — e.g. a busy market / a quiet office].

FOREGROUND (~18s):
In the foreground, on the [left / right], [person/object] is [-ing verb] [detail].
Next to [him / her / it], I can see [object], which is [adjective].

CENTRE & PEOPLE (~18s):
In the centre, [person(s)] is/are [-ing verb], and [he/she/they] look(s) [emotion adjective].
Behind them, [person] wearing [clothing] is [-ing verb].

BACKGROUND & IMPRESSION (~16s):
In the background, I can see [object/scenery] and [object].
Overall, the scene looks [adjective] and gives the impression that [interpretation].

Sentence-starter bank for description

Grab one starter from each row as you move through the sweep. Rotating starters (rather than repeating "there is") is exactly what lifts your Listenability and Vocabulary scores.

StageSentence starters
Overview"This picture shows..." / "This is a photo of..." / "It looks like we're in..." / "The scene takes place at..."
Foreground"In the foreground..." / "The first thing I notice is..." / "Closest to me, there is..." / "In front, on the left..."
Position"To the left/right of..." / "In the centre..." / "Next to / beside..." / "Just behind..." / "In the middle of..."
People & actions"A man/woman is [-ing]..." / "Two people are [-ing]..." / "He/She seems to be [-ing]..." / "They look as if they are..."
Details"I can also see..." / "There's a [adjective] [noun]..." / "Interestingly,..." / "One small detail is..."
Background"In the background..." / "Further back..." / "At the very top of the picture..." / "In the distance..."
Impression"Overall, the mood is..." / "The whole scene looks..." / "It gives me the feeling that..." / "Everyone seems..."

Spatial and descriptive vocabulary table

This is the language that separates a vague answer from a vivid one. Pull a preposition of place, an action verb, and an adjective into almost every sentence.

Prepositions of placeAction verbs (present continuous)Adjectives (people & scene)
in the foreground / backgroundis standing / sitting / leaningcrowded / busy / lively
in the centre / middleis walking / running / strollingquiet / calm / peaceful
on the left / rightis holding / carrying / pointingsunny / bright / overcast
next to / besideis talking / chatting / laughingcheerful / relaxed / excited
behind / in front ofis buying / paying / handing overcolourful / vibrant / dull
above / below / on top ofis looking at / staring at / watchingcrowded with / full of
between / amongis wearing / dressed inspacious / open / cramped
in the corner / at the edgeis reaching for / picking uptidy / messy / cluttered
in the distance / far awayis smiling / frowning / wavingmodern / old-fashioned / rustic
all around / scattered acrossis arranging / stacking / sortingwarm / cosy / chilly

Buying time & staying fluent

The enemy in Task 3 is the dead silence at second 45. These natural fillers buy you a beat to think while still producing English — which protects your Listenability score far better than an audible pause. Use them sparingly; one or two is natural, five sounds nervous.

PurposeNatural phrases
Soft start / hesitation"Let me see..." / "So, looking at this picture..." / "Right, the first thing that stands out..."
Linking to the next zone"Moving on,..." / "Then, if I look further back,..." / "Next to that,..."
Adding a detail"Oh, and I can also see..." / "Another thing worth mentioning is..."
Hedging an uncertain guess"It looks as though..." / "I'm not entirely sure, but it seems..." / "It could be..."
Wrapping up"All in all,..." / "So overall,..." / "To sum up the scene,..."

Model answer

A realistic, high-scoring response to a typical picture — a busy farmers' market — timed to fill roughly 60 seconds (about 125 words).

This picture shows a busy outdoor farmers' market on what looks like a bright, sunny weekend morning. In the foreground, on the left, a young woman is paying for a basket of red apples, and the vendor is handing her some change with a smile. Next to her, I can see a wooden crate full of colourful vegetables. In the centre, two children are pointing excitedly at a fruit stand, and behind them a man wearing a green apron is arranging tomatoes into neat rows. To the right, an older couple is strolling slowly and looking at the flowers. In the background, I can see striped white tents and a few leafy trees. Overall, the scene looks lively and cheerful, and everyone seems to be enjoying the morning.

Why this answer scores high

Every sentence is doing a job. Here is the line-by-line breakdown.

  • "This picture shows a busy outdoor farmers' market on what looks like a bright, sunny weekend morning." — A complete overview in one sentence: setting, weather, mood. The listener now has a mental frame. This is pure Content & Coherence.
  • "In the foreground, on the left, a young woman is paying..." — Two position words plus a present-continuous action verb plus a specific object ("basket of red apples"). This is the core formula and it hits Vocabulary and Task Fulfilment at once.
  • "the vendor is handing her some change with a smile" — A second action and a human detail. It shows the picture is alive, not a static list.
  • "Next to her, I can see a wooden crate full of colourful vegetables." — A linking position word and a precise adjective+noun pairing.
  • "In the centre, two children are pointing excitedly..." — A clean transition to the next zone with an emotion adverb ("excitedly"), which lifts Vocabulary.
  • "behind them a man wearing a green apron is arranging tomatoes into neat rows" — Clothing detail plus a vivid, specific verb ("arranging... into neat rows") instead of a flat "is working."
  • "To the right, an older couple is strolling slowly..." — A fourth zone, keeping the sweep moving and the speech filling the full minute.
  • "In the background, I can see striped white tents and a few leafy trees." — Zooms out cleanly, exactly as planned.
  • "Overall, the scene looks lively and cheerful, and everyone seems to be enjoying the morning." — An interpretive close. It signals completeness and earns top Content & Coherence marks.

Notice there is no random jumping, no "there is... there is... there is," and no dead air — the answer simply sweeps from front to back, naming, placing, and animating each thing in turn.

What CLB 7, 9 and 11 look like

  • CLB 7 — The describer covers the main objects and uses some present continuous, but the order is a little random and starters repeat ("there is a... there is a..."). Vocabulary is functional ("a man," "some food," "happy") and there may be one or two noticeable pauses, but the listener still understands the scene.
  • CLB 9 — The tour is clearly organised foreground to background, position language is varied and accurate, and verbs are specific ("arranging," "strolling," "handing over"). The speaker fills the full 60 seconds with only natural, brief pauses and closes with a short overall impression.
  • CLB 11 — Everything at CLB 9 plus precision and texture: exact adjectives ("a striped white tent," "neat rows"), light interpretation ("it looks as though they're regulars"), smooth self-correction, and effortless, natural pacing. The description feels like a confident narration rather than a test answer.

Common mistakes to avoid

  • Jumping around the picture at random. Describing the sky, then a shoe, then a face, then the ground confuses the listener and tanks Content & Coherence. Always sweep in one direction — foreground to background, or left to right.
  • Overusing "there is / there are." Two or three are fine, but a whole answer built on them sounds flat. Rotate the starters from the bank and lead with position words instead.
  • Using the wrong tense. A scene in progress needs present continuous ("a woman is paying"), not the simple past ("a woman paid") or bare present ("a woman pays").
  • Naming objects vaguely. "A thing," "some stuff," "a person doing something" waste your best chance to show Vocabulary. Commit to a specific noun even if you have to guess — "it looks like a basket of apples" is far stronger than "some fruit."
  • Stopping early. Finishing at 40 seconds with 20 seconds of silence reads as incomplete Task Fulfilment. Keep adding small details — clothing, colours, expressions, background scenery — until the timer ends.
  • Forgetting the overview and the close. Diving straight into the first object, or ending mid-sentence, removes the frame that makes you sound organised. Always bookend with a one-line setting and a one-line impression.
Practice this task with instant AI feedback

Use this template on a real, exam-style speaking question and get a CLB estimate with specific fixes on your own answer.

Frequently asked questions

How long do I get for CELPIP Speaking Task 3?

You get 30 seconds to prepare after the picture appears, then 60 seconds to speak into the microphone. There is no live examiner — your response is recorded. Aim to keep talking for the full minute.

Who am I supposed to be describing the picture to?

You describe the scene to a listener who cannot see it — imagine you are on the phone helping a friend picture exactly what is happening. Your job is to be detailed and clear enough that they could sketch it.

What verb tense should I use in Task 3?

Mainly the present continuous (is/are + -ing), because you are describing actions happening right now in the picture, such as 'a man is arranging vegetables.' Use the simple present for unchanging facts like 'the market is busy.'

In what order should I describe the picture?

Use a systematic sweep: a one-line overview, then the foreground, then the centre and the people, then the background, ending with an overall impression. Never jump around the image at random, as that hurts your Content and Coherence score.

What if I run out of things to say before 60 seconds?

Add smaller details rather than stopping: clothing, colours, facial expressions, objects on tables, the weather, or background scenery. You can also give a light interpretation, such as how the people seem to be feeling.

Do I have to describe every single thing in the picture?

No. Cover the main people and objects clearly and in order; you cannot mention everything in 60 seconds. Raters reward organised, detailed coverage of the key elements far more than a rushed list of every object.