Neurocourse

Voice and images: speak and show

ChatGPT does more than read text: you can talk to it by voice like a person and work with images — both examining your photos and drawing new ones from a description. We cover both modes and when each saves time.

Up to this lesson the whole conversation has run in text: you type, it answers in letters. But ChatGPT has two more channels, voice and images, and they're what turns a clever chat window into an assistant that's always at hand. The apps lesson showed where those buttons live — now let's see what they can actually do.

Voice mode: not dictation, a conversation

We switched voice on back in the apps lesson. Now let's go deeper, because two things get mixed up here. Dictation (the microphone icon in the input box) just turns your speech into text — you talk, letters appear, and from there it's an ordinary chat. Voice mode (the headphones/waveform icon) is a full conversation: you speak, the AI answers out loud in a real voice, and you can cut it off mid-sentence — it stops and switches. The difference is texting versus phoning.

Why voice changes the habit

Typing out a detailed request is work. Saying it is natural. In half a minute of talking you lay out considerably more context you'd have typed, and remember the rule from the first-dialogue lesson: more context, sharper answer. It also frees your hands and eyes: at the stove, pushing a pram, standing in a queue, the AI becomes someone you talk to rather than an app you stare at.

What a good voice session looks like

Voice isn't better for everything. Here are the situations where it beats the keyboard by a distance.

  • Rehearsing out loud. "Let's do a mock interview. You're a tough HR manager, ask one question at a time, I'll answer out loud, and review my answers at the end." Live speaking practice you'll never get from text.
  • Brainstorming on the move. You walk and think out loud, the AI picks it up and suggests turns. Ideas get caught while they're fresh.
  • Working something out on the go. "Explain the difference between leasing and a loan, in plain words" — on the train, no screen involved.

Before you read on: when tomorrow will your hands be busy but your head free?

That's where voice mode turns into a habit — cooking, queueing, running, riding the bus.

Images, part one: the AI looks at your photo

Drop an image into the chat (the paperclip, or the camera in the app) and ask about it. This isn't a party trick, it's a working tool:

  • Photograph a baffling error on your screen — "what does this mean and how do I fix it?"
  • Snap a dish in a restaurant with a foreign menu — "what is this, and does it have nuts in it?"
  • Show it a chart from a report — "describe the trend I'm looking at".
  • Photograph the inside of your fridge — "what can I cook for dinner out of this?"

The model reads what's in the picture and answers in words. Screenshots, photos, diagrams, handwritten notes — all fair game.

Images, part two: the AI draws from a description

The flip side is generation: describe it in words and ChatGPT draws the image from nothing. "Draw a logo for a coffee shop: minimalist, a cup and mountains, warm tones" — a picture in seconds. It comes up more often than you'd think: an illustration for a post, a rough logo, a birthday card, a diagram for a presentation, a reference for your designer. Just like with text, iteration (going round again on the answer, from the first-dialogue lesson) is what works: "brighter", "lose the background", "try a different style".

Where images hit their limits

Two honest limits, so you aren't disappointed. First: generated pictures are illustrations, not photographs of real events, and small details (text inside the image and fingers especially) still come out imperfect — check before you publish. Second: when it's reading a photo, the same rule applies as with text back in the 10-jobs lesson — verify the facts. "Identify this mushroom, is it edible" is a bad idea: the price of being wrong is your health, and the model can be wrong. With food, medicine and legal documents the AI assists, but a professional decides.

Common questions about voice and images

Three questions almost everyone trips on.

  • "Can people around me hear it?" Through earphones, no — only you. Through the speaker, yes: everyone nearby hears it, same as a call on speakerphone.
  • "Can I speak with an accent, or badly?" Yes, the model handles real speech, pauses and slips just fine — talk however you talk.
  • "Can I download a generated image and use it?" Download, yes, via the save button; but before publishing it for a business, check the service's terms and remember those imperfect small details.
  • "Why didn't the AI spot something in my photo?" Usually quality: dark, blurry or angled shots read badly — retake it square-on in good light.

Do this now

Two minutes, two taps. One: turn on voice mode and ask out loud for an explanation of anything you're curious about — just listen to what that feels like. Two: photograph something near you (a label, a receipt, a page of a book) and ask about it. You've just used all three channels: text, voice, image.

Practice · 5 tasks

Short questions on the lesson — with an explanation for every answer.