---
title: "Prompting with files, photos, screenshots and voice"
description: "Assistants can see, read and listen Most major assistants now accept more than typed text. You can upload PDFs, spreadsheets and slides; take a photo…"
url: https://optimizeall.com/learn/prompt-engineering-foundations/prompting-with-files-photos-and-voice
updated: 2026-10-05
---

Prompt Engineering Foundations · Modern assistant features: modes, memory, research and media · lesson 9 of 16 · 12 min

# Prompting with files, photos, screenshots and voice

## Assistants can see, read and listen

Most major assistants now accept more than typed text. You can upload PDFs, spreadsheets and slides; take a photo with your phone; share a screenshot; or simply talk. The same prompting principles apply: say what the material is, what you want done with it, and what good output looks like.

## Files: PDFs, documents and spreadsheets

When you upload a file, the assistant reads its text and, in many products, looks at page images too, which helps with tables and charts. Good habits:

- **Say what the file is and why you uploaded it.** "This is our 2026 staff handbook. I need to answer employee questions about annual leave."
- **Ask for page or section references** so you can check answers.
- **Narrow the job:** "Using only section 4" is more reliable than "Summarise everything".
- **For spreadsheets,** ask the assistant to calculate using its analysis tools rather than estimate, and check a few numbers yourself.

```text
Attached is our supplier contract (PDF, 18 pages). I need to know:
1. The notice period to end the contract.
2. Any penalties for late payment.
3. Whether prices can change during the contract.
For each, quote the exact clause and give the page number. If the
contract doesn't say, write "Not stated". Keep the answer under 200 words.
```

## Photos and screenshots

A phone photo or screenshot can save a lot of typing:

- **Receipts and invoices:** "Extract the date, supplier, total and currency into a table. Write UNREADABLE for anything you can't read clearly."
- **Error messages:** "This is a screenshot of an error in our booking software. What does it mean and what should I try first?"
- **Whiteboards and handwritten notes:** "Turn this whiteboard photo into a tidy action list with owners."
- **Products and places:** "What kind of plant is this and how often should I water it indoors in Dubai's climate?"
- **Design feedback:** "Here are two versions of our flyer. Which is easier to read from two metres away, and why?"

Tips: crop to what matters, take photos flat and in good light, and label multiple images ("Image 1 is the old menu, image 2 the new one"). Before sharing, blur or crop out anything private: faces of people who have not agreed, names, addresses, card numbers.

## Voice mode

Voice mode lets you talk to the assistant and hear its reply. It is excellent for:

- practising interviews, presentations or a new language;
- brainstorming while walking or driving (hands-free and safely);
- getting quick explanations when typing is awkward.

Prompting by voice works best when you set the scene first: "I'm going to practise a sales call. You play a busy restaurant owner who is sceptical about social media ads. Keep your replies short. After five minutes, give me feedback." Some assistants can also look through your phone camera during a voice conversation and talk about what they see; check the feature and its privacy settings before using it with other people around.

## Worked example: a clinic manager in Karachi

A clinic manager photographs a handwritten weekly rota, uploads the staff leave policy PDF, and asks:

```text
Image: this week's handwritten staff rota.
File: our leave policy.
1. Type the rota into a table (Day | Morning | Evening). Mark anything
   unreadable as "?".
2. Check it against the policy: flag anyone scheduled on approved leave
   days or working more than 5 shifts.
Quote the policy rule for each issue.
```

The assistant produces the table and flags two issues. The manager confirms the two "?" entries with the staff directly and corrects one misread name. Total time: ten minutes instead of forty.

## What assistants still get wrong

- **Blurry or tiny text** may be misread or even invented. Ask for UNREADABLE instead of guesses.
- **Long documents:** details in the middle can be missed. Ask for page references and check.
- **Charts:** assistants describe trends well but can misread exact values. Use the underlying data when you have it.
- **Counting** many small objects in a photo is unreliable.

## Hands-on: three multimodal tasks

Try one task of each type this week:

```text
1. FILE: upload a policy, contract or report and ask three questions,
   requiring a quote and page number for each answer.
2. PHOTO: photograph a receipt or handwritten note and ask for a table,
   with UNREADABLE for anything unclear.
3. VOICE: use voice mode to practise a 3-minute conversation you have
   coming up, then ask for feedback on clarity and confidence.
```

## How to know you are doing it well

For files and photos: every important answer can be traced to a page or part of the image, and you checked it. For voice: you finish the session with specific feedback you can act on, not just a pleasant chat.

## Video lecture: Prompting with files, photos, screenshots and voice

Video: [Prompting with files, photos, screenshots and voice](https://www.youtube-nocookie.com/embed/2VzBqIaBE-U) — Assistants can see, read and listen Most major assistants now accept more than typed text. You can upload PDFs, spreadsheets and slides; take a photo…

12 chapters · about 8 minutes · captions and full transcript below.

1. Files, photos, screenshots and voice
2. Why multimodal prompting?
3. Prompting with files
4. Spreadsheets and long files
5. Photos and screenshots
6. Worked example: a clinic rota
7. Voice mode
8. Privacy and limits
9. Practice: a checkable file task
10. Example 1: reading tiny print
11. Example 2: monthly receipts (illustrative)
12. Recap

## Lecture transcript

### Files, photos, screenshots and voice

Your assistant can read a forty-page contract, decode a blurry photo of a receipt, and hold a spoken conversation while you walk. But only if you ask the right way. In this lecture you will learn how to prompt with files, photos, screenshots and voice, the checks that stop the AI inventing details, and how to protect privacy when you share images. By the end, you will save hours of typing and still trust the results.

### Why multimodal prompting?

Why does this matter? Because retyping is slow and error-prone, and a lot of the information you need is already sitting in a PDF, a photo or your own voice. An assistant that can read the actual document works from the real source, not your summary of it. But there is a catch. It can also misread a blurry number or miss a line on page thirty. So the skill in this lesson is two-sided: get the material in easily, and get answers out in a form you can check quickly.

### Prompting with files

Start with files. When you upload a PDF, a document or a spreadsheet, say what it is and why you shared it. This is our staff handbook, and I need to answer employee questions about annual leave. Then narrow the job. Using only section four is more reliable than summarise everything. And always ask for page or section references, so you can check. For example: what is the notice period to end this contract? Quote the exact clause and give the page number, and if the contract does not say, write not stated.

### Spreadsheets and long files

For spreadsheets, ask the assistant to calculate using its analysis tools rather than estimate. Many assistants can run calculations on the actual file. Then check a few numbers yourself. And with long documents, remember that details in the middle can be missed. Page references are your safety net: if the answer matters, open the page and read it.

### Photos and screenshots

Photos and screenshots save lots of typing. Photograph a receipt and ask for the date, supplier, total and currency in a table, writing unreadable for anything unclear. Screenshot an error message and ask what it means and what to try first. Snap a whiteboard and ask for a tidy action list with owners. Or upload two versions of a flyer and ask which is easier to read from two metres away. Crop to what matters, take photos flat and in good light, and label multiple images, like image one is the old menu, image two the new one.

### Worked example: a clinic rota

Here is a worked example. A clinic manager in Karachi photographs a handwritten weekly rota and uploads the staff leave policy. She asks the assistant to type the rota into a table, marking anything unreadable with a question mark, and then to check it against the policy, flagging anyone scheduled on approved leave or working more than five shifts, quoting the policy rule for each issue. In ten minutes she has a clean table and two flagged problems. She confirms the unclear entries with staff, and corrects one misread name. Forty minutes of work, done in ten, and checked.

### Voice mode

Voice mode lets you talk and listen. It shines for practising interviews, presentations and new languages, for brainstorming hands-free, and for quick explanations. The key is setting the scene: I am going to practise a sales call; you play a busy restaurant owner who is sceptical about social media ads; keep replies short; after five minutes give me feedback. Some assistants can also look through your camera during a voice chat. Check the privacy settings before using that with other people around.

### Privacy and limits

Images often hold more personal data than text. Faces, names on screens, addresses on parcels, card numbers. Before sharing, crop or blur anything private, especially other people's details, and follow your workplace policy. And remember the common failures: tiny or blurry text can be misread or even invented, charts are summarised well but exact values can be wrong, and counting many small objects is unreliable. Ask for unreadable markers and check what matters.

### Practice: a checkable file task

Let us practise with a file. Upload a policy, a contract or a report, and ask three questions. For each one, require a quote and a page number, and say what to do if the document does not answer it. Then open the pages it cites and check. You will usually find most answers are right, and occasionally one is attached to the wrong page or reads a table incorrectly. That is exactly why page references matter: they turn a long document from something you hope is right into something you can verify in minutes.

### Example 1: reading tiny print

A simple example to start. You photograph the back of a medicine box because the print is tiny, and ask: what are the dosage instructions for adults? Here is the key habit. Add: quote the exact text you read, and if any part is unreadable, say so. The assistant quotes the line, and you compare it with the box yourself. For health information, the AI is a reading aid, not the authority. The same pattern works for a parking sign in another language, a warranty card or a train timetable. Ask it to quote what it sees, then you check.

### Example 2: monthly receipts (illustrative)

Now a business scenario, with illustrative numbers. Nadia is office manager at a twenty-person accounting firm in Birmingham. Each month she receives about sixty supplier receipts as photos and PDFs. She uploads them in batches of ten and asks for a table: date, supplier, amount in pounds, VAT amount, category, and a readability column, with unreadable instead of guesses. Then she asks for anything over five hundred pounds to be flagged. The assistant produces the table in minutes. Nadia checks every flagged and unreadable row against the original, which is about one in eight. What used to take her most of an afternoon becomes around forty minutes, including checking. Illustrative figures, but a very common win.

### Recap

To recap. For files, say what they are, narrow the question and demand quotes with page numbers. For photos, crop, label and ask for unreadable instead of guesses. For voice, set the scene and ask for feedback. And protect privacy before you share. Try this now: this week, do one file task, one photo task and one voice practice session, and note what you had to check or correct. Next, we will try creating images and short videos with a prompt.

## Key takeaways

- Say what the file or image is, why you are sharing it and what output you need.
- Ask for page references or UNREADABLE markers so answers can be checked instead of guessed.
- Crop, label and de-identify images before sharing; blurry text and exact chart values are common failure points.
- Voice mode is ideal for practice, brainstorming and hands-free help when you set the scene first.

## Try it

This week, complete one file task (with page-referenced answers), one photo task (with UNREADABLE markers) and one voice practice session, and note what you had to check or correct.

- [Previous: Set context once: custom instructions, projects and memory](https://optimizeall.com/learn/prompt-engineering-foundations/context-once-projects-and-memory)
- [Next: Your first image and video prompts](https://optimizeall.com/learn/prompt-engineering-foundations/first-image-and-video-prompts)
- [All lessons of Prompt Engineering Foundations](https://optimizeall.com/learn/prompt-engineering-foundations)
