---
title: "File uploads and data analysis — Mastering ChatGPT (OpenAI)"
description: "What ChatGPT does with your files You can upload PDFs, Word documents, slides, spreadsheets, CSVs, images and more, or pull files from connected apps…"
url: https://optimizeall.com/learn/mastering-chatgpt/file-uploads-and-data-analysis
updated: 2026-10-05
---

Mastering ChatGPT (OpenAI) · Files, data, images and voice · lesson 6 of 19 · 16 min

# File uploads and data analysis

## What ChatGPT does with your files

You can upload PDFs, Word documents, slides, spreadsheets, CSVs, images and more, or pull files from connected apps such as Google Drive, OneDrive or SharePoint (where your plan and admin allow). For documents, ChatGPT reads and reasons over the text and visuals. For data files, it can **write and run Python code** in a secure sandbox to clean, calculate, chart and export results. That matters: numbers computed with code are far more reliable than numbers "estimated" in prose.

File size and count limits depend on plan and change over time; if an upload fails, check the help centre.

## The six-step data workflow

**1. Describe the data before analysing it.**

```text
Load the attached CSV. Before any analysis, report: number of rows, each column
with its data type and an example value, missing values per column, duplicate
rows, and anything that looks like an outlier or a data-entry error.
Don't analyse yet.
```

**2. Clean deliberately.** Decide how to handle duplicates, blanks and odd values yourself: "Remove exact duplicate rows. Treat blank 'Region' as 'Unknown'. Exclude test orders where email contains '@test'." Ask ChatGPT to report how many rows each rule affected.

**3. Define metrics explicitly.** "Engagement rate = (likes + comments + shares + saves) / impressions." "Revenue = gross sales minus refunds, excluding VAT." Ambiguous metrics produce confident, wrong comparisons.

**4. Analyse with the question in mind.** "Which content format had the highest median engagement rate in Q3, by platform? Show the table and the code."

**5. Visualise simply.** "One bar chart per platform, sorted descending, labelled axes, no 3D." Ask for an exportable file (PNG, XLSX, CSV) when you need to share it.

**6. Verify.** Spot-check two numbers yourself in Excel or Sheets. Ask: "Show me the rows behind the top result." Ask about sample size: "How many posts are in each group?"

## Reading results critically

- **Small samples:** "Stories performed best" based on three posts is an anecdote, not a trend.
- **Correlation vs causation:** Tuesday posts may perform better because of what was posted on Tuesdays, not the day itself.
- **Survivorship:** analysing only campaigns that ran to completion hides the ones cancelled early.
- **Units and currencies:** mixed AED/SAR/GBP columns must be converted with a stated rate and date.

Ask ChatGPT directly: *"What are the three biggest reasons this conclusion might be wrong?"*

## Documents: summarise, extract, compare

For long documents, use orient → quote → extract:

```text
1. Give me the document's structure with page numbers.
2. Answer my questions only from the document, quoting the supporting sentence
   and page for each point; say "not stated" if it isn't there.
3. Extract every deadline, deliverable, owner and payment term into a table
   with page references.
```

To compare two versions of a contract or proposal, upload both, label them, and ask for material changes (price, scope, liability, dates) separately from cosmetic ones.

## Worked example: social performance review for a UAE brand

A social media manager exports six months of Instagram and TikTok analytics. ChatGPT reports 14 duplicate rows and a column where some impressions were recorded as text. She sets cleaning rules, defines engagement rate, and asks for median (not mean) engagement by format and platform. The analysis shows carousels outperform reels on Instagram for her account. She asks for the post count behind each figure (carousels: 22 posts, reels: 41) and the code, spot-checks two values in Sheets, and presents the finding as "a pattern worth testing", with a four-week experiment plan, rather than a law.

## Privacy first

- Remove names, emails, phone numbers and order IDs you do not need before uploading.
- Use your organisation's approved workspace (Business or Enterprise) for company data.
- Delete working files from chats when finished, where appropriate.

## Hands-on

Export a real, non-sensitive dataset (for example your own post analytics or a public dataset). Run all six steps, spot-check two numbers manually, and save the prompts as a reusable "data analysis" recipe in your prompt library or a Project.

## Pitfalls

- Asking for conclusions before inspecting data quality.
- Undefined metrics and mixed currencies.
- Charts that look authoritative but rest on a handful of rows.
- Uploading full customer exports when three columns would do.

## How to measure success

Every figure you share has been spot-checked, sample sizes are stated alongside findings, and your recommendations are framed as tests when the evidence is thin.

## Video lecture: File uploads and data analysis

Lecture coming soon · 16 chapters · about 9 minutes. Read the full transcript below.

1. Data analysis you can trust
2. Why data needs a method
3. Why code matters
4. Step 1: describe first
5. Steps 2 and 3
6. Steps 4 and 5
7. Step 6: verify
8. Read results critically
9. Simple example
10. Worked example: a UAE brand’s social review
11. Privacy first
12. Try this now
13. Common mistakes
14. Watch me do it, part 1
15. Watch me do it, part 2
16. Recap and next step

## Lecture transcript

### Data analysis you can trust

Uploading a spreadsheet to ChatGPT and asking what does this tell me feels like magic. It is also one of the easiest ways to present a confident, wrong number to your boss. In this lecture you will learn a six step workflow that turns ChatGPT into a genuinely reliable analyst, and the questions that expose weak conclusions before anyone else does.

### Why data needs a method

Why do we need a method for data? Because a chart always looks authoritative, even when the data behind it is a mess. Duplicate rows, text stored as numbers, a metric defined three different ways. Those problems do not announce themselves. They quietly change the answer. Think of it like cooking. The dish can look beautiful and still be made with the wrong ingredient. The six step method makes sure you check the ingredients before you serve the result.

### Why code matters

ChatGPT handles two kinds of files differently. Documents it reads and reasons over. Data files, like spreadsheets and CSVs, it can analyse by writing and running Python code in a secure sandbox. That is the key. Numbers computed with code are far more reliable than numbers estimated in a paragraph. You can also pull files from connected apps like Google Drive or SharePoint, where your plan and admin allow.

### Step 1: describe first

Step one, describe the data before analysing it. Ask for the number of rows, each column with its type and an example, missing values, duplicate rows and anything that looks like an outlier or data entry error. And say, do not analyse yet. This single step catches most of the problems that ruin an analysis.

### Steps 2 and 3

Step two, clean deliberately. You decide the rules. Remove exact duplicates. Treat blank regions as unknown. Exclude test orders. And ask how many rows each rule affected. Step three, define your metrics. Engagement rate equals likes, comments, shares and saves, divided by impressions. Revenue is gross sales minus refunds, excluding VAT. Ambiguous metrics produce confident, wrong comparisons.

### Steps 4 and 5

Step four, analyse with a specific question. Which content format had the highest median engagement rate in the third quarter, by platform? Show the table and the code. Median is often more honest than the average when a few viral posts skew everything. Step five, visualise simply. Sorted bars, labelled axes, no three D effects, and ask for an exportable file when you need to share it.

### Step 6: verify

Step six, verify. Spot check two numbers yourself in Excel or Sheets. Ask to see the rows behind the top result. And always ask how many items are in each group. A chart that says Stories performed best looks very different once you learn it is based on three posts.

### Read results critically

Read every result critically. Small samples are anecdotes, not trends. Tuesday posts may perform better because of what you posted on Tuesdays, not the day itself. Analysing only completed campaigns hides the ones cancelled early. And mixed currency columns, dirhams, riyals and pounds, must be converted with a stated rate and date. A great habit is to ask, what are the three biggest reasons this conclusion might be wrong?

### Simple example

A simple example. Upload a small sales file with fifty rows and ask, before any analysis, to describe the data. ChatGPT reports two duplicate rows and one blank region. You tell it to remove the duplicates and label the blank as unknown. Then ask for total sales by month, with the code. Pick one month and add it up yourself in a spreadsheet. It matches. You now trust the result, and you know exactly why. That is the whole method in miniature.

### Worked example: a UAE brand’s social review

A social media manager in the UAE exports six months of Instagram and TikTok analytics. ChatGPT finds fourteen duplicate rows and a column where some impressions were stored as text. She sets cleaning rules, defines engagement rate and asks for the median by format and platform. Carousels beat reels on Instagram for her account. She checks the counts, twenty two carousels and forty one reels, spot checks two values in Sheets, and presents it as a pattern worth testing, with a four week experiment plan, not as a law. Four weeks later, she reruns exactly the same cleaned workflow on the new export. Because the cleaning rules and the metric definition were saved as a recipe, the comparison is like for like. The carousel lead holds on Instagram but shrinks, and on TikTok nothing changes. Her report says so plainly, with the post counts beside each figure, and her manager approves a small shift in the content mix rather than a dramatic one.

### Privacy first

Before uploading, remove names, emails, phone numbers and order numbers you do not need. Use your organisation's approved workspace for company data, not a personal account. And delete working files from chats when you are finished, where appropriate. If a colleague later asks for the source data, you will be glad you kept only what the analysis needed.

### Try this now

Try this now. Export a real dataset that contains no personal data, your own social analytics, website traffic by page, or a public dataset. Upload it and work through all six steps, describe, clean, define metrics, analyse against one clear question, visualise simply, and verify. Spot check two numbers yourself in a spreadsheet, and ask how many rows sit behind the top result. Save the prompts as a data analysis recipe in your prompt library or a Project, so next month's analysis takes half the time.

### Common mistakes

Let's name the common mistakes, because you will see them in other people's reports. Asking for conclusions before checking data quality. Using the average when one viral post or one huge order distorts everything, so ask for the median too. Mixing currencies without a stated conversion rate and date. And presenting a result based on a handful of rows as if it were a trend. If you avoid those four, your analysis will already be better than most.

### Watch me do it, part 1

Let me run the UAE social review. I upload the six month export and paste the describe first prompt. Rows, columns with types and examples, missing values, duplicates and outliers, and do not analyse yet. ChatGPT runs code and reports fourteen duplicate rows and an impressions column where some values are stored as text with commas. I set the cleaning rules. Remove exact duplicates, convert impressions to numbers, and exclude two test posts published from our staging account. I ask how many rows each rule affected, and it reports fourteen, thirty one and two. Now I trust the table I am about to analyse.

### Watch me do it, part 2

Next, the metric. Engagement rate equals likes, comments, shares and saves, divided by impressions. Then the question. Median engagement rate by format and platform, with the number of posts in each group, and show the code. On Instagram, carousels lead, from twenty two posts, against reels from forty one. I ask for a simple sorted bar chart. Then my check. I open the cleaned file in Google Sheets, filter Instagram carousels, and calculate the median myself. It matches. I check one reel figure too. Both match, so the finding goes into my report, framed as a pattern to test over four weeks.

### Recap and next step

Recap. Describe the data first, clean deliberately, define metrics, analyse against a clear question, visualise simply and verify. Question every conclusion and protect personal data. Your next step: run the six steps on a real, non sensitive dataset, spot check two numbers, and save the prompts as your reusable data analysis recipe.

## Key takeaways

- ChatGPT can compute with code on uploaded data, making numbers more reliable than mental estimates.
- Describe columns, clean deliberately, define metrics, then analyse and visualise.
- Verify by reviewing steps and spot-checking numbers; watch sample sizes and causation claims.
- Remove unnecessary personal data before uploading exports.

## Try it

Export a real, non-sensitive dataset (for example, your own post analytics). Run the six-step workflow and spot-check two numbers manually.

- [Previous: Projects: organise recurring work](https://optimizeall.com/learn/mastering-chatgpt/projects-in-chatgpt)
- [Next: Image generation and vision](https://optimizeall.com/learn/mastering-chatgpt/image-generation-and-vision)
- [All lessons of Mastering ChatGPT (OpenAI)](https://optimizeall.com/learn/mastering-chatgpt)
