Mastering Claude (Anthropic) · Documents, vision and research · lesson 4 of 20 · 15 min
Long documents and file analysis
What Claude can read
Claude reads PDFs (including their images, charts and tables), Word documents, spreadsheets and CSVs, plain text, code, and more. With code execution (available on Free and paid plans), Claude can also run Python on data files to calculate, clean and chart, rather than estimating in its head. Current models have very large context windows, so a 200-page tender or a year of meeting notes fits in a single conversation. Limits on file size and count vary by plan and change over time; check the help centre if an upload fails.
Large context is a capability, not a guarantee. Models can still miss a clause on page 143 or blend two similar sections. The workflow below is designed around that reality.
The Orient → Quote → Extract → Challenge workflow
1. Orient. Before detailed questions, ask for the map:
Give me the structure of the attached document: sections with page numbers,
the purpose of each section in one line, and anything that looks like an
appendix, schedule or definitions section I should know about.
2. Quote. For any answer you will rely on, require evidence:
Answer using only the attached document. After each point, quote the exact
supporting sentence and give the section or page. If the document does not
answer the question, say "not stated" rather than inferring.
3. Extract. Turn prose into tables you can check and reuse:
Extract every deliverable, deadline, owner, payment milestone and penalty into
a table: Item | Type | Date/Trigger | Owner | Amount | Section ref | Ambiguity?
Mark anything ambiguous and explain why in a notes column.
4. Challenge. Ask Claude to attack its own reading:
Now act as the other party's lawyer. Which three clauses would you exploit, and
which of my extracted items might be misread? Point to the exact wording.
Comparing versions and multiple documents
Claude is excellent at "what changed?" work. Upload both versions and ask for a change log grouped by material (price, scope, liability, dates) versus cosmetic changes. For several documents, name them in your prompt ("Doc A = 2025 contract, Doc B = 2026 draft") so answers can cite which file each point came from.
Spreadsheets and data files
For CSV and Excel files, ask Claude to use code execution and show its method:
Load the attached CSV. First report: row count, column names and types,
missing values, duplicates and obvious outliers. Don't analyse yet.
Then calculate monthly revenue by channel, show the code you ran, and
produce a table plus one chart. State any assumptions.
Spot-check two numbers by hand or in your spreadsheet. If a result looks odd, ask to see the rows that produced it.
If you live in Excel, Claude's Microsoft 365 integration (Claude for Excel, available on paid plans) lets you work inside the workbook itself; the same verification discipline applies.
Worked example: tender response for an agency in Riyadh
A marketing agency receives a 96-page government tender in English with an Arabic annex. The bid manager:
- Uploads both files and asks for the structure, including evaluation criteria and weightings.
- Extracts every mandatory requirement into a compliance matrix with section references, flagging any that differ between the English and Arabic versions.
- Asks Claude to list deadlines and submission formats as a checklist.
- Challenges: "Which requirements are we most likely to miss?" Claude points out a local-content certificate required in the annex but not the main body.
- Assigns each row to a team member and verifies every "mandatory" row against the original pages before submission.
Time to a complete compliance matrix fell from roughly two days of manual reading to an afternoon of checking (illustrative, from a typical workflow; your mileage depends on document quality).
Hands-on: build a reusable document-analysis prompt set
- Save the four prompts above in your prompt library under "Document analysis".
- Upload a long, non-confidential document (a public annual report or government guidance works well).
- Run all four stages and record one thing Claude found that you would have missed and one thing it got wrong or overstated.
- Add a line to your prompts that fixes the error pattern you saw (for example, "treat footnotes as part of the section they belong to").
Pitfalls
- Scanned PDFs: poor scans cause misread numbers. Check figures against the page image.
- Summaries that flatten nuance: "may" becomes "will". Ask for exact wording on obligations.
- Mixing versions: keep one current version per file name and say which is authoritative.
- Confidential material in personal accounts: use your organisation's approved workspace and share only the sections needed.
How to measure success
Every figure you rely on has a quote and reference; your extraction tables survive a manual spot-check of at least five rows; and you catch errors in the Challenge step before a client or counterparty does.
Video lecture: Long documents and file analysis
Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.
- Long documents without the headache
- Why a workflow, not a question
- What Claude can read
- Step 1: Orient
- Step 2: Quote
- Step 3: Extract
- Step 4: Challenge
- Simple example
- Worked example: a tender in Riyadh
- Comparing versions
- Pitfalls to watch
- Try this now
- Watch me do it, part 1
- Watch me do it, part 2
- Recap and next step
Lecture transcript
Long documents without the headache
A ninety six page tender lands in your inbox on Thursday afternoon, due Monday. You could read it all. Or you could use Claude to read it with you, in a way you can actually trust. In this lecture you will learn a four step workflow, orient, quote, extract and challenge, that turns long documents into verified tables and catches the clause everyone else misses.
Why a workflow, not a question
Why do we need a workflow at all? Why not just ask, what does this contract say? Because in long documents, the expensive detail is usually small. A penalty clause in an annex. A definition that changes the meaning of a deadline. A single-sentence exclusion. A general question gets a general answer that sounds complete and quietly skips exactly those details. The workflow forces the details into the open and attaches a reference to every one, so when someone later asks, where does it say that, you can answer in five seconds.
What Claude can read
Claude reads PDFs, including their images, tables and charts, plus Word documents, spreadsheets, CSV files, text and code. With code execution switched on, it can run real calculations on data files instead of estimating in its head. And the newest models have enormous context windows, so a two hundred page contract fits comfortably. But here is the catch. Fitting is not the same as perfect recall. Models can still miss something buried deep, or blend two similar sections. So we build a workflow around that.
Step 1: Orient
Step one is orient. Before any detailed question, ask Claude for the map. The sections with page numbers, the purpose of each in one line, and anything that looks like an appendix, schedule or definitions section. Why? Because the nasty surprises in contracts and tenders almost always live in the annex or the definitions, and you want to know they exist before you start asking questions.
Step 2: Quote
Step two is quote. For anything you will rely on, instruct Claude to answer only from the document, to quote the exact supporting sentence after each point with its page or section, and to say not stated when the document is silent. This turns every answer into something you can verify in seconds, and it dramatically reduces confident guesses.
Step 3: Extract
Step three is extract. Turn prose into tables. Every deliverable, deadline, owner, payment milestone and penalty, with a section reference and an ambiguity flag. Tables are easier to check, easier to assign to people, and easy to paste into a tracker. For spreadsheets, ask Claude to first report missing values, duplicates and outliers, then run the analysis with code and show you the method.
Step 4: Challenge
Step four is challenge. Ask Claude to attack its own reading. Act as the other party's lawyer, which three clauses would you exploit, and which of my extracted items might be misread? This adversarial pass is where Claude often surfaces the ambiguous wording or the hidden requirement that a single read skims over.
Simple example
Let's try a simple example first. You are renting a small studio for your business and the landlord sends a twelve page agreement. Ask Claude to extract the rent, deposit, notice period, break clause and any fees into a table, with the section number for each, and to flag anything ambiguous. The table comes back in seconds, and one row is flagged. The break clause refers to twelve months in section four and eighteen months in schedule two. That is a question for the landlord before you sign, and you found it in two minutes instead of an evening of reading.
Worked example: a tender in Riyadh
Here is how that played out for a marketing agency in Riyadh bidding on a government tender. Ninety six pages in English plus an Arabic annex. The bid manager asked for the structure and scoring weights, extracted every mandatory requirement into a compliance matrix with references, and flagged differences between the two languages. In the challenge step, Claude pointed out a local content certificate that appeared only in the annex. The team then verified every mandatory row against the original pages before submitting. Claude did the heavy reading. Humans did the checking. The following month the same agency received a much shorter request for proposal. The bid manager reused the four saved prompts and finished the compliance matrix in under an hour. The workflow, not the specific document, was the asset.
Comparing versions
Claude is exceptionally good at one job people dread, which is working out what changed between two versions. Upload both and ask for a change log grouped into material changes, like price, scope, liability and dates, and cosmetic ones, like formatting and wording. When you work with several documents, name them in your prompt. Document A is the twenty twenty five contract, document B is the twenty twenty six draft. Now every answer can cite which file a point came from, and you avoid the classic error of mixing old and new terms.
Pitfalls to watch
Four pitfalls. Scanned PDFs can produce misread numbers, so check figures against the page image. Summaries can flatten nuance, turning may into will, so ask for exact wording on anything that is an obligation. Mixed versions cause confusion, so keep one current version per file name and say which is authoritative. And confidential contracts belong in your organisation's approved workspace, with only the sections you need. Success looks like this. Every figure you rely on has a quote and a reference, and your extraction table survives a manual spot check.
Try this now
Try this now. Download a long public document, a company annual report or a government guidance PDF works well. Upload it and run the four steps in order. Ask for the structure with page numbers. Ask one question that requires quotes and references. Ask for an extraction table of dates, amounts or obligations with section references. Then ask Claude to challenge its own reading. Finally, pick five rows from the table and check them against the pages yourself. Note how many were perfect, and whether anything was missed or softened. That result tells you how much checking your real documents will need.
Watch me do it, part 1
Let me run the four steps on a real style tender. I upload the ninety six page English document and the Arabic annex, and I tell Claude which is which. First, orient. I paste the structure prompt and get a map of sections with page numbers, including evaluation criteria on page thirty one and the annex schedules. Second, quote. I ask, what is the submission deadline and format, answer only from the documents, quote the exact sentence and page, and say not stated if it is missing. Claude quotes the sentence from page twelve, and adds that the annex specifies a separate deadline for the local content certificate. Already, that is something I would have missed.
Watch me do it, part 2
Third, extract. I paste the table prompt, and Claude lists every mandatory requirement, deliverable and deadline with section references, flagging two as ambiguous because the English and Arabic wording differ. Fourth, challenge. I ask it to act as the evaluator and name the three requirements bidders most often miss. It points to the certificate, the bank guarantee format and a page limit on the technical proposal. Now my part. I pick five rows at random and open the PDF at each page. Four are perfect. One cites page forty four when the clause is on forty five, so I correct it. Five minutes of checking, and I trust the matrix enough to assign it to the team.
Recap and next step
So, orient to find the map, quote to make answers verifiable, extract to get tables you can check, and challenge to find what you missed. Watch out for poor scans, softened obligations and mixed versions, and keep confidential files in approved workspaces. Your next step: save the four prompts from the lesson, run them on a long public document, and spot check at least five rows by hand.
Key takeaways
- Orient first: ask for structure, key points and gaps before detailed questions.
- Require quotes and section references so every answer can be verified quickly.
- Convert prose into tables for deliverables, deadlines, changes and risks.
- Share only what is needed and use approved accounts for confidential files.
Try it
Upload a long, non-confidential document (for example, a public annual report). Use the orient, quote, extract workflow and record one thing Claude found that you would have missed.