---
title: "Tokens, context windows, training and inference"
description: "Four terms that explain most AI behavior You do not need to be an engineer to use AI well, but four concepts explain almost every \"why did it do that?\"…"
url: https://optimizeall.com/learn/ai-fundamentals-for-marketers/tokens-context-and-training
updated: 2026-10-05
---

AI Fundamentals for Marketers & Creators · How generative AI actually works · lesson 2 of 16 · 12 min

# Tokens, context windows, training and inference

## Four terms that explain most AI behavior

You do not need to be an engineer to use AI well, but four concepts explain almost every "why did it do that?" moment: **tokens**, **context window**, **training** and **inference**.

## Tokens: how AI reads and counts

Models do not read words the way we do. They break text into **tokens**, small chunks that are often a whole short word, part of a longer word, or a punctuation mark. "Marketing" might be one token; an unusual brand name might be split into several.

Why it matters to you:

- **Pricing and limits** for many AI tools are measured in tokens, not words. As a rough rule of thumb, English text uses somewhat more tokens than words. Languages written in other scripts, such as Urdu or Arabic, often use more tokens for the same meaning, so long prompts in those languages can hit limits sooner.
- **Spelling-level tasks** (counting letters, exact character limits) can trip models up because they "see" tokens, not letters. Always check character counts for ad headlines and SMS yourself or with a counter tool.

## The context window: the model's working memory

The **context window** is how much text the model can consider at once: your instructions, any files you pasted, and the conversation so far. Modern assistants have large windows, enough for long documents, but they are still finite.

Practical consequences:

- In a very long chat, early instructions can get less attention or eventually fall outside the window. If the tool starts ignoring your brand rules, start a fresh chat and re-paste your key instructions.
- The model only knows what is in the context (plus its training). If you want it to use your price list, *paste the price list* or attach it.
- More context is not always better. Irrelevant material can distract the model. Give what matters.

Some tools add **memory** or **project** features that store notes about you between chats. These are convenient, but check what is saved and whether it includes anything confidential.

## Training: how the model learned

**Training** is the expensive, one-off process where the model learns patterns from a massive dataset. After the main training, developers typically fine-tune the model to follow instructions and behave helpfully and safely.

Two consequences:

1. **Knowledge cut-off.** The model's built-in knowledge stops at some point in time. Ask about a platform policy that changed last month and a model without web access may confidently describe the old rule.
2. **Your chats are not training in real time.** The model does not learn from your conversation the moment you type. However, depending on the provider and your account settings, your conversations *may* be used to improve future models. Business and enterprise plans often exclude your data from training by default; consumer plans vary. Check the settings.

## Inference: the model at work

**Inference** is what happens every time you send a prompt: the trained model processes your input and generates a response. Each response is generated fresh, which is why the same prompt can give different answers. Many tools add a degree of randomness (sometimes called temperature) to make writing less repetitive.

For marketers, this variability is a feature: ask for ten hook options and pick the best. For facts and numbers, it is a warning: a different answer each time means you need an authoritative source.

## A worked example

Bilal, a Dubai-based real estate marketer, pastes a 90-page developer brochure into an assistant and asks for Instagram posts. Halfway through a long session the tool starts inventing amenities. Why?

- The conversation grew very long, so earlier details got less attention.
- Bilal asked about "the payment plan" without the relevant page being clearly referenced.
- The model filled the gap with plausible real-estate language.

Fix: start a fresh chat, paste only the relevant sections (amenities list, payment plan page), instruct "Use only the facts in the text below; if something is missing, say 'not in brochure'", and check every claim against the brochure.

## Hands-on: manage the context window like a pro

**1. The fresh-start summary.** When a long chat starts drifting (ignoring rules, repeating itself, mixing up products), ask:

```text
Summarize everything important from this conversation in under 150 words:
the goal, the audience, the brand rules, the decisions made, and the open tasks.
```

Paste that summary into a **new chat** (or into a project) and continue. Quality usually jumps back.

**2. Put stable rules where they persist.** Brand rules, banned claims and tone guidance belong in **project instructions** (ChatGPT or Claude Projects), a **Gem** in Gemini, or custom instructions, so they're applied to every chat instead of scrolling out of view.

**3. Give the model today's facts.** Training has a cut-off. For anything recent (platform policies, prices, trends), turn on web search or paste the source:

```text
Here is the current [platform]'s branded content policy (pasted below, retrieved today).
Answer only from this text: can we use a paid partnership label for a gifted product?
```

**Before:** in hour two of a campaign-planning chat, the assistant forgets the "no discount language" rule and describes an old platform policy from memory.

**After:** a fresh chat seeded with a 120-word summary, the brand rules stored in the project, and the current policy pasted in. The answer follows the rules and quotes the policy.

## Reasoning and cost, in one sentence each

- **Reasoning modes** use more tokens (and time) to think before answering; use them for analysis and planning, not captions.
- **Token costs and limits** apply to plans and APIs; long files, long chats and scripts that use more tokens per word (such as Urdu or Arabic) use them faster.

## Quick reference

- **Token:** the unit of text the model reads and generates.
- **Context window:** what the model can see right now.
- **Training:** how the model learned, which fixes its knowledge cut-off.
- **Inference:** the model generating an answer to your prompt, fresh each time.

## Video lecture: Tokens, context windows, training and inference

Lecture coming soon · 10 chapters · about 8 minutes. Read the full transcript below.

1. Four terms that explain AI behavior
2. Why it matters
3. Tokens
4. The context window
5. Training and inference
6. Example 1: the outdated policy
7. Example 2: the three-hour campaign chat
8. Watch me do it
9. Reasoning modes and limits
10. Recap and try this now

## Lecture transcript

### Four terms that explain AI behavior

Why does an AI assistant forget your brand rules halfway through a long chat? Why does it describe a platform policy that changed months ago? And why does it get a character count wrong for an ad headline, even though it writes beautifully? Four terms explain almost every one of those why did it do that moments: tokens, the context window, training and inference. You don't need to be an engineer to understand them. In this lecture you'll get a simple mental model for each, see two examples of them causing real problems, watch me rescue a drifting chat, and leave with three habits that fix most of these issues.

### Why it matters

Why should a marketer care about these words? Because they explain the behavior you see every day. When you understand tokens, you'll know why character limits need checking and why some languages hit limits sooner. When you understand the context window, you'll know why long chats drift and how to fix it in thirty seconds. When you understand training, you'll know why the model doesn't know last month's policy change. And when you understand inference, you'll know why the same question can get different answers. Fewer surprises. Better results on long projects. And smarter use of your plan's limits.

### Tokens

First, tokens. Models don't read words the way we do. They break text into tokens: small chunks that are often a short word, part of a longer word, or a punctuation mark. Marketing might be one token. An unusual brand name might be three. Why does that matter? Two reasons. Limits and pricing for many AI tools are measured in tokens, not words, and text in scripts like Urdu or Arabic often uses more tokens for the same meaning, so long prompts in those languages can hit limits sooner. And because models see tokens rather than letters, tasks like counting characters can trip them up. So for ad headlines and SMS, always check the count yourself or with a counter tool.

### The context window

Next, the context window. Think of it as the model's whiteboard, or working memory. Everything it can consider right now goes on that whiteboard: your instructions, any files you pasted, and the conversation so far. Modern assistants have very large whiteboards, big enough for long documents. But they're still finite. And even before the board is full, in a very long chat with lots of detours, the model can give less attention to instructions from way back at the start. That's why your brand rules seem to fade in hour two. The model isn't being lazy. The whiteboard is crowded.

### Training and inference

Now training and inference. Training is how the model learned, from a huge amount of data up to a cut off date. It's like a library that stopped acquiring books on a certain day. Anything after that date, the model simply doesn't know, unless you give it search or paste in the information. Inference is the model at work: generating an answer, right now, from its training plus whatever's on the whiteboard. One more thing about inference. There's a bit of randomness in how the model picks each next token. That's why asking the same question twice can give you two different answers, and why two different numbers for the same fact is a strong hint that it doesn't really know.

### Example 1: the outdated policy

Here's a simple example. An influencer manager asks an assistant, with search turned off, about a platform's current rules for labeling paid partnerships. The answer is confident, clear, and out of date, because the platform updated its rules after the model's training cut off. Nothing was invented, exactly. It's just old. The fix is to give the model today's facts. Turn on web search and ask for the official source, or paste the current policy text and say: answer only from this text. Now the answer quotes the current rule. Rule of thumb: for anything that changes, like policies, prices, platform features and trends, never rely on training alone.

### Example 2: the three-hour campaign chat

Now a realistic business case. Hamza, a marketing lead at an electronics retailer in Riyadh, spends an afternoon planning a back to school campaign in one enormous chat. At the start, he sets rules: no discount language, because the brand is positioned on quality, and always mention the two year warranty. By hour two, drafts start saying huge savings, and the warranty disappears. The whiteboard is crowded. Hamza asks for a one hundred and fifty word summary of the goal, audience, brand rules, decisions and open tasks, opens a fresh chat, pastes it, and continues. The next drafts follow the rules again. Then he moves the brand rules into his project instructions, so they apply to every chat from now on.

### Watch me do it

Let me show you the rescue. I'm deep into a long chat about a product launch, and I notice the latest draft uses exclamation marks everywhere, which I banned at the start, and it's calling the product by its old name. That's drift. So I type: summarize everything important from this conversation in under one hundred and fifty words: the goal, the audience, the brand rules, the decisions made and the open tasks. It gives me a tight summary. I check it, fix one detail, the product's new name, and copy it. Then I open a new chat inside my project, paste the summary, and ask for the next draft. Clean. Finally, I open the project instructions and add the two rules that drifted, so they're applied every time.

### Reasoning modes and limits

A quick word on reasoning modes and limits, since you'll see them in most assistants. Reasoning or thinking modes let the model work through a problem step by step before it answers. That uses more tokens and more time, so it's worth it for analysis, planning and tricky questions, like comparing three pricing options, but it's unnecessary for a caption or a hashtag list. And your plan's limits are usually measured in tokens. Very long files, very long chats, and text in scripts that use more tokens per word all use them faster. Match the mode to the job.

### Recap and try this now

Let's recap. Tokens are how models read and count, so check character limits yourself. The context window is the model's working memory, so when a long chat drifts, restart it with a summary, and keep stable rules in project instructions or custom instructions. Training has a cut off, so give the model today's facts with search or pasted sources. And inference has some randomness, so if the same question gets different numbers, treat that as a sign the model doesn't know. Here's your try this now. Find your longest running AI chat. Ask for the one hundred and fifty word summary, start fresh with it, and move your brand rules into a project. Compare the next draft with the last one.

## Key takeaways

- Tokens are the chunks models read and count; limits and pricing are usually measured in tokens, not words.
- The context window is the model's working memory: instructions, files and conversation, and it is finite.
- Training gives general knowledge with a cut-off; inference is the model generating an answer from your context now.
- Restart drifting chats with a summary, store stable rules in projects, and give the model today's facts.

## Try it

Take a long document you use at work, identify the two or three sections an AI would actually need for a task, and write a prompt that includes only those sections plus the instruction to use only the supplied facts.

- [Previous: What generative AI is (and is not)](https://optimizeall.com/learn/ai-fundamentals-for-marketers/what-generative-ai-is)
- [Next: Hallucinations: why AI gets things confidently wrong](https://optimizeall.com/learn/ai-fundamentals-for-marketers/why-ai-hallucinates)
- [All lessons of AI Fundamentals for Marketers & Creators](https://optimizeall.com/learn/ai-fundamentals-for-marketers)
