AI Fundamentals for Marketers & CreatorsHow generative AI actually works · Lesson 2 of 16
Tokens, context windows, training and inference
Video lecture
Tokens, context windows, training and inference
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 Four terms that explain AI behavior
Why does an AI assistant forget your brand rules halfway through a long chat? Why does it describe a platform policy that changed months ago? And why does it get a character count wrong for an ad headline, even though it writes beautifully? Four terms explain almost every one of those why did it do that moments: tokens, the context window, training and inference. You don't need to be an engineer to understand them. In this lecture you'll get a simple mental model for each, see two examples of them causing real problems, watch me rescue a drifting chat, and leave with three habits that fix most of these issues.
0:48 Why it matters
Why should a marketer care about these words? Because they explain the behavior you see every day. When you understand tokens, you'll know why character limits need checking and why some languages hit limits sooner. When you understand the context window, you'll know why long chats drift and how to fix it in thirty seconds. When you understand training, you'll know why the model doesn't know last month's policy change. And when you understand inference, you'll know why the same question can get different answers. Fewer surprises. Better results on long projects. And smarter use of your plan's limits.
1:31 Tokens
First, tokens. Models don't read words the way we do. They break text into tokens: small chunks that are often a short word, part of a longer word, or a punctuation mark. Marketing might be one token. An unusual brand name might be three. Why does that matter? Two reasons. Limits and pricing for many AI tools are measured in tokens, not words, and text in scripts like Urdu or Arabic often uses more tokens for the same meaning, so long prompts in those languages can hit limits sooner. And because models see tokens rather than letters, tasks like counting characters can trip them up. So for ad headlines and SMS, always check the count yourself or with a counter tool.
2:24 The context window
Next, the context window. Think of it as the model's whiteboard, or working memory. Everything it can consider right now goes on that whiteboard: your instructions, any files you pasted, and the conversation so far. Modern assistants have very large whiteboards, big enough for long documents. But they're still finite. And even before the board is full, in a very long chat with lots of detours, the model can give less attention to instructions from way back at the start. That's why your brand rules seem to fade in hour two. The model isn't being lazy. The whiteboard is crowded.
3:07 Training and inference
Now training and inference. Training is how the model learned, from a huge amount of data up to a cut off date. It's like a library that stopped acquiring books on a certain day. Anything after that date, the model simply doesn't know, unless you give it search or paste in the information. Inference is the model at work: generating an answer, right now, from its training plus whatever's on the whiteboard. One more thing about inference. There's a bit of randomness in how the model picks each next token. That's why asking the same question twice can give you two different answers, and why two different numbers for the same fact is a strong hint that it doesn't really know.
4:00 Example 1: the outdated policy
Here's a simple example. An influencer manager asks an assistant, with search turned off, about a platform's current rules for labeling paid partnerships. The answer is confident, clear, and out of date, because the platform updated its rules after the model's training cut off. Nothing was invented, exactly. It's just old. The fix is to give the model today's facts. Turn on web search and ask for the official source, or paste the current policy text and say: answer only from this text. Now the answer quotes the current rule. Rule of thumb: for anything that changes, like policies, prices, platform features and trends, never rely on training alone.
4:47 Example 2: the three-hour campaign chat
Now a realistic business case. Hamza, a marketing lead at an electronics retailer in Riyadh, spends an afternoon planning a back to school campaign in one enormous chat. At the start, he sets rules: no discount language, because the brand is positioned on quality, and always mention the two year warranty. By hour two, drafts start saying huge savings, and the warranty disappears. The whiteboard is crowded. Hamza asks for a one hundred and fifty word summary of the goal, audience, brand rules, decisions and open tasks, opens a fresh chat, pastes it, and continues. The next drafts follow the rules again. Then he moves the brand rules into his project instructions, so they apply to every chat from now on.
5:40 Watch me do it
Let me show you the rescue. I'm deep into a long chat about a product launch, and I notice the latest draft uses exclamation marks everywhere, which I banned at the start, and it's calling the product by its old name. That's drift. So I type: summarize everything important from this conversation in under one hundred and fifty words: the goal, the audience, the brand rules, the decisions made and the open tasks. It gives me a tight summary. I check it, fix one detail, the product's new name, and copy it. Then I open a new chat inside my project, paste the summary, and ask for the next draft. Clean. Finally, I open the project instructions and add the two rules that drifted, so they're applied every time.
6:36 Reasoning modes and limits
A quick word on reasoning modes and limits, since you'll see them in most assistants. Reasoning or thinking modes let the model work through a problem step by step before it answers. That uses more tokens and more time, so it's worth it for analysis, planning and tricky questions, like comparing three pricing options, but it's unnecessary for a caption or a hashtag list. And your plan's limits are usually measured in tokens. Very long files, very long chats, and text in scripts that use more tokens per word all use them faster. Match the mode to the job.
7:19 Recap and try this now
Let's recap. Tokens are how models read and count, so check character limits yourself. The context window is the model's working memory, so when a long chat drifts, restart it with a summary, and keep stable rules in project instructions or custom instructions. Training has a cut off, so give the model today's facts with search or pasted sources. And inference has some randomness, so if the same question gets different numbers, treat that as a sign the model doesn't know. Here's your try this now. Find your longest running AI chat. Ask for the one hundred and fifty word summary, start fresh with it, and move your brand rules into a project. Compare the next draft with the last one.
Four terms that explain most AI behavior
You do not need to be an engineer to use AI well, but four concepts explain almost every "why did it do that?" moment: tokens, context window, training and inference.
Tokens: how AI reads and counts
Models do not read words the way we do. They break text into tokens, small chunks that are often a whole short word, part of a longer word, or a punctuation mark. "Marketing" might be one token; an unusual brand name might be split into several.
Why it matters to you:
- Pricing and limits for many AI tools are measured in tokens, not words. As a rough rule of thumb, English text uses somewhat more tokens than words. Languages written in other scripts, such as Urdu or Arabic, often use more tokens for the same meaning, so long prompts in those languages can hit limits sooner.
- Spelling-level tasks (counting letters, exact character limits) can trip models up because they "see" tokens, not letters. Always check character counts for ad headlines and SMS yourself or with a counter tool.
The context window: the model's working memory
The context window is how much text the model can consider at once: your instructions, any files you pasted, and the conversation so far. Modern assistants have large windows, enough for long documents, but they are still finite.
Practical consequences:
- In a very long chat, early instructions can get less attention or eventually fall outside the window. If the tool starts ignoring your brand rules, start a fresh chat and re-paste your key instructions.
- The model only knows what is in the context (plus its training). If you want it to use your price list, paste the price list or attach it.
- More context is not always better. Irrelevant material can distract the model. Give what matters.
Some tools add memory or project features that store notes about you between chats. These are convenient, but check what is saved and whether it includes anything confidential.
Training: how the model learned
Training is the expensive, one-off process where the model learns patterns from a massive dataset. After the main training, developers typically fine-tune the model to follow instructions and behave helpfully and safely.
Two consequences:
- Knowledge cut-off. The model's built-in knowledge stops at some point in time. Ask about a platform policy that changed last month and a model without web access may confidently describe the old rule.
- Your chats are not training in real time. The model does not learn from your conversation the moment you type. However, depending on the provider and your account settings, your conversations may be used to improve future models. Business and enterprise plans often exclude your data from training by default; consumer plans vary. Check the settings.
Inference: the model at work
Inference is what happens every time you send a prompt: the trained model processes your input and generates a response. Each response is generated fresh, which is why the same prompt can give different answers. Many tools add a degree of randomness (sometimes called temperature) to make writing less repetitive.
For marketers, this variability is a feature: ask for ten hook options and pick the best. For facts and numbers, it is a warning: a different answer each time means you need an authoritative source.
A worked example
Bilal, a Dubai-based real estate marketer, pastes a 90-page developer brochure into an assistant and asks for Instagram posts. Halfway through a long session the tool starts inventing amenities. Why?
- The conversation grew very long, so earlier details got less attention.
- Bilal asked about "the payment plan" without the relevant page being clearly referenced.
- The model filled the gap with plausible real-estate language.
Fix: start a fresh chat, paste only the relevant sections (amenities list, payment plan page), instruct "Use only the facts in the text below; if something is missing, say 'not in brochure'", and check every claim against the brochure.
Hands-on: manage the context window like a pro
1. The fresh-start summary. When a long chat starts drifting (ignoring rules, repeating itself, mixing up products), ask:
Summarize everything important from this conversation in under 150 words:
the goal, the audience, the brand rules, the decisions made, and the open tasks.Paste that summary into a new chat (or into a project) and continue. Quality usually jumps back.
2. Put stable rules where they persist. Brand rules, banned claims and tone guidance belong in project instructions (ChatGPT or Claude Projects), a Gem in Gemini, or custom instructions, so they're applied to every chat instead of scrolling out of view.
3. Give the model today's facts. Training has a cut-off. For anything recent (platform policies, prices, trends), turn on web search or paste the source:
Here is the current [platform]'s branded content policy (pasted below, retrieved today).
Answer only from this text: can we use a paid partnership label for a gifted product?Before: in hour two of a campaign-planning chat, the assistant forgets the "no discount language" rule and describes an old platform policy from memory.
After: a fresh chat seeded with a 120-word summary, the brand rules stored in the project, and the current policy pasted in. The answer follows the rules and quotes the policy.
Reasoning and cost, in one sentence each
- Reasoning modes use more tokens (and time) to think before answering; use them for analysis and planning, not captions.
- Token costs and limits apply to plans and APIs; long files, long chats and scripts that use more tokens per word (such as Urdu or Arabic) use them faster.
Quick reference
- Token: the unit of text the model reads and generates.
- Context window: what the model can see right now.
- Training: how the model learned, which fixes its knowledge cut-off.
- Inference: the model generating an answer to your prompt, fresh each time.
Key takeaways
- Tokens are the chunks models read and count; limits and pricing are usually measured in tokens, not words.
- The context window is the model's working memory: instructions, files and conversation, and it is finite.
- Training gives general knowledge with a cut-off; inference is the model generating an answer from your context now.
- Restart drifting chats with a summary, store stable rules in projects, and give the model today's facts.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Take a long document you use at work, identify the two or three sections an AI would actually need for a task, and write a prompt that includes only those sections plus the instruction to use only the supplied facts.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.