AI Fundamentals for Marketers & CreatorsHow generative AI actually works · Lesson 2 of 16

Tokens, context windows, training and inference

Article · 12 min · 8 min lecture

Video lecture

Tokens, context windows, training and inference

10 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 10

Four terms that explain AI behavior

  • Tokens
  • Context window
  • Training
  • Inference

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Four terms that explain most AI behavior

You do not need to be an engineer to use AI well, but four concepts explain almost every "why did it do that?" moment: tokens, context window, training and inference.

Tokens: how AI reads and counts

Models do not read words the way we do. They break text into tokens, small chunks that are often a whole short word, part of a longer word, or a punctuation mark. "Marketing" might be one token; an unusual brand name might be split into several.

Why it matters to you:

  • Pricing and limits for many AI tools are measured in tokens, not words. As a rough rule of thumb, English text uses somewhat more tokens than words. Languages written in other scripts, such as Urdu or Arabic, often use more tokens for the same meaning, so long prompts in those languages can hit limits sooner.
  • Spelling-level tasks (counting letters, exact character limits) can trip models up because they "see" tokens, not letters. Always check character counts for ad headlines and SMS yourself or with a counter tool.

The context window: the model's working memory

The context window is how much text the model can consider at once: your instructions, any files you pasted, and the conversation so far. Modern assistants have large windows, enough for long documents, but they are still finite.

Practical consequences:

  • In a very long chat, early instructions can get less attention or eventually fall outside the window. If the tool starts ignoring your brand rules, start a fresh chat and re-paste your key instructions.
  • The model only knows what is in the context (plus its training). If you want it to use your price list, paste the price list or attach it.
  • More context is not always better. Irrelevant material can distract the model. Give what matters.

Some tools add memory or project features that store notes about you between chats. These are convenient, but check what is saved and whether it includes anything confidential.

Training: how the model learned

Training is the expensive, one-off process where the model learns patterns from a massive dataset. After the main training, developers typically fine-tune the model to follow instructions and behave helpfully and safely.

Two consequences:

  1. Knowledge cut-off. The model's built-in knowledge stops at some point in time. Ask about a platform policy that changed last month and a model without web access may confidently describe the old rule.
  2. Your chats are not training in real time. The model does not learn from your conversation the moment you type. However, depending on the provider and your account settings, your conversations may be used to improve future models. Business and enterprise plans often exclude your data from training by default; consumer plans vary. Check the settings.

Inference: the model at work

Inference is what happens every time you send a prompt: the trained model processes your input and generates a response. Each response is generated fresh, which is why the same prompt can give different answers. Many tools add a degree of randomness (sometimes called temperature) to make writing less repetitive.

For marketers, this variability is a feature: ask for ten hook options and pick the best. For facts and numbers, it is a warning: a different answer each time means you need an authoritative source.

A worked example

Bilal, a Dubai-based real estate marketer, pastes a 90-page developer brochure into an assistant and asks for Instagram posts. Halfway through a long session the tool starts inventing amenities. Why?

  • The conversation grew very long, so earlier details got less attention.
  • Bilal asked about "the payment plan" without the relevant page being clearly referenced.
  • The model filled the gap with plausible real-estate language.

Fix: start a fresh chat, paste only the relevant sections (amenities list, payment plan page), instruct "Use only the facts in the text below; if something is missing, say 'not in brochure'", and check every claim against the brochure.

Hands-on: manage the context window like a pro

1. The fresh-start summary. When a long chat starts drifting (ignoring rules, repeating itself, mixing up products), ask:

Summarize everything important from this conversation in under 150 words:
the goal, the audience, the brand rules, the decisions made, and the open tasks.

Paste that summary into a new chat (or into a project) and continue. Quality usually jumps back.

2. Put stable rules where they persist. Brand rules, banned claims and tone guidance belong in project instructions (ChatGPT or Claude Projects), a Gem in Gemini, or custom instructions, so they're applied to every chat instead of scrolling out of view.

3. Give the model today's facts. Training has a cut-off. For anything recent (platform policies, prices, trends), turn on web search or paste the source:

Here is the current [platform]'s branded content policy (pasted below, retrieved today).
Answer only from this text: can we use a paid partnership label for a gifted product?

Before: in hour two of a campaign-planning chat, the assistant forgets the "no discount language" rule and describes an old platform policy from memory.

After: a fresh chat seeded with a 120-word summary, the brand rules stored in the project, and the current policy pasted in. The answer follows the rules and quotes the policy.

Reasoning and cost, in one sentence each

  • Reasoning modes use more tokens (and time) to think before answering; use them for analysis and planning, not captions.
  • Token costs and limits apply to plans and APIs; long files, long chats and scripts that use more tokens per word (such as Urdu or Arabic) use them faster.

Quick reference

  • Token: the unit of text the model reads and generates.
  • Context window: what the model can see right now.
  • Training: how the model learned, which fixes its knowledge cut-off.
  • Inference: the model generating an answer to your prompt, fresh each time.

Key takeaways

  • Tokens are the chunks models read and count; limits and pricing are usually measured in tokens, not words.
  • The context window is the model's working memory: instructions, files and conversation, and it is finite.
  • Training gives general knowledge with a cut-off; inference is the model generating an answer from your context now.
  • Restart drifting chats with a summary, store stable rules in projects, and give the model today's facts.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. An assistant starts ignoring the brand rules you gave at the start of a very long chat. What is the most likely cause and fix?
  2. Why might an AI tool describe an outdated platform policy?
  3. You need an ad headline under 30 characters. What is the safest approach?

Put it into practice

Take a long document you use at work, identify the two or three sections an AI would actually need for a task, and write a prompt that includes only those sections plus the instruction to use only the supplied facts.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.