How Generative AI Actually Works
Generative AI creates new content β text, images, code β but how? This module demystifies tokens, transformers, training, and inference without requiring a maths degree.
How does generative AI create new content?
Generative AI is like the worldβs best autocomplete.
You know how your phone suggests the next word when you type a message? Generative AI does the same thing β but at a massive scale. Itβs been trained on billions of documents, and it predicts the most likely next word, sentence, or paragraph based on what youβve written.
It doesnβt βunderstandβ language the way you do. Itβs incredibly good at recognising patterns β so good that its outputs look like they were written by a person.
The same idea applies to images: image generation models start with random noise and progressively refine it into a clear image, guided by your text description.
The key concepts
Tokens: how AI reads text
AI models donβt read words β they read tokens. A token is a chunk of text, roughly 3-4 characters.
| Text | Tokens |
|---|---|
| βHelloβ | 1 token |
| βMicrosoft Foundryβ | 2 tokens |
| βMediSparkβs diagnostic AIβ | ~5 tokens |
| A 1-page document | ~500-700 tokens |
Why tokens matter:
- Models have a token limit (context window) β how much text they can process at once
- Pricing is based on tokens processed (input + output)
- Longer prompts = more tokens = higher cost
Training vs inference
| Feature | Training | Inference |
|---|---|---|
| When | Before the model is available | When you use the model |
| What happens | Model learns patterns from massive datasets | Model generates responses based on your input |
| Who does it | OpenAI, Microsoft, model providers | You β through the Foundry portal or SDK |
| Cost | Enormous (millions of dollars, weeks of compute) | Per-token pricing (fractions of a cent) |
| Analogy | Teaching a student for years | Asking the student a question |
Large Language Models (LLMs)
The AI models youβll use in this course are called Large Language Models (LLMs). Theyβre βlargeβ because:
- Billions of parameters β the internal numbers the model adjusts during training
- Trained on massive datasets β books, websites, code, scientific papers
- General-purpose β they can write, summarise, translate, code, and reason
Examples of LLMs:
| Model Family | Provider | Used In |
|---|---|---|
| GPT-4o, GPT-4 | OpenAI | Azure OpenAI, Microsoft Foundry |
| Phi-4 | Microsoft | Microsoft Foundry (smaller, efficient) |
| Llama | Meta | Available in Foundry model catalog |
| Mistral | Mistral AI | Available in Foundry model catalog |
What's a transformer?
The Transformer is the architecture behind modern LLMs. The key innovation is self-attention β the ability to look at every word in a sentence and understand how each word relates to every other word.
Before Transformers, AI models processed text one word at a time (left to right). Transformers can process the entire sentence at once, understanding context in both directions.
Example: In βThe bank of the river was steep,β a Transformer understands that βbankβ means riverbank (not a financial bank) because it looks at the surrounding words simultaneously.
You donβt need to understand the maths for the exam β just know that Transformers are what makes modern LLMs possible.
Grounding and hallucination
Grounding: keeping AI honest
Grounding means connecting the AIβs responses to real, verifiable information. Without grounding, AI models generate responses based purely on their training data β which may be outdated or wrong.
Ways to ground AI responses:
- System prompts β tell the model what data sources to use
- RAG (Retrieval-Augmented Generation) β feed relevant documents to the model alongside the userβs question
- Foundry IQ β Microsoftβs built-in knowledge integration for enterprise data
Hallucination: when AI makes things up
Hallucination is when an AI model generates confident-sounding but factually incorrect information. It happens because the model is predicting likely text, not looking up facts.
Priya scenario: Priya asks an AI model: βWhat year was Microsoft Foundry launched?β The model might confidently say β2023β (wrong β it was rebranded from Azure AI Studio to Foundry more recently). This is a hallucination.
Reducing hallucinations:
- Use grounding (RAG, system prompts)
- Lower the temperature setting (less creative = more predictable)
- Add instructions to say βI donβt knowβ when uncertain
Exam tip: Grounding vs hallucination
The exam loves asking about these concepts:
- Grounding = connecting AI to real data sources β reduces hallucination
- Hallucination = AI generating false information confidently
- RAG = a specific technique for grounding (feed documents to the model)
- If a question asks βhow to improve accuracy of AI responsesβ β the answer is usually grounding or RAG
Multimodal models
Modern AI models arenβt limited to text. Multimodal models can work with multiple types of input and output:
| Modality | Input Example | Output Example |
|---|---|---|
| Text | βDescribe this imageβ | Written description |
| Image | A photo of a product | Product classification |
| Audio | A recorded meeting | Transcription |
| Video | A surveillance clip | Activity summary |
GPT-4o (the βoβ stands for βomniβ) is a multimodal model β it can accept text, images, and audio as input and generate text and audio as output.
MediSpark scenario: MediSpark uses a multimodal model to analyse X-ray images. A doctor uploads the image and types βIdentify any abnormalities.β The model processes both the image and the text instruction to generate a diagnostic summary.
π¬ Video walkthrough
Flashcards
Knowledge Check
Priya is using a generative AI model to summarise a research paper. The model produces a summary that includes a statistic not found in the original paper. What is this called?
DataFlow Corp wants to reduce hallucinations in their customer support AI. Which approach is most effective?
Which of the following best describes what happens during AI model training?
Next up: Choosing the Right AI Model β not all models are equal. Learn when to use GPT-4o vs Phi-4 vs image generators.