Transformers and Generative AI
Transformer fundamentals
Transformers
and Generative AI
Text · Code · Image · Video
One architecture helps machines learn relationships across four creative media—but each medium generates its output differently.
The central question
What changes—and what stays the same?
A single prompt can ask an AI system to explain plant growth, write a simulation, visualise the process, or animate the explanation. The outputs look completely different. Underneath, however, they share learned representations and relationships.
Text
Explain plant growth.
Code
Build a simulation.
Image
Visualise the process.
Video
Animate the explanation.
Foundation
Generative AI learns patterns, then samples from them
A discriminative model chooses a label. A generative model produces new content that is plausible according to patterns learned during training.
Discriminative
Input → label
“Is this email spam?”
The output is a category, score or decision.Generative
Prompt → new content
“Write a safer reply.”
The output is constructed through an iterative prediction process.A generative model does not simply retrieve one stored answer. For text, it estimates likely continuations. For diffusion-based images, it estimates how to remove noise step by step. The model samples from these learned possibilities to produce an output.
Why transformers mattered
Earlier recurrent systems processed sequences mainly one position after another. Transformer attention allows many token positions to be processed in parallel during training and makes long-range relationships easier to model. This combination helped models scale to large datasets and much larger capacities.
Interactive comparison: six token positions
Press Replay to compare the three processes.
Hardware acceleration
How GPUs make transformer parallelism practical
The transformer provides parallelisable operations. A GPU supplies many processing units that can execute large groups of those mathematical operations concurrently.
GPUs began as specialised processors for real-time 3D graphics. Rendering a game requires similar calculations to be repeated across many vertices, fragments and pixels. That hardware design was later opened to general-purpose computation through platforms such as CUDA.
Animated matrix workload
Transformer matrix workload
Attention results
Press Replay to send a matrix workload through the GPU.
The animation is a conceptual model. Actual GPUs contain far more processing resources and execute work in groups using sophisticated memory hierarchies, scheduling and specialised numerical units.
Representation
Everything begins by turning content into machine-readable units
The model does not receive meaning directly. It receives numbers that represent pieces of content.
Worked example: from sentence to vectors
“The quick brown fox jumps over the lazy dog”
Press Replay to transform the sentence step by step.
Conceptual example: real models often use subword tokens, model-specific IDs and vectors with hundreds or thousands of dimensions. The repeated word the receives the same base token ID here; positional information is added later so the two occurrences can play different contextual roles.
Tokens vary by medium
- Text: words or subword pieces
- Code: names, operators, punctuation and spacing
- Images: pixel or latent patches
- Video: patches extended across time
Embedding space
Related concepts move into nearby regions
An embedding is a learned numeric representation. Distance and direction across many dimensions can encode useful relationships. This two-dimensional map is a simplified projection.
Press Replay clustering to see learned similarity represented as distance.
Position information
Same words, different order, different meaning
The base word embeddings are combined with positions: dog₁ + bites₂ + man₃.
The vocabulary items are unchanged. Their position-aware representations change because each word now occupies a different place in the sequence.
Conceptual model: actual embeddings usually contain hundreds or thousands of dimensions. A 2D chart cannot preserve every relationship; it only makes the idea visible.
Self-attention
Attention asks: what matters to this token now?
Rather than applying a fixed dictionary rule, the model calculates weighted relationships from the current context.
Interactive attention head
Select a query word
Each word creates a query. That query is compared with the keys of all words in the context. This laboratory visualises one hypothetical head; it is not assigned a fixed linguistic role. Click any word below to inspect an illustrative attention pattern.
“it” → “animal” receives the strongest weight, helping the model connect the pronoun with its likely referent.
Illustrative weights: real values depend on the model, layer, attention head and training context. A high weight is a learned association, not a guaranteed grammatical explanation.
Attention can assign a stronger weight between it and animal than between it and street. This does not eliminate ambiguity, but it gives the model a flexible way to construct a context-sensitive representation.
Query
What am I looking for?
The current token expresses what information would be useful.
Key
What do I contain?
Each candidate token advertises the type of context it can match.
Value
What can I contribute?
The matched information is combined according to the calculated weights.
Result
Selective context
Similarity → weights → weighted combination.
Worked numerical example
Follow the attention equation
For the query word “it”, use a tiny two-dimensional query vector and compare it with three candidate words. The numbers are deliberately small so the complete calculation remains visible.
| Word | Key K | Value V | Q · K | Scaled | Softmax weight | Weighted value |
|---|---|---|---|---|---|---|
| animal | [0.9, 0.1] | [1.0, 0.0] | 0.900 | 0.636 | 41.4% | [0.414, 0.000] |
| street | [0.2, 0.8] | [0.0, 1.0] | 0.200 | 0.141 | 25.2% | [0.000, 0.252] |
| tired | [0.6, 0.4] | [0.6, 0.4] | 0.600 | 0.424 | 33.4% | [0.200, 0.134] |
Simplified demonstration: real transformers use much larger vectors and calculate many queries, keys and values simultaneously. Rounding causes small differences in the displayed totals.
Transformer block laboratory
Watch representations move through one block
Begin with the token representations from the sentence. The block first lets them exchange context, then transforms each updated position independently.
Every head receives the same token sequence through different learned projections.
Qh = XWhQ Kh = XWhK Vh = XWhV
Headh = softmax(QhKhT / √dk)Vh
Teaching labels only: training does not assign heads names such as “syntax” or “reference.” Heads are normally identified by layer and number. The descriptions below merely characterise the hand-crafted patterns shown in this demonstration.
All heads: each colour represents a different learned view of the same sentence.
How to read the map: one colour shows all connections calculated by one head; each individual line is one token-to-token attention relationship; line thickness represents its attention weight. The complete set of same-coloured lines visualises that head’s attention-weight matrix.
Interpret with caution: a real head can mix several patterns; one pattern can be distributed across several heads; some heads are redundant or difficult to interpret. An attention pattern alone does not prove a specific linguistic function.
Head outputs are merged. The original input travels through a residual path.
original x + attention output → layer normalisation
The same small neural network transforms each position separately.
The feed-forward output is added to its input, then normalised.
Press Next stage to begin, or scroll the laboratory into view.
Why the pieces matter: attention exchanges information across positions; the feed-forward network transforms each position; residual connections preserve a direct information path; normalisation stabilises the evolving representations.
Model development
Training teaches prediction; alignment shapes behaviour
Training and prompting are different processes. Training changes model parameters. Ordinary prompting uses those parameters to generate a response.
Model-development laboratory
Watch when the parameters change—and when they stop
Input
Large training corpusDocuments, code and other examples provide many next-token prediction exercises.
Model parameters θ
Prediction errors are used to adjust many learned weights.
Observable result
General language modelIt learns broad statistical patterns and next-token prediction, but may not reliably follow user instructions.
Phase 1 — pretraining updates the parameters while the model learns broad predictive patterns.
Simplified lifecycle: real development pipelines vary and the boundaries may overlap. Adaptation and alignment are additional forms of training and can change parameters. During ordinary inference, the trained parameters are normally frozen; the prompt changes the temporary context, not the stored weights. Alignment can improve behaviour but cannot guarantee truth, safety or correctness.
Text and code
Text generation repeats one simple prediction loop
The model reads the current context, predicts a probability distribution, selects a token, appends it, and repeats.
Decoding controls focus and variation
The following probabilities are illustrative. A decoding rule determines how the next token is selected from a model’s predicted distribution.
- Low temperature: more focused and repeatable
- High temperature: more varied and risky
- Top-p / top-k: restricts the candidate pool
These settings change sampling behaviour, not the knowledge encoded in the model. Low temperature does not guarantee truth.
Code is language with stricter structure
Shared with text
- Token-by-token prediction
- Context from the prompt
- Learned patterns and conventions
Stricter than prose
- Syntax must parse
- Types and APIs must match
- Behaviour must pass tests
What makes a coding agent “agentic”?
The language model supplies predictions. The surrounding software supplies tools, file access, test execution and controlled actions. Agency comes from this workflow—not from next-token prediction alone.
Image and video
Image generation often begins with noise
Many practical image systems learn to reverse a corruption process, gradually transforming random noise into a structured image.
Diffusion concept laboratory
Watch a prompt guide structure out of noise
Current state
Random latent noise
No recognisable scene exists yet. The random starting state provides the material that will be iteratively reshaped.
Step 0 — start from random noise in a latent representation.
The same cup should retain its identity while its position changes between frames.
Spatial relationships must work inside each frame; temporal relationships help keep colour, shape and motion coherent across frames.
Conceptual—not a model trace: this animation makes progressive denoising visible, but a real model predicts mathematical updates in pixel or latent space rather than revealing objects in these exact hand-designed stages. The prompt conditions what should emerge; not every image or video generator uses diffusion.
An image can become a sequence of patches
A Vision Transformer represents image regions as patch tokens. A Diffusion Transformer can operate on latent patches during denoising. Attention then models relationships across visual regions.
Visual patch laboratory
Turn one image into patch tokens
Visual input
Patch-token sequence
Stage 1 — the model receives an image or latent feature map.
What becomes a token: each region is flattened or projected into a numeric vector, then combined with positional information. Attention does not compare the visible rectangles directly; it compares learned query, key and value representations derived from their patch vectors.
Text and vision meet through conditioning
Conditioning gives the denoiser information about what the generated image should contain. The same random starting noise can therefore be guided towards different visual results.
Conditioning laboratory
Keep the noise; change the prompt
1 · Text encoder
Prompt → embeddings2 · Transformer denoiser
Noise + text conditionfixed starting noise3 · Image decoder
Latent → visible pixelswaiting for refined latentStage 1 — encode the chosen prompt into numeric text representations.
Conceptual demonstration: text conditioning may enter through cross-attention or related mechanisms, depending on the architecture. It guides probability rather than placing literal words inside the image, and the same prompt can produce different outputs when the starting noise or sampling choices change.
Video adds time
A video model must make every frame look convincing and keep it believable beside every other frame. Identity, motion, lighting, camera movement and causality must remain sufficiently consistent.
Space–time laboratory
Compare linked frames with independent frames
Temporal context lets later frames compare with earlier frames, helping the same red cup follow a smooth path.
What the comparison means: “temporal context on” is a teaching abstraction, not one universal switch. Video architectures may use temporal attention, 3D convolutions, recurrence, optical-flow or motion conditioning, or space–time latent tokens. The shared goal is to coordinate frames instead of treating each image as unrelated.
Unifying view
The modalities share a logic—not one identical pipeline
Across modalities, models represent content numerically, learn contextual relationships and generate iteratively. The representation and generation mechanism still depend on the medium.
| Modality | Representation | Generation | Attention links |
|---|---|---|---|
| Text | Tokens | Next token | Context |
| Code | Code tokens | Next token + tools | Dependencies |
| Image | Latent patches | Denoising or tokens | Visual regions |
| Video | Space-time latents | Denoising or tokens | Frames |
Fluency and realism are not evidence of truth
Hallucination
Verify facts, quotations and citations.
Bias
Test performance across people and contexts.
Privacy
Protect prompts, personal information and source data.
Copyright
Check usage rights, licences and attribution.
Deepfakes
Disclose synthetic or materially altered media.
Security
Review generated code and automated tool actions.
Closing synthesis
Five ideas connect the generative AI landscape
- Content becomes tokens or latent units.
- Embeddings represent those units numerically.
- Attention learns relationships among them.
- Generation proceeds iteratively.
- Humans must verify purpose, evidence and impact.
Next step
From understanding the technology to using it effectively
Knowing how generative models work is the foundation. The next stage is learning how AI chat applications turn the technology into practical workflows—and how people can use those systems productively and responsibly.
Application
How do AI chat applications implement generative AI?
Explore how a chat interface combines a foundation model with conversation context, system instructions, uploaded information, retrieval, tools, memory and safety controls to produce a useful response.
Technique
How can we produce higher-quality, higher-value and higher-impact output?
Learn to define the objective, supply relevant context, specify constraints, provide examples, divide complex work into stages, refine the response iteratively and evaluate the result against clear criteria.
Assurance
How can we manage hallucination, bias, fabricated content and other risks?
Ground important claims in reliable sources, request evidence, verify facts independently, test for missing perspectives, protect sensitive data, disclose synthetic media and retain human review for consequential decisions.
Further reading
Foundational sources
- Vaswani et al. (2017), Attention Is All You Need.
- Ho, Jain & Abbeel (2020), Denoising Diffusion Probabilistic Models.
- Dosovitskiy et al. (2020), An Image Is Worth 16×16 Words.
- Peebles & Xie (2022), Scalable Diffusion Models with Transformers.
- NVIDIA, CUDA Programming Guide: Introduction to the GPU.
- NVIDIA, GPU and CUDA history.