Transformers and Generative AI

Transformer fundamentals

Transformers
and Generative AI

Text · Code · Image · Video

One architecture helps machines learn relationships across four creative media—but each medium generates its output differently.

The central question

What changes—and what stays the same?

A single prompt can ask an AI system to explain plant growth, write a simulation, visualise the process, or animate the explanation. The outputs look completely different. Underneath, however, they share learned representations and relationships.

T

Text

Explain plant growth.

</>

Code

Build a simulation.

◫

Image

Visualise the process.

▶

Video

Animate the explanation.

By the end: you should be able to explain tokens, embeddings and self-attention; trace text and code generation; describe the relationship between transformers and diffusion; and identify where human verification remains essential.
01

Foundation

Generative AI learns patterns, then samples from them

A discriminative model chooses a label. A generative model produces new content that is plausible according to patterns learned during training.

Discriminative

Input → label

“Is this email spam?”

The output is a category, score or decision.

Generative

Prompt → new content

“Write a safer reply.”

The output is constructed through an iterative prediction process.

A generative model does not simply retrieve one stored answer. For text, it estimates likely continuations. For diffusion-based images, it estimates how to remove noise step by step. The model samples from these learned possibilities to produce an output.

Why transformers mattered

Earlier recurrent systems processed sequences mainly one position after another. Transformer attention allows many token positions to be processed in parallel during training and makes long-range relationships easier to model. This combination helped models scale to large datasets and much larger capacities.

Interactive comparison: six token positions

Earlier recurrent modelSequential processing
T₁T₂T₃T₄T₅T₆
Transformer trainingParallel positions
T₁T₂T₃T₄T₅T₆
Text generationAutoregressive output
G₁G₂G₃G₄G₅G₆

Press Replay to compare the three processes.

processing nowcompletedwaiting
Important nuance: transformer training can process many positions in parallel, but autoregressive text generation still produces tokens sequentially.
GPU

Hardware acceleration

How GPUs make transformer parallelism practical

The transformer provides parallelisable operations. A GPU supplies many processing units that can execute large groups of those mathematical operations concurrently.

GPUs began as specialised processors for real-time 3D graphics. Rendering a game requires similar calculations to be repeated across many vertices, fragments and pixels. That hardware design was later opened to general-purpose computation through platforms such as CUDA.

Architecture versus hardware: the transformer removes much of the sequential dependency during training; the GPU executes the resulting matrix operations efficiently at scale.

Animated matrix workload

Transformer matrix workload

many parallel workersillustrative—not to scale

Attention results

Press Replay to send a matrix workload through the GPU.

Graphics originRepeated calculations rendered many vertices and pixels for real-time 3D scenes and games.
General computingCUDA exposed GPU throughput to scientific and other non-graphics workloads.
Transformer trainingMatrix multiplication and attention calculations map effectively onto massively parallel hardware.

The animation is a conceptual model. Actual GPUs contain far more processing resources and execute work in groups using sophisticated memory hierarchies, scheduling and specialised numerical units.

02

Representation

Everything begins by turning content into machine-readable units

The model does not receive meaning directly. It receives numbers that represent pieces of content.

Worked example: from sentence to vectors

“The quick brown fox jumps over the lazy dog”

Tokenscontent units
Token IDsvocabulary indices
Vectorslearned features

Press Replay to transform the sentence step by step.

Conceptual example: real models often use subword tokens, model-specific IDs and vectors with hundreds or thousands of dimensions. The repeated word the receives the same base token ID here; positional information is added later so the two occurrences can play different contextual roles.

Tokens vary by medium

  • Text: words or subword pieces
  • Code: names, operators, punctuation and spacing
  • Images: pixel or latent patches
  • Video: patches extended across time

Embedding space

Related concepts move into nearby regions

An embedding is a learned numeric representation. Distance and direction across many dimensions can encode useful relationships. This two-dimensional map is a simplified projection.

feature dimension 1 →feature dimension 2 → cat kitten dog Python compiler banana

Press Replay clustering to see learned similarity represented as distance.

Position information

Same words, different order, different meaning

The base word embeddings are combined with positions: dog₁ + bites₂ + man₃.

The vocabulary items are unchanged. Their position-aware representations change because each word now occupies a different place in the sequence.

Conceptual model: actual embeddings usually contain hundreds or thousands of dimensions. A 2D chart cannot preserve every relationship; it only makes the idea visible.

03

Self-attention

Attention asks: what matters to this token now?

Rather than applying a fixed dictionary rule, the model calculates weighted relationships from the current context.

Interactive attention head

Select a query word

Each word creates a query. That query is compared with the keys of all words in the context. This laboratory visualises one hypothetical head; it is not assigned a fixed linguistic role. Click any word below to inspect an illustrative attention pattern.

Current query:it→ which context is useful now?

“it” → “animal” receives the strongest weight, helping the model connect the pronoun with its likely referent.

Illustrative weights: real values depend on the model, layer, attention head and training context. A high weight is a learned association, not a guaranteed grammatical explanation.

Attention can assign a stronger weight between it and animal than between it and street. This does not eliminate ambiguity, but it gives the model a flexible way to construct a context-sensitive representation.

Query

What am I looking for?

The current token expresses what information would be useful.

Key

What do I contain?

Each candidate token advertises the type of context it can match.

Value

What can I contribute?

The matched information is combined according to the calculated weights.

Result

Selective context

Similarity → weights → weighted combination.

Worked numerical example

Follow the attention equation

For the query word “it”, use a tiny two-dimensional query vector and compare it with three candidate words. The numbers are deliberately small so the complete calculation remains visible.

Q(it) = [1.0, 0.0]dₖ = 2√dₖ = 1.414
WordKey KValue VQ · KScaledSoftmax weightWeighted value
animal[0.9, 0.1][1.0, 0.0]0.9000.63641.4%[0.414, 0.000]
street[0.2, 0.8][0.0, 1.0]0.2000.14125.2%[0.000, 0.252]
tired[0.6, 0.4][0.6, 0.4]0.6000.42433.4%[0.200, 0.134]
Choose step 1 or press Next step to begin the calculation.
Final context vector[ hidden until step 4 ]

Simplified demonstration: real transformers use much larger vectors and calculate many queries, keys and values simultaneously. Rounding causes small differences in the displayed totals.

Transformer block laboratory

Watch representations move through one block

Begin with the token representations from the sentence. The block first lets them exchange context, then transforms each updated position independently.

Theanimaldidnotcrossthestreetbecauseitwastired
Multi-head self-attention
Every head receives the same token sequence through different learned projections.
1
What exactly is one attention head? It is one learned attention-calculation path running in parallel with the other heads. Each head has its own three learned projection matrices and applies the same attention operation: Qh = XWhQ   Kh = XWhK   Vh = XWhV
Headh = softmax(QhKhT / √dk)Vh
✓ mathematical operation✓ three learned weight matrices✓ parallel computational pathnot one weightnot one connecting linenot a complete separate network

Teaching labels only: training does not assign heads names such as “syntax” or “reference.” Heads are normally identified by layer and number. The descriptions below merely characterise the hand-crafted patterns shown in this demonstration.

All heads: each colour represents a different learned view of the same sentence.

How to read the map: one colour shows all connections calculated by one head; each individual line is one token-to-token attention relationship; line thickness represents its attention weight. The complete set of same-coloured lines visualises that head’s attention-weight matrix.

Interpret with caution: a real head can mix several patterns; one pattern can be distributed across several heads; some heads are redundant or difficult to interpret. An attention pattern alone does not prove a specific linguistic function.

Concatenate, project, add and normalise
Head outputs are merged. The original input travels through a residual path.
2
Head 1+Head 2+Head 3+Head 4→projected context
Residual route
original x + attention output → layer normalisation
Position-wise feed-forward network
The same small neural network transforms each position separately.
3
animal′
it′
tired′
Second residual connection and normalisation
The feed-forward output is added to its input, then normalised.
4
z₁: animal + contextz₂: it + referent contextz₃: tired + relation context

Press Next stage to begin, or scroll the laboratory into view.

one block shown · often repeated many times

Why the pieces matter: attention exchanges information across positions; the feed-forward network transforms each position; residual connections preserve a direct information path; normalisation stabilises the evolving representations.

04

Model development

Training teaches prediction; alignment shapes behaviour

Training and prompting are different processes. Training changes model parameters. Ordinary prompting uses those parameters to generate a response.

Model-development laboratory

Watch when the parameters change—and when they stop

Input

Large training corpus

Documents, code and other examples provide many next-token prediction exercises.

UPDATING

Model parameters θ

Prediction errors are used to adjust many learned weights.

Observable result

General language model

It learns broad statistical patterns and next-token prediction, but may not reliably follow user instructions.

Optimisation signalReduce next-token prediction error across broad data.

Phase 1 — pretraining updates the parameters while the model learns broad predictive patterns.

Simplified lifecycle: real development pipelines vary and the boundaries may overlap. Adaptation and alignment are additional forms of training and can change parameters. During ordinary inference, the trained parameters are normally frozen; the prompt changes the temporary context, not the stored weights. Alignment can improve behaviour but cannot guarantee truth, safety or correctness.

05

Text and code

Text generation repeats one simple prediction loop

The model reads the current context, predicts a probability distribution, selects a token, appends it, and repeats.

Read context
→
Predict probabilities
→
Choose a token
→
Append + repeat
“The capital of Malaysia is” → Kuala → Lumpur → …

Decoding controls focus and variation

The following probabilities are illustrative. A decoding rule determines how the next token is selected from a model’s predicted distribution.

Kuala72%
the11%
a7%
beautiful4%
  • Low temperature: more focused and repeatable
  • High temperature: more varied and risky
  • Top-p / top-k: restricts the candidate pool

These settings change sampling behaviour, not the knowledge encoded in the model. Low temperature does not guarantee truth.

Code is language with stricter structure

Shared with text

  • Token-by-token prediction
  • Context from the prompt
  • Learned patterns and conventions

Stricter than prose

  • Syntax must parse
  • Types and APIs must match
  • Behaviour must pass tests
Plausible code is not the same as correct code. Code generation is useful for scaffolding, transformation, explanation and tests, but execution and review are the ultimate checks.

What makes a coding agent “agentic”?

Plan
→
Write
→
Run
→
Inspect
→
Revise

The language model supplies predictions. The surrounding software supplies tools, file access, test execution and controlled actions. Agency comes from this workflow—not from next-token prediction alone.

06

Image and video

Image generation often begins with noise

Many practical image systems learn to reverse a corruption process, gradually transforming random noise into a structured image.

Diffusion concept laboratory

Watch a prompt guide structure out of noise

Prompta red cupon a blue tablesunlit window
NOISE · step 0

Current state

Random latent noise

No recognisable scene exists yet. The random starting state provides the material that will be iteratively reshaped.

0% structure visible

Step 0 — start from random noise in a latent representation.

Video extends the problem across time
The same cup should retain its identity while its position changes between frames.
t₁
t₂
t₃
t₄
t₅

Spatial relationships must work inside each frame; temporal relationships help keep colour, shape and motion coherent across frames.

Conceptual—not a model trace: this animation makes progressive denoising visible, but a real model predicts mathematical updates in pixel or latent space rather than revealing objects in these exact hand-designed stages. The prompt conditions what should emerge; not every image or video generator uses diffusion.

An image can become a sequence of patches

A Vision Transformer represents image regions as patch tokens. A Diffusion Transformer can operate on latent patches during denoising. Attention then models relationships across visual regions.

Visual patch laboratory

Turn one image into patch tokens

Visual input

Patch-token sequence

selected query patchrelated context patches

Stage 1 — the model receives an image or latent feature map.

What becomes a token: each region is flattened or projected into a numeric vector, then combined with positional information. Attention does not compare the visible rectangles directly; it compares learned query, key and value representations derived from their patch vectors.

Text and vision meet through conditioning

Conditioning gives the denoiser information about what the generated image should contain. The same random starting noise can therefore be guided towards different visual results.

Conditioning laboratory

Keep the noise; change the prompt

1 · Text encoder

Prompt → embeddings

2 · Transformer denoiser

Noise + text conditionfixed starting noise

3 · Image decoder

Latent → visible pixelswaiting for refined latent

Stage 1 — encode the chosen prompt into numeric text representations.

Conceptual demonstration: text conditioning may enter through cross-attention or related mechanisms, depending on the architecture. It guides probability rather than placing literal words inside the image, and the same prompt can produce different outputs when the starting noise or sampling choices change.

Video adds time

A video model must make every frame look convincing and keep it believable beside every other frame. Identity, motion, lighting, camera movement and causality must remain sufficiently consistent.

Space–time laboratory

Compare linked frames with independent frames

t₁
t₂
t₃
t₄
t₅
Identity ✓Motion ✓Lighting ✓Causality ✓

Temporal context lets later frames compare with earlier frames, helping the same red cup follow a smooth path.

What the comparison means: “temporal context on” is a teaching abstraction, not one universal switch. Video architectures may use temporal attention, 3D convolutions, recurrence, optical-flow or motion conditioning, or space–time latent tokens. The shared goal is to coordinate frames instead of treating each image as unrelated.

07

Unifying view

The modalities share a logic—not one identical pipeline

Across modalities, models represent content numerically, learn contextual relationships and generate iteratively. The representation and generation mechanism still depend on the medium.

ModalityRepresentationGenerationAttention links
TextTokensNext tokenContext
CodeCode tokensNext token + toolsDependencies
ImageLatent patchesDenoising or tokensVisual regions
VideoSpace-time latentsDenoising or tokensFrames

Fluency and realism are not evidence of truth

Hallucination

Verify facts, quotations and citations.

Bias

Test performance across people and contexts.

Privacy

Protect prompts, personal information and source data.

Copyright

Check usage rights, licences and attribution.

Deepfakes

Disclose synthetic or materially altered media.

Security

Review generated code and automated tool actions.

Closing synthesis

Five ideas connect the generative AI landscape

  1. Content becomes tokens or latent units.
  2. Embeddings represent those units numerically.
  3. Attention learns relationships among them.
  4. Generation proceeds iteratively.
  5. Humans must verify purpose, evidence and impact.
Text, code, images and video differ in representation and generation mechanics. The transformer’s unifying contribution is a general mechanism for learning contextual relationships.
→

Next step

From understanding the technology to using it effectively

Knowing how generative models work is the foundation. The next stage is learning how AI chat applications turn the technology into practical workflows—and how people can use those systems productively and responsibly.

Application

How do AI chat applications implement generative AI?

Explore how a chat interface combines a foundation model with conversation context, system instructions, uploaded information, retrieval, tools, memory and safety controls to produce a useful response.

Technique

How can we produce higher-quality, higher-value and higher-impact output?

Learn to define the objective, supply relevant context, specify constraints, provide examples, divide complex work into stages, refine the response iteratively and evaluate the result against clear criteria.

Assurance

How can we manage hallucination, bias, fabricated content and other risks?

Ground important claims in reliable sources, request evidence, verify facts independently, test for missing perspectives, protect sensitive data, disclose synthetic media and retain human review for consequential decisions.

The practical progression: understand the model → design the interaction → evaluate the output → verify before use.

Further reading

Foundational sources

  1. Vaswani et al. (2017), Attention Is All You Need.
  2. Ho, Jain & Abbeel (2020), Denoising Diffusion Probabilistic Models.
  3. Dosovitskiy et al. (2020), An Image Is Worth 16×16 Words.
  4. Peebles & Xie (2022), Scalable Diffusion Models with Transformers.
  5. NVIDIA, CUDA Programming Guide: Introduction to the GPU.
  6. NVIDIA, GPU and CUDA history.
Transformers and Generative AI · Text · Code · Image · Video