Generative AI - How It Works
How Generative AI Works
Follow the journey from large banks of training data to a generated response—and discover how testing, transformers and human feedback improve the system.
Vast training data sets
The model learns patterns and relationships.
Stages of the Generative AI Process
A useful generative AI system is not created in one step. It first learns from large data sets, is evaluated with controlled unseen data, and is improved through reinforcement and human feedback. Once trained and tested, it can receive new data and prompts during inferencing.
Build Capability
Training develops initial capability by exposing the model to vast data sets and allowing it to identify patterns.
Check Performance
Testing uses controlled, unseen information to assess how the model handles material it did not learn from.
Use Feedback
Reinforcement learning and RLHF use assessments and rewards to strengthen preferred response behaviour.
Generate Responses
Inferencing applies the trained and tested model to new data and prompts.
The Use of Data
Good-quality training and testing data is extremely valuable. Training data teaches the model patterns, while separate test data evaluates the resulting performance. The two roles must remain distinct.
Training Data
Enormous banks of information are supplied so the model can learn patterns used to construct responses.
- Pre-training data is the first broad batch, before refinement or fine-tuning.
- Later training data is usually more focused or specific.
- Training data develops and refines model capability.
- Data quality directly affects generated-output quality.
Test Data
Controlled, unseen information is used to assess the model after training.
- It must not have been used in any training capacity.
- It evaluates model performance and output.
- It shows how the model handles unfamiliar information.
- Its purpose is assessment, not teaching.
Pre-training data
| Data type | Position | Purpose | Important characteristic |
|---|---|---|---|
| Pre-training data | First broad batch | Build a broad foundation | Used before refinement |
| Later training data | After pre-training | Focus or refine capability | Often more specific |
| Test data | After training | Assess performance and output | Must remain unseen |
The Role of Transformers
A transformer is a deep-learning architecture that uses attention mechanisms to process context and relationships between words. It helps the model predict likely next words, phrases, sentences and paragraphs. Repeating this process enables connected responses that can run into thousands of words.
Context → Prediction → Repetition
The transformer considers the prompt and the response generated so far. It estimates a likely continuation, adds it to the response, and repeats the operation using the expanded context.
data 88%
servers 34%
screens 13%
Awaiting prediction…
Likely Next Element
The transformer helps predict the next likely word from the available context.
Extended Prediction
The same capability supports connected phrases, sentences and paragraphs.
Thousands of Words
Continuing predictions allow lengthy output to build step by step.
Fluency Is Not Accuracy
A detailed, confident and lengthy response may still contain inaccurate information.
The Role of Feedback
Initial training provides broad capability, but model responses still require refinement. Human feedback shows what a desired response looks like and which generated responses are considered correct. The syllabus distinguishes supervised fine-tuning from reinforcement learning from human feedback.
| Method | What the human does | What the model receives | Core distinction |
|---|---|---|---|
| SFT | Creates a desired response | A prompt-response training example | The human supplies the target answer. |
| RLHF | Checks generated responses | A reward signal for preferred responses | The human judges the model's answer. |
Remember the Distinctions
Learn and Evaluate
Training data develops capability; unseen test data assesses performance.
Broad then Focused
Pre-training is the first broad batch; later data is more focused or specific.
Develop and Use
Training develops the model; inferencing applies it to new data and prompts.
Answer and Reward
SFT supplies a desired response; RLHF rewards a preferred model response.
Predict and Repeat
Continuing predictions create connected words, sentences and paragraphs.
Length Is Not Accuracy
Long, confident output can still contain incorrect information.