Unit 1: Foundations of Generative AI
Text Generation Models, LLMs & Prompting Basics
What are Text Generation Models?
Text generation models are neural networks trained on vast quantities of text data to understand and produce human-like language. At their core, these models learn statistical patterns from billions of sentences across books, websites, code repositories, and other text sources. Through this training process, they develop an ability to predict and generate coherent sequences of words that follow the patterns they have observed.
The foundation of modern text generation lies in the Transformer architecture, which was introduced in the landmark 2017 paper "Attention Is All You Need." Unlike earlier neural network architectures that processed text sequentially, Transformers use a mechanism called self-attention that allows the model to consider all words in a sequence simultaneously. This parallelization dramatically improved both training speed and the quality of generated text, enabling the creation of models with billions of parameters.
The attention mechanism works by assigning "attention weights" to different parts of the input text, determining how much influence each word should have on the generation of subsequent tokens. This allows the model to capture long-range dependencies and contextual relationships that were difficult for earlier architectures to handle. When you interact with a text generation model, you are engaging with a system that has compressed and encoded these patterns into mathematical representations.
Key Capabilities
- Text Completion: Continuing and extending written passages in a coherent manner
- Question Answering: Providing informative responses to natural language queries
- Summarization: Condensing long documents into concise, meaningful summaries
- Translation: Converting text between different languages with contextual accuracy
- Content Generation: Creating original text across creative, technical, and professional domains
- Code Generation: Writing functional code based on natural language descriptions
Limitations
- Pattern-based, not understanding-based: Models replicate statistical correlations without genuine comprehension of meaning
- Hallucination: May generate plausible-sounding but factually incorrect information
- Temporal knowledge cutoff: Training data has a cutoff date; models do not know current events
- Bias amplification: Can reproduce and amplify biases present in training data
The History of Language Models
The evolution of language models spans over seven decades, progressing from rigid rule-based systems to the flexible, generative AI systems we use today. Understanding this history provides crucial context for appreciating the current capabilities and limitations of modern LLMs.
The earliest approaches to natural language processing relied on hand-crafted rules written by linguists and computer scientists. Systems like ELIZA (1966) used pattern matching and substitution templates to simulate conversation. These systems were brittle and could not handle the vast complexity of natural language.
Researchers shifted toward data-driven approaches, using probability distributions and n-gram models to predict word sequences. IBM's statistical machine translation and early search engines demonstrated that large corpora of text could be leveraged to build more robust language systems without explicitly programming every rule.
The rise of deep learning brought Recurrent Neural Networks (RNNs) to language modeling. RNNs process text sequentially, maintaining a hidden state that carries information from previous tokens. While a significant improvement, RNNs struggled with long sequences due to the vanishing gradient problem.
LSTMs introduced gating mechanisms that allowed networks to selectively remember or forget information over longer sequences. This breakthrough improved the ability to generate coherent text over longer passages and became the dominant architecture for several years.
Google's "Attention Is All You Need" paper introduced the Transformer architecture, replacing recurrent processing with self-attention mechanisms. This enabled fully parallel training and dramatically improved performance on language tasks, forming the foundation for all modern large language models.
OpenAI's GPT (2018) and BERT (2018) demonstrated the power of scaling Transformers to billions of parameters. GPT-2, GPT-3, PaLM, LLaMA, and subsequent models showed that larger models trained on more data could exhibit emergent capabilities, including few-shot learning and complex reasoning.
LLMs in the Market
The large language model market has exploded since 2020, with major technology companies and well-funded startups competing to build the most capable and efficient models. These systems are now being integrated into products across virtually every industry, from healthcare and education to finance and software development.
Market Statistics
Key Efficiency Techniques
- Quantization: Reducing model precision from 32-bit to 4-bit or 8-bit floating point, dramatically decreasing memory requirements with minimal quality loss
- LoRA (Low-Rank Adaptation): Fine-tuning only a small fraction of model parameters, making customization accessible without retraining the entire model
- Mixture of Experts (MoE): Activating only relevant sub-networks for each input, achieving large-model quality with efficient computation
Multimodal Capabilities
Modern LLMs are evolving beyond text-only models. Leading systems now accept and generate images, audio, and video alongside text, enabling richer interactions and more versatile applications across creative, analytical, and operational workflows.
AI System Architecture
Understanding the pipeline from user input to generated output is essential for working effectively with LLMs. Each stage of the architecture plays a specific role in transforming your prompt into a coherent, relevant response. The diagram below illustrates the complete flow, including the feedback mechanisms that improve model performance over time.
Stage Breakdown
- Tokenization: Your text is split into subword units called tokens. Common words may be single tokens, while rare words are split into multiple pieces.
- Embedding: Each token is converted into a high-dimensional vector that captures semantic meaning. These embeddings allow the model to understand relationships between concepts.
- Transformer Blocks: The core computation occurs here. Multiple layers of self-attention and feed-forward networks process the embeddings, building increasingly abstract representations.
- Decoding: The model generates output tokens one at a time, selecting each based on probability distributions over the vocabulary.
- Feedback / RLHF: Human feedback during training (Reinforcement Learning from Human Feedback) and usage patterns help refine model behavior over time.
The 5 Principles of Prompting
Effective prompting is both an art and a science. By mastering these five foundational principles, you can consistently obtain higher-quality outputs from any language model. Each principle addresses a different aspect of the interaction between the user and the AI.
Give Direction
Provide clear context about the task, audience, and goals. Tell the model who it should act as and what perspective to adopt. Direction grounds the model's response in the specific domain and constraints you need.
Specify Format
Define the exact structure you want: bullet points, JSON, table, paragraph, code block, or any other format. Explicit formatting instructions eliminate ambiguity and produce more usable outputs.
Provide Examples
Show the model what good output looks like through one or more demonstrations. Examples calibrate the model's understanding of your expectations and establish patterns it can follow consistently.
Evaluate Quality
Review the model's output and provide corrective feedback. Ask the model to critique its own work, identify errors, or suggest improvements before finalizing the response.
Divide Labor
Break complex tasks into smaller, manageable sub-tasks. Rather than asking for everything in a single prompt, use a sequence of focused prompts that each handle one specific part of the overall goal.
Prompt Types & Combinations
Different tasks call for different prompting strategies. Understanding when and how to apply each prompt type allows you to choose the most effective approach for your specific use case. The decision tree below helps guide your selection process.
Zero-shot
Direct instruction with no examples. The model relies entirely on its pre-trained knowledge. Best for straightforward, well-defined tasks.
One-shot
A single example is provided to demonstrate the expected input-output pattern. Useful when you need a specific format or style.
Few-shot
Multiple examples are given to establish a clear pattern. Ideal for classification, extraction, and tasks requiring consistent formatting.
Role-playing
The model adopts a specific persona or expertise. Effective for domain-specific tasks where the model should simulate specialized knowledge.
Chain of Thought
The model is instructed to reason step-by-step before arriving at a conclusion. Dramatically improves performance on math, logic, and complex reasoning tasks.
Self-evaluation
The model critiques and refines its own output. Useful for quality assurance, catching errors, and improving accuracy without external feedback.
Prompt Limitations
Important Limitations to Understand
- Token Limits: Every model has a maximum context window (e.g., 8K, 128K tokens). Inputs exceeding this limit are truncated or cause errors.
- No Current Knowledge: Models are trained on data up to a specific cutoff date and do not have real-time information unless connected to external tools.
- Bias: Models can reflect and amplify societal biases present in their training data, requiring careful evaluation of outputs.
- Cost: Token-based pricing means that longer prompts and more complex tasks consume more computational resources and budget.