Unit 4: Image Generation & Visual AI

Diffusion Models, Visual Prompting & Image-to-Video

Section 1: Diffusion Models

Diffusion models are a class of generative AI that create images by learning to reverse a noise-adding process. They work by taking random noise and progressively refining it through iterative denoising steps until a coherent image emerges.

Key Terms

  • Sampling Steps - The number of iterative denoising operations applied
  • Latent Space - A compressed representation where diffusion occurs
  • Decoder - Converts latent representations back to pixel space
  • ControlNet - Provides conditional control during generation

Diffusion Process Flow

Random Noise → Denoising Steps → Latent Space Refinement → Decoder → Final Image
graph LR A[Random Noise] --> B[Denoising Step 1] B --> C[Denoising Step 2] C --> D[Denoising Step N] D --> E[Latent Space] E --> F[Decoder] F --> G[Final Image] style A fill:#d97706,stroke:#b45309,color:#fff,font-weight:bold style E fill:#f59e0b,stroke:#fbbf24,color:#1e1b4b,font-weight:bold style G fill:#22c55e,stroke:#4ade80,color:#fff,font-weight:bold classDef process fill:#f59e0b,stroke:#fbbf24,color:#1e1b4b,font-size:18px,font-weight:bold classDef decision fill:#d97706,stroke:#b45309,color:#fff,font-size:18px,font-weight:bold classDef success fill:#22c55e,stroke:#4ade80,color:#fff,font-size:18px,font-weight:bold

Diffusion models work by learning to reverse a noise-adding process, progressively refining random pixels into coherent images.

Section 2: Prompt Engineering for Images (5 Principles)

Effective image prompting follows a structured approach to achieve predictable, high-quality results. Master these five principles for consistent image generation.

graph TD A[Image Prompt] --> B[Define Subject] B --> C[Set View/Angle] C --> D[Choose Style] D --> E[Add Modifiers] E --> F[Apply Negative Prompts] F --> G[Generate Image] G --> H{Satisfactory?} H -->|Yes| I[Final Image] H -->|No| A style A fill:#d97706,stroke:#b45309,color:#fff,font-weight:bold style I fill:#22c55e,stroke:#4ade80,color:#fff,font-weight:bold classDef process fill:#f59e0b,stroke:#fbbf24,color:#1e1b4b,font-size:18px,font-weight:bold classDef decision fill:#d97706,stroke:#b45309,color:#fff,font-size:18px,font-weight:bold classDef success fill:#22c55e,stroke:#4ade80,color:#fff,font-size:18px,font-weight:bold

The Subject-View-Style framework provides a structured approach to image prompting, ensuring consistent and predictable results.

Section 3: Negative Prompts & Reverse Engineering

Negative prompts are powerful tools for refining image output by explicitly telling the model what NOT to include. Combined with reverse engineering techniques, you can achieve precise control over generated images.

Negative Prompts

Negative prompts specify elements to exclude from the generation process. They help remove artifacts, unwanted objects, and quality issues.

Artifact Removal

Eliminate common AI artifacts like deformed hands, extra fingers, and blurry areas.

Style Exclusion

Prevent unwanted styles like "cartoon," "anime," or "pixelated" from appearing.

Quality Control

Block low-quality results with prompts like "low resolution, blurry, distorted."

Reverse Engineering

Analyze AI-generated images to infer the prompt structure and techniques used. This helps understand what works and why certain prompts produce specific results.

Reverse Engineering Process

Analyze Image → Identify Elements → Infer Keywords → Reconstruct Prompt → Validate

Negative prompts are essential for fine-tuning image output by explicitly excluding undesirable elements.

Section 4: Model Architecture Comparison

Understanding the architectural differences between major AI platforms helps choose the right tool for specific image generation tasks.

OpenAI

GPT-4 → DALL-E → Diffusion Decoder

Uses GPT-4's language understanding combined with diffusion-based image generation for high-quality results.

Google

Gemini → Multimodal Fusion

Integrates multimodal understanding directly into the generation process for context-aware outputs.

Open Source

Stable Diffusion + LLaMA 3

Latent diffusion models with fine-tunable language models for community-driven innovation.

graph TD subgraph OpenAI A1[GPT-4] --> B1[DALL-E] B1 --> C1[Diffusion Decoder] end subgraph Google A2[Gemini] --> B2[Multimodal Fusion] B2 --> C2[Image Generation] end subgraph OpenSource A3[Stable Diffusion] --> B3[Latent Diffusion] A4[LLaMA 3] --> B4[Fine-tunable] end style A1 fill:#d97706,stroke:#b45309,color:#fff,font-weight:bold style A2 fill:#0891b2,stroke:#0e7490,color:#fff,font-weight:bold style A3 fill:#7c3aed,stroke:#6d28d9,color:#fff,font-weight:bold style A4 fill:#dc2626,stroke:#b91c1c,color:#fff,font-weight:bold classDef process fill:#f59e0b,stroke:#fbbf24,color:#1e1b4b,font-size:18px,font-weight:bold

Section 5: Text-to-Video & Inpainting/Outpainting

Advanced visual AI techniques extend beyond static images to video generation and selective image modification.

Text-to-Video

Text-to-video generation uses diffusion and GAN architectures to create coherent motion sequences from text descriptions. These models understand temporal relationships and maintain consistency across frames.

Inpainting & Outpainting

Inpainting

Selectively fill or modify specific regions of an image while preserving the surrounding content. Perfect for removing objects or fixing damaged areas.

Outpainting

Extend images beyond their original borders by generating new content that seamlessly blends with existing elements.

Prompt Rewriting & Optimization

Advanced techniques for refining prompts to achieve better results:

Inpainting and outpainting extend image generation by selectively modifying or expanding existing images.

Prev: Unit 3 Next: Unit 5