Unit 3: Chain-of-Thought, Agents & ReAct

Advanced Reasoning and AI Agent Frameworks

1

Chain-of-Thought (CoT) Prompting

Definition & Origin

Chain-of-Thought (CoT) prompting is a technique that encourages large language models to show their step-by-step reasoning process before arriving at a final answer. This approach was introduced by Google researchers in 2022 and has since become a fundamental strategy for improving LLM performance on complex tasks.

Why CoT Works

Traditional prompting asks models to jump directly to an answer. CoT instead guides the model to decompose problems into intermediate reasoning steps, mirroring how humans approach complex problems.

Implementation Strategies

Structured Templates

Provide a clear template that forces the model to follow a step-by-step format. For example: "Let's solve this step by step: Step 1... Step 2... Therefore, the answer is..."

Interactive Prompts

Use follow-up questions within the prompt to guide the model through successive reasoning stages, ensuring each step builds on the previous one.

Feedback Loops

Incorporate self-evaluation mechanisms where the model checks its own intermediate steps for consistency before proceeding to the final answer.

Strategies: Explicit vs. Implicit Instructions

Zero-Shot CoT vs. Few-Shot CoT

Feature Zero-Shot CoT Few-Shot CoT
Examples Required None — uses a simple trigger phrase 2-5 worked examples with reasoning steps
Setup Effort Minimal — just add "Let's think step by step" Moderate — requires crafting quality examples
Performance Good improvement over standard prompting Significantly higher accuracy on complex tasks
Token Usage Lower (no examples in context) Higher (examples consume context window)
Best For Quick improvements, resource-constrained scenarios Critical applications requiring high accuracy

Benefits

Limitations

graph TD
    A[Complex Question] --> B[Break into Steps]
    B --> C[Step 1: Identify Key Information]
    C --> D[Step 2: Apply Relevant Knowledge]
    D --> E[Step 3: Reason Through Logic]
    E --> F[Step 4: Draw Conclusion]
    F --> G[Final Answer]
    style A fill:#0891b2,stroke:#0e7490,color:#fff,font-weight:bold
    style G fill:#22c55e,stroke:#4ade80,color:#fff,font-weight:bold
    classDef process fill:#06b6d4,stroke:#22d3ee,color:#fff,font-size:18px,font-weight:bold
    classDef decision fill:#0891b2,stroke:#0e7490,color:#fff,font-size:18px,font-weight:bold
    classDef success fill:#22c55e,stroke:#4ade80,color:#fff,font-size:18px,font-weight:bold
                    
Key Point: CoT prompting makes the model's reasoning process explicit, improving accuracy on complex multi-step problems by up to 40%.
2

AI Agents — Capabilities & Applications

Core Capabilities

Planning & Reflection

AI agents can decompose complex goals into sub-tasks, create execution plans, and reflect on their progress. They evaluate whether their actions align with the overall objective and adjust strategies accordingly.

Tool Access

Agents interact with external tools and APIs — search engines, calculators, code interpreters, databases, and more. This extends their capabilities beyond pure text generation to real-world actions.

Memory

Agents maintain both short-term memory (current conversation context) and long-term memory (past interactions and learned information), enabling coherent, context-aware behavior across sessions.

Applications

Customer Service

Autonomous support agents that resolve issues, escalate when needed, and learn from interactions.

Content Creation

Multi-step content pipelines that research, draft, edit, and publish articles or marketing copy.

Gaming

Dynamic NPCs with adaptive dialogue, decision-making, and memory of player interactions.

Healthcare

Diagnostic assistants that gather patient data, cross-reference medical literature, and suggest treatment plans.

Education

Personalized tutors that adapt to student progress, generate exercises, and provide targeted feedback.

Finance

Trading agents that analyze market data, execute strategies, and manage portfolio rebalancing.

Challenges

Key Point: AI agents combine reasoning with action capabilities, enabling autonomous task completion through tool use and memory systems.
3

Reason and Act (ReAct) Framework

Overview

The ReAct framework was proposed by Yao et al. (2022) and introduces a paradigm that interleaves reasoning traces with actions. Unlike purely reactive systems or purely deliberative ones, ReAct alternates between thinking and doing at each step.

The ReAct Flow

  1. INPUT: The agent receives a question or task from the user.
  2. REASONING: The model generates a thought explaining what it knows, what it needs, and what action to take next.
  3. ACTION: The agent executes an action (e.g., search, calculation, API call).
  4. OBSERVATION: The agent processes the result of the action and feeds it back into the reasoning loop.
  5. RESPONSE: Once the goal is met, the agent produces a final answer.

Advantages

Thoughtful Problem Solving

Each action is preceded by explicit reasoning, ensuring decisions are grounded in context rather than being purely reactive.

External Tool Integration

Actions can invoke search engines, databases, APIs, or calculators — extending the model's capabilities beyond its training data.

Coordinated Think-Act Cycle

The tight coupling of reasoning and action prevents hallucination by grounding responses in observed evidence.

Adaptive Learning

Observations inform subsequent reasoning, allowing the agent to dynamically adjust its strategy based on real-world feedback.

Transparency

The full trace of thoughts and actions provides a complete audit trail, making debugging and trust-building easier.

graph LR
    A[Input] --> B[Reason]
    B --> C[Act]
    C --> D[Observe]
    D --> E{Goal Met?}
    E -->|No| B
    E -->|Yes| F[Respond]
    style A fill:#0891b2,stroke:#0e7490,color:#fff,font-weight:bold
    style B fill:#6366f1,stroke:#818cf8,color:#fff,font-weight:bold
    style C fill:#0ea5e9,stroke:#38bdf8,color:#fff,font-weight:bold
    style D fill:#f59e0b,stroke:#fbbf24,color:#1e1b4b,font-weight:bold
    style F fill:#22c55e,stroke:#4ade80,color:#fff,font-weight:bold
    classDef decision fill:#0891b2,stroke:#0e7490,color:#fff,font-size:18px,font-weight:bold
    classDef process fill:#06b6d4,stroke:#22d3ee,color:#fff,font-size:18px,font-weight:bold
    classDef success fill:#22c55e,stroke:#4ade80,color:#fff,font-size:18px,font-weight:bold
    classDef warning fill:#f59e0b,stroke:#fbbf24,color:#1e1b4b,font-size:18px,font-weight:bold
    classDef info fill:#8b5cf6,stroke:#a78bfa,color:#fff,font-size:18px,font-weight:bold
                    
Key Point: ReAct bridges reasoning and action, allowing LLMs to interact with external tools and environments for complex task completion.
4

OpenAI Functions (LLMs as API)

API Access & HTTP Requests

OpenAI's API allows developers to interact with large language models programmatically via HTTP requests. Understanding the API architecture is essential for building reliable AI-powered applications.

Key Endpoints

/v1/completions

The original completion endpoint. Given a prompt string, it generates a continuation. Best suited for simple text generation tasks.

/v1/chat/completions

The modern chat endpoint. Accepts a list of messages with roles (system, user, assistant) and supports features like function calling, JSON mode, and structured outputs.

Essential Parameters

Parameter Description Example
model Specifies which model to use gpt-4, gpt-3.5-turbo
prompt The input text to generate from "Explain quantum computing"
temperature Controls randomness (0 = deterministic, 2 = max creative) 0.7
max_tokens Maximum tokens in the generated response 1024
stop Sequences that halt generation ["\n\n", "END"]
functions Schema definitions for structured function calling JSON schema object

Best Practices

graph TD
    A[Client Application] --> B[HTTP Request]
    B --> C[API Gateway]
    C --> D[Authentication]
    D --> E[Rate Limiting]
    E --> F[Model Inference]
    F --> G[Response Formatting]
    G --> H[JSON Response]
    H --> A
    style A fill:#0891b2,stroke:#0e7490,color:#fff,font-weight:bold
    style F fill:#6366f1,stroke:#818cf8,color:#fff,font-weight:bold
    style H fill:#22c55e,stroke:#4ade80,color:#fff,font-weight:bold
    classDef process fill:#06b6d4,stroke:#22d3ee,color:#fff,font-size:18px,font-weight:bold
    classDef decision fill:#0891b2,stroke:#0e7490,color:#fff,font-size:18px,font-weight:bold
    classDef success fill:#22c55e,stroke:#4ade80,color:#fff,font-size:18px,font-weight:bold
    classDef info fill:#8b5cf6,stroke:#a78bfa,color:#fff,font-size:18px,font-weight:bold
                    
Key Point: OpenAI Functions enable structured data extraction by defining output schemas, ensuring consistent and parseable API responses.
5

Agent Toolkits & LangChain

OpenAI Functions vs. ReAct — Comparison

Feature OpenAI Functions ReAct
Approach Structured output via schema-defined functions Interleaved reasoning and action loops
Tool Integration Native function calling support Custom tool wrappers with observation feedback
Reasoning Visibility Hidden — model selects functions internally Explicit — full thought trace is visible
Flexibility Limited to defined function schemas Highly flexible — any tool can be invoked
Error Recovery Requires external retry logic Built-in via observation-based adaptation
Best For Structured data extraction, API orchestration Multi-step research, complex problem solving

Agent Toolkits

CSV Agent

Loads and queries CSV files using natural language. Enables data analysis and manipulation without writing SQL or pandas code.

Gmail Toolkit

Reads, sends, and manages Gmail messages. Supports filtering, drafting, and email workflow automation.

Python Agent

Writes and executes Python code in a sandboxed environment. Useful for calculations, data processing, and API interactions.

JSON Agent

Parses and queries nested JSON structures. Extracts information from API responses and configuration files.

SQL Database Toolkit

Connects to SQL databases (PostgreSQL, MySQL, SQLite) and executes natural language queries translated to SQL.

Customising Standard Agents

Summary Matrix — All Frameworks

Framework Type Key Strength Use Case
CoT Prompting Prompting Strategy Step-by-step reasoning Math, logic, multi-step problems
AI Agents Architecture Pattern Autonomous task completion Complex workflows, real-world tasks
ReAct Agent Framework Reason + Act loop Research, multi-tool tasks
OpenAI Functions API Feature Structured outputs Data extraction, API orchestration
LangChain Agents Agent Toolkit Modular tool integration Domain-specific applications
Key Point: LangChain provides modular agent toolkits that combine LLM reasoning with domain-specific tools for real-world applications.
Prev: Unit 2 Next: Unit 4