Chain-of-Thought (CoT) Prompting
Definition & Origin
Chain-of-Thought (CoT) prompting is a technique that encourages large language models to show their step-by-step reasoning process before arriving at a final answer. This approach was introduced by Google researchers in 2022 and has since become a fundamental strategy for improving LLM performance on complex tasks.
Why CoT Works
Traditional prompting asks models to jump directly to an answer. CoT instead guides the model to decompose problems into intermediate reasoning steps, mirroring how humans approach complex problems.
Implementation Strategies
Structured Templates
Provide a clear template that forces the model to follow a step-by-step format. For example: "Let's solve this step by step: Step 1... Step 2... Therefore, the answer is..."
Interactive Prompts
Use follow-up questions within the prompt to guide the model through successive reasoning stages, ensuring each step builds on the previous one.
Feedback Loops
Incorporate self-evaluation mechanisms where the model checks its own intermediate steps for consistency before proceeding to the final answer.
Strategies: Explicit vs. Implicit Instructions
- Explicit CoT: Directly instruct the model to "think step by step" or "show your reasoning process." Example: "Solve this problem step by step."
- Implicit CoT: Embed reasoning demonstrations within the prompt without explicitly telling the model to reason. The model follows the pattern set by examples.
- Demonstrative Examples: Provide worked examples showing the intermediate reasoning steps, then ask the model to solve a similar problem following the same pattern.
Zero-Shot CoT vs. Few-Shot CoT
| Feature | Zero-Shot CoT | Few-Shot CoT |
|---|---|---|
| Examples Required | None — uses a simple trigger phrase | 2-5 worked examples with reasoning steps |
| Setup Effort | Minimal — just add "Let's think step by step" | Moderate — requires crafting quality examples |
| Performance | Good improvement over standard prompting | Significantly higher accuracy on complex tasks |
| Token Usage | Lower (no examples in context) | Higher (examples consume context window) |
| Best For | Quick improvements, resource-constrained scenarios | Critical applications requiring high accuracy |
Benefits
- Improved Accuracy: Studies show up to 40% improvement on multi-step reasoning tasks compared to standard prompting.
- Enhanced Interpretability: The reasoning trace provides transparency into how the model arrived at its conclusion, making errors easier to identify and correct.
- Better Generalization: Models trained or prompted with CoT show improved performance on out-of-distribution problems.
Limitations
- Model-Dependent: CoT effectiveness varies significantly across model sizes — smaller models may produce coherent-looking but incorrect reasoning chains.
- Verbose Outputs: The step-by-step format increases output length, leading to higher token costs and latency.
- Error Propagation: If an early step is incorrect, the entire reasoning chain can be compromised.
graph TD
A[Complex Question] --> B[Break into Steps]
B --> C[Step 1: Identify Key Information]
C --> D[Step 2: Apply Relevant Knowledge]
D --> E[Step 3: Reason Through Logic]
E --> F[Step 4: Draw Conclusion]
F --> G[Final Answer]
style A fill:#0891b2,stroke:#0e7490,color:#fff,font-weight:bold
style G fill:#22c55e,stroke:#4ade80,color:#fff,font-weight:bold
classDef process fill:#06b6d4,stroke:#22d3ee,color:#fff,font-size:18px,font-weight:bold
classDef decision fill:#0891b2,stroke:#0e7490,color:#fff,font-size:18px,font-weight:bold
classDef success fill:#22c55e,stroke:#4ade80,color:#fff,font-size:18px,font-weight:bold
AI Agents — Capabilities & Applications
Core Capabilities
Planning & Reflection
AI agents can decompose complex goals into sub-tasks, create execution plans, and reflect on their progress. They evaluate whether their actions align with the overall objective and adjust strategies accordingly.
Tool Access
Agents interact with external tools and APIs — search engines, calculators, code interpreters, databases, and more. This extends their capabilities beyond pure text generation to real-world actions.
Memory
Agents maintain both short-term memory (current conversation context) and long-term memory (past interactions and learned information), enabling coherent, context-aware behavior across sessions.
Applications
Customer Service
Autonomous support agents that resolve issues, escalate when needed, and learn from interactions.
Content Creation
Multi-step content pipelines that research, draft, edit, and publish articles or marketing copy.
Gaming
Dynamic NPCs with adaptive dialogue, decision-making, and memory of player interactions.
Healthcare
Diagnostic assistants that gather patient data, cross-reference medical literature, and suggest treatment plans.
Education
Personalized tutors that adapt to student progress, generate exercises, and provide targeted feedback.
Finance
Trading agents that analyze market data, execute strategies, and manage portfolio rebalancing.
Challenges
- Ethics & Bias: Agents may amplify existing biases in training data or make ethically questionable decisions when operating autonomously.
- Privacy & Security: Tool access creates attack surfaces; agents handling sensitive data require robust security measures and access controls.
- Interpretability: Understanding why an agent took a particular action sequence remains difficult, especially with multi-step reasoning chains.
- Scalability: Managing hundreds or thousands of concurrent agents requires efficient resource allocation and coordination mechanisms.
Reason and Act (ReAct) Framework
Overview
The ReAct framework was proposed by Yao et al. (2022) and introduces a paradigm that interleaves reasoning traces with actions. Unlike purely reactive systems or purely deliberative ones, ReAct alternates between thinking and doing at each step.
The ReAct Flow
- INPUT: The agent receives a question or task from the user.
- REASONING: The model generates a thought explaining what it knows, what it needs, and what action to take next.
- ACTION: The agent executes an action (e.g., search, calculation, API call).
- OBSERVATION: The agent processes the result of the action and feeds it back into the reasoning loop.
- RESPONSE: Once the goal is met, the agent produces a final answer.
Advantages
Thoughtful Problem Solving
Each action is preceded by explicit reasoning, ensuring decisions are grounded in context rather than being purely reactive.
External Tool Integration
Actions can invoke search engines, databases, APIs, or calculators — extending the model's capabilities beyond its training data.
Coordinated Think-Act Cycle
The tight coupling of reasoning and action prevents hallucination by grounding responses in observed evidence.
Adaptive Learning
Observations inform subsequent reasoning, allowing the agent to dynamically adjust its strategy based on real-world feedback.
Transparency
The full trace of thoughts and actions provides a complete audit trail, making debugging and trust-building easier.
graph LR
A[Input] --> B[Reason]
B --> C[Act]
C --> D[Observe]
D --> E{Goal Met?}
E -->|No| B
E -->|Yes| F[Respond]
style A fill:#0891b2,stroke:#0e7490,color:#fff,font-weight:bold
style B fill:#6366f1,stroke:#818cf8,color:#fff,font-weight:bold
style C fill:#0ea5e9,stroke:#38bdf8,color:#fff,font-weight:bold
style D fill:#f59e0b,stroke:#fbbf24,color:#1e1b4b,font-weight:bold
style F fill:#22c55e,stroke:#4ade80,color:#fff,font-weight:bold
classDef decision fill:#0891b2,stroke:#0e7490,color:#fff,font-size:18px,font-weight:bold
classDef process fill:#06b6d4,stroke:#22d3ee,color:#fff,font-size:18px,font-weight:bold
classDef success fill:#22c55e,stroke:#4ade80,color:#fff,font-size:18px,font-weight:bold
classDef warning fill:#f59e0b,stroke:#fbbf24,color:#1e1b4b,font-size:18px,font-weight:bold
classDef info fill:#8b5cf6,stroke:#a78bfa,color:#fff,font-size:18px,font-weight:bold
OpenAI Functions (LLMs as API)
API Access & HTTP Requests
OpenAI's API allows developers to interact with large language models programmatically via HTTP requests. Understanding the API architecture is essential for building reliable AI-powered applications.
Key Endpoints
/v1/completions
The original completion endpoint. Given a prompt string, it generates a continuation. Best suited for simple text generation tasks.
/v1/chat/completions
The modern chat endpoint. Accepts a list of messages with roles (system, user, assistant) and supports features like function calling, JSON mode, and structured outputs.
Essential Parameters
| Parameter | Description | Example |
|---|---|---|
| model | Specifies which model to use | gpt-4, gpt-3.5-turbo |
| prompt | The input text to generate from | "Explain quantum computing" |
| temperature | Controls randomness (0 = deterministic, 2 = max creative) | 0.7 |
| max_tokens | Maximum tokens in the generated response | 1024 |
| stop | Sequences that halt generation | ["\n\n", "END"] |
| functions | Schema definitions for structured function calling | JSON schema object |
Best Practices
- Security: Never expose API keys in client-side code. Use environment variables and server-side proxies. Implement rate limiting and input validation.
- Efficiency: Use streaming responses for better user experience. Batch similar requests. Cache frequent responses where appropriate.
- Error Handling: Implement retry logic with exponential backoff for rate limits (429) and server errors (500). Handle timeout gracefully.
- Monitoring: Track token usage, latency, error rates, and costs. Set up alerts for anomalous patterns or budget thresholds.
graph TD
A[Client Application] --> B[HTTP Request]
B --> C[API Gateway]
C --> D[Authentication]
D --> E[Rate Limiting]
E --> F[Model Inference]
F --> G[Response Formatting]
G --> H[JSON Response]
H --> A
style A fill:#0891b2,stroke:#0e7490,color:#fff,font-weight:bold
style F fill:#6366f1,stroke:#818cf8,color:#fff,font-weight:bold
style H fill:#22c55e,stroke:#4ade80,color:#fff,font-weight:bold
classDef process fill:#06b6d4,stroke:#22d3ee,color:#fff,font-size:18px,font-weight:bold
classDef decision fill:#0891b2,stroke:#0e7490,color:#fff,font-size:18px,font-weight:bold
classDef success fill:#22c55e,stroke:#4ade80,color:#fff,font-size:18px,font-weight:bold
classDef info fill:#8b5cf6,stroke:#a78bfa,color:#fff,font-size:18px,font-weight:bold
Agent Toolkits & LangChain
OpenAI Functions vs. ReAct — Comparison
| Feature | OpenAI Functions | ReAct |
|---|---|---|
| Approach | Structured output via schema-defined functions | Interleaved reasoning and action loops |
| Tool Integration | Native function calling support | Custom tool wrappers with observation feedback |
| Reasoning Visibility | Hidden — model selects functions internally | Explicit — full thought trace is visible |
| Flexibility | Limited to defined function schemas | Highly flexible — any tool can be invoked |
| Error Recovery | Requires external retry logic | Built-in via observation-based adaptation |
| Best For | Structured data extraction, API orchestration | Multi-step research, complex problem solving |
Agent Toolkits
CSV Agent
Loads and queries CSV files using natural language. Enables data analysis and manipulation without writing SQL or pandas code.
Gmail Toolkit
Reads, sends, and manages Gmail messages. Supports filtering, drafting, and email workflow automation.
Python Agent
Writes and executes Python code in a sandboxed environment. Useful for calculations, data processing, and API interactions.
JSON Agent
Parses and queries nested JSON structures. Extracts information from API responses and configuration files.
SQL Database Toolkit
Connects to SQL databases (PostgreSQL, MySQL, SQLite) and executes natural language queries translated to SQL.
Customising Standard Agents
- prefix / suffix: Add system instructions before or after the default prompt to customize agent behavior and personality.
- max_iterations: Limit the number of reasoning-action loops to control costs and prevent infinite cycles.
- verbose: Enable detailed logging of agent thought processes for debugging and monitoring.
- callback_manager: Hook into agent lifecycle events (tool calls, errors, completions) for logging, analytics, and custom behavior.
Summary Matrix — All Frameworks
| Framework | Type | Key Strength | Use Case |
|---|---|---|---|
| CoT Prompting | Prompting Strategy | Step-by-step reasoning | Math, logic, multi-step problems |
| AI Agents | Architecture Pattern | Autonomous task completion | Complex workflows, real-world tasks |
| ReAct | Agent Framework | Reason + Act loop | Research, multi-tool tasks |
| OpenAI Functions | API Feature | Structured outputs | Data extraction, API orchestration |
| LangChain Agents | Agent Toolkit | Modular tool integration | Domain-specific applications |