Planning and Reflection: How AI Agents Think
What this article covers
- Learn why AI agents without planning wander aimlessly and waste resources
- Understand the key strategies: ReAct, Plan-and-Execute, Tree-of-Thought, and Chain-of-Thought
- Discover how reflection and self-correction work and why they matter
- See concrete examples of the Think-Act-Observe cycle and the Plan-First approach
- Learn which common pitfalls to avoid
Introduction: Planning and Reflection explained
Imagine you ask an AI agent to create a market analysis for a new product. What happens if the agent just starts working without a plan? It jumps between tasks, searching here and there, and critical data ends up missing. This is where planning and reflection come in.
Planning means the agent determines which steps are needed and in what order before taking action. Reflection means the agent checks after each step whether the result is usable or needs refinement. Together, these abilities transform a simple tool into an agent that tackles complex work systematically.
If you’re new to AI agents, start with What is an AI Agent? to build your foundation. The topic of Agent Memory also connects closely to planning and reflection, since an agent must remember its previous steps.
Why do you need planning and reflection?
An agent without planning is like someone tasked with building a house who just starts stacking bricks. No foundation, no blueprint, no concept. The result: a pile of bricks, not a house.
Here’s what happens in practice. You tell the agent to create a market analysis. Without planning, it does this:
- It searches for the product but finds only marketing material, not data
- It switches to competitor analysis, even though it still lacks baseline information
- It tries a tool, unaware it’s using the wrong API
- It jumps back to product search
- After 20 steps, it has consumed tokens but produced no usable analysis
The agent isn’t stupid. It simply lacks a strategy to break down the task and work through it systematically. Planning provides that strategy. Reflection gives the agent the ability to spot errors and correct them before they cascade through subsequent steps.
With planning, the same agent approaches the task like this:
- It breaks the task into subtasks: gather data, analyze competitors, identify trends, write the report
- It works through each subtask in sequence
- After each subtask, it evaluates the results
- If a subtask fails, it considers why and tries a different approach
The difference is dramatic. The planned agent reaches its goal with fewer tokens because it doesn’t jump back and forth. It delivers better results because each step builds on the previous one. And it’s more robust because it catches and corrects errors.
Planning and Reflection in brief
Imagine two chefs. Both must cook a three-course meal.
The first chef reads the recipe and thinks through which ingredients are needed and the order of execution. The sauce goes first since it needs to simmer. The main course comes next. Dessert goes last because it’s served cold. After each step, the chef checks whether the sauce has the right consistency, whether the meat is cooked through. That’s planning and reflection.
The second chef just starts. It throws ingredients in a pan, realizes ten minutes in that it forgot the sauce, runs to the pantry, discovers the cream is empty, tries milk instead, and notices the meat is now burning. The final meal is chaotic and probably inedible.
Planning is the blueprint. Reflection is quality control after each step. Together, they ensure the agent doesn’t just work, but works purposefully.
AI agents use several strategies to achieve this. The most important are ReAct, Plan-and-Execute, Tree-of-Thought, and Chain-of-Thought. All share the principle of forcing the agent to think before and during action.
Who should read this?
Planning and reflection matter for anyone working with AI agents, whether you’re just starting or already experienced.
For beginners: When you first start with AI agents, it’s important to understand that an agent doesn’t simply “work”, its quality depends heavily on how it thinks. If you understand how planning works, you can write better prompts and get better results.
For developers: If you’re building your own agents, you must choose which planning strategy to use. This choice affects how many tokens the agent consumes, how fast it operates, and how reliably it works. Frameworks like LangGraph provide tools to implement planning and reflection.
For decision-makers: If you’re deploying AI agents in your organization, understand that planning and reflection separate useful tools from expensive experiments. An agent that doesn’t reflect can make costly mistakes.
For the curious: Even if you don’t build agents, understanding planning and reflection helps you grasp what AI systems can do and where they hit their limits.
For the bigger picture, check out the article on Agent Systems.
Key terms in planning and reflection
| Term | Meaning |
|---|---|
| Planning | The agent breaks a task into subtasks and determines their order |
| Reflection | The agent checks its own results and learns from them |
| ReAct | Strategy where thinking and acting alternate: Think, Act, Observe |
| Plan-and-Execute | Strategy where a complete plan is created first, then executed |
| Self-Correction | The agent identifies its own errors and fixes them independently |
| Reasoning | The agent’s logical thinking and inference |
| Thought-Action-Observation | The three-step cycle in the ReAct approach |
| Task Decomposition | Breaking a large task into smaller subtasks |
| Replanning | Adjusting the plan when a step fails or new information emerges |
| Chain-of-Thought | Step-by-step thinking where each step builds on the previous one |
Planning strategies for AI agents
Several strategies exist for how an agent can plan. Each has strengths and weaknesses. Here are the four most important.
ReAct: Think, Act, Observe
ReAct is the most well-known strategy for AI agents. The name combines “Reasoning” and “Acting”. The agent operates in a cycle of three steps:
- Think: The agent considers what to do next and why.
- Act: The agent executes an action, such as calling a tool or starting a search.
- Observe: The agent examines the result of the action and considers what it means for the next step.
This cycle repeats until the task is complete. ReAct’s major advantage is flexibility. After each step, the agent can decide whether to continue as planned or change direction.
The downside: on complex tasks, the agent can lose track because it has no fixed plan, just step-by-step progress. This can lead to many steps and high token consumption.
Plan-and-Execute: Plan First, Then Execute
Plan-and-Execute is the counterpart to ReAct. Here, the agent creates a complete plan with all steps before executing any action at all.
- Plan: The agent breaks down the task into subtasks and determines their order.
- Execute: The agent works through each step of the plan sequentially.
The advantage is that the agent has a clear structure from the start. It knows where it’s going and doesn’t waste time experimenting. This is especially helpful for complex tasks that require many steps.
The disadvantage: if a step fails, the agent must adjust the plan. This is called replanning. If that doesn’t work well, the agent gets stuck because the rest of the plan depends on the failed step.
Tree-of-Thought: Exploring Multiple Paths
Tree-of-Thought takes it a step further. Instead of following just one path, the agent explores multiple possible solution paths simultaneously, like a tree with many branches.
- The agent considers multiple possible next steps.
- It pursues each step partway and evaluates how promising it is.
- It picks the best path and goes deeper, or it prunes unworkable branches.
This is especially useful for tasks with multiple solution paths where the best route isn’t obvious from the start. The disadvantage is that Tree-of-Thought consumes many tokens because the agent explores multiple paths at once.
Chain-of-Thought: Step-by-Step Reasoning
Chain-of-Thought is the simplest form of structured thinking. The agent reasons step by step, with each step building on the previous one. It’s like a mathematical proof: the premise first, then the first conclusion, then the next.
Chain-of-Thought isn’t a complete agent strategy like ReAct or Plan-and-Execute, but rather a way of thinking that’s used within other strategies. It helps the agent stay logical and avoid leaps that lead to errors.
Reflection and Self-Correction
Planning is one half. Reflection is the other. Without reflection, an agent can have a perfect plan and still fail because it doesn’t notice when something goes wrong.
Reflection means the agent checks its result after each step or action. It asks itself questions like:
- Does the result match what I expected?
- Is the result usable for the next step?
- Did I make a mistake?
- Do I need to correct something?
When the agent detects an error, it can correct it in two ways:
Self-correction within the step: The agent realizes during execution that something is wrong and corrects itself immediately. For example: it calls an API, gets an error message, recognizes that the parameter was wrong, and tries again with the correct parameter.
Reflection after the step: The agent looks at the completed result of a step and decides if it’s good enough. If not, it figures out what it needs to do differently and executes the step again, possibly with a different strategy.
The reflection cycle looks like this:
- Execute the step
- Check the result
- On error: analyze the cause and correct the step
- On success: move to the next step
- Repeat until the task is complete
This cycle is crucial for an agent’s reliability. An agent that doesn’t reflect can make errors that cascade through all following steps. An agent that reflects catches errors early and fixes them before they cause damage.
Example: An Agent with ReAct
Let’s see how a ReAct agent handles a concrete task. The task is: find the three largest competitors of a particular company and compare their prices.
Step 1: Think
The agent thinks: I need to first find out who the company is and what industry it operates in. Then I can search for competitors.
Step 1: Act
The agent calls a search function with the company name.
Step 1: Observe
The search returns information about the company: it’s a SaaS provider in the project management space.
Step 2: Think
The agent thinks: now I know what industry the company is in. I need to search for the largest competitors in the project management SaaS space.
Step 2: Act
The agent searches for “largest project management SaaS providers”.
Step 2: Observe
The search returns a list of providers. The agent identifies three that offer similar features.
Step 3: Think
The agent thinks: I’ve found three competitors. Now I need to find their prices. I’ll visit the websites of the three providers and look for pricing pages.
Step 3: Act
The agent visits the first competitor’s website and searches for the pricing page.
Step 3: Observe
The website provides pricing information. The agent notes the prices.
Step 4: Think
The agent thinks: first competitor done. Now the second one.
Step 4: Act
The agent visits the second competitor’s website.
Step 4: Observe
The website doesn’t provide direct prices, only “pricing upon request”.
Step 5: Think
The agent thinks: the second competitor has no public pricing. I’ll note that and move to the third competitor. Later I can try to find the prices another way.
Step 5: Act
The agent visits the third competitor’s website.
Step 5: Observe
The website provides pricing information.
Step 6: Think
The agent thinks: I have prices from two competitors and a “pricing upon request” from the third. I can now create a comparison and note that the second provider’s prices aren’t publicly available.
Step 6: Act
The agent creates the comparison report.
Step 6: Observe
The report is complete and contains all available information.
You can see how the agent thinks, acts, and observes at each step. It stays flexible and can respond to unexpected results, like the missing prices from the second competitor.
Example: An Agent with Plan-and-Execute
Now the same task, but with Plan-and-Execute. The agent creates a plan first, then executes it.
Plan:
- Identify the company and determine its industry
- Find the three largest competitors in the industry
- Research prices for each competitor
- Create a comparison report
Execute Step 1: The agent identifies the company as a SaaS provider in project management. Success, moving to the next step.
Execute Step 2: The agent finds three competitors. Success, moving to the next step.
Execute Step 3a: The agent researches the first competitor’s prices. Success, prices found.
Execute Step 3b: The agent researches the second competitor’s prices. Partial success, only “pricing upon request” available. Here the agent must reflect: is that good enough? It decides yes and notes the limitation.
Execute Step 3c: The agent researches the third competitor’s prices. Success, prices found.
Execute Step 4: The agent creates the comparison report with all collected data.
The difference from ReAct: the agent knew from the start which steps were needed. It didn’t have to rethink what comes next at each step. This makes the process more efficient and clearer.
However: if something completely unexpected had happened in Step 3b, for example if the second competitor had no website at all, the agent would need to replan. That’s the point where Plan-and-Execute becomes more complicated than ReAct.
Reflection in Practice
How do you implement reflection in practice? Here’s a simple pattern you’ll find in many frameworks.
After each step, the agent runs through a reflection phase. This phase consists of three parts:
1. Assessment: The agent examines the result and evaluates it. Did it achieve what it was supposed to achieve? Is it complete, correct, and usable?
2. Root cause analysis: If the result isn’t good enough, the agent figures out why. Was the approach wrong? Were the inputs incomplete? Did the tool fail?
3. Correction: The agent decides what to do differently and reruns the step with an adjusted strategy.
In frameworks like LangGraph, you can implement this reflection phase as its own node in the graph. The agent passes through the reflection node after each execution step, and only moves forward to the next step if reflection succeeds.
One important detail: reflection consumes tokens. At every step, the agent thinks harder, consuming compute time and money. You need to weigh whether the extra reliability justifies the cost. For simple tasks, reflection might be overkill. For complex tasks where mistakes are expensive, it’s essential.
Human-in-the-loop approvals also play a role here. At critical steps, the agent can ask a human for confirmation before proceeding. This is a form of external reflection, particularly useful for sensitive operations.
Common Pitfalls with Planning and Reflection
Planning and reflection are powerful, but there are typical mistakes to avoid.
1. Over-planning simple tasks
Not every task needs a detailed plan. If the task is straightforward, planning consumes more time and tokens than it saves. For a simple search, ReAct often works better than Plan-and-Execute because the agent can start immediately.
2. Plans that are too rigid
A plan is a hypothesis, not a law. If the agent clings to a fixed plan while reality shifts, it will fail. Good agents can replan and adjust their strategy when new information emerges.
3. Reflection without consequences
Reflection only helps if the agent acts on its findings. If the agent recognizes a flawed step but continues anyway, the reflection has no value. The agent must have the ability to correct errors, not just identify them.
4. Infinite loops from self-correction
Self-correction can spiral into a loop. The agent tries to fix a step, fails again, corrects again, fails again. You need to set a maximum retry count so the agent eventually gives up and tries a different approach or asks for help.
5. Token explosion from Tree-of-Thought
Tree-of-Thought can consume enormous amounts of tokens if the agent explores too many paths. You need to limit the breadth and depth of the tree, or costs explode.
6. Missing memory
Planning and reflection only work if the agent remembers its previous steps. Without agent memory, the agent forgets what it’s done and starts from scratch repeatedly. Memory is the foundation for meaningful reflection.
7. Too much reflection at every step
If the agent reflects after every tiny step, it spends more time thinking than working. Reflection should be applied at the right points, not reflexively after every action.
8. Ignoring tool errors
Sometimes tools fail and the agent ignores the error and keeps going. This leads to useless results. The agent must recognize tool errors and respond, either through correction or replanning.
Hardware, Costs, and Security in Planning and Reflection
Planning and reflection affect hardware, costs, and security in ways you should understand.
Hardware: Planning and reflection require more compute than simple execution. The agent performs additional reasoning steps that generate and process tokens. If you work locally, you need sufficiently powerful hardware. More on this in the article What is local AI?.
Costs: Every think step and every reflection costs tokens. With cloud-based models, these costs add up fast. A ReAct agent that runs 20 steps with reflection at each one consumes significantly more tokens than a simple agent that runs five steps without reflection. You need to decide whether the extra reliability is worth the cost.
Rule of thumb: the more complex and error-prone the task, the more reflection pays off. For simple, well-defined tasks, a straightforward strategy without extensive reflection often suffices.
Security: Reflection can be a security win. An agent that checks its results makes fewer mistakes that could cause security problems. If the agent wants to delete a file, reflection can ensure it’s the right one. Human-in-the-loop approvals enhance reflection by bringing a human into the loop for critical steps.
Latency: Planning and reflection increase latency, the time it takes for the agent to deliver a result. For tasks that need to be done quickly, this can be a problem. Plan-and-Execute is often faster than ReAct here because the agent moves steadily through steps after the planning phase.
Further Reading and Resources on Planning and Reflection
- AI Agent Fundamentals - Overview of all foundational articles
- What is an AI Agent? - The basics if you’re just starting
- Agent Systems - How multiple agents work together
- Tool-Calling - How agents invoke tools, which enables the act phase
- Agent Memory - How agents remember past steps, the foundation for reflection
- Human-in-the-Loop - When the agent asks a human for approval
- Frameworks - Overview of frameworks for building agents
- LangGraph - A framework that models planning and reflection as graphs
FAQ: Planning and Reflection - Common Questions
What’s the difference between ReAct and Plan-and-Execute?
ReAct works step by step without a fixed plan. The agent reconsiders what comes next at each step. Plan-and-Execute creates a full plan first, then executes it. ReAct is more flexible, Plan-and-Execute is more structured.
Does every agent need planning?
No. For simple tasks with few steps, a basic agent without explicit planning often works fine. Planning becomes important when tasks are complex, involve many steps, or have dependencies.
What exactly is self-correction?
Self-correction means the agent spots and fixes its own errors without human intervention. It can happen within a single step if the agent notices a bad parameter, or after a step if the result isn’t usable.
Doesn’t reflection cost a lot of tokens?
Yes, reflection costs extra tokens because the agent thinks harder. You need to weigh whether the extra reliability is worth it. For complex tasks it usually is, for simple ones it usually isn’t.
What is replanning?
Replanning means the agent adjusts its plan when a step fails or new information appears. It’s especially important in Plan-and-Execute because the original plan is based on assumptions that can change.
How do I prevent infinite loops in self-correction?
Set a maximum retry count. If the agent can’t fix a step after, say, three attempts, it should give up and try a different approach or ask for help.
What is Tree-of-Thought and when do I need it?
Tree-of-Thought is a strategy where the agent explores multiple solution paths simultaneously. You need it for tasks with multiple approaches where the best one isn’t obvious upfront. It’s costly and expensive, so only worthwhile for complex problems.
How does Chain-of-Thought relate to planning?
Chain-of-Thought is a thinking approach where the agent proceeds step by step logically. It’s not a standalone planning strategy but is used within other strategies like ReAct or Plan-and-Execute to make reasoning more structured.
Can an agent reflect without memory?
No, reflection requires memory. The agent must remember its previous steps and results to evaluate them. Without memory, the agent forgets what it’s done and can’t do meaningful reflection.
Which framework is best for planning and reflection?
LangGraph is particularly well-suited because it models planning and reflection as graphs. You can add reflection nodes that run after each step. More in the LangGraph article and the frameworks overview.
Is reflection in AI the same as reflection in humans?
Not quite. In humans, reflection is often a conscious, emotional process. In AI agents, it’s a structured evaluation step where the agent checks results against expectations. The mechanism differs, but the idea is similar: learn from results and improve.
References and Further Reading
- ReAct: Synergizing Reasoning and Acting in Language Models, Yao et al., 2022
- Tree of Thoughts: Deliberate Problem Solving with Large Language Models, Yao et al., 2023
- Chain-of-Thought Prompting Elicits Reasoning in Large Language Models, Wei et al., 2022
- Reflexion: Language Agents with Verbal Reinforcement Learning, Shinn et al., 2023
- LangGraph Documentation, langchain-ai.github.io/langgraph/
- Plan-and-Solve Prompting: Improving Zero-Shot Chain-of-Thought Reasoning by Large Language Models, Wang et al., 2023


