Solving Complex Multi-Step Reasoning in Agentic AI Workflows
By SolutionJet Architecture Team
A single LLM with a few tools handles "look something up and answer" well. It starts to fail when the answer depends on a chain of steps: each tool result changes what to do next, some steps have to be checked before anyone acts on them, and a person may need to sign off halfway through. This article shows when that happens and how to orchestrate it with LangGraph and LangChain models, using one concrete request.
The problem: one request, five dependent decisions
A B2B customer writes: "Reorder 15 Delta-9 widgets for our Austin site, delivered by Friday, at our contract price." (It's the same order flow as in our Prompt Flow order-processing agent, made harder.) To answer it correctly, the system has to:
- Work out which checks the request needs: stock, contract pricing, shipping.
- Run those checks against the ERP, pricing service and carrier API.
- Notice that only 9 are on hand, and re-plan (for example, a split shipment with a backorder).
- Notice that the contract discount is 12%, above the 10% that sales can approve alone, and stop for a manager.
- Only then write a reply, and only from facts the checks returned.
Give this to one tool-calling agent with a long prompt and it usually goes wrong in predictable ways. It re-checks stock but forgets pricing after re-planning, it quotes a delivery date it never looked up, it loops on the same failing call, or it promises the discount because nothing forced it to pause. Each step is easy. The control flow between them is what the prompt can't enforce.
When you need orchestration, and when you don't
Move from a single agent loop to an explicit graph when one or more of these is true:
| Signal | In the example | What the graph gives you |
|---|---|---|
| Later steps depend on earlier results | Short stock changes the plan | Conditional edges and a re-plan loop |
| Independent lookups | Stock, price and ETA don't depend on each other | Parallel branches with Send |
| Rules the model must not own | Discount policy, stock thresholds | A verifier node written in plain code |
| A person must decide | Manager approves >10% discount | interrupt() and resume, possibly hours later |
| Retries and self-correction must be bounded | At most two planning attempts | Attempt counters, RetryPolicy, recursion_limit |
| Restarts and audits | Who approved what, from which data | A checkpointer keyed by thread_id |
If none of these apply (one or two lookups and a direct answer), a plain tool-calling agent is simpler and you should use it. A good rule of thumb: if you can draw the flow on a whiteboard, put it in the graph, not in the prompt. Leave the LLM the parts that need judgment: understanding the request, choosing checks and writing the reply.
The architecture
- Planner (LLM): turns free text into a typed
Plan(item, quantity, site, deadline, which checks to run) via structured output. Code never parses prose. - Parallel checks (tools): one branch per check using
Send. Results are appended to a sharedfindingslist by a reducer. Each branch has its ownRetryPolicy, so a flaky carrier API is retried without re-running the whole agent. - Verifier (code): applies business rules. A failed check sends the run back to the planner with the failure as feedback, at most twice. A discount over policy routes to approval. Otherwise it goes on to the reply.
- Human approval:
interrupt()pauses the run and hands the approver a JSON payload. Nothing is held in memory while the run waits. - Responder (LLM): writes the reply only from
findings, so every number in it came from a system of record. - Checkpointer: saves state after every step under a
thread_id(one per order), which gives you pause/resume, crash recovery and an audit trail.
Building it with LangChain and LangGraph
The whole graph, using the LangGraph 1.x API. The three tool functions are stubs standing in for your ERP, pricing and carrier calls:
from operator import add
from typing import Annotated, Literal
from pydantic import BaseModel
from typing_extensions import TypedDict
from langgraph.checkpoint.memory import InMemorySaver
from langgraph.graph import END, START, StateGraph
from langgraph.types import Command, RetryPolicy, Send, interrupt
Check = Literal["inventory", "pricing", "shipping"]
class Plan(BaseModel):
item: str
quantity: int
site: str
deliver_by: str
checks: list[Check]
class State(TypedDict, total=False):
request: str
plan: dict # Plan.model_dump(): plain JSON in checkpoints
feedback: str # verifier -> planner on a retry
attempts: int
findings: Annotated[list[dict], add] # parallel checks append here
approved: bool
answer: str
# --- tools: in production these call your ERP, pricing and carrier APIs ----
def inventory_check(item, qty):
on_hand = {"Delta-9 widget": 9}.get(item, 0)
return {"check": "inventory", "ok": on_hand >= qty, "on_hand": on_hand,
"backorder_eta": "Tuesday"}
def contract_price(item, qty):
return {"check": "pricing", "ok": True, "unit": 420.0, "discount": 0.12}
def shipping_eta(site, deliver_by):
return {"check": "shipping", "ok": True, "eta": "Thursday", "site": site}
TOOLS = {"inventory": lambda p: inventory_check(p["item"], p["quantity"]),
"pricing": lambda p: contract_price(p["item"], p["quantity"]),
"shipping": lambda p: shipping_eta(p["site"], p["deliver_by"])}
MAX_ATTEMPTS = 2
APPROVAL_DISCOUNT = 0.10
def build_graph(planner, writer, checkpointer=None):
"""planner: chat model with .with_structured_output(Plan) applied.
writer: plain chat model for the customer-facing reply."""
def plan(state: State):
prompt = f"Plan the checks needed for: {state['request']}"
if state.get("feedback"):
prompt += ("\nPrevious attempt failed verification: "
f"{state['feedback']}")
return {"plan": planner.invoke(prompt).model_dump(),
"attempts": state.get("attempts", 0) + 1}
def fan_out(state: State):
# one parallel branch per check the planner asked for
return [Send("run_check", {"plan": state["plan"], "check": c})
for c in state["plan"]["checks"]]
def run_check(branch: dict):
return {"findings": [TOOLS[branch["check"]](branch["plan"])]}
def verify(state: State) -> Command[Literal["plan", "approve", "respond"]]:
# findings accumulate across attempts; judge only this attempt's batch
latest = state["findings"][-len(state["plan"]["checks"]):]
failed = [f for f in latest if not f["ok"]]
if failed and state["attempts"] < MAX_ATTEMPTS:
return Command(goto="plan", update={"feedback": str(failed)})
pricing = next((f for f in latest if f["check"] == "pricing"), None)
if pricing and pricing["discount"] > APPROVAL_DISCOUNT:
return Command(goto="approve")
return Command(goto="respond")
def approve(state: State):
decision = interrupt({"reason": "discount above policy",
"plan": state["plan"],
"findings": state["findings"]})
return {"approved": bool(decision.get("approved"))}
def respond(state: State):
reply = writer.invoke(
f"Customer request: {state['request']}\nFindings: {state['findings']}\n"
f"Approved: {state.get('approved')}\nWrite a short, accurate reply. "
"If stock is short, offer a split shipment.")
return {"answer": reply.content}
g = StateGraph(State)
g.add_node("plan", plan)
g.add_node("run_check", run_check, retry_policy=RetryPolicy(max_attempts=3))
g.add_node("verify", verify)
g.add_node("approve", approve)
g.add_node("respond", respond)
g.add_edge(START, "plan")
g.add_conditional_edges("plan", fan_out, ["run_check"])
g.add_edge("run_check", "verify")
g.add_edge("approve", "respond")
g.add_edge("respond", END)
return g.compile(checkpointer=checkpointer or InMemorySaver())
Wire in a real model and run it. init_chat_model lets you swap providers without touching the graph:
from langchain.chat_models import init_chat_model
from langgraph.types import Command
# Any LangChain chat model works; Azure OpenAI shown here.
llm = init_chat_model("azure_openai:gpt-4o", azure_deployment="gpt-4o")
graph = build_graph(planner=llm.with_structured_output(Plan), writer=llm)
cfg = {"configurable": {"thread_id": "order-1042"}, "recursion_limit": 25}
out = graph.invoke({"request": "Reorder 15 Delta-9 widgets for our Austin site, "
"delivered by Friday, at our contract price."}, cfg)
if "__interrupt__" in out: # paused at the approval node
ticket = out["__interrupt__"][0].value # send this to the approver's queue
# ...minutes or hours later, from any process sharing the checkpointer:
out = graph.invoke(Command(resume={"approved": True}), cfg)
print(out["answer"])
What happens on the example request
- Plan: the planner asks for inventory, pricing and shipping checks.
- Act: the three checks run in parallel: 9 on hand (short by 6, backorder ETA Tuesday), contract discount 12%, carrier ETA Thursday.
- Verify: inventory failed on attempt 1, so the run goes back to the planner with that failure as feedback. The planner can now plan around it, for example by asking for a split shipment or a substitute check.
- Bound: after attempt 2 the verifier stops looping and moves on with what it knows, so the reply can offer the split shipment instead of retrying forever.
- Approve: the 12% discount is above the 10% policy, so the run pauses at
interrupt(). The manager approves later, andCommand(resume=...)continues from the checkpoint. - Respond: "9 units ship Thursday to Austin, the remaining 6 follow Tuesday, at your contract price." Every fact in that sentence came from a tool result.
We ran this graph end to end with scripted models: two planning attempts, six recorded findings, a pause for approval, and a resumed run that produced the reply.
Takeaways
- Put control flow in the graph and judgment in the LLM. Ordering, branching, limits and approvals are code. Interpreting the request and writing the reply are the model's job.
- Make the LLM return structured plans (
with_structured_output), never prose your code has to parse. - Fan out independent tool calls with
Sendand merge them with a reducer. Latency is the slowest check, not the sum of all of them. - Verify with code and bound every loop: an attempt counter in state for self-correction, plus
recursion_limitas a hard stop. - Use
interrupt()for approvals. The node re-runs from the top on resume, so keep side effects (emails, ERP writes) after the interrupt or in a later node. - Checkpoint everything, one
thread_idper business transaction. UseInMemorySaverin development and a durable checkpointer, such as PostgreSQL vialanggraph-checkpoint-postgres, in production. - Retry at the tool, not the agent. A
RetryPolicyon the node that calls the flaky API is cheaper and more predictable than asking the LLM to try again.
On Azure, the same graph runs well as a container (Azure Container Apps or AKS) with Azure OpenAI models and Azure Database for PostgreSQL behind the checkpointer. For the managed alternative, see who Microsoft Foundry is for.