Standard RAG is "retrieve then answer." Agentic RAG is "think about what to retrieve, retrieve…
it, evaluate whether it's sufficient, and retrieve more if needed."
The difference in practice is enormous.
Standard RAG: user asks a complex question. System retrieves top-5 documents by similarity. Feeds them to the LLM. If the right information wasn't in those 5 documents, the answer is wrong or incomplete.
Agentic RAG: user asks a complex question. Agent decomposes it into sub-questions. Retrieves documents for each sub-question. Evaluates whether the retrieved information is sufficient. Reformulates queries and retrieves again if needed. Synthesizes the final answer from all gathered information.
I built an agentic RAG system for a legal research application. Complex legal questions often require information from multiple areas of law. Standard RAG retrieved tangentially relevant cases. The agentic version identified the specific legal principles involved, searched for relevant precedents for each, and synthesized a comprehensive answer.
Quality improvement: about 35% better on complex multi-hop questions. Cost increase: about 3x (more LLM calls per query). Worth it for the use case: absolutely.
The implementation uses LangGraph for the agent loop, with explicit retrieval evaluation steps. The key design choice: clear criteria for "is this retrieval sufficient?" — without that, the agent either gives up too early or loops forever.
This is where RAG is heading. Static retrieve-and-answer is the baseline. Intelligent, iterative retrieval is the future.