A research-level Retrieval Augmented Generation system that retrieves knowledge, generates answers, evaluates its own outputs, and iteratively self-corrects β with confidence scoring, hallucination detection, and memory support.
Standard RAG systems follow a simple pipeline: retrieve β generate β done.
That's a serious limitation.
They have no mechanism to catch:
- Retrieved documents that are completely off-topic
- Answers that hallucinate facts not present in the source
- Responses that only partially address the question
- Queries that were too vague to retrieve useful context
The result? Users receive confidently wrong answers with no signal that something went wrong.
This project solves that. By building a self-correcting feedback loop on top of RAG, the system can detect its own failures and fix them before responding.
Instead of retrieve β generate β done, this system does:
Retrieve β Grade Relevance β Generate β Self-Evaluate β Fix if Needed β Score Confidence β Respond
If the retrieved documents aren't relevant, it rewrites the query and retrieves again. If the generated answer is poorly grounded or incomplete, it revises and retries (up to 2x). Every final answer comes with a confidence score built from three independent graders.
| Grader | What it checks |
|---|---|
| Relevance | Are the retrieved docs actually useful for this question? |
| Grounding | Is the answer supported by the docs, or is it hallucinating? |
| Completeness | Does the answer fully address what was asked? |
The three scores combine into a single weighted confidence percentage shown to the user.
Built a simple LangChain pipeline: PDF upload β chunking β ChromaDB storage β retrieval β LLM generation. No feedback, no evaluation. Just retrieve and generate.
Problem found: System confidently answered questions using irrelevant chunks. No way to know if the answer was trustworthy.
Added a relevance grader that scores retrieved documents before passing them to the LLM. If relevance is below threshold, the query gets rewritten and retrieval is retried.
Problem found: Even with relevant docs, the LLM sometimes fabricated details not present in the source.
Introduced two more independent graders post-generation:
- Grounding grader β checks if every claim in the answer is supported by retrieved docs
- Completeness grader β checks if the answer actually addresses the full question
Each grader runs as a separate LLM call with a focused prompt, scoring 0.0β1.0.
Problem found: The three graders were running inside a linear chain β there was no way to loop back and retry on failure.
Replaced the linear LangChain chain with a LangGraph graph with conditional branching:
- If relevance fails β rewrite query β retrieve again
- If grounding or completeness fails β revise query β regenerate
- After max 2 retries β accept best available answer, flag low confidence
This turned the system from a pipeline into an agent with a correction loop.
Problem found: ChromaDB would throw stale collection errors if a new PDF was uploaded mid-session. Global object initialization was the culprit.
- Lazy loading for ChromaDB and the retriever β fresh objects on every call, no stale state
- Weighted confidence scorer β combines the three grader scores into one user-facing number
- Streamlit UI β PDF upload, question input, score visualization, query history display
- Conversation memory β tracks prior questions in session for context continuity
User Question
β
βΌ
βββββββββββββββββββ
β Retrieve Docs βββββββββββββββββββββββββββ
β (ChromaDB) β β
ββββββββββ¬βββββββββ β
β β Rewrite Query
βΌ β
βββββββββββββββββββ β
β Grade Docs βββββ Not Relevant? βββββββ
β (Relevance) β
ββββββββββ¬βββββββββ
β Relevant
βΌ
βββββββββββββββββββ
β Generate Answer β
β (OpenRouter β
β LLM) β
ββββββββββ¬βββββββββ
β
βΌ
ββββββββββββββββββββββββββββββββ
β Self-Evaluation β
β ββββββββββββββββββββββββ β
β β Relevance Score β β
β β Grounding Score β β
β β Completeness Score β β
β ββββββββββββββββββββββββ β
ββββββββββ¬ββββββββββββββββββββββ
β
ββββββ΄βββββ
β β
Pass Fail
β β
β βΌ
β ββββββββββββββββ
β β Revise Query ββββΊ Retry (max 2x)
β ββββββββββββββββ
β
βΌ
βββββββββββββββββββββββ
β Confidence Score β
β (Weighted Average) β
ββββββββββ¬βββββββββββββ
β
βΌ
βββββββββββββββββββββββ
β Final Response β
β Answer + Scores + β
β Query History β
βββββββββββββββββββββββ
ChromaDB collections throw NotFoundError if the retriever is initialized globally and a new PDF is uploaded mid-session (the collection reference goes stale). By instantiating the client and retriever fresh on every call, this is completely avoided β no restart needed between uploads.
A single "quality score" prompt would conflate three different failure modes. Splitting into relevance, grounding, and completeness means:
- Each grader has a tightly focused prompt β more accurate scoring
- You can diagnose exactly what went wrong (bad retrieval vs hallucination vs incomplete answer)
- Weights can be tuned independently per use case
LangChain is a linear pipeline. LangGraph supports cycles and conditional edges, which is exactly what a correction loop needs. The graph structure makes the retry logic explicit, inspectable, and easy to extend.
Using OpenRouter as the LLM backend makes the system model-agnostic β swapping from GPT-4o-mini to Claude or Mistral is a one-line config change. No vendor lock-in.
Local embedding model β no extra API cost, no latency from an external call, and performant enough for document-level semantic search. Keeps the system usable on a free API budget.
Unlimited retries would create infinite loops on genuinely unanswerable questions. Capping at 2 retries and always returning a scored answer (even a low-confidence one) keeps the UX predictable. The confidence score tells the user when to trust the answer and when to verify externally.
self_correcting_rag/
β
βββ .env # π API keys (never commit this)
βββ .gitignore
βββ requirements.txt
βββ README.md
β
βββ app/
β βββ config.py # LLM + Embeddings setup (OpenRouter + HuggingFace)
β βββ state.py # LangGraph shared state (Pydantic model)
β β
β βββ ingestion/
β β βββ loader.py # PDF loader (PyPDF)
β β βββ splitter.py # Text chunker (RecursiveCharacterTextSplitter)
β β βββ vectorstore.py # ChromaDB vector store (lazy loading)
β β
β βββ retrieval/
β β βββ retriever.py # Similarity search retriever (lazy loading)
β β
β βββ generation/
β β βββ generator.py # LLM answer generation from context
β β
β βββ grading/
β β βββ relevance.py # Are retrieved docs relevant?
β β βββ grounding.py # Is answer grounded in docs?
β β βββ completeness.py # Does answer fully address the question?
β β
β βββ confidence/
β β βββ confidence_scorer.py # Weighted confidence score
β β
β βββ memory/
β β βββ memory.py # Conversation history tracking
β β
β βββ graph/
β β βββ workflow.py # LangGraph agent loop (CORE)
β β
β βββ ui/
β βββ streamlit_app.py # Streamlit frontend
β
βββ data/
βββ chroma_db/ # Persistent vector database (auto-generated)
βββ temp.pdf # Temporary uploaded file (auto-generated)
git clone https://github.com/your-username/self-correcting-rag-langgraph
cd self-correcting-rag-langgraph
python -m venv .venv
source .venv/bin/activate # Mac/Linux
.venv\Scripts\activate # Windowspip install -r requirements.txtCreate a .env file:
OPENROUTER_API_KEY=your_openrouter_key_here
Get your free key at openrouter.ai
streamlit run app/ui/streamlit_app.py --server.fileWatcherType noneOpen http://localhost:8501
langchain
langgraph
langchain-community
langchain-openai
langchain-text-splitters
chromadb
sentence-transformers
faiss-cpu
pydantic
streamlit
pypdf
numpy
torch
transformers
openai
python-dotenv
| Score | π’ High (β₯0.85) | π‘ Medium (0.65β0.85) | π΄ Low (<0.65) |
|---|---|---|---|
| Relevance | Retrieved docs are on-topic | Partially relevant | Off-topic retrieval |
| Grounding | No hallucination | Some unsupported claims | Hallucinated answer |
| Completeness | Fully answered | Partial answer | Incomplete/refused |
| Confidence | Trust the answer | Use with caution | Verify externally |
| Test | Question | Expected Behaviour |
|---|---|---|
| β Grounding | "What was the exact misdiagnosis reduction rate?" | High grounding, correct number cited |
| π― Hallucination trap | "What did GPT-5 achieve in medical exams?" | Low confidence, refusal to fabricate |
| π Completeness | "How much does AI reduce drug timelines specifically?" | Low completeness flagged |
| π Relevance | "What % of US equity trading is algorithmic?" | Low relevance, query rewritten |
| π Contradiction | "What is the consensus AI accuracy for pneumonia?" | Flags conflicting information |
| Error | Fix |
|---|---|
ModuleNotFoundError: langchain.text_splitter |
pip install langchain-text-splitters |
chromadb.errors.NotFoundError |
Delete data/chroma_db/* and re-upload PDF |
RuntimeError: no running event loop |
Add --server.fileWatcherType none to run command |
| Duplicate chunks retrieved | Clear data/chroma_db/ before re-indexing |
| Feature | Standard RAG | This Project |
|---|---|---|
| Retrieval | β | β |
| Generation | β | β |
| Self-correction loop | β | β |
| Hallucination detection | β | β |
| Multi-criteria grading | β | β |
| Query reformulation | β | β |
| Confidence scoring | β | β |
| Conversation memory | β | β |
| LangGraph agent architecture | β | β |
| Layer | Technology |
|---|---|
| LLM | OpenRouter (GPT-4o-mini) |
| Embeddings | HuggingFace all-MiniLM-L6-v2 |
| Vector Store | ChromaDB |
| Agent Framework | LangGraph |
| RAG Framework | LangChain |
| Frontend | Streamlit |
| PDF Parsing | PyPDF |
Built with LangChain Β· LangGraph Β· ChromaDB Β· HuggingFace Β· OpenRouter Β· Streamlit