A simple playground for developers to try Agentic RAG. Upload documents, then chat with them or search inside them.
RAG means "Retrieval Augmented Generation". We first find the useful parts of the documents, then give those parts to the LLM so it can answer properly.
"Agentic" means the LLM decides by itself what to do. For example, it can search again with better words if the first results are not good.
- Upload PDF documents, with duplicate files blocked automatically by their content
- List and delete uploaded documents
- Search documents by meaning, not just keywords, using vector similarity
- Rerank search results with the chat LLM so the best matches come first
- Chat with the documents through an agent that can search again with better words and gives sources for its answers
- Use the documents from MCP clients like Claude Desktop and Claude Code through a read-only MCP server (list and search)
- Use any LLM or embedding provider supported by LiteLLM, changed from settings only
- Plain web UI for documents, search, chat and status, with no build step
- Health checks for the API and the database
- Uploaded files and database data are kept in the project folder, so cleaning up Docker volumes does not remove them
- Python 3.11+, FastAPI
- PostgreSQL with pgvector, SQLAlchemy, Alembic
- LiteLLM (one interface for many LLM providers), pypdf
- MCP Python SDK for the MCP server
- Web UI with Tabler and plain JavaScript (loaded from a CDN, no build step)
- Docker and Docker Compose
git clone https://github.com/stackblogger/agentic-rag-playground.git
cd agentic-rag-playground
cp .env.example .env # set OPENAI_API_KEY in it
docker compose up -d --buildOpen http://127.0.0.1:8000. The app runs the database migrations by itself when it starts.
- Uploaded files are kept in
data/uploads/and database data indata/postgres/, both inside the project folder, so cleaning up Docker volumes does not remove them. DATABASE_URLandUPLOAD_DIRin.envare ignored in Docker, becausedocker-compose.ymlsets them.- Logs:
docker compose logs -f app. Stop:docker compose down.
Python 3.11+ is needed. Only Postgres runs in Docker.
python3 -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
cp .env.example .env # set OPENAI_API_KEY in it
docker compose up -d db # Postgres with pgvector
alembic upgrade head # create the tables
uvicorn agentic_rag.api.main:app --app-dir src --reloadOpen http://127.0.0.1:8000. Port 8000 is used, so the app container must not be running.
- Documents: upload PDFs, see the list, delete a document. The same file cannot be uploaded twice.
- Search: find the chunks with the closest meaning to a query.
- Chat: ask questions and get answers from the documents, with sources.
- Status: check if the API and the database are working.
curl -X POST http://127.0.0.1:8000/documents -F "file=@/path/to/file.pdf"
curl -X POST http://127.0.0.1:8000/chat -H "Content-Type: application/json" \
-d '{"question": "What is this document about?"}'All routes, responses and error codes are in docs/api.md. API docs are also at http://127.0.0.1:8000/docs.
MCP clients like Claude Desktop, Claude Code and Cursor can search and list the uploaded documents through a read-only MCP server. It has two tools, list_documents and search_documents, and runs on the same machine as the client.
PYTHONPATH=src python -m agentic_rag.mcp_serverPostgres must be running. Client setup examples and the tool details are in docs/mcp.md.
The MCP Inspector is a browser page to call the tools by hand. Run this from the project folder:
npx @modelcontextprotocol/inspector -e PYTHONPATH=$PWD/src .venv/bin/python -m agentic_rag.mcp_serverSettings are read from environment variables or the .env file (see .env.example).
| Name | What it is | Default |
|---|---|---|
OPENAI_API_KEY |
API key for OpenAI, read by LiteLLM | empty |
LLM_MODEL |
Model for chat (any LiteLLM model that supports tool calling) | gpt-4o-mini |
EMBEDDING_MODEL |
Model for embeddings | text-embedding-3-small |
DATABASE_URL |
Postgres connection string | Sample in .env.example file |
UPLOAD_DIR |
Folder where uploaded files are kept | data/uploads |
CHUNK_SIZE |
Maximum characters in one chunk | 1000 |
CHUNK_OVERLAP |
Characters shared by two chunks next to each other | 200 |
AGENT_MAX_STEPS |
Maximum searches the chat agent can do before it must answer | 3 |
RERANK_ENABLED |
Turns reranking of search results by the chat LLM on or off | true |
LOG_LEVEL |
How much the app logs: DEBUG, INFO, WARNING or ERROR (DEBUG also shows search queries) |
INFO |
pip install -r requirements-dev.txt
pytestUnit tests need nothing. Integration tests need Postgres running (docker compose up -d db) and use their own agentic_rag_test database, so real data is not touched. LLM and embedding calls are fake in all tests.
- Website: the same docs with a demo, built from the
docs/folder and deployed by GitHub Actions. - Architecture: how the app is built and how upload, search and chat work.
- API: all routes with examples and error codes.
- Database: tables, the search index, embedding size and migration commands.
- MCP: the MCP server tools and how to connect a client.
Please keep changes small, one thing at a time, so commit history stays easy to read. Open an issue first if the change is big.
MIT. See LICENSE.






