diff --git a/README.md b/README.md
index d44e2e3..e128bc9 100644
--- a/README.md
+++ b/README.md
@@ -1,951 +1,538 @@
# QueryCraft AI
-An intelligent, AI-powered database query assistant that transforms natural language into optimized database queries. QueryCraft leverages large language models (LLMs) to understand user intent and generate queries across multiple database systems.
+An AI-powered database query assistant. Describe what you want in plain English and QueryCraft generates the query, explains it, and can run it against your data. It supports SQL, MongoDB, Cypher and several other query languages, and can route each request to a local model (Ollama), OpenRouter, or Google Gemini.
[](https://querycraft.hubzero.in)
+## Contents
+
+- [Overview](#overview)
+- [Features](#features)
+- [Architecture](#architecture)
+- [Tech stack](#tech-stack)
+- [Project structure](#project-structure)
+- [Getting started](#getting-started)
+- [Configuration](#configuration)
+- [Running the project](#running-the-project)
+- [Models and routing](#models-and-routing)
+- [Data sources](#data-sources)
+- [API reference](#api-reference)
+- [Security notes](#security-notes)
+- [Testing](#testing)
+- [Known limitations](#known-limitations)
+- [Roadmap](#roadmap)
+- [Contributing](#contributing)
+- [License](#license)
+
## Overview
-QueryCraft solves the problem of non-technical users struggling to write complex database queries. By accepting natural language input, the system:
+QueryCraft is a two-part application:
-1. Detects the intended query language (SQL, MongoDB, Cypher, GraphQL, CQL, etc.)
-2. Intelligently selects an appropriate LLM based on query complexity
-3. Generates optimized queries with automatic error handling and validation
-4. Maintains conversation context for multi-turn query refinement
-5. Supports direct database execution and result retrieval
+- **Backend** (`querycraft-backend/`): a Node.js/Express API that authenticates users, stores chats in MongoDB, calls the LLM providers, and executes queries against user-supplied data sources.
+- **Frontend** (`querycraft-frontend/`): a Next.js chat interface with a landing page, sign-in, a model picker, syntax-highlighted query cards and a **Run** button that shows results in a table.
-The project consists of a **Node.js/Express backend** that handles LLM orchestration and database connectivity, paired with a **modern Next.js frontend** for an intuitive chat-based user experience.
+When you send a prompt, the backend detects which query language you want, picks or honours a model, sends a guided prompt (plus the recent chat context) to the LLM, and returns the answer in a fixed format: the query in a fenced code block followed by an explanation. **Queries are never run automatically.** You run one explicitly from the query card, against an uploaded file or a connection string you provide.
-## Use Cases
+## Features
-- **Data Analysts**: Quickly explore datasets without writing SQL from scratch
-- **Business Users**: Generate reports by describing data needs in plain English
-- **Developers**: Accelerate database query development and prototyping
-- **DBAs**: Assist in documentation and query validation
-- **Educational Settings**: Learn database query syntax through AI-assisted instruction
-- **Data Science Teams**: Rapidly prototype ETL and data preparation queries
+- **Multiple LLM providers**: local models through Ollama, OpenRouter (DeepSeek R1, Qwen, Grok Code, Mistral and others), and Google Gemini. Choose a model in the UI or let **Auto** decide.
+- **Query language detection**: recognises SQL, MongoDB, Cypher, GraphQL, CQL, Redis, Elasticsearch and DynamoDB from keywords and structure, and steers the LLM toward that language.
+- **Run queries from the chat**: execute the generated query against an uploaded file or a connection string and see rows in a table (see [Data sources](#data-sources)).
+- **File uploads**: query CSV, JSON, SQLite (`.sqlite`, `.db`) or SQL dump (`.sql`) files, up to 100 MB.
+- **Chat history**: chats and queries are stored per user in MongoDB. The last three completed exchanges of a chat are sent back to the model as context.
+- **Authentication**: sign-up and login with bcrypt-hashed passwords and 7-day JWTs.
+- **Rate limiting and security headers**: `express-rate-limit` and Helmet (see [Security notes](#security-notes)).
+- **Landing page with live demo**: an unauthenticated NL-to-query demo endpoint.
+- **UI**: dark-first theme with a light/dark/system preference, voice input (browser speech recognition, where supported), chat management and toast notifications.
-## Key Features
+## Architecture
-- 🤖 **Multi-LLM Support**: Seamlessly integrate Gemini, Mistral, Qwen, Llama, Phi, OpenRouter models
-- 📊 **Multi-Database Support**: SQL, MongoDB, Neo4j, PostgreSQL, MySQL, Cassandra, DynamoDB, Elasticsearch, GraphQL
-- 🔐 **Secure Authentication**: JWT-based user authentication with bcrypt password hashing
-- 💬 **Conversation Management**: Chat-based interface with persistent conversation history
-- 📁 **File Upload Support**: Upload CSV/database files for direct querying
-- 🧠 **Conversation Memory**: Automatic summarization of chat history for context awareness
-- ⚡ **Rate Limiting**: Configurable rate limits protect against abuse
-- 🛡️ **Security Headers**: Helmet.js security headers, XSS protection, CORS support
-- 🎨 **Modern UI**: Responsive design with dark/light theme support, real-time typing indicators
-- 🔍 **Smart Query Language Detection**: Automatically identifies intended query syntax through heuristics
-- 📋 **Response Parsing**: Extracts executable queries from LLM responses with code fence handling
+```mermaid
+flowchart TB
+ subgraph Client["Browser: Next.js frontend"]
+ Intro["IntroPage
landing page + live demo"]
+ Auth["AuthPage
sign in / sign up"]
+ ChatUI["ChatApp
chat, model picker, Run button"]
+ end
-## Architecture / System Design
+ subgraph Server["Express backend (Node.js)"]
+ MW["Middleware
rate limiting, Helmet, CORS, JWT"]
+ Routes["Routes
/api/auth, /api/chat, /api/query, /api/db"]
+ LLM["utils/llm.js
model routing and provider calls"]
+ Exec["controllers/dbController.js
query execution"]
+ MW --> Routes
+ Routes --> LLM
+ Routes --> Exec
+ end
+ Client -->|"HTTP / JSON"| MW
+ LLM --> Ollama["Ollama
local models"]
+ LLM --> OR["OpenRouter"]
+ LLM --> Gemini["Google Gemini"]
+ Routes --> Mongo[("MongoDB
users, chats, queries")]
+ Exec --> Targets["Your data
SQLite / CSV / JSON / SQL files
PostgreSQL, MySQL / MariaDB
MongoDB, Neo4j"]
```
-┌────────────────────────────────────────────────────────────┐
-│ Next.js Frontend │
-│ (React 19, TypeScript, Tailwind CSS, Framer Motion) │
-│ ┌──────────────┬──────────────┬──────────────────────┐ │
-│ │ AuthPage │ IntroPage │ ChatApp │ │
-│ │ (Login/Reg) │ (Landing) │ (Main Interface) │ │
-│ └──────────────┴──────────────┴──────────────────────┘ │
-└────────────────────────────────────────────────────────────┘
- ↕ HTTP/REST
-┌────────────────────────────────────────────────────────────┐
-│ Express Backend (Node.js) │
-│ │
-│ ┌─────────────────────────────────────────────────────┐ │
-│ │ Routes Layer │ │
-│ │ /api/auth │ /api/query │ /api/chat │ /api/db │ │
-│ └─────────────────────────────────────────────────────┘ │
-│ ↓ │
-│ ┌─────────────────────────────────────────────────────┐ │
-│ │ Controllers / Middleware │ │
-│ │ • JWT Authentication • Rate Limiting │ │
-│ │ • Query Language Detection • Response Parsing │ │
-│ │ • Database Query Execution │ │
-│ └─────────────────────────────────────────────────────┘ │
-│ ↓ │
-│ ┌─────────────────────────────────────────────────────┐ │
-│ │ LLM Orchestration │ │
-│ │ • Model Selection Logic (qwen/mistral/llama/...) │ │
-│ │ • OpenRouter API Integration │ │
-│ │ • Google GenAI / Gemini Support │ │
-│ │ • Local LLM Endpoint (Ollama) │ │
-│ │ • Retry Logic & Error Handling │ │
-│ └─────────────────────────────────────────────────────┘ │
-│ ↓ │
-│ ┌─────────────────────────────────────────────────────┐ │
-│ │ Multi-Database Layer │ │
-│ │ ┌──────────────────────────────────────────────┐ │ │
-│ │ │ SQL: SQLite │ PostgreSQL │ MySQL │ │ │
-│ │ ├──────────────────────────────────────────────┤ │ │
-│ │ │ NoSQL: MongoDB │ Neo4j │ Cassandra │ │ │
-│ │ ├──────────────────────────────────────────────┤ │ │
-│ │ │ Search: Elasticsearch │ DynamoDB │ │ │
-│ │ └──────────────────────────────────────────────┘ │ │
-│ └─────────────────────────────────────────────────────┘ │
-│ ↓ │
-│ ┌─────────────────────────────────────────────────────┐ │
-│ │ Data Models (MongoDB) │ │
-│ │ • User • Chat • Query │ │
-│ └─────────────────────────────────────────────────────┘ │
-└────────────────────────────────────────────────────────────┘
- ↕
-┌────────────────────────────────────────────────────────────┐
-│ MongoDB Database │
-│ Stores: Users, Conversations, Query History, Metadata │
-└────────────────────────────────────────────────────────────┘
-```
-### Request Flow
-
-1. **User Query Submission** → Frontend sends natural language prompt + optional database credentials
-2. **Authentication** → JWT middleware validates user session
-3. **Query Language Detection** → Backend analyzes prompt to identify target query language
-4. **Model Selection** → Chooses optimal LLM based on query type and length
-5. **LLM Call** → Routes to appropriate LLM provider (local, OpenRouter, Gemini)
-6. **Response Parsing** → Extracts executable query from LLM response
-7. **Database Execution** → Runs query against specified database
-8. **Result Storage** → Persists query, response, and metadata to MongoDB
-9. **Response to Client** → Returns results with formatting and syntax highlighting
-
-## Tech Stack
-
-### Backend
-- **Runtime**: Node.js 20
-- **Framework**: Express.js 4.18
-- **Database Drivers**:
- - MongoDB: `mongoose` 7.5, `mongodb` 6.21
- - PostgreSQL: `pg` 8.16
- - MySQL: `mysql2` 3.15
- - SQLite: `better-sqlite3` 11.10
- - Neo4j: `neo4j-driver` 6.0
-- **LLM Integration**:
- - `@google/genai` 1.0 (Gemini models)
- - OpenRouter API (via axios)
- - Local Ollama support
-- **Authentication**: `jsonwebtoken` 9.0, `bcryptjs` 2.4
-- **File Handling**: `multer` 2.0, `csv-parser` 3.2
-- **Security**: `helmet` 8.1, `express-rate-limit` 8.0, `xss` 1.0
-- **Utilities**: `cors`, `morgan`, `dotenv`, `uuid`
-- **Development**: `nodemon` 2.0
-
-### Frontend
-- **Framework**: Next.js 15.5
-- **UI Library**: React 19.1
-- **Language**: TypeScript
-- **Styling**:
- - Tailwind CSS 4.1
- - Radix UI components (Dialog, Select, Tabs, Dropdown, etc.)
- - `framer-motion` 12.23 (animations)
-- **Code Display**: `prismjs` 1.30
-- **3D Graphics**: `three.js` 0.180 + `@react-three/fiber` 9.3
-- **Notifications**: `sonner` 2.0
-- **Theme**: `next-themes` 0.4
-- **Utilities**: `tailwind-merge`, `lucide-react` (icons)
-- **Development**: ESLint, TypeScript, Tailwind Autoprefixer
-
-## Project Structure
+MongoDB holds application data only (users, chats, queries). The databases you query are separate and are reached only through `/api/db/execute`.
+
+### Request flow
+
+1. **Prompt**: the frontend sends `POST /api/query` with the prompt, optional `chatId` and optional `model`, plus the JWT.
+2. **Sanitising**: the prompt has null bytes removed, whitespace collapsed and length capped at 4,000 characters. Triple backticks are neutralised before the prompt is embedded in the LLM instructions.
+3. **Chat lookup**: the chat is loaded (and must belong to the user), or a new one is created titled with the first 50 characters of the prompt. The last three completed exchanges (each prompt and the query it produced) are added as context, capped at 1,800 characters.
+4. **Language detection**: keywords such as "Cypher", "Neo4j", "GraphQL" or "Cassandra", or the structure of a pasted query, select the target language. If nothing matches, the LLM chooses between SQL and MongoDB (SQL by default).
+5. **Model selection**: an explicit model is used as given; `auto` (or no model) applies the routing heuristics in [Models and routing](#models-and-routing).
+6. **LLM call**: the query is saved as `pending`, the provider is called, and the record becomes `done` or `failed`.
+7. **Output formatting**: the answer is normalised to `Here's the query`, then the query in a fenced block, then an explanation.
+8. **Run (optional)**: the user clicks **Run** on the query card, which calls `POST /api/db/execute` with an uploaded file ID or a saved connection string.
+
+## Tech stack
+
+**Backend**
+
+- Node.js 20, Express 4
+- MongoDB with Mongoose 7 for application data
+- Query execution drivers: `better-sqlite3`, `pg`, `mysql2`, `mongodb`, `neo4j-driver`, `csv-parser`
+- LLM access: `@google/genai` (Gemini), `axios` (OpenRouter and Ollama)
+- Auth and hardening: `jsonwebtoken`, `bcryptjs`, `helmet`, `express-rate-limit`, `cors`
+- Uploads and logging: `multer`, `morgan`, `response-time`
+- Dev tooling: `nodemon`
+
+**Frontend**
+
+- Next.js 15 (App Router), React 19, TypeScript 5
+- Tailwind CSS 4, Radix UI primitives, `class-variance-authority`, `lucide-react`
+- `framer-motion` (animation), `three` with `@react-three/fiber` and `@react-three/drei` (3D landing-page hero)
+- `prismjs` (syntax highlighting), `sonner` (toasts), `next-themes` (theming)
+- ESLint 9
+
+## Project structure
```
-QueryCraft AI/
+QueryCraft-AI/
+├── .github/workflows/node.js.yml # GitHub Actions workflow
+├── LICENSE
├── README.md
-├── .gitignore
│
-├── querycraft-backend/ # Express.js Backend
-│ ├── index.js # Server entry point, middleware setup
-│ ├── package.json # Dependencies
-│ ├── Dockerfile # Docker configuration
-│ ├── db_files.json # File upload metadata store
-│ │
+├── querycraft-backend/
+│ ├── index.js # App setup, middleware, routes, Mongo connection, health check
+│ ├── Dockerfile
+│ ├── package.json
│ ├── controllers/
-│ │ └── dbController.js # Multi-database query execution logic
-│ │
+│ │ └── dbController.js # File upload metadata + query execution for all data sources
│ ├── middleware/
-│ │ └── auth.js # JWT authentication middleware
-│ │
-│ ├── models/ # MongoDB schemas
-│ │ ├── User.js # User model with bcrypt hashing
-│ │ ├── Chat.js # Chat session model
-│ │ └── Query.js # Query history model
-│ │
-│ ├── routes/ # API endpoints
-│ │ ├── auth.js # Signup, login, profile endpoints
-│ │ ├── query.js # Query language detection & LLM calls
-│ │ ├── chat.js # Chat management endpoints
-│ │ └── db.js # File upload & database execution
-│ │
+│ │ └── auth.js # JWT verification
+│ ├── models/ # Mongoose schemas: User, Chat, Query
+│ ├── routes/
+│ │ ├── auth.js # /api/auth
+│ │ ├── chat.js # /api/chat
+│ │ ├── query.js # /api/query: language detection, model routing, prompt building
+│ │ └── db.js # /api/db: upload and execute
│ └── utils/
-│ ├── llm.js # LLM provider orchestration
-│ ├── responseParser.js # LLM response parsing & SQL extraction
-│ ├── conversationMemory.js # Chat history summarization
-│ └── uploads/ # Temporary file storage
+│ ├── llm.js # Provider calls (Gemini, OpenRouter, Ollama) and model aliases
+│ ├── responseParser.js # Extracts text/SQL from provider responses
+│ └── conversationMemory.js # Chat summariser (not wired into any route yet)
│
-└── querycraft-frontend/ # Next.js Frontend
+└── querycraft-frontend/
├── package.json
- ├── tsconfig.json
- ├── next.config.ts
- ├── tailwind.config.ts
- ├── postcss.config.mjs
- ├── eslint.config.mjs
- │
- ├── public/
- │ └── assets/ # Static assets
- │
+ ├── next.config.ts, tailwind.config.ts, postcss.config.mjs, eslint.config.mjs, tsconfig.json
+ ├── public/ # Static assets
└── src/
- ├── app/
- │ ├── layout.tsx # Root layout
- │ ├── page.tsx # Main page (view router)
- │ └── globals.css # Global styles
- │
+ ├── app/ # layout.tsx, page.tsx (view router: intro / auth / chat), globals.css
├── components/
- │ ├── auth/
- │ │ └── AuthProviderClient.tsx # Auth context provider
- │ │
- │ ├── chat/ # Chat interface components
- │ │ ├── ChatWindow.tsx # Message display
- │ │ ├── ChatInput.tsx # User input area
- │ │ ├── ChatMessage.tsx # Message rendering
- │ │ ├── ChatHeader.tsx # Chat header UI
- │ │ ├── CodeCard.tsx # Code block display
- │ │ ├── TypingIndicator.tsx # AI typing animation
- │ │
- │ ├── layout/ # Layout components
- │ │ ├── Sidebar.tsx # Chat list sidebar
- │ │ ├── ChatSidebar.tsx # Chat-specific sidebar
- │ │ ├── Header.tsx # Top navigation
- │ │ ├── MobileSidebarTrigger.tsx # Mobile menu
- │ │
- │ ├── modals/ # Dialog components
- │ │ ├── DatabaseImportDialog.tsx # DB connection setup
- │ │ ├── SettingsDialog.tsx # User preferences
- │ │
- │ ├── pages/ # Page containers
- │ │ ├── IntroPage.tsx # Landing page
- │ │ ├── AuthPage.tsx # Login/signup
- │ │ ├── ChatApp.tsx # Main chat interface
- │ │
- │ └── ui/ # Radix UI primitives
- │ ├── button.tsx, card.tsx, dialog.tsx
- │ ├── dropdown-menu.tsx, select.tsx
- │ ├── alert.tsx, badge.tsx, tabs.tsx
- │ ├── input.tsx, textarea.tsx, label.tsx
- │ ├── tooltip.tsx, separator.tsx
- │ └── [more UI components]
- │
- ├── hooks/
- │ └── useAutoLogin.ts # Auto-login logic
- │
- ├── lib/
- │ ├── api.ts # API utilities (empty in current version)
- │ └── utils.ts # Helper utilities
- │
- └── types/
- ├── prismjs.d.ts # Prism.js type definitions
- └── react-three.d.ts # Three.js type definitions
+ │ ├── auth/ # AuthProviderClient (auth context, auto-login)
+ │ ├── chat/ # ChatWindow, ChatInput, ChatMessage, ChatHeader (model picker), CodeCard, TypingIndicator
+ │ ├── layout/ # Sidebar, ChatSidebar, Header, MobileSidebarTrigger
+ │ ├── modals/ # DatabaseImportDialog, SettingsDialog
+ │ ├── pages/ # IntroPage, AuthPage, ChatApp
+ │ └── ui/ # Radix-based UI primitives
+ ├── hooks/ # useAutoLogin
+ ├── lib/ # utils.ts (class-name helper)
+ └── types/ # Type declarations for prismjs and react-three
```
-### Key Files Explained
-
-| File | Purpose |
-|------|---------|
-| `querycraft-backend/index.js` | Express app setup, MongoDB connection, middleware initialization, route registration |
-| `querycraft-backend/controllers/dbController.js` | Multi-database query execution (500+ lines), supports SQLite, PostgreSQL, MySQL, MongoDB, Neo4j with HTTP/Bolt fallback |
-| `querycraft-backend/utils/llm.js` | LLM provider routing (Gemini, OpenRouter, Ollama), model normalization, retry logic, prompt optimization |
-| `querycraft-backend/utils/responseParser.js` | JSON parsing, code fence extraction, SQL detection from LLM responses |
-| `querycraft-backend/routes/query.js` | Query language heuristics (SQL/Mongo/Cypher/GraphQL detection), model selection logic |
-| `querycraft-frontend/src/app/page.tsx` | App state management, view routing (intro/auth/chat), auto-login orchestration |
-| `querycraft-frontend/src/components/pages/ChatApp.tsx` | Main chat interface, message state, LLM interaction, database operations |
+Two files appear at runtime and are git-ignored: `querycraft-backend/uploads/` (uploaded files) and `querycraft-backend/db_files.json` (metadata index for those uploads).
-## Installation
+## Getting started
### Prerequisites
-- **Node.js 20+** ([Download](https://nodejs.org/))
-- **MongoDB** (local or cloud, e.g., MongoDB Atlas)
-- **Git**
-- (Optional) **Docker** for containerized deployment
-- (Optional) **Ollama** for local LLM execution ([Download](https://ollama.ai/))
+- **Node.js 20 or newer** and npm
+- **MongoDB**, local or hosted (for example MongoDB Atlas)
+- At least one LLM provider:
+ - [Ollama](https://ollama.com/) for local models, and/or
+ - an [OpenRouter](https://openrouter.ai/) API key, and/or
+ - a [Google Gemini](https://ai.google.dev/) API key
+- Optional: Docker
-### Clone Repository
+### Install
```bash
git clone https://github.com/sultanmaliki/QueryCraft-AI.git
cd QueryCraft-AI
-```
-
-### Backend Setup
-
-```bash
-cd querycraft-backend
-# Install dependencies
-npm install
-
-# Create environment file
-cp .env.example .env # (if available, or create manually)
+cd querycraft-backend && npm install
+cd ../querycraft-frontend && npm install
```
-### Frontend Setup
+## Configuration
-```bash
-cd ../querycraft-frontend
+There are no `.env.example` files in the repository, so create the two env files yourself.
-# Install dependencies
-npm install
-```
+### Backend: `querycraft-backend/.env`
-## Configuration
+`dotenv` reads this file from the directory you start the server in, so run the backend from `querycraft-backend/`.
-### Backend Environment Variables
+| Variable | Default | Purpose |
+|----------|---------|---------|
+| `MONGO_URI` | `mongodb://localhost:27017/querycraft` | MongoDB connection string. The server exits if it cannot connect. |
+| `PORT` | `5001` | HTTP port. Note that `npm start` forces `PORT=5001`. |
+| `JWT_SECRET` | `change_this_in_production` | Secret used to sign and verify JWTs. **Always set your own.** |
+| `LLM_ENDPOINT` | `http://127.0.0.1:11434/api/generate` | Ollama-style endpoint for local models. |
+| `DEFAULT_MODEL` | `mistral:7b-instruct` | Local model tag used when a request names no model or an unknown one. Sent as-is to `LLM_ENDPOINT`. |
+| `GENAI_KEY` or `GOOGLE_GENAI_KEY` | none | Google Gemini API key. Required for Gemini models. |
+| `OPENROUTER_KEY` | none | OpenRouter API key. Required for `or-*` models and the `mistral` aliases. |
+| `OPENROUTER_SITE_URL` | `https://localhost` | Sent as the `HTTP-Referer` header to OpenRouter. |
+| `OPENROUTER_APP_NAME` | `QueryCraft` | Sent as the `X-Title` header to OpenRouter. |
+| `GEMINI_RETRIES` | `3` | Total attempts for transient Gemini errors. |
+| `GEMINI_BASE_DELAY_MS` | `500` | Initial backoff delay between Gemini retries. |
+| `GEMINI_MAX_BACKOFF_MS` | `5000` | Maximum backoff delay. |
+| `SUMMARY_MODEL`, `SUMMARY_OLDEST_COUNT`, `SUMMARY_MAX_TOKENS` | `mistral:7b-instruct`, `15`, `400` | Only read by `utils/conversationMemory.js`, which no route calls yet. |
-Create a `.env` file in `querycraft-backend/`:
+Minimal example:
```bash
-# MongoDB Connection (required)
MONGO_URI=mongodb://localhost:27017/querycraft
+JWT_SECRET=replace-with-a-long-random-string
+GENAI_KEY=your-gemini-key
+OPENROUTER_KEY=your-openrouter-key
+```
-# Server Configuration
-PORT=5001
-NODE_ENV=development
-
-# Authentication
-JWT_SECRET=your-secret-key-change-in-production
-TOKEN_EXPIRY=7d
-
-# LLM Configuration - Choose one or multiple providers
-
-# 1. Local Ollama (default)
-LLM_ENDPOINT=http://127.0.0.1:11434/api/generate
-DEFAULT_MODEL=mistral:7b-instruct
-
-# 2. Google Gemini
-GENAI_KEY=your-google-genai-api-key
-# or
-GOOGLE_GENAI_KEY=your-google-genai-api-key
-
-# 3. OpenRouter (for specialized models)
-OPENROUTER_KEY=your-openrouter-api-key
-OPENROUTER_SITE_URL=https://yourapp.com
-OPENROUTER_APP_NAME=QueryCraft
-
-# Feature Configuration
-SUMMARY_MODEL=mistral:7b-instruct # Model for conversation summarization
-SUMMARY_OLDEST_COUNT=15 # Number of messages to summarize
-SUMMARY_MAX_TOKENS=400 # Max tokens for summary
+Generate a strong secret with:
-# (Optional) Database Test Credentials - for demo purposes
-TEST_SQLITE_PATH=/tmp/test.db
-TEST_POSTGRES_URI=
-TEST_MYSQL_URI=
-TEST_MONGODB_URI=
+```bash
+node -e "console.log(require('crypto').randomBytes(32).toString('hex'))"
```
-### Frontend Environment Variables
-
-Create a `.env.local` file in `querycraft-frontend/`:
+### Frontend: `querycraft-frontend/.env.local`
```bash
-# API Base URL
NEXT_PUBLIC_API_BASE=http://localhost:5001
-# or production: https://api.querycraft.ai
-
-# Optional: Analytics, feature flags, etc.
```
-### Optional: Docker Setup
+Always set this. A few components fall back to the hosted API (`https://apiquerycraft.hubzero.in`) when it is missing, and the sign-in flow does not, so leaving it unset gives a confusing mix of local and hosted behaviour.
-Backend Dockerfile is provided. Build and run:
+### Local models with Ollama
-```bash
-cd querycraft-backend
-docker build -t querycraft-backend .
-docker run -p 5001:5001 --env-file .env querycraft-backend
-```
+The backend sends these tags to Ollama exactly as written, so they must exist in your local Ollama install: `qwen:4b`, `llama3.2:1b` and `phi3:mini-4k-instruct` (used by the **Auto** router and the demo), and `mistral:7b-instruct` (the default). Pull the ones you need with `ollama pull `, or alias an equivalent with `ollama cp`.
-## Running the Project
+## Running the project
-### Development Mode
+### Development
-#### Terminal 1: Backend
+Backend (port 5001):
```bash
cd querycraft-backend
npm run dev
-# Output: Server running on port 5001
```
-#### Terminal 2: Frontend
+Frontend (port 3000):
```bash
cd querycraft-frontend
npm run dev
-# Output: ▲ Next.js 15.5.7
-# - Local: http://localhost:3000
```
-Visit **http://localhost:3000** in your browser.
+Then open .
-### Production Mode
+> **Windows:** the npm scripts use POSIX-style inline variables (`NODE_ENV=development nodemon index.js`), which fail in `cmd` and PowerShell. Run `npx nodemon index.js` (or `node index.js`) directly, or use Git Bash or WSL.
-#### Backend
+### Production
```bash
+# Backend
cd querycraft-backend
npm start
-# PORT=5001 NODE_ENV=production node index.js
-```
-
-#### Frontend
-```bash
+# Frontend
cd querycraft-frontend
npm run build
npm start
-# ▲ Next.js 15.5.7 (Standalone)
-# ▲ Ready on http://0.0.0.0:3000
```
-## API Documentation
-
-### Authentication Endpoints
-
-#### `POST /api/auth/signup`
+### Docker (backend only)
-Register a new user.
-
-**Request:**
-```json
-{
- "name": "John Doe",
- "email": "john@example.com",
- "password": "securePassword123"
-}
-```
-
-**Response:**
-```json
-{
- "user": {
- "_id": "507f1f77bcf86cd799439011",
- "name": "John Doe",
- "email": "john@example.com",
- "role": "user",
- "createdAt": "2025-03-15T10:30:00Z"
- },
- "token": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9..."
-}
+```bash
+cd querycraft-backend
+docker build -t querycraft-backend .
+docker run -p 5001:5001 --env-file .env querycraft-backend
```
-#### `POST /api/auth/login`
+Things to know:
-Authenticate and receive JWT token.
+- From inside a container, `localhost` is the container. Point `MONGO_URI` and `LLM_ENDPOINT` at addresses the container can reach, for example `host.docker.internal`.
+- There is no `.dockerignore`, and the Dockerfile runs `COPY . .` after `npm install`. Build from a checkout without a local `node_modules`, or a host-built `better-sqlite3` binary can overwrite the container's.
+- Uploaded files and `db_files.json` live inside the container filesystem and are lost when the container is removed unless you mount volumes for them.
-**Request:**
-```json
-{
- "email": "john@example.com",
- "password": "securePassword123"
-}
-```
+## Models and routing
-**Response:**
-```json
-{
- "user": { /* user object */ },
- "token": "eyJhbGciOiJIUzI1NiIsInR5cCI6IkpXVCJ9..."
-}
-```
+The model picker in the UI offers **Auto**, Qwen 4B, Gemini, Phi-3 Mini, Llama 1B, Mistral 7B, DeepSeek R1, Grok Code Fast, Qwen3 235B A22B and Qwen3 Coder. The backend only accepts known aliases; an unknown model name falls back to `DEFAULT_MODEL`.
-#### `GET /api/auth/me`
+| Alias(es) | Resolves to | Provider |
+|-----------|-------------|----------|
+| `qwen4`, `qwen:4b` | `qwen:4b` | Ollama (local) |
+| `llama1b`, `llama3.2:1b` | `llama3.2:1b` | Ollama (local) |
+| `phi3`, `phi3-mini`, `phi3-mini-4k-instruct`, `phi3:mini-4k-instruct` | `phi3:mini-4k-instruct` | Ollama (local) |
+| `gemini`, `gemini-2.5-flash` | `gemini-2.5-flash` | Google Gemini |
+| `mistral`, `mistral:7b`, `mistral:7b-instruct` | `mistralai/mistral-7b-instruct:free` | OpenRouter |
+| `or-deepseek-r1` | `deepseek/deepseek-r1` | OpenRouter |
+| `or-qwen2.5-72b-free` | `qwen/qwen-2.5-72b-instruct:free` | OpenRouter |
+| `or-qwen3-235b-a22b` | `qwen/qwen3-235b-a22b:free` | OpenRouter |
+| `or-qwen3-coder` | `qwen/qwen3-coder:free` | OpenRouter |
+| `or-grok-code-fast` | `x-ai/grok-code-fast-1` | OpenRouter |
-Get current user profile (requires authentication).
+**Auto routing** (`model: "auto"` or no model) is a set of heuristics in `routes/query.js`:
-**Headers:**
-```
-Authorization: Bearer
-```
+- The prompt already looks like a query (under 2,000 characters): local Phi-3 Mini.
+- Vector or semantic search wording: DeepSeek R1.
+- MongoDB or SQL requests: local Qwen 4B.
+- Code-like prompts: Grok Code Fast.
+- Long or explanation-heavy prompts, and anything else: Gemini 2.5 Flash.
-**Response:**
-```json
-{
- "user": {
- "_id": "507f1f77bcf86cd799439011",
- "name": "John Doe",
- "email": "john@example.com",
- "role": "user"
- }
-}
-```
+Defaults: `max_tokens` 512 and `temperature` 0.2 (the demo endpoint uses 256 and 0.0). Retries with exponential backoff and jitter apply to Gemini calls only. OpenRouter calls time out after 30 seconds and local calls after 120 seconds.
----
+## Data sources
-### Chat Endpoints
+QueryCraft can generate queries for more languages than it can run.
-All chat endpoints require JWT authentication via `Authorization: Bearer ` header.
+| Data source | Generate | Run from the UI |
+|-------------|:--------:|:---------------:|
+| SQL (generic ANSI-style) | Yes | Yes |
+| SQLite / `.db` files | Yes | Yes (opened read-only) |
+| CSV, JSON, SQL dump uploads | Yes | Yes (loaded into a temporary SQLite database) |
+| PostgreSQL (`postgres://`, `postgresql://`) | Yes | Yes (10 s statement timeout) |
+| MySQL / MariaDB (`mysql://`, `mariadb://`) | Yes | Yes |
+| MongoDB (`mongodb://`, `mongodb+srv://`) | Yes | Yes (`find` with a filter only) |
+| Neo4j / Cypher (`neo4j://`, `bolt://`, `http(s)://`) | Yes | Yes (Bolt, with HTTP fallback) |
+| Cassandra CQL, Redis, Elasticsearch, DynamoDB, GraphQL | Yes | No |
-#### `GET /api/chat`
+How uploaded files are loaded:
-List all chats for the authenticated user.
+- `.sqlite` and `.db` files are opened directly and read-only.
+- `.csv` files become a table named `imported_csv` (all columns are `TEXT`).
+- `.json` files (an array of objects, or a single object) become a table named `imported_json`.
+- `.sql` files are executed into a temporary SQLite database, so the table names come from your script.
+- Files with other extensions are sniffed: content starting with `[` is treated as JSON, anything else as CSV.
-**Response:**
-```json
-[
- {
- "_id": "507f1f77bcf86cd799439012",
- "user": "507f1f77bcf86cd799439011",
- "title": "SQL Analysis Session",
- "createdAt": "2025-03-15T10:00:00Z",
- "updatedAt": "2025-03-15T11:30:00Z"
- }
-]
-```
+Connection strings are saved in your browser's `localStorage` (`qc_conn_default`) and sent to the backend each time you press **Run**.
-#### `POST /api/chat`
+## API reference
-Create a new chat session.
+Base URL: `http://localhost:5001`. Protected endpoints need `Authorization: Bearer `. Most errors have the form `{ "error": "message" }`, and some also include a `message` field with details.
-**Request:**
-```json
-{
- "title": "Customer Query Analysis"
-}
-```
+### Rate limits
-**Response:**
-```json
-{
- "_id": "507f1f77bcf86cd799439013",
- "user": "507f1f77bcf86cd799439011",
- "title": "Customer Query Analysis",
- "createdAt": "2025-03-15T11:45:00Z"
-}
-```
+| Scope | Limit |
+|-------|-------|
+| All endpoints, per IP | 100 requests per minute (`/api/query/demo` is exempt from this one) |
+| `POST /api/query` and `POST /api/query/demo`, per IP | 10 requests per minute |
-#### `GET /api/chat/:id`
+Exceeding a limit returns `429`, with standard `RateLimit-*` headers.
-Get a specific chat with all its queries.
+### Health
-**Response:**
-```json
-{
- "chat": {
- "_id": "507f1f77bcf86cd799439012",
- "user": "507f1f77bcf86cd799439011",
- "title": "SQL Analysis Session",
- "createdAt": "2025-03-15T10:00:00Z"
- },
- "queries": [
- {
- "_id": "507f1f77bcf86cd799439014",
- "user": "507f1f77bcf86cd799439011",
- "chat": "507f1f77bcf86cd799439012",
- "prompt": "Show me all active customers",
- "response": "SELECT * FROM customers WHERE status='active'",
- "model": "mistral:7b-instruct",
- "status": "done",
- "createdAt": "2025-03-15T10:05:00Z"
- }
- ]
-}
-```
+`GET /` returns `{ "status": "QueryCraft backend is up", "mongo": "connected" }` (`mongo` is `disconnected` when the database is unreachable).
-#### `DELETE /api/chat/:id`
+### Auth
-Delete a specific chat and all associated queries.
+**`POST /api/auth/signup`**: body `{ "name", "email", "password" }`. Returns `201` with `{ "user", "token" }`. Errors: `400` if a field is missing, `409` if the email is taken. Emails are lower-cased and trimmed, and no password strength rules are enforced.
-**Response:**
-```json
-{
- "success": true
-}
-```
+**`POST /api/auth/login`**: body `{ "email", "password" }`. Returns `{ "user", "token" }`. Errors: `400` missing fields, `401` invalid credentials.
-#### `DELETE /api/chat`
+**`GET /api/auth/me`** (auth): returns `{ "user": { "_id", "name", "email", "role", "createdAt" } }`.
-Delete all chats for the authenticated user.
+Tokens expire after 7 days. That value is fixed in `routes/auth.js` and is not configurable through the environment.
-**Response:**
-```json
-{
- "success": true
-}
-```
+### Chats (auth)
----
+| Endpoint | Description |
+|----------|-------------|
+| `GET /api/chat` | The user's chats, most recently updated first. |
+| `POST /api/chat` | Create a chat. Body `{ "title"? }` (default `"New Chat"`). Returns `201` with the chat. |
+| `GET /api/chat/:id` | `{ "chat", "queries" }` with queries oldest first. `404` if the chat is not the user's. |
+| `DELETE /api/chat/:id` | Delete a chat and its queries. Returns `{ "success": true }`. |
+| `DELETE /api/chat` | Delete all of the user's chats and queries. Returns `{ "success": true }`. |
-### Query Endpoints
+A query record has `_id`, `user`, `chat`, `prompt`, `response`, `model`, `status` (`pending`, `done` or `failed`), `usage`, `raw` and `createdAt`.
-#### `POST /api/query`
+### Generate a query
-Generate and execute a query from natural language.
+**`POST /api/query`** (auth, 10/min)
-**Request:**
```json
{
- "prompt": "Get all orders from last month with total amount greater than 1000",
"chatId": "507f1f77bcf86cd799439012",
- "model": "or-qwen2.5-72b-free",
- "sourceType": "connection",
- "connectionString": "postgres://user:pass@localhost/mydb"
+ "prompt": "Get all orders from last month with total greater than 1000",
+ "model": "auto",
+ "max_tokens": 512,
+ "temperature": 0.2
}
```
-**Response:**
-```json
+Only `prompt` is required. Without `chatId`, a new chat is created. Response (the `response` string contains a fenced code block, so it is shown here in a four-backtick fence):
+
+````json
{
"queryId": "507f1f77bcf86cd799439015",
- "response": "SELECT * FROM orders WHERE created_at >= NOW() - INTERVAL '1 month' AND total > 1000",
- "model": "or-qwen2.5-72b-free",
+ "chatId": "507f1f77bcf86cd799439012",
+ "model": "qwen:4b",
"status": "done",
- "usage": {
- "prompt_tokens": 45,
- "completion_tokens": 28
- },
- "raw": { /* LLM raw response */ }
+ "createdAt": "2026-03-15T10:05:00.000Z",
+ "updatedAt": "2026-03-15T10:05:02.000Z",
+ "response": "Here's the query\n\n```sql\nSELECT * FROM orders WHERE ...;\n```\n\nExplanation: ..."
}
-```
+````
----
+Errors: `400` empty prompt, `404` chat not found, `500` `{ "error": "LLM request failed", "message": "..." }` (the stored query is marked `failed`).
-### Database Execution Endpoints
+**`POST /api/query/demo`** (no auth, 10/min): body `{ "prompt", "model"?, "max_tokens"?, "temperature"? }`. Returns `{ "status": "ok", "response": "" }`. If the prompt is too ambiguous, `response` is a message asking for clarification instead. Nothing is stored.
-#### `POST /api/db/upload`
+### Upload and execute (no authentication; see [Security notes](#security-notes))
-Upload a CSV or database file for querying.
+**`POST /api/db/upload`**: `multipart/form-data` with a `file` field (max 100 MB).
-**Request:**
-```
-Content-Type: multipart/form-data
-file: [binary file data]
-```
-
-**Response:**
```json
{
"success": true,
"file": {
- "id": "1678876200000-550e8400-e29b-41d4-a716-446655440000-data.csv",
+ "id": "0b2f5c4e-9d1a-4f5e-8a44-5c6f0f1d7a21",
"originalName": "data.csv",
- "path": "/app/uploads/1678876200000-550e8400-e29b-41d4-a716-446655440000-data.csv",
- "uploadedAt": "2025-03-15T12:00:00Z"
+ "path": "/app/uploads/1742032800000-0b2f5c4e-...-data.csv",
+ "size": 20480,
+ "uploadedAt": "2026-03-15T12:00:00.000Z",
+ "ext": ".csv",
+ "type": "csv"
}
}
```
-#### `POST /api/db/execute`
+**`POST /api/db/execute`**: run a query against an uploaded file or a connection string.
-Execute a query against an uploaded file or database connection.
-
-**Request:**
```json
{
"sourceType": "file",
- "fileId": "1678876200000-550e8400-e29b-41d4-a716-446655440000-data.csv",
- "query": "SELECT * FROM data WHERE age > 30",
+ "fileId": "0b2f5c4e-9d1a-4f5e-8a44-5c6f0f1d7a21",
+ "query": "SELECT * FROM imported_csv WHERE CAST(age AS INTEGER) > 30",
"maxRows": 100
}
```
-or
-
```json
{
"sourceType": "connection",
"connectionString": "postgresql://user:pass@localhost:5432/mydb",
- "query": "SELECT * FROM orders LIMIT 50",
- "maxRows": 50
+ "query": "SELECT * FROM orders LIMIT 50"
}
```
-**Response:**
+MongoDB takes a structured `mongo` object instead of `query`:
+
```json
{
- "source": "file",
- "columns": ["id", "name", "age", "email"],
- "rows": [
- { "id": 1, "name": "Alice", "age": 35, "email": "alice@example.com" },
- { "id": 2, "name": "Bob", "age": 42, "email": "bob@example.com" }
- ],
- "rowCount": 2
+ "sourceType": "connection",
+ "connectionString": "mongodb://localhost:27017/shop",
+ "mongo": { "collection": "orders", "filter": { "status": "paid" }, "projection": { "total": 1 }, "limit": 50 }
}
```
----
-
-### Error Responses
+For Neo4j, send a Cypher string as `query`. Credentials come from the connection string, or from optional `user`, `password` and `database` fields (default database `neo4j`).
-All endpoints may return error responses:
+Response (`maxRows` defaults to 1000):
```json
{
- "error": "error_code",
- "message": "Human-readable error message"
+ "source": "temp-sqlite-import",
+ "columns": ["id", "name", "age"],
+ "rows": [{ "id": "1", "name": "Alice", "age": "35" }],
+ "rowCount": 1
}
```
-Common status codes:
-- `400 Bad Request` - Invalid input
-- `401 Unauthorized` - Missing or invalid token
-- `404 Not Found` - Resource not found
-- `409 Conflict` - Resource already exists (e.g., email in use)
-- `429 Too Many Requests` - Rate limit exceeded
-- `500 Internal Server Error` - Server error
-
-## Internal Modules / Core Components
-
-### LLM Orchestration (`utils/llm.js`)
-
-Intelligently routes requests to appropriate LLM providers:
-
-- **Model Normalization**: Converts frontend model aliases to actual provider names
- - `mistral` → `openrouter:mistralai/mistral-7b-instruct:free`
- - `gemini` → `gemini-2.5-flash`
- - `or-deepseek-r1` → `openrouter:deepseek/deepseek-r1`
-
-- **Provider Selection**:
- 1. Gemini → Uses Google GenAI SDK with API key
- 2. OpenRouter → HTTP calls with auth header
- 3. Default → Local Ollama endpoint
-
-- **Retry Logic**: Exponential backoff for transient errors (429, 503, timeouts)
-- **Response Extraction**: Handles multiple response formats (JSON, streaming, structured)
-- **Token Limits**: Adaptive `max_tokens` based on query complexity
-
-### Query Language Detection (`routes/query.js`)
+Errors return `400` with `{ "error": "" }`, for example `file_not_found`, `unsupported_connection_type` or a database error message.
-Smart heuristics identify intended query language:
+## Security notes
-- **SQL**: Detects keywords (SELECT, INSERT, UPDATE, DELETE, JOIN, CREATE, ALTER)
-- **MongoDB**: Recognizes `db.collection.find()`, aggregation syntax
-- **Cypher** (Neo4j): Matches MATCH/MERGE patterns with relationship syntax `->`, `-[]-`
-- **GraphQL**: Identifies `query{`, `mutation{`, `subscription{` blocks
-- **Explicit Mentions**: Prioritizes natural language hints ("use Cypher", "write MongoDB")
-- **Fallback**: Returns null if ambiguous, lets LLM infer from context
+What the project does:
-### Response Parsing (`utils/responseParser.js`)
+- Passwords are hashed with bcrypt (10 rounds) and never returned by the API.
+- JWTs are verified on `/api/chat`, `POST /api/query` and `/api/auth/me`, and chats are always scoped to the authenticated user.
+- Helmet sets security headers, and the two rate limiters above apply to every request.
+- Prompts are cleaned (null bytes, whitespace, length cap) and triple backticks are neutralised before reaching the LLM.
+- Uploaded SQLite files are opened read-only, and Postgres queries have a 10 second statement timeout.
-Extracts executable queries from LLM responses:
+What you need to know before deploying:
-- **Fence Extraction**: Finds SQL within ` ```sql ... ``` ` code blocks
-- **JSON Handling**: Parses JSON-encoded responses, extracts `response`, `output`, `result` fields
-- **Format Flexibility**: Handles various LLM output formats from different providers
-- **Validation**: Detects obvious SQL keywords (SELECT, INSERT, UPDATE, etc.)
-
-### Database Controller (`controllers/dbController.js`)
-
-Unified interface for multi-database query execution:
-
-**SQLite**:
-```javascript
-const db = new Database(filePath);
-return db.prepare(query).all().slice(0, maxRows);
-```
-
-**PostgreSQL**:
-```javascript
-const client = new PgClient(connectionString);
-const result = await client.query(query);
-```
-
-**MySQL**:
-```javascript
-const connection = await mysql.createConnection(connectionString);
-const [rows] = await connection.execute(query);
-```
-
-**MongoDB**:
-```javascript
-const client = new MongoClient(connectionString);
-const db = client.db(databaseName);
-const coll = db.collection(collectionName);
-const docs = await coll.find(queryObj).toArray();
-```
-
-**Neo4j** (with HTTP fallback):
-- Primary: Bolt driver with auto-fallback if routing/network fails
-- Fallback: HTTP transactional endpoint with base64 auth
-- Supports both `neo4j://`, `bolt://`, and `http://` schemes
-
-### Conversation Memory (`utils/conversationMemory.js`)
-
-Automatic context preservation across conversations:
-
-- **Summarization**: Creates bullet-point summaries of oldest K messages (default 15)
-- **Model**: Uses configurable model (default: Mistral)
-- **Storage**: Persists `memorySummary` and `memorySummaryUpdatedAt` to Chat document
-- **Use Cases**: Reduces context length for long conversations, maintains task understanding
-
-### Security & Validation
-
-**Password Security**:
-- 10-round bcrypt hashing via `bcryptjs`
-- Never returns passwords in API responses
-- Password-only validated on login attempt
-
-**JWT Authentication**:
-- Standard Bearer token in Authorization header
-- 7-day expiry (configurable via `TOKEN_EXPIRY`)
-- Secret must be changed in production (`.env` vs hardcoded `JWT_SECRET`)
-
-**Input Sanitization**:
-```javascript
-// SQL injection prevention: Triple-backtick escaping
-s.replace(/```/g, '`` + `` + `');
-
-// XSS prevention: xss package for user-generated content
-const cleanedText = xss(userInput);
-
-// Null byte filtering for prompt sanitization
-s.replace(/\u0000/g, '')
-```
-
-**Rate Limiting**:
-- Global: 100 requests per 60 seconds per IP
-- Demo bypass: `/api/query/demo` endpoints skip rate limiting
-- Configurable via `express-rate-limit`
-
-**Security Headers**:
-- Helmet.js enforces CSP, X-Frame-Options, X-Content-Type-Options, etc.
-- CORS restricted to configured origins (modifiable in `index.js`)
-
-## Security Considerations
-
-1. **Environment Variables**: Never commit `.env` to version control. Use `.gitignore` (included).
-
-2. **JWT Secret**: Generate a strong, random secret in production:
- ```bash
- node -e "console.log(require('crypto').randomBytes(32).toString('hex'))"
- ```
-
-3. **MongoDB Connection**:
- - Use MongoDB Atlas with firewall rules for production
- - Avoid public IP exposure
- - Database-level authentication (username/password)
-
-4. **LLM API Keys**:
- - Store in `.env`, never hardcode
- - Rotate regularly
- - Monitor usage and set spending limits (e.g., OpenRouter)
-
-5. **File Uploads**:
- - 100 MB file limit (configurable via `multer`)
- - Files stored temporarily in `uploads/` directory
- - Consider cleanup job for abandoned files
-
-6. **HTTPS in Production**:
- - Always use TLS/SSL
- - Update `JWT_SECRET` to a strong value
- - Set `NODE_ENV=production` to enable optimizations
-
-7. **CORS Policy**:
- - Default: All origins allowed
- - Production: Restrict to frontend domain
+- **Set `JWT_SECRET`.** If it is unset, tokens are signed with a well-known default string.
+- **`/api/db/upload` and `/api/db/execute` do not require a JWT**, even though the frontend attaches one to execute requests. Anyone who can reach the server can upload files and ask it to connect to any database they name, and the Neo4j HTTP path will send requests to any host in the connection string. Put these routes behind authentication and network controls before exposing the backend publicly.
+- **Generated queries run as written.** Beyond SQLite files, nothing restricts Postgres, MySQL or MongoDB access to read-only, so connect with a read-only database user and review queries before pressing **Run**.
+- **CORS allows all origins** (`cors()` with defaults). Restrict it in `index.js` for production.
+- The `xss` package is installed but not currently used anywhere.
+- Auth tokens (`qc_token`) and saved connection strings (`qc_conn_default`) are kept in the browser's `localStorage`.
+- Uploaded files are never deleted automatically.
+- Use HTTPS in production, and keep LLM keys in environment variables, never in source control (`.env` files are git-ignored).
## Testing
-The repository does not currently include automated test suites. Manual testing workflow:
-
-1. **Authentication**: Sign up → login → token persistence
-2. **Query Generation**: Test each supported database type
-3. **File Upload**: Upload CSV → verify metadata storage
-4. **Multi-turn Conversation**: Create chat → add multiple queries → verify history
-5. **LLM Providers**: Test Gemini, OpenRouter, and Ollama models
-
-## Limitations
-
-1. **No Cursor-Based Pagination**: Chat queries endpoint returns all results; consider pagination for large chat histories.
-
-2. **Fixed Result Limit**: Default 1000 rows for database queries; very large datasets may be truncated.
-
-3. **No Transaction Support**: Individual queries only; multi-statement transactions not supported.
-
-4. **Limited GraphQL Support**: GraphQL endpoint detection works; execution depends on external GraphQL server.
-
-5. **No Query Validation**: Generated queries are sent directly to databases; malformed queries return database errors.
-
-6. **Conversation Memory**: Summary generation is non-blocking; may lag behind real conversations.
-
-7. **No Concurrent User Limits**: Backend designed for small-to-medium user bases; horizontal scaling needed for enterprise.
+There is no automated test suite. The only quality tooling is the frontend linter:
-8. **File Cleanup**: Uploaded files not auto-deleted; manual cleanup or scheduled job needed.
-
-9. **Model Availability**: OpenRouter and Gemini require active API keys; failures if services unavailable.
-
-10. **No Query Explanation**: LLM generates queries but doesn't automatically explain them to users.
-
-## Future Improvements
-
-1. **Advanced UI**:
- - Query result visualization (charts, graphs, maps)
- - Side-by-side query/result comparison
- - Query saved templates and snippets
-
-2. **LLM Enhancements**:
- - In-context learning from user corrections
- - Few-shot examples in system prompt
- - Query optimization suggestions
+```bash
+cd querycraft-frontend
+npm run lint
+```
-3. **Database Features**:
- - Schema introspection modal
- - Query performance analysis
- - Index recommendations
+A quick manual pass: sign up and log in, send a prompt in a new chat, reload to confirm history persists, upload a CSV and press **Run** on a generated query, then try each provider you have configured.
-4. **Scalability**:
- - Redis caching for repeated queries
- - Query result caching layer
- - Horizontal scaling with load balancing
+## Known limitations
-5. **Testing & Validation**:
- - Comprehensive Jest/Mocha test suite
- - Query validation before execution
- - Dry-run mode for previewing results
+- The row limit (`maxRows`, default 1000) is applied after PostgreSQL, MySQL, SQLite and Neo4j return their rows, so add a `LIMIT` to queries over large tables.
+- MongoDB execution supports `find` with a filter only, not aggregation pipelines.
+- CQL, Redis, Elasticsearch, DynamoDB and GraphQL queries can be generated but not run.
+- Chat context is the last three completed exchanges. The summariser in `utils/conversationMemory.js` is not called anywhere, and the `Chat` schema has no fields to store its output.
+- Prompt whitespace is collapsed, so line breaks in a pasted query are flattened before it reaches the model.
+- The upload index (`db_files.json`) is a local file and the rate limiter is in-memory, so the backend is designed for a single instance.
+- Model output can be wrong or unexpected; there is no query validation step.
+- Retries with backoff exist for Gemini only.
-6. **User Experience**:
- - Query history export (CSV, JSON)
- - Shareable query links
- - Team collaboration features
- - Custom prompt templates
+## Roadmap
-7. **Monitoring & Analytics**:
- - Query execution logging and analytics
- - LLM cost tracking per user
- - Performance metrics dashboard
+Ideas grounded in the gaps above:
-8. **Security**:
- - Two-factor authentication (2FA)
- - Query audit trail with timestamps
- - Role-based access control (RBAC)
+- Add authentication and ownership checks to `/api/db/*`, restrict CORS, and cap or validate connection targets.
+- Add automated tests (API and query-detection unit tests) and make the CI workflow useful for this monorepo layout.
+- Wire up conversation summaries (add the fields to the `Chat` schema and call the summariser), and add pagination for long chats.
+- Support MongoDB aggregation and execution for more of the languages that can already be generated.
+- Schema introspection so prompts can include real table and column names.
+- Automatic clean-up of old uploads.
## Contributing
-Contributions are welcome! To contribute:
+Contributions are welcome.
-1. **Fork** the repository: https://github.com/sultanmaliki/QueryCraft-AI
-2. **Create** a feature branch: `git checkout -b feature/your-feature`
-3. **Commit** changes: `git commit -m "Add your feature"`
-4. **Push** to branch: `git push origin feature/your-feature`
-5. **Open** a Pull Request with description
-
-Please ensure:
-- Code follows existing style (no specific linter configured yet)
-- New features include appropriate error handling
-- Environment variables are documented in `.env.example`
+1. Fork the repository and create a branch: `git checkout -b feature/your-feature`.
+2. Make your changes. Run `npm run lint` in `querycraft-frontend/` for frontend changes (the backend has no linter or tests yet).
+3. Document any new environment variable in the [Configuration](#configuration) tables.
+4. Commit, push, and open a pull request describing the change.
## License
-MIT License. See [LICENSE](LICENSE) file for details.
-
----
+MIT. See [LICENSE](LICENSE).
-**Questions or Issues?** Open a GitHub Issue at [QueryCraft-AI/issues](https://github.com/sultanmaliki/QueryCraft-AI/issues) or contact the development team.
+## Links
-**Repository**: [github.com/sultanmaliki/QueryCraft-AI](https://github.com/sultanmaliki/QueryCraft-AI)
+- Repository: [github.com/sultanmaliki/QueryCraft-AI](https://github.com/sultanmaliki/QueryCraft-AI)
+- Issues: [github.com/sultanmaliki/QueryCraft-AI/issues](https://github.com/sultanmaliki/QueryCraft-AI/issues)
-**Authors**:
-- Syed Mohammed Sultan
-- Rifaque Ahmed Akrami
-- Raif
+**Authors:** Syed Mohammed Sultan, Rifaque Ahmed Akrami, Raif
-**Project Status**: Completed
+**Project status:** Completed