Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

PulseWatch

Cloud Observability and Incident Management Platform

Monitor public websites and APIs, measure service reliability, and track incidents from a focused operations dashboard.

Next.js FastAPI Python PostgreSQL Docker CI

Overview

PulseWatch is a full-stack observability platform that performs scheduled health checks against public HTTP services. It records response status and latency, calculates rolling reliability metrics, and automatically opens or resolves incidents when service state changes.

The project uses asynchronous API handlers, background task processing, persistent check history, Redis event distribution, and a responsive operational dashboard.

Dashboard Preview

PulseWatch operations dashboard with uptime, latency, SLO, incidents, and monitor inventory

The dashboard is connected to real public HTTP services. It displays observed response codes and latency rather than sample data.

Features

  • Configurable GET and HEAD monitors for public websites and APIs
  • Automated health checks using Celery workers and Celery Beat
  • Live service state, HTTP response code, and latency monitoring
  • Interactive 24-hour, 7-day, and 30-day latency analytics
  • Fleet uptime, P95 latency, health score, SLO posture, and estimated error budget
  • Dedicated monitor inventory with search, state filters, and CSV export
  • Monitor detail console with recent check history and pause/resume controls
  • Consecutive-failure and recovery confirmation policies to reduce alert noise
  • Global or per-service maintenance windows with automatic incident suppression
  • Automatic incident creation when a service fails
  • Automatic incident resolution when the service recovers
  • Incident acknowledgement, manual resolution, operator attribution, and audit history
  • Separate blocked-probe state for WAF and rate-limit responses
  • Exponential retry/backoff for transient network and upstream failures
  • DNS rebinding protection plus live DNS and TLS certificate diagnostics
  • Telegram, generic webhook, and SMTP email incident notifications
  • Customer-facing public status page and JSON status feed
  • Prometheus-compatible metrics and database/Redis readiness probes
  • Optional operator authentication with scrypt-hashed passwords and HttpOnly sessions
  • Searchable check history stored in PostgreSQL
  • Redis pub/sub event distribution and WebSocket streaming
  • Responsive Next.js operations dashboard
  • URL validation that blocks localhost, private IP, and reserved IP targets
  • Containerized local environment with Docker Compose

System Architecture

flowchart LR
    UI[Next.js Dashboard] --> API[FastAPI API]
    API --> DB[(PostgreSQL)]
    Beat[Celery Beat] --> Queue[(Redis)]
    Queue --> Worker[Celery Workers]
    Worker --> Target[Public Website or API]
    Worker --> DB
    Worker --> Queue
    Queue --> API
    API --> UI
Loading

Monitoring flow

  1. Celery Beat scans for monitors whose check interval has elapsed.
  2. Redis queues a monitoring task for each due endpoint.
  3. A Celery worker performs the HTTP request and measures latency.
  4. The result is persisted in PostgreSQL.
  5. Transient failures are retried with exponential backoff before classification.
  6. A state transition from up to down opens an incident.
  7. A transition from down to up resolves the active incident.
  8. Redis publishes the new result for real-time consumers.

Technology Stack

Layer Technologies
Frontend Next.js 16, React 19, TypeScript, CSS
API FastAPI, Pydantic, Uvicorn
Data PostgreSQL 16, SQLAlchemy 2, asyncpg
Background processing Celery, Celery Beat, Redis
Real-time events Redis Pub/Sub, WebSockets
HTTP monitoring HTTPX
Infrastructure Docker, Docker Compose
Quality Pytest, Ruff, TypeScript compiler, npm audit

Quick Start

Requirements

  • Docker Desktop
  • Docker Compose
  • Git

1. Clone the repository

git clone https://github.com/Adityagithubhack/pulsewatch.git
cd pulsewatch

2. Configure the environment

cp .env.example .env

For local development, the provided defaults work without additional services. Change the PostgreSQL password before deploying publicly.

3. Start PulseWatch

docker compose up -d --build

4. Open the application

5. Check service status

docker compose ps

All six services should be running, with PostgreSQL and Redis reporting healthy.

Configuration

Variable Purpose Default
POSTGRES_DB PostgreSQL database name pulsewatch
POSTGRES_USER PostgreSQL username pulsewatch
POSTGRES_PASSWORD PostgreSQL password change-me
DATABASE_URL Async SQLAlchemy connection URL PostgreSQL container URL
REDIS_URL Celery broker, result backend, and pub/sub URL redis://redis:6379/0
NEXT_PUBLIC_API_URL Browser-accessible FastAPI base URL http://localhost:8000
FRONTEND_ORIGIN Allowed dashboard origin for CORS http://localhost:3001
PROBE_REGION Region label attached to every check local
TELEGRAM_BOT_TOKEN, TELEGRAM_CHAT_ID Optional Telegram incident notifications Disabled
ALERT_WEBHOOK_URL Optional generic JSON webhook destination Disabled
SMTP_*, ALERT_EMAIL_* Optional SMTP email escalation Disabled
AUTH_ENABLED Require login for every mutating operation false
ADMIN_EMAIL, ADMIN_PASSWORD Bootstrap credentials when authentication is enabled Empty

API Reference

Method Route Purpose
GET /health Check API availability
GET /api/endpoints List monitors with their latest check
POST /api/endpoints Create a new monitor
GET /api/endpoints/{id} Read one monitor
PATCH /api/endpoints/{id} Update or pause a monitor
DELETE /api/endpoints/{id} Delete a monitor and its history
POST /api/endpoints/{id}/check Run an immediate health check
GET /api/endpoints/{id}/checks Read recent check history
GET /api/endpoints/{id}/metrics Calculate rolling uptime and latency metrics
GET /api/incidents Read the incident and recovery timeline
PATCH /api/incidents/{id} Acknowledge, annotate, or resolve an incident
GET, POST /api/maintenance List and schedule maintenance windows
GET /api/endpoints/{id}/diagnostics Inspect DNS and live TLS certificate health
GET /api/audit Read the immutable operations activity feed
GET /api/public/status Read the customer-safe public status feed
GET /api/dashboard/summary Read operational dashboard totals
WS /ws/checks Stream live monitoring results

Operational endpoints:

  • Public status page: /status
  • Readiness probe: /health/ready
  • Prometheus metrics: /metrics
  • Secure operator login: /login (when AUTH_ENABLED=true)

Interactive OpenAPI documentation is available at /docs while the API is running.

Validation and Tests

Backend

cd backend
python3 -m venv .venv
source .venv/bin/activate
pip install -r requirements-dev.txt
ruff check app tests
pytest -q

Frontend

cd frontend
npm install
npm run build
npm audit --omit=dev

Current verification status:

  • GitHub Actions CI is green for the v1.0.0 release
  • Backend: Ruff validation and 9 automated tests pass
  • Frontend: Next.js production build and TypeScript validation pass
  • Docker Compose environment verified locally with PostgreSQL and Redis readiness checks

Continuous Integration and Deployment

The GitHub Actions pipeline validates every push and pull request by running:

  • Ruff backend linting
  • Pytest backend tests
  • Next.js production build and TypeScript validation
  • Production dependency audit
  • Docker Compose image builds

A separate guarded workflow can deploy successful main builds to a DigitalOcean server through SSH. Deployment stays disabled until the required secrets and the ENABLE_DEPLOY repository variable are configured.

See DigitalOcean Deployment for the complete server and GitHub configuration. See Operations Guide for authentication, notifications, incident policy, and production endpoints.

Project Structure

pulsewatch/
├── backend/
│   ├── app/
│   │   ├── api/             HTTP routes
│   │   ├── core/            Configuration and database setup
│   │   ├── services/        Probing, checks, metrics, and URL validation
│   │   ├── main.py          FastAPI application and WebSocket stream
│   │   ├── models.py        SQLAlchemy models
│   │   ├── schemas.py       Pydantic request and response schemas
│   │   └── worker.py        Celery tasks and scheduler
│   └── tests/
├── frontend/
│   ├── app/                 Next.js App Router
│   ├── components/          Dashboard UI
│   └── lib/                 API client and TypeScript models
├── .github/workflows/   CI and guarded deployment workflows
├── docs/                Deployment documentation
├── docker-compose.yml
└── .env.example

Roadmap

  • Multi-user workspaces and fine-grained role-based access control
  • OIDC/SSO and managed API-key lifecycle
  • Dedicated remote probe agents for genuine multi-region monitoring
  • Domain-registration expiration monitoring
  • Alert delivery retries, channel health, and escalation policies
  • Alembic migration workflow replacing compatibility migrations

Author

Aditya Singh

About

Production-style observability platform for API and website monitoring with SLO analytics, incident workflows, DNS/TLS diagnostics, FastAPI, Next.js, PostgreSQL, Redis, Celery, and Docker.

Topics

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Contributors

Languages