- Matrix Multiplication Benchmark (CPU vs GPU) rust_gemm.md
- GPU-Accelerated Image Blur (Box Filter) box_filter.md
- Mandelbrot Set Rendering mandelbrot.md
- N-Body Physics Simulation physics.md
- Burn & Cuda: Resnet-18 on CIFAR-10 restnet-18.md
- Custom CUDA Kenernel with
rustacuda(Parallel Reduction) parrallel_reduction.md- Monte Carlo Pi Estimation monte_carlo.md
- Matrix Multiply (GEMM) : Shared Memory & tiling. 50 -100x expected speedup (vs CPU).
- Image Blur : Shared memory for 3x3 neighborhood. 20-50x expected speedup (vs CPU).
- Mandelbrot : Coalesced memory access. 10-30x expected speedup (vs CPU).
- N-Body : Shared memory & tree methods. 30-80x expected speedup (vs CPU).
- Parallel Reduction : Bank conflict avoidence. 50-150x expected speedup (vs CPU).
- Monte Carlo Pi : Minimize atomic ops. 40-100x expected speedup (vs CPU).
- Run it Locally:
```bash
cargo run --release
```
- Build and Run it Docker:
```bash
docker build -t ${PROJECT_TITLE} .
docker run -it ${PROJECT_TITLE}
```