FlashInfer: Kernel Library for LLM Serving
-
Updated
Sep 11, 2026 - Cuda
FlashInfer: Kernel Library for LLM Serving
A GPU cluster manager for high-performance AI model serving (vLLM, SGLang) and on-demand SSH-accessible GPU instances.
Achieve state of the art inference performance with modern accelerators on Kubernetes
SGLang-Omni is a high-performance serving framework for audio models (TTS, ASR) and unified multimodal models.
Best practices, reference architectures, and examples for distributed AI training and inference on AWS.
Every device brings a slice. Together they run the whole model. Peer-to-peer LLM inference across browser tabs: a from-scratch WebGPU engine and a WebRTC runtime that split a 27B model over the devices in a room.
Split FLUX.2 and LTX 2.3 across two GPUs (LAN or same-machine) — NVENC compresses activations live on the wire. Icarus (ComfyUI node) + Daedalus (back-half server).
[Official] Prima.cpp: Scale Your Local AI Beyond One Device.
一個基於 llama.cpp 的分佈式 LLM 推理程式,讓您能夠利用區域網路內的多台電腦協同進行大型語言模型的分佈式推理,使用 Electron 的製作跨平台桌面應用程式操作 UI。
An educational distributed training and inference library for neural nets using local computing
Mixed-vendor GPU inference cluster manager with speculative decoding
Run any model on Intel silicon
Decentralized peer-to-peer LLM inference network. Single Rust binary, BitTorrent-inspired incentives, OpenAI-compatible API.
Code for paper "JMDC: A Joint Model and Data Compression System for Deep Neural Networks Collaborative Computing in Edge-Cloud Networks"
Turn any Mac or GPU into an OpenAI-compatible inference node. One-command setup, automatic HTTPS, model management, and distributed request routing.
An imperative command-line-interface for AI workload orchestration
🧠 One LLM, split across a Mac (Apple MPS) + a Windows PC (NVIDIA CUDA) — heterogeneous pipeline inference over a framework-neutral wire, bit-for-bit identical to single-machine. No torch.distributed, no datacenter.
MiniMax-H3 multi-GPU parallel inference for ComfyUI | 多卡并行加速节点:2-8 GPU Ulysses sequence parallel, bit-identical video+audio generation
Optimize and run consistent AI Functions.
To associate your repository with the distributed-inference topic, visit your repo's landing page and select "manage topics."