Popular repositories Loading
-
GLM-5.3-Flash-NVFP4-TP3-3x-DGX-Spark
GLM-5.3-Flash-NVFP4-TP3-3x-DGX-Spark PublicRun zai-org/GLM-5.3-Flash (NVFP4) on 3x NVIDIA DGX Spark with vLLM, TP=3 + EP, DFlash2 speculative decoding, CUDA graphs. Full recipe with rationale, benchmarks, patches, and what we tried.
Python 2
-
cuda-exl3
cuda-exl3 PublicForked from Zeuss5/cuda-exl3
EXL3 (ExLlamaV3 trellis) CUDA kernels for Blackwell: dense + MoE GEMM and a fused sparse-MLA attention backend, with a vLLM plugin
Cuda
-
GLM-5.3-Flash-EXL3-TP3-3x-DGX-Spark
GLM-5.3-Flash-EXL3-TP3-3x-DGX-Spark PublicRun zai-org/GLM-5.3-Flash as an EXL3 4bpw checkpoint on 3x NVIDIA DGX Spark (TP=3 + expert parallel, DFlash2) with vLLM and cuda-exl3 — plus a measured TP=2 track for two-node setups. Full recipe w…
Python
-
If the problem persists, check the GitHub status page or contact support.