Highlights
- Pro
Pinned Loading
-
roofline-llama
roofline-llama PublicA from-scratch implementation of Llama-3.2-1B in PyTorch, decode-latency benchmarks on three GPUs (T4, L4, A100), three weight-only quantization methods (RTN, GPTQ, AWQ) measured against both, and …
Jupyter Notebook
-
triton-attention-lab
triton-attention-lab PublicFlashAttention-2 forward and Flash-Decoding kernels written from scratch in Triton, with grouped-query attention (GQA) and causal masking, benchmarked against PyTorch SDPA and a naive eager impleme…
Jupyter Notebook
-
LlamaDistill-Sentiment
LlamaDistill-Sentiment PublicLlamaDistill-Sentiment (LLM-NEO v2) is a complete end-to-end machine learning system for distilling a large LLM (Meta-Llama-3-8B) into a compact, efficient student model (Llama-3.2-1B)
Jupyter Notebook
-
MishrTok
MishrTok PublicMishrTok (मिश्र = mixed): 32k BPE tokenizer for code-mixed Hinglish (Roman) + Hindi (Devanagari) — ~14% fewer tokens vs GPT-4o.
Jupyter Notebook
If the problem persists, check the GitHub status page or contact support.