ML Engineer — LLM quantization & inference on consumer hardware
Dallas–Fort Worth, TX · Portfolio · Hugging Face · LinkedIn · ttimmsinternational@gmail.com
I quantize and serve large models on hardware that isn't supposed to run them — 16 GB Blackwell GPUs, Jetson edge boards — and I reproduce every published number against its confidence interval before I call it done.
8 models on Hugging Face · ~3,800 downloads/month.
Open to ML Engineer roles (inference optimization, model compression, edge deployment) — DFW or remote.
Bible AI Assistant — Local Scripture Q&A on a 16 GB card: hybrid RAG over 31k verses feeding an SFT → ORPO → GRPO fine-tune with a verifiable reward — the cited verse must exist in the index and the quote must match. sha256-pinned benchmark protocol, 430 tests, full CI/CD. Primary project — building toward local SOTA on a 5070 Ti.
MoE Pruning + NVFP4 — A 50%-expert-pruned MoE coder model, quantized to fit 16 GB VRAM. SWE-bench Verified 52.0% (26/50, officially graded), HumanEval+/MBPP+ reproduced inside published confidence intervals, CI-checked reproduction pipeline. Model on Hugging Face →
ZAYA1 NVFP4 W4A4 — 4-bit weights and activations on native Blackwell tensor cores: 9.5 tok/s single-stream from a 6.02 GB checkpoint, 2,100+ combined downloads on Hugging Face. Includes a benchmark I retracted and corrected in public once I found the CUDA-graph path corrupting output. Model on Hugging Face →
Godspeed Coding Agent — A coding agent built from scratch: deny-first permission engine, SHA-256 hash-chained audit trail. SWE-bench Lite 34.8% single-shot / 52.2% oracle best-of-5, $0 API spend.
Sovereign Edge — Five-agent personal AI system running entirely on a Jetson Orin Nano — zero cloud dependencies.
Also: Manna Trading (multi-agent trading pipeline) · an open llama.cpp PR fixing an NVFP4 quantizer crash
Python · PyTorch · vLLM / CUTLASS · TRL / Unsloth · NVFP4 · GGUF · GPTQ / AWQ · CUDA · Docker



