ML Engineer — LLM quantization & inference on consumer hardware
Dallas–Fort Worth, TX · Portfolio · Hugging Face · LinkedIn · ttimmsinternational@gmail.com
I quantize and serve large models on hardware that isn't supposed to run them — 16 GB Blackwell GPUs, Jetson edge boards — and I reproduce every published number against its confidence interval before I call it done.
Open to ML Engineer roles — DFW or remote.
MoE Pruning + NVFP4 — A 50%-expert-pruned MoE coder model, quantized to fit 16 GB VRAM. SWE-bench Verified 52.0% (26/50, officially graded), HumanEval+/MBPP+ reproduced inside published confidence intervals, CI-checked reproduction pipeline. Model on Hugging Face →
ZAYA1 NVFP4 W4A4 — 4-bit weights and activations on native Blackwell tensor cores: 9.5 tok/s single-stream from a 6.02 GB checkpoint, 2,100+ combined downloads on Hugging Face. Includes a benchmark I retracted and corrected in public once I found the CUDA-graph path corrupting output. Model on Hugging Face →
Godspeed Coding Agent — A coding agent built from scratch: deny-first permission engine, SHA-256 hash-chained audit trail. SWE-bench Lite 34.8% single-shot / 52.2% oracle best-of-5, $0 API spend.
Sovereign Edge — Five-agent personal AI system running entirely on a Jetson Orin Nano — zero cloud dependencies.
Also: Bible AI Assistant (ORPO fine-tune + hybrid RAG) · Manna Trading (multi-agent trading pipeline) · an open llama.cpp PR fixing an NVFP4 quantizer crash
Python · PyTorch · vLLM / CUTLASS · TRL / Unsloth · NVFP4 · GGUF · GPTQ / AWQ · CUDA · Docker



