Hardware Optimization

Consumer GPU Performance Tuning

Techniques to maximize tokens/second and train 1.5B–14B models on consumer graphics cards.

1. Flash Attention 2 & SDPA

Ensure PyTorch 2.0+ uses scaled dot-product attention (SDPA) or Flash Attention 2 for 2.5x speedups and 50% activation memory reduction.

2. 8-Bit AdamW Optimizer

Using paged_adamw_8bit cuts optimizer state memory footprint by 75% compared to standard FP32 AdamW without any drop in model quality.

3. LoRA Rank & Target Modules

Setting lora_r=16 with lora_alpha=32 provides 98% of the expressive power of full fine-tuning with only 0.2% of the trainable parameters.