Subsystem 02
VRAM Recipe Predictor
Eliminating trial-and-error fine-tuning by simulating memory consumption down to the megabyte before touching GPU memory.
The VRAM Memory Formula
When fine-tuning on consumer hardware, memory exhaustion is non-linear. Peak GPU VRAM is modeled prior to training execution across five critical tensor allocations:
Analytical VRAM Memory Model Peak Memory Equation
VRAM_Peak = M_base + M_lora + M_optimizer + M_activations + M_cuda_context
Where:
M_base = (N_params * B_quant) / 8 # Base weights in NF4 / FP16
M_lora = 2 * rank * d_model * N_targets * 4 # Trainable LoRA adapter matrices (FP32)
M_optimizer = M_lora * 2 # 8-bit AdamW first and second moments
M_activations = b * s * L * (34 * d_model + 5 * s * h) # Backprop activation tensors
M_cuda_ctx = ~850MB to 1.2GB # CUDA driver context & kernels
Hardware Tier Presets & Validation
The engine maintains validated memory envelopes across standard consumer and workstation GPUs:
| Hardware Profile | Available Memory | Recommended Model | LoRA Rank (r) | Context Length |
|---|---|---|---|---|
| RTX 3060 / 4060 | 12 GB VRAM | Qwen2.5-1.5B / Llama-3.2-1B | r=16, alpha=32 | 2,048 tokens |
| RTX 4070 / 4070 Ti | 12–16 GB VRAM | Qwen2.5-3B / Llama-3.2-3B | r=32, alpha=64 | 4,096 tokens |
| RTX 3090 / 4090 | 24 GB VRAM | Qwen2.5-7B / Llama-3.1-8B | r=64, alpha=128 | 4,096 tokens |
| Apple Silicon (M2/M3/M4) | 16–64 GB Unified | Qwen2.5-7B / Mistral-7B | r=16, alpha=32 | 2,048 tokens |
| Cloud A100 / H100 | 80 GB VRAM | Qwen2.5-14B / 32B | r=64, alpha=128 | 8,192 tokens |
Generating a Reproducible Recipe
Run the recipe engine to automatically probe your active GPU driver, detect memory bandwidth, and output a guaranteed zero-OOM moro.yaml configuration:
# Automatically audit GPU and generate optimal training configuration moro recipe generate --model Qwen/Qwen2.5-1.5B --hardware detect --output ./moro.yaml
Generated Recipe File Anatomy (moro.yaml)
name: "customer-support-qwen" version: "0.1.0" model: base: "Qwen/Qwen2.5-1.5B" quantization: "nf4" # 4-bit NormalFloat quantization lora_rank: 16 lora_alpha: 32 lora_dropout: 0.05 target_modules: ["q_proj", "k_proj", "v_proj", "o_proj", "gate_proj", "up_proj", "down_proj"] training: micro_batch_size: 2 gradient_accumulation_steps: 8 learning_rate: 2.0e-4 max_seq_length: 2048 optimizer: "adamw_8bit" gradient_checkpointing: true auto_heal: true # Enables 5-level autonomous OOM recovery