Cookbook 04

Document Summarizer

Fine-tune a model to distill technical reports and executive briefs into standardized key takeaways.

1. Long-Context Configuration

Configure moro.yaml with sequence length expansion and Flash Attention:

model:
  base_model: "Qwen/Qwen2.5-3B"
  quantization: "4bit"
  max_seq_length: 4096

training:
  gradient_checkpointing: true
  micro_batch_size: 1
  gradient_accumulation_steps: 16