Setup & Verification
Installation Guide
Step-by-step installation instructions for Linux, macOS (Apple Silicon), and Windows WSL2.
System Prerequisites
- Python: Version 3.10, 3.11, or 3.12
- RAM: 8GB minimum (16GB+ recommended for dataset compilation)
- Disk Space: 10GB free SSD space for base model weights and cached adapters
- GPU (Optional for training): NVIDIA GPU with 4GB+ VRAM or Apple Silicon Mac (M1/M2/M3/M4)
Installation Methods
Option 1: Standard `pip` (Recommended)
# Basic installation (CLI, data compiler, and evaluation) pip install moroai # With training extensions (PyTorch, BitsAndBytes, PEFT, TRL) pip install "moroai[train]" # Full installation with all optional export and serving plugins pip install "moroai[all]"
Option 2: From Source (Latest Development)
git clone https://github.com/moroai/moro.git cd moro pip install -e ".[train]"
Option 3: Pre-Built Docker Image
docker pull moroai/moro:latest docker run --gpus all -it -p 3000:3000 moroai/moro:latest
Verifying Installation (`moro doctor`)
Run the built-in diagnostic tool to verify all components:
$ moro doctor MoroAI System Diagnostics ───────────────────────────────────────────────── Platform: Darwin / Linux (x86_64 or arm64) Python: 3.11.5 (Active environment) CLI Version: 0.1.0 (moro) Core Dependencies: ✓ typer 0.12.0 ✓ rich 13.7.0 ✓ pydantic 2.6.0 ✓ pyyaml 6.0.1 Training Backend: ✓ torch 2.2.0 (CUDA 12.1 or MPS available) ✓ transformers 4.38.0 ✓ peft 0.9.0 ✓ bitsandbytes 0.42.0 Hardware Status: ✓ GPU: NVIDIA GeForce RTX 4070 (12,288 MB VRAM) ✓ Driver: 545.29.06 | CUDA: 12.1 ✓ Zero-OOM Pre-Flight Engine: ACTIVE System Verdict: ✓ READY FOR DATA CURATION & LOCAL TRAINING
GPU Compatibility & Model Sizes
| GPU Tier | Dedicated VRAM | Supported Target Models | Quantization |
|---|---|---|---|
| Entry Level | 4 GB – 8 GB | 0.5B – 1.5B (Qwen2.5, Llama-3.2) | 4-bit NF4 |
| Mid-Range (RTX 3060/4070) | 12 GB – 16 GB | 1.5B – 3B (Llama-3.2-3B, Qwen2.5-3B) | 4-bit NF4 |
| High-End (RTX 3090/4090) | 24 GB | 7B – 8B (Llama-3.1-8B, Mistral-7B) | 4-bit NF4 or 8-bit |
| Apple Silicon (M2/M3/M4) | 16 GB – 64 GB Unified | 1.5B – 14B (Metal Acceleration) | FP16 / 4-bit MLX |
| Data Center (A100 / H100) | 40 GB – 80 GB | 14B – 70B (Full LoRA Rank 64) | BF16 / FP16 |
Common Installation Troubleshooting
1. CUDA Not Found or PyTorch CPU-Only Installed
# Verify NVIDIA driver is working: nvidia-smi # Reinstall PyTorch with explicit CUDA 12.1 runtime: pip install --upgrade torch --index-url https://download.pytorch.org/whl/cu121
2. Apple Silicon Metal Performance Shaders (MPS) Setup
On macOS with Apple Silicon chips, hardware acceleration is enabled out of the box through PyTorch's Metal backend:
python -c "import torch; print('MPS available:', torch.backends.mps.is_available())"