Setup & Verification

Installation Guide

Step-by-step installation instructions for Linux, macOS (Apple Silicon), and Windows WSL2.

System Prerequisites

  • Python: Version 3.10, 3.11, or 3.12
  • RAM: 8GB minimum (16GB+ recommended for dataset compilation)
  • Disk Space: 10GB free SSD space for base model weights and cached adapters
  • GPU (Optional for training): NVIDIA GPU with 4GB+ VRAM or Apple Silicon Mac (M1/M2/M3/M4)

Installation Methods

Option 1: Standard `pip` (Recommended)

# Basic installation (CLI, data compiler, and evaluation)
pip install moroai

# With training extensions (PyTorch, BitsAndBytes, PEFT, TRL)
pip install "moroai[train]"

# Full installation with all optional export and serving plugins
pip install "moroai[all]"
  

Option 2: From Source (Latest Development)

git clone https://github.com/moroai/moro.git
cd moro
pip install -e ".[train]"
  

Option 3: Pre-Built Docker Image

docker pull moroai/moro:latest
docker run --gpus all -it -p 3000:3000 moroai/moro:latest
  

Verifying Installation (`moro doctor`)

Run the built-in diagnostic tool to verify all components:

$ moro doctor

MoroAI System Diagnostics
─────────────────────────────────────────────────
Platform:         Darwin / Linux (x86_64 or arm64)
Python:           3.11.5 (Active environment)
CLI Version:      0.1.0 (moro)

Core Dependencies:
  ✓ typer         0.12.0
  ✓ rich          13.7.0
  ✓ pydantic      2.6.0
  ✓ pyyaml        6.0.1

Training Backend:
  ✓ torch         2.2.0 (CUDA 12.1 or MPS available)
  ✓ transformers  4.38.0
  ✓ peft          0.9.0
  ✓ bitsandbytes  0.42.0

Hardware Status:
  ✓ GPU: NVIDIA GeForce RTX 4070 (12,288 MB VRAM)
  ✓ Driver: 545.29.06 | CUDA: 12.1
  ✓ Zero-OOM Pre-Flight Engine: ACTIVE

System Verdict:
  ✓ READY FOR DATA CURATION & LOCAL TRAINING
  

GPU Compatibility & Model Sizes

GPU Tier Dedicated VRAM Supported Target Models Quantization
Entry Level 4 GB – 8 GB 0.5B – 1.5B (Qwen2.5, Llama-3.2) 4-bit NF4
Mid-Range (RTX 3060/4070) 12 GB – 16 GB 1.5B – 3B (Llama-3.2-3B, Qwen2.5-3B) 4-bit NF4
High-End (RTX 3090/4090) 24 GB 7B – 8B (Llama-3.1-8B, Mistral-7B) 4-bit NF4 or 8-bit
Apple Silicon (M2/M3/M4) 16 GB – 64 GB Unified 1.5B – 14B (Metal Acceleration) FP16 / 4-bit MLX
Data Center (A100 / H100) 40 GB – 80 GB 14B – 70B (Full LoRA Rank 64) BF16 / FP16

Common Installation Troubleshooting

1. CUDA Not Found or PyTorch CPU-Only Installed

# Verify NVIDIA driver is working:
nvidia-smi

# Reinstall PyTorch with explicit CUDA 12.1 runtime:
pip install --upgrade torch --index-url https://download.pytorch.org/whl/cu121
  

2. Apple Silicon Metal Performance Shaders (MPS) Setup

On macOS with Apple Silicon chips, hardware acceleration is enabled out of the box through PyTorch's Metal backend:

python -c "import torch; print('MPS available:', torch.backends.mps.is_available())"