Local Serving

Deploying to Local Ollama

Export your fine-tuned LoRA adapters directly into Ollama with GGUF quantization in a single command.

Prerequisites

  • Ollama installed and running locally (ollama --version)
  • A completed MoroAI training run with a certified checkpoint

1-Command Deployment

moro release deploy --target ollama --quantize q4_k_m --model-name my-assistant:v1
  

MoroAI executes the following steps behind the scenes:

  1. Merges the low-rank adapter weights into the base model tensors.
  2. Executes llama.cpp quantization into Q4_K_M GGUF format.
  3. Generates a tailored Modelfile with system prompt templates and stop tokens.
  4. Calls the Ollama API to create and register the model tag locally.

Testing Your Model

ollama run my-assistant:v1 "Hello! Summarize our return policy."