CLI Command Group

moro eval

Execute deterministic schema gates, semantic LLM accuracy evaluations, Wasserstein representation drift audits, and adversarial stress testing.

Subcommands

1. `moro eval run`

Executes a structured evaluation test suite against an adapted model checkpoint or base foundation model.

moro eval run --checkpoint ./runs/run_01/checkpoints/best/ --suite ./eval/suite.yaml
  
Flag Type Default Description
--checkpoint Path Required Path to candidate model weights or adapter directory.
--suite Path ./eval/suite.yaml Path to evaluation suite YAML file.
--judge-model String local Model used for semantic scoring: local (Ollama) or HF model ID.

2. `moro eval compare`

Runs identical benchmarks across a base model and adapted checkpoint to measure accuracy delta and drift:

moro eval compare \
    --base Qwen/Qwen2.5-1.5B \
    --candidate ./runs/run_01/checkpoints/best/ \
    --suite ./eval/suite.yaml