CLI Command Group
moro eval
Execute deterministic schema gates, semantic LLM accuracy evaluations, Wasserstein representation drift audits, and adversarial stress testing.
Subcommands
1. `moro eval run`
Executes a structured evaluation test suite against an adapted model checkpoint or base foundation model.
moro eval run --checkpoint ./runs/run_01/checkpoints/best/ --suite ./eval/suite.yaml
| Flag | Type | Default | Description |
|---|---|---|---|
| --checkpoint | Path | Required | Path to candidate model weights or adapter directory. |
| --suite | Path | ./eval/suite.yaml | Path to evaluation suite YAML file. |
| --judge-model | String | local | Model used for semantic scoring: local (Ollama) or HF model ID. |
2. `moro eval compare`
Runs identical benchmarks across a base model and adapted checkpoint to measure accuracy delta and drift:
moro eval compare \
--base Qwen/Qwen2.5-1.5B \
--candidate ./runs/run_01/checkpoints/best/ \
--suite ./eval/suite.yaml