CLI Command Group

moro data

Manage dataset ingestion, normalization, information-theoretic deduplication, MI Guard protection, and quality diagnostics.

Subcommands

1. `moro data build`

Executes the complete epistemic curation pipeline on raw files. Automatically redacts PII, scores samples with PPMI and Resnik IC, applies the MI Guard, and produces clean train/val/eval splits.

moro data build [OPTIONS]
  
Flag Type Default Description
--source, -s Path data/raw/ Raw input file (.jsonl, .csv, .parquet) or directory.
--mi-guard-threshold Float 0.85 PPMI information density threshold to preserve outliers.
--dedup-threshold Float 0.88 Cosine similarity threshold for near-duplicate pruning.
--pii-scan / --no-pii Boolean true Toggle automatic redaction of emails, phones, and API keys.
--val-ratio Float 0.10 Fraction of dataset allocated to validation set.

2. `moro data report`

Generates an in-depth epistemic report of token entropy, PPMI distribution, and MI Guard statistics:

moro data report --compiled-dir data/compiled/