Subsystem 04

Multi-Layer Evaluation Harness

Comprehensive model certification across deterministic rules, semantic judges, distribution drift detection, and adversarial stress tests.

Beyond Validation Loss: The Need for Multi-Layer Certification

In academic research, evaluation often stops at cross-entropy validation loss or standard perplexity. However, in enterprise deployment, an adapted model with a low validation loss can still:

  • Violate strict JSON schema requirements, producing malformed API payloads
  • Hallucinate plausible-sounding legal or medical falsehoods
  • Experience catastrophic forgetting of general reasoning capabilities
  • Break under noisy user input with misspellings or colloquial phrasing

MoroAI enforces four distinct evaluation gates before certifying any checkpoint for production release.

The 4-Layer Validation Gates

Layer 1: Deterministic Syntax & Schema Gates Weight: 100% Mandatory

Validates completions against strict deterministic patterns. Checks include exact JSON Schema validation, regex constraints for mandatory IDs or formatting, and zero-tolerance scans for forbidden tokens or hallucinated URLs.

Layer 2: Semantic LLM Domain Judge Weight: Accuracy Benchmark

Executes reference-guided domain evaluation against ground-truth question-answer pairs. An evaluator model scores factual correctness, reasoning completeness, and tone fidelity on a calibrated 1-5 scale using structured JSON rubrics.

Layer 3: Wasserstein Embedding Drift Detection Weight: General Reasoning Guard

Calculates the Earth Mover's (Wasserstein-1) distance between output embedding representations of the fine-tuned model versus the base foundation model on standard anchor prompts. If drift exceeds $W_1 > 0.05$, the model is flagged for catastrophic forgetting.

Layer 4: Adversarial Perturbation Stress Testing Weight: Robustness Guard

Automatically injects synthetic perturbations into test prompts (character swaps, common keyboard typos, and homoglyphs). The harness verifies that the model's structured completion remains invariant under realistic noisy operator input.

Defining an Evaluation Suite (eval_suite.yaml)

name: "customer-support-eval"
version: "1.0.0"
gates:
  deterministic:
    min_pass_rate: 1.0        # 100% must pass JSON schema
  semantic:
    min_accuracy: 0.90         # 90% domain accuracy
  drift:
    max_wasserstein: 0.05      # No severe representation collapse
  adversarial:
    min_stability: 0.85        # Robust against typos

tests:
  - id: "test_refund_json"
    prompt: "Customer requests refund for order #12345 after 14 days."
    expected_schema:
      type: "object"
      required: ["eligible", "refund_amount", "reason_code"]
    forbidden_terms: ["I think", "maybe", "as an AI"]
  

Running Evaluations via CLI

# Run complete evaluation harness against candidate checkpoint
moro eval run --checkpoint ./runs/run_01/checkpoints/best/ --suite ./eval/eval_suite.yaml

# Compare candidate checkpoint against base foundation model
moro eval compare --base Qwen/Qwen2.5-1.5B --candidate ./runs/run_01/checkpoints/best/