Subsystem 05

Closed-Loop DPO Feedback Flywheel

Transforming live production interactions and human operator edits into preference pairs for continuous local model alignment.

Why Static Models Degrade Over Time

When a model is fine-tuned and deployed into an enterprise environment, its real-world performance is tested by evolving edge cases, novel customer questions, and shifting operational terminology. In typical cloud-based setups, closing this loop requires shipping sensitive conversation transcripts to external annotation vendors.

MoroAI creates a local, sovereign learning flywheel. Human operators naturally correct model answers or submit thumbs-up / thumbs-down ratings during daily operations. The MoroAI Flywheel mines these local signals and formats them into high-signal Direct Preference Optimization (DPO) datasets without data ever leaving the workstation or VPC.

Direct Preference Optimization (DPO) Mechanics

Traditional RLHF requires training a separate reward model followed by complex PPO reinforcement learning, which is notoriously unstable and memory-intensive on consumer GPUs. DPO optimizes the policy model directly on preference pairs using the exact analytical relationship between policy and reward:

DPO Closed-Form Objective Optimization Math
L_DPO(π_θ; π_ref) = -E_(x, y_w, y_l) [
    log σ( β * log( π_θ(y_w | x) / π_ref(y_w | x) ) - β * log( π_θ(y_l | x) / π_ref(y_l | x) ) )
]

Where:
  x     = The user prompt
  y_w   = Chosen response (human operator edit or high-rated completion)
  y_l   = Rejected response (original unedited model completion)
  π_ref = Frozen reference model (initial fine-tuned LoRA checkpoint)
  β     = Temperature hyperparameter controlling deviation penalty (default 0.1)
    

The 4-Step Flywheel Lifecycle

  1. Local Interaction Logging: As queries pass through the local Ollama or vLLM proxy, input prompts and completions are securely persisted to local encrypted SQLite storage (.moro/moro.db).
  2. Preference Pair Extraction: When an operator manually corrects an output in Mission Control or an integrated CRM, the engine constructs a preference pair with the corrected text as chosen and the initial output as rejected.
  3. Confidence Margin Filtering: To eliminate ambiguous feedback, pairs where the operator edit is merely trivial whitespace or where rating confidence is below $\Delta = 0.3$ are automatically pruned.
  4. Automated Offline Alignment: When a user-configured threshold of validated pairs is reached (e.g. 250 pairs), the engine schedules an automated overnight DPO alignment run, verifies the model against the evaluation harness, and generates a new versioned release candidate.

CLI Mining & Alignment Commands

# Mine interaction logs for high-confidence preference pairs
moro flywheel process \
    --interactions ./logs/chat_ops.jsonl \
    --min-margin 0.3 \
    --output ./data/dpo_pairs.jsonl

# Execute local DPO fine-tuning alignment round
moro flywheel align \
    --base-checkpoint ./runs/run_01/checkpoints/best/ \
    --pairs ./data/dpo_pairs.jsonl \
    --beta 0.1 \
    --output ./runs/dpo_aligned/