Generated 2026-08-07 13:40 UTC from training/eval_history.json. Re-run python training/generate_report.py after evaluating new checkpoints to refresh this report.
This report evaluates whether the synthetic mozu-research-corpus transcript dataset carries
enough real signal to train a model to perform a clinically meaningful task, by measuring how
much fine-tuning on it improves an open-source model's output on a held-out test set. The
underlying question is about dataset quality, not model selection: if fine-tuning on this