LLM fine-tuning

LoadBrief

LoadBrief turns a free-text athlete monitoring summary — training load, heart-rate variability and wellness scores — into a structured load-management brief written for an athlete, a coach or a sports scientist. The fine-tuned model reached 0.960 exact risk-classification accuracy, up from 0.000 for the untuned base model. Auditing that number became the real project.

What I built

A rule-based simulator generates the corpus: athlete profiles, monitoring time series, narrative summaries and reference briefs in three audience registers. The model is a LoRA fine-tune of Llama 3 8B Instruct, trained on UVA's Rivanna cluster through eight revisions of the corpus with the same supervised fine-tuning configuration each time. I also explored a GRPO variant, which isn't part of the release.

Evaluation combines rule-based metrics, a composite reward and an LLM-as-judge that I calibrated by scoring the ground-truth briefs themselves. A single verification script re-derives every table in the paper from the released artifacts and exits non-zero if any check fails.

What I found

Rule-generated data isn't self-validating, and neither are the metrics used to score models trained on it. Five findings, each measured against the released artifacts:

  • Both labels are constant per scenario. Risk level and overreaching class are written from the scenario definition rather than computed from the sampled signals, so each of the 19 scenarios maps to exactly one value. A TF-IDF classifier recovers the risk label from the narrative at 0.950, within a point of the fine-tune, so the headline accuracy can't be evidence of clinical reasoning.
  • 16.1% of records contradict themselves. 2,412 pair critically suppressed heart-rate variability with a low-risk header, and 4,021 name one overreaching class in the narrative and a different one in the classification section.
  • Three consistency guards pass every one of those records, each for a different structural reason. Passing all three isn't evidence of consistency.
  • The composite reward is mis-specified against its own reference data. Its 0.701 ceiling comes from components the ground-truth briefs can't earn, and an untuned model already scores 0.422, leaving a usable range of 0.279.
  • The overreaching metric can't read 45% of correct answers. What looked like a capability loss across seven revisions was a change in the generator, not the model.

Why I released it with its defects

Rather than start another fix cycle, which would likely have led to another after it, I released the corpus and model with their defects documented, so the two stay consistent with each other. The repository and dataset card list each defect and how many records it affects.

The paper recommends seven inexpensive checks, each of which would have caught a defect this project carried through multiple training cycles: audit whether declared generation targets are reachable, state where labels come from, validate extraction metrics against reference outputs, decompose composite rewards against ground truth, calibrate model-based judges against ground truth, keep a guard's scope independent of the data it inspects, and pair deterministic with model-based evaluation.

Limitations

LoadBrief is a research artifact, not a medical device. The corpus is synthetic, and the released model reproduces its documented defects, so neither should inform decisions about a real person's training or health.