Tyler
Hobbs

Applied machine learning

I build models, then check whether their numbers mean what they appear to. Four years in proteomics and genomics labs, now finishing an M.S. in Data Science at UVA. My projects tend to end with an audit.

Dec 2026
M.S. Data Science, UVA
4 yrs
Genomics & proteomics labs
3
Research projects shipped
1
Award, UVA Ophthalmology
01About

From the lab
bench to model evaluation

Clinical Laboratory Technician II

Illumina
2026 — Present · Boulder, CO

Clinical genomics workflows under regulated quality systems, where every reported result has to trace back to a validated measurement.

Laboratory Associate II

Standard BioTools (formerly SomaLogic)
2022 — 2026 · Boulder, CO

Ran high-throughput proteomics assays and traced sources of run-to-run variation. Where the habit of separating real effects from noise started.

02Stack

Tools I've
actually shipped with

Languages

What I write day to day.

PythonSQLRBashHTML / CSS

ML & DL frameworks

What I train models with.

PyTorchTensorFlowscikit-learnHugging FaceStable-Baselines3FinRLtidymodels

ML techniques

Methods I've worked in.

Reinforcement learning (PPO)TransformersCNNsSemantic segmentationLoRA / PEFTGRPOFinBERTAttention mechanisms

Data & visualisation

Getting it clean, then showing it.

pandasNumPyPlotly / DashD3.jsMatplotlibPower BITableau

Infrastructure & deployment

Where it runs.

FlaskGitRenderHPC (SLURM)REST APIsGradio

Currently exploring

What I'm learning next.

JavaScriptReactNext.js
03Projects

Three projects,
each ending
in an audit

LLM fine-tuning

LoadBrief

A LoRA fine-tune of Llama 3 8B that writes structured athlete load-management briefs. It scored 0.960 on risk classification — then a bag-of-words baseline scored 0.950, and the paper became about why.

Llama 3 8BLoRALLM-as-judgeRivanna HPC
Base model0.000
TF-IDF baseline0.950
LoRA fine-tune0.960
Exact risk-classification accuracy
Computer vision · Award winner

Automated glaucoma screening

Optic disc and cup segmentation across four public datasets. Strong on benchmarks at 0.844 Dice, it fell to 0.251 on real UVA clinic images. Hybrid training closed part of the gap.

PyTorchU-Net++Domain shiftGradio
Public test0.844
Clinical, zero-shot0.251
Clinical, hybrid0.330
Dice score, public vs. clinical images
Reinforcement learning

Deep RL stock trading

PPO agents with FinBERT news sentiment, deployed to live paper trading. Single-seed backtests looked strong; across 17 runs the seed moved Sharpe more than any design choice, and nothing beat buy-and-hold.

FinRLPPOFinBERTMassive, formerly Polygon.io(Polygon.io)Multi-seed evaluation
30 stocks0.68
35 stocks1.35
40 stocks1.53
Buy & hold1.68
Mean test Sharpe by universe size
04Achievements

Awards, papers
and talks

Most Innovative Analytical Solution

UVA Ophthalmology capstone award, 2026

Award

Label provenance and metric validity in LLM fine-tuning

LoadBrief — how a 0.960 accuracy turned out to be recoverable by a bag-of-words baseline

Paper
Read the paper

Automated glaucoma screening using AI-enhanced ophthalmoscopy

Optic disc and cup segmentation, and the public-to-clinical domain gap

Paper
Read the paper

Deep reinforcement learning for automated stock trading

PPO with news sentiment, and what multi-seed evaluation did to the result

Paper
Read the paper
05Contact

Let's talk about
measurement

Happy to talk about any of this work. The fastest way to reach me is email.