Reinforcement learning

Deep RL stock trading

A deep reinforcement learning system for daily stock trading, built on the FinRL framework with PPO agents, VGG and Cross-Stock Transformer feature extractors, and FinBERT sentiment over Polygon news. Single-seed backtests looked strong. Tested across random seeds and stock universes over the full test year, no configuration reliably beat an equal-weight buy-and-hold portfolio.

What I built

The system trained on daily data from 2020 through 2023 and was tested on 2024. I built it as a one-variable-at-a-time ablation, changing sentiment, data source, architecture, stock universe and starting capital in turn, for 24 models in total. The reward combines a Sharpe term and a sentiment term with drawdown and concentration penalties.

How the evaluation changed

The project went through three rounds of evaluation, each fixing a weakness in the one before it.

  • Original ablation. One seed per model, with test metrics computed only up to each model's peak portfolio value while buy-and-hold was measured over the full year. That window flatters any strategy by construction, so those comparisons aren't valid out-of-sample results.
  • Multi-seed architecture check. Three seeds for each of four architectures. Two of the four didn't reproduce their single-seed Sharpe ratios.
  • Universe-size sweep. 17 runs across five nested universes from 30 to 50 stocks, with every metric computed per run over the full test year. Partway through, I found the numbers I'd recorded were averages rather than per-run results, so I re-evaluated every run.

Results

UniverseRunsSharpe, mean ± stdBuy-and-holdRuns above
30 stocks30.68 ± 0.971.541
35 stocks51.35 ± 0.821.643
40 stocks31.53 ± 0.401.681
45 stocks31.24 ± 0.201.710
50 stocks31.33 ± 0.351.671
  • No universe size beat buy-and-hold on average, and only 6 of 17 runs beat their benchmark.
  • The random seed was the dominant variable. Seeds within one universe size spanned up to 2.0 Sharpe points, while the means across sizes spanned 0.85.
  • Universe size between 30 and 50 stocks was within seed noise.
  • A promising result dissolved. After three seeds, the 35-stock universe averaged 1.89 against a 1.64 benchmark. Two more seeds came in at 0.08 and 0.98, bringing the mean to 1.35.

Running it live

I deployed the 30-stock model to Alpaca paper trading on March 16, 2026, with daily scheduled execution, intraday stop-loss checks and end-of-day logging. Running it live surfaced problems the backtests didn't. Two sell paths in the trade-execution function each capped sales at the shares held, but together they could sell more than that and open short positions, so I fixed the position bookkeeping and confirmed the broker-side no-shorting setting as a backstop. I also added a 10% per-position cap and made the order scanner re-check live buying power before each order. I shut the deployment down in June 2026, once the multi-seed results showed there was no validated model to run.

What I'd carry forward

  • Measure the model and the benchmark over the same full window.
  • Use multiple seeds before comparing designs. In a system like this, one seed is a single draw from a wide distribution.
  • When a result looks promising, add seeds instead of re-running the weak ones.
  • Record per-run metrics, and keep aggregation separate from recording.