A deep reinforcement learning system for daily stock trading, built on the FinRL framework with PPO agents, VGG and Cross-Stock Transformer feature extractors, and FinBERT sentiment over Polygon news. Single-seed backtests looked strong. Tested across random seeds and stock universes over the full test year, no configuration reliably beat an equal-weight buy-and-hold portfolio.
The system trained on daily data from 2020 through 2023 and was tested on 2024. I built it as a one-variable-at-a-time ablation, changing sentiment, data source, architecture, stock universe and starting capital in turn, for 24 models in total. The reward combines a Sharpe term and a sentiment term with drawdown and concentration penalties.
The project went through three rounds of evaluation, each fixing a weakness in the one before it.
| Universe | Runs | Sharpe, mean ± std | Buy-and-hold | Runs above |
|---|---|---|---|---|
| 30 stocks | 3 | 0.68 ± 0.97 | 1.54 | 1 |
| 35 stocks | 5 | 1.35 ± 0.82 | 1.64 | 3 |
| 40 stocks | 3 | 1.53 ± 0.40 | 1.68 | 1 |
| 45 stocks | 3 | 1.24 ± 0.20 | 1.71 | 0 |
| 50 stocks | 3 | 1.33 ± 0.35 | 1.67 | 1 |
I deployed the 30-stock model to Alpaca paper trading on March 16, 2026, with daily scheduled execution, intraday stop-loss checks and end-of-day logging. Running it live surfaced problems the backtests didn't. Two sell paths in the trade-execution function each capped sales at the shares held, but together they could sell more than that and open short positions, so I fixed the position bookkeeping and confirmed the broker-side no-shorting setting as a backstop. I also added a 10% per-position cap and made the order scanner re-check live buying power before each order. I shut the deployment down in June 2026, once the multi-seed results showed there was no validated model to run.