05. I tried reinforcement learning too
A TradeMaster-related experiment, five seeds, and a small gain that did not survive its checks.
Investing on my own · Part 5/22 · Evidence through 2026-09-22. Historical research is distinct from DRY and live execution.
After writing explicit trading rules, it was natural to wonder whether a model could learn when to take risk instead. That was part of my interest in reinforcement learning.
I had already explored TradeMaster-related code. In late July, I asked for a more concrete experiment using the crypto data available to me. I wanted to understand the assumptions in the existing implementation as well as try a newer method.
🎮 Choosing a position
In reinforcement learning, an agent observes a state, chooses an action, and learns from the outcome. Here, the state included past returns, volume, moving-average gaps, and funding. The action represented the amount of market exposure, and the reward included trading costs.
I limited the action space to five target positions: fully short, half short, cash, half long, and fully long. Before considering elaborate orders, I wanted to test a smaller question about how much risk to take.
The experiment used BTC and ETH for training and 34 features constructed from historical information. I trained five versions with different random seeds and examined them together. A single fortunate initialization was not enough.
The existing code needed scrutiny first. Some paths observed a whole bar and traded at that same bar’s close. Another used future multi-bar returns in reward shaping. A named research method did not make those assumptions appropriate for my intended evaluation.
I built an isolated experiment that decided after a completed bar and entered at the following open. I also checked whether a full-position purchase could quietly leave cash negative after fees.
🧪 Trying assets outside the training pair
I separated the data used to choose the model from the final evaluation. Before opening the last basket, I wrote down AVAX, DOT, TRX, and ETC along with the acceptance criteria. The evaluation covered March 2025 through June 2026.
With a cost of 10 basis points per one-times position change, the basket returned +0.41%. Increasing the cost to 25 basis points changed that to -0.72%. A slightly positive baseline was not very reassuring when a modest cost increase reversed it.
The asset breakdown was more informative. AVAX returned -1.03%, DOT -12.25%, TRX stayed in cash at 0%, and ETC returned +13.88%. ETC almost offset DOT’s loss.
That was not broad success across four assets. It was a small aggregate gain supported by one asset. Exposure was also sparse, so the experiment had not demonstrated repeated performance across a large number of independent market episodes.
Removing one seed at a time produced a median return of -1.07%. The lower end of the block-bootstrap interval was -6.45%. These checks made the weak aggregate result harder to dismiss as a minor cost issue.
🤔 What about the better-looking results?
A separate replay on eight major assets returned +5.11% at 10 basis points and +2.75% at 25 basis points. Those numbers could have supported a more upbeat description.
However, some assets and the period had already been inspected during development. That basket could provide additional diagnostics, but it could not be described as an untouched final test.
Four of the six preregistered acceptance criteria failed. The recorded verdict was FAIL for research promotion, paper trading, and live trading. This experiment submitted no orders.
The method itself also had a limited interpretation. It calculated what different historical positions would have earned while assuming that my small trades did not affect the price path. It did not learn how my orders would change the book or how much would actually fill.
That distinction matters when discussing “offline RL.” The experiment was a fitted-Q procedure over counterfactual historical returns, rather than a complete market simulator with observed outcomes for every possible action.
During the work, I asked whether the existing repository had actually been improved or whether a separate experiment had simply been added beside it. The answer was the latter. A useful research prototype is not automatically an integrated trading component.
What I took from the exercise was a more restrained expectation of model names. Whatever method chooses the position, it still has to survive different assets, later periods, and the cost of changing its mind.
This version did not clear that bar. I kept the failed evaluation alongside the more attractive diagnostic results, because both were necessary to understand what had actually happened.
Previous · Series index · Next