Deep reinforcement trading with predictable returns
From MaRDI portal
Abstract: Classical portfolio optimization often requires forecasting asset returns and their corresponding variances in spite of the low signal-to-noise ratio provided in the financial markets. Modern deep reinforcement learning (DRL) offers a framework for optimizing sequential trader decisions but lacks theoretical guarantees of convergence. On the other hand, the performances on real financial trading problems are strongly affected by the goodness of the signal used to predict returns. To disentangle the effects coming from return unpredictability from those coming from algorithm un-trainability, we investigate the performance of model-free DRL traders in a market environment with different known mean-reverting factors driving the dynamics. When the framework admits an exact dynamic programming solution, we can assess the limits and capabilities of different value-based algorithms to retrieve meaningful trading signals in a data-driven manner. We consider DRL agents that leverage classical strategies to increase their performances and we show that this approach guarantees flexibility, outperforming the benchmark strategy when the price dynamics is misspecified and some original assumptions on the market environment are violated with the presence of extreme events and volatility clustering.
Cites work
- 60 years of portfolio optimization: practical challenges and current trends
- \({\mathcal Q}\)-learning
- Algorithms for reinforcement learning.
- Dynamic programming and optimal control. Vol. 1.
- Empirical properties of asset returns: stylized facts and statistical issues
- End-to-end training of deep visuomotor policies
- scientific article; zbMATH DE number 5348356 (Why is no real title available?)
- scientific article; zbMATH DE number 108112 (Why is no real title available?)
- Kernel-based reinforcement learning
- Learning near-optimal policies with Bellman-residual minimization based fitted policy iteration and a single sample path
- Market Microstructure Invariance: Empirical Hypotheses
- Reinforcement learning. An introduction
- Simulation-based optimization of Markov reward processes
- Technical update: Least-squares temporal difference learning
This page was built for publication: Deep reinforcement trading with predictable returns
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6098411)