Exploratory Control with Tsallis Entropy for Latent Factor Models

DOI10.1137/22M153505XMaRDI QIDQ6200515zbMATH OpenWikidataFDO

Authors Ryan Donnelly, Sebastian Jaimungal

Publication date 22 March 2024

Published in SIAM Journal on Financial Mathematics (Search for Journal in Brave)

Full work available at URL https://arxiv.org/abs/2211.07622

zbMATH Keywords

stochastic control reinforcement learning entropy regularization exploratory control

Mathematics Subject Classification ID

Measures of information, entropy (94A17) Optimal stochastic control (93E20)

Abstract: We study optimal control in models with latent factors where the agent controls the distribution over actions, rather than actions themselves, in both discrete and continuous time. To encourage exploration of the state space, we reward exploration with Tsallis Entropy and derive the optimal distribution over states - which we prove is

q

-Gaussian distributed with location characterized through the solution of an FBS

D e l t a

E and FBSDE in discrete and continuous time, respectively. We discuss the relation between the solutions of the optimal exploration problems and the standard dynamic optimal control solution. Finally, we develop the optimal policy in a model-agnostic setting along the lines of soft

Q

-learning. The approach may be applied in, e.g., developing more robust statistical arbitrage trading strategies.

Recommendations

Cites work

This page was built for publication: Exploratory Control with Tsallis Entropy for Latent Factor Models

Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q6200515)