Deterministic Policy Gradient for Learning Equilibrium in Time-Inconsistent Control Problems
Researchers present DPG-FPI, a new model-free reinforcement learning algorithm for time-inconsistent control in continuous time, a major challenge in finance and mathematics. The method reformulates these problems into a two-stage fixed-point structure, enabling deterministic policy gradient methods. Using extended Hamilton-Jacobi-Bellman equations, DPG-FPI learns equilibrium policies without full model knowledge. It outperforms entropy-regularized q-learning in speed and accuracy for financial tasks, but its effectiveness in other fields and scalability remain untested.
What it examines
This paper introduces a new reinforcement learning algorithm to find equilibrium policies in time-inconsistent control problems, common in finance. By reformulating the problem into two stages and using deterministic policy gradients, the method learns optimal strategies without needing full knowledge of the system's details.
What it concludes
The proposed algorithm effectively learns equilibrium policies in complex financial settings, outperforming existing methods in accuracy and stability. It is especially useful for portfolio management and tracking problems with non-standard preferences. Future work may extend this approach to broader applications and address remaining challenges in time-inconsistent decision-making.
Evidence objects
Researchers unveil DPG-FPI, a model-free reinforcement learning algorithm that cracks time-inconsistent control problems in continuous time, a long-standing challenge in finance and mathematics, using a novel two-stage fixed-point structure.
key_findings bullet 1 · key_findings · validation V0
DPG-FPI leverages deterministic policy gradients, extended Hamilton-Jacobi-Bellman equations, and martingale characterizations to learn equilibrium policies and future preference functionswithout full model knowledgeoutperforming entropy-regularized q-learning in speed, accuracy, and robustness.
key_findings bullet 2 · key_findings · validation V0
While DPG-FPI excels in financial tasks like mean-variance portfolio management and non-exponential discounting, its scalability and effectiveness in other domains remain untested, highlighting opportunities and open questions for future research.
key_findings bullet 3 · key_findings · validation V0
This paper tackles time-inconsistent control in quantitative finance, introducing a continuous-time, model-free reinforcement learning algorithm using deterministic policy gradients. Its two-stage reformulation and actor-critic iterations are novel, with theoretical convergence guarantees. The originality lies in methodological advancements, offering compelling impact for mean-variance portfolio optimization and market prediction applications.
key_findings bullet 4 · key_findings · validation V0
Raw abstract and provenance
In this paper, we develop a continuous-time model-free reinforcement learning algorithm to learn deterministic equilibrium policies in general time-inconsistent control problems. Utilizing the extended Hamilton-Jacobi-Bellman system, we recast the original time-inconsistent problem into an equivalent two-stage problem. In the first stage, for given auxiliary functions, we employ the deterministic policy gradient approach to learn an optimal policy in an auxiliary time-consistent control problem. In the second stage, given the updated policy, we exploit the inner fixed point iterations and some martingale characterizations to learn the auxiliary functions. As a theoretical contribution, we provide some mild model assumptions and establish the convergence of inner fixed point iterations. By repeating this actor-critic style of iterations across two stages, our algorithm aims to learn the equilibrium under different sources of time-inconsistency in a unified manner. The superior effectiveness of the proposed algorithm are illustrated in two classical financial applications with time-inconsistency: mean-variance portfolio management and optimal tracking portfolio under non-exponential discounting.
Source row: 594 · abstract type: unknown