← Back
Evidence source 6023Spot Checked

Reinforcement Learning for Risk-Sensitive Investment Management: a Free Energy--Entropy Duality Approach

arXiv2026-06-18Paper
Executive summary

Researchers present a new reinforcement learning method for risk-sensitive investment, using free energy and entropy duality to recast asset allocation as a linear-quadratic-Gaussian (LQG) game. The approach yields clear, interpretable solutions for both short and long-term portfolios, even with incomplete market models. Notably, the RL algorithm learns optimal allocations quickly and accurately, while adversarial robustness is less precise but not critical. Real U.S. equity data supports the method, though model-free cases and transaction costs remain unaddressed.

What it examines

This paper introduces a reinforcement learning approach for continuous-time, risk-sensitive asset allocation, using free energy--entropy duality to reformulate the problem as a stochastic differential game. The method enables learning optimal investment strategies from data, even when model parameters are uncertain or partially known.

What it concludes

The study shows that the proposed actor--critic algorithm can accurately learn optimal investment policies, with clear economic interpretation. This approach is useful for asset managers seeking robust, explainable strategies under uncertainty. Future work may extend these methods to more complex markets or further improve learning efficiency and interpretability.

Extracted from this source

Evidence objects

Evidence 677782% extraction confidence
Researchers unveil a novel reinforcement learning method for risk-sensitive investment, reformulating asset allocation as a linear-quadratic-Gaussian (LQG) game using free energy--entropy duality, enabling explicit, interpretable solutions for various investment horizons.

key_findings bullet 1 · key_findings · validation V0

Evidence 677882% extraction confidence
The approach features continuous-time actor--critic algorithms, fractional Kelly strategies for economic clarity, and decomposes allocations into Kelly, benchmark-tracking, and intertemporal hedging components, offering portfolio managers direct learning from dataeven with partial market models.

key_findings bullet 2 · key_findings · validation V0

Evidence 677982% extraction confidence
Proof-of-concept on U.S. equity data shows rapid, accurate learning of optimal policies, but the models reliance on partial market knowledge and omission of transaction costs may limit real-world applicability despite its theoretical and practical promise.

key_findings bullet 3 · key_findings · validation V0

Evidence 678082% extraction confidence
This paper introduces a novel duality-based reformulation using free energy--entropy concepts for continuous-time risk-sensitive asset allocation, recasting it as a linear-quadratic-Gaussian stochastic differential game. Its actor--critic RL method, guided by analytical solutions and fractional Kelly decompositions, uniquely enhances interpretability and practical relevance, making it compelling and original.

key_findings bullet 4 · key_findings · validation V0

Raw abstract and provenance

This paper develops a reinforcement-learning approach to continuous-time risk-sensitive benchmarked asset allocation in a partly model-based setting. The benchmarked problem does not directly fit the standard Markovian stochastic-control template: the state is uncontrolled, whereas the terminal reward contains a controlled Itô integral. We use free energy-entropy duality to reformulate the problem as a linear-quadratic-Gaussian stochastic differential game under an equivalent probability measure, yielding explicit finite- and infinite-horizon saddle-point solutions. This structure guides a continuous-time $q$-learning actor-critic method: the quadratic value function motivates the critic, while the affine saddle-point controls motivate deterministic actors for the portfolio allocation and adversarial control. The learned allocation admits an economic interpretation through fractional Kelly decompositions. A proof-of-concept implementation calibrated to U.S. equity data shows that the actors learn the optimal policy with high accuracy and reveals a favorable asymmetry: the portfolio actor receives a cleaner learning signal than the auxiliary adversarial actor.

Source row: 1672 · abstract type: unknown