← Back
Finding 6761Emerging EvidenceValidation V0

A study uses deep reinforcement learning (PPO) to upgrade a classic 60/40 portfolio strategy by introducing a unique negative Sharpe regret reward that compares agent choices against an Oracle-optimal allocation.

86%Confidence
1Evidence objects
v1Version
DraftStatus

Evidence trail

Knowledge status

This Finding was extracted from the configured corpus. It is versioned, traceable, and may evolve through editorial review or new corpus evidence.