← Back
Evidence source 6502Spot Checked

When AI Trading Agents Compete: Adverse Selection of Meta-Orders by Reinforcement Learning-Based Market Making

arXiv2025-10-31Paper
Executive summary

Researchers show that reinforcement learning agents, acting as high-frequency market makers, can profit from medium-frequency traders executing large orders in a simulated market. Using a Hawkes process to model the order book, agents trained with Proximal Policy Optimisation and self-imitation learning exploit predictable price moves. Notably, agent profits rise without significantly increasing costs for slower traders. The study highlights the agent’s asymmetric learning and the need for more advanced trading strategies to better reflect real-world conditions.

What it examines

This study uses a simulated market to explore how high-frequency trading agents, trained with reinforcement learning, interact with medium-frequency traders executing large orders. By modeling realistic market dynamics, the research aims to understand and replicate adverse selection, where faster traders profit from the predictable actions of slower traders.

What it concludes

The research shows that AI market makers can profit from predictable trading patterns without always increasing costs for slower traders. These findings can help design better trading strategies and defenses against adverse selection. Future work includes testing more advanced strategies and improving AI agents to handle different market situations.

Extracted from this source

Evidence objects

Evidence 849478% extraction confidence
Reinforcement learning agents, trained with Proximal Policy Optimisation and self-imitation, act as high-frequency market makers, exploiting predictable price moves from medium-frequency traders using simple strategies like TWAP in simulated markets.

key_findings bullet 1 · key_findings · validation V0

Evidence 849578% extraction confidence
Surprisingly, RL agents boost their profitsespecially against buy-side meta-orderswithout significantly increasing slippage costs for slower traders, suggesting smarter liquidity provision rather than direct harm to medium-frequency participants.

key_findings bullet 2 · key_findings · validation V0

Evidence 849678% extraction confidence
The study introduces 'unaware' and 'fully informed' RL agents, finding that fully informed agents adapt better to buy-side flows but struggle with sell-side ones, highlighting asymmetric learning and real-world replication challenges.

key_findings bullet 3 · key_findings · validation V0

Evidence 849778% extraction confidence
This paper uniquely integrates reinforcement learning (PPO, self-imitation) with Hawkes process-based limit order book simulation, advancing endogenous market impact modeling. Its impulse control RL framework and adversarial setup between RL HFTs and MFT meta-orders offer original insights into adverse selection, making it compelling for market microstructure and trading strategy research.

key_findings bullet 4 · key_findings · validation V0

Raw abstract and provenance

Abstract: We investigate the mechanisms by which medium-frequency… ▽ More We investigate the mechanisms by which medium-frequency trading agents are adversely selected by opportunistic high-frequency traders. We use reinforcement learning (RL) within a Hawkes Limit Order Book (LOB) model in order to replicate the behaviours of high-frequency market makers. In contrast to the classical models with exogenous price impact assumptions, the Hawkes model accounts for endogenous price impact and other key properties of the market (Jain et al. 2024a). Given the real-world impracticalities of the market maker updating strategies for every event in the LOB, we formulate the high-frequency market making agent via an impulse control reinforcement learning framework (Jain et al. 2025). The RL used in the simulation utilises Proximal Policy Optimisation (PPO) and self-imitation learning. To replicate the adverse selection phenomenon, we test the RL agent trading against a medium frequency trader (MFT) executing a meta-order and demonstrate that, with training against the MFT meta-order execution agent, the RL market making agent learns to capitalise on the price drift induced by the meta-order. Recent empirical studies have shown that medium-frequency traders are increasingly subject to adverse selection by high-frequency trading agents. As high-frequency trading continues to proliferate across financial markets, the slippage costs incurred by medium-frequency traders are likely to increase over time. However, we do not observe that increased profits for the market making RL agent necessarily cause significantly increased slippages for the MFT agent. △ Less

Source row: 2151 · abstract type: unknown