← Back
Evidence source 4757Spot Checked

CN-Buzz2Portfolio: A Chinese-Market Dataset and Benchmark for LLM-Based Macro and Sector Asset Allocation from Daily Trending Financial News

arXiv2026-03-17Paper
Executive summary

CN-Buzz2Portfolio is a new benchmark for testing how Large Language Models (LLMs) make investment decisions in China using daily news trends. The study finds LLMs perform well in trend-driven markets but falter in low-volatility periods, sometimes losing to simple strategies. The Tri-Stage CPA Agent Workflow compresses and analyzes news to reduce noise. Surprisingly, smaller models can outperform larger ones in noisy markets. The research warns that too much news can harm decision quality and highlights benchmark limitations.

What it examines

This paper introduces CN-Buzz2Portfolio, a benchmark and dataset for testing large language models (LLMs) in Chinese financial markets. It maps daily trending news to macro and sector asset allocation, aiming to evaluate LLMs' reasoning and decision-making abilities in realistic, high-noise market environments.

What it concludes

Results show LLMs can generate logical asset allocations from news, but their performance depends on market conditions. The benchmark helps diagnose reasoning strengths and weaknesses. Applications include improving AI financial advisors and portfolio management tools. Future work should address regime adaptation, market frictions, and expand asset types for broader evaluation.

Extracted from this source

Evidence objects

Evidence 295582% extraction confidence
CN-Buzz2Portfolio debuts as a pioneering benchmark and dataset, evaluating how Large Language Models (LLMs) make asset allocation decisions in Chinas financial market using daily trending news for investment signals.

key_findings bullet 1 · key_findings · validation V0

Evidence 295682% extraction confidence
Surprisingly, LLMs excel in trend-driven markets but falter in sideways or low-volatility conditions, sometimes being outperformed by simpler, news-agnostic strategies; bigger models can even underperform smaller ones due to overfitting.

key_findings bullet 2 · key_findings · validation V0

Evidence 295782% extraction confidence
The Tri-Stage CPA Agent Workflow compresses, analyzes, and allocates news to reduce noise, but the study warns of information overload and notes limitations like focus on long-only ETFs and simplified trading assumptions.

key_findings bullet 3 · key_findings · validation V0

Evidence 295882% extraction confidence
CN-Buzz2Portfolio presents a novel benchmark and dataset for LLM-driven macro and sector asset allocation using Chinese financial news, shifting focus from entity-centric stock picking to market-narrative reasoning. Its Tri-Stage CPA Agent Workflow and rolling-horizon dataset offer unique, impactful tools for evaluating LLMs in emerging market investment management contexts.

key_findings bullet 4 · key_findings · validation V0

Raw abstract and provenance

Abstract: Large Language Models (LLMs) are rapidly transitioning from static Natural Language Processing (NLP) tasks including sentiment analysis and event extraction to acting as dynamic decision-making agents in complex financial environments. However, the evolution of LLMs into autonomous financial agents faces a significant dilemma in evaluation paradigms. Direct live trading is irreproducible and prone… ▽ More Large Language Models (LLMs) are rapidly transitioning from static Natural Language Processing (NLP) tasks including sentiment analysis and event extraction to acting as dynamic decision-making agents in complex financial environments. However, the evolution of LLMs into autonomous financial agents faces a significant dilemma in evaluation paradigms. Direct live trading is irreproducible and prone to outcome bias by confounding luck with skill, whereas existing static benchmarks are often confined to entity-level stock picking and ignore broader market attention. To facilitate the rigorous analysis of these challenges, we introduce CN-Buzz2Portfolio, a reproducible benchmark grounded in the Chinese market that maps daily trending news to macro and sector asset allocation. Spanning a rolling horizon from 2024 to mid-2025, our dataset simulates a realistic public attention stream, requiring agents to distill investment logic from high-exposure narratives instead of pre-filtered entity news. We propose a Tri-Stage CPA Agent Workflow involving Compression, Perception, and Allocation to evaluate LLMs on broad asset classes such as Exchange Traded Funds (ETFs) rather than individual stocks, thereby reducing idiosyncratic volatility. Extensive experiments on nine LLMs reveal significant disparities in how models translate macro-level narratives into portfolio weights. This work provides new insights into the alignment between general reasoning and financial decision-making, and all data, codes, and experiments are released to promote sustainable financial agent research. △ Less

Source row: 406 · abstract type: unknown