← Back
Evidence source 5204Spot Checked

FinMultiTime: A Four-Modal Bilingual Dataset for Financial Time-Series Analysis

arXiv2025-06-06Paper
Executive summary

FinMultiTime aligns four modalities: news, tables, K-line charts and prices across the S&P 500 and HS 300 from 2009 to 2025. The bilingual 112.6 GB dataset spans minute, daily and quarterly resolutions. Experiments show larger data and Transformer models boost predictions to R²≈0.97 on 35 stocks. Small sets treat extra modalities as noise; multimodal fusion in strong learners yields gains. The team provides a reproducible pipeline with GPT-4.1 trend labels, LSA sentiment scoring and continuous updates, establishing benchmark.

What it examines

The paper introduces FinMultiTime, the first large-scale bilingual dataset for financial time-series analysis. It aligns four modalities—stock prices, news text, technical candlestick charts, and financial tables—across S&P 500 and HS 300 markets from 2009 to 2025. The dataset’s high-resolution, multimodal format supports richer and more accurate forecasting models.

What it concludes

Experiments show that dataset scale and quality greatly improve prediction accuracy, and fusing modalities yields further gains in Transformer-based models. The study highlights applications in sentiment analysis, trend forecasting, anomaly detection, and generative financial AI. Future work includes continuous data updates and enhanced multimodal model development.

Extracted from this source

Evidence objects

Evidence 426978% extraction confidence
FinMultiTime offers a bilingual, 112.6 GB dataset aligning news text, financial tables, K-line charts and price series across S&P 500 and HS 300 (2009--2025) at minute, daily and quarterly resolutions across granular timeframes.

key_findings bullet 1 · key_findings · validation V0

Evidence 427078% extraction confidence
Larger, high-quality datasets improve forecasting accuracy, with Transformer models outperforming RNN, LSTM, GRU, CNN and TimesNet to achieve $R^2\approx0.97$ on 35-stock sets, while smaller data treat extra modalities as noise.

key_findings bullet 2 · key_findings · validation V0

Evidence 427178% extraction confidence
The paper establishes a reproducible pipelineweb scraping, XBRL parsing, chart generationintroducing GPT-4.1 trend labels and LSA-based sentiment, coining the large four-modal financial benchmark despite fixed sentiment granularity and tuning limits.

key_findings bullet 3 · key_findings · validation V0

Evidence 427278% extraction confidence
This paper introduces a large bilingual multimodal financial time-series dataset, uniquely spanning diverse markets and modalities. Its originality lies in comprehensive data scale and cross-lingual, multimodal integration. Although not algorithmic, this resource offers fresh foundations for AI-driven trading research, substantially advancing forecasting benchmarks and enabling novel investigations and practical applications.

key_findings bullet 4 · key_findings · validation V0

Raw abstract and provenance

- … (LLMs) have made significant strides in capturing sentiment and other qualitative signals, thereby enhancing the accuracy of financial time-series predictions… involve LLMs …

Source row: 853 · abstract type: snippet