Supervised Learning Models, Statistical Models or Hybrid Models? A Prediction of Clean Energy Stock Based on Fear and Fundamental Factors
Study compares statistical, supervised, and hybrid models, integrating media features and simulated data, for forecasting the S&P Clean Energy Index.
What it examines
This paper presents a forecasting study for the S&P Clean Energy Index using historical data and media sentiment, applying statistical models (ARIMA, VAR), machine learning (LSTM, SAE) and hybrid methods (Prophet-based models). The study aims to address short, medium, and long-term market predictions incorporating fear and fundamental factors.
What it concludes
The research concludes that Prophet-based models, aided by Random Forest simulated data, perform best in predicting trends for the S&P Clean Energy Index. These methods can assist investors and policymakers in decision-making, with potential applications in financial forecasting, clean energy market analysis, and guiding environmental policy and sustainable investments.
Evidence objects
Traditional models like ARIMA and VAR provide a robust linear baseline, but combined with feature enhancements using Prophet and LSTM, they accurately capture non-linear trends and improve forecast reliability remarkably.
key_findings bullet 1 · key_findings · validation V0
Innovative integration of media sentiment with financial indicators, supported by XGBoost, SHAP, and PCA techniques, enhances forecast accuracy while explaining significance, offering novel solutions for data scarcity challenges across markets.
key_findings bullet 2 · key_findings · validation V0
New proposals use simulated future data via Random Forest and LSTM to address gaps. Prophet-based models excel with simulated data, yet LSTM risks overfitting, emphasizing improvements in model granularity significantly.
key_findings bullet 3 · key_findings · validation V0
Presented is a comprehensive study comparing supervised, statistical, and hybrid methods for predicting clean energy stock prices by incorporating fear and fundamental factors. Employing techniques including SHAP validation and PCA, the paper integrates diverse methodologies. Its originality, methodological rigor, and targeted focus deliver truly compelling insights for quantitative finance researchers.
key_findings bullet 4 · key_findings · validation V0
Raw abstract and provenance
This paper explores several time series models for predicting the S&P Clean Energy Index. We begin by identifying factors previously found to influence the clean energy market and use eXtreme Gradient Boosting (XGBoost) to rank and filter feature importance of variables, followed by further validation and variable selection using SHAP values. Next, we simulate future feature data using methods like Random Forest and Long Short-Term Memory (LSTM). For the LSTM-based simulations, the data is generated through a classification-then-prediction approach using the K-Nearest Neighbors (KNN) algorithm. To predict the index’s volatility, we employ statistical models such as AutoRegressive Integrated Moving Average (ARIMA) and Vector Autoregression (VAR). Additionally, we use advanced methods like LSTM, Supervised Autoencoder (SAE), and hybrid models such as Prophet Features and LSTM_Autoregressive(LSTM_AR). Each model’s parameter-tuning process will be explained in detail. Finally, we compared the models’ performance and prediction results, discussing their strengths and suitability for different scenarios. We confirmed that the Prophet-based models performed well on Random Forest simulated data when predicting both the trend and actual values of the S&P clean energy index.
Source row: 1878 · abstract type: unknown