Prediction of Future Stocks Returns Using an Ensemble of Machine Learning Models
Files
Anciaux_16421800_2025.pdf
Open access - Adobe PDF
- 2.58 MB
Details
- Supervisors
- Faculty
- Degree label
- Abstract
- Over the past few decades, it has proven extremely difficult to consistently outperform financial markets. For instance, about 90% of Large Cap Funds fail to beat the major benchmark indices over extended horizons. At the same time, the availability of financial data has exploded and machine learning (ML) methods have advanced significantly. This convergence offers a promising opportunity to revisit traditional approaches to return forecasting and portfolio construction. Drawing on empirical data from large-scale forecasting competitions and recent research on ensemble learning, this thesis examines whether combining multiple machine learning models can improve short-term stock return forecasting and enhance portfolio performance. Specifically, this thesis develops a set of ML classifiers to predict whether individual stocks will outperform their benchmark index over the following week. The models are trained using a walk-forward approach over a long historical period (January 2007-December 2024), allowing for continuous re-estimation to adapt to changing market conditions. Particular emphasis is placed on threshold selection strategies, which transform the continuous results from the combination into binary signals. These thresholds are optimized according to various criteria and are found to have a significant influence on trading performance. The predicted signals are integrated into a portfolio construction framework using various weights allocations techniques. Of the 544 portfolios constructed, all those that outperformed the benchmark did so using ensemble machine learning strategies. The best performing strategy combined a mean-variance portfolio allocation with a filtering-based covariance estimator and Ledoit-Wolf shrinkage, as well with a trimmed ensemble of ML models whose threshold was optimized to maximize precision. This configuration results in a Sharpe ratio of 1.11 and an APY of 19.29%. By comparison, its universe index achieved a Sharpe ratio of 0.74 and an APY of 14.72% over the same period. Moreover, 11 of the 15 best-performing portfolios used dynamic threshold techniques explicitly designed to maximize classification precision. Overall, this research demonstrates empirically the practical potential of ensemble learning and adaptive thresholding in financial forecasting and portfolio management.