Base-Learner Correlation Constrains The Benefit Of Consensus Machine Learning For Radiosonde-Based Convective Forecasting In The Tropics

Dima Soraya, Aulia Khamas Heikmakhtiar, Leli Setyaningrum

Abstract


Ensemble and consensus methods are widely assumed to improve convective forecasting skill, yet the conditions under which that assumption holds are rarely tested. This study evaluates four consensus schemes built on three algorithmically diverse base learners for 0–12 h rain and thunderstorm prediction from radiosonde-derived instability indices at Soekarno-Hatta International Airport (WMO 96749), Indonesia. In total, 5,839 soundings (2016–2025) were converted into twelve instability indices and labelled against 173,865 METAR reports using an exclusive-lower-bound forecast window. Random Forest, Support Vector Machine, and Gradient Boosting were combined through hard voting, soft voting, skill-weighted voting, and regularised-weight optimisation on a chronologically partitioned, leakage-controlled pipeline. Peak critical success index (CSI) reached 0.460 for rain and 0.428 for thunderstorm. Consensus clearly outperformed a climatological random predictor (permutation test, p < 0.0001) but showed no significant advantage over the best single model (McNemar p = 1.0000 and p = 0.2807; paired-bootstrap ΔCSI intervals spanning zero). We attribute this plateau to high inter-model prediction correlation (Pearson r = 0.753–0.971), which leaves little independent error for voting to exploit and which follows from training all learners on an identical feature space. Precipitable water was the dominant predictor for both targets. Ensemble diversity in this problem is therefore limited by information source rather than algorithm choice, and multimodal inputs are more promising than additional combination rules.

Keywords


consensus machine learning; ensemble diversity; radiosonde; thunderstorm nowcasting; forecast verification; tropical convection.

Full Text:

PDF

References


A. Avižienis, J.-C. Laprie, B. Randell, and C. Landwehr, “Basic concepts and taxonomy of dependable and secure computing,” IEEE Trans. Dependable Secure Comput., vol. 1, pp. 11–33, 2004, doi: 10.1109/TDSC.2004.2.

S. Bondyopadhyay and M. Mohapatra, “Determination of suitable thermodynamic indices and prediction of thunderstorm events for Eastern India,” Meteorol. Atmos. Phys., vol. 135, art. 4, 2023, doi: 10.1007/s00703-022-00942-1.

L. Breiman, “Random forests,” Mach. Learn., vol. 45, pp. 5–32, 2001, doi: 10.1023/A:1010933404324.

C. Cortes and V. Vapnik, “Support-vector networks,” Mach. Learn., vol. 20, pp. 273–297, 1995, doi: 10.1007/BF00994018.

T. G. Dietterich, “Ensemble methods in machine learning,” in Multiple Classifier Systems, LNCS 1857. Berlin, Germany: Springer, 2000, pp. 1–15, doi: 10.1007/3-540-45014-9_1.

C. A. Doswell III, “Severe convective storms—An overview,” in Severe Convective Storms, Meteorol. Monogr. 28. Boston, MA, USA: American Meteorological Society, 2001, pp. 1–26, doi: 10.1175/0065-9401-28.50.1.

C. A. Doswell III and D. M. Schultz, “On the use of indices and parameters in forecasting severe storms,” Electron. J. Severe Storms Meteorol., vol. 1, pp. 1–22, 2006.

J. H. Friedman, “Greedy function approximation: A gradient boosting machine,” Ann. Stat., vol. 29, pp. 1189–1232, 2001, doi: 10.1214/aos/1013203451.

A. J. Haklander and A. Van Delden, “Thunderstorm predictors and their forecast skill for the Netherlands,” Atmos. Res., vol. 67–68, pp. 273–299, 2003, doi: 10.1016/S0169-8095(03)00056-5.

International Civil Aviation Organization, Annex 3: Meteorological Service for International Air Navigation. Montreal, QC, Canada: ICAO, 2022.

I. T. Jolliffe and D. B. Stephenson, Eds., Forecast Verification: A Practitioner’s Guide in Atmospheric Science, 2nd ed. Chichester, U.K.: Wiley-Blackwell, 2012.

S. Kapoor and A. Narayanan, “Leakage and the reproducibility crisis in machine-learning-based science,” Patterns, vol. 4, art. 100804, 2023, doi: 10.1016/j.patter.2023.100804.

S. M. Lundberg and S.-I. Lee, “A unified approach to interpreting model predictions,” in Adv. Neural Inf. Process. Syst., vol. 30, 2017, pp. 4765–4774.

P. Markowski and Y. Richardson, Mesoscale Meteorology in Midlatitudes. Chichester, U.K.: Wiley-Blackwell, 2010.

A. McGovern, R. Lagerquist, D. J. Gagne II, G. E. Jergensen, K. L. Elmore, C. R. Homeyer, and T. Smith, “Making the black box more transparent: Understanding the physical implications of machine learning,” Bull. Amer. Meteorol. Soc., vol. 100, pp. 2175–2199, 2019, doi: 10.1175/BAMS-D-18-0195.1.

H. P. Nayak and M. Mandal, “Analysis of stability parameters in relation to precipitation associated with pre-monsoon thunderstorms over Kolkata, India,” J. Earth Syst. Sci., vol. 123, pp. 689–703, 2014.

M. Reichstein, G. Camps-Valls, B. Stevens, M. Jung, J. Denzler, N. Carvalhais, and Prabhat, “Deep learning and process understanding for data-driven Earth system science,” Nature, vol. 566, pp. 195–204, 2019, doi: 10.1038/s41586-019-0912-1.

N. Sagita and T. Takemi, “Exploring the thunderstorm predictors in Indonesia,” SOLA, vol. 21, pp. 51–60, 2025, doi: 10.2151/sola.2025-007.

D. S. Wilks, Statistical Methods in the Atmospheric Sciences, 3rd ed. Oxford, U.K.: Academic Press, 2011.

World Meteorological Organization, Guide to Aeronautical Meteorological Practices, WMO-No. 732. Geneva, Switzerland: WMO, 2023.

Z.-H. Zhou, Ensemble Methods: Foundations and Algorithms. Boca Raton, FL, USA: Chapman & Hall/CRC, 2012.




DOI: http://dx.doi.org/10.52155/ijpsat.v58.2.8536

Refbacks

  • There are currently no refbacks.


Copyright (c) 2026 Dima Soraya

Creative Commons License
This work is licensed under a Creative Commons Attribution 4.0 International License.