بهبود دقت پیش‌بینی بارش ماهانه با استفاده از چارچوب پیش‌پردازش داده و مدل های یادگیری ماشین

نوع مقاله : مقاله پژوهشی

نویسندگان

1 دانشجوی کارشناسی ارشد گرایش مدیریت منابع اب دانشگاه تهران

2 دانشجوی کارشناسی ارشد مهندسی منابع آب دانشگاه تهران دانشکده کشاورزی گروه عمران و ابادانی

3 دانشجوی کارشناسی ارشد زهکشی دانشگاه تهران دانشکده کشاورزی گروه عمران و ابادانی

چکیده

This study examined the effect of data preprocessing on the performance of machine learning models in forecasting monthly precipitation in the Lake Namak basin. CHIRPS data were first validated against ground based observations (NSE = 70%), then cleaned using the Z Score method to remove outliers and subsequently reconstructed using Singular Spectrum Analysis (SSA) and linear interpolation. Statistical analyses indicated that preprocessing substantially reduced the skewness and kurtosis of the time series and enhanced data stability without noticeably altering the variance. Performance comparisons among the Random Forest, XGBoost, SVR, and CatBoost models showed that preprocessing significantly improved forecasting accuracy; notably, the Random Forest model achieved the largest enhancement, with MSE decreasing from 28.7 in the original series to approximately 15.7–16.0 in the corrected series, while SVR exhibited the least sensitivity to data adjustments. Both visual and numerical evaluations further confirmed the superior performance of Random Forest and XGBoost in terms of higher correlation and lower bias. Overall, the findings demonstrate that integrating effective preprocessing techniques with machine learning models—particularly Random Forest—plays a critical role in improving precipitation forecasting accuracy and reducing hydrological uncertainties.

کلیدواژه‌ها


عنوان مقاله [English]

Improvement of Monthly Precipitation Forecasting Accuracy Using a Data Preprocessing Framework and Machine Learning

نویسندگان [English]

  • ali dalir 1
  • ali milandarzadeh 2
  • Alireza Sheykhan 3
1 M.Sc., Department of Irrigation & Reclamation Engineering, Faculty of Agricultural Engineering & Technology, College of Agriculture & Natural Resources, University of Tehran,
2 M.Sc., Department of Irrigation & Reclamation Engineering, Faculty of Agricultural Engineering & Technology, College of Agriculture & Natural Resources, University of Tehran, Karaj, Tehran, Iran.
3 M.Sc., Department of Irrigation & Reclamation Engineering, Faculty of Agricultural Engineering & Technology, College of Agriculture & Natural Resources, University of Tehran, Karaj, Tehran, Iran.
چکیده [English]

در این پژوهش، اثر پیش‌پردازش داده‌ها بر عملکرد مدل‌های یادگیری ماشین در پیش‌بینی بارش ماهانه حوضه دریاچه ارومیه بررسی شد. داده‌های CHIRPS پس از اعتبارسنجی با ایستگاه‌های زمینی 70%= NSEبا استفاده از روش Z-Score از نقاط پرت پاک‌سازی و سپس به کمک روش‌های SSA و درون‌یابی خطی بازسازی شدند. نتایج آماری نشان داد که پیش‌پردازش موجب کاهش قابل‌توجه چولگی و کشیدگی سری‌های زمانی و افزایش پایداری داده‌ها بدون تغییر محسوس در واریانس شد. مقایسه عملکرد مدل‌های Random Forest، XGBoost، SVR و CatBoost نشان داد که پیش‌پردازش داده‌ها به‌طور محسوسی دقت پیش‌بینی را بهبود می‌دهد؛ به‌طوری‌که مدل Random Forest بیشترین بهبود را تجربه کرد و مقدار MSE آن از 7/28 در سری اولیه به حدود 7/15–0/16 در سری‌های اصلاح‌شده کاهش یافت، در حالی که SVR کمترین حساسیت را به اصلاح داده‌ها نشان داد. تحلیل‌های بصری و آماری نیز برتری Random Forest و XGBoost را از نظر همبستگی بالاتر و سوگیری کمتر تأیید کردند. نتایج این مطالعه نشان می‌دهد که ترکیب پیش‌پردازش مؤثر داده‌ها با مدل‌های یادگیری ماشین، به‌ویژه Random Forest، نقش کلیدی در افزایش دقت پیش‌بینی بارش و کاهش عدم‌قطعیت‌های هیدرولوژیکی دارد

کلیدواژه‌ها [English]

  • Data Preprocessing
  • Machine Learning Models
  • Monthly Precipitation Forecasting
  • Lake Namak Basin