การเพิ่มประสิทธิภาพการคัดเลือกคุณลักษณะสำหรับการพยากรณ์โรคไม่ติดต่อเรื้อรังแบบหลายงานจากข้อมูลสุขภาพที่ไม่สมดุล
DOI:
https://doi.org/10.65205/jcct.2026.e3069คำสำคัญ:
การเรียนรู้แบบหลายงาน, ข้อมูลไม่สมดุล, SMOTE, XGBoost, LightGBMบทคัดย่อ
งานวิจัยนี้มีวัตถุประสงค์เพื่อศึกษาและเปรียบเทียบประสิทธิภาพของเทคนิคการคัดเลือกคุณลักษณะ ทั้งในรูปแบบวิธีการ Filter และ Wrapper สำหรับข้อมูลสุขภาพที่เป็นแบบหลายงาน และมีความไม่สมดุลของข้อมูล เพื่อพัฒนาแบบจำลองการพยากรณ์โรคไม่ติดต่อเรื้อรัง (NCDs) โดยการบูรณาการเทคนิคการเลือกคุณลักษณะที่มีประสิทธิภาพสูงสุด ร่วมกับแบบจำลองการเรียนรู้ของเครื่อง และเพื่อประเมินผลการคัดเลือกคุณลักษณะสำคัญของแต่ละเทคนิค โดยใช้ตัวชี้วัดทางสถิติที่เหมาะสมกับข้อมูลไม่สมดุล การดำเนินงานวิจัยประกอบด้วย 6 ขั้นตอนหลัก ได้แก่ 1) การเก็บรวบรวมข้อมูลจากเว็บไซต์ Kaggle ซึ่งประกอบด้วย ข้อมูลจำนวน 253,680 แถว และคุณลักษณะ 19 คุณลักษณะ 2) การเตรียมข้อมูลและการกำหนดเป้าหมาย 3 โรค 3) การคัดเลือกคุณลักษณะโดยใช้ 2 วิธี คือ วิธีแบบกรอง (Information Gain, Gain Ratio, Chi-Square) และวิธีแบบควบรวม (Forward Selection, Backward Elimination) 4) การจัดการความไม่สมดุลของข้อมูลด้วยเทคนิค SMOTE 5) การสร้างแบบจำลองการจำแนกโดยใช้เทคนิค XGBoost, Random Forest และ LightGBM และ 6) การประเมินและปรับค่าพารามิเตอร์ของแบบจำลองเพื่อหาค่าที่เหมาะสมที่สุด ผลงานวิจัยพบว่า เทคนิคการคัดเลือกคุณลักษณะด้วยเทคนิค Gain Ratio ช่วยเพิ่มประสิทธิภาพให้เทคนิค XGBoost โดยลดคุณลักษณะจาก 19 เหลือเพียง 7 คุณลักษณะ ได้แก่ ดัชนีมวลกาย อายุ รายได้ สุขภาพกายใน 30 วัน ประเมินสุขภาพตนเอง ระดับการศึกษา และสุขภาพจิตใน 30 วัน ซึ่งให้ค่าความถูกต้องสูงสุดอยู่ที่ 0.8768 และให้ค่าพื้นที่ใต้เส้นโค้ง (AUC) สูงสุดอยู่ที่ 0.9300 แสดงให้เห็นว่า วิธีนี้มีความเหมาะสมในการพยากรณ์โรคไม่ติดต่อเรื้อรังแบบหลายงานได้อย่างแม่นยำและรวดเร็ว
Downloads
เอกสารอ้างอิง
Jupriyadi, Budiman, A., Hamidi, E. A. Z., Ahdan, S., & Negara, R. M. (2024, July 4-5). Wrapper-Based Feature Selection to Improve the Accuracy of Intrusion Detection System (IDS). 2024 10th International Conference on Wireless and Telematics, 1-5. Institute of Electrical and Electronics Engineers. https://doi.org/10.1109/icwt62080.2024.10674687 DOI: https://doi.org/10.1109/ICWT62080.2024.10674687
Khan, R. A. (2023). Resilience Family of Receiver Operating Characteristic Curves. IEEE Transactions on Reliability, 72(2), 716-726. https://doi.org/10.1109/tr.2022.3194710 DOI: https://doi.org/10.1109/TR.2022.3194710
Leite, Â. (2025). Chronic Illnesses: Varied Health Patterns and Mental Health Challenges. Healthcare, 13(12), 1396. https://doi.org/10.3390/healthcare13121396 DOI: https://doi.org/10.3390/healthcare13121396
Li, W., Peng, Y., & Peng, K. (2024). Diabetes Prediction Model Based on GA-XGBoost and Stacking Ensemble Algorithm. PLOS One, 19(9), e0311222. https://doi.org/10.1371/journal.pone.0311222 DOI: https://doi.org/10.1371/journal.pone.0311222
Noroozi, Z., Orooji, A., & Erfannia, L. (2023). Analyzing the Impact of Feature Selection Methods on Machine Learning Algorithms for Heart Disease Prediction. Scientific Reports, 13, 22588. https://doi.org/10.1038/s41598-023-49962-w DOI: https://doi.org/10.1038/s41598-023-49962-w
Pongshaing, T., & Thongkam, J. (2023). Optimization of Models for Hypertension Treatment Prediction with Factor Selection. Journal of Science and Technology, Ubon Ratchathani University, 25(1), 13-20. (In Thai)
Rainio, O., Teuho, J., & Klén, R. (2024). Evaluation Metrics and Statistical Tests for Machine Learning. Scientific Reports, 14, 6086. https://doi.org/10.1038/s41598-024-56706-x DOI: https://doi.org/10.1038/s41598-024-56706-x
Romsaiyud, W. (2024). Fast Synthesis of the Minority Class Using Generative Adversarial Networks for Imbalanced Data Classification Problems. Journal of Science and Technology Mahasarakham University, 43(2), 108-121. (In Thai)
Rufo, D. D., Debelee, T. G., Ibenthal, A., & Negera, W. G. (2021). Diagnosis of Diabetes Mellitus Using Gradient Boosting Machine (LightGBM). Diagnostics, 11(9), 1714. https://doi.org/10.3390/diagnostics11091714 DOI: https://doi.org/10.3390/diagnostics11091714
Sai, M. J., Chettri, P., Panigrahi, R., Garg, A., Bhoi, A. K., & Barsocchi, P. (2023). An Ensemble of Light Gradient Boosting Machine and Adaptive Boosting for Prediction of Type-2 Diabetes. International Journal of Computational Intelligence Systems, 16(1), 14. https://doi.org/10.1007/s44196-023-00184-y DOI: https://doi.org/10.1007/s44196-023-00184-y
Salhi, A., Henslee, A. C., Ross, J., Jabour, J., & Dettwiller, I. (2023). Data Preprocessing Using AutoML: A Survey. 2023 Congress in Computer Science, Computer Engineering, & Applied Computing, 1619-1623. Institute of Electrical and Electronics Engineers. https://doi.org/10.1109/csce60160.2023.00265 DOI: https://doi.org/10.1109/CSCE60160.2023.00265
Sunggad, S., & Maneerat, P. (2023). Comparison of Feature Selection Methods to Improve Diabetes Predictions. Journal of Science and Technology Thonburi University, 7(2), 12-24. (In Thai)
Teboul, A. (n.d.). Diabetes Health Indicators Dataset. https://www.kaggle.com/datasets/alexteboul/diabetes-health-indicators-dataset
Tepdang, S. (2023). The Classification of Diabetic Patients Using Machine Learning Method by Feature Selection. RMUTSB Academic Journal, 11(1), 29-44. (In Thai)
World Health Organization. (2023). Global Report on Hypertension: The Race Against a Silent Killer. https://www.who.int/teams/noncommunicable-diseases/hypertension-report
ดาวน์โหลด
เผยแพร่แล้ว
รูปแบบการอ้างอิง
ฉบับ
ประเภทบทความ
สัญญาอนุญาต
ลิขสิทธิ์ (c) 2026 วารสารคอมพิวเตอร์และเทคโนโลยีสร้างสรรค์

อนุญาตภายใต้เงื่อนไข Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.





















