Geometric Enhancements to SMOTE: Centroid, Radial, and Distance-Aware Strategies for Minority Oversampling
Abstract
Data imbalance frequently leads classification algorithms to neglect minority classes, even though these classes tend to be more important in real-world scenarios. A wide range of resampling techniques, particularly oversampling and undersampling methods, have been developed to mitigate this issue. Among these, SMOTE and its variants have gained popularity because they are able to generate synthetic minority instances. However, traditional SMOTE approaches often fall short in capturing local data geometry and adapting to the distribution of minority samples. In this study, we propose a unified framework comprising three novel geometry-aware oversampling methods: (1) Centroid Interpolation, which synthesizes data between the centroid of nearest neighbors and neighbor vectors; (2) Angular Sampling, which leverages the average radius and random angle generation to create diverse samples around the minority instance; and (3) Adaptive Step Scaling, which adjusts interpolation ratios based on the relative distances to neighbors. To validate the effectiveness of the proposed methods, extensive experiments were conducted on 22 imbalanced benchmark datasets using Decision Trees (DT) and Random Forest (RF) as classifiers, comparing our techniques against widely-used oversampling strategies including SMOTE, Borderline SMOTE, K-Means SMOTE, SVM SMOTE, and ADASYN. Performance was evaluated using F-Score, G-Mean, Accuracy, and Youden’s Index metrics. The results demonstrate that our proposed methods consistently outperform existing SMOTE variants, significantly improving classification performance across diverse datasets.
Keywords
Oversampling, cross validation, machine learning, data mining.