[1]李金,王彪.结合模糊聚类和集成学习的不平衡数据过采样方法[J].计算机技术与发展,2025,(10):18-27.[doi:10.20165/j.cnki.ISSN1673-629X.2025.0127]
 LI Jin,WANG Biao.Oversampling Method for Imbalanced Data Based on Fuzzy Clustering and Ensemble Learning[J].,2025,(10):18-27.[doi:10.20165/j.cnki.ISSN1673-629X.2025.0127]
点击复制

结合模糊聚类和集成学习的不平衡数据过采样方法()

《计算机技术与发展》[ISSN:1006-6977/CN:61-1281/TN]

卷:
期数:
2025年10期
页码:
18-27
栏目:
大数据与云计算
出版日期:
2025-10-10

文章信息/Info

Title:
Oversampling Method for Imbalanced Data Based on Fuzzy Clustering and Ensemble Learning
文章编号:
1673-629X(2025)10-0018-10
作者:
李金王彪
西安科技大学理学院,陕西西安 710600
Author(s):
LI JinWANG Biao
School of Science,Xi’an University of Science and Technology,Xi’an 710600,China
关键词:
不平衡数据分类类别重叠过采样软聚类集成学习
Keywords:
imbalanced data classificationclass overlapoversamplingsoft clusteringensemble learning
分类号:
TP181
DOI:
10.20165/j.cnki.ISSN1673-629X.2025.0127
摘要:
目前,不平衡数据的处理方法主要致力于解决类分布不平衡问题,通常采用重采样方法来构建更为平衡的数据集。然而,与类分布不平衡相比,类间重叠问题对不平衡数据分类性能产生的不利影响更大。因此,针对不平衡数据中存在的类内不平衡以及类间重叠问题,提出了一种基于模糊聚类和集成学习的不平衡数据过采样方法FCEL。在数据层面,首先运用SMOTE过采样合成新样本;其次利用软聚类和自适应阈值对数据空间进行区域划分;随后对划分的区域进行重采样,生成两个采样子集。在算法层面,首先根据不同的采样子集构建相应的集成模型;其次通过模型选择算法,根据每个样本的分布选择合适的模型。在9个不平衡数据集上进行对比实验,实验结果表明:与现有一些典型方法相比,FCEL方法的Recall、F?、G-mean和AUC这四项指标的平均值至少提升17.67百分点、0.09百分点、7.25百分点和1.21百分点;最多提升30.29百分点、4.62百分点、17.25百分点和4.35百分点,说明该方法能有效地提高少数类样本的分类精度。
Abstract:
At present,the processing methods of imbalanced data mainly focus on solving the problem of class distribution imbalance and usually adopt resampling methods to construct a more balanced dataset. However,compared with class distribution imbalance,the problem of inter-class overlap has a greater adverse impact on the classification performance of imbalanced data. Therefore,addressing at the issues of intra-class imbalance and inter-class overlap in imbalanced datasets,an imbalanced data oversampling method FCEL based on fuzzy clustering and ensemble learning is proposed. At the data level,firstly,SMOTE oversampling is used to synthesize new samples.Then,soft clustering and adaptive threshold are employed to partition the data space into regions. Subsequently,the partitioned regions are resampled to generate two sampling subsets. At the algorithm level,firstly,corresponding ensemble models are constructed based on the different sampling subsets. Furthermore,a model selection algorithm is applied to assign suitable models to each sample according to its distribution. Comparative experiments conducted on 9 imbalanced datasets. The experimental results show that compared with some existing typical methods,the average values of the four indicators Recall,F?,G-mean,and AUC of the FCEL method are increased by at least 17.67 percentage points,0.09 percentage points,7.25 percentage points,and 1.21 percentage points,and at most 30.29 percentage points,4.62 percentage points,17.25 percentage points,and 4.35 percentage points,indicating that the proposed method can effectively improve the classification accuracy of minority class samples.
更新日期/Last Update: 2025-10-10