A link-based cluster ensemble approach for categorical data clustering

Natthakan Iam-On*, Tossapon Boongeon, Simon Garrett, Chris Price

*Awdur cyfatebol y gwaith hwn

Allbwn ymchwil: Cyfraniad at gyfnodolynErthygladolygiad gan gymheiriaid

119 Dyfyniadau(SciVal)

Crynodeb

Although attempts have been made to solve the problem of clustering categorical data via cluster ensembles, with the results being competitive to conventional algorithms, it is observed that these techniques unfortunately generate a final data partition based on incomplete information. The underlying ensemble-information matrix presents only cluster-data point relations, with many entries being left unknown. The paper presents an analysis that suggests this problem degrades the quality of the clustering result, and it presents a new link-based approach, which improves the conventional matrix by discovering unknown entries through similarity between clusters in an ensemble. In particular, an efficient link-based algorithm is proposed for the underlying similarity assessment. Afterward, to obtain the final clustering result, a graph partitioning technique is applied to a weighted bipartite graph that is formulated from the refined matrix. Experimental results on multiple real data sets suggest that the proposed link-based method almost always outperforms both conventional clustering algorithms for categorical data and well-known cluster ensemble techniques.

Iaith wreiddiolSaesneg
Rhif yr erthygl5677529
Tudalennau (o-i)413-425
Nifer y tudalennau13
CyfnodolynIEEE Transactions on Knowledge and Data Engineering
Cyfrol24
Rhif cyhoeddi3
Dynodwyr Gwrthrych Digidol (DOIs)
StatwsCyhoeddwyd - Maw 2012

Ôl bys

Gweld gwybodaeth am bynciau ymchwil 'A link-based cluster ensemble approach for categorical data clustering'. Gyda’i gilydd, maen nhw’n ffurfio ôl bys unigryw.

Dyfynnu hyn