This Title All WIREs
How to cite this WIREs title:
WIREs Data Mining Knowl Discov
Impact Factor: 7.250

Clustering high dimensional data

Full article on Wiley Online Library:   HTML PDF

Can't access this content? Tell your librarian.

Abstract High‐dimensional data, i.e., data described by a large number of attributes, pose specific challenges to clustering. The so‐called ‘curse of dimensionality’, coined originally to describe the general increase in complexity of various computational problems as dimensionality increases, is known to render traditional clustering algorithms ineffective. The curse of dimensionality, among other effects, means that with increasing number of dimensions, a loss of meaningful differentiation between similar and dissimilar objects is observed. As high‐dimensional objects appear almost alike, new approaches for clustering are required. Consequently, recent research has focused on developing techniques and clustering algorithms specifically for high‐dimensional data. Still, open research issues remain. Clustering is a data mining task devoted to the automatic grouping of data based on mutual similarity. Each cluster groups objects that are similar to one another, whereas dissimilar objects are assigned to different clusters, possibly separating out noise. In this manner, clusters describe the data structure in an unsupervised manner, i.e., without the need for class labels. A number of clustering paradigms exist that provide different cluster models and different algorithmic approaches for cluster detection. Common to all approaches is the fact that they require some underlying assessment of similarity between data objects. In this article, we provide an overview of the effects of high‐dimensional spaces, and their implications for different clustering paradigms. We review models and algorithms that address clustering in high dimensions, with pointers to the literature, and sketch open research issues. We conclude with a summary of the state of the art. © 2012 Wiley Periodicals, Inc. This article is categorized under: Technologies > Structure Discovery and Clustering

Clustering: finding groups of data objects based on mutual similarity; dissimilar objects may be separated as noise.

[ Normal View | Magnified View ]

Document clustering: term frequency per document and inverse term frequency across documents form high‐dimensional vectors.

[ Normal View | Magnified View ]

High‐dimensional time series: dimensionality reduction and discretization.

[ Normal View | Magnified View ]

Grids (left) and hyperplanes (right) for clustering in high‐dimensional spaces.

[ Normal View | Magnified View ]

Clusters in different subspace projections.

[ Normal View | Magnified View ]

Data clustered in each axis, but spread out in their combination.

[ Normal View | Magnified View ]

High dimensions: change in density.

[ Normal View | Magnified View ]

Dendrogram: visualizing hierarchies of clusters.

[ Normal View | Magnified View ]

Browse by Topic

Technologies > Structure Discovery and Clustering

Access to this WIREs title is by subscription only.

Recommend to Your
Librarian Now!

The latest WIREs articles in your inbox

Sign Up for Article Alerts