Statistics
Cluster analysis is a common part of data analysis. Its aim is the identification of unknown structure in data and the determination of a partition with groups of objects as similar as possible (so-called clusters). In contrast to the frequent occurrence of mixed-type data in real-world applications, involving numerical as well as categorical features, research tends to concentrate on data containing exclusively numerical features. There are comparatively few methods for clustering mixed-type data, with the k-prototypes algorithm being presumably the most widely recognized. The purpose of this cumulative dissertation is to expand the scope of this clustering algorithm. It addresses aspects that are not treated in Huang's original publication of the k-prototype algorithm, including the validation of the number of clusters, variable selection of data to be clustered, imputation of incomplete data, algorithm initialization, and the integration of an alternative distance measure in the algorithm routine. These issues are covered as they are prevalent in the application of the k-prototypes algorithm on real-world data. In these clustering tasks, the user lacks knowledge about the optimal number of clusters or the most useful variables to determine the cluster partition. In addition, incomplete data often occur and need to be dealt with. The algorithm’s initialization is analyzed to optimize the iterative routine, which was originally published with a random-based choice of initial prototypes. Additionally, the distance-based partitioning algorithm is extended to ordinal data for distance calculation with the change of the algorithm’s distance measure. To conduct the research, simulation studies on artificially generated data are utilized as well as exemplary analyzes on real-world data.
Cluster analysis represents one of the most versatile methods in statistical science. It is employed in empirical sciences for the summarization of datasets into groups of similar objects, with the purpose of facilitating the interpretation and further analysis of the data. Cluster analysis is of particular importance in the exploratory investigation of data of high complexity, such as that derived from molecular biology or image databases. Consequently, recent work in the field of cluster analysis has focused on designing algorithms that can provide meaningful solutions for data with high cardinality and/or dimensionality, under the natural restriction of limited resources. The present thesis aims to develop improved methods for the clustering of high-dimensional datasets, as well as further applications of such algorithms in practical settings.
In the first part of the thesis, a more detailed review of the representative clustering algorithms focused on the analysis of very large or high-dimensional datasets is provided. Subsequently, a newly developed method for this purpose is described and evaluated. The developed algorithm is based on the principles of projection pursuit and grid partitioning, and focuses on reducing computational requirements for large datasets without loss of performance. In the second part of the thesis, a novel method for generating synthetic datasets with variable structure and clustering difficulty, that is aimed at evaluating clustering algorithms is presented. In the third part of the thesis, the applications of cluster analysis to the field of automatic image classification are investigated. A novel system for the semi-supervised annotation of images is described and evaluated. The system is based on a vocabulary of clusters of visual features extracted from images with known classification.