Non-negative Matrix Factorization in the Field of Clustering : comparison of Methods
Files
Wilberz_25451100_2017.pdf
Closed access - Adobe PDF
- 1.58 MB
Erratum.pdf
Closed access - Adobe PDF
- 578.12 KB
Details
- Supervisors
- Faculty
- Degree label
- Abstract
- In this thesis, we attempted to compare the performance of 7 kernels applied to the kernel k-means as well as the performance of four spectral clustering algorithms and the Louvain method to cluster the nodes of a graph. After explaining the theory and the methodology surrounding our study, we assessed the outputs of these clustering algorithms on each of the 16 datasets tested based on three quality measures: the correct classification rate, the adjusted rand index, and the normalized mutual information. Then, multiples analyses such as the Nemenyi test based on a Friedman test have been performed on the collected data to better visualize the performance of the twelve algorithms studied. The results obtained have then enabled us to: 1. Support the assertion of Ivashkin and Chebotarev (Ivashkin & Chebotarev, 2016) that the logarithmic version of the exponential diffusion kernels perform better than the original kernels. 2. Propose to use the “highest value approach” rather than the k-means algorithm to construct a discrete partition from the transformed indicator matrix obtained after performing the non-negative matrix factorization spectral clustering algorithm. 3. Give an overview of the performance of seven kernels tested on the Kernel k-means and 5 other algorithms. 4. Support the conclusions drawn by Felix Sommer, Francois Fouss and Marco Saerens (Sommer, Fouss, & Saerens, 2016) highlighting the high performance obtained by the free energy and the randomized shortest paths kernels. We conclude the thesis with a brief discussion on these findings followed by a short analysis of the limitations of this work and provides recommendations to mitigate these in the future. A chapter has also been dedicated to explain various clustering applications in the managerial world to help the reader to better understand the potential that represent this statistical tool in organizations such as pharmaceutical companies or banks.