Efficient computation of multiple density-based clustering hierarchies

Araujo Neto, Antonio Cavalcante; Sander, Jörg; Campello, Ricardo J.G.B.; Nascimento, Mario A.

Repository landing page

oai:researchonline.jcu.edu.au:51876

Efficient computation of multiple density-based clustering hierarchies

Authors: Antonio Cavalcante Araujo Neto
Jörg Sander
Ricardo J.G.B. Campello
Mario A. Nascimento
Publication date: 1 December 2017
Publisher: Institute of Electrical and Electronics Engineers
Doi

Abstract

HDBSCAN*, a state-of-the-art density-based hierarchical clustering method, produces a hierarchical organization of clusters in a dataset w.r.t. a parameter mpts. While the performance of HDBSCAN* is robust w.r.t. mpts, choosing a "good" value for it can be challenging: depending on the data distribution, a high or low value for mpts may be more appropriate, and certain data clusters may reveal themselves at different values of mpts. To explore results for a range of mpts, one has to run HDBSCAN* for each value in the range independently, which is computationally inefficient. In this paper we propose an efficient approach to compute all HDBSCAN* hierarchies for a range of mpts by replacing the graph used by HDBSCAN* with a much smaller graph that is guaranteed to contain the required information. Our experiments show that our approach can obtain, for example, over one hundred hierarchies for a cost equivalent to running HDBSCAN* about 2 times. In fact, this speedup tends to increase with the number of hierarchies to be computed

Similar works

Full text

ResearchOnline at James Cook University

oai:researchonline.jcu.edu.au:...

Last time updated on 18/04/2020

This paper was published in ResearchOnline at James Cook University.

Having an issue?

Is data on this page outdated, violates copyrights or anything else? Report the problem now and we will take corresponding actions after reviewing your request.