An Empirical Evaluation of k-Means Coresets

Schwiegelshohn, Chris; Sheikh-Omar, Omar Ali

Repository landing page

oai:drops-oai.dagstuhl.de:17022

An Empirical Evaluation of k-Means Coresets

Authors: Chris Schwiegelshohn
Omar Ali Sheikh-Omar
Publication date: 1 January 2022
Publisher: LIPIcs - Leibniz International Proceedings in Informatics. 30th Annual European Symposium on Algorithms (ESA 2022)
Doi

Abstract

Coresets are among the most popular paradigms for summarizing data. In particular, there exist many high performance coresets for clustering problems such as k-means in both theory and practice. Curiously, there exists no work on comparing the quality of available k-means coresets. In this paper we perform such an evaluation. There currently is no algorithm known to measure the distortion of a candidate coreset. We provide some evidence as to why this might be computationally difficult. To complement this, we propose a benchmark for which we argue that computing coresets is challenging and which also allows us an easy (heuristic) evaluation of coresets. Using this benchmark and real-world data sets, we conduct an exhaustive evaluation of the most commonly used coreset algorithms from theory and practice

Similar works

Full text

Open in the Core reader

Download PDF

Dagstuhl Research Online Publication Server

oai:drops-oai.dagstuhl.de:1702...

Last time updated on 07/10/2022

This paper was published in Dagstuhl Research Online Publication Server.

Having an issue?

Is data on this page outdated, violates copyrights or anything else? Report the problem now and we will take corresponding actions after reviewing your request.