Gaussian Motion Data

Synthetic database with concept drift to test evolving clustering algorithms

In the Evolving Clustering literature there is currently a lack of publicly available data sets (synthetic or real) for evaluating algorithms in concept drift situations. The synthetic database presented here can be used to test the ability of the algorithms to follow data drift.

The fact that the data is simulated provides detailed knowledge about the structure underlying the data, enabling a more thorough evaluation of the results. This is particularly important in online clustering for evaluating not only the final result provided by the algorithm, but also the intermediate results.

By having detailed knowledge about the models that generated the data it is possible to accurate assess the performance of the algorithms through all the intermediate states.

The database is composed by several data sets with concept drift which also contains information about the temporal evolution of the models that generated the data.

The data sets have been generated by Gaussian distributions whose mean and/or covariance change over time. In the database both the simulated data and the Gaussians that generated it are provided, hence enabling an accurate evaluation of the partitions through time.

Recommended citation

@article{MARQUEZ201816,
title = "A novel and simple strategy for evolving prototype based clustering",
journal = "Pattern Recognition",
volume = "82",
pages = "16 - 30",
year = "2018",
issn = "0031-3203",
doi = "https://doi.org/10.1016/j.patcog.2018.04.020",
url = "http://www.sciencedirect.com/science/article/pii/S0031320318301547",
author = "David G. Márquez and Abraham Otero and Paulo Félix and Constantino A. García",
keywords = "Evolving clustering, Data stream, Concept drift, Gaussian mixture models, K-means, Cluster evolution"
}