IRIS Institutional Research Information System - AIR Archivio Istituzionale della Ricerca

Common clustering algorithms require multiple scans of all the data to achieve conver-gence, and this is prohibitive when large databases, with data arriving in streams, must be processed. Some algorithms to extend the popular K-means method to the analysis of streaming data are present in literature since 1998 (Bradley et al. in Scaling clustering algorithms to large databases. In: KDD. p. 9–15, 1998; O’Callaghan et al. in Streaming-data algorithms for high-quality clustering. In: Proceedings of IEEE international confer-ence on data engineering. p. 685, 2001), based on the memorization and recursive update of a small number of summary statistics, but they either don’t take into account the specific variability of the clusters, or assume that the random vectors which are processed and grouped have uncorrelated components. Unfortunately this is not the case in many practical situations. We here propose a new algorithm to process data streams, with data having correlated components and coming from clusters with different covariance matrices. Such covariance matrices are estimated via an optimal double shrinkage method, which provides positive definite estimates even in presence of a few data points, or of data having components with small variance. This is needed to invert the matrices and compute the Mahalanobis distances that we use for the data assignment to the clusters. We also estimate the total number of clusters from the data.

A clustering algorithm for multivariate data streams with correlated components / G. Aletti, A. Micheletti. - In: JOURNAL OF BIG DATA. - ISSN 2196-1115. - 4:1(2017), pp. 48.1-48.20. [10.1186/s40537-017-0109-0]

A clustering algorithm for multivariate data streams with correlated components

G. Aletti^Primo;A. Micheletti^Ultimo

2017

Abstract

Common clustering algorithms require multiple scans of all the data to achieve conver-gence, and this is prohibitive when large databases, with data arriving in streams, must be processed. Some algorithms to extend the popular K-means method to the analysis of streaming data are present in literature since 1998 (Bradley et al. in Scaling clustering algorithms to large databases. In: KDD. p. 9–15, 1998; O’Callaghan et al. in Streaming-data algorithms for high-quality clustering. In: Proceedings of IEEE international confer-ence on data engineering. p. 685, 2001), based on the memorization and recursive update of a small number of summary statistics, but they either don’t take into account the specific variability of the clusters, or assume that the random vectors which are processed and grouped have uncorrelated components. Unfortunately this is not the case in many practical situations. We here propose a new algorithm to process data streams, with data having correlated components and coming from clusters with different covariance matrices. Such covariance matrices are estimated via an optimal double shrinkage method, which provides positive definite estimates even in presence of a few data points, or of data having components with small variance. This is needed to invert the matrices and compute the Mahalanobis distances that we use for the data assignment to the clusters. We also estimate the total number of clusters from the data.

Scheda breve

Scheda completa

Scheda completa (DC)

	Parole chiave
	
				Big data; Data streams; Clustering; Mahalanobis distance
			
	Settori scientifico-disciplinari dell'articolo (sola visualizzazione)
	
				Settore MAT/06 - Probabilita' e Statistica Matematica
Settore SECS-S/01 - Statistica
			
	Data di pubblicazione
	
				2017
			
	Rivista in ANCE
	
				JOURNAL OF BIG DATA
			
	DOI
	
				https://dx.doi.org/10.1186/s40537-017-0109-0
			
	Collegato a
	
				http://hdl.handle.net/2434/514034
			
	Tipologia
	
				Article (author)
			
	Appare nelle tipologie:
	
				01 - Articolo su periodico

File in questo prodotto:

File	Dimensione	Formato
Aletti_et_al-2017-Journal_of_Big_Data.pdf accesso aperto Tipologia: Publisher's version/PDF Dimensione 3.43 MB Formato Adobe PDF Visualizza/Apri	3.43 MB	Adobe PDF	Visualizza/Apri

Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2434/541307

Citazioni

ND

16

ND

ND

social impact