Mining@home: Towards a Public ResourceComputing Framework for Distributed Data Mining

Several classes of scientific and commercial applications require the execution of a large number of independent tasks. One highly successful and low-cost mechanism for acquiring the necessary computing power for these applications is the 'public-resource computing', or 'desktop Grid' paradigm, which exploits the computational power of private computers. So far, this paradigm has not been applied to data mining applications for two main reasons. First, it is not straightforward to decompose a data mining algorithm into truly independent sub-tasks. Second, the large volume of the involved data makes it difficult to handle the communication costs of a parallel paradigm. This paper introduces a general framework for distributed data mining applications called Mining@home In particular, we focus on one of the main data mining problems: the extraction of closed frequent itemsets from transactional databases. We show that it is possible to decompose this problem into independent tasks, which however need to share a large volume of the data. We thus introduce a data-intensive computing network, which adopts a P2P topology based on super peers with caching capabilities, aiming to support the dissemination of large amounts of information. Finally, we evaluate the execution of a pattern extraction task on such network.

Mining@home: Towards a Public ResourceComputing Framework for Distributed Data Mining

LUCCHESE, Claudio;C. MASTROIANNI;ORLANDO, Salvatore;D. TALIA

2010

Abstract

Several classes of scientific and commercial applications require the execution of a large number of independent tasks. One highly successful and low-cost mechanism for acquiring the necessary computing power for these applications is the 'public-resource computing', or 'desktop Grid' paradigm, which exploits the computational power of private computers. So far, this paradigm has not been applied to data mining applications for two main reasons. First, it is not straightforward to decompose a data mining algorithm into truly independent sub-tasks. Second, the large volume of the involved data makes it difficult to handle the communication costs of a parallel paradigm. This paper introduces a general framework for distributed data mining applications called Mining@home In particular, we focus on one of the main data mining problems: the extraction of closed frequent itemsets from transactional databases. We show that it is possible to decompose this problem into independent tasks, which however need to share a large volume of the data. We thus introduce a data-intensive computing network, which adopts a P2P topology based on super peers with caching capabilities, aiming to support the dissemination of large amounts of information. Finally, we evaluate the execution of a pattern extraction task on such network.

Scheda breve

Scheda completa

Scheda completa (DC)

	Anno pubblicazione
	
				2010
			
	Titolo della Rivista
	
				CONCURRENCY AND COMPUTATION
			
	N° Volume
	
				22
			
	DOI
	
				https://dx.doi.org/10.1002/cpe.1545
			
	Appare nelle tipologie:
	
				2.1 Articolo su rivista

File in questo prodotto:

File	Dimensione	Formato
minig_home.pdf non disponibili Tipologia: Documento in Post-print Licenza: Accesso chiuso-personale Dimensione 992.91 kB Formato Adobe PDF Visualizza/Apri	992.91 kB	Adobe PDF	Visualizza/Apri

I documenti in ARCA sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/10278/21103

Citazioni

ND

13

8

social impact