Ensembles of label noise filters: a ranking approach

Ensembles of label noise filters: a ranking approach

Author Garcia, Luis P. F. Google Scholar
Lorena, Ana C. Autor UNIFESP Google Scholar
Matwin, Stan Google Scholar
de Carvalho, Andre C. P. L. F. Google Scholar
Abstract Label noise can be a major problem in classification tasks, since most machine learning algorithms rely on data labels in their inductive process. Thereupon, various techniques for label noise identification have been investigated in the literature. The bias of each technique defines how suitable it is for each dataset. Besides, while some techniques identify a large number of examples as noisy and have a high false positive rate, others are very restrictive and therefore not able to identify all noisy examples. This paper investigates how label noise detection can be improved by using an ensemble of noise filtering techniques. These filters, individual and ensembles, are experimentally compared. Another concern in this paper is the computational cost of ensembles, once, for a particular dataset, an individual technique can have the same predictive performance as an ensemble. In this case the individual technique should be preferred. To deal with this situation, this study also proposes the use of meta-learning to recommend, for a new dataset, the best filter. An extensive experimental evaluation of the use of individual filters, ensemble filters and meta-learning was performed using public datasets with imputed label noise. The results show that ensembles of noise filters can improve noise filtering performance and that a recommendation system based on meta-learning can successfully recommend the best filtering technique for new datasets. A case study using a real dataset from the ecological niche modeling domain is also presented and evaluated, with the results validated by an expert.
Keywords Label noise
Noise filters
Ensemble filters
Noise ranking
Recommendation system
Language English
Date 2016
Published in Data Mining And Knowledge Discovery. Dordrecht, v. 30, n. 5, p. 1192-1216, 2016.
ISSN 1384-5810 (Sherpa/Romeo, impact factor)
Publisher Springer
Extent 1192-1216
Origin http://dx.doi.org/10.1007/s10618-016-0475-9
Access rights Closed access
Type Article
Web of Science ID WOS:000382010500009
URI http://repositorio.unifesp.br/handle/11600/51113

Show full item record




File

File Size Format View

There are no files associated with this item.

This item appears in the following Collection(s)

Search


Browse

Statistics

My Account