Evolutionary feature selection for big data classification: a MapReduce approach
Summary: Nowadays, many disciplines have to deal with big datasets that additionally involve a high number of features. Feature selection methods aim at eliminating noisy, redundant, or irrelevant features that may deteriorate the classification performance. However, traditional methods lack enough scalability to cope with datasets of millions of instances and extract successful results in a delimited time. This paper presents a feature selection algorithm based on evolutionary computation that uses the MapReduce paradigm to obtain subsets of features from big datasets. The algorithm decomposes the original dataset in blocks of instances to learn from them in the map phase; then, the reduce phase merges the obtained partial results into a final vector of feature weights, which allows a flexible application of the feature selection procedure using a threshold to determine the selected subset of features. The feature selection method is evaluated by using three well-known classifiers (SVM, Logistic Regression, and Naive Bayes) implemented within the Spark framework to address big data problems. In the experiments, datasets up to 67 millions of instances and up to 2000 attributes have been managed, showing that this is a suitable framework to perform evolutionary feature selection, improving both the classification accuracy and its runtime when dealing with big data problems.
- Feature selection with partition differentiation entropy for large-scale data sets
- Hadoop neural network for parallel and distributed feature selection
- Feature selection: from the past to the future
- An integrated approach to speed up GA-SVM feature selection model
- Initialization of Feature Selection Search for Classification
- 10.1162/153244303322753616
- Advances in instance selection for instance-based learning algorithms
- Computational Methods of Feature Selection
- Data and task parallelism in ILP using mapreduce
- scientific article; zbMATH DE number 6378115 (Why is no real title available?)
- scientific article; zbMATH DE number 3436645 (Why is no real title available?)
- Introduction to machine learning.
- Nearest neighbor pattern classification
- Nonlinear Dimensionality Reduction
- Wrappers for feature subset selection
- Data learning from big data
- A semi-parallel framework for greedy information-theoretic feature selection
- Hadoop neural network for parallel and distributed feature selection
- Fuzzy rule based classification systems for big data with MapReduce: granularity analysis
- Massively parallel feature selection: an approach based on variance preservation
- An integrated approach to speed up GA-SVM feature selection model
- A detailed study of the distributed rough set based locality sensitive hashing feature selection technique
- Feature selection: from the past to the future
- A greedy feature selection algorithm for big data of high dimensionality
This page was built for publication: Evolutionary feature selection for big data classification: a MapReduce approach
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q1665073)