Robust clustering in regression analysis via the contaminated Gaussian cluster-weighted model
From MaRDI portal
Publication:2403302
Abstract: The Gaussian cluster-weighted model (CWM) is a mixture of regression models with random covariates that allows for flexible clustering of a random vector composed of response variables and covariates. In each mixture component, it adopts a Gaussian distribution for both the covariates and the responses given the covariates. To robustify the approach with respect to possible elliptical heavy tailed departures from normality, due to the presence of atypical observations, the contaminated Gaussian CWM is here introduced. In addition to the parameters of the Gaussian CWM, each mixture component of our contaminated CWM has a parameter controlling the proportion of outliers, one controlling the proportion of leverage points, one specifying the degree of contamination with respect to the response variables, and one specifying the degree of contamination with respect to the covariates. Crucially, these parameters do not have to be specified a priori, adding flexibility to our approach. Furthermore, once the model is estimated and the observations are assigned to the groups, a finer intra-group classification in typical points, outliers, good leverage points, and bad leverage points - concepts of primary importance in robust regression analysis - can be directly obtained. Relations with other mixture-based contaminated models are analyzed, identifiability conditions are provided, an expectation-conditional maximization algorithm is outlined for parameter estimation, and various implementation and operational issues are discussed. Properties of the estimators of the regression coefficients are evaluated through Monte Carlo experiments and compared to the estimators from the Gaussian CWM. A sensitivity study is also conducted based on a real data set.
Recommendations
- Seemingly unrelated clusterwise linear regression for contaminated data
- Local statistical modeling via a cluster-weighted approach with elliptical distributions
- Mixtures of multivariate contaminated normal regression models
- The generalized linear mixed cluster-weighted model
- Clustering bivariate mixed-type data via the cluster-weighted model
Cites work
- A likelihood-based constrained algorithm for multivariate normal mixture models
- A maximum likelihood methodology for clusterwise linear regression
- A multivariate generalization of the power exponential family of distributions
- A multivariate linear regression analysis using finite mixtures of \(t\) distributions
- Bayesian mixture labeling by highest posterior density
- Breakdown points for maximum likelihood estimators of location-scale mixtures
- Choosing initial values for the EM algorithm for finite mixtures
- Choosing starting values for the EM algorithm for getting the highest likelihood in multivariate Gaussian mixture models
- Cluster-weighted t-factor analyzers for robust model-based clustering and dimension reduction
- Clustering and classification via cluster-weighted factor analyzers
- Clustering bivariate mixed-type data via the cluster-weighted model
- Computational and Inferential Difficulties with Mixture Posterior Distributions
- Concomitant variables in finite mixture models
- Constrained monotone EM algorithms for finite mixture of multivariate Gaussians
- Dealing With Label Switching in Mixture Models
- Decision boundaries for mixtures of regressions
- Estimating the dimension of a model
- Finite mixture and Markov switching models.
- Finite mixture models
- Finite mixtures of unimodal beta and gamma densities and the k-bumps algorithm
- Fixed point clusters for linear regression: Computation and comparison
- Flexible mixture modelling with the polynomial Gaussian cluster-weighted model
- scientific article; zbMATH DE number 3673370 (Why is no real title available?)
- scientific article; zbMATH DE number 3567782 (Why is no real title available?)
- scientific article; zbMATH DE number 1059776 (Why is no real title available?)
- scientific article; zbMATH DE number 194744 (Why is no real title available?)
- scientific article; zbMATH DE number 194758 (Why is no real title available?)
- scientific article; zbMATH DE number 951459 (Why is no real title available?)
- scientific article; zbMATH DE number 3320125 (Why is no real title available?)
- Hypothesis Testing for Mixture Model Selection
- Identifiability of models for clusterwise linear regression
- Linear grouping using orthogonal regression
- Local statistical modeling via a cluster-weighted approach with elliptical distributions
- Maximum likelihood estimation via the ECM algorithm: A general framework
- Mixture Models, Outliers, and the EM Algorithm
- Mixtures of multivariate contaminated normal regression models
- Model based labeling for mixture models
- Model-based clustering
- Model-based clustering via linear cluster-weighted models
- Model-Based Gaussian and Non-Gaussian Clustering
- Model-based time-varying clustering of multivariate longitudinal data with covariates and outliers
- Multilevel cluster-weighted models for the evaluation of hospitals
- Multivariate response and parsimony for Gaussian cluster-weighted models
- Parsimonious mixtures of multivariate contaminated normal distributions
- Robust cluster analysis and variable selection
- Robust clusterwise linear regression through trimming
- Robust Estimation of the Mean and Covariance Matrix from Data with Missing Values
- Robust fitting of mixture regression models
- Robust fitting of mixtures using the trimmed likelihood estimator
- Robust linear clustering
- Robust mixture regression model fitting by Laplace distribution
- Robust mixture regression using the \(t\)-distribution
- Root selection in normal mixture models
- Serial and parallel implementations of model-based clustering via parsimonious Gaussian mixture models
- Statistical analysis of finite mixture distributions
- Statistical theory in clustering
- The distribution of the likelihood ratio for mixtures of densities from the one-parameter exponential family
- The generalized linear mixed cluster-weighted model
- The Identification of Multiple Outliers
- The multivariate leptokurtic-normal distribution and its application in model-based clustering
- Trimmed \(k\)-means: An attempt to robustify quantizers
Cited in
(54)- The robust EM-type algorithms for log-concave mixtures of regression models
- Asymmetric clusters and outliers: mixtures of multivariate contaminated shifted asymmetric Laplace distributions
- Least squares moment identification of binary regression mixture models
- Weighted likelihood latent class linear regression
- Mixtures of factor analyzers with covariates for modeling multiply censored dependent variables
- Multivariate hidden Markov regression models: random covariates and heavy-tailed distributions
- Matrix normal cluster-weighted models
- Estimating the covariance matrix of the maximum likelihood estimator under linear cluster-weighted models
- Multivariate cluster-weighted models based on seemingly unrelated linear regression
- Model-based clustering via new parsimonious mixtures of heavy-tailed distributions
- Robust fitting of mixture models using weighted complete estimating equations
- Extending finite mixtures of \(t\) linear mixed-effects models with concomitant covariates
- Modeling frequency and severity of claims with the zero-inflated generalized cluster-weighted models
- Mixture of multivariate \(t\) nonlinear mixed models for multiple longitudinal data with heterogeneity and missing values
- A note on the consistency of the maximum likelihood estimator under multivariate linear cluster-weighted models
- On fractionally-supervised classification: weight selection and extension to the multivariate t-distribution
- Mixtures of multivariate contaminated normal regression models
- Estimation of regression contour clusters -- an application of the excess mass approach to regression
- Model-based clustering
- Local statistical modeling via a cluster-weighted approach with elliptical distributions
- On the contaminated exponential distribution: a theoretical Bayesian approach for modeling positive-valued insurance claim data with outliers
- Robust model-based clustering with mild and gross outliers
- Clustering bivariate mixed-type data via the cluster-weighted model
- Multiple scaled contaminated normal distribution and its application in clustering
- The multivariate tail-inflated normal distribution and its application in finance
- Fitting insurance and economic data with outliers: a flexible approach based on finite mixtures of contaminated gamma distributions
- A new look at the inverse Gaussian distribution with applications to insurance and economic data
- Model-based clustering via skewed matrix-variate cluster-weighted models
- Dichotomous unimodal compound models: application to the distribution of insurance losses
- A Selective Overview and Comparison of Robust Mixture Regression Estimators
- Identifying subgroups of age and cohort effects in obesity prevalence
- Multivariate Contaminated Normal Censored Regression Model: Properties and Maximum Likelihood Inference
- Merging components in linear Gaussian cluster-weighted models
- Seemingly unrelated clusterwise linear regression for contaminated data
- Robust mixture regression modeling based on the normal mean-variance mixture distributions
- Parsimonious mixture-of-experts based on mean mixture of multivariate normal distributions
- Flexible mixture regression with the generalized hyperbolic distribution
- Editorial: Journal of Classification Vol. 41-3. Special issue for IFCS 2022
- Matrix-variate hidden Markov regression models: fixed and random covariates
- Parsimonious seemingly unrelated contaminated normal cluster-weighted models
- Skew multiple scaled mixtures of normal distributions with flexible tail behavior and their application to clustering
- Bayesian analysis of heavy-tailed Heckman selection models using Hamiltonian Monte Carlo
- Robust mixture of regression models using the symmetric α-stable distribution
- Detecting anomalies in European trade data using directed weighted multilayer dynamic networks
- Cluster weighted models for functional data
- Heavy-tailed matrix-variate hidden Markov models
- Multivariate generalized hidden Markov regression models with random covariates: physical exercise in an elderly population
- Finite Mixtures of Multivariate Contaminated Normal Censored Regression Models
- Matrix-variate cluster-weighted bilinear factor analyzers
- Copula-based mixtures of regression models for multivariate response data
- Mixtures of regressions using matrix-variate heavy-tailed distributions
- Heckman Selection-Contaminated Normal Model
- Skew dimension-wise scaled mixtures of normal distributions and their applications in model-based clustering
- Cluster validation for mixtures of regressions via the total sum of squares decomposition
This page was built for publication: Robust clustering in regression analysis via the contaminated Gaussian cluster-weighted model
Report a bug (only for logged in users!)Click here to report a bug for this page (MaRDI item Q2403302)