Categorical variables collected through surveys often share a high degree of information, leading to substantial redundancy. If not properly accounted for, redundant information may result in misleading outcomes when distance-based techniques are applied, such as data visualization and profiling via Multidimensional Scaling (MDS). This issue typically arises when additive dissimilarity coefficients are employed in datasets characterized by moderate to strong associations among variables. In this work, we propose new dissimilarities for categorical data that explicitly account for the association structure of the data. These dissimilarities are subsequently combined with a robust distance for numerical variables, yielding a flexible and robust metric for multivariate heterogeneous data. The performance of the proposed metrics is evaluated under adverse scenarios involving underlying correlation structures and outlier contamination, and is compared with that of the classical Gower distance using MDS representations and a Nearest Neighbor classifier. In addition, the proposed methodology is illustrated through three real-data applications, addressing both profiling and classification tasks. The results show that the proposed distances are effective in isolating outlying units and in improving classification accuracy.
New distances for mixed-type data able to cope with redundant information / A. Grané, S.S.. - In: ASTA. ADVANCES IN STATISTICAL ANALYSIS. - ISSN 1863-818X. - (2026 Jun 22). [Epub ahead of print] [10.1007/s10182-026-00565-6]
New distances for mixed-type data able to cope with redundant information
S. Salini
Secondo
;G. InfanteUltimo
2026
Abstract
Categorical variables collected through surveys often share a high degree of information, leading to substantial redundancy. If not properly accounted for, redundant information may result in misleading outcomes when distance-based techniques are applied, such as data visualization and profiling via Multidimensional Scaling (MDS). This issue typically arises when additive dissimilarity coefficients are employed in datasets characterized by moderate to strong associations among variables. In this work, we propose new dissimilarities for categorical data that explicitly account for the association structure of the data. These dissimilarities are subsequently combined with a robust distance for numerical variables, yielding a flexible and robust metric for multivariate heterogeneous data. The performance of the proposed metrics is evaluated under adverse scenarios involving underlying correlation structures and outlier contamination, and is compared with that of the classical Gower distance using MDS representations and a Nearest Neighbor classifier. In addition, the proposed methodology is illustrated through three real-data applications, addressing both profiling and classification tasks. The results show that the proposed distances are effective in isolating outlying units and in improving classification accuracy.| File | Dimensione | Formato | |
|---|---|---|---|
|
unpaywall-bitstream-453826154.pdf
accesso aperto
Tipologia:
Publisher's version/PDF
Licenza:
Creative commons
Dimensione
4.4 MB
Formato
Adobe PDF
|
4.4 MB | Adobe PDF | Visualizza/Apri |
Pubblicazioni consigliate
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.




