Categorical variables collected through surveys often share a high degree of information, leading to substantial redundancy. If not properly accounted for, redundant information may result in misleading outcomes when distance-based techniques are applied, such as data visualization and profiling via Multidimensional Scaling (MDS). This issue typically arises when additive dissimilarity coefficients are employed in datasets characterized by moderate to strong associations among variables. In this work, we propose new dissimilarities for categorical data that explicitly account for the association structure of the data. These dissimilarities are subsequently combined with a robust distance for numerical variables, yielding a flexible and robust metric for multivariate heterogeneous data. The performance of the proposed metrics is evaluated under adverse scenarios involving underlying correlation structures and outlier contamination, and is compared with that of the classical Gower distance using MDS representations and a Nearest Neighbor classifier. In addition, the proposed methodology is illustrated through three real-data applications, addressing both profiling and classification tasks. The results show that the proposed distances are effective in isolating outlying units and in improving classification accuracy.

New distances for mixed-type data able to cope with redundant information / A. Grané, S.S.. - In: ASTA. ADVANCES IN STATISTICAL ANALYSIS. - ISSN 1863-818X. - (2026 Jun 22). [Epub ahead of print] [10.1007/s10182-026-00565-6]

New distances for mixed-type data able to cope with redundant information

S. Salini
Secondo
;
G. Infante
Ultimo
2026

Abstract

Categorical variables collected through surveys often share a high degree of information, leading to substantial redundancy. If not properly accounted for, redundant information may result in misleading outcomes when distance-based techniques are applied, such as data visualization and profiling via Multidimensional Scaling (MDS). This issue typically arises when additive dissimilarity coefficients are employed in datasets characterized by moderate to strong associations among variables. In this work, we propose new dissimilarities for categorical data that explicitly account for the association structure of the data. These dissimilarities are subsequently combined with a robust distance for numerical variables, yielding a flexible and robust metric for multivariate heterogeneous data. The performance of the proposed metrics is evaluated under adverse scenarios involving underlying correlation structures and outlier contamination, and is compared with that of the classical Gower distance using MDS representations and a Nearest Neighbor classifier. In addition, the proposed methodology is illustrated through three real-data applications, addressing both profiling and classification tasks. The results show that the proposed distances are effective in isolating outlying units and in improving classification accuracy.
association; Dbstats; mixed-type data; Horseshoe effect; MDS; redundant information
Settore STAT-01/A - Statistica
22-giu-2026
22-giu-2026
Article (author)
File in questo prodotto:
File Dimensione Formato  
unpaywall-bitstream-453826154.pdf

accesso aperto

Tipologia: Publisher's version/PDF
Licenza: Creative commons
Dimensione 4.4 MB
Formato Adobe PDF
4.4 MB Adobe PDF Visualizza/Apri
Pubblicazioni consigliate

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2434/1260575
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus 1
  • ???jsp.display-item.citation.isi??? 1
  • OpenAlex 1
social impact