Improved visualization of high-dimensional data using the distance-of-distance transformation

PLoS Comput Biol. 2022 Dec 20;18(12):e1010764. doi: 10.1371/journal.pcbi.1010764. eCollection 2022 Dec.

Abstract

Dimensionality reduction tools like t-SNE and UMAP are widely used for high-dimensional data analysis. For instance, these tools are applied in biology to describe spiking patterns of neuronal populations or the genetic profiles of different cell types. Here, we show that when data include noise points that are randomly scattered within a high-dimensional space, a "scattering noise problem" occurs in the low-dimensional embedding where noise points overlap with the cluster points. We show that a simple transformation of the original distance matrix by computing a distance between neighbor distances alleviates this problem and identifies the noise points as a separate cluster. We apply this technique to high-dimensional neuronal spike sequences, as well as the representations of natural images by convolutional neural network units, and find an improvement in the constructed low-dimensional embedding. Thus, we present an improved dimensionality reduction technique for high-dimensional data containing noise points.

Publication types

  • Research Support, Non-U.S. Gov't

MeSH terms

  • Algorithms*
  • Neural Networks, Computer*
  • Neurons / physiology

Grants and funding

This project was supported by a BMBF Grant to M.V. (Computational Life Sciences, project BINDA, 031L0167). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript.