Learning  Realistic Patterns from Visually Unrealistic Stimuli:  Generalization and Data Anonymization

Konstantinos Nikolaidis; Stein Kristiansen; Thomas Plagemann; Vera Goebel; Knut Liestøl; Mohan Kankanhalli; Gunn Marit Traaen; Britt Overland; Harriet Akre; Lars Aakerøy; Sigurd Steinshamn

doi:10.1613/jair.1.13252

PDF

Published: Dec 3, 2021

DOI: https://doi.org/10.1613/jair.1.13252

Keywords:

machine learning, knowledge representation, neural networks, information retrieval

Konstantinos Nikolaidis

a:1:{s:5:"en_US";s:18:"University of Oslo";}

Stein Kristiansen

Thomas Plagemann

Vera Goebel

Knut Liestøl

Mohan Kankanhalli

Gunn Marit Traaen

Britt Overland

Harriet Akre

Lars Aakerøy

Sigurd Steinshamn

Abstract

Good training data is a prerequisite to develop useful Machine Learning applications. However, in many domains existing data sets cannot be shared due to privacy regulations (e.g., from medical studies). This work investigates a simple yet unconventional approach for anonymized data synthesis to enable third parties to benefit from such anonymized data. We explore the feasibility of learning implicitly from visually unrealistic, task-relevant stimuli, which are synthesized by exciting the neurons of a trained deep neural network. As such, neuronal excitation can be used to generate synthetic stimuli. The stimuli data is used to train new classification models. Furthermore, we extend this framework to inhibit representations that are associated with specific individuals. We use sleep monitoring data from both an open and a large closed clinical study, and Electroencephalogram sleep stage classification data, to evaluate whether (1) end-users can create and successfully use customized classification models, and (2) the identity of participants in the study is protected. Extensive comparative empirical investigation shows that different algorithms trained on the stimuli are able to generalize successfully on the same task as the original model. Architectural and algorithmic similarity between new and original models play an important role in performance. For similar architectures, the performance is close to that of using the original data (e.g., Accuracy difference of 0.56%-3.82%, Kappa coefficient difference of 0.02-0.08). Further experiments show that the stimuli can provide state-ofthe-art resilience against adversarial association and membership inference attacks.

Issue

Vol. 72 (2021)

Section

Articles

Article Sidebar

Main Article Content

Abstract

Article Details