Skip to main navigation Skip to search Skip to main content

Identifying Crying and Screaming in Audio Recordings of Autistic Children Using Transfer Learning

Research output: Contribution to journalArticlepeer-review

Abstract

Crying and screaming (CS) events are important indicators of distress and emotional dysregulation in children with autism spectrum disorder (ASD) that are difficult for researchers and clinicians to quantify. This study aimed to develop an algorithm for identifying and quantifying crying and screaming segments in behavioral assessment audio recordings of 206 children (80% with ASD). We trained a multilayer perceptron (MLP) model to classify audio segments containing CS versus speech. We compared classification accuracy across multiple MLP models trained with different input embeddings extracted from YAMNet, VGGish, OpenL3, Whisper (tiny, base, small, and medium), and Wav2Vec2 models, as well as the eGeMAPs feature set. Integrating classification probabilities of Whisper-medium and eGeMAPs achieved the highest performance, yielding Matthew’s correlation coefficient of 0.56 ± 0.10. This demonstrates strong performance for this challenging, highly imbalanced real-world task, yielding an improvement of 1.8% over the next single-embedding model, when tested on independent recordings from children not included in the training set. The number and duration of algorithm-identified CS segments per child were highly correlated with the number [concordance correlation coefficient (CCC) = 0.79, Pearson’s r(204) = 0.83] and duration [CCC = 0.81, r(204) = 0.85] of manually annotated segments. These findings demonstrate that the accuracy of an algorithm outperforms alternative embedding-based methods and provides a scalable and objective tool for quantifying distress and emotional dysregulation in ASD children that will be of high utility for basic and clinical research purposes.

Original languageEnglish
JournalIEEE Transactions on Computational Social Systems
DOIs
StateAccepted/In press - 1 Jan 2026

Keywords

  • Audio event identification
  • autism
  • classification
  • cry
  • embeddings
  • hand-crafted features
  • pretrained models
  • scream

ASJC Scopus subject areas

  • Modelling and Simulation
  • Social Sciences (miscellaneous)
  • Human-Computer Interaction

Fingerprint

Dive into the research topics of 'Identifying Crying and Screaming in Audio Recordings of Autistic Children Using Transfer Learning'. Together they form a unique fingerprint.

Cite this