Abstract
Crying and screaming (CS) events are important indicators of distress and emotional dysregulation in children with autism spectrum disorder (ASD) that are difficult for researchers and clinicians to quantify. This study aimed to develop an algorithm for identifying and quantifying crying and screaming segments in behavioral assessment audio recordings of 206 children (80% with ASD). We trained a multilayer perceptron (MLP) model to classify audio segments containing CS versus speech. We compared classification accuracy across multiple MLP models trained with different input embeddings extracted from YAMNet, VGGish, OpenL3, Whisper (tiny, base, small, and medium), and Wav2Vec2 models, as well as the eGeMAPs feature set. Integrating classification probabilities of Whisper-medium and eGeMAPs achieved the highest performance, yielding Matthew’s correlation coefficient of 0.56 ± 0.10. This demonstrates strong performance for this challenging, highly imbalanced real-world task, yielding an improvement of 1.8% over the next single-embedding model, when tested on independent recordings from children not included in the training set. The number and duration of algorithm-identified CS segments per child were highly correlated with the number [concordance correlation coefficient (CCC) = 0.79, Pearson’s r(204) = 0.83] and duration [CCC = 0.81, r(204) = 0.85] of manually annotated segments. These findings demonstrate that the accuracy of an algorithm outperforms alternative embedding-based methods and provides a scalable and objective tool for quantifying distress and emotional dysregulation in ASD children that will be of high utility for basic and clinical research purposes.
| Original language | English |
|---|---|
| Journal | IEEE Transactions on Computational Social Systems |
| DOIs | |
| State | Accepted/In press - 1 Jan 2026 |
Keywords
- Audio event identification
- autism
- classification
- cry
- embeddings
- hand-crafted features
- pretrained models
- scream
ASJC Scopus subject areas
- Modelling and Simulation
- Social Sciences (miscellaneous)
- Human-Computer Interaction
Fingerprint
Dive into the research topics of 'Identifying Crying and Screaming in Audio Recordings of Autistic Children Using Transfer Learning'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver