ProjectsSocial Robotics
Back to projects
Social Robotics

Social Robotics

Fine-grained speech emotion recognition for companion robots. IEEE CAI 2023.

PythonTensorFlowPyTorchDockerROS

Context

Social companion robots need to read human emotion in real time, not just recognise it as a label. Existing speech emotion recognition systems flatten emotion into six coarse classes, happy, sad, angry, neutral and so on, which is useful for demos and useless in an actual interaction. A companion robot doesn't need to know that the user is sad. It needs to know whether they're quietly sad or in crisis, mildly frustrated or seething, so it can respond with appropriate intensity.

So we built a fine-grained system that classifies both the emotion and its intensity from a single utterance.

Speech as an image

Each audio clip is turned into a three-channel image: a log-mel spectrogram stacked with harmonic and percussive decompositions of the same signal. Three complementary acoustic views of one utterance, arranged so a network built for photographs can read them. The harmonic channel carries the pitched, voiced part of speech, the percussive channel carries onsets and attack, and the log-mel channel carries the overall energy distribution. Stacking them where a photograph would keep red, green and blue means a pretrained network reads all three at once, in the same receptive field, rather than through three separate towers that have to be fused afterwards.

That representation is what unlocked transfer learning. We compared VGG19, InceptionV3, and InceptionResNetV2 as backbones, all fine-tuned from ImageNet weights, an approach that beat training an audio-native model from scratch on datasets of this size. A network trained on millions of photographs has already learned edges, textures and repeated structure, and a spectrogram is made of exactly those things: formant bands are edges, a rolled consonant is a texture, a rising intonation is a diagonal. InceptionResNetV2 carried the best accuracy-to-latency trade-off for the robotics deployment, which is the trade that actually decided it rather than the top line on a leaderboard.

Intensity as a first-class output

On top of the emotion classifier sits a separate head predicting intensity, normal versus strong expression of the same emotion, trained jointly with a composite loss rather than bolted on as a second stage. Joint training produced better calibrated intensity predictions than a pipeline did, which follows from what the two heads share: the acoustic evidence for anger and the acoustic evidence for how angry are largely the same evidence, and a second-stage model that only sees the first stage's label has thrown most of it away.

The output changed what the robot could do with a prediction, and that was the point of the whole exercise. A sad-strong response is not a sad-normal response. One calls for stepping closer and staying quiet, the other for a light remark, and a system that returns only the word "sad" leaves the robot guessing between them. CREMA-D and RAVDESS supplied the emotion and intensity supervision, with GoEmotions used to widen the fine-grained vocabulary in adjacent experiments.

On RAVDESS the fine-grained emotion and intensity classification came in at 95.85 ± 1.38% accuracy.

Running on the robot

On device
Utterance
Three-channel image
log-mel, harmonic, percussive
InceptionResNetV2
fine-tuned from ImageNet
Emotion plus intensity
Robot response
matched to the intensity, not just the class

The system was integrated with the MiRo-E social companion robot and deployed to an NVIDIA Jetson Nano for on-device inference, containerised with Docker. Edge inference was the whole point rather than an optimisation: a cloud round-trip breaks the illusion of presence that a companion robot depends on, and a robot that pauses to think over the network stops reading as alive. It also decides the backbone, which is why the accuracy-to-latency trade beat raw accuracy in that comparison. A model that scores a point higher and misses the beat of the conversation is the worse model on this hardware, and the Jetson Nano is not a machine you can talk out of that constraint.

MiRo-E social companion robot

Publication

Affective Computing for Social Companion Robots Using Fine-grained Speech Emotion Recognition

Ahuja, S., Shabani, A., 2023 IEEE Conference on Artificial Intelligence (CAI), Santa Clara, June 2023.

DOI: 10.1109/cai54212.2023.00146