Time-continuous Speech Emotion Recognition Benchmark

Evaluation on MSP-Podcast Test2 set

Benchmark for Continuous Speech Emotion Recognition

Benchmark for Continuous Speech Emotion Recognition

This portal serves as a continuous evaluation platform for time-continuous Speech Emotion Recognition (SER) systems. Researchers and developers are invited to submit their models and compare performance using a standardized benchmark based on the MSP-Conversation corpus. The dataset contains naturalistic conversation recordings, annotated continuously over time by multiple raters along three emotional attributes.

Benchmark Task

Participants can evaluate their models on the following task:

  • Time-Continuous Emotional Attribute Prediction: Time-continuous regression of three emotional attributes — arousal (calm to active), valence (negative to positive), and dominance (weak to strong) — tracked throughout each conversation, on a scale of -100 to 100.

Guidelines

Access to the MSP-Conversation corpus requires signing the academic license agreement (Don't forget to sign at the end of third page). Interested users should email the signed agreement to Prof. Carlos Busso ().

Participants may use standard pre-trained models such as wav2vec2.0, HuBERT, and other general-purpose self-supervised models. However, using models pre-trained on emotion-specific datasets is not allowed. To ensure comparability, training should be limited to the MSP-Conversation corpus only.

Submissions must follow the required CSV format and should be uploaded via the submission portal. Once submitted, results will appear on the public leaderboard.

Evaluation

Submissions are automatically evaluated using the time-continuous Concordance Correlation Coefficient (CCC), averaged across arousal, valence, and dominance.

This platform is open year-round to promote ongoing progress and transparent comparison in continuous speech emotion recognition research.

Submission Instructions

Before submission, please read and follow the instructions carefully.
Only registered team submissions will be accepted. To register please visit the overview tab.
Each registered email is permitted a maximum of one submission per two weeks per task. For the first submission, participants may choose any preferred team name. However, it is important to use the same team name for subsequent submissions, as any different name will result in rejection.

The submission portal is open year-round.

Task: Time-Continuous Emotional Attributes

  • Predictions are time-continuous estimates of the following attributes, on a scale from -100 to 100:
    • Arousal (-100 calm, 100 active)
    • Valence (-100 negative, 100 positive)
    • Dominance (-100 weak, 100 strong)
  • Values outside this range will create an error
  • Metric: average of the time-continuous Concordance Correlation Coefficient (CCC)
  • MSP-Conversation annotations include a constant 3-second reaction lag — the time evaluators need to perceive the emotional content and react — which is compensated for during ground-truth processing. As a result, the last 3 seconds of each conversation part are not annotated and should not be included in your submission. See the MSP-Conversation paper for details.
  • Example submission file:
  • Conversation_part, Central_Time, Arousal, Valence, Dominance
    MSP-Conversation_0002_1, 0.0, 31.4742, 40.5867, 34.3325
    MSP-Conversation_0002_1, 0.1, 30.9517, 40.1933, 36.5408
    MSP-Conversation_0002_1, 0.2, 29.1192, 39.4083, 36.1400
    MSP-Conversation_0002_1, 0.3, 31.8550, 36.7067, 37.1150
    ...
  • Sample file (link pending — to be updated once the sample file is uploaded to Google Drive)

Submission Form


Submissions are not open yet. When the challenge starts, your file will be checked and scored immediately and the results will appear on the right.

Results
Submit a file and your scores will appear here.

Leaderboard

The leaderboard will be available soon.

Interspeech 2025 Challenge FAQ

FAQs - Submission

Registering

Do all team members need a license?
Yes, every institution should have a license.

Should the license be signed?
Yes, at the end of the third page.

Submission

What metrics will be used?
Macro-F1 score for task 1, and concordance correlation coefficient (CCC) for task 2.

What is the file format for submission?
Submissions must be in .csv format, detailed in the submission section of the website.

Will we have access to the leaderboard during submission?
Yes, the leaderboard will update shortly after each submission.

Admin Login

Add a new registered email to the allowed list.