Skip to main content Start main content

Successful Conclusion of HumOmni 2026 Challenge: Breakthroughs in Human-Centric Omni-Model Evaluation

31 Aug 2026


We are pleased to announce the successful conclusion of the HumOmni 2026 Challenge, co-organised by PolyU COMP, Huawei, and Peking University from April to July 2026. The competition was the first benchmark dedicated to Human-Centric Omni-Model Evaluation, focusing on contextualised affective speech generation and proactive multimodal interaction in realistic environments, attracting nearly 25 teams from Hong Kong, mainland China, and the United States.

 

The challenge featured two tracks designed to develop next-generation models in two key areas:

  • Track 1: Empathy-Aware Speech Evaluation (EmpathyEval) – Evaluates how multimodal systems understand human context and paralinguistic cues to produce appropriate affective spoken responses.

  • Track 2: Proactive Multimodal Interaction (ProactivEval) – Assesses proactive multimodal systems on their ability to decide when to respond and what to say during streaming video understanding.

 

The award ceremony was held on the PolyU campus on 28 August 2026, bringing together researchers, industry leaders, and students for an engaging program featuring keynote speeches, winning team presentations, recruitment opportunities, and the award presentation. Distinguished speakers delivered insightful talks on:

 

  • “From world reconstruction, generation, to interaction” by Dr QI Xiaojuan, Associate Professor at The University of Hong Kong and a member of the Deep Vision Lab.

  • “Multimodal Large Language Models with Ascend NPUs” by Dr HONG Lanqing, Senior Researcher at Huawei Noah's Ark Lab, Hong Kong.

 

During the sharing sessions, winning teams presented their novel methodologies, including our MPhil student, LAM Ping Him, who proudly secured two awards: 1st Place in Track 2 (ProactivEval) and 2nd Place in Track 1 (EmpathyEval). Ping Him successfully developed lightweight AI solutions across two distinct challenges:

 

Track 1 - Voice Empathy Evaluation

Built an audio-focused system that evaluates emotional nuance in spoken conversations rather than relying on text. Using WavLM-large to extract voice patterns and track delivery metrics over time, the model accurately judges empathy from tone, pace, and warmth.

 

Track 2 - Proactive Video Commentary

Created a training-free filtering framework that enables an AI assistant to speak only when relevant visual changes occur. It uses a Perception Gate to skip static frames and cut compute costs, paired with a Prompt Ensemble to analyse new scenes from multiple perspectives.

 

His outstanding performance highlights the Department's commitment to fostering research excellence and innovation in AI and showcases the strength of our students on the international stage.

 

Click here to read more about HumOmni 2026.



Your browser is not the latest version. If you continue to browse our website, Some pages may not function properly.

You are recommended to upgrade to a newer version or switch to a different browser. A list of the web browsers that we support can be found here