Skip to main content Start main content

Multilingual AI cognitive screening for Chinese-Speaking community enables early detection using speech-based biomarkers in English and Chinese

17 Sep 2026

Research and Innovation

Cognitive impairment, including mild cognitive impairment (MCI) and Alzheimer's disease, is a growing challenge in ageing societies. Early detection is crucial, as timely intervention can improve care planning, monitoring and quality of life. Yet cognitive decline is often identified only after symptoms have become more pronounced, limiting opportunities for support.
 
Current diagnostic tools remain imperfect. Screening tests such as the Mini Mental State Examination (MMSE) can be influenced by language, education and culture, while neuroimaging and specialist assessments are costly and not always accessible. This has made speech an increasingly attractive digital biomarker. Because speech reflects multiple cognitive processes, subtle changes in fluency, pausing and vocal delivery may provide a non-invasive and scalable route to earlier detection.
 
Prof. Hualou LIANG, Chair Professor of Neuroscience and Artificial Intelligence at The Hong Kong Polytechnic University, and his research team advance this line of work in a multilingual context, showing how AI-based speech analysis may support more practical and inclusive cognitive screening.
 
Their study was conducted as part of the INTERSPEECH 2024 TAUKADIAL Challenge, a task focused on detecting mild cognitive impairment and predicting cognitive scores from spontaneous speech. The research titled, “Multilingual prediction of cognitive impairment with Large Language Models and speech analysis” was published in Brain Sciences. The significance of this challenge lies in its multilingual scope. Previous work in speech-based dementia research has been heavily focused on English datasets, leaving uncertain whether reported findings can generalise across languages and cultures. This study directly addresses the gap by examining both English and Mandarin Chinese speech, thereby making a meaningful contribution to the development of more globally relevant cognitive assessment tools.

INTERSPEECH 2024 is a major international conference in speech science and technology, organised by the International Speech Communication Association. It brings together researchers and industry experts working on speech processing, artificial intelligence, spoken language systems and related applications. The conference is known for its challenge tracks, which provide shared tasks and benchmark datasets to advance the field through transparent comparison. One of these was the TAUKADIAL Challenge, focused on multilingual detection of cognitive impairment and cognitive score prediction from spontaneous speech.
 
The study’s dataset was collected through picture description tasks, a well-established format in cognitive assessment because it elicits spontaneous yet structured speech. Participants were required to describe three pictures. Across both languages, the study included 169 participants, comprising individuals with MCI and those with normal cognition, balanced by age and sex to reduce demographic bias. The study addressed two tasks. The first was MCI classification, in which the model distinguished cognitively healthy speakers from those with mild cognitive impairment. The second was MMSE score prediction, a regression task aimed at estimating an individual’s cognitive score directly from their speech. Together, these tasks reflect the need for both screening and monitoring, suggesting that speech-based AI could support not only binary risk detection but also more continuous assessment of cognitive status.
 
At the centre of the approach is Whisper, specifically whisper-large-v3, used as a foundation model for acoustic feature extraction. This is a notable methodological choice. Rather than relying on traditional handcrafted features alone, the team used Whisper's encoder to generate 1280-dimensional embeddings from the speech signal. These embeddings provide a rich learned representation of the audio recording and reflect the growing importance of foundation models in speech technology. Their use here signals an important shift in cognitive assessment research: from manually engineered markers towards large-scale, transferable representations capable of capturing subtle acoustic patterns across languages.

One of the most revealing elements of the study is its examination of between-language transfer. When models trained on one language were applied to the other, performance declined markedly. This finding has important implications. It shows that, although Whisper provides multilingual embeddings, the speech signatures of cognitive impairment are not fully language-independent. They remain shaped by phonetic, linguistic and cultural factors. Multilingual capability does not remove the need for localisation. On the contrary, effective deployment may depend on combining shared representation learning with language-specific adaptation.

The practical significance of this study is reinforced by its standing in the challenge itself. Among all participating teams in the INTERSPEECH 2024 TAUKADIAL Challenge, the proposed model ranked second for MCI classification and first for MMSE prediction. These rankings are impressive not merely as competition results, but because they demonstrate that a relatively streamlined system based on spontaneous speech and acoustic embeddings can achieve state-of-the-art performance against an international benchmark. The work therefore strengthens the case for speech as a clinically useful, low-burden digital biomarker.
 
More broadly, the study points towards a future in which cognitive screening may become more frequent, more remote and more inclusive. It helps define the next stage of research, including multimodal approaches that combine acoustic, linguistic and possibly non-verbal behavioural cues.
 
Source: Innovation Digest 8

 


Your browser is not the latest version. If you continue to browse our website, Some pages may not function properly.

You are recommended to upgrade to a newer version or switch to a different browser. A list of the web browsers that we support can be found here