Accessibility settings

Published on in Vol 9 (2026)

Preprints (earlier versions) of this paper are available at https://preprints.jmir.org/preprint/93279, first published .
Close-up of an elderly woman's face with gray hair, looking thoughtfully to the side.

Multimodal Dementia Prediction With Large Language Models: Cross-Attention Over Text, Audio, and Image

Multimodal Dementia Prediction With Large Language Models: Cross-Attention Over Text, Audio, and Image

Authors of this article:

Felix Agbavor1 Author Orcid Image ;   Hualou Liang1, 2, 3, 4 Author Orcid Image

Felix Agbavor   1 , PhD ;   Hualou Liang   1, 2, 3, 4 , PhD

1 School of Biomedical Engineering and Science, Drexel University, Philadelphia, PA, United States

2 Division of Artificial Intelligence and the Humanities, The Hong Kong Polytechnic University, Kowloon, China (Hong Kong)

3 Department of Language Science and Technology, The Hong Kong Polytechnic University, Kowloon, China (Hong Kong)

4 Departments of Data Science and Artificial Intelligence, The Hong Kong Polytechnic University, Kowloon, China (Hong Kong)

Corresponding Author:

  • Hualou Liang, PhD
  • Division of Artificial Intelligence and the Humanities
  • The Hong Kong Polytechnic University
  • HHB717, 7/F, 8 Hung Lok Road, Hung Hom
  • Kowloon
  • China (Hong Kong)
  • Phone: 852 2766-7697
  • Email: hualou.liang@polyu.edu.hk