Automated Early Dementia Detection for Descriptive Narrative Tasks: An LLM-Driven System for Feature Extraction and Clinical Scoring
Early assessment of Mild Cognitive Impairment (MCI) requires methods that capture clinically meaningful discourse patterns without treating automated outputs as clinical diagnoses. Picture-description tasks such as the Cookie Theft Picture (CTP) elicit spontaneous discourse in which the relevant signal lies not only in how much speech is produced, but in how the transcript relates to the visual scene. However, many automated approaches reduce transcripts to aggregate linguistic or distributional scores, offering limited insight into what was mentioned, omitted, distorted, vague, or inferentially elaborated.
This thesis investigates a rubric-grounded LLM-as-a-judge methodology for transforming CTP transcripts into interpretable semantic and discourse features. Using the Delaware subset of DementiaBank, 112 scorable CHAT-format recordings from 99 participants were analyzed across healthy-control and MCI groups. The empirical evaluation extracted evidence-linked judgments for information units, main concepts, empty speech, higher-order inference, and spatial distribution, and compared the resulting rubric-grounded features with classical linguistic, reference-based, embedding-based, and combined transcript representations under participant-grouped cross-validation.
The results indicate that subtle MCI-related discourse differences are difficult to capture through quantitative feature aggregation alone, as MCI narratives often remain close to healthy-control descriptions. The findings therefore support methodological viability and clinical inspectability rather than diagnostic superiority. The main value of rubric-grounded extraction lies in providing an evidence-linked qualitative layer that makes omissions, vague expressions, incomplete concepts, and higher-order inferences inspectable.
Building on these findings, the thesis further motivates a modular workflow that decomposes rubric assessment into specialized nodes, adds deterministic verification, and preserves evidence traces for auditability. Finally, it explores how these outputs can be translated into clinician-facing visual artifacts that support review, correction, and interpretation in a human-in-the-loop workflow.
Research Questions:
- Can an LLM-as-a-judge methodology extract diagnostically meaningful features from descriptive narrative tasks for early dementia detection, and how does its classification performance compare to traditional linguistic metrics?
- How can a modular system architecture make this methodology robust and auditable for clinical deployment?
- How can the system's outputs be translated into actionable visual artifacts for clinician decision-making?
| Attribute | Value |
|---|---|
| Title (de) | Automated Early Dementia Detection for Descriptive Narrative Tasks: An LLM-Driven System for Feature Extraction and Clinical Scoring |
| Title (en) | Automated Early Dementia Detection for Descriptive Narrative Tasks: An LLM-Driven System for Feature Extraction and Clinical Scoring |
| Project | AssistD |
| Type | Master's Thesis |
| Status | started |
| Student | Marcel Seitz |
| Advisor | Alexandre Mercier |
| Supervisor | Prof. Dr. Florian Matthes |
| Start Date | 26.04.2026 |
| Sebis Contributor Agreement signed | yes |
| Checklist filled | yes |
| Submission date | 26.10.2026 |