Automated Early Dementia Detection for Procedural Narrative Tasks: An LLM-Driven System for Feature Extraction and Clinical Scoring
Early and reliable detection of mild cognitive impairment (MCI) and dementia remains an open challenge in clinical neuropsychology. Procedural discourse tasks such as the Peanut Butter and Jelly (PBJ) sandwich task elicit spontaneous speech in which the clinically relevant signal lies not only in how much is said, but in what procedural steps were mentioned, omitted, sequenced correctly, or expressed with lexical precision. Yet many automated approaches reduce transcripts to aggregate scores, offering limited insight into the underlying cognitive-linguistic evidence.
This thesis investigates a rubric-grounded LLM-as-a-judge methodology for transforming PBJ transcripts into interpretable procedural and discourse features, applied to the Delaware AphasiaBank dataset of CHAT-format recordings from Control and MCI participants. Three research questions structure the work: whether this methodology can extract diagnostically meaningful procedural-temporal features and how it compares to traditional linguistic metrics for MCI classification (RQ1, methodological viability); how a modular architecture can make PBJ assessment robust and auditable for clinical deployment (RQ2, engineering deployment); and how extracted PBJ evidence can be translated into actionable visual artifacts for clinician decision-making (RQ3, clinical translation).
To address RQ1, a lean batch pipeline scores transcripts across five rubrics: Main Concept Analysis (accuracy and completeness of ten canonical procedural propositions), Logical Action Sequencing (whether those concepts appear in the correct procedural order), Correct Information Units (token-level density of on-task informative speech), Coherence (global thematic unity and local utterance-level continuity), and Discourse-Pragmatic features (higher-order communicative adequacy). Together these dimensions capture not only what steps a participant produced, but how fluently, coherently, and pragmatically appropriately they were communicated. A monolithic and a separate-rubrics extraction mode are compared across GPT-4o and GPT-4.5. The findings are expected to support methodological viability and clinical inspectability rather than diagnostic superiority, with the main value of rubric-grounded extraction lying in providing an evidence-linked layer that makes omissions, sequencing errors, and discourse breakdowns directly inspectable.
RQ2 is addressed through a modular LangGraph system that decomposes rubric assessment into specialised, independently testable nodes with deterministic verification steps and preserved evidence traces for auditability. RQ3 is addressed via a clinician-facing dashboard translating feature vectors and MCI-risk signals into structured visual summaries that support review, correction, and interpretation in a human-in-the-loop workflow.
Together, these contributions advance the feasibility of scalable, transparent, and clinician-ready automated cognitive assessment grounded in established neuropsychological discourse methodology.
| Attribute | Value |
|---|---|
| Title (de) | Automated Early Dementia Detection for Procedural Narrative Tasks: An LLM-Driven System for Feature Extraction and Clinical Scoring |
| Title (en) | Automated Early Dementia Detection for Procedural Narrative Tasks: An LLM-Driven System for Feature Extraction and Clinical Scoring |
| Project | AssistD |
| Type | Master's Thesis |
| Status | started |
| Student | Philip Werz |
| Advisor | Alexandre Mercier |
| Supervisor | Prof. Dr. Florian Matthes |
| Start Date | 26.04.2026 |
| Sebis Contributor Agreement signed | yes |
| Checklist filled | yes |
| Submission date | 26.10.2026 |