An evaluation,
and a second look, a year later
Many caregivers don't live with the person they care for. The person with dementia is home; the caregiver is at work, or across town. Background research described the pattern: leaving work mid-shift when something felt wrong, and a low background dread on the quiet days. The not-knowing was its own cost.
In 2024 my team evaluated a caregiver monitoring prototype. In 2026 I re-analysed the same data alone, and it reframed everything. This case study documents both passes.
A five-second promise,
and my job in evaluating it
The prototype was a mobile app built on a single promise, call it the five-second rule: within seconds of opening it, the caregiver should know where the person with dementia is. Every feature either delivered that or supported it.

The project was a team of four; research, concepts, and prototyping were shared. In the final week I unified our high-fidelity screens under one iOS-based system. The evaluation itself was mine, I designed the eight tasks, wrote the questionnaires, ran every session and interview, and authored the report.

An evaluation with targets,
not just observations
The evaluation tested all eight features against targets set in advance. For a safety-critical app, success can't be defined after the fact: if a caregiver can't check location in under fifteen seconds, or set a geofence without help, calling those tasks "usable" in hindsight means nothing. The targets forced a definition of success before a single participant arrived.
Four participants worked the Figma prototype remotely, screen-shared. Eight tasks ran against the targets. Three are worth watching.

Three tasks worth watching
The SUS was 81.5; six of eight tasks passed target; three of four participants said they'd recommend the app. On its own terms, the prototype did well. But three tasks are worth watching, one that passed cleanly, one that failed on the numbers, and one that passed the numbers while the stopwatch disagreed. The distance between those last two is the rest of this case study.

The other five tasks echoed these at smaller scale, two shared Task 4's discoverability problem, three passed as cleanly as Task 1.
Eight problems,
three underlying causes
The evaluation was finished; the data wasn't. Task 8 kept coming back to me. The metrics answer one question, can the user do the thing, and Task 8 passed it: three of four completed it unassisted, P4 rated it easy.
But there's a second question the metrics never ask: does the user's confidence match reality. P4 said easy and spent 1.27 minutes finding the alert. In a low-stakes app that gap is minor friction. In an app whose whole promise is finding the person in seconds, it's the failure the numbers can't see.
One participant said the task was easy. She spent 1.27 minutes finding the alert.
Observation, 2024 evaluation
Once I had that lens, Task 8 stopped looking like a one-off. The same gap ran through the evaluation, and grouped by the kind of gap, eight scattered task problems collapsed into three root causes.
A report, a brief,
and a redesign in progress
The 2024 phase produced a graded, mixed-methods report. The 2026 phase produced what the report couldn't: a compact analytical brief, three root causes, each tied to specific tasks, each with a design implication clear enough to build from. Those three causes are the brief for the redesign I'm on now.
Reflection
- 01The 2024 evaluation was rigorous in method but small in sample: four participants, all female, none with real dementia caregiving experience. The findings held internally, but before any of them generalise they need users who actually live the problem. Rigour of method is not rigour of sample.
- 02Coming back to my own data a year later, with different questions, changed what I could see. That wasn't luck; it was distance. In future evaluations I want to build the "come back later with fresh questions" step in on purpose, not by accident.
- 03The distinction I ended up drawing, between whether the user can do the thing and whether their confidence matches reality, isn't specific to this app. Any interface with real-world stakes benefits from asking both. That's the lens I take forward.
