Speech screening prototype
Need to talk?

Shown at KERICE 2026International Islamic University Malaysia

Low-Cost Edge-AI Prototype for Speech-Based Depression Screening in Community Healthcare

A small device that asks a few simple questions, listens to how a person speaks, and gives health workers an early screening indication. It runs offline, and every recording stays on the device.

  • Offline
  • On-device
  • Private by design
  • Under $300

A screening aid, not a diagnosis. The decision always stays with the clinician.

The problem

Depression often goes unnoticed

Hard to reach

In Malaysia's B40 households (the lower-income 40 per cent) and in rural communities, depression often goes unnoticed and people are not referred for care.

Stigma

Stigma keeps many people from asking for help. A quiet, private first step can make it easier to start the conversation.

Recordings in the cloud

Most tools that listen for signs of depression run in the cloud. That puts sensitive recordings on someone else's server, and rules them out where the internet connection is poor.

What is missing is a screening aid that is small, offline and private, so people can trust it where stigma is high.

How it works

It listens to how people speak, not what they say

Depression often shows in speech as slower replies, longer pauses and a flatter voice. The device measures those patterns from the sound alone.

One question and answer, drawn as soundThe device asks a question. About one second later the person starts to answer. Their answer contains a longer pause before they continue speaking.Device asksPerson answers1.0 sbefore the answer startspause0 s5 s10 s15 s
Illustration, not a real recording.
  • Slowest replies
  • Usual time before answering
  • Pitch range
  • Voice steadiness
Lightweight classifier 33 parameters
Screening indication for referral to a clinician

All of this happens on the device

  1. 1

    The device asks a question

    Each question appears on the screen. No interviewer is needed to run a session, and because the device asks each question itself, it knows exactly when the question ends.

  2. 2

    It notices when the answer starts

    A small voice detector (Silero) hears the moment the person starts to reply. The gap in between is their time before answering.

  3. 3

    It measures how the person speaks

    Pauses, how far the pitch rises and falls, and how steadily the voice is sounded. It listens to the sound of speech and never to the words, so it is designed to work across languages.

  4. 4

    A lightweight classifier combines them

    A 33-parameter classifier turns the measurements into a screening indication, to help decide who may benefit from being seen. The decision always stays with the clinician.

  5. 5

    Nothing leaves the device

    No internet connection is needed. Recordings are processed on the device itself and are never sent to a server.

The prototype

A screening aid that fits on a clinic desk

A Raspberry Pi 5 (8 GB) with a 3.5-inch touchscreen, in a small matte case. Built from off-the-shelf parts for under $300.

Listening

A flat microphone on the table

A USB table microphone lies beside the device and picks up the person's voice. The device asks each question on its own, so no interviewer is needed to run a session.

Inside

Everything happens on the board

The voice detector, the measurements and the classifier all run on the Raspberry Pi 5. No cloud server, no internet connection. Recordings never leave the unit.

  • Case
  • 3.5-inch touchscreen
  • Raspberry Pi 5 (8 GB)
  • USB table microphone

Raspberry Pi OS (Debian 13), Python 3.13.

On the screen

What the health worker sees

After a session, the device shows a short screening summary: an indication, and the four speech measures behind it compared with a typical range.

Screening summary· 11-minute sessionScored on this deviceLowFollow-upFollow-up suggestedFollow-up suggestedSpeech patterns resemblethose of people withdepression symptoms.Slowest replies2.8 stypical rangelonger than typicalUsual time before answering1.03 stypical rangelonger than typicalPitch range173 Hztypical rangenarrower than typicalVoice steadiness0.65typical rangehigher than typicalA screening aid, not a diagnosis. The decision stays with the clinician.
Screening summary· 11-minute sessionScored on this deviceTime before answering1234567891011typicalPitch movementtypical rangethis personVoice steadiness0.65typicalthis personWhole sessiondevice asksperson answers0 min11 minA screening aid, not a diagnosis. The decision stays with the clinician.

Illustrative example. Values are group medians from the DAIC-WOZ study data, not a real person's result.

Read the screen as text

Follow-up suggested. Speech patterns resemble those of people with depression symptoms.

Slowest replies
2.8 s (typical about 2.2 s, longer than typical)
Usual time before answering
1.03 s (typical about 0.90 s, longer than typical)
Pitch range
173 Hz (typical about 211 Hz, narrower, a flatter voice)
Voice steadiness
0.65 (typical about 0.57, higher than typical)

A screening aid, not a diagnosis. The decision stays with the clinician.

Key numbers

Small, fast and private

11.7 s

to score an 11-minute interview

on the Raspberry Pi 5, about 55 times faster than real time, once the interview is recorded.

under $300

of hardware

Off-the-shelf parts: a Raspberry Pi 5, a small touchscreen and a table microphone.

33

-parameter classifier

Its whole model file is 3.6 KB. No graphics card and no cloud needed.

0.74

AUROC, cross-validated

How well the scores separate people with and without depression symptoms: 0.5 is a coin toss, 1.0 is perfect. Measured on 189 participants.

  • 9.0 s of the 11.7 s is voice-activity detection, a separate small neural network (Silero) that finds when someone is speaking.
  • 602 MB peak memory on the device.
  • Tested on DAIC-WOZ clinical interviews: 189 participants, counted as depressed at a PHQ-8 questionnaire score of 10 or more.

Results

In the range of published speech-only systems

On the same clinical interviews, our 33-parameter classifier is in the range of published systems that also use speech alone, including some far larger ones.

The score shown is F1 for the depressed group: it balances how many people with depression symptoms are found against how many of the people flagged really have them. Higher is better.

We compare only with speech-only systems. Systems that read transcripts need a written record of every word, and can pick up clues from the interviewer's questions.

F1 for the depressed group, DAIC-WOZ, speech only
  • Valstar et al. 2016 challenge baseline 0.41 test set
  • Ma et al. 2016 0.52 dev set
  • Al Hanai et al. 2018 0.63 dev set
  • Wu et al. 2023 large speech models 0.70–0.80 dev set
  • Ours 33-parameter classifier 0.53–0.65 cross-validated · test set

Dev set: the development data used while building a model, which usually scores higher than an unseen test set. Test set: held-out data. Ours: 0.53 cross-validated, 0.65 on the held-out test set.

Video

The project in three minutes

Subtitles are built into the video.

Read the transcript

How can you tell when someone is quietly struggling? Around one in twenty adults lives with depression. Even in high-income countries, only about one in three people with depression receive treatment.

In Malaysia's rural and B40 communities, stigma keeps many people from asking for help. Community clinics rarely have the time or the staff to screen everyone. Most tools that could help also run in the cloud, far from where coverage is poor.

So we asked a different question. What if there were a way to spot the signs of depression from the voice alone, on a portable, low-cost device?

This is our prototype. A Raspberry Pi 5 with a small touchscreen and an external microphone, built for under three hundred dollars.

The person answers a few short, simple questions on the screen. Our model listens in two ways.

First, temporal features, the timing of speech. How long each answer takes to begin, how often the person pauses, and how much of the conversation is theirs. Second, prosodic features, from acoustic analysis of the voice. How far the pitch rises and falls, how steadily the voice is sounded, and how loud it is. Depression often shows as slower replies, longer pauses and a flatter voice.

Its strongest cues come from both families: the slowest replies, how steadily the voice is sounded, and how far the pitch moves. Because it listens to the sound of speech and never to the words, it is designed to work across languages.

Everything runs on the device itself. It needs no internet, and no recording ever leaves the unit. We've tested it on the device, and it runs live, in real time.

Our model file is just 3.6 kilobytes, about a thousand times smaller than one phone photo. Because it is so small, it runs on a low-cost edge device, with no graphics card and no cloud.

Tested on recorded clinical interviews, it is competitive with published speech systems, some of them thousands of times larger.

It gives health workers an early indication of who may benefit from being seen. The decision always stays with the clinician.

Our next phase is to test and validate it with Malay speakers, in a clinical pilot at IIUM Kuantan. Listen early. Refer sooner.

Live demo

See the prototype running

A recording of the device running a screening session from start to finish.

Watch the demo

Opens on the video host's site in a new tab.

Next steps

A clinical pilot in Malay

The device listens to the sound of speech, never to the words, so it is designed to work across languages. The next step is to test and validate it with Malay speakers, in a clinical pilot at IIUM Kuantan.

Listen early. Refer sooner.

Open source

pyvoicebox

The voice measures are computed with pyvoicebox, a Python port of VOICEBOX, the speech-processing toolbox by Mike Brookes at Imperial College London. The port is free to use.

Team

International Islamic University Malaysia

  • Muhammad Fahreza AlghifariElectrical and Computer Engineering, Kulliyyah of Engineering[email protected]
  • Teddy Surya GunawanElectrical and Computer Engineering, Kulliyyah of Engineering
  • Mira KartiwiInformation Systems, Kulliyyah of Information and Communication Technology
  • Ali SophianMechatronics Engineering, Kulliyyah of Engineering
  • Jamalludin Ab RahmanCommunity Medicine, Kulliyyah of Medicine

Funded by the Prototype Development Research Grant Scheme (PRGS25-029-0073), Ministry of Higher Education Malaysia.