Strider1

Soundskrit Strider: How Can Four Directional MEMS Microphones Enable Real-Time Sound Tracking?

What Is the Sound Tracking Demo Kit (Strider)?

Soundskrit’s Sound Tracking Demo Kit (Strider) is a development platform designed to demonstrate Soundskrit’s Sound Localization performance. The system mainly consists of the Strider Board, PARDI Board, and Camera Board. The Strider Board carries four Directional MEMS Microphones for capturing sound information from different directions. The PARDI Board interfaces the MEMS microphone signals to a PC over USB, while the Camera Board connects directly to the PC through a separate USB cable. By combining audio information with the camera image, users can intuitively observe the direction and position of the detected sound source.

Strider Board: A Four-SKR0610 Sound Localization Array

The Strider Board is the core hardware used for sound detection in the system. The PCB carries four SKR0610 microphones, each fitted with a special mesh that gives it a Cardioid directional response. The four microphones are arranged in a square, one at each corner of the PCB, with approximately 47 mm spacing between them. By analyzing the differences among the signals captured by the four directional microphones, the backend algorithm can estimate the direction of arrival of the sound. Unlike a conventional microphone system that simply captures sound, this configuration is designed to help the system understand:Which direction is the sound coming from?

Visualizing the Sound Source Directly

To make Sound Localization performance easier to evaluate, Soundskrit integrates the localization function into the Soundskrit Demo Kit Interface. When the system detects a sound source, a green dot appears on the screen to indicate the detected sound position. If the sound source is outside the camera’s Field of View, a green arrow appears instead to indicate the direction of the sound source. In other words, even if the speaker is outside the camera image, the system can still use sound to tell you:

“Which side of the scene the sound is coming from.”

With this visual approach, users can quickly evaluate DOA (Direction of Arrival) accuracy and compare performance under different distances, angles, and environmental conditions.

AI VAD: Track Speech Only or All Sounds?

In addition to DOA-based sound localization, Strider also integrates AI Voice Activity Detection (VAD).The purpose of VAD is to determine whether the detected sound is speech.
Users can enable or disable VAD directly from the Demo Kit Interface:

VAD Enabled: The system focuses on speech and localizes the speaker.

VAD Disabled: The system is not limited to speech and can localize other sounds in the environment.
This allows developers to decide whether the system should “find where a person is speaking” or “find where a sound is coming from”, depending on the application. This is an important distinction for applications such as smart cameras, conference systems, robots, and other sound-event detection systems.

Localizing Up to 3 Sound Sources Simultaneously

In addition to single-source localization, Strider also supports multiple sound sources. Users can set the Speaker Count in the interface from 1 to 3, allowing the system to track a single source or localize multiple simultaneous sources. For example, when Speaker Count is set to 1 and VAD is enabled, the system reports the dominant or strongest speech source. When configured for multiple sources, the system can display several detected sound sources at the same time, allowing developers to evaluate localization performance in multi-speaker or multi-source environments.

How Does Strider’s DOA Algorithm Work?

Strider’s processing chain first captures the raw audio signals from the four Cardioid microphones and then analyzes the differences among the signals received by each microphone. The overall process can be understood as:

4 Directional MEMS Microphones → Audio Signal Analysis → VAD Detection → DOA Estimation → Sound Source Coordinates

The DOA algorithm uses the signals captured by the four Cardioid microphones to determine the direction of arrival and then outputs the position information of the sound source. The Soundskrit Demo Kit Interface then converts these results into a green dot or directional arrow on the screen, turning abstract DOA data into an easily understood visual sound-source position.

Extending Sound Localization to More Applications

Sound localization itself is not the final goal. What matters more is what the system can do after it knows where the sound is coming from.
The DOA information output by Strider can support functions such as:
Voice Tracking: Track the location of the current speaker.
Beam Steering: Steer the main pickup direction toward the sound source.
Camera Tracking: Allow the camera to find or track a target based on sound direction.
Human-Machine Interaction: Allow robots or AI devices to know which direction a user is speaking from.

Therefore, DOA can be considered an important foundation for spatial audio awareness. It allows a device not only to hear sound, but also to understand where that sound is located in space.

For more information about Soundskrit MEMS, please visit AATC’s product pages: SKR0710 SKR0610 SKR0410

If you have any questions, please feel free to contact AATC at: Soundskrit@aatc.tw