Home > News & Updates > Arduino News > SPEECH RECOGNITION ON AN ARDUINO NANO?

SPEECH RECOGNITION ON AN ARDUINO NANO?

Summary of SPEECH RECOGNITION ON AN ARDUINO NANO?


During quarantine, [Peter] developed a speech recognition system using an Arduino Nano. The project demonstrates low-level optimization to achieve 9ksps sampling rates and uses digital bandpass filters instead of Fourier transforms due to processing constraints. Training occurs on a PC which generates templates compiled directly onto the microcontroller for real-time command execution. This offline approach offers a privacy-focused alternative to cloud-based AI for simple voice commands.

Parts used in the Speech Recognition on an Arduino Nano:

  • Arduino Nano
  • MAX9814 microphone amplifier
  • Custom PC program

Like most of us, [Peter] had a bit of extra time on his hands during quarantine and decided to take a look back at speech recognition technology in the 1970s. Quickly, he started thinking to himself, “Hmm…I wonder if I could do this with an Arduino Nano?” We’ve all probably had similar thoughts, but [Peter] really put his theory to the test.

The hardware itself is pretty straightforward. There is an Arduino Nano to run the speech recognition algorithm and a MAX9814 microphone amplifier to capture the voice commands. However, the beauty of [Peter’s] approach, lies in his software implementation. [Peter] has a bit of an interplay between a custom PC program he wrote and the Arduino Nano. The learning aspect of his algorithm is done on a PC, but the implementation is done in real-time on the Arduino Nano, a typical approach for really any machine learning algorithm deployed on a microcontroller. To capture sample audio commands, or utterances, [Peter] first had to optimize the Nano’s ADC so he could get sufficient sample rates for speech processing. Doing a bit of low-level programming, he achieved a sample rate of 9ksps, which is plenty fast for audio processing.

To analyze the utterances, he first divided each sample utterance into 50 ms segments. Think of dividing a single spoken word into its different syllables. Like analyzing the “se-” in “seven” separate from the “-ven.” 50 ms might be too long or too short to capture each syllable cleanly, but hopefully, that gives you a good mental picture of what [Peter’s] program is doing. He then calculated the energy of 5 different frequency bands, for every segment of every utterance. Normally that’s done using a Fourier transform, but the Nano doesn’t have enough processing power to compute the Fourier transform in real-time, so Peter tried a different approach. Instead, he implemented 5 sets of digital bandpass filters, allowing him to more easily compute the energy of the signal in each frequency band.

The energy of each frequency band for every segment is then sent to a PC where a custom-written program creates “templates” based on the sample utterances he generates. The crux of his algorithm is comparing how closely the energy of each frequency band for each utterance (and for each segment) is to the template. The PC program produces a .h file that can be compiled directly on the Nano. He uses the example of being able to recognize the numbers 0-9, but you could change those commands to “start” or “stop,” for example, if you would like to.

[Peter] admits that you can’t implement the type of speech recognition on an Arduino Nano that we’ve come to expect from those covert listening devices, but he mentions small, hands-free devices like a head-mounted multimeter could benefit from a single word or single phrase voice command. And maybe it could put your mind at ease knowing everything you say isn’t immediately getting beamed into the cloud and given to our AI overlords. Or maybe we’re all starting to get used to this. Whatever your position is on the current state of AI, hopefully, you’ve gained some inspiration for your next project.

Source: SPEECH RECOGNITION ON AN ARDUINO NANO?

Quick Solutions to Questions related to Speech Recognition on an Arduino Nano:

  • What hardware components are required for this project?
    The project utilizes an Arduino Nano and a MAX9814 microphone amplifier.
  • How did Peter optimize the Nano for audio processing?
    He performed low-level programming to optimize the ADC and achieved a sample rate of 9ksps.
  • Why were digital bandpass filters used instead of a Fourier transform?
    The Nano lacks sufficient processing power to compute a Fourier transform in real-time.
  • How does the algorithm analyze spoken utterances?
    It divides samples into 50 ms segments and calculates energy across five different frequency bands.
  • Where is the machine learning training process performed?
    The learning aspect and template creation occur on a custom PC program before compiling to the Nano.
  • Can the recognized commands be customized beyond numbers?
    Yes, users can change commands from numbers like 0-9 to phrases such as start or stop.
  • What is a potential benefit of using this offline method?
    It keeps voice data local, avoiding the need to send information to the cloud.
  • What is the primary use case suggested for this technology?
    It is suitable for small hands-free devices like a head-mounted multimeter requiring single word commands.

About The Author

Ibrar Ayyub

I am an experienced technical writer holding a Master's degree in computer science from BZU Multan, Pakistan University. With a background spanning various industries, particularly in home automation and engineering, I have honed my skills in crafting clear and concise content. Proficient in leveraging infographics and diagrams, I strive to simplify complex concepts for readers. My strength lies in thorough research and presenting information in a structured and logical format.

Follow Us:
LinkedinTwitter
Scroll to Top