/* ---- Google Analytics Code Below */
Showing posts with label Speech. Show all posts
Showing posts with label Speech. Show all posts

Sunday, June 25, 2023

Speech AI Spotlight

 Would expand potential capabilities for AR ... makes me broaden they potential usefulness.

Speech AI Spotlight: Visualizing Spoken Language and Sounds on AR Glasses

Jun 23, 2023

By Sirisha Rella

Audio can include a wide range of sounds, from human speech to non-speech sounds like barking dogs and sirens. When designing accessible applications for people with hearing difficulties, the application should be able to recognize sounds and understand speech.

Such technology would help deaf or hard-of-hearing individuals with visualizing speech, like human conversations and non-speech sounds. Combining speech and sound AI together, you can overlay the visualizations onto AR glasses, making it possible for users to see and interpret sounds that they wouldn’t be able to hear otherwise. 

According to the World Health Organization, about 1.5B people (nearly 20% of the global population) live with hearing loss. This number could rise to 2.5B by 2050.

Cochl, an NVIDIA partner based in San Jose, is a deep-tech startup that uses sound AI technology to understand any type of audio. They are also a member of the NVIDIA Inception Program, which helps startups build their solutions faster by providing access to cutting-edge technology and NVIDIA experts.

The platform can recognize 37 environmental sounds, and the company went one step further by adding cutting-edge speech-to-text technology. This gives a truly complete understanding of the world of sound.

AR glasses to visualize any sound

AR glasses have the potential to greatly improve the lives of people with hearing loss as an accessible tool to visualize sounds. This technology can help enhance their communication abilities and make it easier for them to navigate and participate in the world around them.  ... ' 

Monday, February 06, 2023

AI with a Modular Future

Intro to Whisper and its architecture, transcribing speech  ... 

Whispers of A.I.'s Modular Future, By The New Yorker, February 1, 2023      in CACM

What’s so unusual about Whisper is that OpenAI open-sourced it, releasing not just the code but a detailed description of its architecture.

One day in late December, I downloaded a program called Whisper.cpp onto my laptop, hoping to use it to transcribe an interview I'd done. I fed it an audio file and, every few seconds, it produced one or two lines of eerily accurate transcript, writing down exactly what had been said with a precision I'd never seen before. As the lines piled up, I could feel my computer getting hotter. This was one of the few times in recent memory that my laptop had actually computed something complicated—mostly I just use it to browse the Web, watch TV, and write. Now it was running cutting-edge A.I.

Despite being one of the more sophisticated programs ever to run on my laptop, Whisper.cpp is also one of the simplest. If you showed its source code to A.I. researchers from the early days of speech recognition, they might laugh in disbelief, or cry—it would be like revealing to a nuclear physicist that the process for achieving cold fusion can be written on a napkin. Whisper.cpp is intelligence distilled. It's rare for modern software in that it has virtually no dependencies—in other words, it works without the help of other programs. Instead, it is ten thousand lines of stand-alone code, most of which does little more than fairly complicated arithmetic. It was written in five days by Georgi Gerganov, a Bulgarian programmer who, by his own admission, knows next to nothing about speech recognition. Gerganov adapted it from a program called Whisper, released in September by OpenAI, the same organization behind ChatGPT and DALL-E. Whisper transcribes speech in more than ninety languages. In some of them, the software is capable of superhuman performance—that is, it can actually parse what somebody's saying better than a human can.

What's so unusual about Whisper is that OpenAI open-sourced it, releasing not just the code but a detailed description of its architecture. They also included the all-important "model weights": a giant file of numbers specifying the synaptic strength of every connection in the software's neural network. In so doing, OpenAI made it possible for anyone, including an amateur like Gerganov, to modify the program. Gerganov converted Whisper to C++, a widely supported programming language, to make it easier to download and run on practically any device. This sounds like a logistical detail, but it's actually the mark of a wider sea change. Until recently, world-beating A.I.s like Whisper were the exclusive province of the big tech firms that developed them. They existed behind the scenes, subtly powering search results, recommendations, chat assistants, and the like. If outsiders have been allowed to use them directly, their usage has been metered and controlled.

From The New Yorker

Wednesday, January 11, 2023

Advanced Speech Simulation

Security is slipping everywhere ...

Microsoft’s new AI can simulate anyone’s voice with 3 seconds of audio

Text-to-speech model can preserve speaker's emotional tone and acoustic environment.

In Ars Technica,   BENJ EDWARDS - 1/9/2023 M  ...  '

On Thursday, Microsoft researchers announced a new text-to-speech AI model called VALL-E that can closely simulate a person's voice when given a three-second audio sample. Once it learns a specific voice, VALL-E can synthesize audio of that person saying anything—and do it in a way that attempts to preserve the speaker's emotional tone.

Meta’s AI-powered audio codec promises 10x compression over MP3

Its creators speculate that VALL-E could be used for high-quality text-to-speech applications, speech editing where a recording of a person could be edited and changed from a text transcript (making them say something they originally didn't), and audio content creation when combined with other generative AI models like GPT-3.  ... ' 

Sunday, October 30, 2022

AI in Your Ears, Enhancing Speech with Neural 'Clearbuds'

New to me and interesting, out of ACM

AI In Your Ears  By R. Colin Johnson

Commissioned by CACM Staff, August 25, 2022

"ClearBuds" is the code-name of the first "end-to-end hardware-software neural-network based binaural system using wireless synchronized earbuds," according to hardware engineer Maruchi Kim at the University of Washington.

Kim and his colleagues demonstrated a prototype of their speech-enhancing/noise-reducing devices at the ACM International Conference on Mobile Systems, Applications, and Services (ACM MobiSys2022, held in Portland, OR June 27-July 1 ).

The "first" claimed by the researchers is the pairing of binaural (dual) microphones—one in each ear's ClearBud—with two neural networks in an app on a smartphone, resulting in a superior user-experience of voice isolation and noise cancellation during telephone conversations, according to test subjects.

"While neither dual mics nor neural network software is unique or innovative, the combination has value since it reportedly provides an experience that the users liked," said Fan Gang Zeng, a professor of otolaryngology and director of the Hearing and Speech Laboratory at the University of California, Irvine. A researcher in auditory science and technology who was not involved with the research, Zeng added, "Also, there is no technical barrier for others to develop or use the same combo."

To assist other researchers and even commercial telephony equipment providers to use the ClearBud approach, the researchers open-sourced their hardware, software, and neural network architectures. Details are provided in their paper, as well as in their audio demonstrations (which also contain links to the open-source hardware, including the printed circuit board layout, the software code for binaural transmission over Bluetooth, and the code and architectures of the neural networks).

Thursday, January 06, 2022

Closer to Conversation

 Better Conversation in Context

Model Moves Computers Closer to Understanding Human Conversation

Piotr Zelasko at the Johns Hopkins Center for Language and Speech Processing has developed a machine learning model that can differentiate speech functions in transcripts of conversations generated by language understanding (LU) systems. The model performs dialogue act recognition by identifying words' underlying intent, and assigns them to categories like "Statement," "Question," or "Interruption" in the final transcript. Zelasko sought to ensure his system could understand ordinary conversation, which may help with such tasks as summarization, intent recognition, and detection of key phrases. Zelasko said LU systems no longer need contend with "huge, unstructured chunks of text, which they struggle with when trying to classify things such as the topic, sentiment, or intent of the text. Instead, they can work with a series of expressions, which are saying very specific things."

Complete Article

Thursday, June 03, 2021

AppTech Partners with Intel for AI Enabled Speech

A new partnership emerges.  In Cision PRNewswire

AppTek Partners with Intel to Foster the Development of Next Generation AI-Enabled Speech and Language Technologies

MCLEAN, Va., June 2, 2021 /PRNewswire/ -- AppTek, a leader in Artificial Intelligence (AI), Machine Learning (ML), Automatic Speech Recognition (ASR), Neural Machine Translation (NMT), Text-to-Speech (TTS) and Natural Language Processing / Understanding (NLP/U) technologies, announced a partnership with Intel to accelerate and enhance performance benchmarks for the company's award-winning AI-enabled ASR and NMT technologies as part of the Intel AI Builders program. Intel's AI Builders is an  ... 

Tuesday, November 24, 2020

Capturing Words from Silent Speech

An Indication of how AI can become very powerful, even replacing the need for actual speech with mouthed muscle activity.  A kind of pattern recognition.

UC Berkeley researchers detect ‘silent speech’ with electrodes and AI

Khari Johnson  @kharijohnson in VentureBeat

UC Berkeley researchers say they are the first to train AI using using silently mouthed words and sensors that collect muscle activity. Silent speech is detected using electromyography (EMG), with electrodes placed on the face and throat. The model focuses on what researchers call digital voicing to predict words and generate synthetic speech.

Researchers believe their method can enable a number of applications for people who are unable to produce audible speech and could support speech detection for AI assistants or other devices that respond to voice commands.  ... " 

Wednesday, November 27, 2019

New Alexa Emotions

New complexity in voice expression is announced for Alexa.  Need to see this in useful context, but intrigued by possibilities.

Use New Alexa Emotions and Speaking Styles to Create a More Natural and Intuitive Voice Experience      By Catherine Gao

We’re excited to introduce two new Alexa capabilities that will help create a more natural and intuitive voice experience for your customers. Starting today, you can enable Alexa to respond with either a happy/excited or a disappointed/empathetic tone in the US. Emotional responses are particularly relevant to skills in the gaming and sports categories. Additionally, you can have Alexa respond in a speaking style that is more suited for a specific type of content, starting with news and music. Speaking styles are curated text-to-speech voices designed to create a more delightful customer experience for specific content. For example, the news speaking style makes Alexa’s voice sound similar to what you hear from TV news anchors and radio hosts. To learn more, check out our technical documentation for emotions here and speaking styles here.

How Alexa Emotions Work
Alexa emotions use Neural TTS (NTTS) technology, Amazon’s text-to-speech technology that enables more natural sounding speech. For example, you can have Alexa respond in a happy/excited tone when a customer answers a trivia question correctly or wins a game. Similarly, you can have Alexa respond in a disappointed/empathetic tone when a customer asks for the sports score and their favorite team has lost. Early customer feedback indicates that overall satisfaction with the voice experience increased by 30% when Alexa responded with emotions. Check out the following examples and compare them to the neutral tone:  .... ' 

Monday, September 09, 2019

Voice Applications, Why and How

Good, non device specific look at voice applications.   And an overview of what people are doing and why and where to start. And really not just about AI,  think assistance in context.

Got speech? These guidelines will help you get started building voice applications
Speech adds another level of complexity to AI applications—today’s voice applications provide a very early glimpse of what is to come.     By Ben Lorica, Yishay Carmiel  in O'Reilly Media ....