/* ---- Google Analytics Code Below */
Showing posts with label Speech to Text. Show all posts
Showing posts with label Speech to Text. Show all posts

Saturday, March 25, 2023

Capturing What is Said

The value of good speech to text. 

Capturing What is Said,  By Esther Shein

Commissioned by CACM Staff, March 23, 2023

A very basic flow chart for the conversion of speech to text.

New AI-enabled capabilities for speech-to-text systems include taking actions based on a transcript, prompting someone to ask a follow-up question, and summarizing a conversation at the end of a call, said Christine McAllister at Forrester Research.

ChatGPT and generative artificial intelligence (AI) may be having a moment, but don't underestimate the value of speech-to-text transcription, sometimes referred to as automatic speech recognition (ASR) software, which continues to improve.

ASR technology converts human speech into text using machine learning and AI. There are two types: synchronous transcription, which is typically used in chatbots, and asynchronous, where transcription occurs after the fact to capture customer/agent conversations, notes Cobus Greyling, chief evangelist at HumanFirst, which makes a productivity suite for natural language data.

ASR made some waves in recent months with the announcement of Whisper from OpenAI, the organization that created ChatGPT. Whisper was trained on 680,000 hours of multilingual and supervised data collected from the Web. OpenAI claims that large and diverse dataset has improved the accuracy of the text it produces; the company says Whisper also can transcribe text from speech in multiple languages.

"What that means is that it's extremely accurate—right off the top—without much tuning or training,'' says Christina McAllister, a senior analyst at research and advisory company Forrester Research. "The large language model aspect, which is based on huge amounts of data, is what's new and is the most innovative aspect of the ASR market today,'' she says.

Because of its ability to transcribe meetings and interviews more efficiently and accurately, one of the broadest enterprise use cases for speech-to-text is in customer call centers. The next phase in the development of ASR is to use artificial intelligence to analyze call center conversations for customer sentiment and to validate compliance in regulated industries, according to Annette Jump, a vice president analyst at Gartner.

The benefits of ASR in the call center context are its ability to identify customer problems early and to improve customer satisfaction by resolving issues sooner, says Jump.

Other use cases include generating closed captions for movies, television, video games, and other forms of media. ASR is widely used in healthcare by physicians to convert dictated clinical notes into electronic medical records.

Speech vendors typically leverage a third-party ASR engine so they don't have to build their own, McAllister says. That frees them up so they can "do all the rest of their magic from the transcript point forward,'' she says.

Some of the new AI capabilities for speech-to-text systems include taking actions based on a transcript, prompting someone when it's appropriate to ask a follow-up question, and summarizing a conversation at the end of a call, McAllister says.

One frequently used AI-powered speech-to-text transcription service is Otter.ai, which has added capabilities aimed at improving meetings, including integration with collaboration tools such as Zoom and Microsoft Outlook.  ... ' 


Thursday, February 21, 2019

Google Makes more Speech Services Available

Impressive array of cognitive speech services, in 120 languages!   Now broadly available with demonstrations at the link.

Cloud Speech-to-Text

Speech-to-text conversion powered by machine learning and available for short-form or long-form audio.

Powerful speech recognition

Google Cloud Speech-to-Text enables developers to convert audio to text by applying powerful neural network models in an easy-to-use API. The API recognizes 120 languages and variants to support your global user base. You can enable voice command-and-control, transcribe audio from call centers, and more. It can process real-time streaming or prerecorded audio, using Google’s machine learning technology.

Some of the Betas in particular are indicative of future direction of capabilities:

Cloud Speech-to-Text features
Speech-to-text conversion powered by machine learning.
Automatic Speech Recognition
Automatic Speech Recognition (ASR) powered by deep learning neural networking to power your applications like voice search or speech transcription.
Global Vocabulary
Recognizes 120 languages and variants with an extensive vocabulary.
Phrase Hints
Speech recognition can be customized to a specific context by providing a set of words and phrases that are likely to be spoken. This is especially useful for adding custom words and names to the vocabulary and in voice-control use cases.
Real-time Streaming or Prerecorded Audio Support
Audio input can be streamed from an application’s microphone or sent from a prerecorded audio file (inline or through Google Cloud Storage). Multiple audio encodings are supported, including FLAC, AMR, PCMU, and Linear-16.
Auto-Detect Language BETA
When you need to support multilingual scenarios, you can now specify two to four language codes and Cloud Speech-to-Text will identify the correct language spoken and provide the transcript.
Noise Robustness
Handles noisy audio from many environments without requiring additional noise cancellation.
Inappropriate Content Filtering
Filter inappropriate content in text results for some languages.
Automatic Punctuation BETA
Accurately punctuates transcriptions (e.g., commas, question marks, and periods) with machine learning.
Model Selection BETA
Choose from a selection of four pre-built models: default, voice commands and search, phone calls, and video transcription.
Speaker Diarization BETA
Know who said what - you can now get automatic predictions about which of the speakers in a conversation spoke each utterance.
Multichannel Recognition BETA
In multiparticipant recordings where each participant is recorded in a separate channel (e.g., phone call with two channels or video conference with four channels), Cloud Speech-to-Text will recognize each channel separately and then annotate the transcripts so that they follow the same order as in real life.  .... " 

Wednesday, November 30, 2016

Amazon Delivers Serious AI Services

   More tools being added to the box.  Previously some from IBM and Google.  Can remarkable AI be far behind?  In Computerworld.

AWS comes out swinging with A.I. services
Trying to entice enterprises and catch up with cloud rivals, AWS releases three artificial intelligence services   By Sharon Gaudin  

 .... At its re:Invent conference today, Amazon Web Services CEO Andy Jassy announced three artificial intelligence (A.I.) services that will be available to enterprise users this year, with more expected in 2017.

The three services being rolled out are Amazon Rekognition for image recognition; Amazon Polly for text-to-speech services; and Amazon Lex, the technology inside its smart device Alexa, offering speech recognition services.

Now enterprise AWS customers can use A.I. to, for example, search for images that show a mountain next to a lake, or a city with specific architecture. ... " 

via Walter Riker. 

Tuesday, October 18, 2016

Microsoft Announces Speech Recognition Breakthrough

Continued advances in Cognitive interaction.  While speech recognition is good today,  making it better will remove one barrier to interacting nimbly with spoken language.  'Understanding' semantically, with common sense implied,  is still not easy.

In SiliconAngle: 
In an announcement that could be the death knell for stenographers everywhere, Microsoft Corp says that it has made a major breakthrough in the field of speech recognition. According to the company, it has developed an artificial intelligence that is capable of understanding conversational speech at a rate comparable to real humans.  

 In a research paper published through Cornell University on Monday, Microsoft’s Artificial Intelligence and Research team outlined the test they used to determine the effectiveness of their new system. According to the researchers, they tested their AI with the NIST 2000, an evaluation created by the National Institute of Standards and Technology that is specifically designed to determine the accuracy of speech recognition software. ... " 

Thursday, February 25, 2016

Making AI more Human

This direction makes sense, have not looked at the details yet.    In TechRepublic:  Towards more human Watson APIs .  That's good    "  .... IBM Watson, Big Blue's cognitive computing division, announced Tuesday that it was expanding the availability of Watson APIs with three new tools: Tone Analyzer, Emotion Analysis, and Visual Recognition. IBM Watson is also updating their Text to Speech (TTS) feature and rebranding it as Expressive TTS .... " 

Friday, February 06, 2015

Watson Adds Services to Developer Cloud

In SiliconAngle:  Notable to me is the speech to text capability.  But the added analytics are also good directons.  I am in the midst of working with startups to evaluate Watson services.  So far these capabilities look very good, but are incomplete.

" ... The eight services available now becomes thirteen, thanks to the addition of five new services including Concept Insights, Speech-to-Text Translation, Text-to-Speech Translation, Tradeoff Analytics and Visual Recognition. IBM says each of the new services can be embedded into desktop or mobile apps via APIs, offering advanced functionality that developers would find too laborious and time-consuming to build by themselves. ... " 

Watson updated Cloud Services store.