/* ---- Google Analytics Code Below */
Showing posts with label Captioning. Show all posts
Showing posts with label Captioning. Show all posts

Monday, May 23, 2022

Advances in Image Captioning

 Another example of an area we worked in, Webinar, details at the link:

AI advances in image captioning: Describing images as well as people do

Image captioning is an interesting problem in the intersection between computer vision and natural language processing, and it has attracted great attention from their respective research communities. Recent image captioning models have achieved impressive results on the tasks where large amounts of paired image-caption training data is available. However, they generalize poorly to images in the wild, where there are a wide variety of visual objects that are unseen in the caption corpora for training. This raises the challenge of Novel Object Captioning (NOC), that is, generating captions to describe novel objects unseen in paired image-caption training data, which is especially pertinent in real-world applications.

This webinar will focus on some of the recent vision-language pretraining (VLP) approaches for image captioning. We will cover our latest approaches, including object-semantics aligned pretraining (OSCAR) and visual-vocabulary pretraining (VIVO). We will also discuss their key principles and how we address the core challenges in image caption generation. Join us to learn how our discovery leads to a new image captioning framework that achieves state-of-the-art performance on the nocaps benchmark (developed to evaluate NOC at scale) and surpasses human CIDEr scores on nocaps for the first time.

Visual-vocabulary pretraining (VIVO) conducts pretraining with vision data only. As the method does not need paired image-caption data, it opens the possibility of leveraging large amounts of images, paired with either human-labeled or machine-generated tags. By using VIVO pretraining, the performance of the captioning model, especially on novel objects, has been substantially improved..... ' 

Thursday, March 18, 2021

Chrome Announces Captioning for Audio and Video

Conversion from audio.  Considerable thing. 

Chrome can now caption audio and video  in the Google Blog

Captions make online content more accessible. If you’re in a noisy environment, trying to keep the volume down, or are part of the 466 million people in the world who are deaf or hard of hearing, having captions lets you follow along to whatever content you are watching — whether it’s viral feta pasta videos, breaking news or a scientist discussing their latest research.

Unfortunately, captions aren’t always available for every piece of content. Now with Live Caption on Chrome, you can automatically generate real-time captions for media with audio on your browser. It works across social and video sites, podcasts and radio content, personal video libraries (such as Google Photos), embedded video players, and most web-based video or audio chat services.  ... "

Sunday, December 20, 2020

Amazon Group Calling and Captioning

 Had just heard of this, ways to make group calls more efficient is very useful.  About to try.

Amazon’s Echo devices gain group-calling and call captioning features  By Kelly Hodgkins in DigitalTrends

Can’t get together with your family this year for Christmas? You are not alone. Instead of gathering for a large holiday celebration, many people are opting to meet virtually. To help people stay connected while apart, Amazon added new audio and video calling features to Alexa and its Echo devices starting today.

With Amazon’s group-calling feature, people can connect to seven friends and family members on a single hands-free video or audio call. All it takes is a single person to create and name a group using Alexa. Once the group is formed, other friends and family members can join by telling Alexa to “Call my (Group Name).” This group-calling feature is available on compatible Echo devices, including the Echo, Echo Dot, and Echo Show.  

Amazon also added a new closed captioning feature that is available in both audio and video calls. During the call, Amazon will convert the other person’s voice to text in near real time. Captioning works with U.S. English and is available on all Echo Show devices.  ... ' 

Wednesday, October 14, 2020

Microsoft says its AI can describe images 'as well as people do'

 Recall some portion of this claim being made, but had not heard much new from Microsoft,  Know sight impaired people who could use the capability.     We also worked on the general idea of  'captioning', which turns out to be tough to do generally well.  

Microsoft says its AI can Describe Images 'as Well as People Do'  By Devindra Hardawar, @devindra  in Engdget 

It’s a new milestone for AI that could genuinely help the visually impaired. 

Describing an image accurately, and not just like a clueless robot, has long been the goal of AI. In 2016, Google said its artificial intelligence could caption images almost as well as humans, with 94 percent accuracy. Now Microsoft says it’s gone even further: Its researchers have built an AI system that’s even more accurate than humans — so much so that it now sits at the top of the leaderboard for the nocaps image captioning benchmark. Microsoft claims its two times better than the image captioning model it’s been using since 2015. 

And while that’s a notable milestone on its own, Microsoft isn’t just keeping this tech to itself. It’s now offering the new captioning model as part of Azure's Cognitive Services, so any developer can bring it into their apps. It’s also available today in Seeing AI, Microsoft's app for blind and visually impaired users that can narrative the world around them. And later this year, the captioning model will also improve your presentations in PowerPoint for the web, Windows and Mac. It’ll also pop up in Word and Outlook on desktop platforms.

Saturday, August 24, 2019

More on Imminent Audible Captioning

Recently have been reading/hearing long texts in Audible and Kindle,  usually non-fiction.   So gaining an appreciation for the difference in the way we utilize recorded knowledge in varying contexts.  Having the option of having it read to us is good, but seeing it in context provides better retention and integration.     And also has implications for accessibility.   Which leads to this new article in ArsTechnica, considerable additional discussion there:

Seven of the nation's top book publishers sued Amazon subsidiary Audible on Friday, asking federal courts to block the company from releasing a new feature called Audible Captions that's due out next month. The technology does exactly what it sounds like: display text captions on the screen of your phone or tablet as the corresponding words are read in the audio file.

The publishers argue that this is straight-up copyright infringement. In their view, the law gives them the right to control the distribution of their books in different formats. Audio is a different format from text, they reason, so Audible needs a separate license.  ......

The caption feature "is not and was never intended to be a book," Audible explained in an online statement following the lawsuit. "Listeners cannot read at their own pace or flip through pages as they could with a print book or eBook." Instead, the purpose is to allow "listeners to follow along with a few lines of machine-generated text as they listen to the audio performance."

"We disagree with the claims that this violates any rights and look forward to working with publishers and members of the professional creative community to help them better understand the educational and accessibility benefits of this innovation," Audible added.  ... "

Friday, July 26, 2019

Captioning Audible Books

The idea below is a kind of captioning of audio from a a book being read to you.    Gives you two streams, audio and visual, for a given book.    I would find this useful, sometimes one or the other works better, also seems its useful for accessibility as well.  Sometimes I want to 'reread' a section of text I may not have understood audibly, and that's usually easier to visually scan than hear.  Also good for things that are best shown rather than described, like pictures or charts or equations.   Had thought a number of times that captions would be useful.   A recent book on Leonardo kept pointing me to a pdf for painting illustrations, those could have been made accessible in captions.

But it seems some publishers believe its giving too much away:

Publishers are pissed about Amazon’s upcoming Audible Captions feature
Some are asking for their books to be withheld from the feature
By Andrew Liptak  @AndrewLiptak  in TheVerge ... 

Tuesday, September 18, 2018

Efficiency for Machine Learning Automation

 And more elements of AI automation.  Here from MIT.   The details of the data construction for this is also described,  which is always enlightening.

Machine-learning system tackles speech and object recognition, all at once

Model learns to pick out objects within an image, using spoken descriptions.   By Rob Matheson | MIT News Office

MIT computer scientists have developed a system that learns to identify objects within an image, based on a spoken description of the image. Given an image and an audio caption, the model will highlight in real-time the relevant regions of the image being described.

Unlike current speech-recognition technologies, the model doesn’t require manual transcriptions and annotations of the examples it’s trained on. Instead, it learns words directly from recorded speech clips and objects in raw images, and associates them with one another.

The model can currently recognize only several hundred different words and object types. But the researchers hope that one day their combined speech-object recognition technique could save countless hours of manual labor and open new doors in speech and image recognition.

Speech-recognition systems such as Siri and Google Voice, for instance, require transcriptions of many thousands of hours of speech recordings. Using these data, the systems learn to map speech signals with specific words. Such an approach becomes especially problematic when, say, new terms enter our lexicon, and the systems must be retrained.

“We wanted to do speech recognition in a way that’s more natural, leveraging additional signals and information that humans have the benefit of using, but that machine learning algorithms don’t typically have access to. We got the idea of training a model in a manner similar to walking a child through the world and narrating what you’re seeing,” says David Harwath, a researcher in the Computer Science and Artificial Intelligence Laboratory (CSAIL) and the Spoken Language Systems Group. Harwath co-authored a paper describing the model that was presented at the recent European Conference on Computer Vision.

In the paper, the researchers demonstrate their model on an image of a young girl with blonde hair and blue eyes, wearing a blue dress, with a white lighthouse with a red roof in the background. The model learned to associate which pixels in the image corresponded with the words “girl,” “blonde hair,” “blue eyes,” “blue dress,” “white light house,” and “red roof.” When an audio caption was narrated, the model then highlighted each of those objects in the image as they were described. .... "

Related article.

Wednesday, August 22, 2018

Google Explores Podcasts

In the past I ran an internal podcast in a large company. Have always been interested in podcasts as an information channel, but also one that that had issues, for example it is a serial channel, and takes time to ingest that way.  Just recently have started to explore Podcasts again. Especially through voice interfaces.   Below work Google is doing in the space.  Note 'discovering' and efficient playing.    How can AI be used to analyze some of the content and its interaction with other online knowledge?  Captioning a good example.  Both challenges in the channel when using Podcasts.

Google is developing an experimental podcast app called Shortwave    By Russell Brandom  @russellbrandom in TheVerge

An experimental unit within Google has been quietly developing a new app for discovering and playing podcasts. Called Shortwave, the new app was revealed by a trademark filing embedded below, which describes it as “allow[ing] users to search, access, and play digital audio files, and to share links to audio files.”

Nothing in the trademark filing specifies the kind of audio being accessed, but a Google representative said the focus of the app was on spoken word content. There is little public information about the app, although Google has played with smart captioning, translation, and other AI-assisted features in previous podcast products. ... " 

Thursday, November 16, 2017

Caption Generation Tutorial

 Jason Brownlee on caption generation using neural networks, an old problem we addressed.  I like the explanations Jason does.  Get on his instructive solutions mailings.

" ...  Hi, this week we have a gentle intro to neural nets for automatically generating captions for photos ...   Neural network models can automatically describe the contents of a photograph. Discover more in this gentle introduction:

How to Generate Textual Descriptions for Photographs with Deep Learning

Image captioning requires careful preparation of the photograph and text data. Discover how to prepare data for this problem in this tutorial:

How to Prepare a Photo Caption Dataset for Training a Deep Learning Model   .... "

Wednesday, June 14, 2017

Extracting Tagging Data from Imagery

An example of how advanced deep learning methods can be used to extract information from image data.   The image data is already captured and stored.  And its dynamic.   Could also be used with other imagery, like from store shelf images.   Note the common existence of noise in such recordings.  Also the integrated normalization included in the tagging.  Thinking other possibilities.  

Enhancing Google Maps with Deep Learning and Street View  by Srini Penchikala
" ... The deep learning model also automatically labels new Street View imagery, normalizes the text to be consistent with the naming conventions and ignores extraneous text that's not relevant for the data analytics. This allows the team to create new addresses directly from images without even knowing the name of the street or the location of the addresses. For example, when a Street View car drives on a newly built road, the model can analyze the captured images, extract the street names and numbers, and properly create and locate the new addresses automatically on Google Maps. .... " 

Monday, April 10, 2017

Captioning Images with Tensorflow

 Below, Nicely done and instructive piece on captioning using neural methods.   Captioning as a means of multi tagging of images.  Now this we could have used when captioning copy images and documents for retrieval, analysis and re-use.   How accurate in the given context?  How much editing might be required?   In O'Reilly:

Caption this, with TensorFlow
How to build and train an image caption generator using a TensorFlow notebook.  By Raul Puri, Daniel Ricciardelli  March 28, 2017

Image caption generation models combine recent advances in computer vision and machine translation to produce realistic image captions using neural networks. Neural image caption models are trained to maximize the likelihood of producing a caption given an input image, and can be used to generate novel image descriptions. For example, the following are possible captions generated using a neural image caption generator trained on the MS COCO data set.     ... " 

Tuesday, April 04, 2017

Microsoft Captions in a Camera App

I was involved in a project, and later a startup that looked at the process of 'captioning', or selectively describing what is on an image.  A difficult problem.  Microsoft is providing a camera app to this that would be interesting to test. I previously commented on their work in captioning, see the tag link, which also points to related methods in Open Source TensorFlow.

Microsoft launches Sprinkles, a silly camera app powered by machine learning   by Sarah Perez

Microsoft is getting into the “fun camera app” game with a new iOS application called Sprinkles, which has now earned a featured spot in the “New apps we love” section of the App Store. The gist with Sprinkles, clearly aimed at a teen audience, is to offer a variety of traditional photo decorating tools like stickers, emoji and captions, but leverages Microsoft’s machine learning and A.I. capabilities to do things like detect faces, determine the photo subject’s age and emotion, figure out your celebrity look-a-like, suggest captions, and more. ... " 

Thursday, September 22, 2016

Image Captioning by Tensorflow is Now Open Source

I have mentioned before this is a problem we addressed for several AI oriented applications.  We called it 'image recognition'. Now the general solution is open source.   Some samples images in the article, and they are impressive. The general solution of this captioning problem is an important one.

Show and Tell: image captioning open sourced in TensorFlow
Thursday, September 22, 2016
Posted by Chris Shallue, Software Engineer, Google Brain Team

In 2014, research scientists on the Google Brain team trained a machine learning system to automatically produce captions that accurately describe images. Further development of that system led to its success in the Microsoft COCO 2015 image captioning challenge, a competition to compare the best algorithms for computing accurate image captions, where it tied for first place.

Today, we’re making the latest version of our image captioning system available as an open source model in TensorFlow. This release contains significant improvements to the computer vision component of the captioning system, is much faster to train, and produces more detailed and accurate descriptions compared to the original system. These improvements are outlined and analyzed in the paper Show and Tell: Lessons learned from the 2015 MSCOCO Image Captioning Challenge, published in IEEE Transactions on Pattern Analysis and Machine Intelligence.  .... " 

Wednesday, August 03, 2016

Shoe Recognition by Disney

How much better is this than face recognition?    Positioned as less personal I assume.  Link to demographics and style.

Disney may track park visitors using shoe recognition technology
Walt Disney was recently awarded a patent that describes identifying and tracking its visitors based on their footwear in order to provide customers with customized experiences.

As a guest enters a theme park, cameras would capture their shoe’s color scheme as well as brand and model. Further, sensors placed in the ground would recognize and store shoe size and tread pattern. Individuals could then be recognized by other cameras and sensors as they move about the park.

According to the patent, amusements parks, theme parks other sporting and entertainment venues would benefit from a better understanding of guests’ footpaths. For theme parks this would include understanding the “common guest paths from ride to ride.”   .... ' 

Tuesday, July 19, 2016

Nest Camera with AI

In CWorld:  Nest camera detects people in its images.  Continue to see more examples of the use of AI pattern matching methods in consumer systems, like face recognition for system access.

Wednesday, July 06, 2016

Machine Learning Images from Mobile Video

Continued interest in fast image analysis.  Has implications for the perception of images, and thus space and what exists within it.

Google acquires image recognition startup Moodstocks to bolster machine learning efforts by Maria Deutscher, 

   "... Google acquires image recognition startup Moodstocks to bolster machine learning efforts  by Maria Deutscher   " ... The Paris-based startup has developed an image recognition tool specifically optimized to process media captured via mobile devices. ... " 

Wednesday, June 01, 2016

Urban Intelligence

A system called Placemeter takes as input video feeds and analyzes images to determine what is going on outside your doors.  Possible retail measures?    Video content analysis (VCA).  

" .... In its first iteration, the service, which relies on real-time video feeds, was able to quantify the overall number of objects that it saw and distinguish between pedestrians and vehicles.

Now, the service is getting significantly smarter. By default, Placemeter can now distinguish between five different objects: people, bicycles, motorcycles, cars and large vehicles (think trucks, delivery vans, etc.). ... "        In TechCrunch.

Thursday, April 14, 2016

Bots Auto Captioning Photos

An area I worked in both in the enterprise, and with a startup.   Google and Facebook have done work in this area.  Now Microsoft is also doing such captioning, apparently with some problems.  Bad problems.   The implications of error have come of much in recent bot / AI examples.   It is interesting that this is being characterized as a bot, we never used that word.  More of an image recognition and tagging problem.  We need a better definition of the bot concept.  To me it has an element of interaction, rather than just processing.

Thursday, March 03, 2016

Aesthetics and Deep Learning

Understanding aesthetics with Deep Learning. A test of using deep learning neural nets to determine the aesthetics of photographic images.

Wednesday, February 24, 2016

Push on Wearables in the Enterprise

Push for wearables in the enterprise.    Mostly watches it seems, which are not much different than smartphone devices.    Devices that engage more deliberately with users, usually for specialty applications,  are also starting to emerge.  We examined data gathering in maintenance and shelf compliance and inventory.  Real time image analysis will provide better leverage.  AI/Cognitive will allow stronger engagement.  The article is an overview.   In CWorld