/* ---- Google Analytics Code Below */
Showing posts with label visual. Show all posts
Showing posts with label visual. Show all posts

Wednesday, February 08, 2023

Improving Ways to Search Visually

Have tried this a number of times, still learning, like the idea, but it is hard to separate different elements of form effectively.   Any good guidance out there?    Will test. 

From images to videos, how AI is helping you search visually

Feb 08, 2023

With our next generation of AI-powered technology, we’re making it more visual, natural and intuitive to explore information.

Elizabeth Reid, VP, Search

Our products at Google have a singular goal: to be as helpful to you as possible, in moments big and small. And we’ve long believed that artificial intelligence can supercharge how we deliver on that goal.

Since the early days of Search, AI has helped us with language understanding, making results more helpful. Over the years, we've deepened our investment in AI and can now understand information in its many forms — from language understanding to image understanding, video understanding and even understanding the real world.

Today, we’re sharing a few new ways we’re applying our advancements in AI to make exploring information even more natural and intuitive.

If you can see it, you can search it

Cameras have become a powerful way to explore and understand the world around you. In fact, Lens is now used more than 10 billion times per month as people search what they see using their camera or images.

With Lens, we want to connect you to the world’s information, one visual at a time. You can already use Lens to search from your camera or photos, right from the Search bar. Now, we’re introducing a major update to help you search what’s on your mobile screen.

In the coming months, you’ll be able to use Lens to “search your screen” on Android globally. With this technology, you can search what you see in photos or videos across the websites and apps you know and love, like messaging and video apps — without having to leave the app or experience.

Say your friend sends you a message with a video of them exploring Paris. If you want to learn more about the landmark you spot in the background, you can simply long-press the power or home button on your Android phone (which invokes your Google Assistant) and then tap “search screen.” Lens identifies it as Luxembourg Palace — and you can click to learn more.

Mix and match ways to search

With multisearch, you can search with a picture and text at the same time — opening up entirely new ways to express yourself. Today, multisearch is available globally on mobile, in all languages and countries where Lens is available.

We recently took multisearch even further by adding the ability to search locally. You can take a picture and add “near me” to find what you need, whether you’re looking to support neighborhood businesses or just need to find something in a hurry. This is currently available in English in the U.S., and in the coming months, we'll be expanding globally.

And sometimes, you might already be searching when you find something that catches your eye and inspires you. In the next few months, you’ll be able to use multisearch globally on any image you see on the search results page on mobile.

For example, you might be searching for “modern living room ideas'' and see a coffee table that you love, but you’d prefer it in another shape — say, a rectangle instead of a circle. You’ll be able to use multisearch to add the text “rectangle” to find the style you’re looking for.

We’re creating search experiences that are more natural and visual — but we’ve only scratched the surface. In the future, with the help of AI, the possibilities will be endless. .... 

Sunday, October 09, 2022

Detailing Neural Activity

Rare Electrical Recordings of Human Brain Detail Neural Activity

By New York University, October 5, 2022

An international team of neuroscientists used medical data to record human neural activity of visual processing in exceptional detail.

The researchers examined volunteer epilepsy patients who had been implanted with electrodes in order to measure seizure-associated neural activity. Recordings made as patients watched pictures on a laptop showed computational models that were designed to explain neural responses in non-human primates are applicable to human brains. The results indicate that the models can accurately predict changes in human neural activity for various changes in a visually presented image, according to the team's published report.

"We found that both human and animal brains seem to be using a similar 'toolkit' of neural calculations to make sense of the continuous stream of inputs arriving from our senses," says Iris Groen, an assistant professor at the University of Amsterdam, the Netherlands. ... 

A computer model can predict rapid fluctuations in neural activity in the human visual cortex. ... 

From New York University

View Full Article     

Sunday, December 29, 2019

Assembling, Using, Visual Knowledge

We also looked at ways to assemble visual information, for use in training, communication, archiving, data.    And at patterns in that data that wold make it more useful.   One thing that came up was how to augment it to make it most useful to the largest group.  A recognition of it as an asset.  And the need for related metadata.

NBA Teams Enhancing Fan Experience with High-Tech Replays
ABC News
Charles Odum
November 15, 2019

Six National Basketball Association (NBA) teams are implementing 360-degree video replays in their arenas to augment the fan experience. The technology allows fans to review shots and gameplay by changing the angle, similar to video-game players' use of all-angle replays. The teams partnered with technology provider Intel to install 38 5K video cameras in their arenas, which work in concert to bring the replays to in-game video boards, TV broadcasts, and fans' devices via social media. The replays not only enhance the fan experience, but also can help coaches and scouts refine player assessments. Joe Abercrombie with the NBA's Atlanta Hawks called the technology “the wave of the future,” adding that it is “one more thing to give people a reason to come” to watch games at the arenas.

See:  Northeastern University Institute for Experiential AI  ...  

Wednesday, November 13, 2019

AI Agents for Design

Intrigued by integration of Design and AI.  Our early looks were at how to use a vast history of advertising to create new and effective marketing designs.

AI agents imitate engineers to construct effective new designs using visual cues like humans do  by Carnegie Mellon University

Trained AI agents can adopt human design strategies to solve problems, according to findings published in the ASME Journal of Mechanical Design.

Big design problems require creative and exploratory decision making, a skill in which humans excel. When engineers use artificial intelligence (AI), they have traditionally applied it to a problem within a defined set of rules rather than having it generally follow human strategies to create something new. This novel research considers an AI framework that learns human design strategies through observation of human data to generate new designs without explicit goal information, bias, or guidance.

The study was co-authored by Jonathan Cagan, professor of mechanical engineering and interim dean of Carnegie Mellon University's College of Engineering, Ayush Raina, a Ph.D. candidate in mechanical engineering at Carnegie Mellon, and Chris McComb, an assistant professor of engineering design at the Pennsylvania State University.

"The AI is not just mimicking or regurgitating solutions that already exist," said Cagan. "It's learning how people solve a specific type of problem and creating new design solutions from scratch." How good can AI be? "The answer is quite good." ... " 

Saturday, June 15, 2019

Learning from Camera Feeds

Technically interesting.  Learning by seeing, like we do.

Model-Based Reinforcement Learning from Pixels with Structured Latent Variable Models

By Marvin Zhang and Sharad Vikram  in Bair
  
Imagine a robot trying to learn how to stack blocks and push objects using visual inputs from a camera feed. In order to minimize cost and safety concerns, we want our robot to learn these skills with minimal interaction time, but efficient learning from complex sensory inputs such as images is difficult. This work introduces SOLAR, a new model-based reinforcement learning (RL) method that can learn skills – including manipulation tasks on a real Sawyer robot arm – directly from visual inputs with under an hour of interaction. To our knowledge, SOLAR is the most efficient RL method for solving real world image-based robotics tasks.  ... "

Saturday, March 16, 2019

Ask Developer Console for Interaction Insights

Just got a chance to look at this more closely.

Gain Interaction Insights Using New Analytics in the ASK Developer Console  By BJ Haberkorn

Today we added interaction path analysis to the Analytics tab on the Alexa Skills Kit (ASK) Developer Console. Interaction path analysis shows aggregate skill usage patterns in a visual format, including which intents your customers use, in what order. This enables you to verify if customers are using the skill as expected, and to identify interactions where customers become blocked or commonly exit the skill. You can use insights gained from interaction path analysis to make your flow more natural, fix errors, and address unmet customer needs.

View Interaction Paths over Multiple Time Intervals
As shown in the example below, interaction path analysis provides a visual representation of the flow of users from the invocation of your skill to subsequent intents. In this example, most customers moved from LaunchRequest to Intent1. A smaller segment invoked Intent2 instead. Interaction path analysis shows both custom intents and built-in intents, such as   ... "

Saturday, April 14, 2018

Modeling the Visual Intelligence of Dogs

Quite a remarkable claim.   Can many visuals replace the ability to query a system about about what it knows.  Could this work with humans, entities like firms and their data?  Examining.

Via the University of Washington and Allen Institute for AI.  Who used neural networks to understand the behavior of dogs.   And suggest that animals provide data to train AI systems, including robotics.  What other visual agents?  And how accurate is the model?  Examining.

Who Let The Dogs Out? Modeling Dog Behavior From Visual   in arXIV
Kiana Ehsani, Hessam Bagherinezhad, Joseph Redmon, Roozbeh Mottaghi, Ali Farhadi

We introduce the task of directly modeling a visually intelligent agent. Computer vision typically focuses on solving various subtasks related to visual intelligence. We depart from this standard approach to computer vision; instead we directly model a visually intelligent agent. Our model takes visual information as input and directly predicts the actions of the agent. Toward this end we introduce DECADE, a large-scale dataset of ego-centric videos from a dog's perspective as well as her corresponding movements. Using this data we model how the dog acts and how the dog plans her movements. We show under a variety of metrics that given just visual input we can successfully model this intelligent agent in many situations. Moreover, the representation learned by our model encodes distinct information compared to representations trained on image classification, and our learned representation can generalize to other domains. In particular, we show strong results on the task of walkable surface estimation by using this dog modeling task as representation learning. ..."

Sunday, January 07, 2018

Advances in Sensory Substitution

Been reading about this for many years,  are real advances here?   Article looks at the history, technology and likely advances.

Feeling Sounds, Hearing Sights   By Gregory Mone 

Communications of the ACM, Vol. 61 No. 1, Pages 15-17

 In a 2016 video, Saqib Shaikh, a Microsoft Research software engineer, walks out of London's Clapham Station Underground stop, turns, and crosses a street, then stops suddenly when he hears an unexpected noise. Shaikh, who lost his sight when he was seven years old and walks with the aid of the standard white cane, reaches up and swipes the earpiece of his glasses.

The video then shifts to the view from his eyewear, a pair of smart glasses that capture high-quality still images and videos. That simple swipe instructed the glasses, an experimental prototype designed by a company called Pivothead, to snap a still photo. Microsoft software analyzed the picture, then translated the findings into auditory feedback. Through the smart glasses, which include a small speaker, Shaikh hears the results from an automated voice: "I think it's a man jumping in the air doing a trick on a skateboard."

The Pivothead smart glasses and Microsoft AI technology belong to a broader class of what have become known as sensory substitution technologies, apps and devices that collect visual, auditory, and in some cases haptic stimuli, and feed the information to the user through another sensory channel. While the utility of these devices has long been debated in the vision- and hearing-impaired communities, recent advances suggest that sensory substitution technologies are finally starting to deliver on their promise.  .... "

Wednesday, July 20, 2016

Fabric Software for Visual Programming

New way to visually program.  I built software for years, so understand the embedded power,  but always believed there had to be a better way.  Can visual methods be precise enough to construct algorithms, and reduce errors and document its own process?   Visuals shown are impressive: 

" ...  Fabric have just announced that it will be introducing a new visual programming tool which the company believes will be make creation both easier and more dynamic. SIGGRAPH 2016 will mark the debut of the software.

An addition to the existing Canvas visual programming suite, Blocks which will ship with Fabric Engine 2.3 will expand on the existing functionality and Blocks allows users to create solutions they’d previously only been able to envision through writing code.

“Blocks are Canvas graph containers for user-provided functionality within a preset – the “block” inside a “for” loop is literally a Canvas graph. With Blocks, technical directors can create complex presets (either graphically or with code) exposing only the controls that make sense for non-technical users to modify. Users can then modify the preset by filling in the content of the blocks – visually — without needing to know all the details of the preset.” Explains the company in their announcement.... " 

Monday, January 20, 2014

NantMobile and Macy's

A former colleague of mine now runs NantMobile.  In the press it is reported that they are working with Macy's    Its interesting too that the company is owned by major Chinese retailer BJ Hualian.   And more on the Macy's work. ( NantMobile, A NantWorks LLC company, The Beijing Hualian Group )

" ... NantMobile’s core product, iD Browser, is a mobile recognition platform that allows people to browse the physical world around them, unlocking digital experiences, coupons, content and information from featured brands that they know, like and trust.

Today consumers are limited to browsing the web. Why should we limit them? Imagine if consumers could browse the world, any time, any place, unlocking the physical world with digital content related to your brand. Now the physical world is browsable.

With your mobile device, the iD Browser, and our patented recognition technology, we are transforming the world wide web into the web enabled world. Get your branded presence up and running on the iD Browser today.NantMobile’s core product, iD Browser, is a mobile recognition platform that allows people to browse the physical world around them, unlocking digital experiences, coupons, content and information from featured brands that they know, like and trust. ...  " 

Thursday, June 17, 2010

Visual Experience Management

From Saffron. Intriguing linkage of associative memory and visualization. Will examine it more closely.

"‘Visual Experience Management’ Makes Big Data “Pop” With Meaning
McLean, VA & Cary, NC, May 19, 2010 –- Centrifuge Systems, Inc., a leading provider of next-generation Business Intelligence (BI) software, and Saffron Technology, Inc., a data analytics software firm providing associative–memory-based Experience Management solutions, today jointly announced they are partnering to deliver a graphical, highly interactive streaming data analytics and Business Intelligence solution for enterprises worldwide. Both Saffron and Centrifuge were recently recognized as “Cool Vendors” by a leading analyst firm.

The integrated solution combines Saffron Natural Intelligence Platform Version 8.0 and Centrifuge 2.0. The offering gives customers in business and government the ability to apply associative memory to analytics and business intelligence — both via Saffron’s sense-making and decision support capabilities, and Centrifuge’s powerful visualization solutions ... "