/* ---- Google Analytics Code Below */
Showing posts with label vision. Show all posts
Showing posts with label vision. Show all posts

Thursday, July 20, 2023

Computer Vision That Works More Like a Brain Sees More Like People Do

Towards better vision. 

ACM TECHNEWS

Computer Vision That Works More Like a Brain Sees More Like People Do

By MIT News, July 13, 2023

When deep Learning computer vision systems establish efficient ways to solve visual problems, they end up with artificial circuits that work similarly to the neural circuits that process visual information in our own brains.

Researchers made a computer vision model more robust by training it to work like a part of the brain that humans and other primates rely on for object recognition.

Credit: iStock

James DiCarlo and colleagues at the Massachusetts Institute of Technology trained an artificial neural network to function more like the human and primate brain's inferior temporal (IT) cortex to improve computer vision.

The researchers constructed a computer vision model based on neural data from primate vision-processing neurons, and tasked it to recognize objects.

DiCarlo said this made the artificial neural circuits process visual information differently.

The researchers found the biologically informed model IT layer aligned better with the IT neural data than a similarly-sized network model that lacked neural-data training.

They also discovered the neurally aligned model was more resilient against adversarial attacks for assessing computer vision and artificial intelligence systems.

From MIT News

View Full Article      

Sunday, May 07, 2023

Meta Announcement

Meta-AI

COMPUTER VISION

DINOv2: State-of-the-art computer vision models with self-supervised learning

Today, we are open-sourcing DINOv2, the first method for training computer vision models that uses self-supervised learning to achieve results that match or surpass the standard approach used in the field.

Meta AI has built DINOv2, a new method for training high-performance computer vision models.

DINOv2 delivers strong performance and does not require fine-tuning. This makes it suitable for use as a backbone for many different computer vision tasks.

Because it uses self-supervision, DINOv2 can learn from any collection of images. It can also learn features, such as depth estimation, that the current standard approach cannot.

We are open-sourcing our model and sharing an interactive demo.

Today, we are open-sourcing DINOv2, the first method for training computer vision models that uses self-supervised learning to achieve results that match or surpass the standard approach used in the field...'

Wednesday, March 22, 2023

Humans get X-Ray Vision in Augmented Reality

 Seen this in previous worlds we experimented with.   Concept interesting. 

MIT researchers invented an augmented reality headset that gives humans X-ray vision. The invention, dubbed X-AR, combines wireless sensing with computer vision to enable users to see hidden items. X-AR can help users find missing items and guide them toward these items for retrieval. This new technology has many applications in retail, warehousing, manufacturing, smart homes, and more.

For more information, check out:

Website: https://Xar.media.mit.edu

Paper: https://www.mit.edu/~fadel/papers/XAR-paper.pdf

Instagram: @mit_sk_lab (https://instagram.com/mit_sk_lab?igshid=YmMyMTA2M2Y=)

Authors: Tara Boroushaki, Maisy Lam, Laura Dodds, Aline Eid, Fadel Adib

Video Production: Maisy Lam, Jimmy Day

UI design: Maisy Lam, Yuechen Wang

Funding: NSF, Sloan Foundation, MIT Media Lab  .... 

Saturday, May 14, 2022

New form of Machine Learning Vision

Machine Learning Vision

MIT Advances Unsupervised Computer Vision with ‘STEGO’

By Oliver Peckham

Training machine learning models often means working with labeled data. For computer vision tasks, this might look, for instance, like an hour of camera footage from a car, meticulously sectioned by humans to designate roads, road signs, vehicles, pedestrians and so forth. But labeling even this small amount of data could take hundreds of hours for a human, bottlenecking the training process. Now, researchers from MIT’s Computer Science & Artificial Intelligence Laboratory (CSAIL) are introducing a new, state-of-the-art algorithm for unsupervised computer vision tasks that operates without any human labels.

The model is called STEGO, short for “Self-supervised Transformer with Energy-based Graph Optimization.” STEGO is a semantic segmentation algorithm, the process of labeling the pixels in an image. Historically, semantic segmentation has been easiest for discrete objects like people or vehicles and harder for more amorphous, blended elements of the environment like clouds or bushes—or cancers.

“If you’re looking at oncological scans, the surface of planets, or high-resolution biological images, it’s hard to know what objects to look for without expert knowledge. In emerging domains, sometimes even human experts don’t know what the right objects should be,” explained Mark Hamilton, a research affiliate of MIT CSAIL, software engineer at Microsoft, and lead author of the paper describing STEGO, in an interview with MIT’s Rachel Gordon. “In these types of situations where you want to design a method to operate at the boundaries of science, you can’t rely on humans to figure it out before machines do.”

STEGO is built on top of the DINO algorithm, itself trained on 14 million images. The researchers tested STEGO on a variety of test cases, including the incredibly diverse COCO-Stuff image dataset. The researchers reported that STEGO doubled the performance of prior unsupervised computer vision models on the COCO-Stuff benchmark, and performed similarly well on tasks like driverless car datasets and space imagery datasets.  ... ' 

Monday, April 04, 2022

Now, Contact Lenses, with Apps?

 Took a look at the idea some decades ago, are they now here effective and practical?  If so, can see the available Apps idea being a very big thing. 

Looking Through Mojo Vision’s Newest AR Contact Lens 

With batteries, motion sensors, and a microLED on board, Mojo lets me try out its first apps. By  Tekla S. Perry,   30 Mar 2022,  In EEEE Expert

This month Mojo Vision unveiled its latest AR contact lens. Still a prototype, the device has clinical testing and further development ahead before it can apply for the U.S. Food and Drug Administration (FDA) approval needed to sell to consumers. But Mojo’s engineers are steadily ticking off engineering milestones.

Last week I got to literally peek through Mojo’s newest lens. Here’s what I saw, and what Steve Sinclair, Mojo senior vice president of product and marketing, had to say about the company’s progress so far and the challenges that remain.

First, the demo. I did not put the lens in my eye—this prototype is still in safety testing, and fitting a contact lens requires an eye exam. Instead, I held a lens very close to one eye and peeked through. I was able to move around freely, but because holding instead of wearing the device means that it cannot track eye movements, Mojo has temporarily incorporated a tiny crosshair into the user interface to help with alignment. The nature of the demo also meant that images I saw were flat; with a lens in both eyes, the images will appear in 3D. ....  ' 

Sunday, January 09, 2022

Navigational Apps for the Blind

New Sensory Applications

Navigational Apps for the Blind Could Have Broader Appeal

The New York Times, Amanda Morris, January 4, 2022

New apps designed specifically for blind and low-vision people could also have mainstream appeal. Leveraging improvements in mapping technology and smartphone cameras, these apps can provide indoor navigation, detailed descriptions of the surrounding environment, and more warnings about obstacles. For example, MapInHood, released only in Toronto so far, offers information about sidewalk traffic, construction hazards, accessible curb cuts, and locations of benches, among other things. It also can help users avoid stairs or steep slopes, which would benefit disabled individuals as well as those carrying suitcases or pushing strollers. Meanwhile, GoodMaps is creating indoor navigational tools for airports, train stations, office buildings, malls, and hospitals. Said GoodMaps' José Gaztambide, "You as a sighted person are going to be able to enter more and more buildings and find your way around more quickly than ever before because of the work we're doing of enabling accessible navigation."   .... ' 

Tuesday, September 14, 2021

Seeing the way we do

New ways of perceptive seeing, now with Texture and Shape

GLOM: Teaching Computers to See the Way(s) We Do  By John Delaney,  Commissioned by CACM Staff   September 14, 2021

At the virtual Collision technology  earlier this year deep learning pioneer Geoffrey Hinton explained how he conceived of a new type of neural network that, he said, would be able to perceive things the way people do.

Hinton, an emeritus distinguished professor in the department of computer science of the Faculty of Arts & Science at Canada's University of Toronto, and also an Engineering Fellow at Google, is responsible for some of the biggest breakthroughs in deep learning and neural networks. He was honored as co-recipient of the 2018 ACM A.M. Turing Award, along with Yoshua Bengio and Yann LeCun, for conceptual and engineering breakthroughs that have made deep neural networks a critical component of modern computing.

In Hinton's Collision talk, he pointed out that the representations used by most neural networks performing object classification are produced by convolutional neural networks, which work well at classifying objects such as images or words, even winning competitions such as the ImageNet Large Scale Visual Recognition Challenge, but they perceive images in a very different way than people do, which can sometimes lead to "crazy errors."

"They use lots of texture information, which people are insensitive to," Hinton said, "but they fail to use a lot of shape information, which people are very sensitive to."  .... '

Saturday, May 22, 2021

IoT and Vision AI with NVIDIA AMA

Just brought to my attention:

AI, Robotics, and IoT video with NVIDIA, here is one episode: 

Everything you needed to know about IoT and vision AI with NVIDIA AMA video. Watch now.

Join the NVIDIA Jetson team for the latest episode of our AMA-style live stream, Jetson AI Labs.  This episode, we'll be talking all things IoT with guest panelist Paul DeCarlo, Principal Cloud Developer Advocate from Microsoft - along with our hosts Dustin Franklin, Dana Sheahen, and Jim Benson from JetsonHacks.  ... 

The stream will begin on Thursday, February 25, 2021 at 10am Pacific time:     https://www.youtube.com/watch?v=HQBqZEcMIrM

Please enter your questions in the live chat window, and we look forward to talking with you!  ... 

Friday, March 05, 2021

Facebook's Seer Recognizes Images

Thinking the implications of this in business practice.

Facebook’s new AI model SEER can teach itself to recognize images  By Mike Wheatley in SiliconAngle

Researchers at Facebook Inc. have created an artificial intelligence-based image recognition model called SEER  (Technical) that’s able to describe what it’s seeing without being trained first on a labeled dataset.

Facebook said that SEER, which is an acronym for “Self-SupeERvised,” is a breakthrough that could lead to a “revolution” in computer vision.

The SEER model, outlined in a paper released March 2, was fed 1 billion publicly available images without annotations or labels from Instagram. It then worked through the dataset, learning as it progressed, and was eventually able to achieve extremely high accuracy in tasks such as object detection.

Self-supervised AI learning is already established in the AI field. It refers to AI systems that can learn directly from whatever information they are given without being trained on carefully labeled datasets that can teach them how to perform a given task, such as recognizing an object in a photo or translating a piece of text.  ... ' 

Sunday, December 20, 2020

Identification of Plants and Weeds

Have been testing a plant recognition system for some time.   Effectiveness varies, but continues to improve.

Deere's Farm Version of Facial Recognition Coming to Fields in 2021   By CNBC

Agriculture giant Deere & Co. plans to roll out a system next summer that combines machine vision and machine learning to improve the identification of individual plants and weeds.

Deere's Jahmy Hindman said neural network models could be trained to only spray weeds in crop fields, killing everything except genetically modified plants designed to survive chemical applications.

Said Hindman, "We are interested in being able to manage each plant over the course of its life, minimizing inputs and maximizing productivity."

The technology would take pictures of plants, and a machine cruising the field would make the decision to spray in just seconds.

Jeffries' Stephen Volkmann said, "See-and-spray is one of several advanced farming technologies that seem to be moving closer to an inflection point," but commercialization of full plant recognition technology is still a few years away. ... '

Saturday, October 31, 2020

Illusory Perceptions

 In our early work in this area we thought we found such 'illusions'. Based on the text here, these were not the same thing as mentioned here,  but we named them 'illusions', inspired by thoughts of biomimicry.  Is this a hint we are getting closer to brain models?

AI Also Has Illusory Perceptions

RUVID/Network of Valencian Universities for the Promotion of Research, Development, and Innovation

October 16, 2020

Researchers at Spain’s Universitat de València (UV) and Pompeu Fabra University have found that convolutional neural networks (CNN) are affected by visual illusions, much like the human brain. The researchers trained CNNs for simple tasks and found they were susceptible to visual illusions of brightness, although the illusions may not coincide with biological illusory perceptions. Said UV's Jesús Malo, "This is one of the factors that leads us to think that it is not possible to establish analogies between the simple concatenation of artificial neural networks and the much more complex human brain." The researchers warned in a separate study about the use of CNNs to study human vision. Said Malo, "In addition to the intrinsic limitations of these artificial networks to model vision, the non-linear behavior of flexible architectures can be very different from that of the biological visual system."

Thursday, September 17, 2020

The Sciences of Reflection

Another example of advanced sensory analysis that can improve 'seeing' in multiple complex  environments. 

Research reflects how AI sees through the looking glass   by Cornell University

AI learns to pick up on unexpected clues to differentiate original images from their reflections, the researchers found. Credit: Cornell University Things are different on the other side of the mirror.

Text is backward. Clocks run counterclockwise. Cars drive on the wrong side of the road. Right hands become left hands.

Intrigued by how reflection changes images in subtle and not-so-subtle ways, a team of Cornell University researchers used artificial intelligence to investigate what sets originals apart from their reflections. Their algorithms learned to pick up on unexpected clues such as hair parts, gaze direction and, surprisingly, beards—findings with implications for training machine learning models and detecting faked images.

"The universe is not symmetrical. If you flip an image, there are differences," said Noah Snavely, associate professor of computer science at Cornell Tech and senior author of the study, "Visual Chirality," presented at the 2020 Conference on Computer Vision and Pattern Recognition, held virtually June 14-19. "I'm intrigued by the discoveries you can make with new ways of gleaning information."   Zhiqui Lin is the paper's first author; co-authors are Abe Davis, assistant professor of computer science, and Cornell Tech postdoctoral researcher Jin Sun.

Differentiating between original images and reflections is a surprisingly easy task for AI, Snavely said—a basic deep learning algorithm can quickly learn how to classify if an image has been flipped with 60% to 90% accuracy, depending on the kinds of images used to train the algorithm. Many of the clues it picks up on are difficult for humans to notice.

For this study, the team developed technology to create a heat map that indicates the parts of the image that are of interest to the algorithm, to gain insight into how it makes these decisions.

They discovered, not surprisingly, that the most commonly used clue was text, which looks different backward in every written language. To learn more, they removed images with text from their data set, and found that the next set of characteristics the model focused on included wrist watches, shirt collars (buttons tend to be on the left side), faces and phones—which most people tend to carry in their right hands—as well as other factors revealing right-handedness. ... "

Thursday, May 14, 2020

Sony AI Intelligent Image Sensors

What appears to be a considerable means of creating IOT intelligence based on such chips.

Sony Says It Created World’s First Image Sensor With Built-in AI
By Takashi Mochizuki and Vlad Savov in Bloomberg
May 14, 2020, 3:00 AM  

Sony Corp. touted on Thursday the world’s first image sensors with built-in artificial intelligence, promising to make data-gathering tasks much faster and more secure. Calling it the first of its kind, Sony said the technology would give “intelligent vision” to cameras for retail and industrial applications.

The new sensors are akin to tiny self-contained computers, incorporating a logic processor and memory. They’re capable of image recognition without generating any images, allowing them to do AI tasks like identifying, analyzing or counting objects without offloading any information to a separate chip. Sony said the method provides increased privacy while also making it possible to do near-instant analysis and object tracking.  ...  "

See also Sony AI Initiatives:   https://www.sony.net/SonyInfo/sony_ai/

Friday, January 24, 2020

Show Devices Can Now Recognize

Just brought to my attention, Alexa 'Show' devices, that is,  those that have an embedded camera,  can now be asked to recognize things.   Apparently only in the domain of' 'pantry items'.  Designed as an assist for the blind or those that need vision assistance.    This might raise some privacy issues despite that the action is requested.

Alexa can now recognize objects
A new Show and Tell feature on Echo Show devices can recognize some objects.   By Molly Price in CNet

If facial recognition concerns were on your radar, get ready to worry about soup cans, too. Amazon today announced a new feature for its Echo Show devices. The feature, called Show and Tell, is focused on recognizing household pantry items when you hold them in front of the camera.... "

Friday, January 17, 2020

Seeing in Higher Dimensions

'Seeing' is constructing useful models about spaces from sensors to understand and navigate them.   Predict current and future states. A good studey of the idea.

An Idea From Physics Helps AI See in Higher Dimensions
The laws of physics stay the same no matter one’s perspective. Now this idea is allowing computers to detect features in curved and higher-dimensional space.

The new deep learning techniques, which have shown promise in identifying lung tumors in CT scans more accurately than before, could someday lead to better medical diagnostics.

Olena Shmahalo/Quanta Magazine
John Pavlus  Contributing Writer

January 9, 2020

Computers can now drive cars, beat world champions at board games like chess and Go, and even write prose. The revolution in artificial intelligence stems in large part from the power of one particular kind of artificial neural network, whose design is inspired by the connected layers of neurons in the mammalian visual cortex. These “convolutional neural networks” (CNNs) have proved surprisingly adept at learning patterns in two-dimensional data — especially in computer vision tasks like recognizing handwritten words and objects in digital images.

But when applied to data sets without a built-in planar geometry — say, models of irregular shapes used in 3D computer animation, or the point clouds generated by self-driving cars to map their surroundings — this powerful machine learning architecture doesn’t work well. Around 2016, a new discipline called geometric deep learning emerged with the goal of lifting CNNs out of flatland.... "

Sunday, December 08, 2019

Humans Seeing Through Animal Eyes

Applications?

Humans Closer to Seeing Through the Eyes of Animals
University of Exeter (U.K.)
December 3, 2019

Researchers at Australia's University of Queensland (UQ) and the University of Exeter in the U.K. have developed a software framework designed to significantly improve humans' ability to analyze complex visual information as animals would. The Quantitative Color Pattern Analysis (QCPA) framework is a kit of digital image processing techniques and analytical tools. Its use of digital photos means it can adapt to nearly any habitat, using technologies ranging from commercially available cameras to full-spectrum imaging systems. UQ's Karen Cheney said the framework is sufficiently flexible to explore the color patterns and natural surroundings of many organisms. Said Cheney, "We're helping people—wherever they are—to cross the boundaries between human and animal visual perception."  .... '

Friday, November 22, 2019

AI Birdwatching

Seeing and  imaging continue to be the place where AI/Machine learning make the most progress. Here another tagging example.   Currently working on a horticultural example with plants.

This AI birdwatcher lets you 'see' through the eyes of a machine
by Robin A. Smith, Duke University

It can take years of birdwatching experience to tell one species from the next. But using an artificial intelligence technique called deep learning, Duke University researchers have trained a computer to identify up to 200 species of birds from just a photo.

The real innovation, however, is that the A.I. tool also shows its thinking, in a way that even someone who doesn't know a penguin from a puffin can understand.

The team trained their deep neural network—algorithms based on the way the brain works—by feeding it 11,788 photos of 200 bird species to learn from, ranging from swimming ducks to hovering hummingbirds.

The researchers never told the network "this is a beak" or "these are wing feathers." Given a photo of a mystery bird, the network is able to pick out important patterns in the image and hazard a guess by comparing those patterns to typical species traits it has seen before.

Along the way it spits out a series of heat maps that essentially say: "This isn't just any warbler. It's a hooded warbler, and here are the features—like its masked head and yellow belly—that give it away."

Duke computer science Ph.D. student Chaofan Chen and undergraduate Oscar Li led the research, along with other team members of the Prediction Analysis Lab directed by Duke professor Cynthia Rudin.

They found their neural network is able to identify the correct species up to 84% of the time—on par with some of its best-performing counterparts, which don't reveal how they are able to tell, say, one sparrow from the next.

Rudin says their project is about more than naming birds. It's about visualizing what deep neural networks are really seeing when they look at an image.  .... " 

Monday, October 28, 2019

Vehicles Seeing Around Corners

Interesting advances continue in the sensor space.  Another example of sensors inferring related information.

Helping autonomous vehicles see around corners
By sensing tiny changes in shadows, a new system identifies approaching objects that may cause a collision.   Rob Matheson | MIT News Office

To improve the safety of autonomous systems, MIT engineers have developed a system that can sense tiny changes in shadows on the ground to determine if there’s a moving object coming around the corner.  

Autonomous cars could one day use the system to quickly avoid a potential collision with another car or pedestrian emerging from around a building’s corner or from in between parked cars. In the future, robots that may navigate hospital hallways to make medication or supply deliveries could use the system to avoid hitting people.

In a paper being presented at next week’s International Conference on Intelligent Robots and Systems (IROS), the researchers describe successful experiments with an autonomous car driving around a parking garage and an autonomous wheelchair navigating hallways. When sensing and stopping for an approaching vehicle, the car-based system beats traditional LiDAR — which can only detect visible objects — by more than half a second.  ... "

Monday, October 21, 2019

In Context Facial Recognition Improved

More work in this space, apparently particularly useful when seeing alternative views to ID.  The UK's massive coverage of CCTV cameras would benefit considerably.  Probably also will increase the objection to using the approach for first level policing,  since the results will still not be perfect.    But improved in-context recognition should be useful.

AI research achieves world-leading technology for visual recognition of people  by University of Surrey in Techxplore

AI is increasingly being used to help human operators handle massive amounts of images from CCTV and other security sources. Person re-identification (ReID) is a method in which an AI is able to recognize images of the same person taken from different cameras or on different occasions. This helps to track suspects across a CCTV network covering large public space, such as an underground network. ReID is challenging for machines as they have to consider and differentiate the same person under different light sources, poses and changes in appearance such as their clothes.


In a paper to be presented at this year's International Conference on Computer Vision in Seoul, South Korea, the most prestigious conference in visual AI, experts from Surrey's Centre for Vision, Speech and Signal Processing (CVSSP) detail how they have developed a unique system called OSNet that has outperformed many popular identification systems already in use. ... " 

Monday, September 30, 2019

AI Improving Biomedical Imaging

It is notable how modern AI is doing best in 'vision' spaces.    As opposed to what I would call conversational interaction and process logic.   Not what we would have expected in the earlier applications of AI.  Is this because the training data is more available and concise, or because the underlying deep learning models are closer to the underlying human intelligence?   Or both?

Artificial intelligence improves biomedical imaging in TechXplore
by Fabio Bergamin, ETH Zurich

ETH researchers use artificial intelligence to improve quality of images recorded by a relatively new biomedical imaging method. This paves the way towards more accurate diagnosis and cost-effective devices.

Scientists at ETH Zurich and the University of Zurich have used machine learning methods to improve optoacoustic imaging. This relatively young medical imaging technique can be used for applications such as visualizing blood vessels, studying brain activity, characterizing skin lesions and diagnosing breast cancer. However, quality of the rendered images is very dependent on the number and distribution of sensors used by the device: the more of them, the better the image quality. The new approach developed by the ETH researchers allows for substantial reduction of the number of sensors without giving up on the resulting image quality. This makes it possible to reduce the device cost, increase imaging speed or improve diagnosis. .... "