/* ---- Google Analytics Code Below */
Showing posts with label crowd sourcing. Show all posts
Showing posts with label crowd sourcing. Show all posts

Saturday, February 13, 2021

Weaknesses in ML

I like explanations of how a solution works in real world contexts.   Lots of work to still do.

Uncovering Unknown Unknowns in Machine Learning  Google Blog., Thursday, February 11, 2021

Posted by Lora Aroyo and Praveen Paritosh, Research Scientists, Google Research

The performance of machine learning (ML) models depends both on the learning algorithms, as well as the data used for training and evaluation. The role of the algorithms is well studied and the focus of a multitude of challenges, such as SQuAD, GLUE, ImageNet, and many others. In addition, there have been efforts to also improve the data, including a series of workshops addressing issues for ML evaluation. In contrast, research and challenges that focus on the data used for evaluation of ML models are not commonplace. Furthermore, many evaluation datasets contain items that are easy to evaluate, e.g., photos with a subject that is easy to identify, and thus they miss the natural ambiguity of real world context. The absence of ambiguous real-world examples in evaluation undermines the ability to reliably test machine learning performance, which makes ML models prone to develop “weak spots”, i.e., classes of examples that are difficult or impossible for a model to accurately evaluate, because that class of examples is missing from the evaluation set.

To address the problem of identifying these weaknesses in ML models, we recently launched the Crowdsourcing Adverse Test Sets for Machine Learning (CATS4ML) Data Challenge at HCOMP 2020 (open until 30 April, 2021 to researchers and developers worldwide). The goal of the challenge is to raise the bar in ML evaluation sets and to find as many examples as possible that are confusing or otherwise problematic for algorithms to process. CATS4ML relies on people’s abilities and intuition to spot new data examples about which machine learning is confident, but actually misclassifies.  ... 

Saturday, January 09, 2021

Picture Pile Platform to Crowdsource Data for Image Classification

Towards more classification examples for AI models.

Crowdsourced Image Classification Will Train AI Models  By International Institute for Applied Systems Analysis,    January 8, 2021

A Proof of Concept Grant from the European Research Council for the Picture Pile Platform will provide users with an opportunity to set up and run crowd-sourced image classification campaigns, and then make the images available to train AI algorithms.

ERC Proof of Concept Grants provide top-up funding to ERC grantees to explore the potential of scientific discoveries and bring the results closer to market.

"The new platform will address the gap that currently exists in the market for a platform that allows users to build their own tailored, quality controlled crowdsourcing campaigns to collect image classifications in an efficient, engaging, and fair way, and then possibly make the data collected openly and freely available," says Steffen Fritz, strategic initiatives program director at the International Institute for Applied Systems Analysis, who will lead the project.

While existing image databases can be used to train machine learning algorithms to perform computer vision tasks, there is a lack of datasets containing more specific features of interest. The Picture Pile Platform will address this by building upon the existing Picture Pile crowdsourcing application that allows users to classify or help sort through piles of pictures, which can be very high resolution satellite images, geo-tagged photographs, or any other images that require sorting. After a pile has been sorted, the image classifications can be made publicly available with FAIR (Findable, Accessible, Interoperable, and Reusable) metadata so that they can be freely used by anyone.

The gamified version of the Picture Pile annotation tool is accessible as an online version, and as a mobile app in IOS and Android versions.

From International Institute for Applied Systems Analysis  ... '

Monday, October 12, 2020

IBM Enables AI Enabled Debate

This could lead to AI being able to make 'arguments' as parts of business process.  Could be used broadly, for example in courts and as parts of smart contract testing.   Or to test advertising pitches in context.  Truly a newly emergent part of AI.  An improved model for crowd sourcing?

IBM showcases latest A.I. advancements on Bloomberg's "That's Debatable" TV show

Software, which IBM hopes to sell to businesses, distills 'key points' from thousands of individual comments  ... "

More from IBM on Project Debater  ....    Related Blog on Project Debater

Wednesday, September 23, 2020

Participation Washing?

Stated term was new to me,  but a important in any kind of data that is crowd sourced.  Should always be a consideration.  A fix?  Not that either.

Participation-washing could be the next dangerous fad in machine learning
Many people already participate in the field’s work without recognition or pay.  by Mona Sloane in Technology Review
 
 ...  Now, machine-learning researchers and scholars are looking for ways to make AI more fair, accountable, and transparent—but also, recently, more participatory.

One of the most exciting and well-attended events at the International Conference on Machine Learning in July was called “Participatory Approaches to Machine Learning.” This workshop tapped into the community’s aspiration to build more democratic, cooperative, and equitable algorithmic systems by incorporating participatory methods into their design. Such methods bring those who interact with and are affected by an algorithmic system into the design process—for example, asking nurses and doctors to help develop a sepsis detection tool.

This is a much-needed intervention in the field of machine learning, which can be excessively hierarchical and homogenous. But it is no silver bullet: in fact, “participation-washing” could become the field's next dangerous fad. That’s what I, along with my coauthors Emanuel Moss, Olaitan Awomolo, and Laura Forlano, argue in our recent paper “Participation is not a design fix for machine learning.”  ... " 

Monday, March 23, 2020

Crowdsourcing Creativity

Spent much time on broad the idea, like the direction of the process.

Crowdsourcing Plot Lines to Help the Creative Process

Penn State News
Jessica Hallman
March 13, 2020

Researchers at the Pennsylvania State University (Penn State) College of Information Sciences and Technology have launched a crowdsourced system that provides writers with story ideas from the online crowd to facilitate the creative process. The Heteroglossia system lets authors share sections of their story drafts using a text editor, and online workers are tasked with brainstorming plot ideas from the perspective of fictional characters they are assigned. Penn State's Ting-Hao (Kenneth) Huang said human workers currently power the system, but artificial intelligence could be incorporated into the platform in the future. "I believe if we learn how to help creative writing or creative processes in general, we can learn more about how to build systems that can be creative." 

Tuesday, February 04, 2020

IBM Upgrades Debate AI Tool to Better Derive Evidence

Intriguing approach to mining information to support a goal directed conversation.   A key aspect to making conversational systems more powerful.  Note also the crowdsourcing integrated here to grade evidence.  Noting that the report here does not mention 'Watson', it seems IBM is using their AI trademark much less these days.

IBM's Debating AI Just Got a Lot Closer to Being a Useful Tool
By MIT Technology Review via CACM

The IBM Debater system taking part in a debate at the University of Cambridge last year.
IBM upgraded the neural networks used by its Project Debater system, to improve the quality of evidence the argument-mining system uncovers.

IBM upgraded the neural networks used by its Project Debater system to improve the quality of evidence the argument-mining system uncovers.

One new add-on for the debating system is BERT (Bidirectional Encoder Representations from Transformers), a network designed by Google for natural language processing and answering queries.

IBM Research scientists trained the AI on 400 million documents from the LexisNexis database, providing a natural language dataset of roughly 10 billion sentences; the researchers combined the dataset with claims about several hundred different topics, then had crowdsourced workers label the sentences based on the quality of their evidence for or against specific claims.

A supervised learning algorithm digested this data, allowing BERT to manage queries on a wide range of subjects and to yield more relevant sentences compared to previous systems.

Project Debater was 95% accurate for the top 50 sentences across 100 distinct topics, according to IBM researcher Noam Slonim,  ... " 

Thursday, August 22, 2019

Crowd Sources Autonomous Vehicle Training

From Ideaconnection, offer of sale or licensing of  a US patent:

Crowd-sourced Autonomous Vehicle Training

Definitions: When driving, human drivers encounter non-events (expected events like traffic lights turning red) as well as events (unexpected events like an unaccompanied child standing at the edge of the road).

Events are captured from crowd-sourced participants driving their own cars during their routine lives. The determination that a particular scenario is an unexpected event is carried out automatically (without manual input) by sensing eye, foot and hand positions and movements, and comparing it to a map. This arrangement can not only sense the driver’s actions to change speed or direction in response to such unexpected events, it can also detect the driver’s intentions to do so. It can also determine various driving attributes and mental components of the driver. This in turn allows drivers to be scored and ranked in each geographical region, so that expert drivers can be identified. 

Signatures for events extracted from the driving patterns of such expert drivers are more reliable, and form better training routines for improving autonomous vehicle software. A database of thousands of such signatures, and their variations, are obtained by crowd-sourced expert drivers. These signatures are then used by autonomous vehicles to identify potential or real events, and react to these events in a human-like manner. 

Companies developing AVs need not wait for a history of “millions of miles tested”, but instead can rely on extremely quick, cheap and highly reliable, human-like training sub-routines extracted from crowd-sourced drivers.   .... " 

Monday, June 24, 2019

Alexa Expands the Ability to 'Announce'

At first this seemed like a trivial thing, you can send information to a group of people.  Or other internal devices, or external things.  So this is like a specific form of communications,  I can reply, or wait for more information.   A means to effect crowd sourcing.   Crowd sourcing  about specific needs?   We do it in conversation all the time.   Something here that could be expanded.

Alexa's intercom-like broadcasts come to more non-Echo devices
You could send announcements through your thermostat.

By Jon Fingas, @jonfingas in Engadget

Amazon has slowly been expanding the circle of devices that can use Alexa Announcements, but now it's throwing the gates wide open. The company has made the intercom-like feature available to any device with Alexa support built-in -- you could use your thermostat or fridge to tell the kids that dinner is ready. In theory, you won't have to visit a specific room like you might today.

The feature requires device makers to implement it, and not every product will necessarily qualify. They'll have to support converting MP3 files 45 seconds or longer to an Alexa-ready format. There may still be gaps in Announcements support even if your home is full of Alexa devices. Still, this could make the broadcasting tool far more flexible in the long run. ... " 

Thursday, January 03, 2019

Alexa Asks for Answers

A kind of admission that its hard to create general answers to questions?   Further relating to my last post about this kind of approach.   Do note its similarity to what you get with Google search, linking to previous questions with similar keywords and structure.     I have noted recently an increased number of answers to questions.   Attribution indication also interesting.

Amazon Alexa is beta testing crowdsourced answers
Customers in the invitation-only program can start contributing answers to Alexa today  By Dami Lee@dami_lee in TheVerge.

Amazon announced recently that it’s beta testing Alexa Answers, an invitation-only program for users to add responses for questions that Alexa can’t answer. Amazon says that in the last month of its internal Alexa Answers beta program, 100,000 responses have been added and served to customers “millions of times.”

Starting today, selected customers who have received email invitations can participate via the Alexa Answers website. Users can browse through various topic categories, like science and geography, and choose to answer questions that have been asked by other customers that Alexa doesn’t know the answers to. Those answers may then be served to other Alexa customers, and their answer will be attributed to “an Amazon customer.”    ... " 

Thursday, September 06, 2018

Wal-Mart Pilots Crowd sourced Grocery Delivery

What appears to be a novel effort with independently managed delivery drivers.

Walmart pilots crowdsourced grocery delivery platform
Spark Delivery store-to-door service uses independent drivers  By Russell Redman  in SupermarketNews

Walmart has begun testing a crowdsourcing-based service that enlists drivers using their own vehicles to provide last-mile delivery of groceries ordered online.

Dubbed Spark Delivery, the service leverages delivery logistics platform Bringg to engage with independent drivers to pick up grocery orders at Walmart stores and deliver them directly to customers. The drivers are recruited and managed by Delivery Drivers Inc. (DDI), a national firm specializing in last-mile contractor management.  .... " 

Friday, August 31, 2018

Crowdsourcing Training Data for Corn Identification

A classic way to get training data.

Researchers use crowdsourcing to speed up data analysis in corn plants

Iowa State University News Service   By Fred Love

Iowa State University researchers used crowdsourcing to train a computer model to identify the tassels of corn plants from a vast number of photographic images. The crowdsourcing effort produced similar results to those of trained plant scientists and yielded an algorithm that the researchers say will greatly reduce the time it takes to derive useful metrics from massive datasets. The researchers used Amazon Mechanical Turk to find participants for the study, who received instructions to identify tassels in dozens of images of corn by drawing a square around them, and then used those labeled images to train a computer to identify tassels in similar corn images. The researchers said this approach could generate similar results for other types of plants. ... "

Wednesday, August 01, 2018

Superforecasting Update

Just received their latest newsletter, of interest:

Goodjudgment:  How Can Superforecasting     Improve your Decisions?

Good Judgment’s co-founder Philip Tetlock literally wrote the book on state-of-the-art crowd-sourced forecasting.

Now, the training, techniques, and talent that helped the Good Judgment Project win a massive government-sponsored forecasting competition can help your organization manage strategic uncertainty. .... " 

In this newsletter, we share the Wall Street Journal's view on Superforecasting, an update from our partnership with Eurasia Group, Superforecasters view on EU Article 7 sanctions, Superforecaster Frederic Bush's article in Two Plus Two on Superforecasting in Poker, and reveal the GJ Open crowd's tip for the World Cup winner.  .... " 

See also, IARPA's sponsorship of work in superforecasting.

Friday, June 01, 2018

Crowd Flower is now Figure Eight

Formerly Crowdflower has been renamed Figure Eight, with a new web presence.  Impressed by what I have seen there so far.   Putting humans in the loop of machine learning ...

" .... Figure Eight is the essential Human-in-the-Loop AI platform for data science and machine learning teams. The Figure Eight software platform trains, tests, and tunes machine learning models to make AI work in the real world.

At Figure Eight, we believe that AI’s three essential ingredients are training data, machine learning, and human-in-the-loop technology. We pride ourselfvesin providing every component necessary to make AI work in the real world.

Training data is the fuel for machine learning, where humans guide the algorithms by labeling data for the algorithms to build their knowledge. You can read about how we approach training data in our Definitive Guide to Training Data. Human-in-the-loop processing allows the feedback loop between humans and machines to be as optimal as possible, whether that’s finding the right raw data to annotate through Active Learning, or selecting the right annotation interfaces and quality controls for accurate and efficient human input.

The most important problem we face in technology today is how humans and machines and work together to solve tasks. These breakthroughs will be vital in many of the use cases we power today, whether it’s the obvious cases today like autonomous vehicles and AI-powered assistants, to the future technologies that we are supporting in areas like healthcare and agriculture. .... " 

Sunday, May 20, 2018

Crime Matching with GEDMatch

Interesting this has just become apparent, genetic matching starts to work against increasing stored data and matching.  Shows the power of cowdsourced databases.  Other examples?    Technology Review Shows why and how:

Another arrest shows why no one can hide from the genetic detectives
For the second time this year, investigators used a public DNA database to solve a cold case and find a murderer.

The bust: A 55-year-old truck driver, William Talbott, was arrested today in Washington State after being fingered in a 30-year-old double murder.

How they found him: According to Buzzfeed, investigators located Talbott’s family members after uploading old crime scene DNA to GEDMatch, a crowdsourced database that genealogists use to compare DNA and build family trees.  ...  "

It further comes to mind that this is akin to:

 ' ...   "The Selfish Ledger,” was shared internally within Google. The video examines the possibility of a dystopian world where our use of devices such as smartphones creates a sort of digital DNA, which, like physical DNA, could exist within the context of future generations. ..."

More on that and links to the video on my post here.   Will the crimes of the past always match the crimes considered in the future?

Thursday, May 03, 2018

Alibaba Crowdsourcing

Further joint research in Asia:

NTU Singapore Partners With Alibaba to Set Up Joint Research Institute for AI Technologies

Priyankar Bhunia

Researchers at Nanyang Technological University (NTU) in Singapore and China's Alibaba Group have launched the Alibaba-NTU Singapore Joint Research Institute to create and test artificial intelligence (AI) solutions to address societal challenges. The institute seeks to combine NTU's human-centered AI technology with Alibaba's natural-language processing, computer-vision, machine-learning, and cloud computing technologies. NTU and Alibaba also will collaborate on a crowdsourcing platform to connect researchers and industry partners around the world within an AI-focused research and development ecosystem. The technologies will be tested on the NTU Smart Campus to demonstrate their effectiveness before they are distributed. The overall goal is to deploy AI solutions over the next five years in a range of scenarios to help people live healthier, smarter, and happier lives. "Using AI technologies, we can address fundamental societal challenges such as aging population, which is a huge issue for cities with a rapidly aging population such as Singapore," says NTU professor Subra Suresh. .... "

Tuesday, April 10, 2018

Barnes & Noble Crowdsources Reading with Browsery

Been a while that I have heard much from B&N, so this is interesting, can it recreate the bookstore virtually?   Can this alter the nature of reading, book selection?  With come expert comment:

Barnes & Noble’s crowdsourcing app engages readers and earns solid reviews  By Tom Ryan

Barnes & Noble has launched Browsery, an app that uses crowdsourcing to help readers discover new books.

The app basically features a bunch of questions that support “browsing, community, and conversation.” Customers can “like”, comment on or contribute answers to questions about books posted by the Browsery community. They may also like or comment on the answers of others or ask questions of their own.

For example, the Biography & Memoir section includes questions such as:

“Favorite memoirs that include recipes?”
“What are some interesting biographies for a film-lover?”
“What celebrity made you laugh the hardest?”

Launched in late March, one question has already drawn 84 responses. Clicking through enables users to “agree” with a suggestion and offer the reason why. The authors with the most “agrees” to the “laugh the hardest” question included Tina Fey, Amy Poehler and Trevor Noah. .... " 

Thursday, March 15, 2018

Human in the Loop Machine Learning

Attended a very good webinar today in the DSC series.  Strongly recommend joining DSC and taking advantage of their free resources.

This Webinar answers the question you will have as a data scientist.  Where will I get the data to train my models, when its mostly held by people?

Now renamed:  Figure Eight   https://www.figure-eight.com/ 

Robert Munro, CTO of Figure Eight  answers in this recorded Webinar: 

"    ... Curious about what human-in-the-loop machine learning actually looks like? Join CrowdFlower and learn how to effectively incorporate Active Learning, Transfer Learning, and Annotation Quality in your ML projects to achieve better results. 

Join us in this latest Data Science Central webinar, where we will cover the following topics:

When to use the human-in-the-loop as an effective strategy for machine learning projects

How to set up an effective interface to get the most out of human intelligence

How to ensure high-quality, accurate training data sets

How to use ML models from different domains to improve your own labeling

​This webinar will include an end-to-end look at setting up and running a job that generates high-quality training data, and shows how to incorporate that training data into human-in-the-loop machine learning systems that you can run in your own environment.

Speaker: Robert Munro, Chief Technology Officer -- CrowdFlower
Hosted by: Bill Vorhies, Editorial Director -- Data Science Central .... " 

Tuesday, March 06, 2018

Human in the Loop Machine Learning

Planning to Attend ....

Practical Human-in-the-Loop Machine Learning
Join us for this latest DSC Webinar on March 15th, 2018
Register Now! 

Curious about what human-in-the-loop machine learning actually looks like? Join CrowdFlower and learn how to effectively incorporate Active Learning, Transfer Learning, and Annotation Quality in your ML projects to achieve better results. 

Join us in this latest Data Science Central webinar, where we will cover the following topics: 
When to use the human-in-the-loop as an effective strategy for machine learning projects

How to set up an effective interface to get the most out of human intelligence

How to ensure high-quality, accurate training data sets

How to use ML models from different domains to improve your own labeling

​This webinar will include an end-to-end look at setting up and running a job that generates high-quality training data, and shows how to incorporate that training data into human-in-the-loop machine learning systems that you can run in your own environment.

Speaker: Robert Munro, Chief Technology Officer -- CrowdFlower

Hosted by: Bill Vorhies, Editorial Director -- Data Science Central

Title: Practical Human-in-the-Loop Machine Learning
Date: Thursday, March 15th, 2018,  Time: 9:00 AM - 10:00 AM PT

Monday, February 12, 2018

Towards Crowd Sourcing Conversational Agents

Intriguing thought, not sure it actually is the same as what we did and called a 'Concierge' model, linking appropriately to other smart agents or people.   More like Facebook's attempt to include humans within the now defunct Facebook M.   Interesting experiment with the examples.  Once again using Amazon's  Mechanical Turk to crowd source.  And again our own experience was that the careful setup is key.

Crowd Workers, AI Make Conversational Agents Smarter 
Carnegie Mellon News,   By Byron Spice

Researchers at Carnegie Mellon University (CMU) have developed Evorus, a chatbot system that recruits crowd workers on demand from Amazon Mechanical Turk to answer questions from users, with the crowd workers voting on the best answer. Evorus also tracks the questions that have been asked and answered, and over time, it will begin to suggest these answers for subsequent questions. 

The researchers also developed a process by which the artificial intelligence (AI) can help to approve a message with less crowd worker involvement. During a five-month deployment, Evorus worked with 80 users and 181 conversations, and its automated responses to questions were chosen 12 percent of the time, crowd voting was reduced by almost 14 percent, and the cost of crowd work for each reply to a user's message dropped by 33 percent. The researchers will present Evorus in April at the ACM Conference on Human Factors in Computing Systems (CHI 2018) in Montreal, Canada. .... " 


Friday, November 10, 2017

Google Maps Show Checkout Line Length

Adding line lengths to maps of stores.  Saw a similar approach done at a Kroger, but just for their store.  The data will be crowd sourced, which can lead to incorrect numbers.   With expert discussion:

Has Google solved the problem of long lines at grocery checkouts?   by Matthew Stern

Grocers have tried plenty of tech solutions over the years to shorten the time customers spend in line. But the latest wait-shortening technology is coming not from an individual grocer, but Google’s crowdsourced data.

Google Maps is implementing functionality that uses anonymous crowdsourced location data to let users see estimated wait times at grocery stores, according to Thrillist. Google will roll the technology out before Thanksgiving. The tool for determining the wait at grocery stores is the second of two data-driven crowd size-gauging tools. The first, which launched Tuesday, allows users to determine what the average wait time will be for a table at a given restaurant.

Grocer strategies for limiting customer frustration with lines have taken many forms. Some have tried distraction, installing digital signage to make the time in line pass more quickly. Others, like Hy-Vee, have tried data-driven “traffic lights” that gauge the business of each line and signal each customer to the fastest checkout.  ... "