/* ---- Google Analytics Code Below */
Showing posts with label Training Data. Show all posts
Showing posts with label Training Data. Show all posts

Thursday, March 30, 2023

What Training/Testing Data?

Interesting, was wondering which training data was used where, and the consequences.  Is that readily testable?  

ARTIFICIAL INTELLIGENCE/TECH/GOOGLE

Google denies Bard was trained with ChatGPT data

The Information published a report Wednesday including allegations from a former Google AI researcher that the company used a rival’s responses to train its own chatbot. Google denies that Bard uses that data.

By SEAN HOLLISTER

Mar 29, 2023, 10:10 PM EDT|19 Comments / 19 New

DeepMind reportedly lost a yearslong bid to win more independence from Google

OpenAI co-founder on company’s past approach to openly sharing research: ‘We were wrong’

AI chatbots compared: Bard vs. Bing vs. ChatGPT

Google’s Bard hasn’t exactly had an impressive debut — and The Information is reporting that the company is so interested in changing the fortunes of its AI chatbots, it’s forcing its DeepMind division to help the Google Brain team beat OpenAI with a new initiative called Gemini. The Information’s report also contains the potentially staggering thirdhand allegation that Google stooped so low as to train Bard using data from OpenAI’s ChatGPT, scraped from a website called ShareGPT. A former Google AI researcher reportedly spoke out against using that data, according to the publication.

But Google is firmly and clearly denying the data was used: “Bard is not trained on any data from ShareGPT or ChatGPT,” spokesperson Chris Pappas tells The Verge.  ... ' 

Saturday, February 26, 2022

Synthetic vs Real Training Data?

The real world is messy and needs direction.

Are You Still Using Real Data to Train Your AI?

500+IEEE Spectrumby Eliza Strickland / 3d//keep unread//hide

It may be counterintuitive. But some argue that the key to training AI systems that must work in messy real-world environments, such as self-driving cars and warehouse robots, is not, in fact, real-world data. Instead, some say, synthetic data is what will unlock the true potential of AI. Synthetic data is generated instead of collected, and the consultancy Gartner has estimated that 60 percent of data used to train AI systems will be synthetic. But its use is controversial, as questions remain about whether synthetic data can accurately mirror real-world data and prepare AI systems for real-world situations.

Nvidia has embraced the synthetic data trend, and is striving to be a leader in the young industry. In November, Nvidia founder and CEO Jensen Huang announced the launch of the Omniverse Replicator, which Nvidia describes as “an engine for generating synthetic data with ground truth for training AI networks.” To find out what that means, IEEE Spectrum spoke with Rev Lebaredian, vice president of simulation technology and Omniverse engineering at Nvidia.

Rev Lebaredian on...

What Nvidia hopes to achieve with Omniverse

Why today’s real-world data isn’t good enough

Why autonomous vehicles need synthetic data

Overfitting, algorithmic bias, and adversarial attacks

The Omniverse Replicator is described as “a powerful synthetic data generation engine that produces physically simulated synthetic data for training neural networks.” Can you explain what that means, and especially what you mean by “physically simulated”?  .... '