/* ---- Google Analytics Code Below */
Showing posts with label Berkeley. Show all posts
Showing posts with label Berkeley. Show all posts

Tuesday, June 20, 2023

IBM Quantum Computer Beats Supercomputer in Benchmark

More advances, depends of course on the type of problem, link to details below.

An IBM Quantum Computer Beat a Supercomputer in a Benchmark Test

By Shelly Fan, June 20, 2023 in Singularity Hub

Quantum computers may soon tackle problems that stump today’s powerful supercomputers—even when riddled with errors.

Computation and accuracy go hand in hand. But a new collaboration between IBM and UC Berkeley showed that perfection isn’t necessarily required for solving challenging problems, from understanding the behavior of magnetic materials to modeling how neural networks behave or how information spreads across social networks.

The teams pitted IBM’s 127-qubit Eagle chip against supercomputers at Lawrence Berkeley National Lab and Purdue University for increasingly complex tasks. With easier calculations, the Eagle matched the supercomputer’s results every time—suggesting that even with noise, the quantum computer could generate accurate responses. But where it shone was in its ability to tolerate scale, returning results that are—in theory—far more accurate than what’s possible today with state-of-the-art silicon computer chips.

At the heart is a post-processing technique that decreases noise. Similar to looking at a large painting, the method ignores each brush stroke. Rather, it focuses on small portions of the painting and captures the general “gist” of the artwork.   ... .' 

Tuesday, May 30, 2023

Researchers from UC Berkeley Introduce Gorilla LLM

 And more implementations. 

Researchers from UC Berkeley Introduce Gorilla: A Finetuned LLaMA-based Model that Surpasses GPT-4 on Writing API Calls

By Tanya Malhotra

A recent breakthrough in the field of Artificial Intelligence is the introduction of Large Language Models (LLMs). These models enable us to understand language more concisely and, thus, make the best use of Natural Language Processing (NLP) and Natural Language Understanding (NLU). These models are performing well on every other task, including text summarization, question answering, content generation, language translation, and so on. They understand complex textual prompts, even texts with reasoning and logic, and identify patterns and relationships between that data.

Though language models have shown incredible performance and have developed significantly in recent times by demonstrating their competence in a variety of tasks, it still remains difficult for them to use tools through API calls in an efficient manner. Even famous LLMs like GPT-4 struggle to generate precise input arguments and frequently recommend inappropriate API calls. To address this issue, Berkeley and Microsoft Research researchers have proposed Gorilla, a finetuned LLaMA-based model that beats GPT-4 in terms of producing API calls. Gorilla helps in choosing the appropriate API, improving LLMs’ capacity to work with external tools to carry out particular activities.   .... ' 


Monday, May 08, 2023

OpenLLaMA is a fully open-source LLM, now ready for business

New LLM Connections are open for business, have been trying for a while, though are still a bit choppy.

OpenLLaMA is a fully open-source LLM, now ready for business

OpenLLaMA is an open-source reproduction of Meta’s LLaMA language model and can be used commercially.

Since the unveiling of Meta’s LLaMA family of large language models and the subsequent leak, the development of open-source chatbots has exploded. Models such as Alpaca, Vicuna, and OpenAssistant use Meta’s models as the basis for their various forms of instruction tuning.

However, LLaMA models are licensed for research use only, which prevents commercial use of those models.

OpenLLaMA reproduces Meta’s language models

Alternatives based on other freely available models do not match the quality of Meta’s models, as LLaMA follows Deepmind’s Chinchilla scaling laws and has been trained on particularly large amounts of data.

THE DECODER Newsletter

Researchers at Berkeley AI Research want to replicate Meta’s LLaMA models in the OpenLLaMA project. The team is using Together’s RedPajama dataset for the project. The open-source platform also announced its intention to reproduce the LLaMA models in April, releasing the 1.2 trillion parameter dataset as a first step.

The Berkeley team is now releasing an early version of the 7-billion-parameter OpenLLaMA model, which has so far been trained on 300 billion of 1.2 trillion tokens. Performance is already said to be approaching the level of LLaMA, and the team is confident that the fully trained OpenLLaMA will be competitive with Meta’s original .... ' 

Tuesday, April 25, 2023

Researchers From Google AI and UC Berkeley Propose an AI Approach That Teaches LLMs to Debug

Researchers From Google AI and UC Berkeley Propose an AI Approach That Teaches LLMs to Debug its Predicted Program via Few-Shot Demonstrations

By Aneesh Tickoo -April 14, 2023  in MarketTech

Producing accurate code in a single effort for many programming jobs can be challenging. With several applications, including code synthesis from natural languages, programming by examples, and code translation, code creation has long been a problem. Recent big language models, in particular, have substantially improved over earlier deep neural networks. One line of research has developed reranking techniques to choose the best candidate from multiple samples, typically requiring tens of samples. These techniques were inspired by observations that correct code is much more likely to be predicted when various programs are sampled from the model.

It makes intuitive sense that a programmer’s first piece of code is usually inaccurate. Humans often examine the code, check into the execution outcomes, and then make adjustments to fix implementation flaws rather than entirely rejecting faulty code. Previous research has suggested deep learning algorithms to correct the anticipated code, which shows considerable performance improvements on various coding jobs. Nevertheless, these methods call for extra training for the code repair model.

Prior studies suggest that large language models are not yet able to correct code in the absence of external feedback, such as unit tests or human instructions, despite some recent studies showing that these models have the potential to generate feedback messages to critique and refine their outputs for some natural language and reasoning domains. In this study, researchers from Google Research and UCB offer SELF-DEBUGGING, using few-shot prompting to educate the huge language model on debugging its own projected code. SELFDEBUGGING commands the model to run the code, then create a feedback message based on the code and the execution outcome without needing extra model training.  ... ' 

Wednesday, April 12, 2023

Rage Against Intelligent Machines

Good overview of the topic with links to 

ACM NEWS

Rage Against the Intelligent Machines

By Paul Marks,  Commissioned by CACM Staff, April 11, 2023

If the launch of ChatGPT in November 2022 was the point at which generative artificial intelligence (AI) began to make an appreciable impact on the public consciousness, the final week of March 2023 was the start of a multi-faceted fightback against AI, one that could have deep ramifications for the freedom firms have to roll out machine intelligences into the public domain.

The AI counter-offensive that week involved a number of high-profile organizations questioning the risks inherent in the largely unregulated way emerging Large Language Models (LLMs)—like OpenAI's ChatGPT and GPT-4, Microsoft's Bing Chat and Google's Bard systems—are being fielded.

At issue, they say, is the way LLMs are being unleashed without prior, transparent, and auditable assessment of their risks, such as aiding and abetting cybercrime, their propensity for simply fabricating facts people might rely on, reinforcing dangerous disinformation, and exhibiting overt and offensive societal biases. Some are calling for LLM development to be halted while measures to make them safe are thrashed out.

This was not just an argument amongst AI cognoscenti. News of the spat even reached the White House, with President Biden reiterating on April 5 that artificial intelligence providers, like all technology companies, "have a responsibility to make sure their products are safe before making them public." 

First out of the gate, on March 27, was Europol, the joint criminal intelligence organization of the 27 nations of the European Union, which published a report  the "diverse range of criminal use cases" it predicts products like ChatGPT could be used in.

Europol's digital forensics experts found the LLM's ability to quickly produce convincing written text in many languages would serve to hide the telltale typos and grammatical errors that are normally a giveaway with phishing messages,  and so boost the success of phishing campaigns.

Europol also said the ability to write messages in anybody's writing style is a gift to fraudsters impersonating employees to entice their colleagues to download malware, or to move large amounts of cash, as has happened in so-called "CEO fraud" cases. In addition, terrorist groups could prompt LLMs to help them generate text to promote and defend disinformation and fake news, lending false credibility to their propaganda campaigns, Europol says.

Worst, perhaps, is that the code-writing capabilities of LLMs could be misused by criminals with "little to no knowledge of coding" to write malware or ransomware. "Critically, the safeguards preventing ChatGPT from providing potentially malicious code only work if the model understands what it is doing. If prompts are broken down into individual steps, it is trivial to bypass these safety measures," Europol says in its report.

Of particular worry, says the organization, is that LLMs are far from a done deal: they are constantly being improved,  so their potential criminal exploitation could happen ever-faster and at greater scale.

 "The Europol report seems exactly correct. I agree things look grim," says Gary Marcus, a professor of psychology and neural science at New York University, and an AI entrepreneur and commentator. "Perhaps coupled with mass AI-generated propaganda, LLM-enhanced terrorism could in turn lead to nuclear war, or to the deliberate spread of pathogens worse than Covid-19," Marcus later said in his newsletter.

OpenAI did not respond to questions on Europol's findings, and neither did the U.S. Cybersecurity and Infrastructure Security Agency (CISA), part of the Department for Homeland Security.

However, two days later, on March 29, the AI fightback moved up another notch, when Marcus was one of more than 1,000 initial signatories to an open letter to AI labs calling on them to "immediately pause for at least 6 months the training of AI systems more powerful than GPT-4".

Drafted by the Future of Life Institute, in Cambridge, MA, which campaigns against technologies posing existential risks, the letter urged that "AI labs and independent experts should use this pause to jointly develop and implement a set of shared safety protocols for advanced AI design and development that are rigorously audited and overseen by independent outside experts."

"These protocols should ensure that systems adhering to them are safe beyond a reasonable doubt. This does not mean a pause on AI development in general, merely a stepping back from the dangerous race to ever-larger unpredictable black-box models with emergent capabilities."

The letter was signed by some of the leading specialists in AI, including deep neural networking pioneer (and ACM A.M. Turing Award recipient) Yoshua Bengio of the Quebec AI Institute (MILA) in Montreal, Canada, and Stuart Russell, head of the Center for Human Compatible AI at the University of California, Berkeley. ... '   (Much more at the link) 

Tuesday, January 10, 2023

UC Berkely Does Skypilot

 Interesting direction, 

UC Berkeley Launches SkyPilot to Help Navigate Soaring Cloud Costs

By Jaime Hampton, Datanami

Runaway cloud computing costs can stifle machine learning and data science projects, and many organizations are using multiple public clouds for different purposes to save money. However, a multi-cloud approach can add significant complexity, since not everyone is a cloud infrastructure expert.

To address this, researchers at U.C. Berkeley’s Sky Computing Lab have launched SkyPilot, an open source framework for running ML and Data Science batch jobs on any cloud, or multiple clouds, with a single cloud-agnostic interface.

SkyPilot uses an algorithm to determine which cloud zone or service provider is the most cost-effective for a given project. The program considers a workload’s resource requirements (whether it needs CPUs, GPUs, or TPUs) and then automatically determines which locations (zone/region/cloud) have available compute resources to complete the job before sending it to the least expensive option to execute.

The solution automates some of the more challenging aspects of running workloads on the cloud. SkyPilot’s makers say the program can reliably provision a cluster with automatic failover to other locations if capacity or quota errors occur, it can sync user code and files from local or cloud buckets to the cluster, and it can manage job queueing and execution. The researchers claim this comes with substantially reduced costs, sometimes by more than 3x.

SkyPilot developer and postdoctoral researcher Zongheng Yang said in a blog post that the growing trend of multi-cloud and multi-region strategies led the team to build SkyPilot, calling it an “intercloud broker.” He notes that organizations are strategically choosing a multi-cloud approach for higher reliability, avoiding cloud vendor lock-in, and stronger negotiation leverage, to name a few reasons.

To save costs, SkyPilot leverages the large price differences between cloud providers for similar hardware resources. Yang gives the example of Nvidia A100 GPUs, and how Azure currently offers the cheapest A100 instances, but Google Cloud and AWS charge a premium of 8% and 20% for the same computing power. For CPUs, some price differences can be over 50%.

Specialized hardware is also a reason to shop around, as many cloud providers are now offering custom options for different workloads. For example, Google Cloud offers TPUs for ML training, AWS has Inferentia for ML inference and Graviton processors for CPU workloads, and Azure provides Intel SGX codes for confidential computing. Scarcity of these specialized resources is also a reason for using multiple clouds, as high-end GPUs are frequently unavailable with long wait times.  ... 


Monday, December 12, 2022

ESnet Launches Next-Generation Network to Enhance Collaborative Science

Better Networks Built for Research.

ACM TECHNEWS

ESnet Launches Next-Generation Network to Enhance Collaborative Science

By Berkeley Lab News Center, October 18, 2022

“ESnet6 represents a transformational change in the way networks are built for research, with improved capacity, resiliency, and flexibility,” said ESnet executive director Inder Monga.

The Energy Sciences Network (ESnet) has rolled out the latest generation of the U.S. Department of Energy's (DOE) high-performance science network, dubbed ESnet6.

ESnet links all DOE's national laboratories, DOE-funded researchers, and its scientific instruments and supercomputing centers so that data can move quickly between them.

ESnet6 features bandwidth of over 46 terabits per second, a dedicated fiber optic cable footprint spanning 15,000 miles, and network backbone links ranging from 400 gigabits per second to 1 terabit per second, among other things.

Said ESnet's Inder Monga, "ESnet6 provides the foundation for the future of the DOE mission science as we enter an age where discoveries will rely on the integration of scientific experimental facilities, supercomputers, and global science teams operating together as if they are colocated: one instrument in one location. ESnet6 interconnects all of these resources to create a holistic science discovery system."

From Berkeley Lab News Center

View Full Article     


Monday, November 21, 2022

Low Cost, Omni Purpose Robotics

Robust Robotics Solutions 

Low-Cost Robot Ready for Any Obstacle

Carnegie Mellon University News

Aaron Aupperle, November 16, 2022

Scientists at Carnegie Mellon University (CMU) and the University of California, Berkeley, have enabled a low-cost and relatively small legged robot to adapt to obstacles. The robot uses its vision and an onboard computer to quickly adjust to new situations and master difficult terrain. The researchers trained it using 4,000 robot clones as they walked and climbed in a simulator, giving the machine six years of experience in one day. The simulator also retained motor skills acquired in training in a neural network that the team copied to the actual robot. "This system uses vision and feedback from the body directly as input to output commands to the robot's motors," explained CMU's Ananye Agarwal. "This technique allows the system to be very robust in the real world."  ... ' 

Saturday, November 19, 2022

Low Cost Legged Robotics

 Low cost,  especially useful for testing out proposed uses,  sounds good. 

Low-Cost Robot Ready for Any Obstacle   

By Carnegie Mellon University News, November 18, 2022

A robotic system designed by researchers at Carnegie Mellon University's School of Computer Science and the University of California, Berkeley, enables small, low-cost legged robots to maneuver in challenging environments.

Scientists at Carnegie Mellon University (CMU) and the University of California, Berkeley, have enabled a low-cost and relatively small legged robot to adapt to obstacles.

The robot uses its vision and an onboard computer to quickly adjust to new situations and master difficult terrain.  The researchers trained it using 4,000 robot clones as they walked and climbed in a simulator, giving the machine six years of experience in one day.  The simulator also retained motor skills acquired in training in a neural network that the team copied to the actual robot.

"This system uses vision and feedback from the body directly as input to output commands to the robot's motors," explained CMU's Ananye Agarwal. "This technique allows the system to be very robust in the real world."

From Carnegie Mellon University News

View Full Article    

Wednesday, August 31, 2022

Robot Dogs Learning Tough Terrain

Been following this for some time, and potential uses. 

Robot Dog Learns to Walk Tough Terrain in 20 Minutes

New Scientist, Alex Wilkins,  August 26, 2022

Researchers at the University of California, Berkeley (UC Berkeley) developed a machine learning algorithm that enabled a robot dog to learn to navigate difficult terrain in only 20 minutes. The Q-learning algorithm does not need a model of the target terrain. As a result, said UC Berkeley's Sergey Levine, "We don't need to understand how the physics of an environment actually works, we just put the robot into an environment and turn it on." The algorithm teaches the robot by rewarding it for each successful action until reaching its ultimate goal. The researchers demonstrated that the robot was able to walk on terrains it had not previously encountered, including grass, a layer of bark, a memory foam mattress, and a hiking trail, after about 20 minutes of training on each. ... ' 

Tuesday, November 09, 2021

Neural Network Representations

Interesting, technical.  See link for useful supporting visuals.

How should we compare neural network representations?

Frances Ding and Jacob Steinhardt    Nov 8, 2021  Berkeley 

Cross-posted from Bounded Regret.  

To understand neural networks, researchers often use similarity metrics to measure how similar or different two neural networks are to each other. For instance, they are used to compare vision transformers to convnets [1], to understand transfer learning [2], and to explain the success of standard training practices for deep models [3]. Below is an example visualization using similarity metrics; specifically we use the popular CKA similarity metric (introduced in [4]) to compare two transformer models across different layers:

Figure 1. CKA (Centered Kernel Alignment) similarity between two networks trained identically except for random initialization. Lower values (darker colors) are more similar. CKA suggests that the two networks have similar representations.

Unfortunately, there isn’t much agreement on which particular similarity metric to use. Here’s the exact same figure, but produced using the Canonical Correlation Analysis (CCA) metric instead of CKA:

Figure 2. CCA (Canonical Correlation Analysis) similarity between the same two networks. CCA distances suggest that the two networks learn somewhat different representations, especially at later layerss.

In the literature, researchers often propose new metrics and justify them based on intuitive desiderata that were missing from previous metrics. For example, Morcos et al. motivate CCA by arguing that similarity metrics should be invariant to invertible linear transformations [5]. Kornblith et al. disagree about which invariances a similarity metric should have, and instead argue that metrics should pass an intuitive test - given two trained networks with the same architecture but different initialization, layers at the same depth should be most similar to each other - and their proposed metric, CKA, performs the best on their test [4].

Our paper, Grounding Representation Similarity with Statistical Testing, argues against this practice. To start, we show that by choosing different intuitive tests, we can make any method look good. CKA does well on a “specificity test” similar to the one proposed by Kornblith et al., but it does poorly on a “sensitivity test” that CCA shines on.

To move beyond intuitive tests, our paper provides a carefully-designed quantitative benchmark for evaluting similarity metrics. The basic idea is that a good similarity metric should correlate with the actual functionality of a neural network, which we operationalize as accuracy on a task. Why? Accuracy differences between models are a signal that the models are processing data differently, so intermediate representations must be different, and similarity metrics should notice this.

Thus, for a given pair of neural network representations, we measure both their (dis)similarity and the difference between their accuracies on some task. If these are well-correlated across many pairs of representations, we have a good similarity metric. Of course, a perfect correlation with accuracy on a particular task also isn’t what we’re hoping for, since metrics should capture many important differences between models, not just one. A good similarity metric is one that gets generally high correlations across a couple of functionalities.

We assess functionality with a range of tasks. For a concrete example, one subtask in our benchmark builds off the observation that BERT language models finetuned with different random seeds will have nearly identical in-distribution accuracy, but widely varying out-of-distribution accuracy (for example, ranging from 0 to 60% on the HANS dataset [6]). Given two robust models, a similarity metric should rate them as similar, and given one robust and one non-robust model, a metric should rate them as dissimilar. Thus we take 100 such BERT models and evaluate whether (dis)similarity between each pair of model representations correlates with their difference in OOD accuracy.  ..... ' 


Saturday, October 09, 2021

Compressing Data for Humans in the Loop

Compressing data for teleoperation. Human in the loop requires images that can be used to adapt to changing reaction times and decision making in coordination with very remote devices.  Technical. 

 PICO: Pragmatic Compression for Human-in-the-Loop Decision-Making   by Siddharth Reddy,    Berkeley Bair

Imagine remotely operating a Mars rover from a desk on Earth. The low-bandwidth network connection can make it challenging for the teleoperation system to provide the user with high-dimensional observations like images. One approach to this problem is to use data compression to minimize the number of bits that need to be communicated over the network: for example, the rover can compress the pictures it takes on Mars before sending them to the human operator on Earth. Standard lossy image compression algorithms would attempt to preserve the image's appearance. However, at low bitrates, this approach can waste precious bits on information that the user does not actually need in order to perform their current task. For example, when deciding where to steer and how much to accelerate, the user probably only pays attention to a small subset of visual features, such as obstacles and landmarks. .... '

Tuesday, July 13, 2021

Training Game Agents with ML

From the Google AI Blog.  Below just the intro.  Something we proposed wayback, but now have seem several interesting examples.  Certainly you can consider any interaction with data as a 'game' with goals.  So have the game be trained for outcome achievements based upon relevant data and evolving contexts.   Here might be a useful start.     Berkeley is also doing things of interest. 

Quickly Training Game-Playing Agents with Machine Learning

Tuesday, June 29, 2021

Posted by Leopold Haller and Hernan Moraldo, Software Engineers, Google Research

In the last two decades, dramatic advances in compute and connectivity have allowed game developers to create works of ever-increasing scope and complexity. Simple linear levels have evolved into photorealistic open worlds, procedural algorithms have enabled games with unprecedented variety, and expanding internet access has transformed games into dynamic online services. Unfortunately, scope and complexity have grown more rapidly than the size of quality assurance teams or the capabilities of traditional automated testing. This poses a challenge to both product quality (such as delayed releases and post-launch patches) and developer quality of life.   ... ' 

Friday, July 09, 2021

Challenge for Learning from Human Feedback using Minecraft

Berkeley Bair challenge competition here using a common gaming environment.   Been a long time since I looked at Minecraft.  Short extract of the idea below, more complete look at the link.   Seems a novel look at a broader look at contextual learning. 

 BASALT: A Benchmark for  Learning from Human Feedback   by Rohin Shah    Jul 8, 2021

TL;DR: We are launching a NeurIPS competition and benchmark called BASALT: a set of Minecraft environments and a human evaluation protocol that we hope will stimulate research and investigation into solving tasks with no pre-specified reward function, where the goal of an agent must be communicated through demonstrations, preferences, or some other form of human feedback. Sign up to participate in the competition!

Motivation

Deep reinforcement learning takes a reward function as input and learns to maximize the expected total reward. An obvious question is: where did this reward come from? How do we know it captures what we want? Indeed, it often doesn’t capture what we want, with many recent examples showing that the provided specification often leads the agent to behave in an unintended way.

Our existing algorithms have a problem: they implicitly assume access to a perfect specification, as though one has been handed down by God. Of course, in reality, tasks don’t come pre-packaged with rewards; those rewards come from imperfect human reward designers.

For example, consider the task of summarizing articles. Should the agent focus more on the key claims, or on the supporting evidence? Should it always use a dry, analytic tone, or should it copy the tone of the source material? If the article contains toxic content, should the agent summarize it faithfully, mention that toxic content exists but not summarize it, or ignore it completely? How should the agent deal with claims that it knows or suspects to be false? A human designer likely won’t be able to capture all of these considerations in a reward function on their first try, and, even if they did manage to have a complete set of considerations in mind, it might be quite difficult to translate these conceptual preferences into a reward function the environment can directly calculate.  ...................

Conclusion

We hope that BASALT will be used by anyone who aims to learn from human feedback, whether they are working on imitation learning, learning from comparisons, or some other method. It mitigates many of the issues with the standard benchmarks used in the field. The current baseline has lots of obvious flaws, which we hope the research community will soon fix.

Note that, so far, we have worked on the competition version of BASALT. We aim to release the benchmark version shortly. You can get started now, by simply installing MineRL from pip and loading up the BASALT environments. The code to run your own human evaluations will be added in the benchmark release.

If you would like to use BASALT in the very near future and would like beta access to the evaluation code, please email the lead organizer, Rohin Shah, at rohinmshah@berkeley.edu.

This post is based on the paper “The MineRL BASALT Competition on Learning from Human Feedback”, accepted at the NeurIPS 2021 Competition Track. Sign up to participate in the competition!   

Thursday, January 28, 2021

New AI to Reason about Uncertainty

Good thoughts here.   We spend muct time trying to reduce uncertainty to something useful, like risk, both within rule based logic and prediction.

Approach to AI Offers More Certainty in the Face of Uncertainty, By Radboud University (Netherlands), January 28, 2021   Technical 

Researchers at the Netherlands' Radboud University and Eindhoven University of Technology, and the Universities of Austin and California, Berkeley, have formulated a new artificial intelligence (AI) method for reasoning about uncertainty.

The uncertain partially observable Markov decision processes (uPOMDPs) are basically real-world models that calculate event probabilities, so AI can make better and safer decisions faster.

Radboud's Nils Jansen said the uPOMDP approach "allows us to take all our calculations and theoretical information and use it in the real world on a more consistent, regular basis."

He added that systes like autonomous cars could use this method to explain errors in more detail, in order to account for them when calculating. Said Jansen, "This means they have more specific examples of what could go wrong, and make better and more adequate adjustments to avoid those specific risks."

From Radboud University (Netherlands) ... 

Thursday, January 14, 2021

Offline Reinforcement Learning

Technical, but relatively understandable, 

Offline Reinforcement Learning: How Conservative Algorithms Can Enable New Applications

Aviral Kumar and Avi Singh    Dec 7, 2020  in BAIR Berkeley

Deep reinforcement learning has made significant progress in the last few years, with success stories in robotic control, game playing and science problems. While RL methods present a general paradigm where an agent learns from its own interaction with an environment, this requirement for “active” data collection is also a major hindrance in the application of RL methods to real-world problems, since active data collection is often expensive and potentially unsafe. An alternative “data-driven” paradigm of RL, referred to as offline RL (or batch RL) has recently regained popularity as a viable path towards effective real-world RL. As shown in the figure below, offline RL requires learning skills solely from previously collected datasets, without any active environment interaction. It provides a way to utilize previously collected datasets from a variety of sources, including human demonstrations, prior experiments, domain-specific solutions and even data from different but related problems, to build complex decision-making engines. .... ' 


Thursday, January 07, 2021

Mining Private Data: Does GPT-2 Know Your Phone Number?

Schneier has mentioned this as a privacy issue: 

Just reading the Berkeley BAIR paper that develops this  (see short excerpt below)   

And the source technical paper:   Extracting Training Data from Large Language Models

By Nicholas Carlini, Florian Tramer, Eric Wallace, Matthew Jagielski, Ariel Herbert-Voss, Katherine Lee, Adam Roberts, Tom Brown, Dawn Song, Ulfar Erlingsson, Alina Oprea, Colin Raffel

It has become common to publish large (billion parameter) language models that have been trained on private datasets. This paper demonstrates that in such settings, an adversary can perform a training data extraction attack to recover individual training examples by querying the language model.

We demonstrate our attack on GPT-2, a language model trained on scrapes of the public Internet, and are able to extract hundreds of verbatim text sequences from the model's training data. These extracted examples include (public) personally identifiable information (names, phone numbers, and email addresses), IRC conversations, code, and 128-bit UUIDs. Our attack is possible even though each of the above sequences are included in just one document in the training data.

We comprehensively evaluate our extraction attack to understand the factors that contribute to its success. For example, we find that larger models are more vulnerable than smaller models. We conclude by drawing lessons and discussing possible safeguards for training large language models.

Subjects: Cryptography and Security (cs.CR); Computation and Language (cs.CL); Machine Learning (cs.LG)   ..... 

Does GPT-2 Know Your Phone Number?

Berkeley BAIR: Eric Wallace, Florian Tramèr, Matthew Jagielski, and Ariel Herbert-Voss   Dec 20, 2020

Most likely not.

Yet, OpenAI’s GPT-2 language model does know how to reach a certain Peter W--- (name redacted for privacy). When prompted with a short snippet of Internet text, the model accurately generates Peter’s contact information, including his work address, email, phone, and fax:

In our recent paper, we evaluate how large language models memorize and regurgitate such rare snippets of their training data. We focus on GPT-2 and find that at least 0.1% of its text generations (a very conservative estimate) contain long verbatim strings that are “copy-pasted” from a document in its training set.

Such memorization would be an obvious issue for language models that are trained on private data, e.g., on users’ emails, as the model might inadvertently output a user’s sensitive conversations. Yet, even for models that are trained on public data from the Web (e.g., GPT-2, GPT-3, T5, RoBERTa, TuringNLG), memorization of training data raises multiple challenging regulatory questions, ranging from misuse of personally identifiable information to copyright infringement.

Monday, December 07, 2020

Moving AI Recognized Things

Interesting generalization of a set of common tasks using AI identification.

Robotics Researchers Propose AI That Locates, Safely Moves Items on Shelves

Venture Beat   By Kyle Wiggers in CACM

Two new robotics studies detail methods for locating occluded objects on shelves and solving "contact-rich" manipulation tasks. Researchers at the University of California, Berkeley developed the Lateral Access maXimal Reduction of occupancY support Area (LAX-RAY) system, which predicts an object's location even when only a portion of it is visible. LAX-RAY achieved 87.3% accuracy in a simulation, which translated to about 80% for a real-world robot. Meanwhile, Google developed the Contact-aware Online COntext Inference (COCOI), which uses video footage and readings from a robot-mounted touch sensor to encode dynamics information into a representation, which then permits a reinforcement learning algorithm to plan with “dynamics-awareness,” increasing its robustness in difficult environments.  ... ' 

Tuesday, November 24, 2020

Capturing Words from Silent Speech

An Indication of how AI can become very powerful, even replacing the need for actual speech with mouthed muscle activity.  A kind of pattern recognition.

UC Berkeley researchers detect ‘silent speech’ with electrodes and AI

Khari Johnson  @kharijohnson in VentureBeat

UC Berkeley researchers say they are the first to train AI using using silently mouthed words and sensors that collect muscle activity. Silent speech is detected using electromyography (EMG), with electrodes placed on the face and throat. The model focuses on what researchers call digital voicing to predict words and generate synthetic speech.

Researchers believe their method can enable a number of applications for people who are unable to produce audible speech and could support speech detection for AI assistants or other devices that respond to voice commands.  ... " 

Sunday, November 22, 2020

Learning Long Horizon Planning

 Another interesting piece from Berkeley - BAIR. Now looking at more complex planning.  Long range planning that machines are not necessarily good at.   But humans are also not so good at problems that require many option planning and analysis.  COuld this be something where machines and humans could collaborate well.    Had some supply chain planning models that might have used these directions.  Byond the intro this article is technical.

Learning State Abstractions for Long-Hoprizon Planning

By Scott Emmons*, Ajay Jain*, Michael Laskin*, Thanard Kurutach, Pieter Abbeel, Deepak Pathak 

Many tasks that we do on a regular basis, such as navigating a city, cooking a meal, or loading a dishwasher, require planning over extended periods of time. Accomplishing these tasks may seem simple to us; however, reasoning over long time horizons remains a major challenge for today’s Reinforcement Learning (RL) algorithms. While unable to plan over long horizons, deep RL algorithms excel at learning policies for short horizon tasks, such as robotic grasping, directly from pixels. At the same time, classical planning methods such as Dijkstra’s algorithm and A∗ search can plan over long time horizons, but they require hand-specified or task-specific abstract representations of the environment as input.

To achieve the best of both worlds, state-of-the-art visual navigation methods have applied classical search methods to learned graphs. In particular, SPTM [2] and SoRB [3] use a replay buffer of observations as nodes in a graph and learn a parametric distance function to draw edges in the graph. These methods have been successfully applied to long-horizon simulated navigation tasks that were too challenging for previous methods to solve.  ... "