/* ---- Google Analytics Code Below */
Showing posts with label autoML. Show all posts
Showing posts with label autoML. Show all posts

Tuesday, March 01, 2022

Automating Data Science: Prospects and Challenges

 Long, insightful piece, good insights,  somewhat technical, but useful.   A favorite topic we experimented with since very early days.

Home/Magazine Archive/March 2022 (Vol. 65, No. 3)/Automating Data Science/Full Text

 By Tijl De Bie, Luc De Raedt, José Hernández-Orallo, Holger H. Hoos, Padhraic Smyth, Christopher K. I. Williams

Communications of the ACM, March 2022, Vol. 65 No. 3, Pages 76-87 10.1145/3495256

--> Given the complexity of typical data science projects and the associated demand for human expertise, automation has the potential to transform the data science process.

Key insights: 

• Automation in data science aims to facilitate and transform the work of data scientists, not to replace them.

• Important parts of data science are already being automated, especially in the modeling stages, where techniques such as automated machine learning (AutoML) are gaining traction.

• Other aspects are harder to automate, not only because of technological challenges, but because open ended and context-dependent tasks require human interaction.

Introduction

Data science covers the full spectrum of deriving insight from data, from initial data gathering and interpretation, via processing and engineering of data, and exploration and modeling, to eventually producing novel insights and decision support systems. Data science can be viewed as overlapping or broader in scope than other data-analytic methodological disciplines, such as statistics, machine learning, databases, or visualization

To illustrate the breadth of data science, consider, for example, the problem of recommending items (movies, books or other products) to customers. While the core of these applications can consist of algorithmic techniques such as matrix factorization, a deployed system will involve a much wider range of technological and human considerations. These range from scalable back-end transaction systems that retrieve customer and product data in real time, experimental design for evaluating system changes, causal analysis for understanding the effect of interventions, to the human factors and psychology that underliehow customers react to visual information displays and make decisions.

As another example, in areas such as astronomy, particle physics, and climate science, there is a rich tradition of building computational pipelines to support data-driven discovery and hypothesis testing. For instance, geoscientists use monthly global landcover maps based on satellite imagery at sub-kilometer resolutions to better understand how the earth’s surface is changing over time [50]. These maps are interactive and browsable, and they are the result of a complex data-processing pipeline, in which terabytes to petabytes of raw sensor and image data are transformed into databases of automatically detected and annotated objects and information. This type of pipeline involves many steps, in which human decisions and insight are critical, such as instrument calibration, removal of outliers, and classification of pixels.

The breadth and complexity of these and many other data science scenarios means that the modern data scientist requires broad knowledge and experience across a multitude of topics. Together with an increasing demand for data analysis skills, this has led to a shortage of trained data scientists with appropriate background and experience, and significant market competition for limited expertise. Considering this bottleneck, it is not surprising that there is increasing interest in automating parts, if not all, of the data science process. This desire and potential for automation is the focus of this article.

As illustrated in the examples above, data science is a complex process, driven by the character of the data being analyzed and by the questions being asked, and is often highly exploratory and iterative in nature. Domain context can play a key role in these exploratory steps, even in relatively well-defined processes such as predictive modeling (e.g., as characterized by CRISP-DM [5]) where, for example, human expertise in defining relevant predictor variables can be critical.  .... '

Tuesday, April 20, 2021

MBRL Tuning for Partially Understood Environments

Below is very technical,  but I do like some of the background statements such as 'solving tasks in a partially understood environment ...'.    And the idea of optimizing agents to resolve elements of understanding.  (Which exemplifies the situations we are often in).    So I am not saying I understand this yet, but working through it now for broader application.  As part of my broader study of practical reinforcement learning.

The Importance of Hyperparameter Optimization for Model-based Reinforcement Learning

Nathan Lambert, Baohe Zhang, Raghu Rajan, André Biedenkapp    Apr 19, 2021  From BAIR  Berkeley

Model-based reinforcement learning (MBRL) is a variant of the iterative learning framework, reinforcement learning, that includes a structured component of the system that is solely optimized to model the environment dynamics. Learning a model is broadly motivated from biology, optimal control, and more – it is grounded in natural human intuition of planning before acting. This intuitive grounding, however, results in a more complicated learning process. In this post, we discuss how model-based reinforcement learning is more susceptible to parameter tuning and how AutoML can help in finding very well performing parameter settings and schedules. Below, left is the expected behavior of an agent maximizing velocity on a “Half Cheetah” robotic task, and to the right is what our paper with hyperparameter tuning finds.

MBRL

Model-based reinforcement learning (MBRL) is an iterative framework for solving tasks in a partially understood environment. There is an agent that repeatedly tries to solve a problem, accumulating state and action data. With that data, the agent creates a structured learning tool – a dynamics model – to reason about the world. With the dynamics model, the agent decides how to act by predicting into the future. With those actions, the agent collects more data, improves said model, and hopefully improves future actions.  ... " 

Monday, April 05, 2021

B to B Machine Learning

Type of companies and their goals and implications.

Defensible Machine Learning    via O'Reilly  in Basecase.vc

By Ankur Goyal, Alana Anderson  March 3, 2021

B2B machine learning (ML) companies are an enigma: they have the opportunity to revolutionize how we do business, but they look & feel quite different from their traditional SaaS counterparts and have proven difficult to scale. In this post, we aim to demystify the challenges of building an enduring ML company by providing a simple framework for the three ways to do so. For each, we outline common attributes of successful companies and walk through potential pitfalls. These learnings are based on our experience building Impira and evaluating companies at Base Case Capital.

For aspiring entrepreneurs thinking about problems in this area, we hope this gives you a starting point to understand the landscape, frame your idea, and avoid the challenges that may arise.

Three types of ML companies:

Enterprises who seek to solve problems with machine learning have two options: (1) develop ML models in-house or (2) purchase software that has embedded ML. Startups can play a significant role in both scenarios. The first option involves providing in-house users, traditionally referred to as ML practitioners, with ML systems that enable them to work better, faster, and smarter to develop and operate models. The second option is usually packaged as a Powered-by-ML solution, which solves a specific business challenge with the help of machine learning.

There is an emerging third group, AutoML solutions, which allow non-ML practitioners to develop their own models. This is somewhat a hybrid of the first two, and correspondingly represents a massive opportunity, but with significant product and technical risks. In fact, many of the common pitfalls for both ML systems and Powered-by-ML solutions can be overcome with AutoML. We’ll discuss this in depth below. .." 

Friday, December 11, 2020

Google AI Describes AutoML for Time Series

 Most our careers in the big enterprise involved working with time series.   Sales, Shipments delivered, Advertising Dollars,  marketing spends ... forecast plans and predictions.   A favorite quote was 'the forecast is wrong', but how wrong?   And Why?  And what are the risks involved?  So if we could do forecasts better, more data and intelligence based?  How might we do it?     

Using AutoML for Time Series Forecasting      In the GoogleBlog.

Friday, December 4, 2020

Posted by Chen Liang and Yifeng Lu, Software Engineers, Google Research, Brain Team

Time series forecasting is an important research area for machine learning (ML), particularly where accurate forecasting is critical, including several industries such as retail, supply chain, energy, finance, etc. For example, in the consumer goods domain, improving the accuracy of demand forecasting by 10-20% can reduce inventory by 5% and increase revenue by 2-3%. Current ML-based forecasting solutions are usually built by experts and require significant manual effort, including model construction, feature engineering and hyper-parameter tuning. However, such expertise may not be broadly available, which can limit the benefits of applying ML towards time series forecasting challenges.

To address this, automated machine learning (AutoML) is an approach that makes ML more widely accessible by automating the process of creating ML models, and has recently accelerated both ML research and the application of ML to real-world problems. For example, the initial work on neural architecture search enabled breakthroughs in computer vision, such as NasNet, AmoebaNet, and EfficientNet, and in natural language processing, such as Evolved Transformer. More recently, AutoML has also been applied to tabular data.

Today we introduce a scalable end-to-end AutoML solution for time series forecasting, which meets three key criteria:

Today we introduce a scalable end-to-end AutoML solution for time series forecasting, which meets three key criteria:

Fully automated: The solution takes in data as input, and produces a servable TensorFlow model as output with no human intervention.

Generic: The solution works for most time series forecasting tasks and automatically searches for the best model configuration for each task.

High-quality: The produced models have competitive quality compared to those manually crafted for specific tasks.

We demonstrate the success of this approach through participation in the M5 forecasting competition, where this AutoML solution achieved competitive performance against hand-crafted models with moderate compute cost... .' 

Saturday, November 28, 2020

A New, Free, MIT Introduction to Machine Learning AI Course

We worked with MIT in a number of ways, with their supply chain group,  their Media Laboratory and with a number of their researchers.  Impressive group.  I still provide publicity for some of their research. We introduced early AI to the enterprise.  I have now been asked to provide a review of their latest. A new, Free Introduction to Machine Learning AI Course.    See more more about it below.   If you also take it please provide some of your own thoughts about it to me.  - Franz  

MIT Released a New, Free Machine Learning Course

Don’t miss this!  By Frederik Bussler from Medium

MIT Open Learning Library just released an Introduction to Machine Learning course    . The free 13-week course covers machine learning algorithms, supervised and reinforcement learning, and more.

While you don’t need these skills to deploy AI, given no-code AI companies like Obviously.AI   , it can help if you want to work on the cutting-edge and build new architectures.

Why MIT?

MIT is a world-famous organization, and for a good reason. This prestigious university boasts graduates like Buzz Aldrin, the second person to walk on the moon (fun fact: Buzz Lightyear was named after Buzz Aldrin).

You’ll be part of a community of learners. Dreamers. Doers. Indeed, part of the Open Learning Library’s mission is “Extending MIT’s knowledge to the world.” You get all this for free.  ....

Wednesday, August 05, 2020

Opinions on State of the Art of Automated Machine Learning

Very useful piece here,  the individual comments at the link are most interesting, the key takeaways are mostly obvious. 

Anthony Alford, Francesca Lazzeri

Key Takeaways:

-Automated Machine Learning (AutoML) is important because it allows data scientists to save time and resources, delivering business value faster and more efficiently
-AutoML is not likely to remove the need for a "human in the loop" for industry-specific knowledge and translating the business problem into a machine learning problem
-Some important research topics in the area are feature engineering, model transparency, and addressing bias
-There are several commercial and open-source AutoML solutions available now for automating different parts of the machine learning process
-Some limitations of AutoML are the amount of computational resources required and the needs of domain-specific applications  ... "  

(Below at the link lots of individual, often technical comments on the progress and state of automation.  Well worth scanning at least)  ...

Tuesday, July 14, 2020

On Automated Machine Learning

Emphasizing the automated, inevitable that such method will be more broadly integrated with general IT analytics.  But will also require automated updating of their use in context.

AutoML: Not A Magic Bullet, But A Powerful Business Tool  by 7wData

When AI was first introduced into business processes, it was transformative, enabling companies to leverage the vast amounts of accumulated data to improve planning and decision making. It soon became apparent, however, that integrating AI into business processes at scale required significant resources. First, companies had to recruit highly sought-after (and highly paid) data scientists to create the data models behind AI. Second, the process of building and training the machine learning models that accelerated the data analysis process required a significant expenditure of time and energy. This, in turn, led to the development of automated machine learning (AutoML), techniques that essentially automate core aspects of the machine learning process including model selection, training, and evaluation.

In effect, AutoML seeks to trade machine (processing) time for human time. This automation brings many benefits. First and foremost, it decreases labor costs. It also reduces human error, automates repetitive tasks, and enables the development of more effective models. By reducing the technical expertise required to create an ML model, AutoML also lowers the barriers to entry, enabling business analysts to leverage advanced modeling techniques — without assistance from data scientists. And by relieving data scientists from repetitive tasks of the machine learning process, AutoML frees these costly resources to pursue higher-value projects.  ... " 

Saturday, February 29, 2020

Replacing Data Scientists With AutoML?

Excerpt from a current KDNuggets article, which links further to a poll that asks the questions of practitioners.  My answer is yes. AutoML will replace the current needs for data science analysis.  Within a decade.  Of course the needs are likely to expand as well, so there will always be research and new requirements emerging.  And interpretation for specific context needs.  Just as there are needs for statisticians and analytics specialists for the same purposes.

When Will AutoML (Automated Machine Learning) Replace Data Scientists (if ever)?

Soon after tech giants Google and Microsoft introduced their AutoML services to the world, the popularity and interest in these services skyrocketed. We first review AutoML, compare the platforms available, and then test them out against real data scientists to answer the question: will AutoML replace us?

Introduction of AutoML:
One cannot introduce AutoML without mentioning the machine learning project’s life cycle, which includes data cleaning, feature selection/engineering, model selection, parameter optimization, and finally, model validation. As advanced as technology has become, the traditional data science project still incorporates a lot of manual processes and remains time-consuming and repetitive. ... "

Tuesday, June 18, 2019

More on MIT Open Source AutoML: ATMSeer

More on the topic, and some additional background information and links.  Very powerful concept that that should continue to expand.  Automation is the word,   See also Google AutoML, at tag below.

MIT Researchers Open-Source AutoML Visualization Tool ATMSeer    by  Anthony Alford   ... 

A research team from MIT, Hong Kong University, and Zhejiang University has open-sourced ATMSeer, a tool for visualizing and controlling automated machine-learning processes.
Solving a problem with machine learning (ML) requires more than just a dataset and training. For any given ML tasks, there are a variety of algorithms that could be used, and for each algorithm there can be many hyperparameters that can be tweaked. Because different values of hyperparameters will produce models with different accuracies, ML practitioners usually try out several sets of hyperparameter values on a given dataset to try to find hyperparameters that produce the best model. 

This can be time-consuming, as a separate training job and model evaluation process must be conducted for each set. Of course, they can be run in parallel, but the jobs must be setup and triggered, and the results recorded. Furthermore, choosing the particular values for hyperparameters can involve a bit of guesswork, especially for ones that can take on any numeric value: if 2.5 and 2.6 produce good results, maybe 2.55 would be even better? What about 2.56 or 2.54?

Enter automated machine learning, or AutoML. These are techniques and tools for automating the selection and evaluation of hyperparameters (as well as other common ML tasks such as data cleanup and feature engineering). Both Google Cloud Platform and Microsoft Azure provide commercial AutoML solutions, and there are several open-source packages such as auto-sklearn and Auto-Keras.  ...."

Friday, May 31, 2019

Advances in Automated Machine Learning

Automated machine learning is inevitable.  How good will it be, and how much human oversight needs to be applied to ensure confidence in their results is important.  This article is a  good overview of work underway.  Its no only about searching for the right model,  in context of needed goals, its also about maintaining the interaction between data and solutions.   Not too unlike the use of any kind of analytics optimization, which has been studied for years.

Cracking open the black box of automated machine learning
Interactive tool lets users see and control how automated model searches work.
 By Rob Matheson | MIT News Office 

Researchers from MIT and elsewhere have developed an interactive tool that, for the first time, lets users see and control how automated machine-learning systems work. The aim is to build confidence in these systems and find ways to improve them.

Designing a machine-learning model for a certain task — such as image classification, disease diagnoses, and stock market prediction — is an arduous, time-consuming process. Experts first choose from among many different algorithms to build the model around. Then, they manually tweak “hyperparameters” — which determine the model’s overall structure — before the model starts training.

Recently developed automated machine-learning (AutoML) systems iteratively test and modify algorithms and those hyperparameters, and select the best-suited models. But the systems operate as “black boxes,” meaning their selection techniques are hidden from users. Therefore, users may not trust the results and can find it difficult to tailor the systems to their search needs.

In a paper presented at the ACM CHI Conference on Human Factors in Computing Systems, researchers from MIT, the Hong Kong University of Science and Technology (HKUST), and Zhejiang University describe a tool that puts the analyses and control of AutoML methods into users’ hands. Called ATMSeer, the tool takes as input an AutoML system, a dataset, and some information about a user’s task. Then, it visualizes the search process in a user-friendly interface, which presents in-depth information on the models’ performance.

“We let users pick and see how the AutoML systems works,” says co-author Kalyan Veeramachaneni, a principal research scientist in the MIT Laboratory for Information and Decision Systems (LIDS), who leads the Data to AI group. “You might simply choose the top-performing model, or you might have other considerations or use domain expertise to guide the system to search for some models over others.”

In case studies with science graduate students, who were AutoML novices, the researchers found about 85 percent of participants who used ATMSeer were confident in the models selected by the system. Nearly all participants said using the tool made them comfortable enough to use AutoML systems in the future.  ... " 

Also discusses Auto-Tuned Models ATMs ... "

Wednesday, December 19, 2018

The Goal of Automating AI

Good, lengthy and somewhat technical piece from O'Reilly.  Worth reading.   Quite a challenge out there.   This will happen, and relatively soon.   This is also why I think teaching everyone coding will not be a useful thing to do.  It will be specialist work for a few. It will all be automated.   Key will be to have people work with data sources and apply results.  This is worth a read:

Deep Automation in Machine Learning
We need to do more than automate model building with autoML; we need to automate tasks at every stage of the data pipeline.     By Ben Lorica, Mike Loukides in O'Reilly

In a previous post, we talked about applications of machine learning (ML) to software development, which included a tour through sample tools in data science and for managing data infrastructure. Since that time, Andrej Karpathy has made some more predictions about the fate of software development: he envisions a Software 2.0, in which the nature of software development has fundamentally changed. Humans no longer implement code that solves business problems; instead, they define desired behaviors and train algorithms to solve their problems. As he writes, “a neural network is a better piece of code than anything you or I can come up with in a large fraction of valuable verticals.” We won’t be writing code to optimize scheduling in a manufacturing plant; we’ll be training ML algorithms to find optimum performance based on historical data.

If humans are no longer needed to write enterprise applications, what do we do? Humans are still needed to write software, but that software is of a different type. Developers of Software 1.0 have a large body of tools to choose from: IDEs, CI/CD tools, automated testing tools, and so on. The tools for Software 2.0 are only starting to exist; one big task over the next two years is developing the IDEs for machine learning, plus other tools for data management, pipeline management, data cleaning, data provenance, and data lineage. ... "