/* ---- Google Analytics Code Below */
Showing posts with label Automating Data Science. Show all posts
Showing posts with label Automating Data Science. Show all posts

Monday, March 28, 2022

Automating Data Science?

 Good to have consistent,  easily understood efforts in a library of solutions. 

Automating Data Science

By Tijl De Bie, Luc De Raedt, José Hernández-Orallo, Holger H. Hoos, Padhraic Smyth, Christopher K. I. Williams

Communications of the ACM, March 2022, Vol. 65 No. 3, Pages 76-87   10.1145/3495256

Data science covers the full spectrum of deriving insight from data, from initial data gathering and interpretation, via processing and engineering of data, and exploration and modeling, to eventually producing novel insights and decision support systems.

 Data science can be viewed as overlapping or broader in scope than other data-analytic methodological disciplines, such as statistics, machine learning, databases, or visualization

To illustrate the breadth of data science, consider, for example, the problem of recommending items (movies, books, or other products) to customers.

 While the core of these applications can consist of algorithmic techniques such as matrix factorization, a deployed system will involve a much wider range of technological and human considerations.

 These range from scalable back-end transaction systems that retrieve customer and product data in real time, experimental design for evaluating system changes, causal analysis for understanding the effect of interventions, to the human factors and psychology that underlie how customers react to visual information displays and make decisions.

As another example, in areas such as astronomy, particle physics, and climate science, there is a rich tradition of building computational pipelines to support data-driven discovery and hypothesis testing. For instance, geoscientists use monthly global landcover maps based on satellite imagery at sub-kilometer resolutions to better understand how the Earth's surface is changing over time.50 These maps are interactive and browsable, and they are the result of a complex data-processing pipeline, in which terabytes to petabytes of raw sensor and image data are transformed into databases of a6utomatically detected and annotated objects and information. This type of pipeline involves many steps, in which human decisions and insight are critical, such as instrument calibration, removal of outliers, and classification of pixels.

The breadth and complexity of these and many other data science scenarios means the modern data scientist requires broad knowledge and experience across a multitude of topics. Together with an increasing demand for data analysis skills, this has led to a shortage of trained data scientists with appropriate background and experience, and significant market competition for limited expertise. Considering this bottleneck, it is not surprising there is increasing interest in automating parts, if not all, of the data science process. This desire and potential for automation is the focus of this article.

Tuesday, March 01, 2022

Automating Data Science: Prospects and Challenges

 Long, insightful piece, good insights,  somewhat technical, but useful.   A favorite topic we experimented with since very early days.

Home/Magazine Archive/March 2022 (Vol. 65, No. 3)/Automating Data Science/Full Text

 By Tijl De Bie, Luc De Raedt, José Hernández-Orallo, Holger H. Hoos, Padhraic Smyth, Christopher K. I. Williams

Communications of the ACM, March 2022, Vol. 65 No. 3, Pages 76-87 10.1145/3495256

--> Given the complexity of typical data science projects and the associated demand for human expertise, automation has the potential to transform the data science process.

Key insights: 

• Automation in data science aims to facilitate and transform the work of data scientists, not to replace them.

• Important parts of data science are already being automated, especially in the modeling stages, where techniques such as automated machine learning (AutoML) are gaining traction.

• Other aspects are harder to automate, not only because of technological challenges, but because open ended and context-dependent tasks require human interaction.

Introduction

Data science covers the full spectrum of deriving insight from data, from initial data gathering and interpretation, via processing and engineering of data, and exploration and modeling, to eventually producing novel insights and decision support systems. Data science can be viewed as overlapping or broader in scope than other data-analytic methodological disciplines, such as statistics, machine learning, databases, or visualization

To illustrate the breadth of data science, consider, for example, the problem of recommending items (movies, books or other products) to customers. While the core of these applications can consist of algorithmic techniques such as matrix factorization, a deployed system will involve a much wider range of technological and human considerations. These range from scalable back-end transaction systems that retrieve customer and product data in real time, experimental design for evaluating system changes, causal analysis for understanding the effect of interventions, to the human factors and psychology that underliehow customers react to visual information displays and make decisions.

As another example, in areas such as astronomy, particle physics, and climate science, there is a rich tradition of building computational pipelines to support data-driven discovery and hypothesis testing. For instance, geoscientists use monthly global landcover maps based on satellite imagery at sub-kilometer resolutions to better understand how the earth’s surface is changing over time [50]. These maps are interactive and browsable, and they are the result of a complex data-processing pipeline, in which terabytes to petabytes of raw sensor and image data are transformed into databases of automatically detected and annotated objects and information. This type of pipeline involves many steps, in which human decisions and insight are critical, such as instrument calibration, removal of outliers, and classification of pixels.

The breadth and complexity of these and many other data science scenarios means that the modern data scientist requires broad knowledge and experience across a multitude of topics. Together with an increasing demand for data analysis skills, this has led to a shortage of trained data scientists with appropriate background and experience, and significant market competition for limited expertise. Considering this bottleneck, it is not surprising that there is increasing interest in automating parts, if not all, of the data science process. This desire and potential for automation is the focus of this article.

As illustrated in the examples above, data science is a complex process, driven by the character of the data being analyzed and by the questions being asked, and is often highly exploratory and iterative in nature. Domain context can play a key role in these exploratory steps, even in relatively well-defined processes such as predictive modeling (e.g., as characterized by CRISP-DM [5]) where, for example, human expertise in defining relevant predictor variables can be critical.  .... '

Monday, January 04, 2021

Lopsided Digital Transformation

Look at how automation of digitization can make a big difference.   And we are continuing to move to that. 

Digital Transformation Is Lopsided, Even Within the Same Organization    By ZDNet

A study of organizations in Harvard Business Review found that some corporate departments can be more digitally advanced than their counterparts at another location, or even down the hall. The IT department itself often finds itself on the wrong side of progress.

A Harvard Business Review study of organizations found "a growing divide between teams with ready access to automation and [artificial intelligence (AI)] tools and teams without," with the latter trailing "in terms of productivity and AI skill development."

Accenture's Ramnath Venkataraman and colleagues determined that two-thirds of investigated companies relied on a less-than-optimal blend of cloud-based and on-premises enterprise solutions.

Exacerbating the problem is information technology (IT) departments being mired in maintenance and upkeep, while other departments have access to cutting-edge systems.

The Accenture team also learned that some coders might devote 60% of their time executing automatable tasks.

The researchers suggested greater inter-enterprise collaboration to address these challenges, by de-siloing walled-off departments and encouraging open access to shadow IT systems. ... ' 

From ZDNet

Monday, July 20, 2020

Automatically Generating Reinforcement Learning Algorithms

Can be expected that many current machine learning techniques will move towards automation.  Papers mentioned are worth looking at.

DeepMind’s AI automatically generates reinforcement learning algorithms
Kyle Wiggers in VentureBeat

 In a study  published on the preprint server Arxiv.org, DeepMind researchers describe a reinforcement learning  algorithm-generating technique that discovers what to predict and how to learn it by interacting with environments. They claim the generated algorithms perform well on a range of challenging Atari video games, achieving “non-trivial” performance indicative of the technique’s generalizability.

Reinforcement learning algorithms — algorithms that enable software agents to learn in environments by trial and error using feedback — update an agent’s parameters according to one of several rules. These rules are usually discovered through years of research, and automating their discovery from data could lead to more efficient algorithms, or algorithms better adapted to specific environments.  ... " 

Friday, September 20, 2019

Apple Overton Leading to Code Automation?

Increasingly moving towards automating many aspects of coding.   In fact robot assistants that 'observe' the coding process could readily insure that secure, robust and repeatable methods were used when building AI systems.    They could also make sure that the most important methods were shared, maintained and updated as new research dictated. 

On the data side, that the data was properly selected, prepared and delivered with needed metadata to support explainable results.  That's why I am not a believer in just training everyone in low level coding.  People are not good at these skills.    Train them in problem solving supported by prefabricated AI systems and results visualization methods, because ultimately the classic methods will be built, solved, updated and delivered by automation.

Apple ‘Overton’: Automating Low-Code Machine Learning     By Nick Kolakowski

Apple has struggled in recent years to establish a robust artificial intelligence (A.I.) practice. This partially stems from the company’s ironclad privacy policies—it’s more difficult to analyze datasets for insights when internal rules prevent the company from using every piece of user data it can vacuum up. Nonetheless, Apple’s newest projects show that it’s powering ahead anyway—including one platform that, if it’s ever released, could change how you use A.I. and machine learning (ML).

(It’s worth remembering how, in a 2015 speech, Apple CEO Tim Cook accused tech giants such as Facebook and Google of “gobbling up everything they can learn about you and trying to monetize it,” which he framed as “wrong.” It seems unlikely that Apple’s stance on data and privacy will change during Cook’s tenure.)

According to a just-released paper with the dry-but-mysteriously-compelling title “Overton: A Data System for Monitoring and Improving Machine Learned Products,” a group of Apple researchers describe their work on a machine-learning platform (named—you guessed it—“Overton”) designed to “support engineers in building, monitoring, and improving production machine learning systems.”   ......... '

Abstract of paper mentioned above:    https://arxiv.org/pdf/1909.05372.pdf   (technical)

 ... We describe a system called Overton, whose main design goal is to support engineers in building, monitoring, and improving production machine learning systems. Key challenges engineers face are monitoring fine-grained quality,diagnosing errors in sophisticated applications, and handling contradictory or incomplete supervision data. Overton automates the life cycle of model construction, deployment, and monitoring by providing a set of novel high-level,declarative abstractions. Overton’s vision is to shift developers to these higher-level tasks instead of lower-level machine learning tasks. In fact, using Overton, engineers can build deep-learning-based applications without writing any codein frameworks like TensorFlow. For over a year, Overton has been used in production to support multiple applications in both near-real-time applications and back-of-house processing. In that time, Overton-based applications have answered billions of queries in multiple languages and processed trillions of records reducing errors 1.7 − 2.9× versus  production systems. .... "

Wednesday, April 17, 2019

Automating Machine Learning with Azure

This example was sent to me, a straight forward example of using Azure.   Always looking for useful examples of better automating at least initial tests of a machine learning example.  Big proponent of quick, early,  cheap tests of even complex modeling efforts.      Look for ways to get the idea in front of other analysts, decision makers, data providers.   Even a sketch on a board is worthwhile, but showing something interactive gives a clear taste of the potential value.

How to forecast energy demand with Azure Machine Learning | Azure Makers Series
Use Azure Machine Learning to create a model and apply it to a real-world scenario: predicting energy demand and expected load on energy grids - a critical business operation for energy companies. The same principles apply across use cases, so you can adjust for your organization’s critical operations and needs.  

GitHub Repo: https://github.com/FrancescaLazzeri/A... 

Create your Azure free account: https://aka.ms/J1PxFdcK2tY

Follow Francesca on Twitter: https://twitter.com/frlazzeri

Friday, July 13, 2018

Analytics Data Catalogs, Approaches, not new

Yes, we know this, and just because some call it AI, does not mean we won't have to gather the data consistently and continually to solve real problems. 

Analytics Industrial Revolution- From The Occult to the Ordinary
By  Snehamoy (Sneh) Mukherjee In Linkedin

Senior Director - Delivering Data Science, Big Data, Machine Learning and Analytics projects for Fortune 500 companies

There is a quiet revolution taking place in the Analytics industry that has the potential to completely turn the industry on its head and the way work gets done in this space. Doomsday pundits have already summoned the evil spirit called AI to put an end to the misery of our uneven paychecks, to be replaced by an Universal Basic Income and some of us have reconciled ourselves to that cruel fate. But before the Apocalypse happens, there is another subtle and continuous tectonic movement happening right under our feet, which if gone unnoticed for long, can catch us in a tidal wave of upheaval in the analytics /machine learning industry.

The typical Analytics (often very eruditely rechristened by brilliant marketers and/or the academia as Machine Learning and a lot would break their heads to prove that the two are different) or a Machine Learning project gets delivered in the following atypical manner in most firms: -

·        Data is pulled from one/multiple tables from a database(s) (by someone who may either be from the client side or by the analytics vendor). It may be a onetime data extraction or if data is needed on a periodic basis, an ETL (Extract Transfer Load or in some cases ELT) process is created to do a batch fetching and processing of files (e.g. weekly transaction data from retail stores, monthly/weekly call data in telecom firms, weekly/daily transaction data in banks) .... " 

Friday, March 02, 2018

Operationalizing Analytics

Good points, barriers yes, but not as difficult as it is stated.  We did it years ago.  Depends too upon what you mean by 'analytics', I define it broadly as using computational means to regularly improve decisions.  Embedded or not. Automated as needed.

Why Operationalizing Analytics is So Difficult by Lisa Morgan in InformationWeek

Plenty of companies have plenty of data and plenty of analytics tools, but they fall short when it comes to converting analytics results into action.

Today's businesses are applying analytics to a growing number of use cases, but analytics for analytics' sake has little, if any, value. The most analytically astute companies have operationalized analytics, but many of them, particularly the non-digital natives, have faced several challenges along the way getting the people, processes and technology aligned in a way that drives value for the business.

Here are some of the hurdles that an analytics initiative might encounter. .... " 

Saturday, August 05, 2017

Watson Machine Learning Now Publicly Available

I note that this also includes 'visual' machine learning,  using methods that look like the Clementine system they acquired long ago for System Modeler.   Good direction.   Explored some of this via their Bluemix services, but this takes it further, for both professionals and for non professional use.

I contractually examined several examples of enterprise use in real projects, impressive plug-in power.  Natural links to Watson Analytics?  Looking at the new usability here and cost.

Watson Machine Learning is now Generally Available
Today we are excited to announce the general availability of the IBM Watson Machine Learning service. Over the past 12 months we've got feedback from hundreds of beta users of the Watson Machine Learning (WML) service. During the beta period, we’ve been actively collecting feedback provided via email, Slack, and targeted surveys. The WML product team has been actively engaged in those conversations and wherever possible we’ve worked to incorporate your feedback in to the service. With today’s announcement, we are now opening this service to the general public and rolling out a number of new features. Read on to learn more…

What is WML and why are we building it?
WML is a Bluemix service that enables users to perform two fundamental operations of machine learning. .... 

Training: this is the process of refining an algorithm so that it can 'learn' from a dataset. The output of this operation is called a model. A model encompasses the learned coefficients of mathematical expressions.

Scoring: the operation of predicting an outcome using a trained model. The output of the scoring operation is another dataset containing predicted values.   ... " 

More details.


Automated Machine Learning Tools

William Vorhies surveys some 'automated' data machine learning systems.  Agree these are for professionals,  but can likely decrease the effort needed to produce such models.  But their emergence can only mean that these techniques will ultimately be automated more fully. ...

Automated Machine Learning for Professionals
Posted by William Vorhies  

Summary:  There are a variety of new Automated Machine Learning (AML) platforms emerging that led us recently to ask if we’d be automated and unemployed any time soon.  In this article we’ll cover the “Professional AML tools”.  They require that you be fluent in R or Python which means that Citizen Data Scientists won’t be using them.  They also significantly enhance productivity and reduce the redundant and tedious work that’s part of model building.   ...   " 

Wednesday, July 26, 2017

Automated Machine Learning

Excellent piece on this topic. Am in the process of preparing a talk to a Columbia University Group on exactly this topic. Useful detail at the link. I think it is inevitable we will see highly automated processes of these types, for professionals and even end users.  Here a good list of work underway and implications for professionals doing the coding.

Automated Machine Learning for Professionals   Posted by William Vorhies  
Summary:  There are a variety of new Automated Machine Learning (AML) platforms emerging that led us recently to ask if we’d be automated and unemployed any time soon.  In this article we’ll cover the “Professional AML tools”.  They require that you be fluent in R or Python which means that Citizen Data Scientists won’t be using them.  They also significantly enhance productivity and reduce the redundant and tedious work that’s part of model building. ... " 

Friday, July 07, 2017

Driverless AI

The term was new to me, but not the concept.   Makes sense,  if it is really AI then it should manage itself.  Easier/cheaper than having room fulls of scientists building systems.   But here also aspects like testing, maintaining and connecting to corporate systems creates most of the difficulty.   How close is this?

H2O.ai’s Driverless AI automates machine learning for businesses
 Driverless AI is the latest product from H2O.ai aimed at lowering the barrier to making data science work in a corporate context. The tool assists non-technical employees with preparing data, calibrating parameters and determining the optimal algorithms for tackling specific ...   In Techcrunch.

(Update):And further, in Datanami:
H2O.ai Boasts New AI Product Like ‘Kaggle Grandmaster in a Box’  by Alex Woodie

Monday, May 29, 2017

Automating Aspects of Machine Learning

A topic I brought up as a key part of the future of machine learning at our recent UC analytics summit.  And as we might expect, Google is working on it.  From the recent Google Research Blog.   This kind of problem will need lots of data to explore it, so expect companies like Google, Amazon, Apple  and Microsoft to have the data to do it.  Continuing to watch this.  As the internal architecture of AI is fine tuned by ML methods, it will become more powerful. 

Posted by Quoc Le & Barret Zoph, Research Scientists, Google Brain team ... 

At Google, we have successfully applied deep learning models to many applications, from image recognition to speech recognition to machine translation. Typically, our machine learning models are painstakingly designed by a team of engineers and scientists. This process of manually designing machine learning models is difficult because the search space of all possible models can be combinatorially large — a typical 10-layer network can have ~1010 candidate networks! For this reason, the process of designing networks often takes a significant amount  .... " 

As an example, they show the complexity and effort needed to develop their Googlenet Architecture.

Monday, April 03, 2017

SAS Visual Analytics

Just saw a demonstration of SAS Visual Analytics.  Nicely done:    " ... See the big picture – and underlying connections. ... Visually explore critical drivers for making better decisions. Find out why something happened. Examine all options and uncover opportunities hidden deep in your data. Automatically highlight key relationships, outliers, clusters and more, revealing vital insights that inspire action.   ... " .    An automation direction addressing other competitors in the space.

Friday, January 27, 2017

Allowing End Users to Build Models

Very good general question that we have been discussing for years, well, really decades.   Started way back in the spreadsheet era, when you empowered users to build 'models' of any kind.    Afterwards we were sometimes brought in to clean up the mess.     I would go back to a risk analysis.  What are the risk implications of giving a wrong answer?  How wrong?  Depending on the answer to that, what kind of model checks the answer?  Who/what decides the answer is right?  Need the results be formally checked at all? How can this cautionary checking be properly automated?

Discussion in Linkedin:
Is it risky to let users create their own predictive models?    by Michael Surkan

Friday, January 20, 2017

AI Building AI

It has been long suggested that this would occur.  This is still relatively hard, but we have seen apps like Watson Analytics show how it could be done.  Though that is about data science, and this further embodies the logic and process of AI.  In Technology Review:

 AI Software Learns to Make AI Software
Google and others think software that learns to learn could take over some work done by AI experts.
by Tom Simonite .... "

Wednesday, January 11, 2017

Transfer Learning for AI Projects

Had always thought that intelligence was about learning, so this concept struck me.   Note mention of improbable events and model correctness maintenance,  always of concern in such studies.  Technical.

'Transfer learning' jump-starts new AI projects
Machine learning, once implemented, tends to be specific to the data and requirements of the task at hand. Transfer learning is the act of abstracting and reusing those smarts

'Transfer Learning' Jump-Starts New AI Projects  in InfoWorld by James Kobielus

Abstracting and reusing knowledge gleaned from a machine-learning application in other, newer apps--or "transfer learning"--is supplementing other learning methods that constitute the backbone of most data science practices. Among the technique's practical uses is productivity acceleration modeling, which is viable when prior work can be reused without extensive revision in order to speed up time to insight. Another transfer-learning application involves the method helping scientists produce machine-learning models that exploit relevant training data from prior modeling projects.

This technique is particularly appropriate for addressing projects in which prior training data can easily become obsolete, which is a problem that frequently occurs in dynamic problem domains. A third area of data science in which transfer learning could yield benefits is risk mitigation. In this situation, transfer learning can help scientists leverage subsets of training data and feature models from related domains when the underlying conditions of the modeled phenomenon have radically changed. 

This can help researchers ameliorate the risk of machine-learning-driven predictions in any problem domain vulnerable to extremely improbable events. Transfer learning also is critical to data scientists' efforts to create "master learning algorithms" that automatically obtain and apply fresh contextual knowledge via deep neural networks and other forms of artificial intelligence. ... " 

Saturday, October 29, 2016

Conversational Analytics Using AI

An obvious thought, place bot style AI into a conversational interaction about doing analytics for given problems and data.    The caution is still that we still don't know how to manage complex conversations,  just simplistic ones.  Perhaps as a kind of conversationally managed automation of data science.   This effort was new to me.  Plan to take a look.  Drastin unveils the World's first conversational analytics product powered by Artificial Intelligence.   See the Drastin site.

Wednesday, October 26, 2016

Automating Big Data Analysis

Automating big-data analysis

With new algorithms, data scientists could accomplish in days what has traditionally taken months.
Larry Hardesty | MIT News Office

Last year, MIT researchers presented a system that automated a crucial step in big-data analysis: the selection of a “feature set,” or aspects of the data that are useful for making predictions. The researchers entered the system in several data science contests, where it outperformed most of the human competitors and took only hours instead of months to perform its analyses.

This week, in a pair of papers at the IEEE International Conference on Data Science and Advanced Analytics, the team described an approach to automating most of the rest of the process of big-data analysis — the preparation of the data for analysis and even the specification of problems that the analysis might be able to solve.

The researchers believe that, again, their systems could perform in days tasks that used to take data scientists months.

“The goal of all this is to present the interesting stuff to the data scientists so that they can more quickly address all these new data sets that are coming in,” says Max Kanter MEng ’15, who is first author on last year’s paper and one of this year’s papers. “[Data scientists want to know], ‘Why don’t you show me the top 10 things that I can do the best, and then I’ll dig down into those?’ So [these methods are] shrinking the time between getting a data set and actually producing value out of it.” ...  '

Sunday, October 23, 2016

End of the Data Scientist Era? Not Right Away.

In Forbes.  Insightful article by Margaret Harrist of Oracle.   An interesting view of how data science is moving towards becoming automated.     Agree this is happening,  though the time required and dynamics it will require to become common will vary considerably by business domain.    It is pointed out there are three broad skills sets involved:

" -  The business-savvy data scientist, typically hired by lines of business.

   - The programmer data scientist who’s adept with statistical analysis toolkits, often hired to work in the IT department.

   - The algorithm expert who can build hypothesis and statistical models, frequently hired by startups and marketing agencies.  ... "

True its  rare to see all of these in one person.  And also inefficient to have them in one person.  Saw this at several enterprise interactions.  Companies try to hire programmers/algorithm experts, with no business saavy.  Difficult to establish ongoing expertise linked to the business.  We are evolving towards systems that will operate by the business process and data experts, not programmers.  Timing will vary, and there will still always be need for some 'science' to develop new techniques.   But probably not as many coders as before.