/* ---- Google Analytics Code Below */
Showing posts with label Brownlee. Show all posts
Showing posts with label Brownlee. Show all posts

Wednesday, April 28, 2021

EBook on Ensemble Learning

I see that Jason Brownlee has a new book on Ensemble Learning,  a method  I recently mentioned here,  “Ensemble Learning Algorithms With Python“   at the link a considerable look at it. Have not examined it myself as yet.  

…so What is Ensemble Learning?

Ensemble learning algorithms combine the predictions of two or more models.

The idea of ensemble learning is closely related to the idea of the “wisdom of crowds“. This is where many different independent decisions, choices or estimates are combined into a final outcome that is often more accurate than any single contribution.

This is the core idea behind major aspects of modern society, such as a scientific peer review, a jury of peers, and seeking a second opinion. It is an alternative to seeking out and taking the advice of an expert.

In applied machine learning, it means combining the predictions from multiple models trained on your dataset, instead of seeking the single best performing model. ... 

Sunday, December 20, 2020

Occams Razor and Ensemble Learning

Thoughtful piece.  Much more at the link.   Here the intro:

Ensemble Learning Algorithm Complexity and Occam’s Razor   by Jason Brownlee on December 21, 2020 in Ensemble Learning 

Occam’s razor suggests that in machine learning, we should prefer simpler models with fewer coefficients over complex models like ensembles.

Taken at face value, the razor is a heuristic that suggests more complex hypotheses make more assumptions that, in turn, will make them too narrow and not generalize well. In machine learning, it suggests complex models like ensembles will overfit the training dataset and perform poorly on new data.

In practice, ensembles are almost universally the type of model chosen on projects where predictive skill is the most important consideration. Further, empirical results show a continued reduction in generalization error as the complexity of an ensemble learning model is incrementally increased. These findings are at odds with the Occam’s razor principle taken at face value.

In this tutorial, you will discover how to reconcile Occam’s Razor with ensemble machine learning.

After completing this tutorial, you will know:

heuristic that suggests choosing simpler machine learning models as they are expected to generalize better. The heuristic can be divided into two razors, one of which is true and remains a useful tool and the other that is false and should be abandoned.

Ensemble learning algorithms like boosting provide a specific case of how the second razor fails and added complexity can result in lower generalization error.

Let’s get started.  ...  "

Tuesday, November 17, 2020

Tutorial: Random Forest for Time Series

 Am a long  time proponent of ensemble methods.  Here Jason Brownlee provides a nice tutorial on an often powerful method everyone should know.  As usual, well done, minimal tech.

Random Forest for Time Series Forecasting   by Jason Brownlee

by Jason Brownlee  in Time Series

Random Forest is a popular and effective ensemble machine learning algorithm.

It is widely used for classification and regression predictive modeling problems with structured (tabular) data sets, e.g. data as it looks in a spreadsheet or database table.

Random Forest can also be used for time series forecasting, although it requires that the time series dataset be transformed into a supervised learning problem first. It also requires the use of a specialized technique for evaluating the model called walk-forward validation, as evaluating the model using k-fold cross validation would result in optimistically biased results.

In this tutorial, you will discover how to develop a Random Forest model for time series forecasting.

After completing this tutorial, you will know:

Random Forest is an ensemble of decision trees algorithms that can be used for classification and regression predictive modeling.  Time series datasets can be transformed into supervised learning using a sliding-window representation.How to fit, evaluate, and make predictions with an Random Forest regression model for time series forecasting.

Let’s get started.   ....  

Wednesday, July 01, 2020

Book: Data Preparation for Machine Learning

Just saw this announcement from Jason Brownlee. Have read some of his previous works, nicely done.   This intro starts where it should, at data preparation and understanding.   Much more,  including examples at the link.

Data Preparation for Machine Learning
Data Cleaning, Feature Selection, and Data Transforms in Python
Data Preparation for Machine Learning
By Jason Brownlee
$37 USD

Data preparation involves transforming raw data in to a form that can be modeled using machine learning algorithms.

Cut through the equations, Greek letters, and confusion, and discover the specialized data preparation techniques that you need to know to get the most out of your data on your next project.

Using clear explanations, standard Python libraries, and step-by-step tutorial lessons, you will discover how to confidently and effectively prepare your data for predictive modeling with machine learning.

About this Ebook:

Read on all devices: English PDF format EBook, no DRM.
Tons of tutorials: 30 step-by-step lessons, 398 pages.
Foundations: intuitions feature selection, scaling, more.
Working code: 168 Python (.py) code files included.
Clear, Complete End-to-End Examples.
Convinced?
Click to jump straight to the packages.  ... " 

Monday, January 27, 2020

When can you Selectively look at Less than all of the Data?

I pass along Jason Brownlee's links from time to time, have found them very useful.   Subscribe to his stream, buy his books.  Here its about sampling and when you can look at less than all of the data. 

Jason @ ML Mastery jason@machinelearningmastery.com 
Thu, Jan 23, 1:12 PM (3 days ago)    to Franzdill

Hi, this week we have a tutorial on undersampling algorithms for imbalanced classification, a tutorial on combining oversampling and undersampling, and a tour of data sampling methods.

Discover how to delete examples from your dataset to improve performance:
>> Undersampling Algorithms for Imbalanced Classification

Discover specialized techniques that perform both oversampling and undersampling:
>> Combine Oversampling and Undersampling for Imbalanced Classification

Discover a suite of data sampling techniques available for imbalanced classification:
>> Tour of Data Sampling Methods for Imbalanced Classification

See his new book
Imbalanced Classification with Python
Better Metrics, Balance Skewed Classes, Cost-Sensitive Learning  ... "

Wednesday, November 20, 2019

Types of Machine Learning

Once again an excellent piece by Jason Brownlee,  a intro to a number of kinds of machine learning.   Abstract intro below, More at the link, and do subscribe:

14 Different Types of Learning in Machine Learning
by Jason Brownlee on November 11, 2019 in Start Machine Learning

Machine learning is a large field of study that overlaps with and inherits ideas from many related fields such as artificial intelligence.

The focus of the field is learning, that is, acquiring skills or knowledge from experience. Most commonly, this means synthesizing useful concepts from historical data.

As such, there are many different types of learning that you may encounter as a practitioner in the field of machine learning: from whole fields of study to specific techniques.

In this post, you will discover a gentle introduction to the different types of learning that you may encounter in the field of machine learning.

After reading this post, you will know:

Fields of study, such as supervised, unsupervised, and reinforcement learning.
Hybrid types of learning, such as semi-supervised and self-supervised learning.
Broad techniques, such as active, online, and transfer learning.
Let’s get started.
Types of Learning
Given that the focus of the field of machine learning is “learning,” there are many types that you may encounter as a practitioner.

Some types of learning describe whole subfields of study comprised of many different types of algorithms such as “supervised learning.” Others describe powerful techniques that you can use on your projects, such as “transfer learning.”

There are perhaps 14 types of learning that you must be familiar with as a machine learning practitioner; they are:

Learning Problems

1. Supervised Learning
2. Unsupervised Learning
3. Reinforcement Learning
Hybrid Learning Problems

4. Semi-Supervised Learning
5. Self-Supervised Learning
6. Multi-Instance Learning
Statistical Inference

7. Inductive Learning
8. Deductive Inference
9. Transductive Learning
Learning Techniques

10. Multi-Task Learning
11. Active Learning
12. Online Learning
13. Transfer Learning
14. Ensemble Learning
In the following sections, we will take a closer look at each in turn.

Did I miss an important type of learning?
Let me know in the comments below.

Learning Problems ..... " 

Saturday, October 26, 2019

Cross Entropy

What is this?  New to me as a term.  I record here for my own refernce and pass it out to others.   In particular the suggestion here is that classification may be improved or even optimized this way.   See Jason's other publications, some mentioned here.

A Gentle Introduction to Cross-Entropy for Machine Learning   by Jason Brownlee

Cross-entropy is commonly used in machine learning as a loss function.

Cross-entropy is a measure from the field of information theory, building upon entropy and generally calculating the difference between two probability distributions. It is closely related to but is different from KL divergence that calculates the relative entropy between two probability distributions, whereas cross-entropy can be thought to calculate the total entropy between the distributions.

Cross-entropy is also related to and often confused with logistic loss, called log loss. Although the two measures are derived from a different source, when used as loss functions for classification models, both measures calculate the same quantity and can be used interchangeably.

In this tutorial, you will discover cross-entropy for machine learning.

After completing this tutorial, you will know:

How to calculate cross-entropy from scratch and using standard machine learning libraries.
Cross-entropy can be used as a loss function when optimizing classification models like logistic regression and artificial neural networks.
Cross-entropy is different from KL divergence but can be calculated using KL divergence, and is different from log loss but calculates the same quantity when used as a loss function.
Discover bayes opimization, naive bayes, maximum likelihood, distributions, cross entropy, and much more in my new book, with 28 step-by-step tutorials and full Python source code.

Let’s get started: ... " 

Sunday, October 13, 2019

Information Entropy and Data

With a background in physics this is a great topic.  It links the universe to information technologies in interesting ways.  Even includes a hint at the nature of 'surprise'.  At very minimum impress your friends.

 Gentle Introduction to Information Entropy   by Jason Brownlee  in Probability

Information theory is a subfield of mathematics concerned with transmitting data across a noisy channel.

A cornerstone of information theory is the idea of quantifying how much information there is in a message. More generally, this can be used to quantify the information in an event and a random variable, called entropy, and is calculated using probability.

Calculating information and entropy is a useful tool in machine learning and is used as the basis for techniques such as feature selection, building decision trees, and, more generally, fitting classification models. As such, a machine learning practitioner requires a strong understanding and intuition for information and entropy.

In this post, you will discover a gentle introduction to information entropy.

After reading this post, you will know:

Information theory is concerned with data compression and transmission and builds upon probability and supports machine learning.

Information provides a way to quantify the amount of surprise for an event measured in bits.
Entropy provides a measure of the average amount of information needed to represent an event drawn from a probability distribution for a random variable.

Discover bayes opimization, naive bayes, maximum likelihood, distributions, cross entropy, and much more in my new book, with 28 step-by-step tutorials and full Python source code.

Let’s get started.    .... " 

Sunday, September 15, 2019

Certainty is Unusual

Jason Brownlee does his usual good job of explaining important concepts.   Here largely non technical.  And this is perhaps the most important.  Often the hardest to explain to decision makers, despite the fact that they deal with the problem every day.  Risk must always be considered.   Heartily recommend you subscribe.  See the 'Brownlee' tag below for other tutorials from Jason I have mentioned.

What Is Probability?  by Jason Brownlee  

Uncertainty involves making decisions with incomplete information, and this is the way we generally operate in the world.

Handling uncertainty is typically described using everyday words like chance, luck, and risk.

Probability is a field of mathematics that gives us the language and tools to quantify the uncertainty of events and reason in a principled manner.

In this post, you will discover a gentle introduction to probability.

After reading this post, you will know:

Certainty is unusual and the world is messy, requiring operating under uncertainty.
Probability quantifies the likelihood or belief that an event will occur.
Probability theory is the mathematics of uncertainty.

Let’s get started. ... 

Thursday, July 04, 2019

Intro to GAN's by Jason Brownlee

Good intro, have just passed this on to a group ...

A Gentle Introduction to Generative Adversarial Networks (GANs)
by Jason Brownlee on June 17, 2019 in Generative Adversarial Networks  Follow him, good understandable content. 

Generative Adversarial Networks, or GANs for short, are an approach to generative modeling using deep learning methods, such as convolutional neural networks.

Generative modeling is an unsupervised learning task in machine learning that involves automatically discovering and learning the regularities or patterns in input data in such a way that the model can be used to generate or output new examples that plausibly could have been drawn from the original dataset.

GANs are a clever way of training a generative model by framing the problem as a supervised learning problem with two sub-models: the generator model that we train to generate new examples, and the discriminator model that tries to classify examples as either real (from the domain) or fake (generated). The two models are trained together in a zero-sum game, adversarial, until the discriminator model is fooled about half the time, meaning the generator model is generating plausible examples.

GANs are an exciting and rapidly changing field, delivering on the promise of generative models in their ability to generate realistic examples across a range of problem domains, most notably in image-to-image translation tasks such as translating photos of summer to winter or day to night, and in generating photorealistic photos of objects, scenes, and people that even humans cannot tell are fake.

In this post, you will discover a gentle introduction to Generative Adversarial Networks, or GANs. .... " 

Tuesday, March 19, 2019

Jason Brownlee Reviews Stanford CNN Course

Thoughtful piece on a seminal course.  Useful if you are considering taking a deeper dive.

Stanford Convolutional Neural Networks for Visual Recognition Course (Review) by Jason Brownlee 

The Stanford course on deep learning for computer vision is perhaps the most widely known course on the topic.

This is not surprising given that the course has been running for four years, is presented by top academics and researchers in the field, and the course lectures and notes are made freely available.

This is an incredible resource for students and deep learning practitioners alike.

In this post, you will discover a gentle introduction to this course that you can use to get a jump-start on computer vision with deep learning methods.

After reading this post, you will know:

The breakdown of the course including who teaches it, how long it has been taught, and what it covers.

The breakdown of the lectures in the course including the three lectures to focus on if you are already familiar with deep learning.

A review of the course, including how it compares to similar courses on the same subject matter.
Let’s get started.  ...." 

Saturday, January 26, 2019

Ensemble Models for Deep Learning

Ensemble models are now commonly used in all sorts of analytics.  You use the results of multiple models and combine the results.  Jason Brownlee shows how this can be done for deep learning methods.  Good tutorial explanation.

How to Create a Random-Split, Cross-Validation, and Bagging Ensemble for Deep Learning in Keras  by Jason Brownlee in Better Deep Learning

Ensemble learning are methods that combine the predictions from multiple models.

It is important in ensemble learning that the models that comprise the ensemble are good, making different prediction errors. Predictions that are good in different ways can result in a prediction that is both more stable and often better than the predictions of any individual member model.

One way to achieve differences between models is to train each model on a different subset of the available training data. Models are trained on different subsets of the training data naturally through the use of resampling methods such as cross-validation and the bootstrap, designed to estimate the average performance of the model generally on unseen data. The models used in this estimation process can be combined in what is referred to as a resampling-based ensemble, such as a cross-validation ensemble or a bootstrap aggregation (or bagging) ensemble.

In this tutorial, you will discover how to develop a suite of different resampling-based ensembles for deep learning neural network models.   ... " 

Sunday, November 18, 2018

Use Weight Regularization to Reduce Overfitting of Deep Learning Models

Overfitting is a classic problem with all models.  It means you are finding the solution to a particular set of data, rather than a generalized problem.  This should be found afterward in testing against new data,  but can be dangerously misleading.   Jason Brownlee discusses approaches to reduce overfitting.   Follow Jason, lots of good nuggets.

Use Weight Regularization to Reduce Overfitting of Deep Learning Models   by Jason Brownlee  in Better Deep Learning

Neural networks learn a set of weights that best map inputs to outputs.
A network with large network weights can be a sign of an unstable network where small changes in the input can lead to large changes in the output. This can be a sign that the network has overfit the training dataset and will likely perform poorly when making predictions on new data.

A solution to this problem is to update the learning algorithm to encourage the network to keep the weights small. This is called weight regularization and it can be used as a general technique to reduce overfitting of the training dataset and improve the generalization of the model.

In this post, you will discover weight regularization as an approach to reduce overfitting for neural networks. .... "

Wednesday, November 14, 2018

Statistics for Machine Learning Models

Nicely done look at the problem, by the always  useful Jason Brownlee.   Follow his writing subscribe to his tutorials.

 Statistics for Evaluating Machine Learning Models  by Jason Brownlee  

Tom Mitchell’s classic 1997 book “Machine Learning” provides a chapter dedicated to statistical methods for evaluating machine learning models.

Statistics provides an important set of tools used at each step of a machine learning project. A practitioner cannot effectively evaluate the skill of a machine learning model without using statistical methods. Unfortunately, statistics is an area that is foreign to most developers and computer science graduates. This makes the chapter in Mitchell’s seminal machine learning text an important, if not required, reading by practitioners.

In this post, you will discover statistical methods recommended by Mitchel to evaluate and compare machine learning models. .. " 

Thursday, August 30, 2018

Deep Learning for Time Series Forecasting

Like Jason's style of clear motivations  and short tutorials.  You can get free samples of his writing below.

Jason Brownlee's New Book: 
Deep Learning for Time Series Forecasting
Predict the Future with MLPs, CNNs and LSTMs in Python
Deep Learning for Time Series Forecasting
$37 USD

Deep learning methods offer a lot of promise for time series forecasting, such as the automatic learning of temporal dependence and the automatic handling of temporal structures like trends and seasonality.

In this new Ebook written in the friendly Machine Learning Mastery style that you’re used to, finally cut through the math, research papers and patchwork descriptions about time series forecasting with deep learning algorithms.

With clear explanations, standard Python libraries, and step-by-step tutorial lessons you’ll discover how to develop deep learning models for your own time series forecasting projects.

About this Ebook:

Read on all devices: PDF format Ebook, no DRM.
Tons of tutorials: 5 parts, 25 step-by-step lessons, 575 pages.
Real-world projects: 2 large end-to-end tutorial projects.
Many datasets: Univariate, multivariate, multi-step, and more.
Working code: 131 Python (.py) code files included.
Clear, Complete End-to-End Examples.
Convinced? ....  "

Sunday, August 26, 2018

What and Why are ARCH and GARCH?

When you do time series forecasting you almost always get changes in variance over time.  Sometimes enough to invalidate your decisions and conclusions.  We used these  methods in key ways to produce better results over time.  Somehow I rarely hear these methods mentioned recently.  Here Jason Brownlee provides a good Python based intro.  Fairly non-technical, but coding based.

How to Model Volatility with ARCH and GARCH for Time Series Forecasting in Python by Jason Brownlee   in Time Series

A change in the variance or volatility over time can cause problems when modeling time series with classical methods like ARIMA.

The ARCH or Autoregressive Conditional Heteroskedasticity method provides a way to model a change in variance in a time series that is time dependent, such as increasing or decreasing volatility. An extension of this approach named GARCH or Generalized Autoregressive Conditional Heteroskedasticity allows the method to support changes in the time dependent volatility, such as increasing and decreasing volatility in the same series.

In this tutorial, you will discover the ARCH and GARCH models for predicting the variance of a time series.

After completing this tutorial, you will know:

The problem with variance in a time series and the need for ARCH and GARCH models.
How to configure ARCH and GARCH models.
How to implement ARCH and GARCH models in Python.
Let’s get started.   .... "

Sunday, August 12, 2018

What is Machine Learning?

Straightforward introduction.

How to Think About Machine Learning    by Jason Brownlee 
Machine learning is a large and interdisciplinary field of study.

You can achieve impressive results with machine learning and find solutions to very challenging problems. But this is only a small corner of the broader field of machine learning often called predictive modeling or predictive analytics.

In this post, you will discover how to change the way you think about machine learning in order to best serve you as a machine learning practitioner.

After reading this post, you will know:

What machine learning is and how it relates to artificial intelligence and statistics.
The corner of machine learning that you should focus on.
How to think about your problem and the machine learning solution to your problem.

Let’s get started. ... ." 

Wednesday, August 08, 2018

Choosing a Neural Network

Another excellent piece from Jason, suggest you join up with his service:

Jason Brownlee writes:   What neural network is appropriate for your predictive modeling problem?

It can be difficult for a beginner to the field of deep learning to know what type of network to use. There are so many types of networks to choose from and new methods being published and discussed every day.

To make things worse, most neural networks are flexible enough that they work (make a prediction) even when used with the wrong type of data or prediction problem.

In this post, you will discover the suggested use for the three main classes of artificial neural networks.

After reading this post, you will know:

Which types of neural networks to focus on when working on a predictive modeling problem.
When to use, not use, and possible try using an MLP, CNN, and RNN on a project. ... To consider the use of hybrid models and to have a clear idea of your project goals before selecting a model.

Let’s get started.  ... "

Sunday, April 15, 2018

Getting the Most from your Data

Once more, Jason Brownlee makes excellent points.   It starts and ends with the data.   Its the biggest asset, and also the largest obstacle.  Frame it and test it.  Is there enough, does it support the methods you can use,  is it clean?  Here is his intro, much more at the link:


How to Get the Most From Your Machine Learning Data
by Jason Brownlee on April 16, 2018 in Machine Learning Process

The data that you use, and how you use it, will likely define the success of your predictive modeling problem.

Data and the framing of your problem may be the point of biggest leverage on your project.

Choosing the wrong data or the wrong framing for your problem may lead to a model with poor performance or, at worst, a model that cannot converge.

It is not possible to analytically calculate what data to use or how to use it, but it is possible to use a trial-and-error process to discover how to best use the data that you have.

In this post, you will discover to get the most from your data on your machine learning project.

After reading this post, you will know:

The importance of exploring alternate framings of your predictive modeling problem.
The need to develop a suite of “views” on your input data and to systematically test each.
The notion that feature selection, engineering, and preparation are ways of creating more views on your problem.

Let’s get started.... "

Sunday, April 08, 2018

How to Think about Machine Learning

Nicely and non-technically put:

How to Think About Machine Learning  Intro:   by Jason Brownlee  

Machine learning is a large and interdisciplinary field of study.

You can achieve impressive results with machine learning and find solutions to very challenging problems. But this is only a small corner of the broader field of machine learning often called predictive modeling or predictive analytics.

In this post, you will discover how to change the way you think about machine learning in order to best serve you as a machine learning practitioner.

After reading this post, you will know:

What machine learning is and how it relates to artificial intelligence and statistics.
The corner of machine learning that you should focus on.
How to think about your problem and the machine learning solution to your problem.
Let’s get started. ....   "