Decision trees were a favorite method in the enterprise, in part because the solution could be easily understood. Now think of them in relation to Logistic regression for use in classification:
" .... By Lalit Sachan in Blogon 05/10/2015 ...
Classification is one of the major problems that we solve while working on standard business problems across industries. In this article we’ll be discussing the major three of the many techniques used for the same, Logistic Regression, Decision Trees and Support Vector Machines [SVM].
All of the above listed algorithms are used in classification [ SVM and Decision Trees are also used for regression, but we are not discussing that today!]. Time and again I have seen people asking which one to choose for their particular problem. Classical and the most correct but least satisfying response to that question is “it depends!”. Its downright annoying, I agree. So I decided to shed some light on it depends on what. ... " ..... '
Showing posts with label Regression Tree. Show all posts
Showing posts with label Regression Tree. Show all posts
Sunday, November 22, 2015
Tuesday, November 03, 2015
Cross Industry Standard Process for Data Mining
CRISP-DM ( Cross Industry Standard Process for Data Mining) which we plugged into as far back as 07, seems to have disappeared, along with its web site. Data mining also appears to have been folded into the study and methods of data science. In particular the decision tree aspects. That is too bad, because the decision oriented pieces were particularly understandable to executives. We found the tree methods indispensable. A WP article still covers the basics . The six fundamental aspects of the process are:
1.) Business Understanding
2.) Data Understanding
3.) Data Preparation
4.) Modeling
5.) Evaluation
6.) Deployment
With lots of details included in each segment. Each element makes excellent points about what is to be done, what resources are needed, who is responsible and where the results go. I see that SPSS Modeler (formerly Clementine) still embraces the concept. Does anyone still teach this?
1.) Business Understanding
2.) Data Understanding
3.) Data Preparation
4.) Modeling
5.) Evaluation
6.) Deployment
With lots of details included in each segment. Each element makes excellent points about what is to be done, what resources are needed, who is responsible and where the results go. I see that SPSS Modeler (formerly Clementine) still embraces the concept. Does anyone still teach this?
Thursday, August 27, 2015
Viz Description of Machine Learning
From R2D3:

Very nicely designed scrolling and animated visual of specific machine learning methods. An introductory, intuitive and clear look, with no need for math. This appears to be only the first edition of this, there is more to machine learning. How about drilling through, if desired, to some of the math, to link it to the graphics? Animate data gathering methods? Testing methods?
I particularly like the visual description of building a decision tree and then visually simulating results of the prediction. I used regression trees many times for applications. Would have been a very effective way to show it someone, with their own data. Look forward to more. Check out what they have done so far. Link above.
" .... In machine learning, computers apply statistical learning techniques to automatically identify patterns in data. These techniques can be used to make highly accurate predictions.
Keep scrolling. Using a data set about homes, we will create a machine learning model to distinguish homes in New York from homes in San Francisco .... "
Very nicely designed scrolling and animated visual of specific machine learning methods. An introductory, intuitive and clear look, with no need for math. This appears to be only the first edition of this, there is more to machine learning. How about drilling through, if desired, to some of the math, to link it to the graphics? Animate data gathering methods? Testing methods?
I particularly like the visual description of building a decision tree and then visually simulating results of the prediction. I used regression trees many times for applications. Would have been a very effective way to show it someone, with their own data. Look forward to more. Check out what they have done so far. Link above.
" .... In machine learning, computers apply statistical learning techniques to automatically identify patterns in data. These techniques can be used to make highly accurate predictions.
Keep scrolling. Using a data set about homes, we will create a machine learning model to distinguish homes in New York from homes in San Francisco .... "
Saturday, February 14, 2015
Predictive Regression Trees
Regression trees were a favorite method for prediction in the enterprise. Easy to use and also easy to explain to a decision maker. Investigating their use in open source solutions. Here is one R library example. Technical warning, but an indication of the variety of solutions that are available. This was not the case a decade ago. We are moving quickly, but the space is thus more complicated. More tech depth here.
Wednesday, December 05, 2012
IBM/DemandTec Assortment Optimization
Early in my career I was given the task of how to optimize assortments under the constraints of the market shelf. On the surface the problem looks easy. You have a set of possible products you want to place on shelf. You have the value of each of the products to the manufacturer or retailer expressed by profit or historical demand, you have the constraints of space on the shelf. You may also have design constraints that tell you what products 'should' be placed next to others based on consumer behavior understanding. Plus a lot of additional and sometimes subtle constraints. All this can be expressed mathematically, using well known methods that have been known for a long time. But it turns out the resulting expression is complex and difficult to solve. We explored how to do this and found it daunting. So I was very impressed when I first saw an example of how DemandTec did assortment optimization for retailers. Their method cuts to the chase to increase profitability for retailers based on historical demand while also including the shopper in the model. They write:
" ... No more guesswork. The best assortment for each location. The best assortment is the one shoppers expect to see on your shelves, but also the one that delivers the best profits. So why make guesses? IBM DemandTec Assortment Optimization offers a collaborative approach for defining localized merchandise assortments through a complete understanding of what drives buying decisions – incorporating shopper demand, space, productivity, and profitability. .... Now you can offer items that meet unique demands or targeted shopper segments. You'll know exactly which items you can replace or substitute for other items in your portfolio. Most importantly, you'll understand the implications of various assortments on each customer segment. With one common view of the shopper, their decision trees, and the marginal additional value of each SKU to the equation, manufacturers and retailers can work together more effectively to delight shoppers and build customer loyalty. ... "
" ... No more guesswork. The best assortment for each location. The best assortment is the one shoppers expect to see on your shelves, but also the one that delivers the best profits. So why make guesses? IBM DemandTec Assortment Optimization offers a collaborative approach for defining localized merchandise assortments through a complete understanding of what drives buying decisions – incorporating shopper demand, space, productivity, and profitability. .... Now you can offer items that meet unique demands or targeted shopper segments. You'll know exactly which items you can replace or substitute for other items in your portfolio. Most importantly, you'll understand the implications of various assortments on each customer segment. With one common view of the shopper, their decision trees, and the marginal additional value of each SKU to the equation, manufacturers and retailers can work together more effectively to delight shoppers and build customer loyalty. ... "
Sunday, October 09, 2011
Viewing Assortment Optimization Through Shopper Decision Trees
We were always looking at ways to analytically describe which products should be included in retail and where they should be placed.
It can be analytically described, but the resulting number of combinations are daunting. We did work in this area for parts of categories, using IBM's behemoth MPSX package, as early as the late 60's.
When I arrived much later on I was asked to look at the systems involved. You could not solve all of the problem, but you could satisfactorily do very useful things for parts of the shelf and aisle. Then use them as building blocks to understand the entire store.
It was understood early on that such planning needed to be done collaboratively. In particular between humans acting as designers, and computers that could handle the sheer quantity of information involved. It was also clear that there was a need to have long term design stability, so as to not confuse the shopper and effectively support brand loyalty.
This turned out to be done best my having people adjust the computer generated design. Then re-evaluate the results. You also needed to think about how the shopper reacted when their favorite item was moved or disappeared. Where would they go? What was next? Think of it as a decision tree of choices. That is what we ultimately considered.
The complete implications of any particular assortment choice was often not considered completely from the perspective of the shopper.
This was where I particularly enjoyed DemandTec's overview, and their view of " ... the shopper, their decision trees, and the marginal additional value of each SKU to the equation, manufacturers and retailers can work together more effectively to delight shoppers and build customer loyalty ... ".
Ultimately it is about the shopper and how they interact with shelf and purchase. The right assortment produces value for retailer, manufacturer and shopper.
See their overview, and their Assortment optimization brief.
Monday, March 21, 2011
Thinking Too Much
In HBS Working Knowledge: Are We Thinking Too Little, or Too Much? With a retail example. I agree that it is a context thing, but the recent writing about the value of snap 'blink' decision making is troubling. It depends on the nature of the decision.
" ... "And then there's this whole stream of research about ways in which you should think more carefully in more logical ways—creating decision trees that map out 'if you want to do this, then you should do this and not that,' making lists of the pros and cons and making a decision based on which list is longer, and so on." - However, there has been little research that considers the notion that overthinking a decision might actually lead to the wrong outcome. Nor have researchers come up with a model that explores how to determine when we're overthinking a decision—even though logic tells us that there certainly is such a thing ... '
" ... "And then there's this whole stream of research about ways in which you should think more carefully in more logical ways—creating decision trees that map out 'if you want to do this, then you should do this and not that,' making lists of the pros and cons and making a decision based on which list is longer, and so on." - However, there has been little research that considers the notion that overthinking a decision might actually lead to the wrong outcome. Nor have researchers come up with a model that explores how to determine when we're overthinking a decision—even though logic tells us that there certainly is such a thing ... '
Monday, March 08, 2010
Statistics and Behavioral Modeling
Intriguing article, by Steve Miller. I have now dealt with these 'two cultures' a number of times in various enterprises. This article and the included link are worth the read ... about the application of traditional statistics in behavioral realms ...' ... My take on the article was less literal. I think the author's point that traditional statistical models might not be up to the task of predicting the realities of human behavior is quite valid. In fact, one of the giants of the statistical world, the late UC Berkeley professor Leo Breiman, originator of Classification and Regression Trees (CART) and Random Forests, said as much in a provocative 2001 article, Statistical Modeling: The Two Cultures, in which he criticized the statistics status quo. The abstract for this paper is telling:
"There are two cultures in the use of statistical modeling to reach conclusions from data. One assumes that the data are generated by a given stochastic data model. The other uses algorithmic models and treats the data mechanism as unknown. The statistical community has been committed to the almost exclusive use of data models. This commitment has led to irrelevant theory, questionable conclusions, and has kept statisticians from working on a large range of interesting current problems. Algorithmic modeling, both in theory and practice, has developed rapidly in fields outside statistics. It can be used both on large complex data sets and as a more accurate and informative alternative to data modeling on smaller data sets. If our goal as a field is to use data to solve problems, then we need to move away from exclusive dependence on data models and adopt a more diverse set of tools.” ... '
Sunday, December 27, 2009
Quantum Machine Learning
Google Research blog on quantum machine learning. Machine learning is the essence of artificial intelligence ... make a machine that learns and you can set it loose and wait for intelligent results. Trouble is, up to now its worked in only the most focused domains. Our own experiments in this domain took very simple tables of results and created editable decision trees. This worked. New ideas here are always welcome.
Sunday, August 10, 2008
Classification Trees
Classification or decision trees are a method to deliver decisions by making choices progressing through a set of predefined paths in the form of a tree . Expert systems can be constructed from extracting rules from experts. They can also be constructed statistically by taking a number of examples and using them to build a system that makes decisions as consistent as possible with the examples provided. As part of our AI development system, meant to be used by internal analysts to construct expertise-based systems, we included a system that used data and made rules. In general, a classification system is one that produces a discrete value and a regression tree produces continuous values. This is a simple form of machine learning. The Wikipedia has a nice overview. Here is a good survey of software that does this.
Subscribe to:
Posts (Atom)
