Interesting podcast piece in K@W. Some useful cautions. The notorious p-hack is brought up again. Its often used because is so easy to apply. Simplicity Bias It cannot be used alone.
What Marketers Are Doing Wrong in Data Analytics
Podcasts Research North America leveraging-customer-analytics-featured-image
Wharton's Ron Berman explains why most marketers 'p-hack' and why it could lead to wrong results.
(Podcast at the link)
Companies gather and analyze data to fine-tune their operations, whether it’s to help them figure out which webpage design works best for customers or what features to include in their product or service to boost sales. Marketers, in particular, use data analytics to answer questions like this: To put people in a shopping mood, is it better to make the webpage banner blue or yellow? Or do these colors not matter? Getting the answer right could mean the difference between higher sales or losing to the competition.
But new Wharton research shows that 57% of marketers are incorrectly crunching the data and potentially getting the wrong answer — and perhaps costing companies a lot of money. “We expected business experimenters [to make this error], but I was nevertheless surprised that so many of them do so,” said Wharton marketing professor Christophe Van den Bulte, who coauthored the study. Wharton marketing professor Ron Berman, another of the study’s authors, agreed: “This was a pretty common phenomenon that we observed.” (Listen to a podcast interview with Berman about the research at the top of this page.)
Their paper, “p-Hacking and False Discovery in A/B Testing,” which was popularly downloaded and widely cited in social media, looked at the A/B testing practices of marketers who used the online platform Optimizely before the platform added safeguards against potential mistakes. In A/B testing, two or more versions of a webpage are tested to see which one resonates more with users. For example, half of a company’s customers would see webpage version A and the other half version B. “Imagine one version says something about the brand of your product and the other version says something about the technical abilities of your product,” Berman said. “You want to determine which one makes consumers respond better, to buy more of your products.” ... "
Showing posts with label P-Values. Show all posts
Showing posts with label P-Values. Show all posts
Wednesday, August 22, 2018
Saturday, August 12, 2017
Debating Statistical Significance
A considerable look, both technical and non-technical about statistical significance. Have it has been used, and how that is being re-considered. The original title says this is a nerdy debate, I disagree, it is very important. Having replicable significance is essential.
The case for, and against, redefining “statistical significance.”
Updated by Brian Resnick
There’s a huge debate going on in social science right now. The question is simple, and strikes near the heart of all research: What counts as solid evidence?
The answer matters because many disciplines are currently in the midst of a “replication crisis” where even textbook studies aren’t holding up against rigorous retesting. The list includes: ego depletion, the idea that willpower is a finite resource; the facial feedback hypothesis, which suggested if we activate muscles used in smiling, we become happier; and many more.
Scientists are now figuring out how to right the ship, to ensure scientific studies published today won’t be laughed at in a few years.
One of the thorniest issues with this question is statistical significance. It’s one of the most influential metrics to determine whether a result is published in a scientific journal. .... "
The case for, and against, redefining “statistical significance.”
Updated by Brian Resnick
There’s a huge debate going on in social science right now. The question is simple, and strikes near the heart of all research: What counts as solid evidence?
The answer matters because many disciplines are currently in the midst of a “replication crisis” where even textbook studies aren’t holding up against rigorous retesting. The list includes: ego depletion, the idea that willpower is a finite resource; the facial feedback hypothesis, which suggested if we activate muscles used in smiling, we become happier; and many more.
Scientists are now figuring out how to right the ship, to ensure scientific studies published today won’t be laughed at in a few years.
One of the thorniest issues with this question is statistical significance. It’s one of the most influential metrics to determine whether a result is published in a scientific journal. .... "
Monday, March 14, 2016
Caution with the Statistical P Value
Statistical measures like P-Values and R Squares are dragged out to prove a number of things. But caution should be considered. This Nature article does a good job of explaining the needed cautions:
" ... Scientific method: Statistical errors
P values, the 'gold standard' of statistical validity, are not as reliable as many scientists assume.
by Regina Nuzzo ... "
The final quote in the article brings us back to how any study, analytic or statistical should be considered. It is about the process involved.
" .... Statistician Richard Royall of Johns Hopkins Bloomberg School of Public Health in Baltimore, Maryland, said that there are three questions a scientist might want to ask after a study: 'What is the evidence?' 'What should I believe?' and 'What should I do?' One method cannot answer all these questions, Goodman says: “The numbers are where the scientific discussion should start, not end. .... "
" ... Scientific method: Statistical errors
P values, the 'gold standard' of statistical validity, are not as reliable as many scientists assume.
by Regina Nuzzo ... "
The final quote in the article brings us back to how any study, analytic or statistical should be considered. It is about the process involved.
" .... Statistician Richard Royall of Johns Hopkins Bloomberg School of Public Health in Baltimore, Maryland, said that there are three questions a scientist might want to ask after a study: 'What is the evidence?' 'What should I believe?' and 'What should I do?' One method cannot answer all these questions, Goodman says: “The numbers are where the scientific discussion should start, not end. .... "
Saturday, March 05, 2016
Mismeasure of Uncertainty
Stephen Few thoughtfully reviews Willful Ignorance: The Mismeasure of Uncertainty, by Herbert Weisberg.
Modern science relies heavily on an approach to the assessment of uncertainty that is too narrow. Scientists rely on statistical measures of significance to establish the merits of their findings, often without fully understanding the limitations of those statistics and the original intentions for their use. P-values and even confidence intervals are cited as stamps of approval for studies that are meaningless and of no real value. Researchers strive to reach significance thresholds as if that were the goal, rather than the addition of useful knowledge. In his book Willful Ignorance: The Mismeasure of Uncertainty, Herbert I. Weisberg, PhD, describes this impediment to science and suggests solutions. ... "
Modern science relies heavily on an approach to the assessment of uncertainty that is too narrow. Scientists rely on statistical measures of significance to establish the merits of their findings, often without fully understanding the limitations of those statistics and the original intentions for their use. P-values and even confidence intervals are cited as stamps of approval for studies that are meaningless and of no real value. Researchers strive to reach significance thresholds as if that were the goal, rather than the addition of useful knowledge. In his book Willful Ignorance: The Mismeasure of Uncertainty, Herbert I. Weisberg, PhD, describes this impediment to science and suggests solutions. ... "
Friday, October 02, 2015
Questioning Statistical Foundations. Implications for Business?
In the DSC: Some interesting points made, though I can already hear my statisticians moaning. All statistics, all modeling. is meant to be used predictively in business. Back to the question of what you can reasonably predict in a given context. And the risks of a prediction. Let's combine all our resources to do the best job. Diving deeper. Look forward to comments at the DSC link above:
" .... Interesting article by Regina Nuzzo, posted in Nature.com. Indeed, it's not just p-values that are being questioned, but even the Fisher-Neyman-Pearson (FNP) paradigm and the concept of maximum likelihood estimates (MLE). ...
Here's an extract published on the American Statistical Association's website :
Over the same period, but especially since the 1990s, there has been an increasing disconnect between the traditional Fisher-Neyman-Pearson (FNP) math statistics course and the demands for complex analysis in many application areas. The failure of classical maximum likelihood methods to deal effectively with complex models and the success of MCMC-based methods has led to a similar situation: The undergraduate FNP course does not prepare students for these models, and Bayesian MCMC retraining courses are needed to prepare graduates for these applications. .... "
" .... Interesting article by Regina Nuzzo, posted in Nature.com. Indeed, it's not just p-values that are being questioned, but even the Fisher-Neyman-Pearson (FNP) paradigm and the concept of maximum likelihood estimates (MLE). ...
Here's an extract published on the American Statistical Association's website :
Over the same period, but especially since the 1990s, there has been an increasing disconnect between the traditional Fisher-Neyman-Pearson (FNP) math statistics course and the demands for complex analysis in many application areas. The failure of classical maximum likelihood methods to deal effectively with complex models and the success of MCMC-based methods has led to a similar situation: The undergraduate FNP course does not prepare students for these models, and Bayesian MCMC retraining courses are needed to prepare graduates for these applications. .... "
Subscribe to:
Posts (Atom)