/* ---- Google Analytics Code Below */
Showing posts with label Significance. Show all posts
Showing posts with label Significance. Show all posts

Saturday, August 12, 2017

Debating Statistical Significance

A considerable look, both technical and non-technical about statistical significance. Have it has been used, and how that is being re-considered.    The original title says this is a nerdy debate, I disagree, it is very important.  Having replicable significance is essential.

The case for, and against, redefining “statistical significance.” 
Updated by Brian Resnick

 There’s a huge debate going on in social science right now. The question is simple, and strikes near the heart of all research: What counts as solid evidence?

The answer matters because many disciplines are currently in the midst of a “replication crisis” where even textbook studies aren’t holding up against rigorous retesting. The list includes: ego depletion, the idea that willpower is a finite resource; the facial feedback hypothesis, which suggested if we activate muscles used in smiling, we become happier; and many more.

Scientists are now figuring out how to right the ship, to ensure scientific studies published today won’t be laughed at in a few years.

One of the thorniest issues with this question is statistical significance. It’s one of the most influential metrics to determine whether a result is published in a scientific journal.  .... " 

Monday, December 05, 2016

Standardizing Predictive Influence Measures

Determining influence is a very frequent need for analytics in business.  It can often guide us towards data collection.   It is often a prelude to further analysis.  Always thought it would be useful to have a more standard way of measuring this.   Now the I-Score.  Reading the paper and plan to follow.

In PhysOrg News:  Researchers at Princeton, Columbia and Harvard have created a new method to analyze big data that better predicts outcomes in health care, politics and other fields.

The study appears this week in the journal Proceedings of the National Academy of Sciences.  (Abstract of technical paper) 

In previous studies, the researchers showed that significant variables might not be predictive and that good predictors might not appear statistically significant. This posed an important question: how can we find highly predictive variables if not through a guideline of statistical significance? Common approaches to prediction include using a significance-based criterion for evaluating variables to use in models and evaluating variables and models simultaneously for prediction using cross-validation or independent test data.

In an effort to reduce the error rate with those methods, the researchers proposed a new measure called the influence score, or I-score, to better measure a variable's ability to predict. They found that the I-score is effective in differentiating between noisy and predictive variables in big data and can significantly improve the prediction rate. For example, the I-score improved the prediction rate in breast cancer data from 70 percent to 92 percent. The I-score can be applied in a variety of fields, including terrorism, civil war, elections and financial markets.  ... " 

Wednesday, October 26, 2016

Judea Pearl on Engines of Evidence

A favorite researcher on the topic.  How do we understand how evidence models results?  In the Edge: 

Engines of Evidence,  A Conversation With Judea Pearl
A new thinking came about in the early '80s when we changed from rule-based systems to a Bayesian network. Bayesian networks are probabilistic reasoning systems. An expert will put in his or her perception of the domain. A domain can be a disease, or an oil field—the same target that we had for expert systems. 

The idea was to model the domain rather than the procedures that were applied to it. In other words, you would put in local chunks of probabilistic knowledge about a disease and its various manifestations and, if you observe some evidence, the computer will take those chunks, activate them when needed and compute for you the revised probabilities warranted by the new evidence.

It's an engine for evidence. It is fed a probabilistic description of the domain and, when new evidence arrives, the system just shuffles things around and gives you your revised belief in all the propositions, revised to reflect the new evidence.         

JUDEA PEARL, professor of computer science at UCLA, has been at the center of not one but two scientific revolutions. First, in the 1980s, he introduced a new tool to artificial intelligence called Bayesian networks. This probability-based model of machine reasoning enabled machines to function in a complex, ambiguous, and uncertain world. Within a few years, Bayesian networks completely overshadowed the previous rule-based approaches to artificial intelligence.

Leveraging the computational benefits of Bayesian networks, Pearl realized that the combination of simple graphical models and probability (as in Bayesian networks) could also be used to reason about cause-effect relationships. The significance of this discovery far transcends its roots in artificial intelligence. His principled, mathematical approach to causality has already benefited virtually every field of science and social science, and promises to do more when popularized. 

He is the author of Heuristics; Probabilistic Reasoning in Intelligent Systems; and Causality: Models, Reasoning, and Inference. He is the winner of the Alan Turing Award.  .... " 

Saturday, March 05, 2016

Mismeasure of Uncertainty

Stephen Few thoughtfully reviews Willful Ignorance:  The Mismeasure of Uncertainty, by Herbert Weisberg.

Modern science relies heavily on an approach to the assessment of uncertainty that is too narrow. Scientists rely on statistical measures of significance to establish the merits of their findings, often without fully understanding the limitations of those statistics and the original intentions for their use. P-values and even confidence intervals are cited as stamps of approval for studies that are meaningless and of no real value. Researchers strive to reach significance thresholds as if that were the goal, rather than the addition of useful knowledge. In his book Willful Ignorance: The Mismeasure of Uncertainty, Herbert I. Weisberg, PhD, describes this impediment to science and suggests solutions. ... " 

Tuesday, December 08, 2015

Statistical Significance at Play

In Insight@Kellogg. What would the frequentists say then, Kenneth?  .... Blinded by Statistical Significance ... Putting too much stock in an arbitrary threshold may lead to bad decisions. ... Based on the research of Blakeley B. McShane and David Gal. 

Wednesday, September 23, 2015

Future of Excel in Business Intelligence

Via Paulina Gibson from Investintech:

" .... An interview with 27 Excel experts, where they talk about the future of Excel in Business Intelligence. ... "    Some good thoughts here.    Every enterprise uses Excel, so you can't ignore it as an existing consolidation place for data in the enterprise.  So also a place where business intelligence will live.   It will take some time to evolve away from this situation.   And effort's like Microsoft's Power BI will further slow the change.

(Update) I have added the comments on the interviews of my Excel Expert, Walter Riker below:
....

Friday, August 21, 2015

Significance and Uncertainty

In Flowingdata: Short lesson in significance.   Is the p value being misused to determine publishable significance for correlation?  ( I don't know the truth of that statement, but will follow up )  Is this like grade inflation?   We want more people to graduate/publish?   Become satisfied consumers of our product.

Thursday, August 20, 2015

Ode to A&P

In Adage: An Ode to A&P, a Once-Proud Grocery Brand  ...
The Sun Sets on a Chain With Personal Significance

An Ode to A&P, a Once-Proud Grocery Brand
To many people, the A&P brand today probably signifies high prices and middle-of-the-road quality, but it wasn't always that way.   .... 

Friday, July 24, 2015

Statistical Significance

Kaiser Fung takes a core value principle look at the meaning of statistical significance.  His statement is very true, but is it actually useful in practice, except as a general caution?   What does it tell me to do with the models I have derived?   Give me something  I can work with.

Wednesday, December 03, 2014

Big Data and Patents

How Big Data Impacts the World of Patents .... Patents and Intellectual Property are gradually gaining significance around the world. This is leading to a bottleneck–large databases and ever growing information. Read on to find how Big Data analytics can show a way out. .... "   By Mark van Rijmenam

Wednesday, May 14, 2014

Morphological Approaches to Engineering Design

A long time correspondent,  Tom Ritchey, sends a note about an analytical approach that is little used in  enterprise, but should be much better known.  Using the structure, or morphology, of a business decision problem to improve it.  This is a mostly logical rather than numerical approach.    Below a link to a not overly technical example paper on the subject.  Also see my previous writings on General Morphology Analysis.

Short description of method.

SweMorph | Swedish Morphological Society
An interesting example of the early application of General Morphological Analysis to engineering is the work of the British Norris Brothers, who in the 1960s worked with the speed racer Donald Campbell to achieve several world records. In this paper, Asunción Álvarez (Inplanta, Madrid) describes their work and its significance within its historical context as an early example of the “morphological approach” to actual engineering design.

You can download: “The Norris Brothers Ltd. morphological approach to engineering design – an early example of applied morphological analysis”, Acta Morphologica Generalis, Vol.3 No.2 (2014),

Sunday, February 02, 2014

On Fair Use

In TechDirt:   Good extended piece on recent hearings on copyright and fair use.  A concept often made use of in this blog.   " .... It is twenty years since the Supreme Court of the United States handed down its landmark decision on copyright law and the defence of fair use in the "Pretty Woman" case, Campbell v. Acuff Rose Music. Inspired by the jurisprudence of Justice Story and Justice Leval, Justice Souter developed a doctrine of transformative use. His Honour stressed that "the goal of copyright, to promote science and the arts, is generally furthered by the creation of transformative works." Justice Souter observed: "Such works thus lie at the heart of the fair use doctrine's guarantee of breathing space within the confines of copyright, and the more transformative the new work, the less will be the significance of other factors, like commercialism, that may weigh against a finding of fair use."  ... " 

Saturday, August 24, 2013

Games Matter in the Progress of AI

In IBM Research News:  We used a number of game approaches to teach AI.  But business decisions are much harder than games.   So games are pointers to AI, but rarely directly applicable.

" ... Dr. Gerald Tesauro, the IBM Research scientist who taught Watson how to make wagers when its Jeopardy!, has been named an Association for the Advancement of Artificial Intelligence (AAAI) Fellow. His development of TD-Gammon, “a self-teaching neural network that learned to play backgammon at human world championship level,” and work applying machine learning across disciplines from computer virus recognition to computer chess, and other fields made him an ideal candidate for the association’s title.

You’ve worked on machines that play Jeopardy!, chess and backgammon.  What is the significance of machines that can play games?
         
Dr. Gerald Tesauro: 
In the early decades of AI, algorithms were not ready to tackle the ambiguous, ill-defined nature of real-world problems. Researchers therefore proposed that complex board games like chess and backgammon could serve as an ideal testing ground for AI algorithms (the so-called "Drosophila of AI"). Tasks such as playing grandmaster-level chess may be incredibly complex, but they can be precisely specified for the computer.

By working in these domains, researchers made enormous progress in search, learning, and simulation techniques, to the point where the best computers now surpass the best humans in virtually all classic board games. As a result, AI is now moving on to tackle real-world ambiguity head-on.  ... " 

Saturday, July 27, 2013

Mathematics of Poker

Some recent exposure to card game theory led me to this WSJ article.  " ... This growth over the past decade has been accompanied by a profound change in how the game is played. Concepts from the branch of mathematics known as game theory have inspired new ideas in poker strategy and new advice for ordinary players. Poker is still a game of reading people, but grasping the significance of their tics and twitches isn't nearly as important as being able to profile their playing styles and understand what their bets mean. ... " 

Friday, September 16, 2011

Jumping the Shark

I had not known the origin of the term  'jumping the shark',   a perceived instantaneous loss of significance, which has been increasingly used of late.  Discovered its origins in mindless network TV in this Wikipedia article.  Appropriate.

Saturday, April 02, 2011

Statistical Significance

The Number Guy blogs about statistical significance.  A frequently asked question:  Is it significant?  And it turns out often a troublesome concept.  Don't let the professionals fool you, but it is not as easy as what was taught in stat class.  If often hinges on, surprise, the context of the decision problem you are attempting to formulate.

Monday, February 14, 2011

Significance of Watson and Jeopardy

In KurzweilAI:  a good piece on the artificial intelligence of the upcoming Jeopardy challenge using the IBM Watson computer.  Ray Kurzweil on the meaning of computer machine challenges in games and Jeopardy to AI.  I have followed the game playing aspects of AI experimentation since their beginning. Although a powerful experiment in the right direction, it has never met the expectations of delivering the providing the  broad problem solving needs of the enterprise.

Update:  I saw the first half hour of this competition.  Far too small a sample to make much of a conclusion about the generality of the technique.  Watson did not have a connection to the Web, so no outside help.  Though very large databases could have been downloaded and indexed ahead of time. Watson was very impressive, getting what I perceived to be very hard questions and missing some easier ones.  Most impressive the interpretation of natural language.  Would like to see many more examples to see how this might be used say in an interface within the enterprise to answer questions that included both internal and expernal information, using informal business language.

Friday, December 17, 2010

Bradley Efron Says it Simply

A very good interview with Stanford statistics prof Bradley Efron in Significance Magazine.  I was sent a PDF, but I see it is regrettably behind a pay wall. There are some tempting quotes in Numbers Rule Your World.

Monday, June 08, 2009

Public Relations 2.0 Examined

Good detailed article in PR 2.0 by Brian Solis: The State of PR, Marketing, and Communications: You are the Future. ' ... Contrary to popular belief, Social Media isn’t killing PR, but the business of PR IS in a state of paramount crisis. It’s not without merit however. Perhaps up until now, we have been our own worst enemy. The Social Web, the democratization of content and the wisdom of the crowds is merely amplifying PR’s weaknesses and expediting the declination of a broken business model. As is, many of us are collectively contributing to its perceived insignificance and irrelevance ... '

Tuesday, September 16, 2008

A Move from Statistical Significance?

In a recent post by correspondent Steve King: Are Statistically Significant Research Methods Passe?. " ... A growing number of surveys are being based on informal groups of respondents rather than statistically significant population samples ... ". Speaking about the use of questionnaires. I comment at his post as well, where I agree it is a growing trend. It is also like the trend of design versus quantitative analyses. Like the 'Blink' hypothesis. A form of generate-and-test, but do not forget the test. Classic statistically sound methods are still important. It is OK to use roughly exploratory methods to understand the landscape, but not to make the big decisions. See related previous post, which suggests that chat is more important than carefully directed instruments.