/* ---- Google Analytics Code Below */
Showing posts with label Outliers. Show all posts
Showing posts with label Outliers. Show all posts

Sunday, May 02, 2021

Simple Examples of Anomaly/Outlier Detection

 A common need,  nicely and simply put here in KDNuggets.  Very classic and should be done with most every set of data you are seriously working with.  With more tech and coding at the link:

Four Techniques for Outlier Detection

Tags: DBSCAN, Knime, Outliers, Python

There are many techniques to detect and optionally remove outliers from a dataset. In this blog post, we show an implementation in KNIME Analytics Platform of four of the most frequently used - traditional and novel - techniques for outlier detection.   By Maarit Widmann, Moritz Heine, Rosaria Silipo, Data Scientists at KNIME

Anomalies, or outliers, can be a serious issue when training machine learning algorithms or applying statistical techniques. They are often the result of errors in measurements or exceptional system conditions and therefore do not describe the common functioning of the underlying system. Indeed, the best practice is to implement an outlier removal phase before proceeding with further analysis.

But hold on there! In some cases, outliers can give us information about localized anomalies in the whole system; so the detection of outliers is a valuable process because of the additional information they can provide about your dataset.

There are many techniques to detect and optionally remove outliers from a dataset. In this blog post, we show an implementation in KNIME Analytics Platform of four of the most frequently used - traditional and novel - techniques for outlier detection.  .... " 

Monday, February 24, 2020

Improving Machine Learning

Treatment and analysis of outliers and noise are key to improve these methods.   The claim here is that this will be improved.  Remains to be seen how well this will work.  Technical.

Mathematicians propose new way of using neural networks to work with noisy, high-dimensional data   by RUDN University 

Mathematicians from RUDN University and the Free University of Berlin have proposed a new approach to studying the probability distributions of observed data using artificial neural networks. The new approach works better with so-called outliers, i.e., input data objects that deviate significantly from the overall sample. The article was published in the journal Artificial Intelligence.  ... "

Saturday, June 01, 2019

Managing Bad Data

Self correcting data is a good idea, but want to see some examples.   Its similar to removing outliers. It depends on the context and often metadata involved.   Have seen it improperly used. 

Researchers Develop AI Tool Better Able to Identify Bad Data
University of Waterloo News

An international team of researchers led by Alireza Heidari and Ihab Ilyas at the University of Waterloo in Canada has developed an artificial intelligence-powered system to manage data quality. The HoloClean tool sifts out bad data and corrects errors prior to processing. The new system also can automatically generate bad examples without tainting source data, so the system can learn to identify and correct errors on its own. Once HoloClean is trained, it can independently differentiate between errors and correct data, and determine the most likely value for missing data if an error exists. Ilyas said the work “deviates from the old way of manually trying to clean the data, which was expensive, didn’t scale, and does not meet the current needs for cleaning the data.” ... '

Tuesday, March 27, 2018

Visualizing Outliers

Nicely done, non-technical piece

Visualizing Outliers  by Nathan Yau in Flowingdata
Visualizing data that looks like it came straight out of Statistics 101 text book is nice and all — for teaching and learning purposes. You gotta learn to stand before you can run a marathon. Once you’re ready for the real data though, which is fuzzier and more irregular, you run into data points that don’t quite fit in with the rest. The outliers.

There are various ways to incorporate outliers into your visualization, but you have to understand them first.

Why is the outlier there in the first place? Maybe it’s a recording error or a kink in methodology. For example, PornHub claimed that a disproportionate percentage of traffic came from Kansas. However, location was based on IP addresses, and any locations that could not be identified defaulted to the center of the country. That spot was in Kansas.

Sometimes outliers might be an exception or something extraordinary. We see this in sports a lot, like when Stephen Curry broke the single-season three-point record. Or when Usain Bolt ran faster than everyone.

In one case, the outlier is noise relative to the rest of the data. In another the outliers deserve a closer look.

With your own data, figure out which is which first. Then decide if the outlier belongs in the background or foreground. The visualization options below will be much more useful. ... " 

Wednesday, February 08, 2017

Outlier Detection in Time Data

Outlier detection and analysis is an important topic for any kind of data.  Outliers, depending on their context, can be useful, or dangerous in the development of models.    New edited:  Book: Outlier Detection for Temporal Data    by Manish Gupta, Jing Gao, Charu Aggarwal, Jiawei Hao.  Review posted by Vincent Granville

Saturday, July 16, 2016

The Importance of Carefully Considering Outliers

Good short piece by Rick Delgado on the concept of outliers.  Very important that both the analyst and the decision maker understands this idea as it relates to collected data and decision problem.  Frequently brought up here (See below)  and in every consulting connection.  The article is non technical.

Tuesday, July 12, 2016

Is Machine Learning Driving AI?

Always thought ML was part of AI.   In Datanami: An interesting distinguishing point is made

" .... The authors quotes Jaime Carbonell, a professor of computer science at Carnegie Mellon University, as noting” “Machine learning essentially is the engine that is driving modern artificial intelligence.”

Machine learning “often deals with unbalanced data sets in which the ultimate focus of decision making is precisely the outlier cases,” that is, the extreme cases that contrast sharply with typical data and provide opportunities for the “most important learning opportunities.”

Ignoring the rare cases, or outliers, means “you cold miss everything [in a data set] that is interesting,” Carbonell asserted. .... " 

Sunday, May 29, 2016

Mapping Entrepreneurs

Lone entrpreneurs?  Bain maps them.  Surprising truths.  From the WSJ.  

" ... Far from being disconnected outliers, entrepreneurs thrive best when they operate in tight-knit networks, nurture fellow risk takers, and trade know-how and capital.

The Endeavor organization together with my colleagues at Bain & Co. mapped this social network of entrepreneurs across multiple generations and multiple continents—even in some of the harshest terrain for innovation, such as Buenos Aires, Istanbul and Mexico City. For example, they surveyed more than 200 Argentine entrepreneurs, asking such questions as:

Who inspired you?
Who invested in your company?
Who mentored you?   .. " 

Friday, September 19, 2014

Countering Fraud With Big Data

Obvious application.  Having more data lets you more easily extract outliers and patterns of interest.  Find correlations that don't need to be causation.  This has been done long before data was big, or even commonly available.  It can also lead you to patterns that I was reminded of by a practitioner of compliance earlier this week.  Data can be too good, track other indicators too well.  Indicating a manipulation to influence a result.

Tuesday, September 09, 2014

Visual Control of Big Data

Visual Control of Big Data - MIT    Very interesting new development.

" ... The Database Group at MIT’s Computer Science and Artificial Intelligence Laboratory has released a data-visualization tool that lets users highlight aberrations and possible patterns in the graphical display; the tool then automatically determines which data sources are responsible for which. ... 

It could be, for instance, that just a couple of faulty sensors among dozens are corrupting a very regular pattern of readings, or that a few underperforming agents are dragging down a company’s sales figures, or that a clogged vent in a hospital is dramatically increasing a few patients’ risk of infection. ...  

  ... The idea of provenance tracking is not new, but Wu’s system is particularly well suited to the task of tracking down outliers in data visualizations. Rather than simply telling the user the million data entries that were used to compute the outliers, it first identifies those that most influenced the outlier values, and summarizes those data entries in human readable terms. ... " 

Thursday, August 07, 2014

Text Analytics for Everyone

In UXMag: Our own early experience with text analytics was that it can be readily done with today's software, but coming to specific conclusions can be more difficult.  Using some of the visual methods shown can provide starting points for discussions about the data to understand it better.  We found in particular that descriptive outliers were useful to re-calibrate our understanding of reactions to products and fixtures.

Wednesday, February 26, 2014

A Mainstream Internet of Things

In MIIT Technology Review:  Short piece on the mainstreaming of the IOT.  I agree, starting in areas were sensors are needed to monitor and look for outliers.  The analytics will quickly combine this data with contexts and other data points, leading to simulations of states in the real world. It makes much sense, and it can start very simple.

Thursday, December 05, 2013

Product Development Using Structured Analogies

Just received this from SAS.  Note that it is an ad and you have to provide information to get the background paper, have not done that yet.  Click here for more.  I am intrigued because we experimented with a similar approach a decade ago, and analogies are powerful ways to generate new product ideas.  May be something useful here.

" ... New Product Forecasting Using Structured Analogies

Learn about a new patent-pending approach that may be helpful in certain new product forecasting situations. Make manual overrides to the statistical forecasts, and get a better sense of the risks and uncertainties in new product forecasts through visualization of past new product introductions. ...  They describe it further:

SAS has a new patent-pending approach to NPF that combines the use of analogies with structured judgment. This "structured analogy" process for new product forecasting has six main steps:

Query step: Find a set of candidate products that have similar attributes to the new product.

Filter step: Manually remove inappropriate or outlier products from the set of candidate products.

Cluster step: Cluster the candidate products according to their sales pattern, and manually select the most appropriate cluster to serve as the surrogate products.

Model step: Select the most appropriate statistical model for the cluster of surrogate products, and extract the statistical model features.

Forecast step: Use the extracted statistical model features to forecast the new product. ... " 

Monday, December 31, 2012

Automatic Monitoring of Surveillance Cameras

Out of work at MIT.  A classic problem.  Think of this also as an event analysis situation where we are seeking changes and outliers and need to search possibilities quickly.   " ... Now a system being developed by Christopher Amato, a postdoc at MIT’s Computer Science and Artificial Intelligence Laboratory (CSAIL), can perform this analysis more accurately and in a fraction of the time it would take a human camera operator. “You can’t have a person staring at every single screen, and even if you did the person might not know exactly what to look for,” Amato says. “For example, a person is not going to be very good at searching through pages and pages of faces to try to match [an intruder] with a known criminal or terrorist.”  .. Existing computer vision systems designed to carry out this task automatically tend to be fairly slow, Amato says. “Sometimes it’s important to come up with an alarm immediately, even if you are not yet positive exactly what it is happening,” he says. “If something bad is going on, you want to know about it as soon as possible.” ... '

Tuesday, June 26, 2012

Global Food Supply Chain Improvement and Protection

I just got a chance to interview N2N Global.  A business intelligence and analytics provider for food supply chains.   Aimed at Food Safety, Quality Assurance and supply and product traceability systems. Their CEO Ernesto Nardone is a former IBMer who worked in the area of using expertise-based systems that led to some of IBM's recent AI applications like Watson.

N2N is using IBM's Cognos to do Business intelligence and related basic analytics to utilize the huge amounts of data, millions of records per month, for the Food Coop Cherry Central.   Cherry Central has a space on Facebook which succinctly describes the work done.   See also this press release for more detail, which emphasizes how this relates to consumer food safety.

In particular I was interested in how this data can be connected to the decision process.  They are doing this now by looking at outlier analysis, reporting and alerting.  A likely good next step is to include process models that specify set of rules based on the knowledge of food experts, and also including quantitative analytical processes that can find patterns, predictively forecast future situations and bring together this intelligence to make next steps safer and more efficient.  Bring on the tool set to make this happen.  That leads to smarter commerce.

Cherry Central and N2N won IBM's Engine of the Week for midsize businesses.    A part of the Midmarket initiative. " ..... This honor is given to midsize businesses that have transformed themselves using insight and IBM technology.... " .  Their transformation has started with the automating the large amount of paper that was once needed to run any complex business, large or small.  It will be instructive to continue to follow their effort as food knowledge is included to make the enterprise truly intelligent.

Monday, April 16, 2012

Visualizing Risk

A study in visualizing risk, from Forbes.   Very complex dashboards of data are shown.  This gets back to using the human visual system as a trend and outlier recognition engine.   While the eye cannot directly determine the complexity of a statistical inference, it can select among possible changes, especially in near real time, to determine where the deep dive should take place.  So why not connect the eye to the algorithm in new and fruitful ways?

Saturday, June 04, 2011

Crowds Need Independence for Wisdom

More evidence that crowds need to retain independence to maintain the 'outlier canceling' effect that can make them useful to solve problems. Obvious, but important to remember. Our own experience was that it was easy to have influential individuals skew the results. Influential could mean many things, including well spoken, being well known or unknown, management role or apparent group role. Bottom line, anonymous worked better. Avoid even clarifying discussions.

Tuesday, May 10, 2011

Messy Analytics

The conclusion of a three part series on practical aspects of using analytical methods by Frank Buytendijk.   I had missed the first two parts, but now plan to go back and read them.  " ... First, when you do statistical analysis, resist the temptation to remove the outliers. Improbable scores or data are usually filtered out of the dataset because it is noise "messing up" the model. However, the outliers might actually represent the most interesting bits. They could be the early warning signal for a black swan coming or could represent new business opportunities that others – following best practices –neatly filter out. If the model is your lens, you won't see any change coming. You won't get any weird new ideas. What you see is what you've always seen. All the model does is confirm your hypothesis. Outliers deserve extra attention.  ... "

Monday, April 11, 2011

Business Intelligence With the Cloud and Internet as Data

Always intriguing Recorded Future blog spins some ideas about the Internet as data.  I have had the opportunity to commission and examine some very large scale econometric based simulation models that used the Internet in part as a source of data. Simulations, Statistical explorations, integration of human intelligence are all possible.   Its an idea worth looking at.  Yet using the Internet also requires the careful examination of the quality of data involved.  That carefully done, there are some clear possibilities to be explored.  Recorded Future itself is an example of how this can be done.

" ... the next compelling step is when we realize that the big breakthrough is not to put traditional BI software in the cloud but to realize that the most compelling data source in itself is the Internet. The amount of true business intelligence we can extract from the “open internet” – in everything from government filings, mainstream news, blogs, twitter, etc. is staggering. And don’t think about this as navigating our way to the right article (i.e. glorified Google News) but real analysis – find patterns, trends, clusters, outliers, anomalies, etc. ... "

Friday, June 04, 2010

Tour Through a Visualization Zoo

From CACM: A Tour Through the Visualization Zoo,
A survey of powerful visualization techniques, from the obvious to the obscure...
The goal of visualization is to aid our understanding of data by leveraging the human visual system's highly tuned ability to see patterns, spot trends, and identify outliers. Well-designed visual representations can replace cognitive calculations with simple perceptual inferences and improve comprehension, memory, and decision making. By making data more accessible and appealing, visual representations may also help engage more diverse audiences in exploration and analysis. The challenge is to create effective and engaging visualizations that are appropriate to the data ... '
.

I like this, but I repeat ... keep it simple when you visualize!
.