/* ---- Google Analytics Code Below */
Showing posts with label Data Quality. Show all posts
Showing posts with label Data Quality. Show all posts

Thursday, July 07, 2022

Data Observability Stack

Noting requirements of a 'Data Observability Stack', see link.

IBM acquires Databand to bolster its data observability stack

IBM today announced that it acquired Databand, a startup developing an observability platform for data and machine learning pipelines. In a statement, IBM general manager for data and AI Daniel Hernandez said that folding Databand into IBM's broader portfolio would help the latter's customers better identify and fix data issues including errors, pipeline failures and poor quality. The plan is to expand Databand's observability capabilities for integrations across open source and commercial ...'

Thursday, April 16, 2020

Cyber Warranties

Warranties for data?

Cyber Warranties: Market Fix or Marketing Trick?
By Daniel W. Woods, Tyler Moore

Communications of the ACM, April 2020, Vol. 63 No. 4, Pages 104-107  10.1145/3360310

When buying a second-hand car you are at the mercy of the dealer. The dealer knows which cars were treated well by past owners and which are likely to break down within a few months. When buying an information security product, the vendor has a better idea of how effective the product truly is. In both cases, the seller has information the buyer lacks.


Economists refer to this phenomenon as a market with asymmetric information. Akerlof1 suggested this leads to a "market for lemons" dominated by lower quality goods (aka lemons in the case of used cars). Consumers cannot differentiate between lemons and quality used cars. Akerlof's model suggests only lemons would be sold in such a market.

Car dealers offer warranties to overcome this problem. If the used car breaks down within, say, six months, the dealer must pay for its repair. This discourages dealers from selling lemons with lengthy warranties. Consequently, the length of the warranty provides information about how likely the vehicle is to break down.

Returning to information security, vendors have started attaching cyber warranties to information security products with no additional fee. Will cyber warranties better align incentives in the market for information security products? Or are they marketing tricks riddled with coverage exclusions hidden in the fine print of the terms and conditions?

Might Cyber Warranties Remedy the Market for Lemons?
A natural first question to ask is why warranties might succeed in addressing the market for lemons where other mechanisms have failed. Akerlof identified possible solutions including brand reputation, certification, liability laws, and warranties.

Linking brand reputation to the effectiveness of products is difficult because they appear to be working until an attack succeeds, which happens infrequently. Reputation systems are further limited by commercial sensitivity preventing information from being pooled across organizations. Vendors instead signal quality by speaking at conferences, publishing security research, and through marketing activities. The latter can lead to (arguably deceptive) claims about product functionality that may not reflect reality.

External experts could certify the effectiveness of the product. Past history shows certification firms face incentives to skimp on assessment. A framework for certifying computer systems as secure "motivated the vendor to shop around for the evaluation contractor who would give his product the easiest ride."2 Even if such incentives were overcome, there are difficulties in using laboratory experiments to establish real world security. .... " 

Sunday, March 29, 2020

Data Resources: Our World in Data

As part of a larger project that is looking at Data Sources, Open Source Data, Data as an Asset, Data Quality, Data for Machine Learning, Semantic Data, Knowledge Mapping, Metadata and related topics.  This looks to be a great resource, just examining

Specific Data of the Coronavirus/COVID-19  (Updated frequently)

And via the Center For Disease Conrol:  https://www.cdc.gov/coronavirus/2019-ncov/index.html

Our World in Data:  (Used widely for teaching, research etc) 

About:

Research and data to make progress against the world’s largest problems

Poverty, disease, hunger, climate change, war, existential risks, and inequality: The world faces many great and terrifying problems. It is these large problems that our work at Our World in Data focuses on.

Thanks to the work of thousands of researchers around the world who dedicate their lives to it, we often have a good understanding of how it is possible to make progress against the large problems we are facing. The world has the resources to do much better and reduce the suffering in the world.

We believe that a key reason why we fail to achieve the progress we are capable of is that we do not make enough use of this existing research and data: the important knowledge is often stored in inaccessible databases, locked away behind paywalls and buried under jargon in academic papers. 

The goal of our work is to make the knowledge on the big problems accessible and understandable. As we say on our homepage, Our World in Data is about Research and data to make progress against the world’s largest problems.  ... " 

Wednesday, April 18, 2018

Your Data Has to be Good. Define Good.

Very good piece, covers lots of topics.  Worth reading.   Without enough quality data, you have nothing.  Key elements of process, like getting everyone involved early and often.  Biases mentioned, but the principle kinds of biases are not enumerated, and are often dependent on the business domain.  I like to count through likely biases and specifically test for some of the worst.

If Your Data Is Bad, Your Machine Learning Tools Are Useless    Thomas C. Redman in the HBR

Poor data quality is enemy number one to the widespread, profitable use of machine learning. While the caustic observation, “garbage-in, garbage-out” has plagued analytics and decision-making for generations, it carries a special warning for machine learning. The quality demands of machine learning are steep, and bad data can rear its ugly head twice — first in the historical data used to train the predictive model and second in the new data used by that model to make future decisions.

To properly train a predictive model, historical data must meet exceptionally broad and high quality standards. First, the data must be right: It must be correct, properly labeled, de-deduped, and so forth. But you must also have the right data — lots of unbiased data, over the entire range of inputs for which one aims to develop the predictive model. Most data quality work focuses on one criterion or the other, but for machine learning, you must work on both simultaneously.  .... " 

Monday, April 09, 2018

Free Talk on How Small are our Big Data

The 27th Annual E. Leonard Arnoff/Milton J. Schloss Memorial Lecture on the Practice of Business Analytics

Wednesday 04/18, 2018, 6:00-6:45PM
Kingsgate Marriott Conference Center (Ballroom on the 2nd Floor)
University of Cincinnati
Free & Open to the Public

“How Small Are Our Big Data: Turning the 2016 Surprise into a 2020 Vision”
Xiao-Li Meng, Harvard University 
Whipple V. N. Jones Professor of Statistics
Dean of the Graduate School of Arts and Sciences

Xiao-Li Meng is well known for his depth and breadth in research, his innovation and passion in pedagogy, and his vision and effectiveness in administration, as well as for his engaging and entertaining style as a speaker and writer. Meng has received numerous awards and honors for the more than 150 publications; he has delivered more than 400 research presentations and public speeches. His interests range from the theoretical foundations of statistical inferences (e.g., the interplay among Bayesian, frequentist, and fiducial perspectives) to statistical methods and computation (e.g., posterior predictive p-value; EM algorithm; Markov chain Monte Carlo) to applications in natural, social, and medical sciences and engineering (e.g., complex statistical modeling in astronomy and astrophysics, assessing disparity in mental health services, and quantifying statistical information in genetic studies).

Abstract: The term “Big Data” emphasizes data quantity, not quality.  However, much of the current measures of statistical uncertainties and errors are adequate only when the data are of perfect quality.  We show that once we take into account the data quality, the effective sample size of a Big Data set can be vanishingly small.  Without understanding this phenomenon, Big Data can do more harm than good due to drastically inflated precision assessments that cause gross overconfidence. This overconfidence leads us to be caught by surprise when reality unfolds, as we all experienced during the 2016 US Presidential election. Data from the Cooperative Congressional Election Study (conducted by Stephen Ansolabehere, Douglas River and others, and analyzed by Shiro Kuriwaki), are used to assess the data quality in 2016 US election polls, with the aim of gaining a clearer vision for the 2020 election and beyond.

Sponsor: Department of Operations, Business Analytics, and Information Systems, Lindner College of Business, University of Cincinnati    More.

Saturday, May 27, 2017

Common Big Data Quality Issues

In KDNuggets, a good overview of quality issues, some related to the 'Big' part, but generally applicable to any kind of non-trivial data.  I have found that data quality errors are mostly dealt with afterwards in practice.   If discovered at all.  Worth a closer look.

Friday, March 31, 2017

Reminder of the Need for Data Quality

Am often reminded in visits to the enterprise that data quality is not often enough considered carefully.  You don't have much if you don't have quality.  Consider also the quality and nature of meta data involved.  It is even more important as you move beyond BI to heavily leveraged analytics. Bigger risks emerge.  Having a process model of your data can alert you to issues about quality and change.

Data Quality in BI: It’s More Than Putting Lipstick on a Pig!   Published by Pat Hennel
 Data Quality and BIData quality is one of the biggest challenges that enterprises face when it comes to business intelligence. If the data isn’t accurate, inferior reporting and poor business decisions that can have potentially serious consequences on the entire organization can occur.

When first examining the quality of data as you implement a business intelligence (or BI) solution, there are a number of things that need to be considered and several questions that you need to ask yourself. For example, as Paul Dorsett shared in one of his blog posts, Self-Service BI: Fill It Up!: 
... " 

Sunday, May 08, 2016

Shapes Constraint Language

Brought to my attention.  Technical.  Note the mention of data quality use, using the concept of 'shape',  a recent concern.

" .... The Shapes Constraint Language (SHACL) is an evolving specification produced by the W3C RDF Data Shapes Working Group. TopQuadrant is actively supporting the development of SHACL in the W3C Working Group towards its becoming an official Standard Recommendation alongside RDF, OWL and SPARQL. ....

So what is SHACL? First and foremost SHACL is based on RDF and covers similar ground like RDF Schema and the Web Ontology Language (OWL). SHACL can be used to describe the structure of data – be it stored in RDF or JSON or similar formats. SHACL provides an RDF vocabulary for classes, properties and almost arbitrary integrity constraints that instances need to fulfil. SHACL is not limited to classes and instances, but also includes a more generic concept called “Shape” that can be overlaid on any existing data. Schemas created with SHACL can be shared on the web to communicate the intended structure of your data to other people or tools. SHACL tools can improve Data Quality. .... " 

Sunday, August 09, 2015

Defining Data Quality

Good set of definitions it is worth considering.

" .. To tackle any problem in a systematic and effective way, you must be able to break it down into parts. After all, understanding the problem is the first step to finding the solution.  From there, you can develop a strategic battle plan.

When starting a data quality improvement program, it’s not enough to count the amount of records that are incorrect, or duplicated, in your database. Quantity only goes so far. You also need to know what kind of errors exist to allocate the correct resource.

In this interesting blog by Jim Barker, the different types of data quality are broken down into two parts. In this article, we’ll look closely at defining these ‘types’, and how we can use this to our advantage when developing a budget. ... "

Wednesday, April 01, 2015

Dysfunctional Teams and Healthy Ecosystems

In FastCompany: An interesting view.   Somewhat over segmented.  Have seen some of these examples.   It then suggests healthier ecosystems.  Good idea, but also harder to implement in many cultures.   And where are the rewards to make these systems operate beyond just 'leadership'?

" ... In nature, one quality of healthy ecosystems is their vibrant and interdependent biodiversity, enabling adaptation to environmental shifts and threats. Healthy leadership teams, especially in today’s dynamic world, display divergent perspectives that they respect and value. They avoid getting stuck because they can evaluate options based on data and members’ individual knowledge, and ultimately find places of agreement. They see conflicting opinions as dilemmas to grapple with rather than fights to win. ... "

A Struggle with Data Quality

Organizational struggle with data quality.   I will add,  the worst form of quality is the complete lack of required data and metadata needed for high leverage future decisions.  I have seen this all too often. 

I hate the unnecessary slideshow format, so I avoid linking to them, but this is a favorite topic.  My apologies for the format.

Thursday, February 19, 2015

Stanford DeepDive

Brought to my attention.  Another example of the automation of analytics.

" .... DeepDive is a new type of system that enables developers to analyze data on a deeper level than ever before. DeepDive is a trained system: it uses machine learning techniques to leverage on domain-specific knowledge and incorporates user feedback to improve the quality of its analysis. .... 

 DeepDive is targeted to help user extract relations between entities from data and make inference about facts involving the entities. DeepDive can process structured, unstructured, clean, or noisy data and outputs the results into a database.

Users should be familiar with SQL and Python in order to build applications on top of DeepDive or to integrate DeepDive with other tools. A developer who would like to modify and improve DeepDive must have some basic background knowledge listed in the documentation below. .... 

Thursday, February 12, 2015

Business Impacts of Effective Data

Measuring the result is important, or how would you know you had any?  The means and techniques of measurement are very important. And then what is the value of the data assets being used.   In SyBase.

" ... In a study of over 150 Fortune 1000 firms from every major industry or vertical, we explored issues associated entirely with the lifeblood of today’s enterprises: data. The quality of data, the ability for that data to be accessed wherever and whenever it’s needed, and the relevance of that data in addressing a specific problem were areas of focus in the study – in essence, effective data, and the business implications of greater access to effective data.

The findings, being publicized now for the first time, definitely demonstrate the often dramatic impacts that even marginal investments in information technology can have when that technology addresses data  quality, usability, and intelligence, whether it be using mobility or remote access solutions, analytics or business intelligence solutions, or a combination of the two ... " 

Saturday, November 08, 2014

Why the IBM and Twitter Data Deal

It has been presented as a means to understand the pulse of the planet.   A sensory coup previously impossible.  Connecting the firehose of information created in twitter with the analytics and cognitive capabilities of IBM.  It certainly makes me think.  Is the data represented by Twitter, and other sources, clean enough, precise enough to provide quality predictions?  Good introductory interviews in Fortune,  which don't answer these questions, but address the motivation and possible first steps.  Watching.

Wednesday, October 29, 2014

Water Canary

Brought to my attention as a sensor and data gathering example: The Water Canary.

" ... The Water Canary is an inexpensive water-testing device that makes it possible to collect real-time water quality data from the field. With the push of a button, anyone can measure water quality and share that information with the world. .. 

By placing real-time water quality information within reach, the devices make it possible to quickly identify invisible threats so that appropriate actions can be taken to protect people and ecosystems and prevent hazards from erupting into full-scale emergencies. ... "  

Sunday, September 28, 2014

Is Your Content Compelling?

Could be quite valuable for a number of areas.  This is being experimented with every day in advertising and online.  The data to confirm or deny it is considerable.  Have not read as yet.

  In Fastcocreate:   " ... RivetedA book from cognitive science professor Jim Davies, presents a unified theory of compellingness. ....  What makes for compelling art? Any creator who has given half a thought to paying the rent, or achieving immortality, has considered what makes art sell. We know that the notion of quality--the idea that "the best" art and marketing and media reaches the most people--is insufficient to explain what gives some creations mass appeal. So why do people--large number of people--find books, ads, movies and art works compelling? How can we know, ahead of time, what will pique our curiosity and sustain our interest?   ... " 

Friday, August 22, 2014

Local Short Course on Web Analytics


Just reminded of this, I know the people involved and it will be quality.  Useful for anyone involved in the space.

The UC Center for Business Analytics is pleased to announce its next short course:
Fundamentals of Web Analytics September 11-12, 2014

With over 500 websites created and 2 million Google searches made every minute of every day, there is no shortage of data online.  Understanding how to use all this data to uncover meaningful insights is a challenge faced by many digital marketers and analysts today.  In this 2 day course, we will review the basic principles of analysis particularly in the digital space.  We will discuss web analytics architecture, implementation and overall strategy with the goal of not only providing the best practices and examples of how web analytics can be done right but to empower attendees to be more data-driven with business decisions, particularly in the digital space.

Course Fee:  $595
Location:   Kingsgate Marriott at the University of Cincinnati  151 Goodman Drive, Cincinnati, Ohio 45219
For more information and registration ... 

Friday, June 20, 2014

Building Knowledge Graphs for Assistants

The MindMeld API:  With an interesting mission.  Leading to more intelligent voice?  " .... Social physics is a big data science that models how networks of people behave and uses these network models to create actionable intelligence. It is a quantitative science that can accurately predict patterns of human behavior and guide how to influence those patterns to (for instance) increase decision making accuracy or productivity within an organization. Included in this course is a survey of methods for increasing communication quality within an organization, approaches to providing greater protection for personal privacy, and general strategies for increasing resistance to cyber attack. ... " 

Friday, June 06, 2014

Data Science Research Center

Launched this week by Data Science Central.  Of interest.  Examining it now.  " .... Data Science Central launched this week the Data Science Research Center, a public online resource for practitioners to read, download or publish high quality papers. .... ",  The site already contains a number of sample papers and white papers.   Will describe what I find of interest.

Thursday, May 22, 2014

Real Time Business Intelligence

An important consideration is often what the size of the time step needs to be.  I often have to ask, What does your 'real-time' look like?    In TDWI.    " ... "Real time" of course, means different things depending on users and their circumstances, so expectations need to be managed carefully. Faster always sounds better, and users, being human, are susceptible to the hype of instant gratification. Although Hadoop-related technologies could offer a less-expensive alternative, costs usually rise as organizations push their technology stacks in the direction of real-time data access, analytics, and presentation. Users also need to be informed that data quality will likely fall below standards because there's usually no time to run data quality, profiling, and cleansing processes. ... "