/* ---- Google Analytics Code Below */
Showing posts with label unstructured. Show all posts
Showing posts with label unstructured. Show all posts

Sunday, October 04, 2015

Clinical Machine Learning at NYU

Examining more Closely.
Via Principal Investigator David Sontag at NYU:

 Clinical machine learning
Our group is particularly interested in machine learning problems motivated by clinical medicine. We work on algorithms for electronic phenotyping in electronic health records, natural language processing from clinical text, disease progression modeling, and predictive analytics on health insurance claims. Our aim is to develop robust methodologies that work directly with the unstructured data found in electronic medical records, and which generalize between institutions without significant manual effort. We collaborate closely with the Emergency Medicine Informatics Research Lab at Beth Israel Deaconess Medical Center and with Independence Blue Cross. .... " 

See my report on his September 17 talk.  Includes detailed slides.

Sunday, September 20, 2015

On the Mess of Unstructured Analytics

Dealing with the mess.  Agree that this should not be.  But we have more different kinds of unstructured data today than we ever had.  In face we dealt with it since the content analytics days.   Some very good points are made here.  But I warn that organizational changes are harder to make than statistical choices.

"   .... There’s a reason for this: Because the data is poorly managed to begin with. Businesses are treating analytics as a separate business function from data governance, when it’s actually fundamentally dependent on it. Analysis occurs downstream; so by neglecting initial information governance infrastructure and practices, the enterprise is essentially sampling tiny random buckets of data from a whitewater river of information.

Many firms struggle to manage or even understand what sort of unstructured content they even have, let alone begin to effectively manage it. History is partially to blame, to be sure; most attempts at managing unstructured content were hastily prompted by waves of regulatory and legal reform that demanded immediate action. A reactive response was triggered, and many of those initial “band-aid” information management fixes remain in place today. Simply scratching beneath the surface often reveals a tangled mess of siloed    .... " 

Saturday, July 11, 2015

Contract Analytics

Recent post in Prism Legal.   Seems a natural text analytics application, with the potential to link to trend analysis and structured data from a number of legal contexts.

" ... IBM Watson for Contract Analytics at Legal OnRamp

In a prior post on IBM Watson, I noted that Legal OnRamp uses Watson “to process and understand high volumes of contracts.” I describe here more about this and provide context about Watson and contract analytics.

Legal OnRamp Uses IBM Watson to Analyze Contracts

OnRamp describes itself as offering “hosted software and services to help legal departments deliver results more efficiently with higher quality and lower cost.” I recently spoke with Paul Lippe, CEO (and former GC of Synopsys, an AI software company for computer-aided design), and other management to learn more about their use of IBM Watson.  ... " 

Friday, February 13, 2015

Text Signals as Sentiment

An area that a number of groups are looking at.

" ... In this burgeoning era of big data, a substantial majority of all the data is unstructured. Much of this unstructured data is textual, such as the data in reports, articles, emails, tweets, and even conversations or support calls recorded in textual transcripts. Because some of this information is perishable, the capability to process it quickly—in many cases, in real time or near-real time—is becoming quite important to enterprises. This processing requires text analytics capabilities.

What does the capability to perform text analytics mean? One simple example is processing a social media feed, such as Twitter, to extract any tweet that mentions a specific element of data such as a company name or a product. This simple approach to matching keywords can provide a quick glimpse into the presence of a product name in the public’s mind. Further, using Twitter-based metadata, for example, creates the possibility to divide this public perception into regions. ... "

Sunday, January 18, 2015

Structural Data Objects

Interesting piece.   I recall working on something like this for the construction of a competitive analysis database.    Though it was not called a structural data object.   Maintenance of an unexpectedly varying structure became an issue.  Detailed piece.

Wednesday, June 25, 2014

No Such Thing as Unstructured Data? Define it With Care.

Right,  the term is broadly used.  It really means in many cases 'not structured enough for the purpose at hand'.  In a recent architecture project fo BI it meant 'not in a corporately managed database'.  In other cases it means 'free text data'.  And in other examples image data.   Or  'not in an established database field'.  It is a good idea to understand what it means in your example.  TDWI makes some related points.

Wednesday, June 18, 2014

Sparseness of Linguistic Big Data

Interesting comment that we can confirm, but it's not unlike many events, statistically rare events are often problematic.  In Language Log:    At first I disagreed, we generate huge amounts of textual content every day.  But consider:   " ... Big data is at its best when analyzing things that are extremely common, but often falls short when analyzing things that are less common. For instance, programs that use big data to deal with text, such as search engines and translation programs, often rely heavily on something called trigrams: ..... " 

Thursday, May 01, 2014

Business Relevancy of Corporate Data

By Data Guru Bill Inmon.  An instructive, non technical view of corporate data in multiple parts.   Well worth a read.    " ... There is a tremendous difference in business relevancy when looking at the different types of data found in the corporation. With structured data, nearly all data is business relevant – or at least potentially relevant. With unstructured repetitive data, hardly any data is business relevant. And with unstructured nonrepetitive data, some moderate percentage of data is business relevant. ... " .   An interesting statement worth pondering.  Hardly any data of one type may be relevant, but what is may be very valuable.

Thursday, April 17, 2014

Intro to Unstructured Data

Good short piece on the increasingly important topic, by Radhika Subramanian.   The article also points to a free eBook on the topic.  Have not looked at that yet.   Its also important to understand two conundrums on this. One is that text is not the only unstructured data.  And second, that text has structure, notable from the increasing amount of natural language processing being used.  Language contains structure.

Monday, January 13, 2014

Visualizing Unstructured Data

Good overview of the topic.  Examples like 'word clouds' are typical.  And semantic nets that we used to represent the interaction between terms.   Concept maps can also be loaded with information based on semantic extracts of text databases.  Remember, too, that unstructured is more than just text.  It can,for example,  include images and audio data.  Also, unstructured data includes tags and related structured data.  Or geographic tags.  So it is usually semi-structured when it is finally leveraged.

Wednesday, December 11, 2013

Ontology and Big Data

An interesting case of combining Ontologies and big data, in a slide show. Using health data as an example. Ontologies are shared vocabularies as they relate to the use of terms to represent knowledge.  Good look at better organizing unstructured data.

Wednesday, November 13, 2013

DB Pedia

Can a semi-structured text database like the Wikipedia be used to extract structured data?  This is the goal of the DbPedia: 

" ... DBpedia is a crowd-sourced community effort to extract structured information from Wikipedia and make this information available on the Web. DBpedia allows you to ask sophisticated queries against Wikipedia, and to link the different data sets on the Web to Wikipedia data. We hope that this work will make it easier for the huge amount of information in Wikipedia to be used in some new interesting ways. Furthermore, it might inspire new mechanisms for navigating, linking, and improving the encyclopedia itself. ... "

Is this effort still active?  How has it been used to provide real value?

Wednesday, October 23, 2013

More on Store Surveillance

The idea of doing 'visual analytics', has been growing, a new effort:

" ... Video analytics startup Prism Skylabs announced today that it has raised $15 million in Series B funding. ....  the company says it can provide graphics showing footpaths through the store, heat maps of customer interest, and customer counts and conversion. It supposedly works with more than 80 customers.... " 

Note this an example of what is called unstructured data, which often means text, but can mean other kinds of sensor acquisition, like video.

Thursday, July 25, 2013

Deriving Visual Structure

Every line in the Novel: The Great Gatsby, as an infographic.  The future of writing?  Or a better understanding of structure in unstructured data?   Here mostly a time line.

Thursday, June 06, 2013

What Big Data Means

Bob Inmon provides a useful Venn diagram to describe what Big Data is and how it relates to other kinds of data.  Text is certainly the most important part of unstructured big data.  But what about other kinds of data, like images, audio and brain scans?  Increasingly large and useful areas.

Sunday, March 17, 2013

Mining Suicide Notes

Locally a  good example of predictive text analysis.   Mining a large corpus of suicide notes at Cincinnati's Children's Hospital, to model and predict when suicide attempts actually happened.  Plan to dig deeper into this example. Text examples often depend strongly on the size and generality of the text used.  I also like the fact that, based on the article, the analytics have been built directly into the decision process,

Saturday, February 02, 2013

Building Better Products

Thoughtful piece in GigaOM:   About the process of coming up with new products and extensions.  Starting with knowing what has been done before.    This article derived from experiences at Google and Microsoft. There is almost no excuse for doing that well anymore, with search and data being universal. But especially in cases where multiple people are examining a space, collaborating on product development can be chancy.  See for example recent ways to address collaborative interaction with structured, unstructured and tacit knowledge, in Zakta.com.

Wednesday, October 31, 2012

The Marketing and Video

In eCommerceTimes:  A look at new content marketing.  Video is getting more powerful as a tool.  We have seen this for some time.  The more attention gathering, the more likely the connection.  Video, though is still an unstructured medium, and needs adequate text and tagging support.

Friday, October 05, 2012

Google's Virtual Brain

Google has long been known for working on techniques that allow the automatic analysis of unstructured data like images and video.  A classic element of human intelligence.   Here is more on what they are up to.  It brings up new kinds of worries about privacy, but this progress is inevitable.

Monday, September 03, 2012

IKEA Embraces Augmented Reality

A nice example of the use of augmented reality for a catalog.   A very good place to try the idea first.  Much more at the link.   I plan to try this at the nearby IKEA.  Via Metaio.  " ... Speaking of free download- both iOS and Android users can download the amazing IKEA Catalogue app for free(!) to discover all 43 pieces of augmented and activated content when the publication finally hits mail and IKEA stores around the world. Hybrid data comes from structured sources like enterprise applications and unstructured sources (mainly text), like news feeds, social media, and documents.... "