/* ---- Google Analytics Code Below */
Showing posts with label Data Marts. Show all posts
Showing posts with label Data Marts. Show all posts

Friday, July 13, 2018

Analytics Data Catalogs, Approaches, not new

Yes, we know this, and just because some call it AI, does not mean we won't have to gather the data consistently and continually to solve real problems. 

Analytics Industrial Revolution- From The Occult to the Ordinary
By  Snehamoy (Sneh) Mukherjee In Linkedin

Senior Director - Delivering Data Science, Big Data, Machine Learning and Analytics projects for Fortune 500 companies

There is a quiet revolution taking place in the Analytics industry that has the potential to completely turn the industry on its head and the way work gets done in this space. Doomsday pundits have already summoned the evil spirit called AI to put an end to the misery of our uneven paychecks, to be replaced by an Universal Basic Income and some of us have reconciled ourselves to that cruel fate. But before the Apocalypse happens, there is another subtle and continuous tectonic movement happening right under our feet, which if gone unnoticed for long, can catch us in a tidal wave of upheaval in the analytics /machine learning industry.

The typical Analytics (often very eruditely rechristened by brilliant marketers and/or the academia as Machine Learning and a lot would break their heads to prove that the two are different) or a Machine Learning project gets delivered in the following atypical manner in most firms: -

·        Data is pulled from one/multiple tables from a database(s) (by someone who may either be from the client side or by the analytics vendor). It may be a onetime data extraction or if data is needed on a periodic basis, an ETL (Extract Transfer Load or in some cases ELT) process is created to do a batch fetching and processing of files (e.g. weekly transaction data from retail stores, monthly/weekly call data in telecom firms, weekly/daily transaction data in banks) .... " 

Machine Learning Data Catalogs

Makes sense, also connecting the data to its actual meaning,  semantic ontologies,  is also good to do in the same place.      It should be a broader aspect of governance.   It is also an important fundamental aspect of interpretability to know where the data is coming from, what its stability and credibility are.

How the Machine Learning Catalogs Stack Up  
Alex Woodie in Datanami

You can’t do anything with data – let alone use it for machine learning – if you don’t know where it is. In the age of big data, this is not a trivial matter. It is also the main driver that’s propelling the rise of machine learning data catalogs, which the analysts at Forrester recently ranked and sorted. Just a word of warning: the name at the top of the list might surprise you.

According to Michelle Goetz’s June 21 Forrester Wave report, the percentage of analytic decision makers managing more than 1 petabyte of data (either structured, semi-structured, or unstructured) has essentially tripled from 2016 to 2017. That rapid growth has exposed all manner of problems in company’s existing data management and analytic endeavors.

Two of the biggest challenges that companies face today, Goetz writes, are gathering and managing data in a governed manner on the one hand, and managing the business processes that surround the data analytics activities on the other.

“For EA [enterprise analytics] professionals, relying on people and manual processes to provision, manage, and govern data simply does not scale,” the Forrester analyst writes. “Enterprises are waking up to this fact and turning to data catalogs to democratize access to data, enable tribal data knowledge to curate information, apply data policies, and activate all data for business value quickly.”  .... "