/* ---- Google Analytics Code Below */
Showing posts with label Data Scientist. Show all posts
Showing posts with label Data Scientist. Show all posts

Friday, June 07, 2019

Seeking Great Data Analytics

Good piece, which even includes the rare 'decision' aspects included.    But implies the whole process is still very haphazard.  Suppose you wanted the automate the whole process, where would you start?

What Great Data Analysts Do — and Why Every Organization Needs Them
By Cassie Kozyrkov ... '

Monday, March 25, 2019

Bootstrapping Techniques

Useful and good introduction.     Was used long before AI techniques.

The bootstrap. The Swiss army knife of any data scientist
Applications of the bootstrap techniques in R
By Gianluca Malato

Bootstrap in a nutshell
Bootstrap is a technique made in order to measure confidence intervals and/or standard error of an observable that can be calculated on a sample.

It relies on the concept of resampling, which is a procedure that, starting from a data sample, simulates a new sample of the same size, considering every original value with replacement. Each value is taken at the same probability of the others (which is 1/N).  .... "

Thursday, February 14, 2019

Doing Pre Data Science

Nicely put piece,  that I have passed it along before to others,  worth considering again.   Especially good as it includes the business parts more explicitly than you typically.   Prework is a good idea.  But also make sure you have the right team doing the work, especially the business stakeholders, and it should be pre, during and followup work that gets the same attention.

How Do You Win the Data Science Wars?  You Cheat By Doing The Necessary Pre-work!
Posted by Bill Schmarzo  in DSC

I’m sure that most data scientists have experienced that moment when they realize that the folks around them have no idea what they do.  That moment when someone walks up to them and says “I’ve got some data.  Can you do some data science on it?”

Many organizations started their data science journey by hiring a “data scientist” and asking him or her to perform magic on the data.  And while there are countless problems with that approach, companies quickly learned that 1) not everyone who calls themselves a data scientist is a data scientist (I can call myself young and dashing, but that don’t make it so) and 2) there is no magic when it comes to data science. Sorry, but as Chris Rock famously said:

There is no sex in the champagne room…

Data Science is very hard work, requiring experience and expertise in gathering (scraping in some cases) data from a wide variety of poorly documented and hard-to-access data sources; dealing with the incompleteness, inaccuracies, vagueness and poor documentation about the data; massaging, twisting and torturing that data into some useful form; and trying a seemingly endless number of analytic transformations, enrichments and algorithms in an attempt to find those combinations of variables and metrics that might yield a better predictor of performance (see Figure 1).  ... "

Friday, February 08, 2019

Who is a Citizen Data Scientist?

When I saw the term at first, I wondered too.  In part part because of the 'scientist' part.    Does this assume the particular 'method' we know, or does it just mean more casual use?   Its usually also a  journalist as well, because they want to leverage some piece with it.   Not sure I agree that everyone is talking about it, unless you are the target of such a piece.   Does it require open data science?

Who, Who, Who…Are You Citizen Data Scientist?   By Mike Gualtieri, Vice President, Principal Analyst,   Forrester

Ugh. Everyone is talking about the citizen data scientist, but no one can define it (perhaps they know   one when they see one). Here goes. The simplest definition of a citizen data scientist is: non-data scientist. That’s not a pejorative. It just means that citizen data scientists nobly desire to do data science but are not formally schooled in all the ins and outs of the data science lifecycle. For example, a citizen data scientist may be quite savvy about what enterprise data is likely to be important to create a model but may not know the difference between GBM, random forester, and SVM. Those algorithms are data scientist geek to many of them. The citizen data scientist’s job is not data science. Rather they use it as a tool to get their job done.  Here is my definition of the enterprise citizen data scientist:

A business person who aspires to use data science techniques such as machine learning to discover new insights and create predictive models to improve business outcomes.  ..... " 

Monday, January 28, 2019

What exactly is Artificial Intelligence?

I believe I mentioned this piece before, but worth pointing to again.

What Exactly is Artificial Intelligence and Why is it Driving me Crazy

Posted by William Vorhies 

Summary:  Advanced analytic platform developers, cloud providers, and the popular press are promoting the idea that everything we do in data science is AI.  That may be good for messaging but it’s misleading to the folks who are asking us for AI solutions and makes our life all the more difficult.

Arrgh!  Houston we have (another) problem.  It’s the definition of Artificial Intelligence (AI) and specifically what’s included and what isn’t.  The problem becomes especially severe if you know something about the field (i.e. you are a data scientist) and you are talking to anyone else who has read an article, blog, or comic book that has talked about AI.  Sources outside of our field, and a surprising number written by people who say they are knowledgable are all over the place on what’s inside the AI box and what’s outside.

This deep disconnect and failure to have a common definition wastes a ton of time as we have to first ask literally everyone we speak to “what exactly do you mean by that when you say you want an AI solution”.

Always willing to admit that perhaps the misperception is mine, I spent several days gathering definitions of AI from a wide variety of sources and trying to compare them.  ... " 

Wednesday, December 19, 2018

CSI Talk: Teaching Data Science

From last week's CSIG Talk, 

Speaker:  Dr. Mine Cetinkaya-Rundel, Duke University:
Suppose our goal is to educate the new generation of data scientists working on machine learning and artificial intelligence problems, and especially those who are not intimidated by learning new computing technologies. Where do we start their education at the college level? Which topics do we cover in their first course, and which topics do we postpone till later? In this talk, we propose an introductory data science course that places a heavy emphasis on exploratory data analysis and modeling as well as collaboration, effective communication of findings, and ethical considerations as a welcoming and horizon broadening introduction to the discipline at large.

Mine Çetinkaya-Rundel is the Director of Undergraduate Studies and Associate Professor of the Practice in the Department of Statistical Science at Duke University as well as Data Scientist and Professional Educator at RStudio. She is also the creator and maintainer of datasciencebox.org and she teaches the popular Statistics with R MOOC on Coursera as well as numerous courses on DataCamp. ... 

Data Science in a Box: http://datasciencebox.org     

Good inclusion of data visualization, exploratory analysis, decision making, cautions to bias ...

Slides from talk.
Recording of talk.

Tuesday, December 11, 2018

Newsletters from the HBR, Sloan

I see that HBR has set up some 15 technology and management topic newsletters.  This is similar to what MIT has done with a number of topics.  These newsletters are mores scannable than blogs.  Also can be more up to the minute topical.  A look into the future?  Unclear if this is free, but appears to be.   Interesting to this group, this newsletter:

Managing Data Science
An eight-week newsletter on making analytics and AI work for your organization.  ... 

Monday, December 03, 2018

Difference Between Business Intelligence and Data Science

Nicely done piece.  Good charts at the links below.   Though in some ways I have to ask if there should be a difference in practice?   Depending on Goals and potential value?   I always say you should start any 'advanced'  analytics/science with a descriptive examination.    Or you should have the experts of that description closely accessible to your team.  Attitude should be the same:  Achieve the business goal.

Updated: Difference Between Business Intelligence and Data Science

Posted by Bill Schmarzo in DSC.

I'm reposting this blog (with updated graphics) because I still get many questions about the difference between Business Intelligence and Data Science. Hope this blog helps.

I recently had a client ask me to explain to his management team the difference between a Business Intelligence (BI) Analyst and a Data Scientist.  I frequently hear this question, and typically resort to showing Figure 1 (BI Analyst vs. Data Scientist Characteristics chart, which shows the different attitudinal approaches for each)...  " 

Saturday, November 17, 2018

Mindmap for Managers about Data Science

Nicely done,   Though would have liked more connection from the map to details.  But worth a scan.

Intro to Data Science for Managers [Mindmap]  Posted by Igor Bobriakov in DSC

Data science has become an integral part of many modern projects and businesses, with an increasing number of decisions now based on data analysis. The data science industry is experiencing an acute shortage of talents, not only of data scientists but also of managers, having some understanding of analytics and data science. As a manager, you can ultimately become the company's expert in data usage, creating opportunities for the evolution of your organization. Whether you are working with a team of data scientists, as a part of a data-driven business, or you are interested in implementing data science solutions — you shall have some data knowledge and understand its organizational capabilities.   ... " 

Friday, November 16, 2018

Which Data Science Project?

We  dealt with exactly this, with AI projects and with and any new tech projects.    Below article very nicely put piece worth reading.    Even more straight forwardly put:  Start simply, where you have data, know your goals, can define current process and measures.   Increase credibility to leverage new projects.

How to Decide Which Data Science Projects to Pursue   By Hilary Mason in the HBR

In 2018, every organization has a data strategy. But what makes a great one?

We all know what failure looks like. Resources are invested, teams are formed, time goes by — but nothing comes of it. No one can necessarily say why; it’s always Someone Else’s Fault.

It’s harder to tell the difference between a modest success and excellence. Indeed, in data science they can they look very similar for perhaps a year.  After several years, though, an excellent strategy will yield orders of magnitude more valuable results.

Both mediocre and excellent strategies begin with a series of experiments and investments leading to data projects. After a few years, some of these projects work out and are on their way to production.

In the mediocre strategy, one or two of these projects may even have a clear ROI for the business. Typically, these projects will be some kind of automation for cost savings, or applying machine learning to an existing process to improve its efficiency or performance. This looks a lot like success, and it may suffice, but it’s missing out on the unique advantages of an excellent data strategy.

In an excellent strategy, more data projects have worked out, and they were surprisingly cost-effective to develop. Further, the process of building the first few projects inspires new project ideas. In an excellent strategy, the projects will include automation and efficiency and performance improvements, but they will also include projects and ideas for new revenue generation and entirely new businesses driven by your unique data assets. The data teams work well together, build on each other’s work, and collaborate smoothly with their business partners. There’s a clear vision of what the machine-learning driven future of the business can look like, and everyone is working together to achieve it.  .... "

Wednesday, November 07, 2018

Kinds of Data Scientists

I like the distinction, often useful.  But would add a distinction by problem domain, which often adds a further understanding of data types, sources, resources and restrictions that can be essential.  And a knowledge of ethics can also be domain specific.  Good general and non-technical piece ...

The Kinds of Data Scientist   By Yael Garten in the HBR

In 2012, HBR dubbed data scientist “the sexiest job of the 21st century”. It is also, arguably, the vaguest. To hire the right people for the right roles, it’s important to distinguish between different types of data scientist. There are plenty of different distinctions that one can draw, of course, and any attempt to group data scientists into different buckets is by necessity an oversimplification. Nonetheless, I find it helpful to distinguish between the deliverables they create. One type of data scientist creates output for humans to consume, in the form of product and strategy recommendations. They are decision scientists. The other creates output for machines to consume like models, training data, and algorithms. They are modeling scientists.   .... " 

Monday, August 27, 2018

Enabling Reliable Data science and ML Projects

By my own experience, agreed.  Enterprise Data Science is rarely done very carefully, leading it open to error.

Enabling reliable, secure collaboration on data science and machine learning projects    A conversation with Paul Taylor, chief architect in Watson Data and AI, and IBM fellow.

By Frank Kane   O'Reilly Conference Jupyter

Machine learning researchers often prototype new ideas using Jupyter, Scala, or R Studio notebooks, which is a great way for individuals to experiment and share their results. But in an enterprise setting, individuals cannot work in isolation—many developers, perhaps from different departments, need to collaborate on projects simultaneously, and securely. I recently spoke with IBM’s Paul Taylor to find out how IBM Watson Studio is scaling machine learning to enterprise-level, collaborative projects.

First, a bit of background about Taylor. He has enjoyed a distinguished career at IBM over the past 17 years, where he started off working on Db2 and Informix, and working with big data and unstructured data well before those fields exploded. He has held many titles working in different technology areas as a distinguished engineer, chief architect, master inventor, CTO, and this year was appointed as an IBM Fellow.

Today, Taylor leads the technology of IBM Watson data and AI components, where he is exploring the convergence of data, AI, and public cloud with IBM Watson Studio. Watson Studio provides a suite of tools for data scientists, application developers, and subject matter experts to collaborate and work with data to conduct analytics and data science, and to build, train, and deploy models at scale.

Frank Kane: Why is better collaboration in data science important? What sorts of opportunities do you see it creating for real-world developers and businesses?

Paul Taylor: A lot of times I go in to talk to C-suite folks who are running the data science teams. They're in a real challenge because, traditionally, many of those clients and the scientists are using their own little tools, and they may be very sophisticated tools, or they may be very naïve ones. They're all working in silos, and they're using their own tools in their own way.   ... " 

Friday, August 03, 2018

Data Science Reports

Figure Eight has their Data Science Report out.  A 33 page report.
 Also some other interesting eBooks are made available.  Registration information required. 

" ... Machine learning projects are proliferating, and more and more data - 2.5 quintillion bytes of data each day - is required to power them.

Understanding what practitioners think about the technology they're pioneering is important. To that end, for the fourth straight year, we've taken the pulse of the data scientist community and are excited to share the results with you in our annual Data Scientist Report.

In this year's report, we cover:

Job satisfaction. We found out that data scientists don't just like their job, they love it.
The data and tools that data scientists are working with in 2018.
Ethical issues around building and deploying AI.
Algorithmic bias. Are we still in denial? 

Download the report today.

Best regards,   The Figure Eight Team .... 

Tuesday, July 31, 2018

Futuretext: Data Science for the IOT

Looks to be of interest.  Will be reviewing as it proceeds,  From Futuretext:

" ... After a bit of delay, here is the methodology section of our book Data Science for Internet of Things co-authored by Ajit Jaokar, Jean Jacques Bernard and Sukanya Mandal
I have included the overall structure of the book also. I will send sections in a variable order as we write them. Leading up to my teaching in Oxford Uni in the fall - we should be able to complete most of it.

Book chapter URL:

Also,   My colleague Cheuk and I are launching an Enterprise AI workshop in London in September ... Details below. Initially only 10 places. Saturdays Or Remote
http://www.opengardensblog.futuretext.com/archives/2018/07/enterprise-ai-workshop-saturdaysremote-only-10-places.html  (Overview contains some interesting thoughts about approaches for the enterprise and AI philosophy)

I hope you find the book useful. Its not easy to write it since AI for IoT (Data Science for Internet of Things) is a complex domain . Comments on book section welcome. We may add a section on implementation also (in AWS and Azure) both of which we include in my teaching.

To sign up for the book chapters and ongoing emails https://my.sendinblue.com/users/subscribe/js_id/31hme/id/1
futuretext.ai   ....     Ajit  ....." 

Wednesday, July 18, 2018

Causality and Data Science

Of interest, reviewing.

Causal Data Science  By Adam Kelleher
Physicist; Data @ BuzzFeed; Adjunct Prof. at Columbia

I started a series of posts aimed at helping people learn about causality in data science (and science in general), and wanted to compile them all together here in a living index. This list will grow as I post more:  ... " 

Saturday, July 07, 2018

Saturday Data Science Reading from DSC

Good Saturday Analytics and Data Science Reading from DSC

Always interesting, at various levels of complexity, join the DSC

Posted by Vincent Granville 
Monday newsletter published by Data Science Central. Previous editions can be found here. The contribution flagged with a + is our selection for the picture of the week. ... "

Thursday, July 05, 2018

Demystifying Data Science Free online Conference.

Looks to be of interest with good industry and technical speakers.

A FREE Live Online Conference for Aspiring Data Scientists & Data-Curious Business Leaders
28 Speakers • 2 Days • FREE
July 24 - July 25, 2018     10am - 5pm ET

The #DemystifyDS Experience
Demystifying Data Science is designed to be equal parts informative and interactive. All registrants will have access to the presentation recordings after the conference - but you have to attend live for the full experience!

More information and registration.

Saturday, June 16, 2018

Top 20 Python Data Science Libraries from DSC

Extensive sets with good descriptions. 

Top 20 Python libraries for data science in 2018

Posted by Igor Bobriakov

Python continues to take leading positions in solving data science tasks and challenges. Last year we made a blog post overviewing the Python’s libraries that proved to be the most helpful at that moment. This year, we expanded our list with new libraries and gave a fresh look to the ones we already talked about, focusing on the updates that have been made during the year.

Our selection actually contains more than 20 libraries, as some of them are alternatives to each other and solve the same problem. Therefore we have grouped them as it's difficult to distinguish one particular leader at the moment. .... " 

Sunday, June 10, 2018

Systems Dynamics and Data Science

Interesting connection, though I have never heard of systems dynamics directly invoked in a business process model in a very long time.   Embedded yes in things like process and systems control, but not in high level process.  Should  be considered more often, probably.  But business process usually happens many levels above these possibilities.  Do like that this makes me think of it again.

Developed at MIT’s Sloan School of Management in 1950s system dynamics is a methodological approach to model the behavior of complex systems, where change in one component leads to change in others (like the dominos effect with feedback loops added). This approach is widely applied in industries such as healthcare, disease research, public transportation, business management and revenue forecasting. The most famous application of system dynamics probably is in Limits to Growth.  ... " 

Saturday, June 09, 2018

Free Books on Data Science and Machine Learning

KDNuggets has been very good in pointing these out.    Very high quality.  Good price (free).  We live in amazing times.  Books that cover breaking technology insights.  Have you looked at the cost of required texts for college tech courses?   And covering specific problem solving methods that could save you millions.  For free.  And there is open source software out there you can use too.  Can't beat our times. Prototype your ideas, then sell them to roll them out. ...

...  Summer, summer, summertime. Time to sit back and unwind. Or get your hands on some free machine learning and data science books and get your learn on. Check out this selection to get you started.     By Matthew Mayo, in KDnuggets.

It's time for another collection of free machine learning and data science books to kick off your summer learning season. Because that's a thing. Right?

If, after reading this list, you find yourself wanting more free quality, curated books, check the previous iteration of this series or the related posts below.   ... "