/* ---- Google Analytics Code Below */
Showing posts with label DSC. Show all posts
Showing posts with label DSC. Show all posts

Wednesday, February 02, 2022

On Easy Data Animation

Very nice piece by Vincent Granville,  worth a read. 

Data Animation: Much Easier than you Think!

By Vincent Granville, February 2, 2022 at 1:36 am  in DSC

In this article, I explain how to easily turn data into videos. Data animations add value to your presentations. They also constitute a useful tool in exploratory data analysis. Of course, as in any data visualization, carefully choice of what to put in your video, and how to format it, is critical. The technique described here barely requires any programming skills. You may be able to produce your first video in less than one hour of work, and even with no coding at all. I also believe that data camps and machine learning courses should include this topic in their curriculum.

Examples include convergence of a 2D series related to the Riemann Hypothesis, and fractal supervised classification.

Preparing a video

To produce a data video, there are three steps:

Step 1: Prepare a standard data set (for instance, summary data in Excel, or raw data in text format) with one extra column indicating the frame number. If your video has 20 frames, that column indicates the frame number: an integer between 1 and 20.

Step 2: Create the frames. In our example, it consists of 20 images, typically in PNG format, and named (say) input001.png, input002.png and so on. The number of frames can be as large as a few thousands or a small as 10. The production of the PNG images is typically automated.

Step 3: Turn your images into a video, say an mp4 file. Some people think that this is the most difficult part, but actually it is the easiest one. It can be done with a few clicks, even without writing a single line of code, using free online platforms designed for that purpose. Here I illustrate how to do it in R with just two lines of code.

If your plan is to create a small presentation with 10 frames, a video may not be the best medium. You can still do it with a video by choosing a long duration (say 5 seconds) for each frame. However, a slide presentation may be a better alternative. My videos typically contain between 100 and 1,000 frames, with a frequency anywhere from 4 to 12 frames per second.   .... '  

Thursday, September 02, 2021

Ethical AI

 Via DSC, below the intro, much more at the link:

The Ethical AI Application Pyramid

Posted by Bill Schmarzo  In the blog “How Can Your Organization Manage AI Model Biases?”, I wrote:

“In a world more and more driven by AI models, Data Scientists cannot effectively ascertain on their own the costs associated with the unintended consequences of False Positives and False Negatives. Mitigating unintended consequences requires the collaboration across a diverse set of stakeholders in order to identify the metrics against which the AI Utility Function will seek to optimize.”

I’ve been fortunate enough to have had some interesting conversations since publishing that blog, especially with an organization who is championing data ethics and “Responsible AI” (love that term). As was so well covered in Cathy O’Neil’s book “Weapons of Math Destruction”, the biases built into many of the AI models that are being used to approve loans and mortgages, hire job applicants, and accept university admissions are yielding unintended consequences that severely impact both individuals and society.

AI models only optimize against the metrics against which it has been programmed to optimize. If the AI model yields unintended consequences, that’s not the AI models fault. It’s the fault of the data science team and the operational stakeholders who are responsible for defining the AI Utility Function against which the AI model will judge model progress and success (see Figure 1).  ... " 

Monday, May 10, 2021

AI's for Art Forgery?

Some good points made here, in their current state these are not really even forgery-perfect.   But they could be.  And not to say they could not do other kinds of counterfeiting that don't require so much testing and identity accounting.   

AI: The Next Great Art Forger

Posted by Stephanie Glen in DSC

AI develops new “art” using image analysis.

GANs do not create, they repaint.

The result is a pastiche, a poor copy of the real thing.

AI art is created with algorithms that enable AI to learn a specific aesthetic by analyzing thousands of images; The algorithm then attempts to generate new images based on that learning [1].  Original pieces can also be created by GANs, which pit two neural networks against each other. The result is “art” that is difficult to differentiate from human-made artwork. One such piece, Portrait of Edmond Belamy, sold for a staggering $432, 500 when it went under the hammer at Christie’s Prints & Multiples sale at Christie’s on 23-25 October last year [2].

But do these AI-generated pastiches qualify as real art? Probably not. Many in the art and AI communities agree that these cannot be called art, at least in the traditional sense. Even if you could stretch the definition of art to include AI-generated images, they are of poor quality and no better than a factory produced knock off.    ..." 

Monday, November 30, 2020

Four Stages of Robotic Process Automation (RPA)

 A good intro look at RPA from DSC.  Below the intro, more at the link.   A good place to start with AI oriented automation.  Also not emphasized enough.  I would include the 'process' part, and ask: how does the process modeled integrate with key business process?     I would also include:  "What are the risks involved of the automation under current and future context"?

The Four Stages of Robotic Process Automation (RPA)   Posted by Amit Dua in DSC

“Robotic process automation is not a physical [or] mechanical robot,” says Chris Huff, chief strategy officer at Kofax. In fact, there is no presence or involvement of any robots in the automation software, as its name says.

RPA or Robotic Process Automation is an amalgamation of three factors:

Robotic – entities impersonating human activities and processes known as ‘Robots.’

Process – multiple small activities that lead to one outcome or result.

Automation – tasks done by machines and robots instead of humans.

Robotic Process Automation is a set of software robots that run on a virtual or physical machine. It automates most of the mundane and boring tasks of the business. RPA bots are capable of impersonating almost all human-computer interactions, which further tackles those business tasks, error-free. Besides, this technology works at a much faster pace and volume..... "

Wednesday, November 04, 2020

Examples of Machine Learning for Law and Compliance

Good general overview of a space we have worked for a while,  yes good opportunities. But beyond ML methods.

AI/ML Applications in Law and Compliance    Posted by William Vorhies in DSC

Summary:  Some industries are a clear slam-dunk for AI/ML applications and some less so.  The legal, regulatory, and compliance businesses (law firms, internal legal departments, and the contract review and regulatory compliance departments of heavily regulated industries) fall in this last category.  This is a review of seven companies found by TopBots to be successful; pointing to opportunities others can follow.

Remember just a few years ago when we were looking forward to now or a little beyond and imagining what applications AI/ML would have in different industries.  Some of those prognostications were slam dunks as they applied to customer propensity or using machine vision to count whatever widgets you were interested in.

What struck me as really speculative at the time were applications that impacted fields dense in laws, regulations, and complex situations where human knowledge and intuition had long dominated. ... ' 

Wednesday, October 07, 2020

Free Book on Computational Agents

The use of an agent to represent objects, IOTs or people is very useful.  Closer, but not the same as classic optimization models.  The book and comments here are good introductions.   Technical. 

Free book - Artificial Intelligence: Foundations of Computational Agents   Posted by ajit jaokar   in DSC       

Comments:  There are many excellent free books on Python – but Artificial Intelligence: Foundations of Computational Agents is about  a subject not commonly covered

I found the book useful as a introduction to Reinforcement Learning

As the title suggests, the book is about computational agents

An agent observes the world and carries out actions in the environment. The agent maintains an internal state that it updates. Also, the environment takes in actions of the agents, and in turn updates its internal state and returns the percepts. In this implementation, the state of the agent and the state of the environment are represented using standard Python variables, which are updated as the state changes.  .... " 

This structure can be used to model many interesting problems and is the focus of the book. Ultimately, it leads to Reinforcement Learning. ... " 

Wednesday, September 30, 2020

Analyics and Data Scence Convergiing

I would say that good analytics also requires just the same care about data use and results curation.  It has always been that way.  Automating aspects of analytics requires more training for decision makers.   Otherwise good piece about the trend.

For Better or Worse Analytics and Data Science are Converging  Posted by William Vorhies  in DSC 

Summary:  Analytic Platforms are rapidly being augmented with features previously reserved for data scientists.  They are presented as easy to use but require substantial data literacy and advanced DS skills for the most complex.  Business users and analysts can pursue more complex problems on their own, but need good oversight.

Data Science Platform developers and Analytics Platform developers have been circling each other for years.  The DSP folks see the analyst market presenting a much larger customer base.  The ASP folks need to keep adding higher value features to maintain their shares.  Gartner says these two ecosystems are swirling closer and closer like binary stars headed for collision.

In Gartner’s report “Augmented Analytics Is the Future of Analytics” published late last year they show the various elements that used to be unique to Data Science Platforms that are finding their way into Analytic Platforms.

The emphasis here is definitely on ‘Augmented’ as in fairly sophisticated features easy enough for a non-data scientist to use.  Not all vendors of these augmented platforms provide the same features though Gartner proposes that they will rapidly converge.

Whether you’re a CXX exec in charge of this integration or a data scientists asked to participate there are some risks as well as advantages here.  ... "

Monday, September 14, 2020

Data Science Fails If it Looks too Good to be True

Not sure if I completely agree.  Have seen very good results come out of an analytic solution.  I agree that if it makes recommendations very different from current practice, or suggests buying into high risk, depends on unknown future states or or high investments, it deserves very close examination.    But if it simply has different methods, results or valuation.  Why not?  Hype bothers me too, but much value started there.

DSC Podcast

Data Science Fails – If It Looks Too Good To Be True...

You’ve probably seen amazing AI news headlines such as: AI can predict earthquakes. Using just a single heartbeat, an AI achieved 100% accuracy predicting congestive heart failure. AI can diagnose covid19 in seconds from a chest scan. A new marketing model is promising to increase the response rate tenfold. It all seems too good to be true. But as the modern proverb says, “If it seems too good to be true, it probably is”.

In this latest Data Science Central podcast, https://dsc.news/3fhbOt9  we look behind the hype to show whether there is substance to these claims, and then show you how to avoid these types of data science fails.
Speaker: Colin Priest, VP of AI Strategy - DataRobot
Hosted by: Sean Welch, Host and Producer - Data Science Central
https://dsc.news/3fhbOt9

via DataRobot

Friday, September 04, 2020

Google Machine Learning Tutorial

Had the honor to meet Kirk D Bourne some time ago.   Here he passes along a link to the Google Tutorial on Machine Learning certainly worth a look: 

Kirk Borne   @KirkDBorne  

★★★★This Google Tutorial (100+ slides total) on #MachineLearning is the best: https://dy.si/bv6mP 

@jason_mayes  (author, very nicely done ....)

#abdsc #BigData #DataScience #DataMining #NeuralNetworks #AI #DeepLearning #TensorFlow #Mathematics #Algorithms #DataScientists #ReinforcementLearning

Thursday, September 03, 2020

Comparing Kinds of AI based Learning

Nicely done piece that compares different forms of artificial intelligence based learning methods.   And what they are typically meant to do.   Not overly technical.

Clear The Confusion-Artificial Intelligence vs Machine Learning vs Deep Learning
Posted by Varun Bhagat in DSC

Raise your hand, if you are stuck in confusion between Artificial Intelligence, Machine Learning and Deep Learning. Are you one of these? If yes, then you have clicked the right page. Here, you will get your answers.

However, we use these terms interchangeably; the truth is they are different but a part of the same branch, i.e. computer science. Take a look at the image and understand the difference.

As you can see in the image,  there are three concentric circles; Deep Learning is a subset of Machine Learning, which is also a subset of Artificial Intelligence.

See how the technologies of the same departments are competing with each other and serving us differently.    ... "

Improving Performance Engineering with Machine Learning

Had never heard this specifically positioned this way, nice idea.   Needs some more detail to explain, how it has been done, but a great start.   We did a form of this with business process modelling, and the integration with machine learning to determine how elements of the process performed.   Not sure if this is quite the same thing.  Like the anomaly reference.

Machine Learning: How it Improves Performance Engineering in DSC    Posted by Ryan Williamson  

Enterprise software, as well as other kinds, remains a complicated endeavor, thus necessitating the use of modern means to gauge, analyze, and adapt their performance. And one of the most popular technologies in the performance engineering market right now is machine learning. Since it has demonstrated an unparalleled ability to not only help foresee performance issues and fix them. When used in the right manner — this combination can also help performance engineering teams to steer clear of any issues at all completely. It is because machine learning comes equipped with the ability to interpret and analyze data in real-time, thus delivering valuable insights about the system’s performance.

However, if you are going to truly leverage machine learning’s abilities in the context of performance engineering, it is first essential to understand the basics. Through this article, let’s discuss the kind of performance anomalies one can encounter.

Point anomalies: This is when there is only one data point that is distinct from the entire set.
Contextual anomalies: In this scenario, the anomaly is contextual, i.e., exists only in a particular context.
Collective anomalies: This refers to a data set that exhibits signs of an anomaly.  ....  " 

Sunday, August 02, 2020

Google Tutorial and Presentation on AI and Machine Learning

Link via DSC to a considerable presentation on Machine Learning from Google. Below a quick intro with a clue to some common uses of what we are now calling AI.

Google Tutorial on Machine Learning
Posted by Capri Granville  in DSC

This presentation was posted by Jason Mayes, senior creative engineer at Google, and was shared by many data scientists on social networks. Chances are that you might have seen it already. Below are a few of the slides. The presentation provides a list of machine learning algorithms and applications, in very simple words. It also explain the differences between AI, ML and DL (deep learning.)  ...



              ......

Much, much more at the link above.   Really worth it if only to scan the slides ...

You can check out the whole presentation (96 slides) here.     ....   I see this is perhaps dated being from 2017,  but appears to be very useful and mostly non-technical. 

Tuesday, June 16, 2020

Podcast: Do we Need Data Scientists in a World of Automation?

New Podcast below, by SAS and DSC.    Good topic.  The question should be:  How many and how should the be involved in an enterprise?  Just like Computer Scientists, Economists, Statisticians ... and other technical fields.  Machine Learning can't do it all.

Do We Need Data Scientists in Today’s World of Automation?

Machine learning is said to be an important driver of the future of intelligent systems, automatically analyzing data and distilling new knowledge, actionable insights and compelling decisions. But why show market trends that investments in data scientists – those golden people that train machine learning models – have never been higher if machine learning can be fully automated? Why do we need data scientists if machine learning is designed to do it all?

In today’s Data Science Central podcast, Véronique Van Vlasselaer, Data & Decision Scientist at SAS, will discuss what machine learning automation entails, and how valuable human input in the machine learning process is.

Speaker: Véronique Van Vlasselaer, Data & Decision Scientist - SAS
Hosted by: Rafael Knuth, Contributing Editor - Data Science Central

Monday, May 25, 2020

Free Fundamentals of Machine Learning Book

Free 185 page PDF:  Comprehensive Guide to  Machine learning   Technical, but parts are useful generally.

Free Book: A Comprehensive Guide to Machine Learning     (Berkeley University)  via Capri Granville in DSC

By Soroush Nasiriany, Garrett Thomas, William Wang, Alex Yang. Department of Electrical Engineering and Computer Sciences, University of California, Berkeley. Dated June 24, 2019. This is not the same book as The Math of Machine Learning, also published by the same department at Berkeley, in 2018, and also authored by Garret Thomas....  

Monday, May 18, 2020

Embracing Responsible AI from Pilot to Production

A topic rarely addressed well, we discovered early on it had to be carefully done.  Upcoming DSC webinar:

Embracing Responsible AI from Pilot to Production
Join us for the latest DSC Webinar on May 27th, 2020
register-now: https://dsc.news/2LFMEHz

On average, 80% of AI projects fail to make it to production. But it IS possible to successfully launch AI, at scale, that is built responsibly and works for everyone. How you scale from pilot to production is critical to ensuring AI success, while continuing to be a good corporate citizen through responsible productization. In this latest Data Science Central webinar, we'll talk about the framework for scaling AI pilots to production with a focus on ethical responsibilities and bias mitigation at each step.

We'll look at:

The five-step AI development cycle
Ways to control for unwanted bias across data, models, and run time at the production layer
Explainability and why it is key for moving AI pilots to production that delivers core business value

Speakers:
Lukas Biewald, Founder & CEO -- Weights & Biases
Alyssa Simpson-Rochwerger, VP of AI & Data -- Appen

Hosted by: Sean Welch, Host and Producer -- Data Science Central 
Title: Embracing Responsible AI from Pilot to Production
Date: Wednesday, May 27th, 2020
Time: 9 AM - 10 AM PDT 
Space is limited so please register early:
https://dsc.news/2LFMEHz

Reserve your Webinar seat now 

After registering you will receive a confirmation email containing information about joining the Webinar.  ... 

Thursday, May 14, 2020

Can Personality Determine How we Interact with Future AI?

Very thoughtful piece in DSC.  I am suspicious of using the Myers-Briggs model here.  Read the complete article there, and join DSC for much more.

Could your personality decide how you can collaborate with AI in future jobs?
Posted by ajit jaokar in DSC

I am currently exploring this idea in a paper
Could your personality decide how you can collaborate with AI in future jobs?
I am a strong believer in a mode where humans and AI work together and collaborate 

I also believe that if we can master this skill, we can bring manufacturing jobs back from offshore destinations

In the book "Humans + Machines – Reinventing work in the age of AI": by Daugherty and Wilson (2018) propose a model for humans and AI working together based on the idea of identifying the ‘missing middle'. The missing middle represents areas where humans and AI collaborate and comprises of two sections:   Where humans complement machines and  where AI gives humans superpowers.   (Complete article at link) 


Sunday, March 08, 2020

Monitoring a Complex System Over time for Patterns

Brought to my close attention because this is in the astronomy space.  A long time academic interest.  But then it came to mind that any system be observed in this way over time.  Then examine the data for anomalies, patterns, trends.   Probably with different observable parameters, and goals and needs. Thinking that further.Will be attending for inspiration.

Via DSC:

The Vera C. Rubin Observatory, currently under construction in Chile, will conduct a vast astronomical survey of our dynamic Universe starting in 2022. They plan to collect 500 petabytes of image data by observing the skies continuously for 10 years and produce nearly instant alerts for objects that change in position or brightness every night. In addition to astronomical data, their dataset will include DevOps, IoT, and real-time monitoring data.

In this latest Data Science Central webinar, Dr. Angelo Fausti will demonstrate:

How a time-series database has the versatility to address their needs
How they created a solution to enhance visibility across their organization and improve actionable insights
How they pull software development and sensor data from their telescope, camera and observatory IoT devices

Speaker:
Dr. Angelo Fausti, Software Engineer -- Vera C. Rubin Observatory

Hosted by: Rafael Knuth, Contributing Editor -- Data Science Central 

Title: 500 Petabytes of Data to Understand the Universe Better
Date: Wednesday, March 18th, 2020
Time: 9:00 AM - 10:00 AM PDT 

Space is limited so please register early:
Reserve your Webinar seat now 
After registering you will receive a confirmation email containing information about joining the Webinar.    .... '

Saturday, March 07, 2020

Case for Open Data in Fight Against COVID-19

The case for open data for AI in the fight against COVID-19
Posted by ajit jaokar  in DSC ....

COVID-19
   2019 Novel Coronavirus COVID-19 (2019-nCoV) Data Repository by Johns Hopkins CSSE
This is the data repository for the 2019 Novel Coronavirus Visual Dashboard operated by the Johns Hopkins University Center for Systems Science and Engineering (JHU CSSE). Also, Supported by ESRI Living Atlas Team and the Johns Hopkins University Applied Physics Lab (JHU APL). .... 

Note also the comments ...

Tuesday, February 25, 2020

Using Business Rules and Expertise

Via DSC, what looks to be a good podcast on this topic.  It has been a favorite approach of mine since the beginning.  Narrow machine learning methods can be very valuable, but to deliver them they have to be part of existing or proposed tasks or businesses.  Operationally embedded.   That requires real-life decision rules.  Access information to the podcast at the link below.  More on this topic to follow. 

Data Science Fails: Ignoring Business Rules & Expertise 

Nowadays, we have unprecedented access to data, plus the computing power and advanced algorithms to find correlations. We look at a cautionary case study of a cancer center that embarked on an ambitious plan to use AI to eradicate cancer. When AI is being asked to make decisions with significant consequences, such as life and death healthcare recommendations, it needs to be trustworthy. But if you don't follow best practices, if you don't include the knowledge of subject matter experts, and if you don't enforce business rules, your AI project will not be successful.

In this latest Data Science Central podcast, learn four AI governance practices that can help you achieve AI success.

Speaker: Colin Priest, VP of AI Strategy - DataRobot
Hosted by: Sean Welch, Host and Producer - Data Science Central .... 

Tuesday, December 17, 2019

Articles about Bayesian Methods and Networks

Just had a chance to reference this resource in DSC, good intro with a mix of technical and general introduction.

15 Great Articles about Bayesian Methods and Networks
Posted by Vincent Granville on March 15, 2019 

This resource is part of a series on specific topics related to data science: regression, clustering, neural networks, deep learning, decision trees, ensembles, correlation, Python, R, Tensorflow, SVM, data reduction, feature selection, experimental design, cross-validation, model fitting, and many more. To keep receiving these articles, sign up on DSC.  .... "