/* ---- Google Analytics Code Below */
Showing posts with label KDNuggets. Show all posts
Showing posts with label KDNuggets. Show all posts

Thursday, April 13, 2023

Open Source Alternatives to ChatGPT and Bard

Nicely done piece which illustrates a number of existing tools , Open-Source examples, Several I have not heard of.  Instructive.

8 Open-Source Alternatives to ChatGPT and Bard   in KDNuggets

Discover the widely-used open-source frameworks and models for creating your ChatGPT like chatbots, integrating LLMs, or launching your AI product.

By Abid Ali Awan, KDnuggets on April 6, 2023 in Natural Language Processing   ... '

Wednesday, October 06, 2021

Data Science Books

In KDNuggets:

Data Science books to start reading:   Good selection,   I would pick one or more that makes the most sense based on your current background and go from there. 

By Przemek Chojecki, CEO Contentyze

Data science is undoubtedly one of the hottest career choices right now. Companies (many of whom have data science departments) are hiring data scientists around the board. It is a considerable thing to become a data scientist. It is also a fantastic opportunity to hone your expertise if you are already a statistician and want to step through the ranks.

This article discusses the most popular data science books for any level.  .... '

Sunday, August 15, 2021

Exploratory Data Analyses

 Am a big proponent of exploratory data analysis, and making it initially as simple, visual and goal oriented as possible.   Here a piece from KDNuggets that shows one approach.  Link through.    Have read KDN for years, long before AI became big and we did what was called optimization and analytics. 

EDA is a fundamental early process for any Data Science investigation. Typical approaches for visualization and exploration are powerful, but can be cumbersome for getting to the heart of your data. Now, you can get to know your data much faster with only a few lines of code... and it might even be fun!   KNIME Fall Data Talks Bringing Business and Data Science Together

By Francois Bertrand, coder & designer for data visualization and games.

Sunday, May 02, 2021

Simple Examples of Anomaly/Outlier Detection

 A common need,  nicely and simply put here in KDNuggets.  Very classic and should be done with most every set of data you are seriously working with.  With more tech and coding at the link:

Four Techniques for Outlier Detection

Tags: DBSCAN, Knime, Outliers, Python

There are many techniques to detect and optionally remove outliers from a dataset. In this blog post, we show an implementation in KNIME Analytics Platform of four of the most frequently used - traditional and novel - techniques for outlier detection.   By Maarit Widmann, Moritz Heine, Rosaria Silipo, Data Scientists at KNIME

Anomalies, or outliers, can be a serious issue when training machine learning algorithms or applying statistical techniques. They are often the result of errors in measurements or exceptional system conditions and therefore do not describe the common functioning of the underlying system. Indeed, the best practice is to implement an outlier removal phase before proceeding with further analysis.

But hold on there! In some cases, outliers can give us information about localized anomalies in the whole system; so the detection of outliers is a valuable process because of the additional information they can provide about your dataset.

There are many techniques to detect and optionally remove outliers from a dataset. In this blog post, we show an implementation in KNIME Analytics Platform of four of the most frequently used - traditional and novel - techniques for outlier detection.  .... " 

Sunday, February 07, 2021

Top Vision Papers via KDNuggets

 Great often inspirational, leading edge AI stuff.  Technical details.  From KDNuggets.  Most have good embedded video explanations.

The top 10 computer vision papers in 2020 with video demos, articles, code, and paper reference. By Louis (What's AI) Bouchard, Montrealer, explaining AI stuff on YouTube and Medium

Given with everything that happened in the world this year, we still had the chance to see a lot of amazing research come out. Especially in the field of artificial intelligence and more precisely computer vision. More, many important aspects were highlighted this year, like the ethical aspects, important biases, and much more. Artificial intelligence and our understanding of the human brain and its link to AI is constantly evolving, showing promising applications in the soon future, which I will definitely cover.

Here are my top 10 of the most interesting research papers of the year in computer vision, in case you missed any of them. In short, it is basically a curated list of the latest breakthroughs in AI and CV with a clear video explanation, link to a more in-depth article, and code (if applicable). Enjoy the read, and let me know if I missed any important papers in the comments, or by contacting me directly on LinkedIn!

The complete reference to each paper is listed at the end of this article.  ... "

Friday, January 15, 2021

Twelve Innovative Companies

 From KDNuggets:

We bring you industry predictions from 12 innovative companies - what key trends they expect in 2021 in AI, Analytics, Data Science, and Machine Learning?   By Gregory Piatetsky, KDnuggets.

Earlier we published: AI, Analytics, Machine Learning, Data Science, Deep Learning Research Main Developments in 2020 and Key Trends for 2021  

 Main 2020 Developments and Key 2021 Trends in AI, Data Science, Machine Learning Technology.    

Here is last part in our 2021 Predictions series - the predictions from the industry. We received many submissions, and to keep this article size manageable, we limited this to 12 companies: Alluxio, Alteryx, Diamanti, Dremio, Indicative, Lexalytics, Luminoso, MathWorks, MobiDev, Qlik, SAS, and Splice Machine. ... " 

Saturday, January 09, 2021

Machine Learning Algorithms You Should Know

A nice piece which covers the basic concepts of the most used data science algorithms.  Not enough to know how to use them, but good for knowing when and why.  I often reviewed this level of understanding with executives who were involved in the results.  I would also not say 'all',  it depends on your domain.


Tags: Algorithms, Decision Tree, Explained, Gradient Boosting, K-nearest neighbors, Machine Learning, Naive Bayes, Regression, SVM

Many machine learning algorithms exits that range from simple to complex in their approach, and together provide a powerful library of tools for analyzing and predicting patterns from data. If you are learning for the first time or reviewing techniques, then these intuitive explanations of the most popular machine learning models will help you kick off the new year with confidence. 
 
By Terence Shin, Data Scientist 

Tags: Algorithms, Decision Tree, Explained, Gradient Boosting, K-nearest neighbors, Machine Learning, Naive Bayes, Regression, SVM    ... " 

Saturday, December 05, 2020

Last Year in AI, Analytics, Machine Learning and Data Science ....

Good end of the year piece from KDNuggets that was instructive.

AI, Analytics, Machine Learning, Data Science, Deep Learning Research Main Developments in 2020 and Key Trends for 2021

Tags: 2021 Predictions, AI, Ajit Jaokar, Analytics, Brandon Rohrer, Daniel Tunkelang, Data Science, Deep Learning, Machine Learning, Pedro Domingos, Predictions, Research, Rosaria Silipo

2020 is finally coming to a close. While likely not to register as anyone's favorite year, 2020 did have some noteworthy advancements in our field, and 2021 promises some important key trends to look forward to. As has become a year-end tradition, our collection of experts have once again contributed their thoughts. Read on to find out more.

By Matthew Mayo, KDnuggets.

To the chagrin of absolutely no one, 2020 is finally drawing to a close. It has been a rollercoaster of a year, one defined almost exclusively by the COVID-19 pandemic. But other things have happened, including in the fields of AI, data science, and machine learning as well. To that end, it's time for KDnuggets annual year end expert analysis and predictions. This year we posed the question:

What were the main developments in AI, Data Science, Machine Learning Research in 2020 and what key trends do you see for 2021?

Last year's noted main developments and predictions included continued advancements in many research areas, NLP in particular. While there can be debate as to whether 2020's big NLP advancement was as formidable as some may have originally thought (or continue to think), there is no doubt that there was a continued and intense focus on NLP research in 2020. It should not be difficult to surmise that this continues into 2021 as well.  ... "

Friday, October 02, 2020

AI Papers to Read in 2020:

Reading suggestions to keep you up-to-date with the latest and classic breakthroughs in AI and Data Science.   Beyond just healthcare, where the writers main applications are.

AI Papers to read in 2020 via KDNuggets  By Ygor Rebouças Serpa, developing explainable AI tools for the healthcare industry

Artificial Intelligence is one of the most rapidly growing fields in science and is one of the most sought skills of the past few years, commonly labeled as Data Science. The area has far-reaching applications, being usually divided by input type: text, audio, image, video, or graph; or by problem formulation: supervised, unsupervised, and reinforcement learning. Keeping up with everything is a massive endeavor and usually ends up being a frustrating attempt. In this spirit, I present some reading suggestions to keep you updated on the latest and classic breakthroughs in AI and Data Science.

Although most papers I listed deal with image and text, many of their concepts are fairly input agnostic and provide insight far beyond vision and language tasks. Alongside each suggestion, I listed some of the reasons I believe you should read (or re-read) the paper and added some further readings, in case you want to dive a bit deeper into a given subject.

Before we begin, I would like to apologize to the Audio and Reinforcement Learning communities for not adding these subjects to the list, as I have only limited experience with both.

Here we go.  ...  '   (Detail at the link) ....

Saturday, February 29, 2020

Replacing Data Scientists With AutoML?

Excerpt from a current KDNuggets article, which links further to a poll that asks the questions of practitioners.  My answer is yes. AutoML will replace the current needs for data science analysis.  Within a decade.  Of course the needs are likely to expand as well, so there will always be research and new requirements emerging.  And interpretation for specific context needs.  Just as there are needs for statisticians and analytics specialists for the same purposes.

When Will AutoML (Automated Machine Learning) Replace Data Scientists (if ever)?

Soon after tech giants Google and Microsoft introduced their AutoML services to the world, the popularity and interest in these services skyrocketed. We first review AutoML, compare the platforms available, and then test them out against real data scientists to answer the question: will AutoML replace us?

Introduction of AutoML:
One cannot introduce AutoML without mentioning the machine learning project’s life cycle, which includes data cleaning, feature selection/engineering, model selection, parameter optimization, and finally, model validation. As advanced as technology has become, the traditional data science project still incorporates a lot of manual processes and remains time-consuming and repetitive. ... "

Wednesday, February 26, 2020

Gartner Magic Quadrant for Data Science and Machine Learning

KD Nuggets publishes and analyzes the most recent Gartner quadrant analysis. While I am skeptical of this approach, it does have a useful list of participants which can fill in the gaps.   Clip at link below to get to the 'Magic Quadrant'.   Some of the included analysis by KDN is more interesting, with  short, general, non-technical descriptions of what many companies are doing.

The Gartner 2020 Magic Quadrant for Data Science and Machine Learning Platforms has the largest number of leaders ever. We examine the leaders and changes and trends vs previous years.
By Gregory Piatetsky, KDnuggets.

Gartner has released last week its highly-anticipated report and magic quadrant (MQ) for Data Science and Machine Learning Platforms (DSML) and you can get copies from several vendors - see a list at the bottom of this blog. In previous years, the MQ name kept changing but the 4 leaders remained the same. Now the name has remained the same as in 2019 MQ and 2018 MQ reports, reflecting a more mature understanding of the DSML field, but the contents, especially the leader quadrant, have changed dramatically, reflecting accelerating progress and competition in the field.

The 2020 MQ report went back to evaluating 16 vendors (down from 17 last year), placed as usual in 4 quadrants, based on completeness of vision (vision for short) and ability to execute (ability for short).

We note that the report included only vendors with commercial products, and did not consider open-source platforms like Python and R, even though those are very popular with Data Scientists and Machine Learning professionals.   ... )

Saturday, November 16, 2019

Via KDNuggets: 10 Free Books on AI

Via the always interesting KDNuggets.    Which we often pored long ago for information about developments in 'Knowledge Discovery',  before the term changed to something sexier.

10 Free Must-read Books on A  (good descriptions of these at the link) 

Artificial Intelligence continues to fill the media headlines while scientists and engineers rapidly expand its capabilities and applications. With such explosive growth in the field, there is a great deal to learn. Dive into these 10 free books that are must-reads to support your AI study and work.... " 

By Matthew Dearing, KDnuggets.

Sunday, October 27, 2019

Data Management for Data Science

In Kdnuggets a good description and visualization of data management needed for data science.

Everything a Data Scientist Should Know About Data Management

For full-stack data science mastery, you must understand data management along with all the bells and whistles of machine learning. This high-level overview is a road map for the history and current state of the expansive options for data storage and infrastructure solutions. By Phoebe Wong and Robert Bennett.

To be a real “full-stack” data scientist, or what many bloggers and employers call a “unicorn,” you have to master every step of the data science process — all the way from storing your data, to putting your finished product (typically a predictive model) in production. But the bulk of data science training focuses on machine/deep learning techniques; data management knowledge is often treated as an afterthought. Data science students usually learn modeling skills with processed and cleaned data in text files stored on their laptop, ignoring how the data sausage is made.   ... " 

Sunday, October 20, 2019

Data Visualization in Quadrants with a Story

Clever piece in KDNuggets, about how to take survey data and take some basic, well known frameworks, and add a story to make your point.  Non-technical.   Not sure how wow-viral the point is, but this gives an example of how to proceed to make simple data description useful for anyone.  Nice thoughtful piece.

The 4 Quadrants of Data Science Skills and 7 Principles for Creating a Viral Data Visualization

As a data scientist, your most important skill is creating meaningful visualizations to disseminate knowledge and impact your organization or client. These seven principals will guide you toward developing charts with clarity, as exemplified with data from a recent KDnuggets poll. 

By Jose Berengueres, Professor & Angel Investor.

I teach CS and Design Thinking. But today, I am on a mission to show how to do great charts because I dislike confusing charts. You may think of me as the Marie Kondo of charts or the Cole Knaflic of Dubai, and you might not be entirely wrong. In this post, I will share 7 principles* to go from Aha-charts to Wow-charts. Let’s get our hands dirty with a dataset from a recent KDnuggets poll.

The poll had just two questions:

Which skills/knowledge areas do you currently have? and
Which skills do you want to add or improve?
KDnuggets received 1,500 answers, and we will use the aggregates by skill [1].  ...  "   

Friday, September 06, 2019

Build Your Own Voice Assistant

A look at building your own Voice Assistant from KDNuggets, instructive about how relatively little it takes to set up the basics using Python.   Of course setting it in a complete ecosystem requires quite a few additional details.   Below just the intro, more at the link:

Hone your practical speech recognition application skills with this overview of building a voice assistant using Python.    By Nagesh Chauhan, Big data developer at CirrusLabs

Introduction
Who doesn't want to have the luxury to own an assistant who always listens for your call, anticipates your every need, and takes action when necessary? That luxury is now available thanks to artificial intelligence-based voice assistants.

Voice assistants come in somewhat small packages and can perform a variety of actions after hearing your command. They can turn on lights, answer questions, play music, place online orders and do all kinds of AI-based stuff.   

Voice assistants are not to be confused with virtual assistants, which are people who work remotely and can, therefore, handle all kinds of tasks. Rather, voice assistants are technology based. As voice assistants become more robust, their utility in both the personal and business realms will grow as well.

What is a Voice Assistant?

A voice assistant or intelligent personal assistant is a software agent that can perform tasks or services for an individual based on verbal commands i.e. by interpreting human speech and respond via synthesized voices. Users can ask their assistants’ questions, control home automation devices, and media playback via voice, and manage other basic tasks such as email, to-do lists, open or close any application etc with verbal commands.

Let me give you the example of Braina (Brain Artificial) which is an intelligent personal assistant, human language interface, automation and voice recognition software for Windows PC. Braina is a multi-functional AI software that allows you to interact with your computer using voice commands in most of the languages of the world. Braina also allows you to accurately convert speech to text in over 100 different languages of the world.  .... " 

Monday, June 17, 2019

Data Science Behind Top Machine Learning Tools

KDNuggets examines top data science machine learning tools.  With considerable data visualizations at the link.

Tags: Anaconda, Apache Spark, Big Data Software, Deep Learning, Excel, Keras, Poll, Python, R, RapidMiner, scikit-learn, Software, SQL, Tableau, TensorFlow

We identify the 6 tools in the modern open-source Data Science ecosystem, examine the Python vs R question, and determine which tools are used the most with Deep Learning and 
By Gregory Piatetsky, KDnuggets.

Recently we reported the results of 20th annual KDnuggets Software Poll:
Python leads the 11 top Data Science, Machine Learning platforms: Trends and Analysis.
As we have done before (see 2017 data science ecosystem, 2018 data science ecosystem), we examine which tools were part of the same answer - the skillset of the user. We note that this does not necessarily mean that all tools were used together on each project, but having knowledge and skills to used both tools X and Y makes it more likely that both X and Y were used together on some projects. The results we see are consistent with this assumption.

The top tools show surprising stability - we see essentially the same pattern as last year.

First, we selected the tools with at least 20% of the vote. There were 11 such tools - exactly the same list of 11 tools as last year, although the order has changed a little. Keras moved up from n. 10 to n. 8, and Anaconda moved up from n. 6 to n. 5. Tableau and SQL moved down a little.

The cutoff for this group of 11 is a natural one, since there is a big gap between n. 11 (Apache Spark, with 21%) and n. 12 (Microsoft Power BI, 13%).

We used the same Lift measure as in our 2017 analysis and 2018 analysis.
We then grouped together the tools with the strongest association, starting with Tensorflow and Keras, until we arrived to the figure 1 below. We made the patterns easier to see by showing only associations with abs(Lift1) > 15%.     ... "

Wednesday, January 16, 2019

KDNuggets: Books on Language Processing

I have read KDNuggets long before deep learning was discovered.    Love it.  The most recent post was useful for a project:

KDnuggets™ News 19:n03, Jan 16: Top 10 Books on NLP and Text Analysis; End To End Guide For Machine Learning Projects

Also: Why Vegetarians Miss Fewer Flights - Five Bizarre Insights from Data; 4 Myths of Big Data and 4 Ways to Improve with Deep Data; The Role of the Data Engineer is Changing; How to solve 90% of NLP problems: a step-by-step guide ... 

This week, check out a collection of top 10 NLP and text analysis books, see a step to step guide on the process that you can follow to implement a successful data science project, find out why vegetarians miss fewer flights (???), read about the fundamental misconception that bigger data produces better machine learning results, and find out how the role of the data engineer is changing. ... " 

Sunday, June 10, 2018

Tutorial on Association Rules

Good piece in KDNuggets,  the approach is useful because it is transparent and thus easily visualized. Here is a further technical definition,  from the approach of set mining and rule generation.  Note this applies in many domains.  We used this to generate initial sample expert rule sets.  The method can be extending to the idea of a knowledge graph.  Below the tutorial introduction:

Association Rules and the Apriori Algorithm: A Tutorial
A great and clearly-presented tutorial on the concepts of association rules and the Apriori algorithm, and their roles in market basket analysis.

The (an Example) Problem

When we go grocery shopping, we often have a standard list of things to buy. Each shopper has a distinctive list, depending on one’s needs and preferences. A housewife might buy healthy ingredients for a family dinner, while a bachelor might buy beer and chips. Understanding these buying patterns can help to increase sales in several ways. If there is a pair of items, X and Y, that are frequently bought together:

Both X and Y can be placed on the same shelf, so that buyers of one item would be prompted to buy the other.

Promotional discounts could be applied to just one out of the two items.
Advertisements on X could be targeted at buyers who purchase Y.
X and Y could be combined into a new product, such as having Y in flavors of X.
While we may know that certain items are frequently bought together, the question is, how do we uncover these associations?

Besides increasing sales profits, association rules can also be used in other fields. In medical diagnosis for instance, understanding which symptoms tend to co-morbid can help to improve patient care and medicine prescription.

Definition

Association rules analysis is a technique to uncover how items are associated to each other. There are three common ways to measure association. .... "

Saturday, June 09, 2018

Free Books on Data Science and Machine Learning

KDNuggets has been very good in pointing these out.    Very high quality.  Good price (free).  We live in amazing times.  Books that cover breaking technology insights.  Have you looked at the cost of required texts for college tech courses?   And covering specific problem solving methods that could save you millions.  For free.  And there is open source software out there you can use too.  Can't beat our times. Prototype your ideas, then sell them to roll them out. ...

...  Summer, summer, summertime. Time to sit back and unwind. Or get your hands on some free machine learning and data science books and get your learn on. Check out this selection to get you started.     By Matthew Mayo, in KDnuggets.

It's time for another collection of free machine learning and data science books to kick off your summer learning season. Because that's a thing. Right?

If, after reading this list, you find yourself wanting more free quality, curated books, check the previous iteration of this series or the related posts below.   ... "

Keras Workflow for Experiments

I just reviewed and discarded several packages we used to experiment with neural nets in the late 90s.  Fairly easy to use, but little iterative architecture capability and data wrangling capabilities.    Had been asked about Keras,  an open source neural network library written in Python.  Have never used it commercially.  Also found a piece in KDNuggets on a  Keras 4 step workflow that's useful.  If you are in the early stages of experimentation, and you are already in Python, this makes sense.    If you are beyond that,  something like Tensorflow looks better to test what kind of muscle might be required.