/* ---- Google Analytics Code Below */
Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

Sunday, March 19, 2023

Technology in Insurance Sector (India)

 Would think that AI would also change things strongly here.  How specific this is to India is unclear.

ACM NEWS

Technology Innovation in the Insurance Sector

By Express Computer (India), March 14, 2023

The insurance sector has experienced a paradigm shift because of big data analytics.

The insurance industry has long been known for its traditional, risk-averse nature. However, the emergence of technology has brought about significant changes in recent years. As consumers become more tech-savvy and demanding, the insurance industry has begun embracing technology in order to maintain its competitiveness and improve its services. A new wave of innovation known as "insurtech" has evolved, which refers to using technology to enhance and streamline insurance services. These companies are disrupting the traditional insurance market with new business models, products, and services. Thus, in order to provide specialized and effective insurance solutions, they use technology like big data, artificial intelligence, and machine learning. Also, these businesses offer their clients a user-friendly and practical digital experience, which is critical in today's fast-paced world.

The insurance industry has benefited from technological innovation by enhancing ease, personalization, transparency, efficiency, profitability, and risk management. To fully utilize the potential of technology in the insurance sector, however, issues including legislation, client uptake, and data privacy and security must be addressed.

From Express Computer

View Full Article

Wednesday, May 11, 2022

Interactive Big Data Journalism

Finding information and useful interaction with it.   Though unclear what a 'big data journalism' means here. 

Interactive Tools May Help People Become Big Data Journalists

By Penn State News, May 6, 2022

Pennsylvania State University (Penn State) researchers say people could use interactive tools to navigate, save, and tailor online content to extract meaning from big data.

The researchers found people more engaged with news websites that offered modality, message, and source interactivity tools than with sites that did not.

Penn State's S. Shyam Sundar said user experience is shaped by how these tools may be combined and how engaged users are in the topic; for example, the most engaging sites boasted a high concentration of modality tools and message interactivity.

"The types of interactivities we are talking about here can help users find information that they find personally meaningful and that they care about," Sundar said.

From Penn State News

View Full Article  

Tuesday, June 08, 2021

Scoping from Big Data to Big Knowledge

 This was brought to my attention by the 'Window Weekly Podcast' this week, in part because it dealt with conversations we had with Linkedin before they were acquired by Microsoft.   This did not lead anywhere, but touched on many aspects of how to handle corporate knowledge effectively.  See also the similarity to another system, called Zakta, which we called  a 'Collaborative Search Engine'.  See Zakta.com   Which uses classifications of kinds of knowledge.   Which we tested early on.  Worth a look.  

Project Alexandria is a research project within Microsoft Research Cambridge dedicated to discovering entities, or topics of information, and their associated properties from unstructured documents. This research lab has studied knowledge mining research for over a decade, using the probabilistic programming framework Infer.NET. Project Alexandria was established seven years ago to build on Infer.NET and retrieve facts, schemas, and entities from unstructured data sources while adhering to Microsoft’s robust privacy standards. The goal of the project is to construct a full knowledge base from a set of documents, entirely automatically.

The Alexandria research team is uniquely positioned to make direct contributions to new Microsoft products. Alexandria technology plays a central role in the recently announced Microsoft Viva Topics, an AI product that automatically organizes large amounts of content and expertise, making it easier for people to find information and act on it. Specifically, the Alexandria team is responsible for identifying topics and rich metadata, and combining other innovative Microsoft knowledge mining technologies to enhance the end user experience. ... '

https://www.microsoft.com/en-us/research/blog/alexandria-in-microsoft-viva-topics-from-big-data-to-big-knowledge/   Originally in MS Research.   ... ' 

Saturday, March 13, 2021

Reducing Complexity of Big Data

Feature detection and extraction to simplify data use in complex data.  Intriguing approach is worth a look,  Simplifying is always a good idea if it provides insight. 

Algorithm Could Reduce Complexity of Big Data

Texas A&M Engineering News, By Stephanie Jones, March 8, 2021

Researchers at Texas A&M University, the University of Texas at Austin, and Princeton University have developed an algorithm that can be applied to large datasets, with the ability to extract and directly order features from most to least salient. Texas A&M's Reza Oftadeh said, "There are many ad hoc ways to extract these features using machine learning algorithms, but we now have a fully rigorous theoretical proof that our model can find and extract these prominent features from the data simultaneously, doing so in one pass of the algorithm." The algorithm adds a new cost function to an artificial neural network to provide the exact location of features directly ordered by their relative performance, allowing it to perform classic data analysis on larger datasets more efficiently.  ... ' 

Wednesday, December 02, 2020

SAS: Four Principles of Analytics

As usual, SAS does a good job of precisely outlining analytics. Here including AI and Big Data with statistics. 

It's no secret that technology's changing. A change accelerating so fast, it's hard to keep up.

Big data and AI are exploding. Industries are reinventing themselves. And breakthroughs seem to redefine our world every day. Yet thankfully, some things do stay the same.

The four principles of analytics are steadfast truths that inform your approach to data and analytics. They're truths because, well, they work – helping you make the best decisions no matter what else changes. Read this blog to learn why:

Analytics follows the data.

Analytics is more than algorithms.

Data and analytics should be available to everyone.

Analytics is a differentiator.

 .... After all, change isn't the only thing that's constant. So are good ideas.

Read the Blog

Friday, October 09, 2020

Getting Serious about Data Science

Fairly obvious, seen many of these situations occur over the years in analytics and data interactions.  And how they are linked to business results.   Good checklist.   Non technical.

Getting Serious About Data and Data Science  To implement successful data programs, companies need to shift goals, muster resources, and align people.

By Thomas C. Redman and Thomas H. Davenport 

Data science, including analytics, big data, and artificial intelligence, is no longer a novel concept. Nor is the important foundation of high-quality data. Both have contributed to impressive business successes — particularly among digital natives — yet overall progress among established companies has been painfully slow. Not only is the failure rate high, but companies have also proved unable to leverage successes in one part of the business to reap benefits in other areas. Too often, progress depends on a single leader, and it slows dramatically or reverses when that individual departs the company. In addition, companies are not seizing the strategic potential in their data. We’d estimate that less than 5% of companies use their data and data science to gain an effective competitive edge. ... " 

Sunday, September 06, 2020

Data and Its World is Getting Bigger

Almost every talk I see starts with some statistics about the growth of data.  Here Datanami updates some of this, with related quantifying of related aspects.  Below I show just the first two, more at the link.  Impressive and even useful at times to scale your usage.  Ready?

10 Big Data Statistics That Will Blow Your Mind   By Alex Woodie in Datanami

 They call it “big data” for a reason–it’s really, really big. But getting your head wrapped around the growth of information digitization is not easy. That’s why we carefully curated these 10 mind-blowing facts about today’s data-geist, and how it’s projected to grow in the future.

1. The Global Datasphere will grow from 33 Zettabytes (ZB) in 2018 to 175 ZB by 2025, a 26% annual compound growth rate (CAGR), per IDC‘s DataAge 2025 report. However, only about 9ZB of that data will actually be stored,  up from about 0.9ZB in 2015. Only about one-third of the data that’s stored will actually be used, the analyst group says.

2. The annual capacity of shipped HDDs, SSDs, and LTO tape drives is projected to amount to about 1,300 exabytes in 2020, and will reach 4,500 exabytes by 2025, with HDDs accounting for the lion’s share of that capacity, according to Coughlin Associates. Per IDC, HDDs will account for more than 80% of enterprise storage needs by 2025, with legacy SSDs accounting for about 15% and newer NVMe-NAND solid state devices accounting for less than 5%.  .... " 

Tuesday, June 02, 2020

Astronomy Methods for Business Analytics?

Brought to attention by some of my astro colleagues, the effort is considerable.  Could businesses also construct such a 'survey' of how they operate?   Which could lead to a determination of where data might be used, needed?

The Vera C. Rubin Observatory, currently under construction in Chile, will conduct a vast astronomical survey of our dynamic Universe starting in 2022. They plan to collect 500 petabytes of image data by observing the skies continuously for 10 years and produce nearly instant alerts for objects that change in position or brightness every night. In addition to astronomical data, their dataset will include DevOps, IoT, and real-time monitoring data.

In this latest Data Science Central webinar, Dr. Angelo Fausti will demonstrate:

●     How a time-series database has the versatility to address their needs
●     How they created a solution to enhance visibility across their organization and improve actionable insights
●     How they pull software development and sensor data from their telescope, camera and observatory IoT devices

Speaker:
Dr. Angelo Fausti, Software Engineer - Vera C. Rubin Observatory

Hosted by:
Sean Welch, Host and Producer - Data Science Central

--- --------------------------------------------------------------------------

LSST Project Mission Statement
LSST’s mission is to build a well-understood system that provides a vast astronomical dataset for unprecedented discovery of the deep and dynamic universe. .... 


Monday, March 23, 2020

AI and Big Data

A frequent question I have received, is what is the difference between Big Data and AI?  My answer is DB is a means of using much more available data to perform analytical methods.  While AI uses a particular set of machine learning methods to find complex patterns in data.   In general AI methods are still less transparent but more powerful in some domains.    They can and we did use them in conjunction.   Sometimes there is little difference, both depend on large, sometimes very complex data.

Evolving Relationship Between Artificial Intelligence and Big Data  in ReadWrite   By Nitin Garg / 11 Jan 2020 / AI / Data and Security / Tech

Find the evolving relationship between big data and artificial intelligence. The growing popularity of these technologies offers engaging audience experience. It encourages newcomers to come up with an outstanding plan.

AI and Big Data help you transform your idea into substance. It helps you make full use of visuals, graphs, and multimedia to give your targeted audience with a great experience. According to Markets And Markets, the worldwide market for AI in accounting assumed to grow. As a result, growth from $666 million in 2019 to $4,791 million by 2024.

The critical component of delivering an outstanding pitch is taking a step further with an incredible plan of assuring success. Big data and Artificial intelligence help you contribute to multiple industries bringing an effective plan. It can directly speak to investors and your targeted audience, covering essential aspects and representing your idea in a nutshell.

According to Techjury, The big data analytics market is set to reach $103 billion by 2023, and in 2019, the big data market is expected to grow by 20%.  .... "

Tuesday, February 25, 2020

Framework to Predict Success of Big Data

In the HBR,  thoughts on predicting the success of Big Data.

Use This Framework to Predict the Success of Your Big Data Project
By Carsten Lund Pedersen, Thomas Ritter

Big data projects that revolve around exploiting data for business optimization and business development are top of mind for most executives. However, up to 85% of big data projects fail, often because executives cannot accurately assess project risks at the outset. We argue that the success of data projects is largely determined by four important components — data, autonomy, technology, and accountability — or, simply put, by the four D.A.T.A. questions. These questions originate from our four-year research project on big data commercialization.  ... " 

Thursday, January 02, 2020

Unlocking Data for Public Policy

Improved governmental data access for policy and regulation decisions.

Unlocking Data to Improve Public Policy
By Justine S. Hastings, Mark Howison, Ted Lawless, John Ucles, Preston White
Communications of the ACM, October 2019, Vol. 62 No. 10, Pages 48-53
10.1145/3335150

There is a growing consensus among policymakers that bringing high-quality evidence to bear on public policy decisions is essential to supporting the effective and efficient government their constituencies want and need. At the U.S. federal level, this view is reflected in a recent Congressional report by the Commission on Evidence-Based Policymaking, which recommends creating a data infrastructure that enables "a future in which rigorous evidence is created efficiently, as a routine part of government operations, and used to construct effective public policy."4

This article describes a new approach to data infrastructure for fact-based policy, developed through a partnership between our interdisciplinary organization Research Improving People's Lives and the State of Rhode Island.13 Together, we constructed RI 360, an anonymized database that integrates administrative records from siloed databases across nearly every Rhode Island state agency. The comprehensive scope of RI 360 has enabled new insights across a wide range of policy areas, and supports ongoing research into improving policies to alleviate poverty and increase economic opportunity for all Rhode Island residents (see the sidebar "Policy Areas in which RI 360 Has Contributed Insights"). Our approach can guide other policymakers and researchers seeking to similarly transform and integrate administrative data to guide and improve policy.

The role of administrative data in policymaking. Administrative data can be collected from the computer systems used by government agencies to run their programs. When transformed into databases that are more suitable for insights, these anonymized records provide new sources of facts for policymakers to benchmark goals and measure the successes and shortcomings of existing and future programs. Often classified as "big data"10 due to their volume, variety, and availability, administrative records are also an increasingly valuable source for empirical social science research.5 Research with administrative records can contribute new data-driven insights to inform important policy decisions (see the side-bar "Recent Data-Driven Insights from Administrative Records"), and add objectivity and scientific rigor to measuring program impact and designing effective program changes. Moreover, scientists can inform how data from administrative systems, which are primarily designed around operational needs and often not suitable for analysis, can be transformed effectively to support research and insights.

Although the idea of guiding policy with data dates back to the 1970s and 1980s, early studies only considered isolated data sources and come from a time when data was scarce. It was not until recently that advances in data collection, storage, and scale provided the opportunity to integrate data across nearly every facet of government. Early case studies and survey studies highlight how the process of data modeling can facilitate negotiation and consensus-building among policymakers,8 but also how the unmet promises of new information technologies prompted frustration among government leaders at that time.9

An important lesson is to engage policymakers and leaders to fully understand their needs, which is why we formed extensive partnerships with state government leaders while building RI 360. Integrated administrative data can support not only academic research, but also the analytics requirements of government itself. Like researchers, government analysts need access to data that has been transformed to provide insights and integrated across programs that serve what are often overlapping populations. For these reasons, RI 360 was selected as the primary data source for the Rhode Island Executive Office of Health and Human Service's Data Ecosystem project, to empower its data analysts and partners with data optimized for insights.   .... "

Tuesday, November 05, 2019

Listening for Indicators of Failure

Not an uncommon thing that is done by human engineers responsible for maintenance.  Now proposed to be more autonomous.

'Listening' to Engine Blades to Stop Failures, Disasters   By Purdue University News

This new method can measure and monitor blade vibration at the same time.
Purdue University researchers have developed a monitoring system that can detect one of the most common causes of premature blade failure in gas turbine engines: rotor forced response vibration.

Purdue University researchers have developed a monitoring system to identify the sound of rotor-forced response vibration, a common cause of premature blade malfunction in gas turbine engines.

The system employs multiple unsteady pressure sensors to listen for specific pressure waves associated with turbine engine blade vibration. Gas turbine blades are typically very lightly damped, and can function like a tuning fork; their blades can produce a specific frequency or tone when they resonate, which the pressure sensors are designed to pick up.

Said Purdue's Nicole Key, "With the help of big data analytics, the blade vibration information can be used to predict the possible engine failure and optimize preventive maintenance schedules."

From Purdue University News
View Full Article

Tuesday, October 29, 2019

Splunk and the Data Problem

Had several interactions with Splunk, in the enterprise,liked it for streaming data problems.  But I would rather say that every problem has an embedded data problem, and that size of that can vary according to the business context. Don't remember Splunk being particularly aimed or useful at that, but seems that are building towards that.  Will take a look.

Every problem’s a data problem, says Splunk. Can its new platform fix them?  By R. Danes in Siliconangle

Splunk Inc. is getting serious about this data platform thing. The company wants to get friendly with data from any source — not just the Splunk index. The idea is that a large, inclusive platform can ultimately get more juice from data — business insights, social impact, etc. — than a hodgepodge of software products.

“We believe, at the heart of every problem, is a data problem,” said Susan St. Ledger (pictured), president of worldwide field operations at Splunk. Think that’s an overstatement? St. Ledger named wildfires, the opioid crises, and human trafficking as examples of the issues people are attacking with data today.  ..." ... '

Sunday, October 27, 2019

Data Management for Data Science

In Kdnuggets a good description and visualization of data management needed for data science.

Everything a Data Scientist Should Know About Data Management

For full-stack data science mastery, you must understand data management along with all the bells and whistles of machine learning. This high-level overview is a road map for the history and current state of the expansive options for data storage and infrastructure solutions. By Phoebe Wong and Robert Bennett.

To be a real “full-stack” data scientist, or what many bloggers and employers call a “unicorn,” you have to master every step of the data science process — all the way from storing your data, to putting your finished product (typically a predictive model) in production. But the bulk of data science training focuses on machine/deep learning techniques; data management knowledge is often treated as an afterthought. Data science students usually learn modeling skills with processed and cleaned data in text files stored on their laptop, ignoring how the data sausage is made.   ... " 

Thursday, May 30, 2019

Business Cards and Big Data

This made me think of other 'lost' sources of data

Here’s How Big Data And Business Card Marketing Go Together

Big data and business card marketing are a match made in heaven, and contrary to popular belief, business card marketing is here to stay.

Many people believe that digital media is rapidly replacing traditional forms of branding. They believe that advances in big data have made business cards, brochures and direct mail marketing obsolete.

Nothing could be further from the truth. We previously published an article on the state of direct mail marketing. We showed that marketers are actually using big data to improve the performance of their direct mail marketing campaigns. ...

We can draw a similar conclusion about the relevance of business cards in 2019. Online marketing did not make business cards go out of style. Data Floq made this point clear in a post they made in 2016. ... "

By Diana Hope

Saturday, May 18, 2019

Egeria for Sharing Metadata

We spent much time on the application of relevant contextual metadata.  One big issue is making sure the quantity and quality of metadata is preserved in context for specific goals.  Maintenance again.

ODPi Announces Egeria for Open Sharing, Exchange and Governance of Metadata

Industry’s First Open Metadata Standard Helps Organizations Better Understand, Manage and Gain Value from Data

Vancouver, BC, Canada – August 27, 2018 – Open Source Summit North America — ODPi, a nonprofit organization accelerating the open ecosystem of big data solutions, today announced Egeria, a new project from ODPi that supports the free flow of metadata between different technologies and vendor offerings. Egeria enables organizations to locate, manage and use their data more effectively.

Last year’s ODPi white paper on “The Year of Enterprise-wide Production Hadoop” found that Data Governance and Security were the biggest blocking factors to enabling enterprises to take big data into true production. Recent data privacy regulations such as GDPR have brought these concerns to the forefront, and enterprises around the globe need a standard for ensuring that data providence and management is clear and consistent across the enterprise. Egeria enables this, as the only open source driven solution designed to set a standard for leveraging metadata in line of business applications, and enabling metadata repositories to federate across the enterprise.

“A consistent view on data across the entire landscape is essential for any organisation that wants to become data driven. Not just where the data is, but also the quality, the ownership, and the full lineage across the entire set of technologies used,” said Ferd Scheepers, chief information architect, ING. “The open metadata standard delivered by Egeria delivers this consistent view across all the technologies, while reducing the cost of metadata capture, and the management challenges of working with various data tool vendors.”

Egeria is built on open standards and delivered via Apache 2.0 open source license. The ODPi Egeria project creates a set of open APIs, types and interchange protocols to allow all metadata repositories to share and exchange metadata. From this common base, it adds governance, discovery and access frameworks for automating the collection, management and use of metadata across an enterprise. The result is an enterprise catalog of data resources that are transparently assessed, governed and used in order to deliver maximum value to the enterprise.

“Egeria’s open source metadata management presents an exciting opportunity to rethink both management and governance of data to provide greater trust and flexibility in how we all share and consume data,” said John Mertic, director of program management, ODPi. “Egeria’s open governance model allows our community and practitioners to develop and evolve the base for use in any offerings and deployments.” .... '

Wednesday, March 27, 2019

McDonald's Acquires Dynamic Yield for Decision Logic

Companies leveraging advanced analytic methods for their data by acquisition.  Have recently stood in front of McDonald's store kiosk and after looking for the coffee I wanted, wondered  what recommendation it should suggest, and what data and context it had at hand to make that recommendation.

McDonald's Bites on BigData with $300 Million Acquisition By Brian Barrett in Wired  

Mention McDonald’s to someone today, and they're more likely to think about Big Mac than Big Data. But that could soon change: The fast-food giant has embraced machine learning, in a fittingly super-sized way.

McDonald’s is set to announce that it has reached an agreement to acquire Dynamic Yield, a startup based in Tel Aviv that provides retailers with algorithmically driven "decision logic" technology. When you add an item to an online shopping cart, it’s the tech that nudges you about what other customers bought as well. Dynamic Yield reportedly had been recently valued in the hundreds of millions of dollars; people familiar with the details of the McDonald’s offer put it at over $300 million. That would make it the company's largest purchase since it acquired Boston Market in 1999.  ... " 

and also:

McDonald’s to use A.I. to tempt you into extra purchases at the drive-thru  By Trevor Mogg in DigitalTrends   ...  "

Friday, March 15, 2019

Big Data's Biggest Challenge

More Broadly, losing site of the goals.  Crucial for any analytical project.

Big Data’s Biggest Challenge: How to Avoid Getting Lost in the Weeds  from K@W. 
Podcast and Transcript

Wharton's Raghuram Iyengar and Evite CEO Victor Cho discuss how firms can optimize  

Companies have access to more data than ever before. But how can they optimize it without getting lost in the weeds – or losing sight of the customer? Evite CEO Victor Cho and Wharton marketing professor Raghuram Iyengar offered advice from their own experiences during a recent conversation with Knowledge@Wharton. Cho was on campus to host a Datathon with the Wharton Customer Analytics Initiative. Penn students from multiple academic majors were given datasets from Evite and asked to come up with solutions based on the data for improving Evite’s platform and increasing revenue. Evite is among the participants in WCAI’s corporate partner program, which seeks to help companies find ways to better use their data through collaborations with academic researchers, student projects and other initiatives.

An edited transcript of the conversation follows:

Friday, January 04, 2019

Frameworks for Big Data Ecosystems

Of interest, below is an abstract, just a bit more at the link, further detail requires signup.  Ecosystems are hard to sell, but ultimately essential.

Framework for Implementing a Big Data Ecosystem in Organizations
By Sergio Orenga-Roglá, Ricardo Chalmeta 

Communications of the ACM, January 2019, Vol. 62 No. 1, Pages 58-65
10.1145/3210752

Enormous amounts of data have been generated and stored over the past few years. The McKinsey Global Institute reports this huge volume of data, which is generated, stored, and mined to support both strategic and operational decisions, is increasingly relevant to businesses, government, and consumers alike,7 as they extract useful knowledge from it.

There is no globally accepted definition of "big data," although the Vs concept introduced by Gartner analyst Doug Laney in 2001 has emerged as a common structure to describe it. Initially, 3Vs were used, and another 3Vs were added later.13 The 6Vs that characterize big data today are volume, or very large amounts of data; velocity, or data generated and processed quickly; variety, or a large number of structured and unstructured data types processed; value, or aiming to generate significant value for the organization; veracity, or reliability of the processed data; and variability, or the flexibility to adapt to new data formats through collecting, storing, and processing.  ... "

Tuesday, November 06, 2018

Introduction to the Community Data License Agreement

CSIG (Cognitive Systems Institute Group) Talk — Thursday Nov 8, 2018 - 10:30-11am US Eastern 

Talk Title: "Introduction to the Community Data License Agreement: An "Open Source" Agreement Specifically Designed for Data and Content Sharing and Analysis in the Big Data world

Speaker: Christopher O'Neill, IBM 

Abstract: This talk will provide a general introduction the Community Data License Agreement ("CDLA") family of agreements, published  by The Linux Foundation in late 2017. Topics will include the general structure of the agreements, with a particular focus on their unique  data analysis terms and other aspects in which they are designed to address the unique characteristics of data from an Intellectual  Property perspective. 

Bio: Christopher O'Neill is Associate General Counsel — Intellectual Property Law at IBM Corporation, based in Armonk, New York. In over  25 years at IBM, Mr. O'Neill has held a variety of positions, both in IBM's product businesses and in its litigation group. In his current role,  Mr. O'Neill has responsibility for a variety of matters, including IP indemnity matters, data rights issues, adversely-held patent matters,  and open source issues. He graduated from New York University School of Law in 1987 and is a member of the New York bar. 

Zoom meeting Link: https://zoom.us/j/7371462221; Zoom Cailin: (415) 762-9988 or (646) 568-7788 Meeting id 7371462221 
Zoom International Numbers: https://zoom.us/zoomconference 
Check http://cognitive-science.info/community/weekly-update/ for recordings & slides, and for any date & time changes 

Join Linkedin Group: https://www.Iinkedin.com/groups/6729452/ (Cognitive Systems Institute) to receive notifications 
Thu, Nov 8, 10:30am US Eastern • https://zoom.us/j/7371462221 
More Details Here : http://cognitive-science.info/community/weekly-update/   (Also slides and talk recording will be placed here) 

More at: https://www.linuxfoundation.org/press-release/2017/10/linux-foundation-debuts-community-data-license-agreement/  #CSIGnews #opentechai #AI @KarolynSchalk @mattganis @jwaup @MishiChoudhary @t_streinz @hyurko