/* ---- Google Analytics Code Below */
Showing posts with label Metadata. Show all posts
Showing posts with label Metadata. Show all posts

Tuesday, March 22, 2022

TopQuadrant: Data Fabric Management Webinar

Had to miss this one, of particular interest, worked with TopQuadrant

 ... Sorry that you were not able to attend our recent webinar: " How Metadata Management Must Evolve to Support Data Fabric".

The recording and slides from the webinar are available here: https://www.topquadrant.com/how-metadata-management-must-evolve-to-support-data-fabric/.

Should you have any follow-up questions or would like to explore all of the capabilities in TopBraid EDG (http://www.topquadrant.com/products/topbraid-enterprise-data-governance/) in more detail please contact us at edg-info@topquadrant.com ... '

Friday, March 11, 2022

Webinar: How Metadata Management Must Evolve to Support Data Fabric

Looks to be useful, plan to attend:

TopQuadrant: How Metadata Management Must Evolve to Support Data Fabric   by Irene Polikoff | Mar 8, 2022 | Metadata Management, Webinars

About This Webinar

On Thu, Mar 17, 2022 11:30 AM EDT

If you have not heard the term “data fabric” yet, you will. It is rapidly growing in popularity. Gartner identified data fabric as the top trend for data and analytics in 2021.

You can think of data fabric as a web connecting multiple locations, types, and sources of data – both on-premises and in the public cloud. It is an architectural approach designed to help organizations better deal with the growing number of available data sources and ever-changing application requirements. The backbone of data fabric design is a Knowledge Graph capturing information about data sources in RDF. This is a new type of a data catalog with semantically augmented and enriched metadata.

Join us for this webinar to:

Learn What Is Data Fabric

Understand Why Data Fabric Requires a Knowledge Graph

Get Advice on Moving from Traditional Metadata Management to Metadata Management With Knowledge Graphs

Envision How Tools Participating in Data Fabrics Will Interact With Knowledge Graphs

See these Concepts in Action  ... 

 Register:   https://register.gotowebinar.com/register/7807142108219203083 

Wednesday, February 16, 2022

Metadata – The Magic Behind Data Fabric

Topquadrant talks active metadata.

Metadata – The Magic Behind Data Fabric

by Irene Polikoff | Feb 7, 2022 | Blog from Topquadrant

The main goal of creating an enterprise data fabric is not new. It is the ability to deliver the right data at the right time, in the right shape, and to the right data consumer, irrespective of how and where it is stored. Data fabric is the common “net” that stitches integrated data from multiple data and application sources and delivers it to various data consumers. 

So, what makes the data fabric approach different from previous, more traditional data integration architectures? The key differentiator of a data fabric is its fundamental reliance on metadata to accomplish this goal. Implementing a data fabric means establishing a metadata-driven architecture capable of delivering integrated and enriched data to data consumers.To emphasize this point, Gartner coined the term active metadata. 

Data fabric relies on active metadata. 

Metadata describes different aspects of data. The more comprehensive the sets of metadata we collect, the better they will be able to support our application scenarios. Traditionally, metadata categories have included:

Business metadata – provides the meaning of data through mappings to business terms.

Technical metadata – provides  information on the format and structure of the data such as physical database schemas, data types, data models.

Operational metadata – describes details of the processing and accessing of data such as data sharing rules, performance, maintenance plans, archive and retention rules.

More recently, a new category of metadata became important – Social metadata. It typically includes discussions and feedback on the data from its technical and business users. Business metadata has evolved beyond just mapping to terms to now encompassing ontologies to better assist with interpreting data’s context and meaning.

How is active metadata different from passive metadata? Gartner defines passive metadata as any metadata that is collected. Some Gartner analysts equate active metadata with metadata that is being used. By use, we mean use of the metadata by software (such as software components within the data fabric) in support of a broad range of data integration, analysis, reporting and other data processing scenarios. Other analysts push this concept further and say that active metadata is created by the data fabric by analyzing passive metadata and using the results to recommend or automate tasks.

Irrespective of the exact definition of active metadata, the underlying premise of the data fabric is that the optimal solution for the delivery of the right data in the right shape is to leverage its metadata. For example, the data fabric may use metadata to:  ... ' 

Sunday, February 13, 2022

Plausible Deniability

 Fascinating thoughts on the subject by Bruce Schneier. with considerable comment and discussion, considering about how this would be useful in multiple contexts.  And how it might be best tested in each. Below the intro, more at the link. 

Bunnie Huang’s Plausibly Deniable Database

Bunnie Huang has created a Plausibly Deniable Database.

Most security schemes facilitate the coercive processes of an attacker because they disclose metadata about the secret data, such as the name and size of encrypted files. This allows specific and enforceable demands to be made: “Give us the passwords for these three encrypted files with names A, B and C, or else…”. In other words, security often focuses on protecting the confidentiality of data, but lacks deniability.

A scheme with deniability would make even the existence of secret files difficult to prove. This makes it difficult for an attacker to formulate a coherent demand: “There’s no evidence of undisclosed data. Should we even bother to make threats?” A lack of evidence makes it more difficult to make specific and enforceable demands.   ... ' 


Friday, October 01, 2021

Machines Unlearning

Perhaps yet more important, unlearn in new contexts.  Say for maintenance. With new or changed metadata.  And help us understand the difference.

Now that machines can learn, can they unlearn?

Researchers see if they can remove sensitive data without retraining AI from scratch.  By Tom Simonite, Wired.com

Companies of all kinds use machine learning to analyze people’s desires, dislikes, or faces. Some researchers are now asking a different question: How can we make machines forget?

A nascent area of computer science dubbed machine unlearning seeks ways to induce selective amnesia in artificial intelligence software. The goal is to remove all trace of a particular person or data point from a machine-learning system, without affecting its performance.

If made practical, the concept could give people more control over their data and the value derived from it. Although users can already ask some companies to delete personal data, they are generally in the dark about what algorithms their information helped tune or train. Machine unlearning could make it possible for a person to withdraw both their data and a company’s ability to profit from it.

Although intuitive to anyone who has rued what they shared online, that notion of artificial amnesia requires some new ideas in computer science. Companies spend millions of dollars training machine-learning algorithms to recognize faces or rank social posts, because the algorithms often can solve a problem more quickly than human coders alone. But once trained, a machine-learning system is not easily altered, or even understood. The conventional way to remove the influence of a particular data point is to rebuild a system from the beginning, a potentially costly exercise. “This research aims to find some middle ground,” says Aaron Roth, a professor at the University of Pennsylvania who is working on machine unlearning. “Can we remove all influence of someone’s data when they ask to delete it, but avoid the full cost of retraining from scratch?”

Sunday, August 08, 2021

Telling the Data's Story

 Good thoughts,  cant say we ever did this completely, would have been especially useful with the metadata, which tends to me less well understood.     In some cases when reviewing data needs with decision makers even the need and existence of particular metadata was a useful revelation. 

Data doesn’t speak for itself: Why data storytelling is important  by 7wData  August 7, 2021

Almost all data is a recording of past events - what has happened before. People have to explore and analyze it to find historical truth, in the hopes it can reveal trends, inform the next direction to take, and act as a guide for the actions necessary to improve the future.

For analysts in an organization who can read data presented as is - on dashboards, reports, charts - this traditional analytical process may be enough to make sense. But not everyone can consume or understand data shown upfront, or extract value from it.

Helping everyone understand what’s happening and getting them invested in taking action to get toward your ideal state, then, can only occur when you use the right skills and right analytics tools to not only communicate the data well, but make it memorable.

This need is why data storytelling is so important right now - especially because telling stories with data will be the most widespread way of consuming our analytics by 2025.

In this blog, we want to focus on the many reasons data storytelling is as important as any other modern analytics initiative in helping your users make decisions today.

Firstly, to make data more useful for decision-making for people who aren’t experts, it must be given a clear and compelling voice. Problems and opportunities (the what) may just be numbers on a dashboard at this point; possibly interesting, but not clear for everyone on what to do next.

Combining narrative with data is a great way for organizations to better explain the ‘why’ behind the results, and tell an engaging story of how an insight was discovered or conclusion was drawn, so everyone can connect with and understand why it’s important:

These complex questions aren't so easily conveyed in dashboards or charts alone. The answers require nuance, interpretation and sometimes arguments, for people to ‘get’ it.   ... ' 

Wednesday, January 20, 2021

Data Catalogs vs Discovery

Depends... Internally controlled vs externally, other ...   Metadata also a key component. Governance too.

Data Catalogs Are Dead; Long Live Data Discovery  in Towardsdatascience

Why we need to rethink our approach to metadata management and data governance

By Barr Moses... 

Sunday, September 13, 2020

Finding the New from the Old

Makes the case that you can use older data to make useful conclusions.  Though in our experience the required meta data was often not available.    Physics/astronomy is at least often based on some strong science foundations.

AI Identifies 50 New Planets From Old NASA Data
Jessie Yeung
August 26, 2020

Machine learning artificial intelligence (AI) developed by astronomers and computer scientists at the U.K.'s University of Warwick found 50 new planets by mining old data from the U.S. National Aeronautics and Space Administration (NASA). The researchers educated the algorithm on data collected by the Kepler Space Telescope, teaching it to differentiate real planets from false positives. The AI was then tasked to analyze old datasets of planetary candidates, in which it discovered the 50 previously unknown exoplanets. Warwick's David Armstrong said this is the first time machine learning has been used to rank planetary candidates in a probabilistic framework, and the research suggests the AI could "validate thousands of unseen candidates in seconds."... '

Thursday, August 20, 2020

Problems from Data Intensive Science

Quite interesting thoughts.  Basically managing data for many needs, motivations, coming from many directions.  And please, consider too the metadata, brought us to our knees many times!

Thorny Problems in Data (-Intensive) Science
By Christine L. Borgman, Michael J. Scroggins, Irene V. Pasquetto, R. Stuart Geiger, Bernadette M. Boscoe, Peter T. Darch, Charlotte Cabasse-Mazel, Cheryl Thompson, Milena S. Golshan
Communications of the ACM, August 2020, Vol. 63 No. 8, Pages 30-32 10.1145/3408047

As science comes to depend ever more heavily on computational methods and complex data pipelines, many non-tenure track scientists find themselves precariously employed in positions grouped under the catch-all term "data science." Over the last decade, we have worked in a diverse array of scientific fields, specializations, and sectors, across the physical, life, and social sciences; professional fields such as medicine, business, and engineering; mathematics, statistics, and computer and information science; the digital humanities; and data-intensive citizen science and peer production projects inside and out of the academy.3,7,8,15 We have used ethnographic methods to observe and participate in scientific research, semi-structured interviews to understand the motivations of scientists, and document analysis to illustrate how science is assembled with data and code. Our research subjects range from principal investigators at the top of their fields to first-year graduate students trying to find their footing. Throughout, we have focused on the multiple challenges faced by scientists who, through inclination or circumstance, work as data scientists.

The "thorny problems" we identify are brambly institutional challenges associated with data in data-intensive science. While many of these problems are specific to academe, some may be shared by data scientists outside the university. These problems are not readily curable, hence we conclude with guidance to stakeholders in data-intensive research.   ... '   (full article at the link)

Sunday, May 17, 2020

Data is the Most Important Thing

Or poor associated metadata.

Poor data quality is the leading cause of digital transformation failure. It’s time to prioritise data transformation!   in 7wdata

In a digitally empowered age, businesses across the globe have grand ambitions of leveraging the power of AI, Big Data, and machine learning. Innovation is happening at a rapid pace. Companies are investing millions of dollars in building data lakes, moving to the cloud, hiring data scientists and chief data officers to run their Digital Transformation plans.

Yet, they fail. Spectacularly. Plenty of reports and surveys show that over 85 per cent of big data projects are failing with varying causes. 

Enough has been written lately about how business cultures and unchecked ambitions lead to big data project failures. This piece will focus on how poor data quality is often overlooked and makes for one of the leading cause of digital transformation failure.

Data transformation, the process of transforming raw data into a usable format is often, incorrectly, juxtaposed with digital transformation. Companies assume that because they are implementing data lakes, data centres or new ERPs, (which are all part of digital transformation), they are transforming their data.   ... " 

Saturday, April 25, 2020

Knowledge Graphs versus Property Graphs

This came in the mail, it had been asked in an interaction recently.  Worth a look.  We worked with TopQuadrant.

New White Paper: Knowledge Graphs versus Property Graphs

We are in the era of graphs. Graphs are hot. Why? Flexibility is one strong driver: heterogeneous data, integrating new data sources, and analytics all require flexibility. Graphs deliver it in spades.

The two main graph data models are: Property Graphs and Knowledge (RDF) Graphs. People who want to take advantage of graph-based solutions for data and metadata management want to know what they are, what are their similarities and differences, and what they are each good for.

This white paper covers the following, it:
Describes the two main graph data models: Property Graphs and RDF Graphs and explains the key differences in their terminology and capabilities
Compares their strengths and limitations
Provides guidance on their respective capabilities

Other TopQuadrant resources to explore:
RECORDED WEBINARS, including this most recent one: "Getting Started with Data Governance"
WHITE PAPER COLLECTION, including: "Implementing Data Governance with Knowledge Graphs to Enable Enterprise AI"

Download Now    https://www.topquadrant.com/knowledge-assets/whitepapers/

This email sent by TopQuadrant   www.topquadrant.com

Saturday, April 11, 2020

On Metadata and Cooking

In almost every piece of work I have done in enterprises, there has been a need to deal with metadata.  Metadata can mean a number of things .... like for example 'data about data'.  for example the number of hits in a search defining data.    But in most of the cases I have worked with 'metadata'  means supporting data, and is often data that comes from a different source, that may be hard to find, or may need to be specially created for an effort.

Well done piece in the BiPolar blog by Matthew Roche.  Which compares metadata to recipe based cooking.

BI Polar
Business Intelligence, Data Governance, Mental Health, Diversity, Martial Arts, and Heavy Metal.
Metadata is Not a “Nice to have” 

He also has other posts on metadata I have not read yet.

Thursday, April 02, 2020

Search Tool for Science Metadata

Interesting idea to connect public and private data in science databases.    Have worked on this kind of application, and the key is having the right kind of data to support the goal.

Research Tool Creates Metadata for Scientific Content   By ResoluteAI in CACM

ResoluteAI has launched Nebula, an enterprise search solution for scientific organizations. Nebula tags and categorizes content, connecting public and proprietary information to ResoluteAI's platform of comprehensive scientific databases.

"Nebula deploys domain-specific artificial intelligence on the institutional knowledge that is mission critical to science-driven enterprises," says Steve Goldstein, CEO of ResoluteAI. "Nebula lets our clients find, analyze, and visualize their own data that was — until now — unsearchable in existing storage and knowledge management systems."

Nebula ingests content in any file format — text, video, audio, embedded documents, and more — from any storage system — SharePoint, Azure Blob, S3, Box, Dropbox, and the like — and creates what it calls Precision Metadata that can then be applied to connecting concepts, ideas, and discoveries across an organization.

"We are extremely interested in products that apply artificial intelligence to enable search of our proprietary scientific and technical content, and believe that ResoluteAI provides a practical, user-friendly, and smart solution," says Mike Dale, Digital R&D Director of Unilever. "We are truly excited about Nebula's potential impact on our internal knowledge management, and anticipate significant benefits to our organization as we continue with the deployment."

"Our Foundation solution already creates structured metadata from unstructured content that's available via public databases," says Matt Doherty, CTO & Founder of ResoluteAI. "Nebula applies ResoluteAI's signature search techniques to proprietary enterprise data, analyzing and tagging enormous inventories of legacy content. We allow teams to find anything, anywhere, like an image in a PowerPoint, a phrase in a video, or a graph in a document."

Foundation is ResoluteAI's discovery engine for scientific content — a multi-source research hub that accesses scientific content about patents, companies, grants, clinical trials, publications, and technology transfer opportunities to be searched from a single source. Nebula can be connected to external content through Foundation, bridging the gap between proprietary information retrieval and public content search.  ... " 

Polinode System Adds Features

I see that the Polinode system, which I have followed for some time, has added some interesting new features.   Much more at the link.

Five New Features Including an Integration with Office 365
By Andrew Pitts

A number of new features have recently gone live in Polinode so we wanted to pull them all together and summarize them for you. They are all designed to make organizational network analysis even easier and more powerful!

1. OFFICE 365 INTEGRATION

As a platform, Polinode allows you to conduct both active organizational network analysis (where you ask respondents questions like “who do you turn to for advice?”) as well as passive organizational network analysis (where you use communication data such as email metadata). For quite a few years, organizations have successfully used Polinode for both types of organisational network analysis. Typically, where the platform has been used for passive organizational network analysis, organizations have prepared their own communication data as networks and uploaded that prepared data to Polinode. With this new integration though it’s possible for an admin user to simply provide authorization for Polinode to access Office 365 metadata and the platform will then automatically collect and store this metadata on a rolling daily basis. This data is then available to run queries against with these queries then producing interactive networks, i.e. networks that a user can then calculate different measures of centrality, community detection and brokerage on. The screenshot below illustrates our new Office 365 add-on and how easy it is to connect (note that you can click on any of the images below for larger versions).    .... "

Saturday, March 07, 2020

Leveraging with Sensitive Metadata

Its all about the metadata, that is the context, of information being gathered.

MIT News
Rob Matheson
February 26, 2020

Massachusetts Institute of Technology (MIT) researchers have developed a scalable metadata-protection scheme to shield the information of millions of users of communications networks against possible state-level surveillance. In the Crossroads (XRD) scheme, users send encrypted messages to multiple server chains, with each chain mathematically ensured to have at least one hacker-free server. Each server decrypts and randomly shuffles the messages before sending them to the next server down the line; the final server decrypts the last encryption layer and transmits the message to the target recipient. XRD also uses aggregate hybrid shuffle, a type of cryptographic proof that guarantees servers are properly receiving and shuffling messages to identify malicious activity. MIT's Albert Kwon said, "We want to get to the point where we're sending metadata-protected messages in near-real-time."

Monday, December 23, 2019

Google's Summarization Performance: Pegasus

A kind of AI.   Summarization is useful and powerful concept.  But consider that summarization also exists in a context.   Its output is only useful in a particular context, and that exists based  also on the requester.  And that can also change over time, location, requesters current goals, etc.    And influenced by the metadata involved with its construction.   Still a very useful step forward.

Google Brain’s AI achieves state-of-the-art text summarization performance
Kyle Wiggers in Venturebeat

Summarizing text is a task at which machine learning algorithms are improving, as evidenced by a recent paper published by Microsoft. That’s that’s good news — automatic summarization systems promise to cut down on the amount message-reading done by enterprise workers, which one survey estimates amounts to 2.6 hours each day.

Not to be outdone, a Google Brain and Imperial College London team built a system — Pre-retraining with Extracted Gap-sentences for Abstractive SUmmarization Sequence-to-sequence, or PEGASUS — that leverages Google’s Transformers architecture combined with pre-training objectives tailored for abstractive text generation. They say it achieves state-of-the-art results in 12 summarization tasks spanning news, science, stories, instructions, emails, patents, and legislative bills, and that it shows “surprising” performance on low-resource summarization, surpassing previous top results on six data sets with only 1,000 examples.  .... " 

Sunday, September 15, 2019

Building Knowledge Graphs

Have been looking at means of continuously and coherently connecting company data sources to analytical and AI methods.    Most recently have looked at the idea of 'Knowledge Graphs'.   Of interest, an upcoming webinar by Neo4j on knowledge graphs, here particularly on financial services, but applicable beyond that.   Note in particular the mention of 'intelligent metadata',  which we posed to construct understandable and maintainable data sources.   Will be attending.

Financial Services Companies Make Disparate Data Simple with Knowledge Graphs Webinar

Tuesday, September 24    11:00 am PST

Knowledge graphs are driving industry disruption and business transformation by bringing together previously disparate data, using connections for superior decision support, and adding context for more intelligent applications (including AI). In this session, we’ll walk through the fundamental elements of knowledge graphs including contextual relevance, dynamic self-updating, understandability with intelligent metadata, and the combination of heterogeneous data. ....'

More information and register here.

Tuesday, September 10, 2019

Automating and Optimizing Experiment Data Collection

 Collecting data from processes, and using some automatic method of choosing, pre-analysis, cleansing, visualizing and tagging, associating with metadata .... can be very useful.   Here a more complex example.

SMART Algorithm Makes Beamline Data Collection Smarter
By Lawrence Berkeley National Laboratory

The "data deluge" in scientific research stems in large part from the growing sophistication of experimental instrumentation and optimizing tools — often using machine- and deep-learning methods — to analyze increasingly large data sets. But what is equally important for improving scientific productivity is the optimization of data collection — aka "data taking" — methods.

Toward this end, Marcus Noack, a postdoctoral scholar at Lawrence Berkeley National Laboratory in the Center for Advanced Mathematics for Energy Research Applications (CAMERA), and James Sethian, director of CAMERA and Professor of Mathematics at UC Berkeley, have been working with beamline scientists at Brookhaven National Laboratory to develop and test SMART (Surrogate Model Autonomous Experiment), a mathematical method that enables autonomous experimental decision making without human interaction. A paper describing SMART and its application in experiments at Brookhaven's National Synchrotron Light Source II (NSLS-II) are described in "A Kriging-Based Approach to Autonomous Experimentation with Applications to X-Ray Scattering," published in Scientific Reports.

"Modern scientific instruments are acquiring data at ever-increasing rates, leading to an exponential increase in the size of data sets," says Noack, lead author on the paper. "Taking full advantage of these acquisition rates requires corresponding advancements in the speed and efficiency not just of data analytics but also experimental control."

The goal of many experiments is to gain knowledge about the material that is studied, and scientists have a well-tested way to do this: they take a sample of the material and measure how it reacts to changes in its environment. User facilities such as Brookhaven's NSLS-II and the Center for Functional Nanomaterials offer access to high-end materials characterization tools. The associated experiments are often lengthy, and complicated procedures and measurement time is precious. A research team might only have a few days to measure their materials, so they need to make the most of each step in each measurement.  .... " 

Thursday, August 01, 2019

The Value of Lineage Metadata

Did lots of work with metadata, especially in the corporate laboratory space.   Had never heard the term 'Lineage Metadata', though again it was often considered in our work.   Now would think this is more important than ever, if we are to create useful predictions, and also add some accuracy to future maintenance of any automated systems that emerge.  Lineage predicts changing context.   At very least the lineage can determine what errors might exist in the data, but in reality should provide much more in AI.  Should always be considered.     Lineage can also be considered as part of data asset value, predicting stability of value in context.    Also be included in a semantic representation of data involved.
I like that this is presented here.  Below an excerpt, much more at the link.- FAD

Lineage Metadata: The Fuel for Data Governance in Informationweek
Moshe Kranc is the chief technology officer at Ness Digital Engineering

The best way to achieve data quality is by combining or blending these three techniques: decoded lineage, data similarity lineage and manual lineage mapping.

Enterprises aspire to derive insights from data that can provide a competitive advantage. The most common impediment to achieving this goal is poor data quality. If the data that is being input to a predictive algorithm is “dirty” (with missing or invalid values), then any insights produced by that algorithm cannot be trusted.

To achieve data quality, it’s not enough to clean up the existing historical data. You also need to ensure that all newly generated data is clean by instituting a set of capabilities and processes known collectively as data governance. In a governed data environment, each type of data has a data steward who is responsible for defining and enforcing criteria for data cleanliness. And, each data value has a clearly defined lineage: We know where it came from, what transformations it underwent along the way, and what other data items are derived from this data value.

Data lineage provides an enterprise with many benefits:

The ability to perform impact analysis and root-cause analysis, by tracing lineage backwards (to find all data that influenced the current data) or forwards (to identify all other data that is impacted by the current data) from a given data item;
Standardization of the business vocabulary and terminology, which facilitates clear communication across business units;

Ownership, responsibility and traceability for any changes made to data, thanks to the lineage’s comprehensive record of who made what changes and when.

It sounds great, but where does data lineage information come from? Looking at a specific data value in the database tells us its current value, but it will not provide information about how the data evolved into its current value. What is missing is data about the data (lineage metadata) that automatically remembers the time and source of every change made to every data item, whether the change was made by software or by a human database administrator.

There are three competing techniques for collecting lineage metadata, each of which has its strengths and weaknesses: .... " 

Thursday, July 11, 2019

Integrating Metadata

In the latest ACM, interesting piece on transferring data.   And notable about mentioning the metadata ultimately involved ....

Extract, Shoehorn, and Load  By Pat Helland
Communications of the ACM, July 2019, Vol. 62 No. 7, Pages 32-3310.1145/333113

A lot of data is moved from system to system in an important and increasing part of the computing landscape. This is traditionally known as ETL (extract, transform, and load). While many systems are extremely good at this process, the source for the extraction and the destination for the load frequently have different representations for their data. It is common for this transformation to squeeze, truncate, or pad the data to make it fit into the target. This is really like using a shoehorn to fit into a shoe that is too small. Sometimes it's a needed step. Frequently it's a real pain!

Two major parts of ETL are the extraction and the load. These processes are where the rubber meets the participating data stores.

Extraction pulls data out of a source system. This may be relational data kept in a database. If so, it may be converted to an object relational format where each object transforms the join of multiple relational rows into a cohesive thing. Data is frequently organized as messages when it is sucked out. It's also common for data to be extracted from key-value stores where it is kept in a semi-structured representation.

Load happens when the data is placed into the target system. The target will have its own metadata describing the shape and form of the data in its belly. If the target is an analytics system, then its data will likely be loaded into a relational form.

While it may be counterintuitive, it is frequently useful to take relational data out of a system as objects; convert, massage, and shoehorn the data from one object representation to another; and load it into the target system in relational form.  .... "