/* ---- Google Analytics Code Below */
Showing posts with label Semantic. Show all posts
Showing posts with label Semantic. Show all posts

Friday, March 11, 2022

Webinar: How Metadata Management Must Evolve to Support Data Fabric

Looks to be useful, plan to attend:

TopQuadrant: How Metadata Management Must Evolve to Support Data Fabric   by Irene Polikoff | Mar 8, 2022 | Metadata Management, Webinars

About This Webinar

On Thu, Mar 17, 2022 11:30 AM EDT

If you have not heard the term “data fabric” yet, you will. It is rapidly growing in popularity. Gartner identified data fabric as the top trend for data and analytics in 2021.

You can think of data fabric as a web connecting multiple locations, types, and sources of data – both on-premises and in the public cloud. It is an architectural approach designed to help organizations better deal with the growing number of available data sources and ever-changing application requirements. The backbone of data fabric design is a Knowledge Graph capturing information about data sources in RDF. This is a new type of a data catalog with semantically augmented and enriched metadata.

Join us for this webinar to:

Learn What Is Data Fabric

Understand Why Data Fabric Requires a Knowledge Graph

Get Advice on Moving from Traditional Metadata Management to Metadata Management With Knowledge Graphs

Envision How Tools Participating in Data Fabrics Will Interact With Knowledge Graphs

See these Concepts in Action  ... 

 Register:   https://register.gotowebinar.com/register/7807142108219203083 

Tuesday, June 08, 2021

Big Players Want to Play Semantics for Meaning

 This precisely what we wanted to do, to let our systems understand our meaning of words, rather than just what the meaning they have been assigned by language.  But would that work?  We wanted to assign the meaning that WE had assigned for that assemblage of letters.      Assume this could be one by some preprocessing or postprocessing or meta assignment.  Nevertheless this is still a big step forward  to empower practical use of meaning.    But where is AWS?

Semantics Beats Syntax  By R. Colin Johnson who is a Kyoto Prize Fellow who has worked as a technology journalist for two decades.    In ACM.

IBM, Google, and Microsoft are all poised to release semantic engines (algorithms using the meaning of words) to supplement their current syntax engines (using the spelling of words, such as the search engine BM25). Their common goal is to extend their natural language processing (NLP) capabilities into engines that rival human semantics (our understanding of what language, words/sentences, mean). In line with other contenders, including Amazon, Intel, and Oracle, these semantic engines offer machine-understanding of meaning aimed at enhancing searches, artificial intelligence (AI), human–computer interaction, question answering, and automatic narrative generation (from descriptions/explanations to prose/poetry).

Today's syntax-only engines are blind to the meaning of the keywords used to ascertain results. For example, a human understands that "where Alan Turing was born" means the same thing as "the birthplace of Alan Turing," and "the town where Alan Turing was delivered as a baby." Their syntax is different, having unique keywords "where, born"; "birthplace"; and "town, delivered," respectively. However, each phrase's meaning, or semantics, are identical ("London" is the answer to all three). People understand this immediately, but computers—not so much.

Consequently, all three companies are developing algorithms that understand the meaning of words. Google and Microsoft are both building semantic engines that add metadata to sentences (Google) or words (Microsoft) using clusters of processors running multiple deep neural networks (called transformers, which use massive parallelization).

IBM

For a decade, IBM has been extending NLP to artificial intelligences (AIs)—starting with its Watson (2011), the AI that beat human experts in the television game Jeopardy, and most recently with its Project Debater (2021). During this decade (2014), IBM also developed its Cognos transformer—a metadata-based algorithm for creating engines it calls PowerCubes, which can then be used with its Business Intelligence software. But for its semantic engine, IBM chose instead to augment neural networks with symbolic logic that cuts down the number of examples it requires to learn. Transformers, it discovered when developing Cognos, require larger example sets to learn, according to Forrester Research principal analyst Kjell Carlsson.

"IBM's semantics uses a much more efficient encoding of knowledge, enabling high performing enterprise use-cases to be built with significantly smaller training examples," said Carlsson. "IBM's semantics can also provide explainability of  its conclusions by virtue of symbolic-logic reasoning, governance to block the model from using faulty logic, plus provides fairness that prevents models from learning discriminatory reasoning," said Carlsson.

Transformers, on the other hand, are all neural networks—end-to-end—without meaning instilled by symbolic logic, which also enables easier explanations for why a neural network comes to specific conclusions.

"IBM's neuro-symbolic approach," said Carlsson, "enables higher accuracy with less training data, plus it also enables engineers to 'teach' a model logical-relationships that domain experts know to be true, which is far more efficient than having these relationships be learned by transformers."

In the words of Salim Roukos, an IBM Fellow and the company's Global Leader for Language Research, as well as Chief Technology Officer of the company's Translation Technologies unit, "Large end-to-end neural models require significant amounts of data to perform well in a new domain. IBM is more focused on semantic parsing of human language to enable developers to build text-understanding applications. By leveraging the semantics of human language, very small amounts of data from the application domain are needed to enable understanding."

Saturday, May 23, 2020

Microsoft Project Cortex

Brought to my attention from this weeks Build meetings as about to be launched.   Form of knowledge management.  Had been show very early version of this.  Notion of a semantic Web has been around for a long time, though not used often enough.   Will be following this.

Project Cortex

Today, we’re pleased to introduce Project Cortex, the first new service in Microsoft 365 since the launch of Microsoft Teams. Project Cortex uses advanced AI to deliver insights and expertise in the apps you use every day, to harness collective knowledge and to empower people and teams to learn, upskill and innovate faster.

Project Cortex uses AI to reason over content across teams and systems, recognizing content types, extracting important information, and automatically organizing content into shared topics like projects, products, processes and customers. Cortex then creates a knowledge network based on relationships among topics, content, and people.

New topic pages and knowledge centers—created and updated by AI—enable experts to curate and share knowledge with wiki-like simplicity. And topic cards deliver knowledge just-in-time to people in Outlook, Microsoft Teams, and Office.  ... " 

Saturday, April 25, 2020

Knowledge Graphs versus Property Graphs

This came in the mail, it had been asked in an interaction recently.  Worth a look.  We worked with TopQuadrant.

New White Paper: Knowledge Graphs versus Property Graphs

We are in the era of graphs. Graphs are hot. Why? Flexibility is one strong driver: heterogeneous data, integrating new data sources, and analytics all require flexibility. Graphs deliver it in spades.

The two main graph data models are: Property Graphs and Knowledge (RDF) Graphs. People who want to take advantage of graph-based solutions for data and metadata management want to know what they are, what are their similarities and differences, and what they are each good for.

This white paper covers the following, it:
Describes the two main graph data models: Property Graphs and RDF Graphs and explains the key differences in their terminology and capabilities
Compares their strengths and limitations
Provides guidance on their respective capabilities

Other TopQuadrant resources to explore:
RECORDED WEBINARS, including this most recent one: "Getting Started with Data Governance"
WHITE PAPER COLLECTION, including: "Implementing Data Governance with Knowledge Graphs to Enable Enterprise AI"

Download Now    https://www.topquadrant.com/knowledge-assets/whitepapers/

This email sent by TopQuadrant   www.topquadrant.com

Saturday, September 21, 2019

Structured Signals for Model Training

Technical but interesting point about how to add structured knowledge into otherwise non transparent networks.  Examining further.

Posted by Da-Cheng Juan (Senior Software Engineer) and Sujith Ravi (Senior Staff Research Scientist)

We are excited to introduce  Neural Structured Learning in TensorFlow, an easy-to-use framework that both novice and advanced developers can use for training neural networks with structured signals. Neural Structured Learning (NSL) can be applied to construct accurate and robust models for vision, language understanding, and prediction in general.

Neutral structured learning framework

Many machine learning tasks benefit from using structured data which contains rich relational information among the samples. For example, modeling citation networks, Knowledge Graph inference and reasoning on linguistic structure of sentences, and learning molecular fingerprints all require a model to learn from structured inputs, as opposed to just individual samples. These structures can be explicitly given (e.g., as a graph), or implicitly inferred (e.g., as an adversarial example). Leveraging structured signals during training allows developers to achieve higher model accuracy, particularly when the amount of labeled data is relatively small. Training with structured signals also leads to more robust models. These techniques have been widely used in Google for improving model performance, such as learning image semantic embedding.

Neural Structured Learning (NSL) is an open source framework for training deep neural networks with structured signals. It implements Neural Graph Learning, which enables developers to train neural networks using graphs. The graphs can come from multiple sources such as Knowledge graphs, medical records, genomic data or multimodal relations (e.g., image-text pairs). NSL also generalizes to Adversarial Learning where the structure between input examples is dynamically constructed using adversarial perturbation.  ... " 

See also:  https://www.datanami.com/2019/09/04/google-adds-structured-signals-to-model-training/

See also:  https://venturebeat.com/2019/09/03/google-launches-tensorflow-machine-learning-framework-for-graphical-data/ 

Thursday, August 22, 2019

Shape of Data Implies its Uses

Recall hearing about this approach in training we took.    Making it easier to connect key data to analytics.   Shape constraint languages.

Graph Database ‘Shapes’ Data   By George Leopold in Datanami

A semantic graph database technology vendor is supporting a key specification designed to validate graph-based data against a set of conditions that specify the “shape” of data. The goal is a more agile way of analyzing larger volumes of complex, distributed data.

Franz Inc. said this week its flagship AllegoGraph platform now supports SHACL, the SHApe Constraint Language used to describe “the shape that data should have.” The spec was developed by the World Wide Web Consortium.

The language using a concept known as “triples” to describe data properties. Triples refers to the subject (the thing being described), predicate (properties or relationships of the subject) and object (an intrinsic value such as an integer or text).

SHACL also specifies the number of triples required in a repository along with metadata about the object, including a specified subject and predicate.

Oakland-based Franz said support for SHACL would improve the ability of its graph database to validate data and applications from outside sources while improving interoperability. “Adding SHACL to AllegroGraph helps our customers simplify the complexity of enterprise systems through the ability to loosely combine independent elements, while allowing the overall system to function smoothly,” said CEO Jans Aasman.

The 6.6 version of AllegroGraph includes a data-shaping validation engine used to confirm whether data conforms with desired requirements. SHACL allows a data graph, for instance, to specify the corresponding shapes graph used to describe the link between a given shape and targeted data.  .... "

Friday, August 09, 2019

Knowledge Graphs in the Enterprise

How big enterprises are using Knowledge Graphs.  Very good to understand this if you are planning to use knowledge graphs of your own.    Instructive.   in ACMQueue, CACM    Here a short starting excerpt,  full text at the link:  

Industry-scale Knowledge Graphs: Lessons and Challenges
Five diverse technology companies show how it's done
Natasha Noy, Google; Yuqing Gao, Microsoft; Anshu Jain, IBM Watson; Anant Narayanan, Facebook; Alan Patterson, eBay; Jamie Taylor, Google

Knowledge graphs are critical to many enterprises today: They provide the structured data and factual knowledge that drive many products and make them more intelligent and "magical."

In general, a knowledge graph describes objects of interest and connections between them. For example, a knowledge graph may have nodes for a movie, the actors in this movie, the director, and so on. Each node may have properties such as an actor's name and age. There may be nodes for multiple movies involving a particular actor. The user can then traverse the knowledge graph to collect information on all the movies in which the actor appeared or, if applicable, directed.

Many practical implementations impose constraints on the links in knowledge graphs by defining a schema or ontology. For example, a link from a movie to its director must connect an object of type Movie to an object of type Person. In some cases the links themselves might have their own properties: a link connecting an actor and a movie might have the name of the specific role the actor played. Similarly, a link connecting a politician with a specific role in government might have the time period during which the politician held that role.

Knowledge graphs and similar structures usually provide a shared substrate of knowledge within an organization, allowing different products and applications to use similar vocabulary and to reuse definitions and descriptions that others create. Furthermore, they usually provide a compact formal representation that developers can use to infer new facts and build up the knowledge—for example, using the graph connecting movies and actors to find out which actors frequently appear in movies together.

This article looks at the knowledge graphs of five diverse tech companies, comparing the similarities and differences in their respective experiences of building and using the graphs, and discussing the challenges that all knowledge-driven enterprises face today. The collection of knowledge graphs discussed here covers the breadth of applications, from search, to product descriptions, to social networks:

• Both Microsoft's Bing knowledge graph and the Google Knowledge Graph support search and answering questions in search and during conversations. Starting with the descriptions and connections of people, places, things, and organizations, these graphs include general knowledge about the world.    .... "

• Facebook has the world's largest social graph, which also includes information about music, movies, celebrities, and places that Facebook users care about.

• The Product Knowledge Graph at eBay, currently under development, will encode semantic knowledge about products, entities, and the relationships between them and the external world.

• The Knowledge Graph Framework for IBM's Watson Discovery offerings addresses two requirements: one focusing on the use case of discovering nonobvious information, the other on offering a "Build your own knowledge graph" framework.

The goal here is not to describe these knowledge graphs exhaustively, but rather to use the authors' practical experiences in building knowledge graphs in some of the largest technology companies today as a scaffolding to highlight the challenges that any enterprise-scale knowledge graph will face and where some innovative research is needed. ....'

Thursday, August 01, 2019

The Value of Lineage Metadata

Did lots of work with metadata, especially in the corporate laboratory space.   Had never heard the term 'Lineage Metadata', though again it was often considered in our work.   Now would think this is more important than ever, if we are to create useful predictions, and also add some accuracy to future maintenance of any automated systems that emerge.  Lineage predicts changing context.   At very least the lineage can determine what errors might exist in the data, but in reality should provide much more in AI.  Should always be considered.     Lineage can also be considered as part of data asset value, predicting stability of value in context.    Also be included in a semantic representation of data involved.
I like that this is presented here.  Below an excerpt, much more at the link.- FAD

Lineage Metadata: The Fuel for Data Governance in Informationweek
Moshe Kranc is the chief technology officer at Ness Digital Engineering

The best way to achieve data quality is by combining or blending these three techniques: decoded lineage, data similarity lineage and manual lineage mapping.

Enterprises aspire to derive insights from data that can provide a competitive advantage. The most common impediment to achieving this goal is poor data quality. If the data that is being input to a predictive algorithm is “dirty” (with missing or invalid values), then any insights produced by that algorithm cannot be trusted.

To achieve data quality, it’s not enough to clean up the existing historical data. You also need to ensure that all newly generated data is clean by instituting a set of capabilities and processes known collectively as data governance. In a governed data environment, each type of data has a data steward who is responsible for defining and enforcing criteria for data cleanliness. And, each data value has a clearly defined lineage: We know where it came from, what transformations it underwent along the way, and what other data items are derived from this data value.

Data lineage provides an enterprise with many benefits:

The ability to perform impact analysis and root-cause analysis, by tracing lineage backwards (to find all data that influenced the current data) or forwards (to identify all other data that is impacted by the current data) from a given data item;
Standardization of the business vocabulary and terminology, which facilitates clear communication across business units;

Ownership, responsibility and traceability for any changes made to data, thanks to the lineage’s comprehensive record of who made what changes and when.

It sounds great, but where does data lineage information come from? Looking at a specific data value in the database tells us its current value, but it will not provide information about how the data evolved into its current value. What is missing is data about the data (lineage metadata) that automatically remembers the time and source of every change made to every data item, whether the change was made by software or by a human database administrator.

There are three competing techniques for collecting lineage metadata, each of which has its strengths and weaknesses: .... " 

Wednesday, July 10, 2019

Explainable AI from Kyndi is Funded

Quite a claim, like the components I see mentioned.  More explanation, logic, graphs, semantics ....  we explored and used them all,  is it enough for required transparency?  And how will such AI and explanation be maintained in changing contexts?   And resulting risks?

Kyndi Platform funding:

....  EXPLAINABLE ARTIFICIAL INTELLIGENCE ...

Kyndi is building the first Explainable AI platform for government, financial services, and life sciences. Our solutions are the antithesis of the old AI “black box.”
KYNDI ADDS $20M TO EXPAND TEAM AND ACCELERATE GROWTH
The Series B round was led by Intel Capital, with participation from UL Ventures, PivotNorth Capital, and existing investors.
THE KYNDI AI PLATFORM
Kyndi is an artificial intelligence company that’s building the first Explainable AI product and Intelligent Process Automation software for government, pharmaceutical, and financial services organizations. .... 

The Kyndi AI Platform uses machine learning to streamline regulated business processes and offer auditable AI systems for enterprises and government. Kyndi’s product exists because Deep Learning is a “black box” and cannot be used in regulated industries where organizations are required to explain the reasons for any decision.

Our platform uses a novel approach to AI, unifying probabilistic and logical methods. This enables organizations to analyze massive amounts of data to create actionable knowledge significantly faster and without having to sacrifice explainability. Kyndi’s Explainable AI™ Platform supports the following solutions: Intelligence, Defense, Compliance (i.e., for financial services and healthcare), and Research. Crucially, the Kyndi AI Platform also helps to mitigate the human bias that can arise in the process of extracting knowledge and answers from data.

EXPLAINABILITY
Explainable artificial intelligence achieves the level of trust that is so important for accelerated growth and acceptance of this revolutionary technology.
AI cannot be a “black box,” as it so often is today. Explainable AI™ means that our software’s reasoning is apparent to the user, and that the system can explain its rationale. This visibility allows you to have confidence in the system’s outputs, be aware of any uncertainties, anticipate how the software will work in the future, and know how to improve the system. Such knowledge is essential to confident analysis and decision making. Explainability is at the core of Kyndi’s products and solutions.

NATURAL LANGUAGE PROCESSING (NLP)
With Kyndi’s AI products, knowledge is accumulated and transferred through written forms of natural language.

Kyndi has developed a unique and effective approach to NLP to automate and scale knowledge consumption. First, our NLP solution tokenizes text and identifies parts of speech and sentence structure. Next, we identify named entities with real-world references and compute semantic distances between words using our proprietary Semantic Distance Field Model to show us how strongly any two entities are related. Semantic parsing and Relation Extraction allows us to formally name this relationship where both the entities and relationships create a proto-ontology that encodes and condenses the meaning of a document, a collection of documents, or a whole domain.

With this clear view of ideas, concepts, and relationships, we build a knowledge graph that scales with your business and that lets you query your collection and retrieve the information you want, not just the words you used to ask the question.

KNOWLEDGE GRAPHS
Once Kyndi’s NLP pipeline identifies the structure and contents of your data we create a graph representation, no matter the size.

With each node in the graph identifying an entity, the connections between the nodes are semantic vectors that signify the relationships and significance between entities. The feature-rich graph enables you to conduct a quick and accurate analysis by matching fragments and returning details relevant to your search queries. To achieve this functionality Kyndi has developed industry-leading technology that allows for sub-graph matching based on cognitive signatures.  ..... "

Sunday, July 07, 2019

Innovations in Graph Representation Learning

From the Google AI Blog.  Good introduction to semantic networks.  And some technical information about where they are experimenting.   To model something complex, you first need to have a basic representation, a graph is a good place to start.   In particular because it can be used to explain the complexity to non technicals.

Innovations in Graph Representation Learning
Tuesday, June 25, 2019
Posted by Alessandro Epasto, Senior Research Scientist and Bryan Perozzi, Senior Research Scientist, Graph Mining Team

Relational data representing relationships between entities is ubiquitous on the Web (e.g., online social networks) and in the physical world (e.g., in protein interaction networks). Such data can be represented as a graph with nodes (e.g., users, proteins), and edges connecting them (e.g., friendship relations, protein interactions). Given the widespread prevalence of graphs, graph analysis plays a fundamental role in machine learning, with applications in clustering, link prediction, privacy, and others. To apply machine learning methods to graphs (e.g., predicting new friendships, or discovering unknown protein interactions) one needs to learn a representation of the graph that is amenable to be used in ML algorithms.

However, graphs are inherently combinatorial structures made of discrete parts like nodes and edges, while many common ML methods, like neural networks, favor continuous structures, in particular vector representations. Vector representations are particularly important in neural networks, as they can be directly used as input layers. To get around the difficulties in using discrete graph representations in ML, graph embedding methods learn a continuous vector space for the graph, assigning each node (and/or edge) in the graph to a specific position in a vector space. A popular approach in this area is that of random-walk-based representation learning, as introduced in DeepWalk.   .... " (more useful content follows at the link)

Saturday, April 13, 2019

Natural Language and Intent with Ambiguity

Good simple explanation from Tableau on Intent.   Have used the concept now in several projects, and of course there is ambiguity, beyond dictionary-definition,  in the use of many terms within a company.   The ambiguity in context is important to consider.

Machine learning, natural language meet to understand intent
 By Mark Jewett, VP of Marketing, Tableau

Machine learning and natural language processing promise to better translate human curiosity into pertinent answers. If true, these smart capabilities will broaden the use of analytics and reach people who are less comfortable dealing with data. It will all start with helping machines learn to interpret human intent. The key is semantics.

Sometimes intent is simple and explicit, like asking Siri or Alexa if a flight is delayed. This question has clear intention and a simple response—returning the flight status answers the question. Such simplicity is seldom the case when it comes to data analysis. Questions are usually more nuanced, making it hard to correctly assume what the user is really looking for. Natural language is even more tricky where ambiguous terms are common.

It’s also difficult for a machine to understand our intent within a limited context. The machine has the data itself but doesn’t grasp the bigger picture in the same way a person with domain expertise can. Asking “How are my sales doing in the Northeast?” is a lot more ambiguous than the flight status example above.

Ambiguity isn’t a new challenge in data analysis. Different groups within an organization may have different definitions or calculations for the same words: for example, the term “profitability”. Some organizations use central dictionaries (also called data catalogs) to reduce ambiguity and create consistency across the organization. These tools can help provide users with the context they need to understand more deeply. .... " 

Thursday, February 21, 2019

Knowledge Graphs, Governance, Learning and Much More

Have now been involved in a number of efforts in this area, worth understanding since Google has made an impressive run at this.  The article tells a historical journey I have traveled as well, we might finally be getting to real enterprise value.  And regulations like GDPR are forcing us to take notice of the need to really understand our data. The article is long, but has good points to make.

The Semantic Zoo - Smart Data Hubs, Knowledge Graphs and Data Catalogs   By Kurt Cagle Contributor in Forbes

COGNITIVE WORLDContributor Group
Sometimes, you can enter into a technology too early. The groundwork for semantics was laid down in the late 1990s and early 2000s, with Tim Berners-Lee’s stellar Semantic Web article, debuting in Scientific American in 2004, seen by many as the movement’s birth. Yet many early participants in the field of semantics discovered a harsh reality: computer systems were too slow to handle the intense indexing requirements the technology needed, the original specifications and APIs failed to handle important edge cases, and, perhaps most importantly, the number of real world use cases where semantics made sense were simply not at a large enough scope; they could easily be met by existing approaches and technology.


Semantics faded around 2008, echoing the pattern of the Artificial Intelligence Winter of the 1970s. JSON was all the rage, then mobile apps, big data came on the scene even as Javascript underwent a radical transformation, and all of a sudden everyone wanted to be a data scientist (until they discovered the fact that data science was mostly math). Meanwhile, from the dim recesses of the troughs of despair, semantics was readying itself for its own metamorphosis. Several semantic standards, including the SPARQL query language along with a new update language began seeing implementations by 2015.  Servers became faster and cheaper, and a rise of graphics processor units (GPUs) fueled by the gaming and entertainment industry provided tools for a new class of graph databases.

Meanwhile, the Big Data initiatives that had marked the early part of the 2010s was facing some real problems. The original promise of Hadoop as a map / reduce framework had ended up creating large numbers of data lakes that aggregated content but that sat under-utilized. Data scientists struggled to deal with dirty data that was really no cleaner for having been put in data lakes. JSON databases had grown in popularity, but they were proving hard to query in a consistent fashion, and all too many Hadoop projects ended up becoming large, slow, but cheap data graveyards for regulatory data (the kind of data that must be retained for five years). ..... "

Friday, February 15, 2019

Knowledge Graphs and Governance Webinar

Will be attending.   Ultimately a crucial next step for AI and integration of Semantic data with enterprise data and metadata.

From TopQuadrant:

Join us for a webinar on February 21, 2019 at 11:30 AM ET.

Why are knowledge graphs relevant to data governance? The answer lies in key characteristics. Knowledge graphs are:

Flexible – graphs are the most flexible formal data structures
Evolvable – able to accommodate diverse data and metadata
Semantic – the meaning of the data is stored alongside the data in the graph
Intelligent – semantics of data are explicit and enable data validation and drawing conclusions and new information from the available data.

These qualities make knowledge graphs an ideal and, arguably, the only viable foundation for bridging and connecting enterprise metadata silos – the main goal of data governance.

Join us for this webinar where we will:
Provide a brief history of Knowledge Graphs
Give examples of what they are good for
Show how they provide a powerful platform for integrated data management and governance ... " 

Wednesday, February 06, 2019

UpComing Webinar on Knowledge Graphs

Its important to link knowledge with other forms of semantic databases.   As we use these methods they become easier to link and maintain to AI style applications.   We used TopQuadrant, I will attend.

WEBINAR: What are Knowledge Graphs? Why Are They Key to Successful Data Governance?

Join us for a webinar on February 21, 2019 at 11:30 AM ET

Robert Coyne of TopQuadrant writes:

Why are knowledge graphs relevant to data governance? The answer lies in key characteristics. Knowledge graphs are:

Flexible – graphs are the most flexible formal data structures
Evolvable – able to accommodate diverse data and metadata
Semantic – the meaning of the data is stored alongside the data in the graph
Intelligent – semantics of data are explicit and enable data validation and drawing conclusions and new information from the available data.
These qualities make knowledge graphs an ideal and, arguably, the only viable foundation for bridging and connecting enterprise metadata silos – the main goal of data governance.

Join us for this webinar where we will:

Provide a brief history of Knowledge Graphs
Give examples of what they are good for
Show how they provide a powerful platform for integrated data management and governance 

More info and registration: https://www.topquadrant.com/knowledge-graphs-webinar/ 

Monday, August 13, 2018

Modeling User Journeys

Have had a few explorations into 'User Journey's.  Its a trace of how people travel through coded interaction. From where they begin, what choices they make and where they end up. Ideally measuring how much the journey results in value.  Its a kind of business process model based on a path of interactions.  It can be used to plan, construct and even optimize user interactions.  Here an example by the marketplace Etsy.  with considerable detail regarding integration of machine learning.

Modeling User Journeys via Semantic Embeddings  Posted by Nishan Subedi in O'Reilly

Etsy is a global marketplace for unique goods. This means that as soon as an item becomes popular, it runs the risk of selling out. Machine learning solutions that simply memorize the popular items are not as effective, and crafting features that generalize well across items in our inventory is important. In addition, some content features such as titles are sometimes not as informative for us since these are seller provided, and can be noisy.

In this blog post, I will cover a machine learning technique we are using at Etsy that allows us to extract meaning from our data without the use of content features like titles, modeling only the user journeys across the site. This post assumes understanding of machine learning concepts,  specifically word2vec. ... " 

A good Tensorflow Tutorial on Word2vec.

Friday, July 13, 2018

Machine Learning Data Catalogs

Makes sense, also connecting the data to its actual meaning,  semantic ontologies,  is also good to do in the same place.      It should be a broader aspect of governance.   It is also an important fundamental aspect of interpretability to know where the data is coming from, what its stability and credibility are.

How the Machine Learning Catalogs Stack Up  
Alex Woodie in Datanami

You can’t do anything with data – let alone use it for machine learning – if you don’t know where it is. In the age of big data, this is not a trivial matter. It is also the main driver that’s propelling the rise of machine learning data catalogs, which the analysts at Forrester recently ranked and sorted. Just a word of warning: the name at the top of the list might surprise you.

According to Michelle Goetz’s June 21 Forrester Wave report, the percentage of analytic decision makers managing more than 1 petabyte of data (either structured, semi-structured, or unstructured) has essentially tripled from 2016 to 2017. That rapid growth has exposed all manner of problems in company’s existing data management and analytic endeavors.

Two of the biggest challenges that companies face today, Goetz writes, are gathering and managing data in a governed manner on the one hand, and managing the business processes that surround the data analytics activities on the other.

“For EA [enterprise analytics] professionals, relying on people and manual processes to provision, manage, and govern data simply does not scale,” the Forrester analyst writes. “Enterprises are waking up to this fact and turning to data catalogs to democratize access to data, enable tribal data knowledge to curate information, apply data policies, and activate all data for business value quickly.”  .... " 

Monday, May 21, 2018

Microsoft Buys Semantic Machines

Towards more conversational machines.  We spent many years trying to figure out how analytics, systems and machines could better 'understand' the meaning of data.  Now this will be essential to lead to better conversational interaction.  Note the term 'multiturn' exchanges.  Ultimately its all about the intelligent conversation.

Microsoft snaps up Semantic Machines to build out its conversational AI technology  By Duncan Riley in SiliconAngle
  
Microsoft Corp. Sunday said it has acquired Semantic Machines Inc., a Berkeley, California-based company that has built a conversational artificial intelligence platform that competes with the likes of Google Inc., for an undisclosed sum.

Founded in 2014, Semantic Machines has designed a new, language-independent technology platform that claims to go beyond understanding commands to understanding conversations. Compared with a neurolinguistic programming approach, the company said, it offers a new technology that extracts semantics across “multiturn” natural language exchanges to maintain contextual understanding over time, enabling computers to communicate, collaborate, understand goals and accomplish tasks.

The acquisition for Microsoft is aimed at boosting its existing conversational AI efforts in services such as Microsoft Cognitive, Cortana and the Azure Bot. The technology and the company itself will be used by Microsoft to establish a conversational AI center of excellence in Berkeley “to push forward the boundaries of what is possible in language interfaces.”  ... " 

Wednesday, March 14, 2018

The Semantics of Image Deep Learning

Google once again shows its impressive advanced AI/Deep Learning capabilities.     Which made me recall that it is often the 'semantic', or meaning in context aspects that are most important for an AI or analytic method to be useful.   And that assigning tags also implies we will need to maintain the tags as context changes.  Below is technical, look at the link for some image examples that make this clearer.

Semantic Image Segmentation with DeepLab in Tensorflow

Posted by Liang-Chieh Chen and Yukun Zhu, Software Engineers, Google Research

Semantic image segmentation, the task of assigning a semantic label, such as “road”, “sky”, “person”, “dog”, to every pixel in an image enables numerous new applications, such as the synthetic shallow depth-of-field effect shipped in the portrait mode of the Pixel 2 and Pixel 2 XL smartphones and mobile real-time video segmentation. Assigning these semantic labels requires pinpointing the outline of objects, and thus imposes much stricter localization accuracy requirements than other visual entity recognition tasks such as image-level classification or bounding box-level detection. ... "

Wednesday, May 10, 2017

Google, Linked Data and Ontologies

Very good,  have spent a long time thinking about getting the semantics of data arranged for search and analytics.

Via Dean Allemang
Principal Consultant at Working Ontologist, LLC
One of the best short articles on this subject I have seen. Advice about how we should organize ontologies. 

Google, Linked Data and The Needle In The Haystack
 By Jan Voskuil   CEO at Taxonic

Modern search engines rely heavily on structured metadata for high precision. Unbeknownst to many, Linked Data plays a quintessential role in this. In fact, many people seem to think that modern search engine technology has obviated the need for assigning keywords to content items, such as journal articles, either manually or through an automated process. The opposite is true. I will illustrate this by showing how Google finds needles in a haystack, using a simple do-it-yourself experiment to make a complex point.  .....  " 

Sunday, April 30, 2017

Promoting Mind Mapping

We were very active users of 'mind mapping',  a much simplified form of  'concept mapping'.  in the enterprise  So I much like to promote the idea, its always useful.  Good to see this article and expose others to it.  There are many packages that do this, many are free.  Spodek uses Freeplane, which I have not heard of, but sounds interesting.  Its a way to get organized, and communicate that organization to others.   If you have the slightest interest in organizing your thoughts and work, give it a try.   In Inc: 

The Most Useful (and Fun) Software You Don't Have
If you don't use mindmapping software, you don't know the fun, efficient productivity you're missing     By Joshua Spodek

Mindmaps and mindmapping software are awesome!

Yesterday, I finished creating two new keynote talks, each over 60 slides. Going from idea to complete deck used to mean complicated struggling between paper, blackboard, whiteboard, word processor, and presentation software.

Now I create one mindmap using one piece of software to create the whole presentation. It's faster, easier, and more fun. I'm writing this post because of how fun and simple it was to be so productive.

I rarely like using computers for what I can use paper and pencil, but mindmaps and mindmapping software help organize complex ideas better by every measure I care about. They're simple, effective, and, best of all, fun. ... "