/* ---- Google Analytics Code Below */
Showing posts with label Documents. Show all posts
Showing posts with label Documents. Show all posts

Wednesday, May 13, 2020

Value and Methods of Text Summarization

Ultimately a powerful concept.    We examined its use with group meetings and focus groups, to gather important topic information, a kind of more focused crowdsourcing of information.   Often aimed at specific goals.  Also was sometime used with documents supporting the topics.   Indeed a 'holy grail' if it can be done systematically and well.  Below from KDNuggets:

This article will present the main approaches to text summarization currently employed, as well as discuss some of their characteristics.

The bona fide semantic understanding of human language text, exhibited by its effective summarization, may well be the holy grail of natural language processing (NLP). That statement isn't as hyperbolic as it sounds: as true human language understanding definitely is the holy grail of NLP, and genuine effective summarization of said human language would necessarily entail true understanding, transitivity would back me up on this.

Unfortunately — or perhaps not, depending on your outlook — honest to goodness "understanding" of human language is not something we can currently count on for text summarization. However, the show must go on, and there currently exist an array of actual techniques for summarizing text, some of which stretch back decades. These techniques take different approaches to reaching the same goal, and can be classified into a fairly narrow set of categories for pursuing their shared goal.

This article will present the main approaches to text summarization currently employed, as well as discuss some of their characteristics: ... '

Friday, December 08, 2017

Document Classification with Deep Learning

Once again, a good introduction.  And demonstration of how these systems can be constructed.

Best Practices for Document Classification with Deep Learning   by Jason Brownlee

Text classification describes a general class of problems such as predicting the sentiment of tweets and movie reviews, as well as classifying email as spam or not.

Deep learning methods are proving very good at text classification, achieving state-of-the-art results on a suite of standard academic benchmark problems.

In this post, you will discover some best practices to consider when developing deep learning models for text classification.

After reading this post, you will know:

The general combination of deep learning methods to consider when starting your text classification problems.

The first architecture to try with specific advice on how to configure hyperparameters.

That deeper networks may be the future of the field in terms of flexibility and capability.

Let’s get started. .... ".

Wednesday, September 27, 2017

Reasonet Ingests and Answers Questions

Came up to me in several directions.   ReasoNet.  A machine reading program that allows machines to ingest documents, read them, and then answer questions based on the documents.

" ... Teaching a computer to read and answer general questions pertaining to a document is a challenging yet unsolved problem. In this paper, we describe a novel neural network architecture called the Reasoning Network (ReasoNet) for machine comprehension tasks. ReasoNets make use of multiple turns to effectively exploit and then reason over the relation among queries, documents, and answers ... " 

Reasonet

https://arxiv.org/abs/1609.05284

Reasonet
https://github.com/ymcui/Eval-on-NN-of-RC     Github example

Thursday, April 13, 2017

Semantic Document Processing

Today's call was of interest, and showed how complex document understanding can be:

" ... Today ... Our speaker was Sridhar Iyengar, IBM Distinguished Engineer at the IBM T. J. Watson Research Center, who will be presenting "Semantic PDF Processing & Document Representation."  ....  Slides here.   and Recording here

Please find the schedule of presenters herefor the next several calls.   A link to slides and a recording of each call should be available on the CSIG website (http://cognitive-science.info/community/weekly-update/).   We encourage those who join the calls to add questions and comments to the LinkedIn Discussion Group https://www.linkedin.com/groups/Cognitive-Systems-Institute-6729452  .... 

Thank you!
Dianne Fodell

Saturday, April 08, 2017

Document Review Inquiry

In the Gartner Blog, they mention a process called a Document Review Inquiry .    Which like the author , I had never hear of.  Worth a look.  A means towards curation for Gartner, apparently.  The same approach could lead you towards a way to automate the storage and retrieval of accumulated knowledge.   And then add an advisor to help with the curation.

Thursday, March 16, 2017

AI Driven Automation in Law

In some ways, the most obvious application.   A heavily document and analysis driven world.   Was frequently mentioned in the 80s, but not implemented because the analysis of meaning in documents was very imperfect.

AI automation starts to transform legal profession

Artificial intelligence and automation are making inroads into legal work, but are more likely to support than displace lawyers and will draw IT more into legal service delivery

In February 2016, a London court supported the use of predictive coding software in a legal disclosure process, which often involves lawyers receiving huge volumes of documents from those representing the other side in a case.  ... 

In Pyrrho Investments v MWB Business Exchange, Master Paul Matthews of the Chancery division supported the use of software in scoring documents for relevance. He found there was no evidence that software would be less accurate than manual review and keyword searches. He added that software could provide greater consistency in searching more than 3 million documents that could be involved in the disclosure. A final reason was that both sides had agreed to the use of the software, which would be much cheaper than a manual search – they just wanted the court’s approval. ...  " 

Tuesday, December 20, 2016

Automatic Document Search

Back to this approach, examined for years, obviously useful in legal, but in most any other kinds of research as well.  Now being integrated into word processors and note taking systems.  In Fastcompany. 

Friday, April 24, 2015

MindMeld and Documents

Voice Drive content discovery from MindMeld.   " .... You can index documents by crawling a website or by using our API -- and we’ve just open-sourced a new way to post documents to the MindMeld API! Using our new Node.js poster script, you can programmatically upload a very large collection of documents for your app (we're talking millions of nodes). Once you add your document collection, we build a custom knowledge graph that enables highly accurate natural language understanding for your application domain. ... " 

Wednesday, April 08, 2015

P&G Selling More Beauty Brands

Quite a big sell off.  In BizJournals:  " ... Cincinnati-based P&G (NYSE: PG) sent sale documents to potential bidders for its Wella haircare unit along with cosmetics brands and its fragrance business, according to a Bloomberg report. People familiar with the offerings told Bloomberg the brand sales could bring in a total of as much as $19 billion. P&G bought Wella in 2003 for $7 billion. ... " 

Friday, March 27, 2015

Extracting Data from PDF Documents

Could have used this idea a number of years ago. In projects meant to gather and archive enterprise knowledge.  Data tables are often embedded in PDF documents, and extracting these systematically, in volume, sometimes ends up as a manual task with potential for error.  In CWorld:  Tabula, a free open source tool to do this.    Have not tried.

Wednesday, March 04, 2015

DebateGraph: Collective, Collaborative Wisdom in Action

This week’s presenter for the ISSIP SIG Education & Research Service Evangelist Series was David Price, Ph.D. Organizational Learning (Cambridge University), Judge Business School, Consultant & Public Policy Advisor. He will discuss DebateGraph – Collaborative Wisdom in Action.  The system can be see in action  here.  There you can explore Debategraph using David Price's own map.   (Slide link will be posted here)

Anyone can use this freely, and create maps that are public or private.   External documents can be attached on the web.  Transparency of knowledge is the default. and considered key.    See also Zakta.

Has similarity to concept maps, but with  more problem oriented visualization. Multi lingual interaction?   Euro commission example.

" ... DebateGraph.org is an award-winning cloud-based platform that enables communities of any size to build and share dynamic interactive visualizations of all the ideas, arguments, evidence, options and actions that anyone in the community believes relevant to the issues under consideration, and to ensure  that all perspectives are represented transparently, fairly, and in full in a meaningful, structured, and iterative dialogue.DebateGraph.org is an award-winning cloud-based platform that enables communities of any size to build and share dynamic interactive visualizations of all the ideas, arguments, evidence, options and actions that anyone in the community believes relevant to the issues under consideration, and to ensure that all perspectives are represented transparently, fairly, and in full in a meaningful, structured, and iterative dialogue. ... " 

(Update)  A description of DebateGraph, using the debategraph visualization.

Monday, February 23, 2015

Warning of a Digital Dark Age

Despite methods like the Internet Archive, in the BBC:

" ... Vint Cerf, a "father of the internet", says he is worried that all the images and documents we have been saving on computers will eventually be lost. Currently a Google vice-president, he believes this could occur as hardware and software become obsolete.

He fears that future generations will have little or no record of the 21st Century as we enter what he describes as a "digital Dark Age".. ... "

Saturday, February 21, 2015

What Does the Term Cognitive Assistance Mean?

A Linkedin Discussion started in the CSG group by Frank Stein, Director of Analytics Solution Center at IBM.  I have added some additional examples and comments

What Does the Term "Cognitive Assistance" mean? .... 

A US Government representative asked one of our team, Chuck Howell, "What does the term 'Cognitive Assistance' mean"? What is in scope for cognitive assistance systems? Chuck stitched together this working definition, after a quick back and forth, by combining excerpts from two important documents ....  "

Chuck Howell, Mitre, Scott Kordella, Mitre,
Frank Stein, IBM

Thursday, January 15, 2015

Open Source Application of Natural Language

Natural language, our use of communications like text and speech and documents to communicate with other people,  would be the ideal way to communicate with computers.    Yet there are still complexities of its use that are difficult to implement.  Today as part of the Cognitive Systems Institute weekly talk series,  attended a presentation by Prof Chris Biemann of the University of Darmstadt, that described their work in this area.   Here are the slides.   Good, brief exposition of the problem and application directions.

Some of this can be explored interactively here:  JoBimText :
JoBimText is an open source framework for application of Distributional Semantics using lexicalized features. It is providing a software solution for automatic text expansion using contextualized distributional similarity The project is maintained by the Language Technology group at the TU Darmstadt and IBM Research. .. " 

Friday, November 21, 2014

Linked Data Ontology

Part of an exploration:    Much more at the link.

" ... What is Linked Data?
The Web enables us to link related documents. Similarly it enables us to link related data. The term Linked Data refers to a set of best practices for publishing and connecting structured data on the Web. Data from heterogeneous sources can be combined using typed links. Key technologies that support Linked Data are URIs (a generic means to identify entities or concepts in the world), HTTP(a simple yet universal mechanism for retrieving resources, or descriptions of resources), and RDF (a generic graph-based data model which helps to structure and link conceptual data).  ... "

Tuesday, November 18, 2014

Mail Re-Imagined in Verse

Today from IBM.   Mail that understands you ... Imagine email that works for you instead of email that makes you work. .... Guided by analytics, IBM Verse learns your behaviors to adapt to the way you work, wherever you work. And because it's built for business, it understands you have special security and privacy needs, too. ... " .

You can sign up now.  It describes this as an analytics, not a cognitive App.  Which surprises me. But the classic cognitive function, learning,  is mentioned.

" ... Less clutter, more clarity ... Move to a bright place far beyond the mess of senders, subjects and folders.

With built-in intelligence and a user-first, user-tested design, IBM Verse offers a faster, better way to manage your communications across devices, organize inbound and outbound information, and focus on what you need most. ... " 

All good thoughts.  Would like to try, though for a heavy user changing your E-mail system can be a considerable task.  So many links between documents, contacts, mail, groups, applications.  Link to video demo.  Press release.

Friday, October 31, 2014

Enterprise Message from Microsoft

In Computing Now:   I have started to wonder more about how the typical office utilities and methods, like Office 365, will integrate with the cognitive sciences. Documents are information, will they look the same when converted into knowledge?   What will the tools and services be that will connect these as needed in the cloud?   A technical view at the link that does not address this, but is about the infrastructure involved.

Monday, October 20, 2014

Invention of Google Scholar

Now ten years old.    Contains 160 million documents. " ... Making the world’s problem solvers 10% more efficient ... Ten years after a Google engineer empowered researchers with Scholar, he can’t bear to leave it  .. " 

" ... Some people have never heard of this service, which treats publications from scholarly and professional journals as a separate corpus and makes it easy to find otherwise elusive information. Others have seen it occasionally when a result pops up on their search activity, and may even know enough to use it for a specific task, like digging into medical journals to gather information on a specific ailment. But for a significant and extremely impactful slice of the population: researchers, scientists, academics, lawyers, and students training in those fields — Scholar is a vital part of online existence, a lifeline to critical information, and an indispensable means of getting their work exposed to those who most need it. ....  "

Thursday, October 09, 2014

So What is Watson Anyway?

A good compact video on how IBM Watson works, and also how it is different from the 'rule based' expert systems methods we developed in the 90s.

In particular the natural language ingesting of resources like documents, images and expertise.  Also including learning as a natural part of its process, which should improve maintenance of expertise, and  its use in new contexts.

We should note that these are claims, and quite ambitious ones, so how well this works in practice remains to be seen.   Actively investigating.   Health care application tests are the most prominent ones today.

Tuesday, June 03, 2014

More Cheaply Done

MJ Perry describes an example where an individual has been much more efficient than the government.  Here scanning documents. Many more examples like this exist.  We have infrastructure and technology now that make it very possible.  So why not?