/* ---- Google Analytics Code Below */
Showing posts with label Natural Language NLP. Show all posts
Showing posts with label Natural Language NLP. Show all posts

Tuesday, March 15, 2022

Therapists and AI

ARTIFICIAL INTELLIGENCE

The therapists using AI to make therapy better

By Charlotte Jee and Will Douglas Heaven

Researchers are learning more about how therapy works by examining the language therapists use with clients. It could lead to more people getting better, and staying better.

December 6, 2021

Kevin Cowley remembers many things about April 15, 1989. He had taken the bus to the Hillsborough soccer stadium in Sheffield, England, to watch the semifinal championship game between Nottingham Forest and Liverpool. He was 17. It was a beautiful, sunny afternoon. The fans filled the stands.

He remembers being pressed between people so tightly that he couldn’t get his hands out of his pockets. He remembers the crash of the safety barrier collapsing behind him when his team nearly scored and the crowd surged.

Hundreds of people fell, toppled like dominoes by those pinned in next to them. Cowley was pulled under. He remembers waking up among the dead and dying, crushed beneath the weight of bodies. He remembers the smell of urine and sweat, the sound of men crying. He remembers locking eyes with the man struggling next to him, then standing on him to save himself. He still wonders if that man was one of the 94 people who died that day.

The Trevor Project, America’s hotline for LGBT youth, is turning to a GPT-2-powered chatbot to help troubled teenagers—but it’s setting strict limits.

These memories have tormented Cowley his whole adult life. For 30 years he suffered from flashbacks and insomnia. He had trouble working but was too ashamed to talk to his wife. He blocked out the worst of it by drinking. In 2004 one doctor referred him to a trainee therapist, but it didn’t help, and he dropped out after a couple of sessions.

But two years ago he spotted a poster advertising therapy over the internet, and he decided to give it another go. After dozens of regular sessions in which he and his therapist talked via text message, Cowley, now 49, is at last recovering from severe post-traumatic stress disorder. “It’s amazing how a few words can change a life,” says Andrew Blackwell, chief scientific officer at Ieso, the UK-based mental health clinic treating Cowley.

What’s crucial is delivering the right words at the right time. Blackwell and his colleagues at Ieso are pioneering a new approach to mental-health care in which the language used in therapy sessions is analyzed by an AI. The idea is to use natural-language processing (NLP) to identify which parts of a conversation between therapist and client—which types of utterance and exchange—seem to be most effective at treating different disorders.

The aim is to give therapists better insight into what they do, helping experienced therapists maintain a high standard of care and helping trainees improve. Amid a global shortfall in care, an automated form of quality control could be essential in helping clinics meet demand. 

Ultimately, the approach may reveal exactly how psychotherapy works in the first place, something that clinicians and researchers are still largely in the dark about. A new understanding of therapy’s active ingredients could open the door to personalized mental-health care, allowing doctors to tailor psychiatric treatments to particular clients much as they do when prescribing drugs. ....'

Wednesday, October 06, 2021

Machine Learning can Catch Natural Language Attacks

Particularly like the broad notion of 'honey pot' or attraction approach for threats.  Reading the broader article to see possible reapplication.

Honeypot Security Technique Can Stop Attacks in Natural Language Processing  By Penn State News

A machine learning framework can proactively counter universal trigger attacks—a phrase or series of words that deceive an indefinite number of inputs—in natural language processing (NLP) applications.

Scientists at Pennsylvania State University (Penn State) and South Korea's Yonsei University engineered the DARCY model to catch potential NLP attacks using a honeypot, offering up words and phrases that hackers target in their exploits.

DARCY searches and injects multiple trapdoors into a textual neural network to detect and thresh out malicious content produced by universal trigger attacks.

When tested on four text classification datasets and used to defend against six different potential attack scenarios, DARCY outperformed five existing adversarial detection algorithms.

From Penn State News

Wednesday, January 13, 2021

Salesforce Doing Advanced Metric Analysis for NLP

Good to see interesting AI things in the sales-marketing domain, a place we played early on.

Salesforce researchers release framework to test NLP model robustness

Kyle Wiggers, @Kyle_L_Wiggers, January 13, 2021 6:00 AM in VentureBeat

In the subfield of machine learning known as natural language processing (NLP), robustness testing is the exception rather than the norm. That’s particularly problematic in light of work showing that many NLP models leverage spurious connections that inhibit their performance outside of specific tests. One report found that 60% to 70% of answers given by NLP models were embedded somewhere in the benchmark training sets, indicating that the models were usually simply memorizing answers. Another study — a meta analysis of over 3,000 AI papers — found that metrics used to benchmark AI and machine learning models tended to be inconsistent, irregularly tracked, and not particularly informative.

This motivated Nazneen Rajani, a senior research scientist at Salesforce who leads the company’s NLP group, to create an ecosystem for robustness evaluations of machine learning models. Together with Stanford associate professor of computer science Christopher Ré and University of North Carolina at Chapel Hill’s Mohit Bansal, Rajani and the team developed Robustness Gym, which aims to unify the patchwork of existing robustness libraries to accelerate the development of novel NLP model testing strategies. ... '

Saturday, December 05, 2020

Last Year in AI, Analytics, Machine Learning and Data Science ....

Good end of the year piece from KDNuggets that was instructive.

AI, Analytics, Machine Learning, Data Science, Deep Learning Research Main Developments in 2020 and Key Trends for 2021

Tags: 2021 Predictions, AI, Ajit Jaokar, Analytics, Brandon Rohrer, Daniel Tunkelang, Data Science, Deep Learning, Machine Learning, Pedro Domingos, Predictions, Research, Rosaria Silipo

2020 is finally coming to a close. While likely not to register as anyone's favorite year, 2020 did have some noteworthy advancements in our field, and 2021 promises some important key trends to look forward to. As has become a year-end tradition, our collection of experts have once again contributed their thoughts. Read on to find out more.

By Matthew Mayo, KDnuggets.

To the chagrin of absolutely no one, 2020 is finally drawing to a close. It has been a rollercoaster of a year, one defined almost exclusively by the COVID-19 pandemic. But other things have happened, including in the fields of AI, data science, and machine learning as well. To that end, it's time for KDnuggets annual year end expert analysis and predictions. This year we posed the question:

What were the main developments in AI, Data Science, Machine Learning Research in 2020 and what key trends do you see for 2021?

Last year's noted main developments and predictions included continued advancements in many research areas, NLP in particular. While there can be debate as to whether 2020's big NLP advancement was as formidable as some may have originally thought (or continue to think), there is no doubt that there was a continued and intense focus on NLP research in 2020. It should not be difficult to surmise that this continues into 2021 as well.  ... "

Friday, December 04, 2020

Shrinking BERT Networks to Model Language

Considerable shrinking of neural networks, more likely for applications at the Edge.

A new approach could lower computing costs and increase accessibility to state-of-the-art natural language processing.

Daniel Ackerman | MIT News Office

Researchers at the Massachusetts Institute of Technology (MIT), the University of Texas at Austin, and the MIT-IBM Watson Artificial Intelligence Laboratory identified lean subnetworks within a state-of-the-art neural network approach to natural language processing (NLP). These subnetworks, found in the Bidirectional Encoder Representations from Transformers (BERT) network, could potentially enable more users to develop NLP tools using less bulky and more efficient systems, like smartphones. BERT is trained by repeatedly attempting to fill in words omitted from a passage of writing, using a massive dataset; users can then refine its neural network to a specific task. By iteratively trimming parameters from the BERT model, then comparing the new subnetwork's performance to that of the original model, the team found effective subnetworks that were 40% to 90% leaner, and required no task-specific fine-tuning to identify "winning ticket" subnetworks that executed tasks successfully.  .... 

Tuesday, September 29, 2020

Numenta Research Meeting: A Look at GPT-3

 Also a talk on GPT-3 of interest  ...

Numenta Research Meeting: “Steve Omohundro on GPT-3”    by omohundro  On July 1, 2020, Steve Omohundro gave a talk on GPT-3 and it's implications for artificial intelligence to Numenta's Research Meeting: In this research meeting, guest Stephen Omohundro gave a fascinating talk on GPT-3, the new massive OpenAI Natural Language Processing model. He reviewed the network architecture, training process, and results in the context of […]

Friday, April 24, 2020

Following Legal Analytics: Contracts and Liabilities

Our recent looks at eDiscovery and related AI and analytics alerted us to this.  Accenture's connection would seem to indicate seriousness of these efforts.  It is a very obvious space for advanced natural language processing, analysis and cross referencing with current and future contexts.

Legal analytics: Accenture applies NLP to analyze contracts and liabilities

To find specific information in a million-plus contracts, the global professional services company turned to natural language processing and AI, launching a legal analytics hub in the process.
     
By Thor Olavsrud in CIO

Organizations steeped in text documents have an ally in their quest to streamline business processes. Natural language processing, a branch of AI focused on communication, is helping companies such as Accenture surface high-value information and cut costs by bringing text-based, unstructured communications into the machine learning age.  

With more than a million contracts in its records system and thousands more added monthly, Accenture’s legal organization of about 2,800 professionals was struggling to find specific information across contracts, thanks to a tedious, costly process for which detailed cross-document search capability was limited.   ... " 

Sunday, April 05, 2020

Apple to Improve Siri with Voysis?

Most voice assistants today do a rocky job of interpreting complex conversation and context.  Beyond just simple interpretation, but on to understanding.  Beyond just 'Do this' or 'Do That'.   The basis of conversation is useful and credible response.   I know, I use several different assistant versions in context every day. Will Voysis help?  They claim domain specific voice AI.   Is that sufficiently similar to conversation context specific?

Apple's latest acquisition could help Siri understand what you're saying in Engadget

Voysis focused on AI that could respond to natural language requests.

The battle between AI voice assistants continues to rage on, and now Apple has acquired a tech firm, Voysis, that is all about helping computers understand natural language. As reported by Bloomberg, the firm's now-deleted website said it could produce search results from phrases like "I need a new LED TV, my budget is $1,000."  ... '

See also an article in TechCrunch.

Thursday, March 12, 2020

Debater Improves the way Watson Uses Language

Things that came of of the 'Debater' project, some interesting ways to include in a system the way we construct words and phrases creatively.   The kinds of things I try hard not to use with my assistants.   Avoiding 'complicated word schemes', or even slightly complicated phrases.    Look forward to trying that.  See 'Project Debater' tag below for my previous posts on this.

IBM enhances Watson’s natural language understanding capabilities  By Mike Wheatley from Siliconangle

IBM Corp. said today it has made some big improvements to the natural language processing capabilities of its IBM Watson platform.

The new capabilities, which were born out of IBM Research’s Project Debater, will help Watson understand and analyze some of the most challenging aspects of English language with greater clarity than before, the company said.

IBM’s Project Debater is an artificial intelligence system that was built by the company to debate with humans on a complex range of topics. The system is designed to understand subtleties in the English language such as idioms and colloquialisms that traditional AI has always struggled with.

For example, a phrase such as “hot under the collar” is well-understood by humans, but most AI systems are likely to get the wrong end of the stick, so to speak, since their algorithms can’t detect the true meaning. But those kinds of idioms aren’t a problem for Project Debater, which is capable of much more advanced sentiment analysis.

IBM said it will be integrating Project Debater’s natural language processing capabilities into Watson “throughout the year.” Customers will then be able to add the new capabilities to their existing AI models to help them better exploit natural language.

“Language is a tool for expressing thought and opinion, as much as it is a tool for information,” Rob Thomas, general manager of IBM Data and AI, said in a statement. “This is why we believe that advancing our ability to capture, analyze and understand more from language with NLP will help transform how businesses utilize their intellectual capital that is codified in data.”

IBM said Project Debater adds four distinct capabilities to Watson, including more advanced sentiment analysis. As a result, it said, it can better identify and understand complicated word schemes like idioms and so-called “sentiment shifters” — combinations of words that, taken together, take on new meaning, such as “hardly helpful.”   .... " 

Friday, February 28, 2020

The Hutter Prize

Just informed of this work, via a podcast referenced below.  Still aiming to make some further sense of this in general.  Technical.

The Hutter Prize Site
Being able to compress well is closely related to intelligence as explained below. While intelligence is a slippery concept, file sizes are hard numbers. Wikipedia is an extensive snapshot of Human Knowledge. If you can compress the first 1GB of Wikipedia better than your predecessors, your (de)compressor likely has to be smart(er). The intention of this prize is to encourage development of intelligent compressors/programs as a path to AGI. ... 

Interview with Lex Fridman (26.Feb'20) (Video, Audio, Tweet)  ... " 

In the Wikipedia (The Hutter Prize)
"... The goal of the Hutter Prize is to encourage research in artificial intelligence (AI). The organizers believe that text compression and AI are equivalent problems. Hutter proved that the optimal behavior of a goal seeking agent in an unknown but computable environment is to guess at each step that the environment is probably controlled by one of the shortest programs consistent with all interaction so far.[4] However, there is no general solution because Kolmogorov complexity is not computable. Hutter proved that in the restricted case (called AIXItl) where the environment is restricted to time t and space l, a solution can be computed in time O(t2l), which is still intractable.

The organizers further believe that compressing natural language text is a hard AI problem, equivalent to passing the Turing test. Thus, progress toward one goal represents progress toward the other.[5] They argue that predicting which characters are most likely to occur next in a text sequence requires vast real-world knowledge. A text compressor must solve the same problem in order to assign the shortest codes to the most likely text sequences.  .... " 

Saturday, February 15, 2020

Machines Understanding Language

Have seen a number claims recently of how good machine understanding of human language had advanced.  Here is a contrary view.  Its all hack to the basics of common sense.

Artificial Intelligence / Machine Learning
AI still doesn’t have the common sense to understand human language
Natural-language processing has taken great strides recently—but how much does AI really understand of what it reads? Less than we thought.  ... 
by Karen Hao

Friday, January 10, 2020

SAS CEO on Augmenting People and Processes

Jim Goodnight of SAS helped us through many of the early years of doing analytics in the enterprise.  Good to see him talking about natural language processes and AI, which we never hear of much of in those days, when classic statistics reigned.

AI technologies that matter now: Augmenting People, Processes, and Potential   Natural language processing, machine learning, and computer vision promise to extend human capabilities, and ultimately, improve the world around us.

by Jim Goodnight in Technology Review

Sponsored Content   Provided by SAS
Jim Goodnight is co-founder and CEO at SAS.

From mass surveillance to mind-reading machines, each new day seems to bring another alarming prediction about the potential of artificial intelligence to change the world.

But we don’t give enough attention to the practical AI applications that are in use every day. These real-world applications aren’t creepy or futuristic. You might even call some of them mundane. But they provide practical value to businesses and consumers, and they aren’t leading us to impending doom.

AI does have the potential to change our world. But it’s not going to do that through sentient robots or computers that take control away from humans. The AI applications that we see will more often augment human activity than replace it.

As we continue to witness increased computing power and a more connected world, practical AI technologies like natural language processing, computer vision, and especially machine learning will proliferate and become even more useful. These are the practical applications of AI that will improve our lives..... " 

Friday, January 03, 2020

Baidu and GLUE China Benchmarks

China and langage meaning.  A very key part of ultimately delivering useful conversation.

Baidu has a new trick for teaching AI the meaning of language in 7Wdata

Earlier this month, a Chinese tech giant quietly dethroned Microsoft and Google in an ongoing competition in AI. The company was Baidu, China’s closest equivalent to Google, and the competition was the General Language Understanding Evaluation, otherwise known as GLUE.

GLUE is a widely accepted benchmark for how well an AI system understands human language. It consists of nine different tests for things like picking out the names of people and organizations in a sentence and figuring out what a pronoun like “it” refers to when there are multiple potential antecedents. A language model that scores highly on GLUE, therefore, can handle diverse reading comprehension tasks. Out of a full score of 100, the average person scores around 87 points. Baidu is now the first team to surpass 90 with its model, ERNIE. .... "

Sunday, May 12, 2019

Cisco Moves with MindMeld

Have we seen the emergence of a new general assistant?  I remember the acquisition of Mindmeld, but did not expect this direction.  Don't expect a device to follow,   but likely a business oriented conversational AI and assistance, with the ability to embed in other systems.   Smart homes?  There has been considerable posts about MindMeld in this blog.  Our innovation center did considerable work with Cisco.

Cisco opens up its MindMeld voice AI platform    By Maria Deutscher in SiliconAngle

Back in 2017, Cisco Systems Inc. shelled out $125 million to acquire MindMeld Inc., an early-stage startup that had created a platform for building voice assistants. The offering was one of the first development tools focused specifically on conversational artificial intelligence.

Cisco announced Wednesday that it’s releasing MindMeld under an open-source Apache 2.0 license. The move is meant to lower the adoption barrier for enterprises looking to add voice features to their applications. Just  as important, it will enable outside developers to improve upon to the platform and contribute their code back to the project.

MindMeld provides a set of tools that cover most of the core tasks involved in building a conversational AI. At the heart of the platform is a natural-language processing engine for parsing spoken commands. It can identify the topic a user is talking about, isolate what it is exactly they’re asking for and analyze “entities” such as restaurant names that require special interpretation.

Once a request is parsed, it’s passed on to MindMeld’s question-answering engine. The system can automatically generate replies by drawing on a corpus of information provided in advance by the application developer. MindMeld includes a dedicated module for managing a service’s knowledge repository that helps with tasks such as organizing data and finding alternative answers when an application’s first reply misses the mark..... " 

Saturday, April 13, 2019

Natural Language and Intent with Ambiguity

Good simple explanation from Tableau on Intent.   Have used the concept now in several projects, and of course there is ambiguity, beyond dictionary-definition,  in the use of many terms within a company.   The ambiguity in context is important to consider.

Machine learning, natural language meet to understand intent
 By Mark Jewett, VP of Marketing, Tableau

Machine learning and natural language processing promise to better translate human curiosity into pertinent answers. If true, these smart capabilities will broaden the use of analytics and reach people who are less comfortable dealing with data. It will all start with helping machines learn to interpret human intent. The key is semantics.

Sometimes intent is simple and explicit, like asking Siri or Alexa if a flight is delayed. This question has clear intention and a simple response—returning the flight status answers the question. Such simplicity is seldom the case when it comes to data analysis. Questions are usually more nuanced, making it hard to correctly assume what the user is really looking for. Natural language is even more tricky where ambiguous terms are common.

It’s also difficult for a machine to understand our intent within a limited context. The machine has the data itself but doesn’t grasp the bigger picture in the same way a person with domain expertise can. Asking “How are my sales doing in the Northeast?” is a lot more ambiguous than the flight status example above.

Ambiguity isn’t a new challenge in data analysis. Different groups within an organization may have different definitions or calculations for the same words: for example, the term “profitability”. Some organizations use central dictionaries (also called data catalogs) to reduce ambiguity and create consistency across the organization. These tools can help provide users with the context they need to understand more deeply. .... " 

Monday, April 01, 2019

Machine Learning and Intent

Deriving intent,  goals is still hard.  Especially in multiple component conversations.

Machine Learning, Natural Language Meet to Understand Intent     By Mark Jewett, VP, product marketing, Tableau Software in InformationWeek

Machine learning and natural language capabilities will bring the power of analytics to more people through semantics.

Machine learning and natural language processing promise to better translate human curiosity into pertinent answers. If true, these smart capabilities will broaden the use of analytics and reach people who are less comfortable dealing with data. It will all start with helping machines learn to interpret human intent. The key is semantics.

Sometimes intent is simple and explicit, like asking Siri or Alexa if a flight is delayed. This question has clear intention and a simple response -- returning the flight status answers the question. Such simplicity is seldom the case when it comes to data analysis. Questions are usually more nuanced, making it hard to correctly assume what the user is really looking for. Natural language is even more tricky where ambiguous terms are common.

It’s also difficult for a machine to understand our intent within a limited context. The machine has the data itself but doesn’t grasp the bigger picture in the same way a person with domain expertise can. Asking “How are my sales doing in the Northeast?” is a lot more ambiguous than the flight status example above..... "

Wednesday, January 16, 2019

KDNuggets: Books on Language Processing

I have read KDNuggets long before deep learning was discovered.    Love it.  The most recent post was useful for a project:

KDnuggets™ News 19:n03, Jan 16: Top 10 Books on NLP and Text Analysis; End To End Guide For Machine Learning Projects

Also: Why Vegetarians Miss Fewer Flights - Five Bizarre Insights from Data; 4 Myths of Big Data and 4 Ways to Improve with Deep Data; The Role of the Data Engineer is Changing; How to solve 90% of NLP problems: a step-by-step guide ... 

This week, check out a collection of top 10 NLP and text analysis books, see a step to step guide on the process that you can follow to implement a successful data science project, find out why vegetarians miss fewer flights (???), read about the fundamental misconception that bigger data produces better machine learning results, and find out how the role of the data engineer is changing. ... " 

Tuesday, December 04, 2018

Alibaba has a Better Intelligent Assistant than Google's

In particular the claim that it is more conversationally powerful is interesting.  Such advances could lead to deeper interaction with humans and further engagement.

Alibaba already has a voice assistant way better than Google’s
It navigates interruptions and other tricky features of human conversation to field millions of requests a day   by Karen Hao  in Technology Review excerpt: 

In May, Google made quite the splash when it unveiled Duplex, its eerily humanlike voice assistant capable of making restaurant reservations and salon appointments. It seemed to mark a new milestone in speech generation and natural-language understanding, and it pulled back the curtain on what the future of human-AI interaction might look like.  Still only in Chinese. 

But while Google slowly rolls out the feature in a limited public launch, Alibaba’s own voice assistant has already been clocking overtime. On December 2 at the 2018 Neural Information Processing Systems conference, one of the largest annual gatherings for AI research, Alibaba demoed the AI customer service agent for its logistics company Cainiao. Jin Rong, the dean of Alibaba’s Machine Intelligence and Technology Lab, said the agent was already servicing millions of customer requests a day.

The demo call involved the agent asking a customer where he wanted his package delivered. In the back-and-forth exchange, the agent successfully navigated several conversational elements that demonstrated the breadth of its natural-language capabilities.

Take this exchange at the beginning of the call, translated from Mandarin: ... "

Wednesday, August 08, 2018

Names in Speech Recognition for Assistants

An example of the complexity of natural language understanding in assistants. Accuracy gets more essential in business applications.  Likely technical.  Usually all slides and recordings are placed in the site at the bottom within a few days.

CSIG (Cognitive Systems Institute Group) Talk — Aug 9, 2018 - 10:30-11am US Eastern

Title: Correcting person names for automatic speech recognition (ASR) vendor-agnostic voice assistants

Speaker: Vijay Ramakrishnan, Tue Minh Vo (Cisco)

Abstract: 
ASR systems trained on generic data often mis-transcribe domain-specific words and phrases. For voice assistants, errors  in the ASR transcript cascade to the assistant's natural language understanding (NLU) components. We focus on the  problem of ASR errors in person names and describe a novel method of correcting person names by leveraging a domain- specific language model (LM), and character and phoneme-based information retrieval (IR) techniques.

Short bios:
Vijay Ramakrishnan is a ML/NLP engineer at Cisco's Cognitive Collaboration Group where his team develops  conversational Al products for Cisco's collaboration portfolio. His research interests include deep networks for domain-  specific ASR, empirical methods for NLP and ML for sequence models.

Bio: Minh-Tue Vo is a senior engineer at Cisco's Cognitive Collaboration Group where his team develops conversational Al
products for Cisco's collaboration portfolio. His research interests include deep networks for domain-specificASR,  empirical methods for NLP and ML for sequence models.

Zoom meeting Link: https://zoom.us/j/7371462221
Zoom Cailin: (415) 762-9988 or (646) 568-7788 Meeting id 7371462221
Zoom International Numbers: https://zoom.us/zoomconference
( Check the website in case the date or time changes: http://cognitive-science.info/community/weekly-update/ )

Sunday, June 17, 2018

How Businesses Can Get Inside the Minds of Their Competitors

If we could, perhaps we could make a better use of game dynamics.     Or just simulate their behavior under multiple contexts.

How Businesses Can Get Inside the Minds of Their Competitors

Wharton's Anoop Menon and Jaeho Choi discuss their research on using natural language processing to analyze competitive strategy.

Every business would love to know the minds of its competitors, and what they are likely to do next. Strategy analysts have thus far used simple tools that employ mostly financial and other structured data to try and predict competitors’ moves. But new research at Wharton has shown how natural language processing techniques could be used to parse tomes of unstructured data such as text buried in conference calls or annual reports to more accurately anticipate competitor strategies.

The research opens new pathways to measure and test assumptions firms make in their competitive strategies, and to “visualize how firms are positioned with respect to each other, and then map that on to performance consequences,” says Wharton management professor Anoop Menon. His research paper, “What You Say Your Strategy Is and Why It Matters: Natural Language Processing of Unstructured Texts,” is co-authored with Jaeho Choi, a Wharton doctoral student, and Haris Tabakovic, an associate at The Brattle Group, a Boston-based international arbitration services firm.

For their study, the researchers used natural language processing (NLP) techniques to measure “strategic change, positioning, and focus,” across their sample of 50,506 business descriptions of publicly held companies contained in their 10-K annual reports, from 1997 to 2016.

Menon and Choi shared the main takeaways for business strategy analysis from their research with Knowledge@Wharton.

An edited transcript of the conversation follows.

Knowledge@Wharton: Anoop, could you tell us what led you to explore this topic in your research? What was your objective?

Anoop Menon: The notion that there is a lot of information that is buried in unstructured text has been around for a while. We know that strategy is very complicated, but we tend to measure it using very, “simple metrics” like a few financials here and there. But we all agree and understand there is a huge amount of information that is buried in text like conference calls and annual reports that gets at the meat of the strategy, how the strategists are thinking about competition and product market choices.

Sadly, we currently don’t have a really good technique or set of techniques to get at that information. So that was the starting point. About six or seven years ago, my co-author Haris [Tabakovic] and I came across this burgeoning line of research in computer science about using natural language processing techniques to extract text, but in very different fields – not ours. [There were] some applications to political science but not at all to strategy. We said we should be able to take some of those techniques and get at the information that is buried in the text..... "