/* ---- Google Analytics Code Below */
Showing posts with label Language. Show all posts
Showing posts with label Language. Show all posts

Saturday, July 01, 2023

The Boundary Between Human Language and ChatGPT Is Fuzzier Than You Think

The Boundary Between Human Language and ChatGPT Is Fuzzier Than You Think

By  Brendan H. O'Connor

June 12, 2023  in Singularity Hub

ChatGPT is a hot topic at my university, where faculty members are deeply concerned about academic integrity, while administrators urge us to “embrace the benefits” of this “new frontier.” It’s a classic example of what my colleague Punya Mishra calls the “doom-hype cycle” around new technologies. Likewise, media coverage of human-AI interaction—whether paranoid or starry-eyed—tends to emphasize its newness.

In one sense, it is undeniably new. Interactions with ChatGPT can feel unprecedented, as when a tech journalist couldn’t get a chatbot to stop declaring its love for him. In my view, however, the boundary between humans and machines, in terms of the way we interact with one another, is fuzzier than most people would care to admit, and this fuzziness accounts for a good deal of the discourse swirling around ChatGPT.

When I’m asked to check a box to confirm I’m not a robot, I don’t give it a second thought—of course I’m not a robot. On the other hand, when my email client suggests a word or phrase to complete my sentence, or when my phone guesses the next word I’m about to text, I start to doubt myself. Is that what I meant to say? Would it have occurred to me if the application hadn’t suggested it? Am I part robot? These large language models have been trained on massive amounts of “natural” human language. Does this make the robots part human?

AI chatbots are new, but public debates over language change are not. As a linguistic anthropologist, I find human reactions to ChatGPT the most interesting thing about it. Looking carefully at such reactions reveals the beliefs about language underlying people’s ambivalent, uneasy, still-evolving relationship with AI interlocutors.

ChatGPT and the like hold up a mirror to human language. Humans are both highly original and unoriginal when it comes to language. Chatbots reflect this, revealing tendencies and patterns that are already present in interactions with other humans.  ... '

Sunday, April 16, 2023

Robot Brains: Neurosymbolics

Thoughtful piece that outlines what is necessary for Building Future Intelligence.  Not sure that current results are yet close to this.  Year old.  Technical.   

https://youtu.be/fCoavgGZ64Y

Neurosymbolic Models

86,914 views  Sep 21, 2021  Season One | The Robot Brains Podcast

On the last episode of Season One, our guest is Ilya Sutskever. Ilya is the Co-Founder and Chief Scientist of OpenAI. As a PhD student at Toronto, Ilya was one of the authors on the 2012 AlexNet paper that completely changed the field of AI, resulting in the widespread adoption of deep learning and the avalanche of AI breakthroughs we’ve seen the past 10 years. 

After the AlexNet breakthrough in computer vision, at Google, among many other breakthroughs, Ilya showed that neural networks are unexpectedly great at machine translation, at least at the time it was unexpected, now it’s long become the norm to use neural nets for machine translation. Late 2015 Ilya left Google to co-found OpenAI, where he is Chief Scientist. Some of his breakthroughs include GPT, CLIP, DallE, Codex. Ilya’s academic work, less than 10 years out of his PhD, has ben cited over 250,000 times, reflecting his absolutely mind-blowing influence on the field. 

What's in this episode:

00:00:00 Introductions

00:03:00 Why take a closer look at neural net works originally?

00:08:25 What was going through Ilya's mind during the AlexNet discovery?

00:18:25 Ilya's early years 

00:21:19 How Ilya stayed motivated 

00:29:07 Sam Altman and the beginning of OpenAI

00:36:22 LSTM models and reinforcement learning 

00:56:06 How will our productivity change?

01:00:22 Instruction-following models

01:12:13 Ilya's vision of the future of work

01:16:14 Ilya's advice to be productive 


| SUBSCRIBE TO THE ROBOT BRAINS PODCAST TODAY | 

Website: https://therobotbrains.ai 

Twitter: https://twitter.com/therobotbrains

LinkedIn:https://www.linkedin.com/company/the-...


Host: Pieter Abbeel

Executive Producers: Ricardo Reyes & Henry Tobias Jones  .... '

Wednesday, March 22, 2023

Large language models also Work for Protein Structures

We asked this question long ago, usefully answered?  Very big deal if so.

THE LANGUAGE OF BIOCHEMISTRY —

Large language models also work for protein structures

Training on raw protein sequences allows the AI to make inferences about structure.

JOHN TIMMER - 3/16/2023, 3:01 PM

The success of ChatGPT and its competitors is based on what's termed emergent behaviors. These systems, called large language models (LLMs), weren't trained to output natural-sounding language (or effective malware); they were simply tasked with tracking the statistics of word usage. But, given a large enough training set of language samples and a sufficiently complex neural network, their training resulted in an internal representation that "understood" English usage and a large compendium of facts. Their complex behavior emerged from a far simpler training.

A team at Meta has now reasoned that this sort of emergent understanding shouldn't be limited to languages. So it has trained an LLM on the statistics of the appearance of amino acids within proteins and used the system's internal representation of what it learned to extract information about the structure of those proteins. The result is not quite as good as the best competing AI systems for predicting protein structures, but it's considerably faster and still getting better.  

LLMs: Not just for language

The first thing you need to know to understand this work is that, while the term "language" in the name "LLM" refers to their original development for language processing tasks, they can potentially be used for a variety of purposes. So, while language processing is a common use case for LLMs, these models have other capabilities as well. In fact, the term "Large" is far more informative, in that all LLMs have a large number of nodes—the "neurons" in a neural network—and an even larger number of values that describe the weights of the connections among those nodes. While they were first developed to process language, they can potentially be used for a variety of tasks. .... ' 

Friday, February 17, 2023

Training Data from Audiobooks?

 Ultimately its all about training data.  And the quality of that data. Of course a Google gets lots of language data, but its it the right quality and type?    Authors giving up data that could replace them.  300 audiobooks probably not near enough.  

Audiobook Narrators Fear Apple Used Their Voices to Train AI   By Wired, February 16, 2023

An email to SAG-AFTRA union members seen by WIRED said the two companies had agreed to stop all “use of files for machine learning purposes” for union members affected

Gary Furlong, a Texas-based audiobook narrator, had worried for a while that synthetic voices created by algorithms could steal work from artists like himself. Early this month, he felt his worst fears had been realized.

Furlong was among the narrators and authors who became outraged after learning of a clause in contracts between authors and leading audiobook distributor Findaway Voices, which gave Apple the right to "use audiobooks files for machine learning training and models." Findaway was acquired by Spotify last June.

Some authors and narrators say they were not clearly informed about the clause and feared it may have allowed their work or voices to contribute to Apple's development of synthetic voices for audiobooks. Apple launched its first books narrated by algorithms last month. "It was very disheartening," says Furlong, who has narrated over 300 audiobooks and is one of more than a dozen narrators and authors who told WIRED of their concerns with Findaway's agreement. "It feels like a violation to have our voices being used to train something for which the purpose is to take our place," says Andy Garcia-Ruse, a narrator from Kansas City.

From Wired

View Full Article   

Wednesday, February 08, 2023

Reactions to todays Google/AI Bard Event

Nicely done, worth taking a look, should still be available om the Google Youtube site.

Here my very early impressions.  Based on building systems that did this.

Not too much was specifically said about Bard, except that it would handle chat interactions with their language system.  Just like today you can ask Google any question and it looks for direct and partial matches.   Often very useful, but sometimes irrelevant to what you want.    Usually less useful as your question is more complex.  Cannot usually pin together knowledge from multiple tries.  Will Bard do better at that?  

Also, how can you determine the source of information?   Hint was there would be some sort of button to push to get sources.  Often that is very useful to determine your trust in a result.   How about a way to measure the risk of a result?  Are you sure you want to do that?   Have Legal involved?   Had the need for that too.  Also does the provider of information get an indication that their info was used, giving them incentive to provide more?   Had that problem with internal Company wikis.   

Mapping updates were interesting, for example the ability to have .immersive view. linked to maps.  So you could get deep local understanding of a view, to improve navigation in a a city.    How about a historical view?  Had the cause to use that for city planning immersion.  

Also linked to maps an 'Indoor Live View', which  let you look at internal design,  so a company could provide precise internal navigation for say retail spaces.   It could be as detail as needed.  We experimented with the idea in grocery type designs, even adding virtual ads that could be updated as wanted.  

Further,  any kind of system that deals with language generation needs to consider context.  Language understanding and usage is important.    Who will be using the results?  Children, New employees, Chemical Engineers?  Trainers or trainees?    Also studied this in Wiki applications.   

Also, have examples where such  a system refuses to generate a particular result,   because there is some (imagined?)  corporate problem with it.  (Political, Operational? )  Who gets to decide that we wont go there?  Who owns the output.  The User? Organization? Google?   Who endorses the answer?    As you start to make these things very universally used, these issues will come up.     Will there be a personal data issue?    Likely.  Should the results always be stored in a memory for later reuse?    - FAD 

Monday, September 05, 2022

AI Learns Patterns of Human Language

 Trained and tested from Linguistic textbooks in  languages.  Some surprises in learning cross languages. 

AI Can Learn the Patterns of Human Languages

In MIT News  By Adam Zewe, August 30, 2022

Researchers at Massachusetts Institute of Technology, Cornell University, and McGill University developed an artificial intelligence model that can learn the rules and patterns of human languages automatically, without specific human guidance. The model was trained and tested on problems from linguistic textbooks in 58 different languages that involved word-form changes. The researchers observed that the model could determine a correct set of rules to describe the word-form changes for 60% of the problems. Said Cornell's Kevin Ellis, "One of the things that was most surprising is that we could learn across languages, but it didn't seem to make a huge difference. That suggests two things. Maybe we need better methods for learning across problems. And maybe, if we can't come up with those methods, this work can help us probe different ideas we have about what knowledge to share across problems."  ... 

Tuesday, August 23, 2022

Stories, Dice, and Rocks That Think:

Just  Reading, very interesting, will review further as I progress.  by a correspondent I have often mentioned here.  


Stories, Dice, and Rocks That Think: How Humans Learned to See the Future--and Shape It  ...  
 by Byron Reese 

". . . Byron Reese gets to the heart of what makes humans different from all others." —Midwest Book Review

What makes the human mind so unique? And how did we get this way?   Amazon Description:

This fascinating tale explores the three leaps in our history that made us what we are—and will change how you think about our future.

Look around. Clearly, we humans are radically different from the other creatures on this planet. But why? Where are the Bronze Age beavers? The Iron Age iguanas? In Stories, Dice, and Rocks That Think, Byron Reese argues that we owe our special status to our ability to imagine the future and recall the past, escaping the perpetual present that all other living creatures are trapped in. 

Envisioning human history as the development of a societal superorganism he names Agora, Reese shows us how this escape enabled us to share knowledge on an unprecedented scale, and predict—and eventually master—the future.

Thoughtful, witty, and compulsively readable, Reese unravels our history as an intelligent species in three acts: 

Act I: Ancient humans undergo “the awakening,” developing the cognitive ability to mentally time-travel using language

Act II: In 17th century France, the mathematical framework known as 'probability theory' is born—a science for seeing into the future that we used to build the modern world

Act III: Beginning with the invention of the computer chip, humanity creates machines to gaze into the future with even more precision, overcoming the limits of our brains

A fresh new look at the history and destiny of humanity, readers will come away from Stories, Dice, and Rocks that Think with a new understanding of what they are—not just another animal, but a creature with a mastery of time itself.  ... ' 

Friday, May 20, 2022

Meta AI System

 Plan to test

ARTIFICIAL INTELLIGENCE

Meta has built a massive new language AI—and it’s giving it away for free

Facebook’s parent company is inviting researchers to pore over and pick apart the flaws in its version of GPT-3

By Will Douglas    May 3, 2022

Meta’s AI lab has created a massive new language model that shares both the remarkable abilities and the harmful flaws of OpenAI’s pioneering neural network GPT-3. And in an unprecedented move for Big Tech, it is giving it away to researchers—together with details about how it was built and trained.

“We strongly believe that the ability for others to scrutinize your work is an important part of research. We really invite that collaboration,” says Joelle Pineau, a longtime advocate for transparency in the development of technology, who is now managing director at Meta AI.

Meta’s move is the first time that a fully trained large language model will be made available to any researcher who wants to study it. The news has been welcomed by many concerned about the way this powerful technology is being built by small teams behind closed doors.

“I applaud the transparency here,” says Emily M. Bender, a computational linguist at the University of Washington and a frequent critic of the way language models are developed and deployed.  

“It’s a great move,” says Thomas Wolf, chief scientist at Hugging Face, the AI startup behind BigScience, a project in which more than 1,000 volunteers around the world are collaborating on an open-source language model. “The more open models the better,” he says.  ... 


Monday, May 09, 2022

Beyond Interpretability, Towards Language

Yes, with training in the questions we need to ask to shape the results we need for a given context. Not always easy. 

Beyond Interpretability: Developing a Language to Shape Our Relationships with AI  By Medium, May 9, 2022 in CACM

By Been Kim, Google Research, Brain Team

Artificial intelligence (AI) is more than just a tool; we will be influenced by it, and that will influence the next generation of AI. Without a language to meaningfully communicate with it, we do not understand its decisions and, therefore, will not know what we are creating.

Building a language to communicate with AI will not be easy, but it is the only way to gain control of the way we want to live. Languages shape the way we think. We have an opportunity to shape our own thinking and future machines.

From Medium  

Tuesday, January 11, 2022

On New Open Language Models

Bigscience and open Language Models

Inside BigScience, the quest to build a powerful open language model

Kyle Wiggers  @Kyle_L_Wiggers  in VentureBeat.  

January 10, 2022 9:30 AM

Roughly a year ago, Hugging Face, a Brooklyn, New York-based natural language processing startup, launched BigScience, an international project with more than 900 researchers that is designed to better understand and improve the quality of large natural language models. Large language models (LLMs) — algorithms that can recognize, predict, and generate language on the basis of text-based datasets — have captured the attention of entrepreneurs and tech enthusiasts alike. But the costly hardware required to develop LLMs has kept them largely out of reach of researchers without the resources of companies like OpenAI and DeepMind behind them.

Taking inspiration from organizations like the European Organization for Nuclear Research (also known as CERN),  and the Large Hadron Collider, the goal of BigScience, then, is to create LLMs and large text datasets that will eventually be open-sourced to the broader AI community. The models will be trained on the Jean Zay supercomputer located near Paris, France, which ranks among the most powerful machines in the world.

“From Data to Knowledge”. How the Organization of Data Using LC:NC Can Drastically Reduce the Technical Complexity of Deriving Knowledge From Data._

While the implications for the enterprise might not be immediately clear, efforts like BigScience promise to make LLMs more accessible — and transparent — in the future. With the exception of several models created by EleutherAI, an open AI research group, few trained LLMs exist for research or deployment into production. OpenAI has declined to open source its most powerful model, GPT-3, in favor of exclusively licensing the source code to Microsoft. Meanwhile, companies like Nvidia have released the code for capable LLMs, but left the training of those LLMs to users with sufficiently powerful hardware. ... ' 

Monday, January 03, 2022

Does AI Understand Our Language?

Some good points made here regarding the context of language understanding. 

Does AI Understand Our Language?

By TechTalks, December 27, 2021

If a computer gives you all the right answers, does it mean that it is understanding the world as you do? This is a riddle that artificial intelligence scientists have been debating for decades. And discussions of understanding, consciousness, and true intelligence are resurfacing as deep neural networks have spurred impressive advances in language-related tasks.

Many scientists believe that deep learning models are just large statistical machines that map inputs to outputs in complex and remarkable ways. Others beg to differ, arguing that large language models have a great deal to teach us about "the nature of language, understanding, intelligence, sociality, and personhood."

From TechTalks

View Full Article   


Thursday, October 28, 2021

Ai Modeling Brain Processing Language

Notable is the statement that AI is not attempting to directly mimic the brain.  But the question is always how closely should we use the brain as a working model?   At the link includes much more detail and an explanatory video.

AI Sheds Light on How the Brain Processes Language

MIT News, Anne Trafton, October 25, 2021

Research by Massachusetts Institute of Technology (MIT) neuroscientists suggests the latest predictive language models' underlying mechanism functions similarly to the human brain's language-processing centers. MIT's Nancy Kanwisher said, "The better the model is at predicting the next word, the more closely it fits the human brain." Computer models that perform well on other language tasks do not exhibit this resemblance, implying the brain may drive language processing using next-word prediction. Stanford University's Daniel Yamins said, "Since the AI [artificial intelligence] network didn't seek to mimic the brain directly—but does end up looking brain-like—this suggests that, in a sense, a kind of convergent evolution has occurred between AI and nature."  ... ' 

Saturday, July 31, 2021

On Communication with Animals

Are we closer to communication with parrots, chimps, dolphins?  Will it be a key aspect of AI? 

On Communication By Vinton G. Cerf  in CACM.

Communications of the ACM, August 2021, Vol. 64 No. 8, Page 5  10.1145/3472146

As I write this, summer is upon us in the Northern Hemisphere. I have just attended an online lecture about non-human species communication, sponsored by the Interspecies Internet project (interspecies.io). While the primary objective of the project is to determine experimentally whether it is possible to demonstrate communication between non-human species, there is also considerable interest in understanding the nature of intraspecies communication. The lecturer, Ofer Tchernichovski, explored years of experience with zebra finches. Of particular interest were their songs and how they propagated through generations of "tutors" and "pupils" among families of finches. Among the interesting observations he made was a concern that we sometimes bring preconceived but unwarranted notions to science. For example, consider the way in which we might analyze bird songs. We make audio recordings and spectral Fourier diagrams of the songs. We segment these vocalizations as if they might represent phonemes, but our segmentation could be inappropriately influenced by what we know of human speech.

Linguists have learned a great deal about human speech, how it is produced, and how the phonemes give structure to utterances. Whether we can apply such structural assumptions to bird songs is a matter for research. Tchernichovski points out that an alien arriving on planet Earth, even if it is capable of sensing human speech, might not have any idea how to segment sounds into phonemes and words. Language is a concept that organizes sound into phonemes, words, and sentences representing structures that follow grammatical rules and from which semantic content can be derived. The alien might not have any a priori clue as to how human languages are expressed, parsed, and give rise to semantic meaning. If the alien itself has language, it might adopt a protocol for human language discovery, starting, for example, with self-identification. .... "


Monday, June 14, 2021

How Long Will it be Until we can Talk to Animals?

First I thought this was a silly question.   But it is really a primary question of intelligence.   Its not just a translation, but a need to parse into a useful structure so it can be translated into/from our world.  Peoples world vs An Animal world. In away that mapped translations will provide for useful goals.  It is hard.  Below just the intro.

How long before AI can 'understand' animals?

Scientists are working on it, but it's a rough job.

By James Trew, @itstrew. June 4th, 2021  in Engadget

In this article: gear, animal translation, feature, tomorrow, ai, artificial intelligence

The Regent Honeyeaters of Australasia are forgetting how to talk. The songbird’s habitat has been so severely devastated that its numbers are dwindling. Worse, the ones that remain are so scattered that the adult males are too far apart to teach the young how to sing for a mate — how to speak their own language. The gradual loss of the Honeyeaters’ song, their primary tool for wooing a partner, creates a vicious circle of spiraling decline. 

Humans, on the other hand, cannot shut up. Estimates peg the total number of languages in use today to be around 7,000. In the US, roughly 25 percent of people claim they can converse in a second language. In Europe this number floats around 60 percent. In Asia or Africa, bilingualism is even more common as local tongues and regional dialects live alongside (often multiple) “official” languages. But not one person on this planet can speak Cat or Dog — much less Regent Honeyeater.  ... " 


Thursday, June 03, 2021

AppTech Partners with Intel for AI Enabled Speech

A new partnership emerges.  In Cision PRNewswire

AppTek Partners with Intel to Foster the Development of Next Generation AI-Enabled Speech and Language Technologies

MCLEAN, Va., June 2, 2021 /PRNewswire/ -- AppTek, a leader in Artificial Intelligence (AI), Machine Learning (ML), Automatic Speech Recognition (ASR), Neural Machine Translation (NMT), Text-to-Speech (TTS) and Natural Language Processing / Understanding (NLP/U) technologies, announced a partnership with Intel to accelerate and enhance performance benchmarks for the company's award-winning AI-enabled ASR and NMT technologies as part of the Intel AI Builders program. Intel's AI Builders is an  ... 

Tuesday, April 13, 2021

Advances in Language-Based AI Tasks

Not surprising,  language is communications.  Between us and everything else.  Spent lots of time trying to make that work.

Advances in Language-Based AI Tasks Seen as Dawn of New Era    By AI Trends Staff  

Researchers at Accenture have found that 10% of early adopters of digital technologies have grown at twice the rate of the bottom 25%, and they are using cloud systems—not legacy systems—to enable adoption.  

H. James Wilson, Managing Director, Accenture:

“We expect the trend to accelerate among industry leaders over the coming five years,” stated the authors, H. James Wilson and Paul R. Daugherty, in an account in Harvard Business Review. And more specifically, following the release of the GPT-3 large language model from Open AI, “The 2020s will be about major advances in language-based AI tasks,” the authors suggest. 

Generative pre-trained transformers (GPTs) rely on a transformer, a mechanism that learns contextual relationships between words in a text, state the authors, who are coauthors of the book, Human + Machine:Reimagining Work in the Age of AI (Harvard Business Review Press). 

Despite flaws in GPT-3 including producing nonsense or biased responses and generating plausible but false content, “A new age of AI is upon us,” the authors state 

Microsoft, Google, Alibaba, and Facebook are all working on their own version of “advanced transformers.” The tools will be trained in the cloud and accessible via APIs. “Companies that want to harness the power of next generation AI will shift their compute workloads from legacy to cloud-AI services like GPT-3,” the authors suggest.  

This will enable a new class of enterprise applications that will make the process of synthesizing words and information in language cheaper. Based on an analysis of more than 50 business-relevant proof of concept applications of GPT-3, the authors see three broad categories linked to language understanding: writing, coding and discipline-specific reasoning.  

For example, GPT-3 is capable of converting natural language to programming language. It can plot graphs based on verbal descriptions. One beta tester created a GPT-3 bot that enables people with no accounting skills to generate financial statements.   ... ' 

Saturday, April 03, 2021

Building Multilingual, Multipurpose, Multi Context Wikipedias

A long time user, and supporter of Wikipedia.  Have run into the problem that this describes,   Posts in different languages in the WP vary in their knowledge content.  In fact in too many examples,  there may be useful, detailed posts in one language, but the same topic is uncovered in another language. Within 50 million articles in 300 languages.   We ran into some related problems when we scoped out a wikipedia for a large corporation.  So I expanded the title of this piece, to include other aspects. But the article referenced below uses a 'language' knowledge problem example. Sharing knowledge.   Tough problem, let me know if you solve it effectively.

Building a Multilingual Wikipedia,   By Denny Vrandečić  in CACM

Communications of the ACM, April 2021, Vol. 64 No. 4, Pages 38-41  10.1145/3425778

Wikipedia has more than 50 million articles in approximately 300 languages. The content in these languages is independently created and maintained. The knowledge in Wikipedia is very unevenly distributed over the languages: some languages have more than a million articles, but more than 50 languages have only a few hundred articles or less. More importantly, also the number of contributors is very unevenly distributed: English Wikipedia has more than 418,000 contributors, the second-most active one, Spanish, drops down to 90,000. More than half of language editions have fewer than 10 contributors doing more than four edits per month. To assume that fewer than 10 active contributors can write and maintain a comprehensive encyclopedia in their spare time is optimistic at best.

In order to close these knowledge gaps we are building a multilingual Wikipedia where content is created only once but made available in all languages. The multilingual Wikipedia has two main components: Abstract Wikipedia where the content is created and maintained in a language-independent notation, and Wikifunctions, a project to create, catalog, and maintain functions. For the multilingual Wikipedia, the most important function is one that takes content from Abstract Wikipedia and renders it in natural language, which in turn gets integrated into Wikipedia proper.

This will considerably reduce the effort required to create a comprehensive and maintain a current encyclopedia in many languages. It will allow more people to share more knowledge in more languages than ever before. It will be particularly useful for under-served languages, providing an important way to help improve education and ready access to knowledge in many countries. ... "

Wednesday, February 10, 2021

Infants as a Model for Perceiving Language

Classic kind of approach, , can we have systems learn knowledge  like language?  Have followed the idea for many years. 

Scientists Develop Computational Approach to Understand How Infants Perceive Language  By News-Medical Life Sciences, February 5, 2021

Cognitive scientists and computational linguists have developed a quantitative modeling framework based on large-scale simulation of infants' language learning process.  A multi-institutional team of cognitive scientists and computational linguists has developed a quantitative modeling framework based on large-scale simulation of infants' language learning process.

The approach uses machine learning to enable the systematic linkage of learning mechanisms to testable predictions about infants' attunement to native language. The researchers trained a clustering algorithm on realistic speech input to model infants' language learning process; they fed the program spectrogram-like auditory features sampled at regular intervals obtained from naturalistic speech recordings in American English and Japanese.

This resulted in a candidate model for infants' early phonetic knowledge, which the team queried about observed differences in how Japanese- and English-learning infants discriminate speech sounds, as well as vowel- and consonant-like phonetic categories. The model yielded positive and negative outcomes for these respective queries, suggesting current literature on early phonetic learning requires a dramatic rethink.... ' 

Monday, February 08, 2021

Spelling in the MS AI Research Blog

 Had not thought the idea of capturing spellings in multiple languages was important, but this piece makes the point.

Microsoft details Speller100, an AI system that checks spelling in over 100 languages Kyle Wiggers  @Kyle_L_WiggersFebruary 8, 2021 9:05 AM

In a post on its AI research blog, Microsoft today detailed a new language system, Speller100, that the company claims is one of the most comprehensive ever made in terms of language coverage and accuracy. Comprising a number of machine learning models that can understand speech in over 100 languages collectively, Speller100 now powers spelling correction on Bing.

As Microsoft notes, for a language with very little web presence, it’s challenging to collect an adequate amount of data to train a model. Moreover, models can’t rely solely on training data to learn the spelling of a language. At its core, spelling correction is about building both an error and a language model, and not all errors are the same. For example, non-word errors occurs when a word isn’t in the vocabulary for a given language, while real-word errors occur when the word exists but doesn’t fit in a larger context.  ... "

See:  https://www.microsoft.com/en-us/research/project/speller/ 

Saturday, January 23, 2021

Translating Lost Languages Using ML

 Most interesting, lots more at the link.  Via patterns of association in other languages.   Might this be used in ways to link with associations between other language style 'patterns'?  What are the assumptions regarding the forms of the languages?   Ideas?   Plan to pass this along to people at our  'Language Lab'.  

Translating Lost Languages Using ML

MIT News, By Adam Conner-Simons. October 21, 2020

Researchers at the Massachusetts Institute of Technology (MIT) have developed a machine learning system that can automatically translate a lost language, without advanced knowledge of its relationship to other dialects. The system applies principles based on historical linguistic insights, including the fact that languages generally evolve in certain predictable patterns. MIT's Regina Barzilay and Jiaming Luo developed a decipherment algorithm that can segment words in an ancient language and map them to words in related languages. The algorithm infers relationships between languages, and can assess proximity between languages.