/* ---- Google Analytics Code Below */
Showing posts with label LLMs. Show all posts
Showing posts with label LLMs. Show all posts

Friday, April 28, 2023

Map of Evolutionary Tree LLMs

Good resources,  gives an interesting indication  about how much has been done.

Yann Lecun     VP & Chief AI Scientist at Meta     (Technical)

A survey of LLMs with a practical guide and evolutionary tree.

Number of LLMs from Meta = 7

Number of open source LLMs from Meta = 7

The architecture nomenclature for LLMs is somewhat confusing and unfortunate.

What's called "encoder only" actually has an encoder and a decoder (just not an auto-regressive decoder).

What's called "encoder-decoder" really means "encoder with auto-regressive decoder"

What's called "decoder only" really means "auto-regressive encoder-decoder"

https://lnkd.in/eZKhwmuz    -   Automated map appears below: 

We build an evolutionary tree of modern Large Language Models (LLMs) to trace the development of language models in recent years and highlights some of the most well-known models, in the following figure:   ... 

Wednesday, April 19, 2023

Demystifying LLMs with Amazon distinguished scientists

 Good piece

Demystifying LLMs with Amazon distinguished scientists

April 18, 2023 • 2419 words

Werner, Sudipta, and Dan behind the scenes

Last week, I had a chance to chat with Swami Sivasubramanian, VP of database, analytics and machine learning services at AWS. He caught me up on the broad landscape of generative AI, what we’re doing at Amazon to make tools more accessible, and how custom silicon can reduce costs and increase efficiency when training and running large models. If you haven’t had a chance, I encourage you to watch that conversation.

Swami mentioned transformers, and I wanted to learn more about how these neural network architectures have led to the rise of large language models (LLMs) that contain hundreds of billions of parameters. To put this into perspective, since 2019, LLMs have grown more than 1000x in size. I was curious what impact this has had, not only on model architectures and their ability to perform more generative tasks, but the impact on compute and energy consumption, where we see limitations, and how we can turn these limitations into opportunities.

Diagram of transformer architecture

Transformers pre-process text inputs as embeddings. These embeddings are processed by an encoder that captures contextual information from the input, which the decoder can apply and emit output text.

Luckily, here at Amazon, we have no shortage of brilliant people. I sat with two of our distinguished scientists, Sudipta Sengupta and Dan Roth, both of whom are deeply knowledgeable on machine learning technologies. During our conversation they helped to demystify everything from word representations as dense vectors to specialized computation on custom silicon. It would be an understatement to say I learned a lot during our chat — honestly, they made my head spin a bit.

There is a lot of excitement around the near-infinite possibilites of a generic text in/text out interface that produces responses resembling human knowledge. And as we move towards multi-modal models that use additional inputs, such as vision, it wouldn’t be far-fetched to assume that predictions will become more accurate over time. However, as Sudipta and Dan emphasized during out chat, it’s important to acknowledge that there are still things that LLMs and foundation models don’t do well — at least not yet — such as math and spatial reasoning. Rather than view these as shortcomings, these are great opportunities to augment these models with plugins and APIs. For example, a model may not be able to solve for X on its own, but it can write an expression that a calculator can execute, then it can synthesize the answer as a response. Now, imagine the possibilities with the full catalog of AWS services only a conversation away.

Services and tools, such as Amazon Bedrock, Amazon Titan, and Amazon CodeWhisperer, have the potential to empower a whole new cohort of innovators, researchers, scientists, and developers. I’m very excited to see how they will use these technologies to invent the future and solve hard problems.

The entire transcript of my conversation with Sudipta and Dan is available below.

Now, go build!... '   


Thursday, April 13, 2023

Open Source Alternatives to ChatGPT and Bard

Nicely done piece which illustrates a number of existing tools , Open-Source examples, Several I have not heard of.  Instructive.

8 Open-Source Alternatives to ChatGPT and Bard   in KDNuggets

Discover the widely-used open-source frameworks and models for creating your ChatGPT like chatbots, integrating LLMs, or launching your AI product.

By Abid Ali Awan, KDnuggets on April 6, 2023 in Natural Language Processing   ... '

Sunday, January 08, 2023

Dense Visual Representations and More for Robotics

 Quite an interesting podcast I am about to listen to  ...   ideal applications?.

ACM OPINION

Dense Visual Representations, NeRFs, and LLMs for Robotics

By The Gradient, January 5, 2023

Google Research Scientist Pete Florence.

Pete Florence is a research scientist at Google Research on the Robotics at Google team inside Brain Team in Google Research.

As a Google research scientist, Pete Florence focuses on topics in robotics, computer vision, and natural language, including 3D learning, self-supervised learning, and policy learning in robotics. In an interview, Florence discusses his start in artificial intelligence, Ph.D. work with quadcopters, dense visual representations, NeRFs for robotics, language models for robotics, talking to robots in real time, and more.

From The Gradient

Listen to Podcast