How Large Language Models Work: An Inside Look From Developers

Understanding the fundamentals of how large language models operate is crucial for anyone interested in their potential. Let's delve into how these systems generate human-like text, from the perspective of their creators.

How Large Language Models Work: An Inside Look From Developers

Sometimes, the best explanations for how a technological solution works come directly from the software engineers who built it. To understand how large language models (LLMs) function, we turned to the team responsible for their development and scaling. These engineers walked us through the process that happens under the hood of these impressive systems.

A leading AI research organization aims to safely bring advanced research insights to the world. Within it, a specialized group has focused from the beginning on creating and scaling these technologies, including the large language models now commonly available to both users and developers.

When you ask a language model a question, several key steps occur behind the scenes:

1. Input: The system receives your text from a text input.

2. Tokenization: The input text is divided into so-called "tokens." A token roughly corresponds to a few Unicode characters and can be thought of as a word or part of a word. It's important to note that tokenization doesn't necessarily split text into whole words; tokens can also be sub-words.

3. Embedding Creation: Each token is then converted into a vector of numbers, known as an embedding. Embeddings are at the heart of large language models and represent a multi-dimensional representation of a token. Models are trained to capture the semantic meaning and relationships between words or phrases. For example, the embeddings for "dog" and "puppy" are closer to each other in several dimensions than the embeddings for "dog" and "computer." These multi-dimensional representations help machines understand human language more effectively.

4. Multiplication of Embeddings by Model Weights: These embeddings are then multiplied by an enormous quantity of model weights, numbering in the hundreds of billions. This operation is extremely computationally intensive. Model weights are used to calculate a weighted matrix of embeddings, which is then utilized to predict the most probable next token.

Dominik Medal

Planning a new website or app?

Tell me what you have in mind — I will take a look and advise on the best way forward.

5. Prediction Sampling: After billions of multiplications, the resulting vector of numbers represents the probability of the most likely next token. Sampling is the process of selecting this most probable token and sending it back to the user. Every word the model generates is the result of this iterative process, occurring many times per second.

How is this complex set of model weights, whose values encode a large part of human knowledge, generated? This happens through a process called "pre-training." The goal is to create a model that can predict the next token for all words available on the internet. During pre-training, weights are gradually updated using a mathematical optimization method called gradient descent. We can imagine this like a hiker in dense fog on a mountain, trying to get down. Since they cannot see the entire mountain, they can only assess the steepness of the slope in their immediate vicinity and head in the direction of the steepest descent. The model gradually "measures" the steepness and learns how to adjust its weights to minimize prediction error.

Once we have a finished model, we can run "inference" on it, which is the process of providing a text query to the model. For example, a query might be: "write a short blog post about the benefits of agile development." The model then predicts the most probable next token (word). It performs this prediction based on the previous input and the text generated so far, and this repeats token by token, word by word, until it generates the complete response.

The functioning of these models is not magic, and it's worth understanding. The initial reaction to interacting with these systems is often a feeling that it's something magical. They can generate responses that feel human-like and have access to a vast amount of information. It's important to realize that large language models do not "think" or "understand" like humans. Instead, they generate words based on the statistical probability of what word should follow, considering the input and everything generated so far.

The basic principle is quite simple: start with a massive sample of human text (from the web, books, etc.) and then train a neural network to generate text that is "similar to that sample." Specifically, to be able to start from a "prompt" and continue with text that is "similar to what it was trained on." Although the actual neural network is composed of billions of very simple elements, and its fundamental operation is also very straightforward – essentially, it involves passing the input derived from the text generated so far "once through its elements" for each new word (or part of a word) – it is remarkable and unexpected that this process can produce text that is successfully "similar" to what is available on the web and in books. Even though models "merely" select a "coherent thread of text" from the "statistics of conventional wisdom" they have accumulated, the results are astonishingly human-like.

Dominik Medal

Let's talk about your project

Tell me what you need and together we will work out the best way forward.

Stop scrolling, call me

+420 735 505 585