What Are Large Language Models, and How Do They Actually Work?
The previous article explained neural networks and deep learning. A large language model (LLM) — the technology behind ChatGPT, Claude, Gemini, and similar tools — is a deep neural network built for one specific job: predicting what text comes next.
Tokens, not words
LLMs don't process whole words the way people do. Text is first broken into tokens — chunks that might be a whole short word, part of a longer word, or even just punctuation. "Understanding" might become two tokens, "under" and "standing," for instance. Everything the model reads and writes is really a sequence of these tokens, converted to numbers before any calculation happens.
The actual task: predict the next token
At its core, an LLM does one thing repeatedly: given the tokens so far, predict which token is most likely to come next. That's it. To answer a question, it predicts the next token, adds that token to the text, then predicts the next one after that, and so on — one token at a time — until it produces a full response.
This sounds far too simple to explain conversations, code, or explanations like this one. The reason it works is scale: the model was trained on a huge fraction of publicly available text — books, articles, code, forums — and, to get good at next-token prediction across all of that, it had to internalize an enormous amount about grammar, facts, reasoning patterns, and writing style. Predicting text well, it turns out, requires something that behaves a lot like understanding, even though the training objective never explicitly mentions "understanding" at all.
Why context matters: attention
Predicting the next token well requires weighing which earlier words matter most. In the sentence "The trophy didn't fit in the suitcase because it was too big," figuring out what "it" refers to requires connecting it back to "trophy," not "suitcase." The mechanism that lets LLMs do this — called attention — lets every token in the input look at every other token and learn how relevant each one is to it. This is the specific neural network design (a "transformer") that made today's LLMs practical, and it's why LLMs can handle much longer, more coherent passages of text than earlier language models could.
Where the "training" happens
Building an LLM roughly happens in two stages:
- Pretraining — the model is trained purely on next-token prediction across a massive text corpus. This is where it picks up language, facts, and patterns, but it isn't yet especially good at being a helpful assistant — left alone, a purely pretrained model tends to just continue text in whatever style it started in.
- Fine-tuning / alignment — the model is further trained, often using human feedback on which responses are more helpful or accurate, to behave like a cooperative assistant that answers questions directly rather than just continuing text patterns.
The catch: it predicts plausible, not verified
Because the underlying task is "what token is statistically likely to come next," an LLM has no built-in mechanism for checking whether a plausible-sounding sentence is actually true. It can produce a confident, well-formatted, completely incorrect answer — a fabricated citation, a wrong date, a function that doesn't exist — with exactly the same fluency as a correct one. This behavior has a name, hallucination, and it's a direct consequence of how these models work rather than a rare bug. It's part of why this site's tool and article pages carry an as-is disclaimer: always verify anything that matters.
What's next
Knowing roughly how LLMs generate text also explains why how you ask changes what you get back — the model is continuing a pattern, so the pattern you start it with matters a lot. The next article, Prompt Engineering Basics, covers practical ways to get better, more reliable answers out of AI chatbots.
Related articles
Prompt Engineering Basics: How to Get Better Answers from AI Chatbots
Since an LLM continues whatever pattern you give it, how you ask shapes what you get back. Practical, concrete ways to write better prompts.
Neural Networks Explained: The Idea Behind Deep Learning
Neural networks are the specific kind of machine learning model behind almost every recent AI breakthrough. Here's how they're structured and why 'deep' learning works.
Generative AI Explained: How It Creates Content — and Why It Sometimes Gets Things Wrong
Generative AI covers far more than chatbots — image, audio, and video generators work on a similar underlying principle. Here's how they work, and why AI output still needs a human check.