sangsen
Neural Networks Explained: The Idea Behind Deep Learning
AI

Neural Networks Explained: The Idea Behind Deep Learning

The previous article covered machine learning in general: software that improves from examples instead of hand-written rules. Neural networks are one specific way of building that kind of model — and right now, they're the way behind almost everything people call "AI."

The basic building block

A neural network is made of simple units called neurons, loosely inspired by (but far simpler than) neurons in a brain. Each artificial neuron does something modest: it takes several numbers in, multiplies each by a "weight" that says how important that input is, adds them up, and passes the result through a simple function before sending it onward.

On its own, one neuron can't do much. The power comes from arranging thousands or billions of them into layers, where each layer's output feeds into the next layer's input. An input layer receives the raw data (pixel values for an image, or word fragments for text), one or more hidden layers in the middle transform it, and an output layer produces the final answer (a label, a probability, the next word in a sentence).

Diagram of a neural network's layers: input, hidden, and output, connected by weighted lines

What makes it "learn"

When a neural network is created, its weights start out essentially random, so its first guesses are close to useless. Training adjusts those weights using an algorithm called backpropagation: the network makes a prediction, that prediction is compared to the correct answer, and the difference (the "error") is used to work out how much each individual weight in every layer contributed to the mistake. Each weight then gets nudged slightly in the direction that would have reduced the error. Repeat this millions of times across millions of examples, and the weights gradually settle into values that make correct predictions likely.

No one hand-designs what any individual weight should be — the training process finds them automatically. This is also why neural networks are hard to fully explain from the inside: the "reasoning" is spread across billions of numbers adjusted by a training process, not written as readable logic.

Why "deep"

"Deep learning" simply means a neural network with many hidden layers stacked on top of each other, rather than just one or two. Depth matters because each layer can build on the patterns the previous layer found — in an image-recognition network, early layers might learn to detect edges and simple textures, middle layers combine those into shapes, and later layers combine shapes into recognizable objects. This layered composition is what lets deep networks handle far more complex patterns than earlier, shallower models could.

Why now, not decades ago

The core math behind neural networks and backpropagation has existed since the 1980s. Three things had to come together before deep learning became practical at today's scale:

  • Data — the internet produced enormous amounts of text, images, and other content to train on.
  • Compute — GPUs made it fast (and affordable) to do the huge number of repeated calculations training requires.
  • Technique refinements — better ways of initializing weights, structuring layers, and keeping very deep networks stable during training, developed through years of research.

What's next

Scale one particular kind of neural network — built around a mechanism called "attention," which lets it weigh the relevance of different words to each other — up to billions of parameters and train it on a large fraction of all written text, and you get a large language model: the technology behind ChatGPT and similar tools. That's the subject of the next article, What Are Large Language Models, and How Do They Actually Work?

Related articles