Start with the name
The name comes from a loose analogy to brains: a neural network is built from many simple units ("neurons") connected to each other in layers, similar to how brain cells connect. The analogy is loose — a neural network doesn't work like a real brain — but it explains where the name comes from.
What each "neuron" actually does
Each unit takes in numbers, multiplies them by learned weights, adds a bias, and usually applies a nonlinear activation function before sending a value onward. A compact version is: output = activation(weighted inputs + bias). The activation matters because stacking only linear calculations would still produce a linear function, regardless of how many layers were added.
Here is a tiny example. Suppose the inputs are 2 and 3, the weights are 0.5 and -0.25, and the bias is 0.1. The weighted sum is (2 × 0.5) + (3 × -0.25) + 0.1 = 0.35. A ReLU activation keeps positive values, so this unit outputs 0.35. No single unit is intelligent; useful behaviour comes from many learned transformations working together.
Where the intelligence actually comes from
A neural network usually has an input layer, one or more "hidden" middle layers, and an output layer. Data flows through the layers, getting transformed a little at each step. With enough layers and enough units — modern networks have billions of connections — this simple process can approximate extremely complex patterns: recognizing a face, translating a sentence, predicting the next word in a paragraph.
The capability is distributed across learned parameters and computations rather than stored in one special unit. Different architectures use layers in different ways, so a tidy rule such as "early layers learn grammar and later layers learn meaning" is too broad to treat as a fact.
How the network learns those adjustments
Every connection has a weight, and units commonly have a bias. Training compares predictions with a target using a loss function. Backpropagation calculates how each parameter contributed to that loss, and an optimizer updates the parameters in a direction intended to reduce future error. Repeating this process over many examples can produce a model that generalizes to new inputs, but low training error alone does not prove that it will generalize.
Why this design became so powerful
Neural networks aren't new — the basic idea dates back decades. What changed recently is scale: far more data, far more computing power, and a specific neural network design called a transformer that's especially good at handling long sequences like sentences and paragraphs. That combination is what produced today's large language models.
Neural network vs machine learning vs deep learning
These terms describe different levels of the same landscape:
- Artificial intelligence is the broad goal of making software perform tasks associated with intelligence.
- Machine learning is a way to pursue that goal by learning patterns from data.
- A neural network is one family of machine-learning models.
- Deep learning uses neural networks with multiple learned layers.
Every deep-learning system uses a neural network, but machine learning also includes methods such as decision trees and linear regression. Not every machine-learning problem needs a neural network.
Common types of neural networks
The basic ingredients stay similar, but the connections change to suit the data:
- Feed-forward networks move information from input to output and suit many prediction tasks.
- Convolutional neural networks are designed to detect local patterns and became well known through image recognition.
- Recurrent networks process sequences while carrying information forward, though transformers now replace them in many language tasks.
- Transformers use attention to relate different parts of a sequence and power modern large language models.
The architecture does not guarantee quality. The training data, objective, evaluation and deployment conditions still determine whether the result is useful.
Where to go deeper
The AI Glossary has quick definitions for related terms like parameter and transformer. For a full walkthrough of how this connects to how AI actually learns, see How Does AI Actually Work? and the AI Fundamentals course.
What a layer actually does
Descriptions of neural networks usually mention layers and move on, which leaves the most interesting part unexplained. Here is the useful version.
Each layer takes numbers in and puts different numbers out. What makes it useful is that early layers detect simple things and later layers combine those into complicated things. In an image network, the first layer might respond to edges at various angles. The next combines edges into corners and curves. The next combines those into shapes, then into textures, then eventually into objects.
Nobody manually programs those detectors. Training shapes internal representations that help reduce the chosen loss on the available data. Researchers can inspect activations and identify patterns, but those interpretations depend on the architecture, task, dataset, and individual unit; they are not a universal layer-by-layer dictionary.
Why they need so much data
A network starts with random settings, meaning its first predictions are noise. Training shows it an example, measures how wrong it was, and nudges every setting slightly in the direction that would have been less wrong. Then it does that again. Millions of times.
The optimizer's learning rate controls the size of parameter updates. Updates that are too large can make training unstable, while updates that are too small can make it slow. Large datasets are useful for a different reason too: they expose the model to more of the variation it must handle instead of letting it memorize a narrow sample. Data quality, model size, regularization, and the task all affect how much data is needed.
What they are bad at
Being honest about the limits matters more than the mechanics.
- Explaining themselves. A network can be correct without any accessible reason. This is a genuine problem in medicine, lending, and hiring, where the reason matters as much as the answer.
- Anything unlike their training data. They interpolate confidently and extrapolate badly, and they give no signal about which one they are doing.
- Inheriting bias. Patterns in the training data are learned faithfully, including the ones nobody wanted.
- Knowing when they are wrong. There is no internal flag for uncertainty. Confident and correct look identical from outside, which is the root of most AI failures people actually encounter.
Sources and further learning
- Google Machine Learning Crash Course: nodes and hidden layers
- Google Machine Learning Crash Course: activation functions
- Google Machine Learning Crash Course: backpropagation
- PyTorch tutorial: building models and activation functions
For how these networks became today's chat assistants, see How Does AI Actually Work?, or the AI Fundamentals course for the full picture. If you are deciding what to study next, use the realistic AI learning roadmap.