← Back to the live visualizerprograminds

How neural networks actually work

This page is deliberately separate from the live tool above. Every demo below computes real math live in your browser — real weighted sums, real convolutions, real softmax — but the specific numbers (weights) are illustrative, not trained on anything, since we don't have access to Llama's actual trained parameters (no hosted API exposes that). Where a demo genuinely represents what the live tool's model does architecturally, we say so explicitly.

1 · The building block

Artificial Neuron (ANN)

A neuron takes in some numbers, decides how much each one matters, adds them up, and turns that total into a single output. That's the whole idea — every neural network, no matter how huge, is just millions of this same simple step wired together.

  • Each input (x1, x2) gets multiplied by a weight — a number saying how much that input should count. A bigger weight means more influence on the result.
  • The bias (b) is a small nudge added to the total, letting the neuron lean toward a higher or lower output even before it sees any input.
  • The activation function (sigmoid, here) squashes that total into a smooth 0–1 range, so the output reads like a confidence score instead of an unbounded number.
×w1×w2+bx1x2Σσoutputweighted sumactivation

Drag the sliders below and watch each real number move through every step above.

z = w1·x1 + w2·x2 + b
z = (0.60 × 0.80) + (-0.40 × 0.30) + 0.10
z = 0.460

output = sigmoid(z) = 1 / (1 + e⁻ᶻ)
output = 0.613

2 · Spatial patterns

Convolutional Neural Network (CNN)

Imagine sliding a small stencil across a photo, checking at every position: does the patch underneath match the pattern this stencil is looking for? That's exactly what a CNN's kernel does — a tiny 3×3 grid of numbers (a real one, below) slides across the image and scores one position at a time, the same small stencil reused everywhere instead of a separate neuron wired to every single pixel.

Different kernels look for different things — one might light up on sharp edges, another on smooth color changes — and stacking layers of them lets the network build up from simple patterns (an edge here, a curve there) to whole recognizable shapes.

Try the presets below, or edit the kernel's numbers directly and watch which patterns light up.

Input
3×3 kernel (editable)
Output (real convolution)
3 · Sequences and memory

Recurrent Neural Network (RNN)

Imagine reading a sentence one word at a time, and after every word you update a short mental summary of everything you've read so far. That running summary is the hidden state — an RNN does exactly this. The same small cell runs at every step, taking in the next input plus its own previous summary, and produces an updated summary to carry forward.

That's the whole idea: what happens at step 3 depends on what happened at steps 1 and 2, not just on that step's input alone — something the plain neuron above can't do, since it only ever looks at whatever inputs you hand it right now.

h₀h₁h₂h₃outputx1RNN cellx2RNN cellx3RNN cellsame cell, reused at every step — the weights don't change, only the hidden state does

Step through the sequence below and watch the real hidden-state number change at each step.

h₀ (start)0.000
→
after x1 = 0.5?
→
after x2 = -0.3?
→
after x3 = 0.8?
→
after x4 = 0.2?
→
after x5 = -0.6?
h0 = tanh(Wx·x + Wh·h0 + b)
4 · What our live tool actually uses

Transformer & Self-Attention

Read this: “The trophy didn't fit in the suitcase because it was too big.” To understand what “it” means, you instinctively glance back and weigh which earlier word fits better — trophy, or suitcase? That instinctive glancing-back-and-weighing is exactly what self-attention does: every word looks at every other word in the sentence and scores how relevant it is. High-scoring words shape the result more; low-scoring ones are mostly ignored.

This is the real mechanism behind the live visualizer above — neither CNN nor RNN powers it. Llama (the model behind it) is a Transformer, and self-attention is its core move, repeated across every layer. Type a sentence below: it runs through the same real tokenizer the live tool uses, and real attention scores get computed on the spot for every word pair.

thecatsatonthemat
the
0.67
cat
0.16
0.31
sat
0.18
0.19
0.36
0.17
on
0.18
0.52
the
0.21
0.48
mat
0.27
0.37

Rows = the token “asking”, columns = the tokens it attends to. Each row sums to 1 (it's a real softmax) — brighter cells mean that token contributed more to the row token's updated representation.

the → new vector, |v|=2.37cat → new vector, |v|=0.97sat → new vector, |v|=1.24on → new vector, |v|=2.03the → new vector, |v|=1.30mat → new vector, |v|=1.42

This is the closest we can get to genuine transformer internals without self-hosting a model with direct weight access (Groq's API, like every hosted inference API, never exposes real attention weights) — that's a planned, separate, heavier build, not something any hosted API can provide.