Stacking Neurons Into Layers
Follow a two-input example through a hidden layer and understand why nonlinear activations matter.
A case one threshold cannot separate
This optional reading follows on from the neuron reading, and uses the same Neurons blocks.
Consider two binary inputs, each either zero or one. We want output one when exactly one input is one, and zero when both are the same. This is called exclusive OR, or XOR.
| First input | Second input | Desired output |
|---|---|---|
| 0 | 0 | 0 |
| 1 | 0 | 1 |
| 0 | 1 | 1 |
| 1 | 1 | 0 |
Placed on a two-dimensional grid, the positive cases sit at opposite corners. A single straight boundary cannot separate them from the two negative corners. That is a representational limitation, not a shortage of training repetitions.
Build with me · 1
Define the four allowed cases
Each input is either zero or one, so there are four possible pairs. The desired outputs in this order are false, true, true, false. Exactly one active input is different from at least one active input.
On a two-dimensional grid, the true cases occupy opposite corners. One straight separating boundary cannot put both of them on the true side and the other two on the false side.
Build the four input lists and print them with a loop. Write the intended XOR result beside each before constructing neurons.
What to look for
All four pairs print once.
Make it yours
Explain why [1, 1] must be false for XOR even though at least one input is active.
Build intermediate features
Use two threshold neurons in a first layer. Neuron A has weights [1, 1] and threshold 1, so it fires when at least one input is one. Neuron B has the same weights and threshold 2, so it fires only when both are one.
Their outputs become the next layer's inputs. Use a final neuron with weights [1, -2] and threshold 1. It fires when A is on and B is off.
For input [1, 0], A produces 1 and B produces 0. The final weighted sum is 1×1 + 0×(-2) = 1, so output is 1. For [1, 1], A and B both produce 1; the final total is 1 - 2 = -1, so output is 0. For [0, 0], both produce zero and the final output is zero.
These are hand-designed parameters for a complete worked example. We have not claimed a training algorithm discovered them.
What a hidden layer means
The intermediate outputs are hidden features: values inside the network rather than the original input or final answer. Hidden does not mean secret or impossible to inspect. It means the layer lies between input and output.
Every unit in a dense layer receives the previous layer's outputs, with its own weights and bias. A layer can therefore compute several different combinations of the same information. The following layer combines those results.
For digit recognition, a dense network can take brightness values, compute hidden features, and produce ten class scores. Training adjusts the numerical parameters rather than asking you to name what each hidden unit detects.
Build with me · 5
Combine the hidden features
The final weighted sum is A minus two times B. A contributes support; B vetoes the both-active case strongly enough to move the total below the final threshold of one.
The blocks express that output neuron using arithmetic and a comparison so every operation remains visible. True behaves as one and False as zero in this numeric calculation, exactly matching the binary hidden-feature interpretation.
Set final_total to a minus two times b, then compare it with one. Keep both hidden outputs printed beside the final result.
What to look for
For [1, 0], the final total is 1 and XOR output is True.
Make it yours
Calculate the final total for hidden [1, 1] before loading the next stage.
Build with me · 6
Trace the negative contribution
Both hidden units fire for [1, 1]. The final sum is one minus two, or negative one. It does not reach the threshold of one, so XOR is false.
The negative output weight is essential to this design. Merely stacking another positive sum would not exclude the both-active case. The nonlinear thresholds also matter; a stack of only linear transformations would collapse to a linear transformation.
Change the original pair to [1, 1] and trace a, b, final_total, and the comparison.
What to look for
Hidden A and B are True, Final total is -1, and XOR output is False.
Make it yours
Remove the factor two as a thought experiment. Explain which threshold and total relationships would need checking rather than assuming the new network is equivalent.
Build with me · 7
Trace the all-zero path
Neither hidden unit fires when both inputs are zero. The final sum is zero, which is below the output threshold. This completes the other false case without adding a special if statement for the original pair.
The same fixed parameters process each input. Forward prediction changes values flowing through the network while keeping the weights and thresholds fixed.
Use [0, 0] with the same network and inspect the full trace.
What to look for
Both hidden values are False, Final total is 0, and XOR output is False.
Make it yours
Explain how a training update would differ from this forward computation: which quantities would it change?
Trace every path through the tiny network
For XOR, keep the two hidden decisions and final threshold fixed. This complete table makes it possible to check every allowed input, rather than relying on one successful example.
| Original inputs | A: at least one | B: both | Final total A - 2B | Final output |
|---|---|---|---|---|
| [0, 0] | 0 | 0 | 0 | 0 |
| [1, 0] | 1 | 0 | 1 | 1 |
| [0, 1] | 1 | 0 | 1 | 1 |
| [1, 1] | 1 | 1 | -1 | 0 |
The final neuron receives [A, B], not the original input pair. That change of representation is the useful work of the hidden layer. The network first computes intermediate properties, then combines them to solve the task.
During this forward calculation, values move from input to output and the weights stay fixed. During training, a suitable learning procedure uses mistakes to change parameters. These are separate passes with separate jobs. A diagram showing arrows between layers normally describes the flow of values; it does not mean every prediction is another training update.
Build with me · 8
Verify the complete truth table
The final program evaluates every allowed input. Because the domain has only four cases, checking all of them verifies this manually designed XOR behaviour completely. Larger image or language tasks have too many possible inputs for that kind of exhaustive table.
Dense networks extend the idea to many weighted combinations and nonlinear activations, with training adjusting parameters. More layers can add useful representational power and also more fitting cost or overfitting risk; depth alone does not guarantee better evaluation results.
Put the hidden calculations and output trace inside for each while creating the two neurons only once above the loop.
What to look for
The XOR outputs in case order are False, True, True, False.
Make it yours
Explain the input layer, hidden layer, and output calculation using actual values from one true and one false case.
Nonlinearity is essential
If you stack only linear weighted sums with no nonlinear activation between them, the combined computation is still a linear transformation. Extra layers alone do not create the useful expressive change illustrated by XOR.
The threshold above is nonlinear. Practical networks commonly use ReLU or other activations so the model can learn richer relationships while supporting efficient gradient-based training. "It is the same vote repeated" is an incomplete explanation unless it includes this role of nonlinearity.
More layers or units also increase complexity and fitting cost. They can overfit limited data. Choose architecture using development evidence, compare against simpler models, and reserve a final test. Depth is a tool for representing patterns, not a promise of better predictions.
Full reference solution
This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# --------------------------------------------------------------- class Neuron: """A weighted vote. Multiply each input by how much it counts, add it all up, and fire only if the total clears the bar.""" def __init__(self, weights=None, threshold=0.0, name="neuron"): self.weights = [float(w) for w in (weights or [])] self.threshold = float(threshold) self.name = name def total(self, inputs): values = inputs if isinstance(inputs, (list, tuple)) else [inputs] total = 0.0 for i in range(min(len(values), len(self.weights))): total += float(values[i]) * self.weights[i] return round(total, 4) def fires(self, inputs): return self.total(inputs) >= self.threshold def show(self): print( self.name + ": weights " + str(self.weights) + ", fires at " + str(self.threshold) + " or more" ) def new_neuron(weights=None, threshold=0.0, name="neuron"): return Neuron(weights, threshold, name) # ---------------------------------------------------------------# Your script# --------------------------------------------------------------- at_least_one = new_neuron([1, 1], 1, "at_least_one")both = new_neuron([1, 1], 2, "both")cases = []cases.append([0, 0])cases.append([1, 0])cases.append([0, 1])cases.append([1, 1])for features in cases: print("Original input", features) a = at_least_one.fires(features) b = both.fires(features) final_total = (a - (2 * b)) print("Hidden A", a) print("Hidden B", b) print("Final total", final_total) print("XOR output", (final_total >= 1))Compare this with your version. Different names and personal choices are fine when the program follows the same logic.
Keep your progress
Sign in and every reading, quiz, and exercise you finish is saved.