0%
BuildUnder the hoodabout 29 min, 8 steps

Stacking Neurons Into Layers

Follow a two-input example through a hidden layer and understand why nonlinear activations matter.

A case one threshold cannot separate

This optional reading follows on from the neuron reading, and uses the same Neurons blocks.

Consider two binary inputs, each either zero or one. We want output one when exactly one input is one, and zero when both are the same. This is called exclusive OR, or XOR.

First inputSecond inputDesired output
000
101
011
110

Placed on a two-dimensional grid, the positive cases sit at opposite corners. A single straight boundary cannot separate them from the two negative corners. That is a representational limitation, not a shortage of training repetitions.

Build with me · 1

Define the four allowed cases

Each input is either zero or one, so there are four possible pairs. The desired outputs in this order are false, true, true, false. Exactly one active input is different from at least one active input.

On a two-dimensional grid, the true cases occupy opposite corners. One straight separating boundary cannot put both of them on the true side and the other two on the false side.

Blocks at this stageWorked example
make an empty list calledcases
add
a list of0, 0
tocases
add
a list of1, 0
tocases
add
a list of0, 1
tocases
add
a list of1, 1
tocases
for eachfeaturesin listcases
sayInput pairthen
features

Build the four input lists and print them with a loop. Write the intended XOR result beside each before constructing neurons.

What to look for

All four pairs print once.

Make it yours

Explain why [1, 1] must be false for XOR even though at least one input is active.

Build intermediate features

Use two threshold neurons in a first layer. Neuron A has weights [1, 1] and threshold 1, so it fires when at least one input is one. Neuron B has the same weights and threshold 2, so it fires only when both are one.

Their outputs become the next layer's inputs. Use a final neuron with weights [1, -2] and threshold 1. It fires when A is on and B is off.

For input [1, 0], A produces 1 and B produces 0. The final weighted sum is 1×1 + 0×(-2) = 1, so output is 1. For [1, 1], A and B both produce 1; the final total is 1 - 2 = -1, so output is 0. For [0, 0], both produce zero and the final output is zero.

These are hand-designed parameters for a complete worked example. We have not claimed a training algorithm discovered them.

Build with me · 2

Detect at least one active input

Weights one and one add the two inputs. A threshold of one fires for totals one or two. This hidden unit therefore detects at least one active input, which is useful but not yet XOR.

A hidden feature is an intermediate value inside the network. Hidden means between input and output, not impossible to inspect. We can print it directly.

Blocks at this stageWorked example
make a neuron calledat_least_onewith weights
a list of1, 1
and threshold1
sayOne activethen
at_least_onefires for
a list of1, 0
sayBoth activethen
at_least_onefires for
a list of1, 1
sayNeither activethen
at_least_onefires for
a list of0, 0

Create at_least_one and inspect it on the three distinct total cases.

What to look for

One active and Both active are True; Neither active is False.

Make it yours

Predict the result for [0, 1] and explain why this neuron treats it like [1, 0].

Build with me · 3

Detect the case XOR must exclude

The second unit uses the same weights but raises the threshold to two. It fires only when both inputs are one. This detects the special case that must be excluded from XOR's result.

A layer can compute several different combinations or decisions from the same original inputs. Having the same inputs does not require its units to produce the same features.

Blocks at this stageWorked example
make a neuron calledbothwith weights
a list of1, 1
and threshold2
sayOne activethen
bothfires for
a list of1, 0
sayBoth activethen
bothfires for
a list of1, 1

Create both with threshold 2 and inspect the two cases.

What to look for

One active is False and Both active is True.

Make it yours

Explain how changing only a threshold gave this unit a different job.

Build with me · 4

Pass intermediate outputs forward

For input [1, 0], the first unit is true and the second false. These can be interpreted numerically as one and zero for the next weighted calculation. The later layer receives these hidden outputs, not the original pair.

This change of representation is the useful work of the hidden layer. It separates at-least-one from both so the output can combine them.

Blocks at this stageWorked example
make a neuron calledat_least_onewith weights
a list of1, 1
and threshold1
make a neuron calledbothwith weights
a list of1, 1
and threshold2
setfeaturesto
a list of1, 0
setato
at_least_onefires for
features
setbto
bothfires for
features
sayHidden Athen
a
sayHidden Bthen
b

Store the original features once, then assign a and b from the two neurons using that same feature value.

What to look for

Hidden A is True and Hidden B is False.

Make it yours

Change original features to [1, 1] and predict both hidden results before running.

What a hidden layer means

The intermediate outputs are hidden features: values inside the network rather than the original input or final answer. Hidden does not mean secret or impossible to inspect. It means the layer lies between input and output.

Every unit in a dense layer receives the previous layer's outputs, with its own weights and bias. A layer can therefore compute several different combinations of the same information. The following layer combines those results.

For digit recognition, a dense network can take brightness values, compute hidden features, and produce ten class scores. Training adjusts the numerical parameters rather than asking you to name what each hidden unit detects.

Build with me · 5

Combine the hidden features

The final weighted sum is A minus two times B. A contributes support; B vetoes the both-active case strongly enough to move the total below the final threshold of one.

The blocks express that output neuron using arithmetic and a comparison so every operation remains visible. True behaves as one and False as zero in this numeric calculation, exactly matching the binary hidden-feature interpretation.

Blocks at this stageWorked example
make a neuron calledat_least_onewith weights
a list of1, 1
and threshold1
make a neuron calledbothwith weights
a list of1, 1
and threshold2
setfeaturesto
a list of1, 0
setato
at_least_onefires for
features
setbto
bothfires for
features
setfinal_totalto
a
-
2x
b
sayHidden Athen
a
sayHidden Bthen
b
sayFinal totalthen
final_total
sayXOR outputthen
final_total
>=1

Set final_total to a minus two times b, then compare it with one. Keep both hidden outputs printed beside the final result.

What to look for

For [1, 0], the final total is 1 and XOR output is True.

Make it yours

Calculate the final total for hidden [1, 1] before loading the next stage.

Build with me · 6

Trace the negative contribution

Both hidden units fire for [1, 1]. The final sum is one minus two, or negative one. It does not reach the threshold of one, so XOR is false.

The negative output weight is essential to this design. Merely stacking another positive sum would not exclude the both-active case. The nonlinear thresholds also matter; a stack of only linear transformations would collapse to a linear transformation.

Blocks at this stageWorked example
make a neuron calledat_least_onewith weights
a list of1, 1
and threshold1
make a neuron calledbothwith weights
a list of1, 1
and threshold2
setfeaturesto
a list of1, 1
setato
at_least_onefires for
features
setbto
bothfires for
features
setfinal_totalto
a
-
2x
b
sayHidden Athen
a
sayHidden Bthen
b
sayFinal totalthen
final_total
sayXOR outputthen
final_total
>=1

Change the original pair to [1, 1] and trace a, b, final_total, and the comparison.

What to look for

Hidden A and B are True, Final total is -1, and XOR output is False.

Make it yours

Remove the factor two as a thought experiment. Explain which threshold and total relationships would need checking rather than assuming the new network is equivalent.

Build with me · 7

Trace the all-zero path

Neither hidden unit fires when both inputs are zero. The final sum is zero, which is below the output threshold. This completes the other false case without adding a special if statement for the original pair.

The same fixed parameters process each input. Forward prediction changes values flowing through the network while keeping the weights and thresholds fixed.

Blocks at this stageWorked example
make a neuron calledat_least_onewith weights
a list of1, 1
and threshold1
make a neuron calledbothwith weights
a list of1, 1
and threshold2
setfeaturesto
a list of0, 0
setato
at_least_onefires for
features
setbto
bothfires for
features
setfinal_totalto
a
-
2x
b
sayHidden Athen
a
sayHidden Bthen
b
sayFinal totalthen
final_total
sayXOR outputthen
final_total
>=1

Use [0, 0] with the same network and inspect the full trace.

What to look for

Both hidden values are False, Final total is 0, and XOR output is False.

Make it yours

Explain how a training update would differ from this forward computation: which quantities would it change?

Trace every path through the tiny network

For XOR, keep the two hidden decisions and final threshold fixed. This complete table makes it possible to check every allowed input, rather than relying on one successful example.

Original inputsA: at least oneB: bothFinal total A - 2BFinal output
[0, 0]0000
[1, 0]1011
[0, 1]1011
[1, 1]11-10

The final neuron receives [A, B], not the original input pair. That change of representation is the useful work of the hidden layer. The network first computes intermediate properties, then combines them to solve the task.

During this forward calculation, values move from input to output and the weights stay fixed. During training, a suitable learning procedure uses mistakes to change parameters. These are separate passes with separate jobs. A diagram showing arrows between layers normally describes the flow of values; it does not mean every prediction is another training update.

Build with me · 8

Verify the complete truth table

The final program evaluates every allowed input. Because the domain has only four cases, checking all of them verifies this manually designed XOR behaviour completely. Larger image or language tasks have too many possible inputs for that kind of exhaustive table.

Dense networks extend the idea to many weighted combinations and nonlinear activations, with training adjusting parameters. More layers can add useful representational power and also more fitting cost or overfitting risk; depth alone does not guarantee better evaluation results.

Blocks at this stageWorked example
make a neuron calledat_least_onewith weights
a list of1, 1
and threshold1
make a neuron calledbothwith weights
a list of1, 1
and threshold2
make an empty list calledcases
add
a list of0, 0
tocases
add
a list of1, 0
tocases
add
a list of0, 1
tocases
add
a list of1, 1
tocases
for eachfeaturesin listcases
sayOriginal inputthen
features
setato
at_least_onefires for
features
setbto
bothfires for
features
setfinal_totalto
a
-
2x
b
sayHidden Athen
a
sayHidden Bthen
b
sayFinal totalthen
final_total
sayXOR outputthen
final_total
>=1

Put the hidden calculations and output trace inside for each while creating the two neurons only once above the loop.

What to look for

The XOR outputs in case order are False, True, True, False.

Make it yours

Explain the input layer, hidden layer, and output calculation using actual values from one true and one false case.

Nonlinearity is essential

If you stack only linear weighted sums with no nonlinear activation between them, the combined computation is still a linear transformation. Extra layers alone do not create the useful expressive change illustrated by XOR.

The threshold above is nonlinear. Practical networks commonly use ReLU or other activations so the model can learn richer relationships while supporting efficient gradient-based training. "It is the same vote repeated" is an incomplete explanation unless it includes this role of nonlinearity.

More layers or units also increase complexity and fitting cost. They can overfit limited data. Choose architecture using development evidence, compare against simpler models, and reserve a final test. Depth is a tool for representing patterns, not a promise of better predictions.

Full reference solution

This is the complete worked program. Try building it yourself first, then use this reference to find the first place your version behaves differently. The Python below is generated from these exact blocks; helper functions are included so its behaviour can be inspected.

make a neuron calledat_least_onewith weights
a list of1, 1
and threshold1
make a neuron calledbothwith weights
a list of1, 1
and threshold2
make an empty list calledcases
add
a list of0, 0
tocases
add
a list of1, 0
tocases
add
a list of0, 1
tocases
add
a list of1, 1
tocases
for eachfeaturesin listcases
sayOriginal inputthen
features
setato
at_least_onefires for
features
setbto
bothfires for
features
setfinal_totalto
a
-
2x
b
sayHidden Athen
a
sayHidden Bthen
b
sayFinal totalthen
final_total
sayXOR outputthen
final_total
>=1
PythonHover over a line to see an explanation
# ---------------------------------------------------------------# Building blocks, written out in plain Python.# This part is generated for you. Your script starts further down.# ---------------------------------------------------------------  class Neuron:    """A weighted vote. Multiply each input by how much it counts, add it all    up, and fire only if the total clears the bar."""     def __init__(self, weights=None, threshold=0.0, name="neuron"):        self.weights = [float(w) for w in (weights or [])]        self.threshold = float(threshold)        self.name = name     def total(self, inputs):        values = inputs if isinstance(inputs, (list, tuple)) else [inputs]        total = 0.0        for i in range(min(len(values), len(self.weights))):            total += float(values[i]) * self.weights[i]        return round(total, 4)     def fires(self, inputs):        return self.total(inputs) >= self.threshold     def show(self):        print(            self.name            + ": weights "            + str(self.weights)            + ", fires at "            + str(self.threshold)            + " or more"        )  def new_neuron(weights=None, threshold=0.0, name="neuron"):    return Neuron(weights, threshold, name)  # ---------------------------------------------------------------# Your script# ---------------------------------------------------------------  at_least_one = new_neuron([1, 1], 1, "at_least_one")both = new_neuron([1, 1], 2, "both")cases = []cases.append([0, 0])cases.append([1, 0])cases.append([0, 1])cases.append([1, 1])for features in cases:    print("Original input", features)    a = at_least_one.fires(features)    b = both.fires(features)    final_total = (a - (2 * b))    print("Hidden A", a)    print("Hidden B", b)    print("Final total", final_total)    print("XOR output", (final_total >= 1))

Compare this with your version. Different names and personal choices are fine when the program follows the same logic.

Keep your progress

Sign in and every reading, quiz, and exercise you finish is saved.

Sign in