4. Output: prediction vs target
Black = trained network output. Dashed = desired target. Light gray = input pulse.
Browser-only training: HTML + CSS + JavaScript. No PyTorch, no TensorFlow, no automatic differentiation.
We numerically discretize the continuous-time equation with Euler's method. BPTT then differentiates through this entire sequence of Euler steps.
Input channel x₁ receives a pulse from 1 to 2 s. The desired response appears later on output channel y₁, from 3 to 5 s. Any additional input and output channels are present but set to zero for this task.
h(t) influences h(t+dt), which influences later states. Therefore an output error at a late time can be caused by recurrent activity much earlier in the trial. BPTT follows these causal links backward through the unrolled Euler steps.
Black = trained network output. Dashed = desired target. Light gray = input pulse.
MSE should generally decrease as the weights learn the task.
These four curves are the learned internal recurrent state h₁(t), …, h₄(t).
PCA projects the 4D hidden trajectory into PC1 and PC2. Green = trial start, black = trial end.
Blue = positive weight, orange = negative weight, line thickness = |weight|. These lines change during training because SGD changes the matrices.
These are the current learned parameters. The values update during training, so after training you can inspect exactly what the network learned.
Input → hidden weights
Hidden → hidden recurrent weights
Hidden → output weights
Hidden and output bias terms
The code computes these derivatives explicitly. There is no autograd engine.
The current setup keeps the task fixed while allowing the architecture to change. The delayed-response task remains fixed so architecture changes can be compared cleanly. Later versions can expose task timing, activation, optimizer, multiple trials and task families.