Skip to content

Make eval6 evaluate its RNN; fix multiclass_accuracy on sequences and RNN.__str__ - #15

Merged
vmspereira merged 4 commits into
mainfrom
fix/eval6-and-multiclass-accuracy
Jul 29, 2026
Merged

vmspereira merged 4 commits into
mainfrom
fix/eval6-and-multiclass-accuracy

Conversation

@vmspereira

@vmspereira vmspereira commented Jul 29, 2026 •

Copy link
Copy Markdown
Owner

eval6 trained an RNN and then stopped. No markdown cells at all, nothing after nn.fit(), and it imported multiclass_accuracy without ever calling it — so it never showed that the model had learned anything.

Chasing that turned up two code bugs, one of them significant.

PYTHONPATH=src python -m pytest tests/ -q
479 passed        (with cvxopt; 476 passed, 3 skipped without)
python -m flake8 src tests setup.py
0 violations

Every commit independently green. All nine notebooks execute.

multiclass_accuracy was broken for sequences

It reduced with axis=1, which is the class axis only for 2-D input. For a sequence model scoring (n_samples, seq_len, n_classes), axis=1 is time, so it took the argmax over timesteps and compared the wrong things:

perfect predictions on a (4, 5, 6) target  ->  6.0

An accuracy above 1 is not a fraction, so the result wasn't merely inaccurate — it was meaningless. axis=-1 is the class axis whatever the rank; ravel() makes each scored position contribute one comparison instead of a whole row. 2-D behaviour unchanged.

It was found precisely because eval6 imported it and never called it — the one notebook that would have exercised it. Had it been called, it would have printed a number above 1 beside a model that was in fact perfect.

RNN had no __str__

Every other layer defines one, so NN.__str__ printed a raw repr — which eval6 displayed:

<si.supervised.nn.rnn.RNN object at 0x110e9a7e0>     ->     RNN(10 units, bptt_trunc=5, input_shape=(10, 20))
SoftMax                                                     SoftMax

eval6 rewritten

It now explains the recurrence, shows the data as numbers rather than one-hot rows, plots the loss, evaluates on the held-out split, prints decoded predictions against targets, and feeds predictions back to continue a sequence.

Training drops from 20000 epochs to 60 — about two seconds instead of many minutes, with the loss curve showing where it actually flattens.

The data generator also had a real defect: it built targets by np.roll-ing the inputs and then setting y[:, -1, 1] = 1. roll wraps the first one-hot into the last position, so that assignment added a second 1 — the final target was two-hot, which is incoherent for a softmax objective. Targets are now built directly as start+1 .. start+10.

The part I think is most worth having

The model scores 1.000 at every position, and the notebook now demonstrates why that is a fact about the task, not the model. The target at step t is input[t] + 1 — a function of the current input alone — so a hand-built position-wise shift matrix scores 1.000 too:

a fixed position-wise shift scores: 1.0
the trained RNN scores:            1.0

The counting task exercises the RNN's plumbing but never its memory. So a second task is added that does require it — echo the first element at every position:

input     [3, 1, 0, 9, 2, 6, 2, 4]
predicted [3, 3, 3, 3, 3, 3, 3, 3]

A memoryless predictor scores 0.21 there; the RNN reaches 1.000 by carrying the first value in its hidden state. That then motivates the two genuine limits of recurrence — a single hidden state, and the bptt_trunc window — and points at attention in eval8 as the answer to both.

A wrong claim of my own, removed

My first draft of the notebook asserted accuracy would be lowest at position 0 "for want of history". Running it showed 1.000 everywhere — which is what led to the observation above. Position 0 is not hard, because the target never depended on history in the first place. The claim is gone.

Verification

Sequence-accuracy tests cover 2-D, 3-D and 4-D, that the result never exceeds 1, and that raw probabilities work as well as one-hot. Reverting multiclass_accuracy to axis=1 fails two tests; removing RNN.__str__ fails two more.

scripts/README.md updated: eval6 no longer belongs on the list of notebooks with long training cells.


Update: the RNN's weight gradients were never verified

Added in response to "is the RNN working?" — which turned out to be worth asking.

Its input gradient was fixed and verified earlier this session. Its weight gradients were not: the only test touching U, V and W asserted they change after a backward pass, which says nothing about whether the values are right. Same gap found suite-wide for Dense/LayerNorm/BatchNorm, left open here.

They are correct. Against central differences, comparing what the layer actually applied:

Gradient T=2 T=4 T=6
dE/dV 7.3e-11 1.7e-10 1.4e-10
dE/dU 1.5e-10 2.5e-10 2.1e-10
dE/dW 1.1e-10 3.3e-10 2.8e-10

This matters more for the weights than for the input: each of the three is shared across every timestep, so its gradient is a sum over the whole sequence — the part of BPTT most likely to be wrong, and the part a shape assertion cannot see.

A companion test asserts the other side, so the exact checks cannot pass trivially if truncation were ignored:

bptt_trunc (T=8) max relative error
8 (untruncated) 2.2e-10 — exact
4 6.7e-03 — approximate by design
1 8.2e-02 — approximate by design

Halving grad_U, grad_V or grad_W now fails a test. Before this, all three passed unnoticed.

Suite: 481 passed with cvxopt.

It reduced with axis=1, which is the class axis only for 2-D input. For a
sequence model scoring (n_samples, seq_len, n_classes), axis=1 is TIME, so it
took the argmax over timesteps and compared the wrong things:

    perfect predictions on a (4, 5, 6) target  ->  6.0

An accuracy above 1 is not a fraction, so the result was not merely inaccurate
but meaningless. axis=-1 is the class axis whatever the rank, and ravel() makes
each scored position contribute one comparison instead of a whole row. 2-D
behaviour is unchanged.

Found via eval6.ipynb, which imported this metric and never called it -- the one
notebook that would have exercised it. Had it been called, it would have printed
a number above 1 next to a model that was in fact perfect.

Tests cover 2-D, 3-D and 4-D, that the result never exceeds 1, and that raw
probabilities work as well as one-hot. Reverting to axis=1 fails two of them.
Every other layer defines one, so NN.__str__ -- which joins str(layer) over the
stack -- printed a raw object repr for the RNN:

    <si.supervised.nn.rnn.RNN object at 0x110e9a7e0>
    SoftMax

eval6.ipynb displayed exactly that. It now reads

    RNN(10 units, bptt_trunc=5, input_shape=(10, 20))
    SoftMax

which is the information a reader wants when inspecting a network.
…es not test

eval6 trained an RNN and then stopped. It had no markdown cells at all, nothing
after nn.fit(), and imported multiclass_accuracy without ever calling it -- so it
never showed that the model had learned anything.

It now explains the recurrence, shows the data as numbers rather than one-hot
rows, plots the loss, evaluates on the held-out split, prints decoded
predictions against targets, and feeds predictions back to continue a sequence.

Training drops from 20000 epochs to 60: about two seconds instead of many
minutes, with the loss curve showing where it actually flattens. The generator
now builds the targets directly as start+1..start+10 rather than np.roll-ing the
inputs and overwriting the last row, which produced a TWO-hot final target --
np.roll wrapped the first one-hot into the last position and `y[:, -1, 1] = 1`
then added a second 1, so the last target was incoherent for a softmax objective.

The more useful change is what the notebook now admits. The model scores 1.000 at
every position, and the notebook demonstrates why that is a fact about the TASK,
not the model: the target at step t is input[t] + 1, a function of the current
input alone, so a hand-built position-wise shift matrix scores 1.000 as well. The
counting task exercises the RNN's plumbing but not its memory.

So a second task is added that does need memory -- echo the first element of the
sequence at every position. A memoryless predictor scores 0.21 there; the RNN
reaches 1.000, carrying the first value forward in its hidden state:

    input     [3, 1, 0, 9, 2, 6, 2, 4]
    predicted [3, 3, 3, 3, 3, 3, 3, 3]

which then motivates the two real limits of recurrence -- one hidden state, and
the bptt_trunc window -- and points at attention in eval8 as the answer to both.

An earlier draft of this notebook claimed accuracy would be lowest at position 0
for want of history. Running it showed 1.000 everywhere, which is what led to the
observation above; the claim is gone.
Prompted by the question "is the RNN working?", which turned out to be worth
asking. Its INPUT gradient was fixed and verified earlier; its WEIGHT gradients
never were. The only test touching them asserted that U, V and W CHANGE after a
backward pass, which says nothing about whether the values are right -- the same
gap found suite-wide for Dense, LayerNorm and BatchNormalization, left open here.

They are correct. Against central differences, comparing what the layer actually
applied (recovered as (before - after)/lr with momentum-free SGD):

    dE/dV   7.3e-11 .. 1.4e-10
    dE/dU   1.5e-10 .. 2.5e-10
    dE/dW   1.1e-10 .. 3.3e-10

over sequence lengths 2, 4 and 6. This matters more for the weights than for the
input: each of the three is SHARED across every timestep, so its gradient is a
sum over the whole sequence -- the part of backpropagation-through-time most
likely to be wrong, and the part a shape assertion cannot see.

A companion test asserts the other side: with bptt_trunc shorter than the
sequence the weight gradient is a deliberate approximation (6.7e-03 relative
error at bptt_trunc=4, 8.2e-02 at 1, against 2.2e-10 untruncated). Without it the
exact checks above could pass trivially if truncation were silently ignored.

Halving grad_U, grad_V or grad_W now fails a test. Before this, all three passed
unnoticed.
@vmspereira
vmspereira merged commit 5250380 into main Jul 29, 2026
10 checks passed
@vmspereira
vmspereira deleted the fix/eval6-and-multiclass-accuracy branch July 29, 2026 10:10
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant