Repository navigation
Make eval6 evaluate its RNN; fix multiclass_accuracy on sequences and RNN.__str__ - #15
Merged
Merged
Conversation
It reduced with axis=1, which is the class axis only for 2-D input. For a
sequence model scoring (n_samples, seq_len, n_classes), axis=1 is TIME, so it
took the argmax over timesteps and compared the wrong things:
perfect predictions on a (4, 5, 6) target -> 6.0
An accuracy above 1 is not a fraction, so the result was not merely inaccurate
but meaningless. axis=-1 is the class axis whatever the rank, and ravel() makes
each scored position contribute one comparison instead of a whole row. 2-D
behaviour is unchanged.
Found via eval6.ipynb, which imported this metric and never called it -- the one
notebook that would have exercised it. Had it been called, it would have printed
a number above 1 next to a model that was in fact perfect.
Tests cover 2-D, 3-D and 4-D, that the result never exceeds 1, and that raw
probabilities work as well as one-hot. Reverting to axis=1 fails two of them.
Every other layer defines one, so NN.__str__ -- which joins str(layer) over the
stack -- printed a raw object repr for the RNN:
<si.supervised.nn.rnn.RNN object at 0x110e9a7e0>
SoftMax
eval6.ipynb displayed exactly that. It now reads
RNN(10 units, bptt_trunc=5, input_shape=(10, 20))
SoftMax
which is the information a reader wants when inspecting a network.
…es not test
eval6 trained an RNN and then stopped. It had no markdown cells at all, nothing
after nn.fit(), and imported multiclass_accuracy without ever calling it -- so it
never showed that the model had learned anything.
It now explains the recurrence, shows the data as numbers rather than one-hot
rows, plots the loss, evaluates on the held-out split, prints decoded
predictions against targets, and feeds predictions back to continue a sequence.
Training drops from 20000 epochs to 60: about two seconds instead of many
minutes, with the loss curve showing where it actually flattens. The generator
now builds the targets directly as start+1..start+10 rather than np.roll-ing the
inputs and overwriting the last row, which produced a TWO-hot final target --
np.roll wrapped the first one-hot into the last position and `y[:, -1, 1] = 1`
then added a second 1, so the last target was incoherent for a softmax objective.
The more useful change is what the notebook now admits. The model scores 1.000 at
every position, and the notebook demonstrates why that is a fact about the TASK,
not the model: the target at step t is input[t] + 1, a function of the current
input alone, so a hand-built position-wise shift matrix scores 1.000 as well. The
counting task exercises the RNN's plumbing but not its memory.
So a second task is added that does need memory -- echo the first element of the
sequence at every position. A memoryless predictor scores 0.21 there; the RNN
reaches 1.000, carrying the first value forward in its hidden state:
input [3, 1, 0, 9, 2, 6, 2, 4]
predicted [3, 3, 3, 3, 3, 3, 3, 3]
which then motivates the two real limits of recurrence -- one hidden state, and
the bptt_trunc window -- and points at attention in eval8 as the answer to both.
An earlier draft of this notebook claimed accuracy would be lowest at position 0
for want of history. Running it showed 1.000 everywhere, which is what led to the
observation above; the claim is gone.
Prompted by the question "is the RNN working?", which turned out to be worth
asking. Its INPUT gradient was fixed and verified earlier; its WEIGHT gradients
never were. The only test touching them asserted that U, V and W CHANGE after a
backward pass, which says nothing about whether the values are right -- the same
gap found suite-wide for Dense, LayerNorm and BatchNormalization, left open here.
They are correct. Against central differences, comparing what the layer actually
applied (recovered as (before - after)/lr with momentum-free SGD):
dE/dV 7.3e-11 .. 1.4e-10
dE/dU 1.5e-10 .. 2.5e-10
dE/dW 1.1e-10 .. 3.3e-10
over sequence lengths 2, 4 and 6. This matters more for the weights than for the
input: each of the three is SHARED across every timestep, so its gradient is a
sum over the whole sequence -- the part of backpropagation-through-time most
likely to be wrong, and the part a shape assertion cannot see.
A companion test asserts the other side: with bptt_trunc shorter than the
sequence the weight gradient is a deliberate approximation (6.7e-03 relative
error at bptt_trunc=4, 8.2e-02 at 1, against 2.2e-10 untruncated). Without it the
exact checks above could pass trivially if truncation were silently ignored.
Halving grad_U, grad_V or grad_W now fails a test. Before this, all three passed
unnoticed.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
eval6 trained an RNN and then stopped. No markdown cells at all, nothing after
nn.fit(), and it importedmulticlass_accuracywithout ever calling it — so it never showed that the model had learned anything.Chasing that turned up two code bugs, one of them significant.
Every commit independently green. All nine notebooks execute.
multiclass_accuracywas broken for sequencesIt reduced with
axis=1, which is the class axis only for 2-D input. For a sequence model scoring(n_samples, seq_len, n_classes),axis=1is time, so it took the argmax over timesteps and compared the wrong things:An accuracy above 1 is not a fraction, so the result wasn't merely inaccurate — it was meaningless.
axis=-1is the class axis whatever the rank;ravel()makes each scored position contribute one comparison instead of a whole row. 2-D behaviour unchanged.It was found precisely because eval6 imported it and never called it — the one notebook that would have exercised it. Had it been called, it would have printed a number above 1 beside a model that was in fact perfect.
RNNhad no__str__Every other layer defines one, so
NN.__str__printed a raw repr — which eval6 displayed:eval6 rewritten
It now explains the recurrence, shows the data as numbers rather than one-hot rows, plots the loss, evaluates on the held-out split, prints decoded predictions against targets, and feeds predictions back to continue a sequence.
Training drops from 20000 epochs to 60 — about two seconds instead of many minutes, with the loss curve showing where it actually flattens.
The data generator also had a real defect: it built targets by
np.roll-ing the inputs and then settingy[:, -1, 1] = 1.rollwraps the first one-hot into the last position, so that assignment added a second 1 — the final target was two-hot, which is incoherent for a softmax objective. Targets are now built directly asstart+1 .. start+10.The part I think is most worth having
The model scores 1.000 at every position, and the notebook now demonstrates why that is a fact about the task, not the model. The target at step t is
input[t] + 1— a function of the current input alone — so a hand-built position-wise shift matrix scores 1.000 too:The counting task exercises the RNN's plumbing but never its memory. So a second task is added that does require it — echo the first element at every position:
A memoryless predictor scores 0.21 there; the RNN reaches 1.000 by carrying the first value in its hidden state. That then motivates the two genuine limits of recurrence — a single hidden state, and the
bptt_truncwindow — and points at attention in eval8 as the answer to both.A wrong claim of my own, removed
My first draft of the notebook asserted accuracy would be lowest at position 0 "for want of history". Running it showed 1.000 everywhere — which is what led to the observation above. Position 0 is not hard, because the target never depended on history in the first place. The claim is gone.
Verification
Sequence-accuracy tests cover 2-D, 3-D and 4-D, that the result never exceeds 1, and that raw probabilities work as well as one-hot. Reverting
multiclass_accuracytoaxis=1fails two tests; removingRNN.__str__fails two more.scripts/README.mdupdated: eval6 no longer belongs on the list of notebooks with long training cells.Update: the RNN's weight gradients were never verified
Added in response to "is the RNN working?" — which turned out to be worth asking.
Its input gradient was fixed and verified earlier this session. Its weight gradients were not: the only test touching
U,VandWasserted they change after a backward pass, which says nothing about whether the values are right. Same gap found suite-wide forDense/LayerNorm/BatchNorm, left open here.They are correct. Against central differences, comparing what the layer actually applied:
dE/dVdE/dUdE/dWThis matters more for the weights than for the input: each of the three is shared across every timestep, so its gradient is a sum over the whole sequence — the part of BPTT most likely to be wrong, and the part a shape assertion cannot see.
A companion test asserts the other side, so the exact checks cannot pass trivially if truncation were ignored:
bptt_trunc(T=8)Halving
grad_U,grad_Vorgrad_Wnow fails a test. Before this, all three passed unnoticed.Suite: 481 passed with cvxopt.