Train
Teach both models again on everything collected so far. Every new version sits the same exam, and it's kept only if it reads better than the version it would replace.
1 The data
—samples
—writers
—syllables covered
—last snapshot
2 The models
3 Train a new version
Which models?
What happens
- Snapshot Freeze today's data as a numbered dataset, then split it into train and exam.
- Train Each model learns from scratch on the train part only.
- Exam The new version and the current one read every exam sample.
- Keep or discard The new version is kept only if it scores higher.
It runs in the background on the server and can take several hours. You can close this page and come back.
- Snapshot
- Train
- Exam
- Keep or discard
Show the log
History
No training runs yet.
What do these words mean?
- Sample
- One handwritten syllable, saved as pen strokes with their timing.
- Snapshot
- A frozen copy of the data at training time (d001, d002, …), so every result can be traced back to exactly the data it came from.
- Train / exam split
- About 10% of syllables are kept out of training entirely. They're the exam: questions the model has never seen, so the score shows how well it reads new handwriting, not how well it memorised.
- Validation
- While training, 10% of the train part is held back to check progress after every epoch. It decides when to stop, never what to keep.
- Epoch
- One full pass over all the training samples.
- Loss
- A number for how wrong the model is. Training makes it go down. If training loss keeps falling but validation loss doesn't, the model is starting to memorise (overfitting).
- Early stopping
- Training stops by itself when validation loss hasn't improved for a set number of epochs (the "patience").
- Exact match
- Out of all exam samples, how many the model read completely right. Higher is better.
- Character error rate (CER)
- Out of every 100 characters, how many the model got wrong (missing, extra or swapped). Lower is better; it gives partial credit that exact match doesn't.
- Parameters (0.9m)
- The numbers the model learns. 0.9m = 0.9 million. More isn't always better with small data.
- Version (swalchat-1.2)
- 1 = the recipe (architecture and settings); .2 = the second time that recipe was retrained on more data.