Generalization to new cases on the AI Learning track. Generalization is success on examples that were not used to change the parameters. A perfect score on the training rows can still fail tomorrow. The honest number is the score on held-out rows.
This lesson assumes you already worked through Loss: the score training reduces.
The idea in practice
Hold out a test set and do not tune on it. Use a validation set for choices. Touch the test set once, at the end.
A concrete check
goal = {
'track': 'AI Learning',
'lesson': 'Generalization to new cases',
}
checks = [
'input available at decision time',
'score matches the real decision',
'failure case written down',
]
print(goal['lesson'])
for item in checks:
print('-', item)
Run the sketch locally if you have Python. The printout is a reminder of the checks, not a trained model. Replace the strings with the real inputs from your own example before you treat it as a design.
What usually goes wrong
Tuning prompts, features, or hyperparameters on the test set makes the test number a training number. When this happens, stop adding parameters or tools. Fix the check, the data, or the permission, then run the same example again.
What to write down
- The input you are allowed to use at decision time.
- The output and the score or pass rule.
- One failure you will test on purpose.
- What you will not claim the system can do.
Practice
Describe how you would split 10,000 rows so that a future week stays untouched.
Self-check
- Say Generalization to new cases in one sentence that mentions an input and an output.
- Name the failure mode in this lesson and the check that would catch it.
Done when: you can explain this lesson without the page open, and you have a written failure case.