A Machine Learning model is trained to predict house prices using this small dataset:
| Features | ||||
|---|---|---|---|---|
| House | Size | Bedrooms | Has cat? | LABEL |
| 1 | 100 | 3 | Yes | "Expensive" |
| 2 | 80 | 2 | No | "Cheap" |
| 3 | 150 | 4 | Yes | "Expensive" |
The Model attempts to attach the correct label to some new data:
| Features | ||||
|---|---|---|---|---|
| House | Size | Bedrooms | Has cat? | LABEL |
| 4 | 80 | 1 | Yes | " " |
With your team, discuss which label that the model might attach to this new data. Justify your answer.
The model is trained on too little data.
It memorises the examples it has seen (like remembering the answers to a quiz).
But it fails to generalise — it performs poorly when shown new data.
For example:
If you train a model to predict house prices using only 5 houses, it might learn:
But in reality, that’s just a coincidence in the small dataset.
When it sees new houses, it gives bad predictions — because it memorised, it didn’t understand.
With your team, discuss strategies to reduce the opportunity for overfitting in a Machine Learning model.
Try to think of atleast 2:
One feature representing a person is used to train an unsupervised ML model. The feature is age:
| Age |
|---|
| 16 |
| 17 |
| 23 |
| 37 |
| 42 |
| 45 |
| 18 |
| 47 |
The following clusters are produced and the model is starting to create patterns and somehow make sense of the data points.
Each red data point (a person) is composed of 1 feature - age
There is only one feature used by this model.
We can say that the model processes data in one-dimensional space.
Two features representing a person are now used to train an unsupervised ML model. The features are age and weight:
| Age | Weight |
|---|---|
| 16 | 60 |
| 17 | 65 |
| 23 | 80 |
| 37 | 77 |
| 42 | 45 |
| 45 | 77 |
| 18 | 65 |
| 47 | 64 |
There are 2 features. We can say that the model operates in two-dimensional space.
We can see that the distance between data points (people) has increased.
It is becoming difficult for the model to identify clusters and patterns.
We can now suggest that as the number of features (dimensions) increases, then the distance between the data points also increases.
Here 3 features are used.
It is getting increasingly difficult to form meaningful clusters.
The data is getting sparser as the number of dimensions increases.
Dimensionality reduction means reducing the number of input features (dimensions) used by a machine learning model while keeping as much useful information as possible.
Here are some key reasons for reducing the number of dimensions in a ML model.
| Reason | Explanation | Example |
|---|---|---|
| Less computational complexity | Fewer features > fewer calculations > faster training and prediction. | A dataset with 1,000 features might take hours to process, but with 50 important ones, it trains in minutes. |
| Reduces overfitting | Too many features let the model memorize noise. Removing unimportant ones helps it generalize better. | A model predicting house prices doesn’t need the “wall paint color” feature — that’s just noise. |
| Mitigates the curse of dimensionality | High-dimensional data is sparse. Fewer dimensions mean the model can learn real patterns instead of random gaps. | Instead of learning from a 100-feature space (mostly empty), the model works on 10 features that truly matter. |
| Improves data visualization | Humans can’t visualize data beyond 3D. Reducing to 2D or 3D helps us plot and inspect patterns. | Using PCA, we can see how different types of houses cluster visually. |
| Reduces memory usage and storage needs | Smaller datasets mean less space and faster processing. | Easier to store and load during training. |
Make a new entry in the Journal section of the booklet.
Did you discover anything new today? Do you have questions in your mind that you want answers to?
Make these notes in your journal.
Machine Learning
Supervised Learning
Unsupervised Learning
Labels
Training
Training data
Test data
Features
Machine Learninglabels Supervised learning Unsupervised learning features model training data trainingtest data