Objectives

Students will be able to:

  • uderstand overfitting and how to eliminate it.

Machine Learning

Theme A4

Dimensionality Reduction

Scenario

A Machine Learning model is trained to predict house prices using this small dataset:

Features
House Size Bedrooms Has cat? LABEL
1 100 3 Yes "Expensive"
2 80 2 No "Cheap"
3 150 4 Yes "Expensive"

New Data

The Model attempts to attach the correct label to some new data:

Features
House Size Bedrooms Has cat? LABEL
4 80 1 Yes "         "

With your team, discuss which label that the model might attach to this new data. Justify your answer.

Overfitting

The model is trained on too little data.

It memorises the examples it has seen (like remembering the answers to a quiz).

But it fails to generalise — it performs poorly when shown new data.

For example:

If you train a model to predict house prices using only 5 houses, it might learn:

If a house costs exactly $500,000, it must have 3 bedrooms and 2 bathrooms.

But in reality, that’s just a coincidence in the small dataset.

When it sees new houses, it gives bad predictions — because it memorised, it didn’t understand.

Reducing Overfitting

With your team, discuss strategies to reduce the opportunity for overfitting in a Machine Learning model.

Try to think of atleast 2:

  •  
  •  

1D Space

One feature representing a person is used to train an unsupervised ML model. The feature is age:

Age
16
17
23
37
42
45
18
47

The following clusters are produced and the model is starting to create patterns and somehow make sense of the data points.

 

Each red data point (a person) is composed of 1 feature - age

There is only one feature used by this model.

We can say that the model processes data in one-dimensional space.

2D Space

Two features representing a person are now used to train an unsupervised ML model. The features are age and weight:

Age Weight
16 60
17 65
23 80
37 77
42 45
45 77
18 65
47 64

There are 2 features. We can say that the model operates in two-dimensional space.

 

We can see that the distance between data points (people) has increased.

It is becoming difficult for the model to identify clusters and patterns.

Data Sparsity

We can now suggest that as the number of features (dimensions) increases, then the distance between the data points also increases.

Here 3 features are used.

 

It is getting increasingly difficult to form meaningful clusters.

The data is getting sparser as the number of dimensions increases.

Dimensionality Reduction

Dimensionality reduction means reducing the number of input features (dimensions) used by a machine learning model while keeping as much useful information as possible.

Here are some key reasons for reducing the number of dimensions in a ML model.

Reason Explanation Example
Less computational complexity Fewer features > fewer calculations > faster training and prediction. A dataset with 1,000 features might take hours to process, but with 50 important ones, it trains in minutes.
Reduces overfitting Too many features let the model memorize noise. Removing unimportant ones helps it generalize better. A model predicting house prices doesn’t need the “wall paint color” feature — that’s just noise.
Mitigates the curse of dimensionality High-dimensional data is sparse. Fewer dimensions mean the model can learn real patterns instead of random gaps. Instead of learning from a 100-feature space (mostly empty), the model works on 10 features that truly matter.
Improves data visualization Humans can’t visualize data beyond 3D. Reducing to 2D or 3D helps us plot and inspect patterns. Using PCA, we can see how different types of houses cluster visually.
Reduces memory usage and storage needs Smaller datasets mean less space and faster processing. Easier to store and load during training.

Journal Task

Make a new entry in the Journal section of the booklet.

Did you discover anything new today? Do you have questions in your mind that you want answers to?

Make these notes in your journal.

Glossary

Machine Learning

Supervised Learning

Unsupervised Learning

Labels

Training

Training data

Test data

Features

Tags

Machine Learninglabels Supervised learning Unsupervised learning features model training data trainingtest data