UNIT-2R Deep Learning (1)

UNIT-1 INTRODUCTION
Model Improvement and

Performance
MODULE -01
UNIT-2 INTRODUCTION
• What is computer vision?

• Why Convolutions (CNN)?
• Introduction to CNN,
• Train a simple convolutional neural net,
• Explore the design space for convolutional nets,
• Pooling layer motivation in CNN,
• Design a convolutional layered application,
• Understanding and visualizing a CNN,
• Transfer learning and fine-tuning CNN,
• Image classification,
• Text classification,
• Image classification and hyper-parameter tuning,
• Emerging NN architectures.
UNIT-2 What is computer vision?
• Computer vision is a field of artificial intelligence (AI) that enables computers and systems
to derive meaningful information from digital images, videos and other visual inputs —
and take actions or make recommendations based on that information. If AI enables
computers to think, computer vision enables them to see, observe and understand.
• Computer vision works much the same as human vision, except humans have a head
start. Human sight has the advantage of lifetimes of context to train how to tell objects
apart, how far away they are, whether they are moving and whether there is something
wrong in an image.
• Computer vision trains machines to perform these functions, but it has to do it in much
less time with cameras, data and algorithms rather than retinas, optic nerves and a visual
cortex. Because a system trained to inspect products or watch a production asset can
analyze thousands of products or processes a minute, noticing imperceptible defects or
issues, it can quickly surpass human capabilities.
• Definition: Computer Vision is a field of artificial intelligence that focuses on enabling machines to
interpret and understand visual information from the world, typically through digital images and
videos.
• Image Processing: It involves various techniques for image manipulation, enhancement, and feature
extraction to prepare visual data for analysis.
• Object Detection: Computer Vision can detect and locate objects within images or videos, making it
valuable for tasks like facial recognition, object tracking, and security surveillance.
• Image Classification: It classifies images into predefined categories or labels, allowing for tasks like
identifying animals, recognizing handwritten digits, or diagnosing medical conditions from X-rays.
• Semantic Segmentation: This technique assigns labels to individual pixels in an image, enabling
precise understanding of object boundaries and fine-grained analysis.
• Depth Estimation: Computer Vision can estimate the depth or 3D structure of a scene from 2D
images, which is crucial for applications like autonomous driving and augmented reality.
• Feature Extraction: It involves identifying key patterns, edges, corners, or other distinctive elements
in images, which are then used for various tasks such as object recognition.
• Convolutional Neural Networks (CNNs): CNNs are a fundamental architecture in Computer Vision,
designed to automatically learn and extract features from images through layers of convolutional
and pooling operations.
• Face Recognition: Computer Vision can identify and verify individuals by analyzing facial features,
making it useful for security, access control, and even personalized user experiences.
• Object Tracking: It can track the movement of objects over time in videos, crucial for applications
like autonomous drones, surveillance, and sports analysis.
• Pose Estimation: This technique determines the positions and orientations of human or object
parts within an image or video, commonly used in fields like gesture recognition and robotics.
• OCR (Optical Character Recognition): OCR technology in Computer Vision can recognize and convert
printed or handwritten text into machine-readable text, enabling text analysis in documents or
images.
• Medical Imaging: Computer Vision plays a significant role in medical diagnosis by analyzing medical
images like X-rays, MRIs, and CT scans for disease detection and localization.
• Agriculture: It's used for crop monitoring, disease detection in plants, and automated harvesting by
analyzing images of fields and crops.
• Autonomous Vehicles: Computer Vision is a key component in self-driving cars,

helping them perceive the environment, identify road signs, and avoid obstacles.
• Retail and E-commerce: It's employed for product recommendation, inventory
management, and cashierless checkout systems using computer vision in stores.
• Quality Control: In manufacturing, Computer Vision can inspect products for
defects, ensuring high-quality production.
• Augmented Reality (AR): AR applications overlay digital information on the real
world, often relying on Computer Vision to align and interact with the physical
environment.
• Human-Computer Interaction: Computer Vision enables gesture recognition, gaze
tracking, and other forms of natural interaction between humans and machines.
• Privacy and Ethical Considerations: The use of Computer Vision raises privacy
concerns, and ethical considerations are crucial in its development and deployment,
especially in surveillance and facial recognition applications.
UNIT-2 Introduction to CNN
• CNN stands for Convolutional Neural Network.

• It is a specialized type of ANN that has proven to be highly effective in
various computer vision tasks, such as image classification, object
detection, and image segmentation.
• CNNs are designed to automatically and adaptively learn patterns
and features from input images, making them well-suited for tasks
that involve analyzing visual data.
• Advantages of Convolutional Neural Networks (CNNs):

• Good at detecting patterns and features in images, videos, and audio
signals.
• Robust to translation, rotation, and scaling invariance.
• End-to-end training, no need for manual feature extraction.
• Can handle large amounts of data and achieve high accuracy.
• Disadvantages of Convolutional Neural Networks (CNNs):
• Computationally expensive to train and require a lot of memory.
• Can be prone to overfitting if not enough data or proper regularization
is used.
• Requires large amounts of labeled data.
• Interpretability is limited, it’s hard to understand what the network has
learned.
• Basic Architecture
• Basic Architecture
• There are two main parts to a CNN architecture
• A convolution tool that separates and identifies the various features of the
image for analysis in a process called as Feature Extraction.
• The network of feature extraction consists of many pairs of convolutional or
pooling layers.
• A fully connected layer that utilizes the output from the convolution process
and predicts the class of the image based on the features extracted in
previous stages.
• This CNN model of feature extraction aims to reduce the number of features
present in a dataset.
• It creates new features which summarises the existing features contained in an original
set of features.
• There are many CNN layers as shown in the CNN architecture diagram.
• Convolution Layers
• There are three types of layers that make up the CNN which are the convolutional layers,
pooling layers, and fully-connected (FC) layers.
• When these layers are stacked, a CNN architecture will be formed.
• In addition to these three layers, there are two more important parameters which are the
dropout layer and the activation function which are defined below.
• 1. Convolutional Layer
• This layer is the first layer that is used to extract the various features from the input images.
• In this layer, the mathematical operation of convolution is performed between the input image and
a filter of a particular size MxM.
• By sliding the filter over the input image, the dot product is taken between the filter and the parts
of the input image with respect to the size of the filter (MxM).
• The output is termed as the Feature map which gives us information about the image such as the
corners and edges.
• Later, this feature map is fed to other layers to learn several other features of the input image.
• The convolution layer in CNN passes the result to the next layer once applying the convolution
operation in the input.
• Convolutional layers in CNN benefit a lot as they ensure the spatial relationship between the pixels
is intact.
2. Pooling Layer
• In most cases, a Convolutional Layer is followed by a Pooling Layer.
• The primary aim of this layer is to decrease the size of the convolved feature map to reduce
the computational costs.
• This is performed by decreasing the connections between layers and independently operates
on each feature map.
• Depending upon method used, there are several types of Pooling operations. It basically
summarises the features generated by a convolution layer.
• In Max Pooling, the largest element is taken from feature map.
• Average Pooling calculates the average of the elements in a predefined sized
Image section.
• The total sum of the elements in the predefined section is computed in Sum
Pooling.
• The Pooling Layer usually serves as a bridge between the Convolutional Layer and
the FC Layer.
• This CNN model generalises the features extracted by the convolution layer, and helps the
networks to recognise the features independently.
2. Types of Pooling Layer

Max Pooling
• Max pooling is a pooling operation that selects the maximum element
from the region of the feature map covered by the filter.
• Thus, the output after max-pooling layer would be a feature map
containing the most prominent features of the previous feature map.
Average Pooling
• Average pooling computes the average of the elements
present in the region of feature map covered by the filter.
• Thus, while max pooling gives the most prominent feature in
a particular patch of the feature map, average pooling gives
the average of features present in a patch.
Advantages of Pooling Layer:

• Dimensionality reduction: The main advantage of pooling layers is that they
help in reducing the spatial dimensions of the feature maps. This reduces the
computational cost and also helps in avoiding overfitting by reducing the
number of parameters in the model.
• Translation invariance: Pooling layers are also useful in achieving translation
invariance in the feature maps. This means that the position of an object in
the image does not affect the classification result, as the same features are
detected regardless of the position of the object.
• Feature selection: Pooling layers can also help in selecting the most important
features from the input, as max pooling selects the most salient features and
average pooling preserves more information.
Disadvantages of Pooling Layer:

• Information loss: One of the main disadvantages of pooling layers is that they
discard some information from the input feature maps, which can be
important for the final classification or regression task.
• Over-smoothing: Pooling layers can also cause over-smoothing of the feature
maps, which can result in the loss of some fine-grained details that are
important for the final classification or regression task.
• Hyperparameter tuning: Pooling layers also introduce hyperparameters such
as the size of the pooling regions and the stride, which need to be tuned in
order to achieve optimal performance. This can be time-consuming and
requires some expertise in model building.
Feature Map Size Calculation Formula

Padding
• Padding is a technique used to preserve the spatial dimensions of the
input image after convolution operations on a feature map.
• Padding involves adding extra pixels around the border of the input
feature map before convolution.
• This can be done in two ways:
• Valid Padding: In the valid padding, no padding is added to the input feature
map, and the output feature map is smaller than the input feature map. This
is useful when we want to reduce the spatial dimensions of the feature maps.
• Same Padding: In the same padding, padding is added to the input feature
map such that the size of the output feature map is the same as the input
feature map. This is useful when we want to preserve the spatial dimensions
of the feature maps.
Padding Same Padding

same padding adds additional rows and columns
of pixels around the edges of the input data so
that the size of the output feature map is the
same as the size of the input data. This is
achieved by adding rows and columns of pixels
with a value of zero around the edges of the
input data before the convolution operation.
Valid Padding
Valid padding is used when it is desired to reduce the size
of the output feature map in order to reduce the number
of parameters in the model and improve its computational
efficiency.
Padding
• The most common padding value is zero-padding, which involves
adding zeros to the borders of the input feature map.
• Padding can help in reducing the loss of information at the borders of
the input feature map and can improve the performance of the
model.
• However, it also increases the computational cost of the convolution
operation.
• Overall, padding is an important technique in CNNs that helps in
preserving the spatial dimensions of the feature maps and can
improve the performance of the model.
• Striding
• "striding" refers to the step
size or the distance by which
the convolutional filter or
kernel moves across the input
image during the convolution
operation.
• Striding is an important
parameter that determines
how much the output feature
map's spatial dimensions are
reduced compared to the
input.
• Reasons of using Striding in CNNs

• Downsampling: Striding with a larger value helps downsample the feature
map, reducing its spatial dimensions. This can be beneficial in reducing
computational complexity and memory usage, especially in deep networks.
• Increasing Receptive Field: Larger strides can increase the receptive field of
neurons in deeper layers, allowing them to capture information from a larger
portion of the input image.
• Dimension Reduction: Striding can be used to reduce the spatial dimensions
of the feature map gradually as you move deeper into the network, which can
be helpful in certain tasks like object detection and image classification.
• Example of padding = 1 and striding = 2

• Implementation of Padding and Striding
https://colab.research.google.com/drive/116QNuxPQtxhb-
1zqyWqDIboYh2GzTvDD?authuser=1#scrollTo=Esb12gcaa9V_
3. Fully Connected Layer

• The Fully Connected (FC) layer consists of the weights and biases along with
the neurons and is used to connect the neurons between two different layers.
• These layers are usually placed before the output layer and form the last few
layers of a CNN Architecture.
• In this, the input image from the previous layers are flattened and fed to the
FC layer.
• The flattened vector then undergoes few more FC layers where the
mathematical functions operations usually take place.
• In this stage, the classification process begins to take place.
• The reason two layers are connected is that two fully connected layers will
perform better than a single connected layer.
• These layers in CNN reduce the human supervision
4. Dropout
• When all the features are connected to the FC layer, it can cause overfitting in
the training dataset.
• Overfitting occurs when a particular model works so well on the training data
causing a negative impact in the model’s performance when used on a new
data.
• To overcome this problem, a dropout layer is utilised wherein a few neurons
are dropped from the neural network during training process resulting in
reduced size of the model.
• On passing a dropout of 0.3, 30% of the nodes are dropped out randomly
from the neural network.
• Dropout results in improving the performance of a machine learning model as
it prevents overfitting by making the network simpler.
• It drops neurons from the neural networks during training.
4. Dropout
5. Activation Functions
• Finally, one of the most important parameters of the CNN model is the
activation function.
• They are used to learn and approximate any kind of continuous and complex
relationship between variables of the network.
• It decides which information of the model should fire in the forward direction
and which ones should not at the end of the network.
• It adds non-linearity to the network.
• There are several commonly used activation functions such as the ReLU,
Softmax, tanH and the Sigmoid functions.
• Each of these functions have a specific usage. For a binary classification CNN
model, sigmoid and softmax functions are preferred and for a multi-class
classification, generally softmax us used.
UNIT-2 Expected Questions on Padding and Striding
1. What is padding in convolutional neural networks, and why is it used?
2. Explain the concept of "valid" and "same" padding in CNNs. How do they affect the spatial dimensions of the output
feature maps?
3. How does zero-padding impact the size of the output feature map during convolution?
4. In the context of CNNs, what is the purpose of using padding when applying convolutional filters to an input image?
5. What is striding in CNNs, and how does it affect the spatial dimensions of the output feature maps?
6. How can you calculate the size of the output feature map when applying a convolutional filter with a specific size and
stride to an input image with a given size?
7. In what scenarios might you use larger strides during convolution, and what are the advantages and disadvantages of
doing so?
8. How does the choice of padding and striding influence the receptive field of neurons in a CNN?
9. Explain the trade-off between using padding and striding in CNNs in terms of preserving spatial information versus
reducing computational complexity.
10. Can you describe a situation where you would use both padding and striding in a CNN layer, and what would be the
expected impact on the output feature map?
11. What is the relationship between the size of the input image, the size of the convolutional filter, the amount of
padding, and the stride value in determining the size of the output feature map?
12. How do padding and striding affect the performance of a CNN in tasks like image classification or object detection?

UNIT-2 Why Convolutions (CNN)?
• Introduction to CNN,

UNIT-2R Deep Learning (1)

Uploaded by

Document Information

Copyright

Available Formats

Share this document

Share or Embed Document

Sharing Options

Did you find this document useful?

Is this content inappropriate?

Copyright:

Available Formats

UNIT-2R Deep Learning (1)

Uploaded by

Copyright:

Available Formats

UNIT-1 INTRODUCTION

Model Improvement and

• What is computer vision?

• Autonomous Vehicles: Computer Vision is a key component in self-driving cars,

• CNN stands for Convolutional Neural Network.

• Advantages of Convolutional Neural Networks (CNNs):

2. Types of Pooling Layer

Advantages of Pooling Layer:

Disadvantages of Pooling Layer:

Feature Map Size Calculation Formula

Padding Same Padding

• Reasons of using Striding in CNNs

• Example of padding = 1 and striding = 2

• Implementation of Padding and Striding

3. Fully Connected Layer

• Train a simple convolutional neural net,

You might also like