Skip to main content

What Is Image Segmentation?

What Is Image Segmentation?

Image segmentation is the process of dividing a digital image into labeled pixel groups for object detection, image processing, and other computer vision tasks. Organizations have extensive collections of image data, such as single images, video sequences, views from multiple cameras, or three-dimensional data. Image segmentation techniques parse an image's complex visual data into specifically shaped segments to identify object boundaries and background regions. They automate image analysis and classification to support many AI use cases like medical image diagnosis, machinery maintenance, and self-driving vehicles.

What are the applications of image segmentation?

Image segmentation allows a machine to understand images—the scenes, objects, and the image's meaning. For example, you can use image segmentation to

  • Identify abnormalities in medical image analysis that the human eye may overlook.

  • Alter an image based on a particular object within the scene.

  • Notice places of interest, particularly if they do not align with the natural landscape in satellite images.

It also has applications in self-driving cars. You can navigate the roads safely by determining what in the visual scene, such as traffic lights, people, speed signs, etc., requires braking, turning, or accelerating.

How does image segmentation work?

Image segmentation works by processing an image to define two or more parts, showing the boundaries between different regions. Regions are distinct image areas that contain background information and foreground objects. For example, objects such as a person or a car are well-defined shapes distinctly separate from background elements like a field or a group of people without clear boundaries.

In the segmentation process, regions are classified into segments based on criteria such as:

  • Visual similarity—including color, texture, or proximity to other similar pixels

  • Semantic characteristics, such as the type of object (e.g., vehicle, animal)

  • Image setting like a candy store or a park.

Each pixel in an image is labeled to denote the segment to which it belongs.

Masking

Each segment, especially those representing objects, is overlaid with a mask that isolates it from the rest of the image. This mask is a binary matrix that identifies which pixels within the image belong to the object (marked as '1') and which do not (marked as '0'). This allows other algorithms to take image segmentation output and precisely apply tagging or editing operations to specific regions while keeping the objects intact and separate from their environments. A colored overlay may visually denote masks.

Traditional image segmentation techniques

Data scientists were using traditional image segmentation techniques decades before current advancements in machine learning and artificial intelligence. These techniques use more primitive digital image processing methods and are less accurate in their outcomes. However, they can sometimes give results faster than modern methods and are still a viable option for single-object images. We give an overview below.

Thresholding

Thresholding produces binary maps of an image based on a threshold value. Pixels scored above the threshold value are set to 1, and the rest are set to 0.

Edge detection

Edge-based segmentation detection finds the edges of objects by variations in the brightness of pixels.

Watersheds

Watersheds greyscale an image and determine regions based on the variations in the mapping.

Region-based segmentation

The technique uses seed pixels within an image, where the pixels expand to capture similar pixels for segmentation purposes.

Clustering-based segmentation

Clustering algorithms use unsupervised learning, such as K-means clustering, to cluster pixels into different K classes based on similar characteristics between the pixels.

How does deep learning improve image segmentation?

Deep learning uses neural networks to advance the field of image segmentation. Most techniques specifically use convolutional neural networks (CNN) that contain mathematical filters called convolutions to self-identify complex features in a given dataset. Compared to traditional image segmentation techniques, deep learning techniques bring accuracy, efficiency, and the ability to handle complex and diverse image data.

For example, deep learning models can:

  • Learn hierarchical features from raw images and generalize to new, unseen datasets.

  • Discern intricate patterns and textures that are critical for distinguishing between different objects.

  • Generate more accurate output even for image variations like changes in lighting, scale, and orientation.

  • Scale up as needed to handle large datasets in commercial and research settings.

Humans have helped label images to train deep-learning models for practical use.

Deep learning image segmentation techniques

We explain some popular methods below.

Fully convolutional networks

A fully convolutional network has a series of convolutional layers in an encoder neural network fed by visual data. The encoder processes the input image to extract and condense useful information into feature maps representing various aspects of the input. It identifies segments and downsamples the data to remove fuzziness. Downsampling reduces the input image's spatial dimensions (width and height). The data is then processed through decoder layers to add segmentation labels.

U-Nets

U-Nets advanced the field by adding skip connections that directly linked the encoder and decoder layers. The connections reduce loss when downsampling and thus obtain more accurate levels of segmentation. U-Nets can better localize and delineate object boundaries by combining the high-level, semantically rich features from the deeper layers with the detailed spatial information from the early layers.

Other advanced techniques

Even more advanced types of deep learning image segmentation architecture include:

  • Mask R-CNNs, which identify objects, their bounding boxes, and exact pixel masks

  • Transformers, based on transformer models, similar to natural language processing architectures

  • Other models, such as DeepLab, SegNet, and more

The details of these techniques are out of the scope of this article.

What are the types of deep learning image segmentation?

Deep learning image segmentation can be classified into three broad categories.

Semantic segmentation

Semantic segmentation is a way to define different classes of objects or areas of interest in an image by labeling pixels of the same semantic class in the same color mask. For instance, a field scene may have a grass field, two trees in the distance, a person, and the sky. Each of these will be colored in a different color, pixel by pixel, for complete pixel-based segmentation of an image.

In semantic segmentation, there is no more context given other than these color-mapped parts; it is simply the entire image segmented into these new color groups. There is also no distinction between objects and background; both are equally reported.

Instance segmentation

Instance segmentation is concerned with segmenting an image by picking out everything, even where some things partly obscure others of the same class. Semantic segmentation labels every pixel of an image with a class, irrespective of the object instances, while instance segmentation identifies each instance of particular objects separately. For example, semantic segmentation views a gathering of people as a whole, whereas instance segmentation draws boundaries around the different people. Instance segmentation differs from object detection in that it doesn't just give a bounding box approximation; it gives pixel-based segmentation boundaries of all elements within a scene.

Panoptic segmentation

Panoptic segmentation integrates aspects of both semantic and instance segmentation. It performs semantic segmentation and instance segmentation in parallel with a feature pyramid network (FPN) to capture features at multiple scales. Each distinct object instance is usually given a unique identifier, while similar objects (e.g., all cars) share the same category label but are differentiated by instance.

What is the difference between image segmentation, object detection, and image classification?

Image segmentation, object detection, and image classification are all specific image processing techniques involved in computer vision activities.

  • Image segmentation shows the different focal areas within an image by producing a pixel-based mask.

  • Object detection identifies and labels objects within an image (e.g., cat, couch, person)

  • Image classification labels the image as a whole (e.g., living room)

These three image-processing techniques tell a machine all about an image. When a machine consumes and uses this data, it can make decisions based on images, much like a human would.

How can AWS support image segmentation?

Amazon SageMaker is a fully managed service that combines a broad set of tools to enable high-performance, low-cost machine learning for any use case. SageMaker JumpStart provides pre-trained, open-source models for various problem types to help you start with machine learning. Using JumpStart APIs, you can quickly fine-tune and deploy a trained image segmentation model from MXNet.

SageMaker Ground Truth offers the most comprehensive set of human-in-the-loop capabilities, allowing you to harness the power of human feedback across the ML lifecycle to improve the accuracy and relevancy of models. You can use semantic segmentation labeling jobs to get human workers to classify pixels in the image into predefined labels or classes. Ground Truth supports both single and multi-class semantic segmentation labeling jobs.

Some labeling job tasks contain images with many objects that need to be segmented. Ground Truth provides an auto-segmentation tool that uses a machine learning model to automatically segment individual objects with minimal worker input. Workers can review and refine the auto-segmented output to reduce segmentation time and increase accuracy.

For customers looking for a fully managed service, Amazon Rekognition helps you add computer vision APIs to applications without the associated time and cost of building machine learning models. Without worrying about image segmentation algorithms, you can analyze millions of images and videos in seconds. It includes capabilities for every computer vision use case, from content moderation and face search to video segment detection and custom labels.

Get started with image segmentation and computer vision on AWS by creating a free account today.

Browse all cloud computing concepts

Browse all cloud computing concepts content here:

Loading
Loading
Loading
Loading
Loading

Did you find what you were looking for today?

Let us know so we can improve the quality of the content on our pages