This project demonstrates a high-performance image segmentation algorithm that combines CPU parallelism (OpenMP) with GPU acceleration (CUDA) for efficient k-means clustering of image pixels.
Ensure you have a GPU available in your environment. You can check this by running:
nvidia-smiInstall the necessary dependencies:
sudo apt-get update
sudo apt-get install -y libopencv-dev
sudo apt-get install -y libomp-devCreate the source file for the hybrid OpenMP+CUDA implementation. The source code includes the k-means clustering algorithm implemented in CUDA and OpenMP.
Compile the CUDA and OpenMP program using nvcc and g++.
Run the segmentation with different cluster counts:
./kmeans lenna.png output_5.png 5
./kmeans lenna.png output_8.png 16Visualize the original image and the segmented results using Python and OpenCV:
import cv2
import matplotlib.pyplot as plt
def display_image(title, path):
img = cv2.cvtColor(cv2.imread(path), cv2.COLOR_BGR2RGB)
plt.figure(figsize=(8, 8))
plt.imshow(img)
plt.title(title)
plt.axis('off')
plt.show()
# Display original image
display_image("Original Image", "lenna.png")
# Display segmented images
display_image("3 Clusters", "output_3.png")
display_image("5 Clusters", "output_5.png")
display_image("8 Clusters", "output_8.png")Compare the performance with different cluster counts and plot the results:
import time
cluster_counts = [2, 4, 8, 16, 32]
times = []
for k in cluster_counts:
start_time = time.time()
!./kmeans lenna.png benchmark_{k}.png {k} > /dev/null 2>&1
elapsed = time.time() - start_time
times.append(elapsed)
print(f"{k} clusters: {elapsed:.2f} seconds")
# Plot results
plt.figure(figsize=(10, 5))
plt.plot(cluster_counts, times, 'o-', markersize=8)
plt.xlabel('Number of Clusters')
plt.ylabel('Execution Time (seconds)')
plt.title('Performance vs. Cluster Count')
plt.grid(True)
plt.show()This project demonstrates:
- A hybrid OpenMP+CUDA implementation of k-means image segmentation.
- How to compile and run the program in a GPU-enabled environment.
- Visualization of segmentation results with different cluster counts.
- Performance benchmarking across different configurations.
The hybrid approach combines:
- GPU acceleration for compute-intensive distance calculations.
- CPU parallelism for other tasks like image loading and centroid initialization.
Try experimenting with:
- Different images.
- Various cluster counts.
- Adjusting the convergence threshold.
- Modifying the block size for CUDA kernels.