Skip to content

About

A from-scratch deep learning image compression prototype using a convolutional autoencoder, latent quantization, entropy modeling, and Huffman coding.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

9 Commits

Folders and files

Repository files navigation

Learned Image Compression Using Deep Learning

Live Demo

A from-scratch research prototype for learned image compression using a convolutional autoencoder, latent-space quantization, entropy modeling, and Huffman coding.

The project explores how a neural network can learn a compact representation of an image and how that representation can be transformed into an actual compressed binary bitstream.

Status: Research prototype / baseline implementation
Focus: Understanding and implementing the complete learned image-compression pipeline from first principles


Overview

Traditional image codecs such as JPEG use hand-designed transforms, quantization, and entropy coding.

This project investigates a learned alternative in which a neural network learns an image representation directly from data.

The complete pipeline is:

                    COMPRESSION

Input Image
     │
     ▼
┌──────────────┐
│ CNN Encoder  │
└──────────────┘
     │
     ▼
Latent Representation
     │
     ▼
Quantization
     │
     ▼
Discrete Latent Symbols
     │
     ▼
Huffman Coding
     │
     ▼
Compressed Binary Stream

Decompression follows the reverse process:

                 DECOMPRESSION

Compressed Binary Stream
          │
          ▼
    Huffman Decoding
          │
          ▼
 Quantized Latent Symbols
          │
          ▼
    Dequantization
          │
          ▼
      CNN Decoder
          │
          ▼
 Reconstructed Image

Project Goal

The main question explored in this project is:

Can a convolutional neural network learn a compact image representation that allows an image to be compressed into a small binary representation while maintaining useful reconstruction quality?

The project therefore studies three connected ideas:

  1. Learning a compact latent representation.
  2. Reducing that representation through quantization.
  3. Entropy-coding the resulting discrete symbols.

Key Features

  • Convolutional image encoder
  • Convolutional image decoder
  • Learned latent representation
  • Latent-space quantization
  • Gaussian entropy-model experiment
  • Huffman entropy coding
  • Actual binary .bin compressed stream
  • End-to-end compression and decompression
  • MSE, PSNR, SSIM, and BPP evaluation
  • Rate-distortion experiments
  • JPEG baseline comparison
  • DIV2K-based experimentation
  • Reusable Python source modules
  • Colab research notebook
  • Local Python project structure

1. Method

1.1 Encoder

The encoder transforms an image into a learned latent representation.

The encoder mapping is:

y = fθ(x)

where:

  • x is the input image
  • fθ is the encoder
  • y is the latent representation
  • θ represents the encoder parameters

For a 256 × 256 RGB image, the current encoder produces:

Input
3 × 256 × 256

      ↓

Conv2D + ReLU

64 × 128 × 128

      ↓

Conv2D + ReLU

128 × 64 × 64

      ↓

Conv2D + ReLU

192 × 32 × 32

      ↓

Conv2D + ReLU

128 × 16 × 16

      ↓

Conv2D

96 × 16 × 16

Therefore, the current latent representation has dimensions:

96 × 16 × 16

The encoder reduces the spatial resolution while increasing the number of learned feature channels.


1.2 Decoder

The decoder reconstructs an approximation of the original image.

The decoder mapping is:

x̂ = gφ(y)

where:

  • y is the latent representation
  • gφ is the decoder
  • x̂ is the reconstructed image
  • φ represents the decoder parameters

The current decoder progressively increases the spatial resolution:

96 × 16 × 16

      ↓

128 × 32 × 32

      ↓

192 × 64 × 64

      ↓

128 × 128 × 128

      ↓

64 × 256 × 256

      ↓

3 × 256 × 256

2. Reconstruction Learning

Before compression can be evaluated, the autoencoder must first learn to reconstruct images.

The reconstruction distortion is measured using Mean Squared Error:

D(x, x̂) = (1/N) Σᵢ (xᵢ − x̂ᵢ)²

The baseline reconstruction objective is:

L_reconstruction = D(x, x̂)

The network learns encoder and decoder parameters that minimize this reconstruction error.


3. Latent Quantization

The encoder produces continuous-valued latent features.

For compression, these values are converted into discrete symbols using quantization.

The current implementation uses:

ŷ = round(y / q)

where q is the quantization step.

During reconstruction, the quantized latent is converted back to the original scale:

y_dequantized = q · ŷ

A smaller q preserves more information, while a larger q produces stronger quantization and generally increases distortion.


4. Entropy Modeling

After quantization, the latent representation consists of discrete symbols.

The project includes a simple Gaussian entropy model that estimates the probability mass of quantized latent values.

For a discrete symbol s with probability p(s), entropy is:

H(S) = −Σₛ p(s) log₂ p(s)

An estimated coding rate can be written as:

R = −Σᵢ log₂ p(ŷᵢ)

and normalized by image pixels:

BPP = R / (H × W)

The Gaussian entropy model is used as an experimental probability and rate-estimation component.

The actual binary demonstration uses Huffman coding over the quantized latent symbols.


5. Huffman Coding

The final prototype converts quantized latent symbols into an actual binary stream using Huffman coding.

The process is:

Quantized Latent
      ↓
Symbol Frequency Counting
      ↓
Huffman Tree
      ↓
Binary Codes
      ↓
Encoded Bitstream
      ↓
.bin File

The decoding process is lossless with respect to the quantized latent symbols.

Therefore:

Quantized Latent
      ↓
Huffman Encode
      ↓
Binary Bitstream
      ↓
Huffman Decode
      ↓
Same Quantized Latent

The image compression itself remains lossy because information is discarded during the learned transformation and quantization stages.


6. Rate-Distortion Optimization

A central idea in image compression is the trade-off between bitrate and reconstruction quality.

The general rate-distortion objective is:

L = R + λD

where:

  • R represents the coding rate
  • D represents reconstruction distortion
  • λ controls the rate-distortion trade-off

A system may accept greater distortion in exchange for a lower bitrate, or spend more bits to preserve more image information.

The project investigates this relationship using multiple quantization strengths.


7. Dataset

The project uses the DIV2K high-resolution image dataset.

The dataset is used for:

  • Model development
  • Training
  • Validation
  • Compression experiments

The training pipeline extracts random:

256 × 256

image crops.

The dataset is not included in this repository because of its size.


8. Training Pipeline

The current workflow is:

DIV2K Image
     ↓
Random 256 × 256 Crop
     ↓
RGB Image
     ↓
Normalize to [0, 1]
     ↓
Mini-batch
     ↓
CNN Encoder
     ↓
Latent Representation
     ↓
CNN Decoder
     ↓
Reconstructed Image
     ↓
MSE Loss
     ↓
Backpropagation

After the reconstruction baseline is established, the latent representation is quantized and entropy-coded.


9. Evaluation Metrics

9.1 Mean Squared Error

MSE = (1/N) Σᵢ (xᵢ − x̂ᵢ)²

Lower values indicate smaller pixel-wise reconstruction error.


9.2 Peak Signal-to-Noise Ratio

The images are normalized to the range [0, 1].

Therefore:

PSNR = 10 log₁₀(1 / MSE)

Higher values indicate better reconstruction fidelity.


9.3 Structural Similarity

SSIM measures structural similarity between the original and reconstructed images.

Higher values indicate greater structural similarity.


9.4 Bits Per Pixel

The bitrate is normalized by image area:

BPP = coded bits / (H × W)

Lower BPP indicates a more compact representation.


9.5 Compression Ratio

For raw 8-bit RGB data:

CR = original raw size / compressed size

The compression ratios reported by this prototype are measured relative to raw RGB data, not against an already compressed JPEG or PNG file.


10. Results

10.1 Current Learned Compression Baseline

A full validation evaluation produced the following current baseline:

Metric Result
BPP 1.2898
PSNR 23.10 dB
SSIM 0.6250
MSE 0.007518

These values represent the current learned-compression baseline on the DIV2K validation set.

They should be interpreted as experimental baseline results rather than state-of-the-art results.


10.2 Quantization Study

The current validation experiment produced the following operating points:

Quantization Step BPP PSNR SSIM
0.05 2.2148 22.70 dB 0.6053
0.10 1.2837 23.26 dB 0.6258
0.20 1.0483 22.72 dB 0.5980
0.50 1.0232 21.78 dB 0.5852
1.00 1.0358 18.58 dB 0.4919

As quantization becomes more aggressive, reconstruction quality decreases substantially.

The BPP values in this experiment come from the current experimental probability/rate-estimation setup and should not be interpreted as a production codec bitrate.


10.3 Actual Binary Compression

The project also generates an actual Huffman-coded .bin file.

For one unseen test image, one run produced:

Metric Result
Compressed size 10,254 bytes
BPP 1.2517
Compression ratio 19.17×
PSNR 21.02 dB
SSIM 0.6274

The compression ratio is measured against raw 8-bit RGB data for a 256 × 256 image.

The Huffman decoder was verified to recover the exact same quantized latent symbols.


11. JPEG Baseline

JPEG was evaluated on the same DIV2K validation images.

JPEG Quality BPP PSNR SSIM
10 0.3961 27.34 dB 0.8091
25 0.6377 31.34 dB 0.8946
50 0.9719 33.34 dB 0.9238
75 1.4307 35.05 dB 0.9475
90 2.3661 38.53 dB 0.9707

12. Learned Compression vs JPEG

The current learned baseline does not outperform JPEG in this experiment.

For example:

Learned baseline
≈ 1.29 BPP
≈ 23.10 dB PSNR

JPEG
≈ 0.97 BPP
≈ 33.34 dB PSNR

JPEG therefore provides substantially better rate-distortion performance for the current experimental setup.

This result is useful because it establishes a working learned-compression baseline and identifies the areas that require further improvement.


13. Rate-Distortion Plot

The current rate-distortion comparison is stored at:

experiments/plots/rate_distortion_comparison.png

The plot compares the learned compression baseline and JPEG using BPP and PSNR.

Rate-Distortion Comparison


14. End-to-End Demonstration

The prototype can process an unseen image through the complete pipeline:

Input Image
     ↓
CNN Encoder
     ↓
Latent Representation
     ↓
Quantization
     ↓
Huffman Coding
     ↓
Compressed .bin File
     ↓
Huffman Decoding
     ↓
Quantized Latent
     ↓
CNN Decoder
     ↓
Reconstructed Image

A demonstration image and associated metrics are included in the project outputs.


15. Project Architecture

learned-image-compression/
│
├── README.md
├── LICENSE
├── requirements.txt
├── .gitignore
│
├── notebooks/
│   └── image_compression.ipynb
│
├── src/
│   ├── __init__.py
│   ├── model.py
│   ├── entropy.py
│   ├── huffman.py
│   ├── metrics.py
│   └── inference.py
│
├── experiments/
│   ├── results/
│   │   ├── final_results.json
│   │   ├── final_comparison.json
│   │   └── final_demo_metrics.json
│   │
│   └── plots/
│       └── rate_distortion_comparison.png
│
└── demo/

16. Source Modules

src/model.py

Contains:

  • Encoder
  • Decoder
  • Autoencoder

src/huffman.py

Contains:

  • Huffman tree construction
  • Code generation
  • Encoding
  • Decoding
  • Bit-to-byte conversion

src/entropy.py

Contains the experimental Gaussian entropy model used for probability and rate estimation.

src/metrics.py

Contains:

  • MSE
  • PSNR
  • SSIM

src/inference.py

Contains reusable utilities for:

  • Image loading
  • Latent encoding
  • Quantization
  • Huffman compression
  • Latent decoding
  • Reconstruction

17. Installation

Clone the repository:

git clone https://github.com/doitmuna/learned-image-compression.git
cd learned-image-compression

Create a virtual environment:

python -m venv .venv

Activate it on Windows:

.venv\Scripts\activate

Install the dependencies:

pip install -r requirements.txt

18. Running the Project

The complete experimental development notebook is available at:

notebooks/image_compression.ipynb

The notebook contains the development and evaluation workflow used to build the prototype.

The dataset itself is not included in the repository.


19. Reproducibility

To reproduce the experiments:

  1. Install the dependencies from requirements.txt.
  2. Download the required DIV2K training and validation images.
  3. Place the images in the directories expected by the notebook.
  4. Open notebooks/image_compression.ipynb.
  5. Run the preprocessing, training, evaluation, and compression experiments.

Large datasets, trained checkpoints, and generated binary files are intentionally excluded from Git version control.


20. What the Project Demonstrates

This project demonstrates the path from neural image representation to an actual compressed bitstream:

Image
  ↓
Learned Representation
  ↓
Quantization
  ↓
Entropy Modeling
  ↓
Entropy Coding
  ↓
Binary Stream

It also demonstrates an important practical lesson in compression research:

A smaller representation does not automatically imply better compression.

A useful compression system must jointly consider:

  • Representation quality
  • Quantization
  • Probability modeling
  • Entropy coding
  • Reconstruction quality
  • Actual coded size

21. Limitations

The current implementation is intentionally a baseline research prototype.

Important limitations include:

  • The encoder and decoder are relatively small compared with modern learned image-compression architectures.
  • The final training schedule is limited compared with a full-scale research run.
  • The entropy model is simplified.
  • The Huffman codebook is generated from the symbols being encoded.
  • The binary format is a prototype format rather than a standardized image codec.
  • The current pipeline primarily operates on 256 × 256 images/crops.
  • The probability model does not use a hyperprior or autoregressive context model.
  • The current learned model does not outperform JPEG in the evaluated experiment.
  • The current BPP experiments should not be interpreted as a production codec benchmark.
  • The implementation is optimized for understanding the compression pipeline rather than maximum compression performance.

22. Future Work

Several extensions can make the system substantially stronger.

Stronger Analysis and Synthesis Transforms

Use deeper and more expressive encoder and decoder architectures.

Hyperprior-Based Entropy Modeling

Introduce a learned hyperprior to improve probability estimation of the latent representation.

A future architecture could follow:

Image
  ↓
Analysis Transform
  ↓
Latent y
  ↓
Hyperprior
  ↓
Probability Model
  ↓
Entropy Coding
  ↓
Bitstream

Context-Based Probability Modeling

Use neighboring latent values to improve symbol probability prediction.

Improved Rate-Distortion Training

Train the model directly with:

L = R + λD

for multiple values of λ.

This can produce dedicated operating points across a broader bitrate range.

Better Entropy Coding

Replace the current prototype Huffman implementation with a more sophisticated learned entropy coder.

Perceptual Objectives

Investigate perceptual losses and additional quality metrics beyond MSE, PSNR, and SSIM.

Arbitrary Image Resolution

Extend inference beyond the current 256 × 256 workflow.

Larger-Scale Training

Train the final architecture for substantially longer schedules using larger compute resources.

Broader Benchmarking

Compare against additional traditional and learned image codecs using standardized evaluation protocols.


23. Research Takeaway

The main outcome of this project is not that the current model beats JPEG.

It does not.

Instead, the project demonstrates the construction of a learned image-compression system from basic components:

Image
  ↓
Encoder
  ↓
Latent Representation
  ↓
Quantization
  ↓
Entropy Coding
  ↓
Compressed Bitstream
  ↓
Entropy Decoding
  ↓
Decoder
  ↓
Reconstructed Image

The experiments show how representation learning, quantization, entropy, bitrate, and reconstruction quality interact in a practical compression pipeline.


24. Author

Munna Kumar Sah

GitHub: doitmuna


25. License

This project is licensed under the MIT License.

See the LICENSE file for details.

About

A from-scratch deep learning image compression prototype using a convolutional autoencoder, latent quantization, entropy modeling, and Huffman coding.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages