A from-scratch research prototype for learned image compression using a convolutional autoencoder, latent-space quantization, entropy modeling, and Huffman coding.
The project explores how a neural network can learn a compact representation of an image and how that representation can be transformed into an actual compressed binary bitstream.
Status: Research prototype / baseline implementation
Focus: Understanding and implementing the complete learned image-compression pipeline from first principles
Traditional image codecs such as JPEG use hand-designed transforms, quantization, and entropy coding.
This project investigates a learned alternative in which a neural network learns an image representation directly from data.
The complete pipeline is:
COMPRESSION
Input Image
│
▼
┌──────────────┐
│ CNN Encoder │
└──────────────┘
│
▼
Latent Representation
│
▼
Quantization
│
▼
Discrete Latent Symbols
│
▼
Huffman Coding
│
▼
Compressed Binary Stream
Decompression follows the reverse process:
DECOMPRESSION
Compressed Binary Stream
│
▼
Huffman Decoding
│
▼
Quantized Latent Symbols
│
▼
Dequantization
│
▼
CNN Decoder
│
▼
Reconstructed Image
The main question explored in this project is:
Can a convolutional neural network learn a compact image representation that allows an image to be compressed into a small binary representation while maintaining useful reconstruction quality?
The project therefore studies three connected ideas:
- Learning a compact latent representation.
- Reducing that representation through quantization.
- Entropy-coding the resulting discrete symbols.
- Convolutional image encoder
- Convolutional image decoder
- Learned latent representation
- Latent-space quantization
- Gaussian entropy-model experiment
- Huffman entropy coding
- Actual binary
.bincompressed stream - End-to-end compression and decompression
- MSE, PSNR, SSIM, and BPP evaluation
- Rate-distortion experiments
- JPEG baseline comparison
- DIV2K-based experimentation
- Reusable Python source modules
- Colab research notebook
- Local Python project structure
The encoder transforms an image into a learned latent representation.
The encoder mapping is:
y = fθ(x)
where:
- x is the input image
- fθ is the encoder
- y is the latent representation
- θ represents the encoder parameters
For a 256 × 256 RGB image, the current encoder produces:
Input
3 × 256 × 256
↓
Conv2D + ReLU
64 × 128 × 128
↓
Conv2D + ReLU
128 × 64 × 64
↓
Conv2D + ReLU
192 × 32 × 32
↓
Conv2D + ReLU
128 × 16 × 16
↓
Conv2D
96 × 16 × 16
Therefore, the current latent representation has dimensions:
96 × 16 × 16
The encoder reduces the spatial resolution while increasing the number of learned feature channels.
The decoder reconstructs an approximation of the original image.
The decoder mapping is:
x̂ = gφ(y)
where:
- y is the latent representation
- gφ is the decoder
- x̂ is the reconstructed image
- φ represents the decoder parameters
The current decoder progressively increases the spatial resolution:
96 × 16 × 16
↓
128 × 32 × 32
↓
192 × 64 × 64
↓
128 × 128 × 128
↓
64 × 256 × 256
↓
3 × 256 × 256
Before compression can be evaluated, the autoencoder must first learn to reconstruct images.
The reconstruction distortion is measured using Mean Squared Error:
D(x, x̂) = (1/N) Σᵢ (xᵢ − x̂ᵢ)²
The baseline reconstruction objective is:
L_reconstruction = D(x, x̂)
The network learns encoder and decoder parameters that minimize this reconstruction error.
The encoder produces continuous-valued latent features.
For compression, these values are converted into discrete symbols using quantization.
The current implementation uses:
ŷ = round(y / q)
where q is the quantization step.
During reconstruction, the quantized latent is converted back to the original scale:
y_dequantized = q · ŷ
A smaller q preserves more information, while a larger q produces stronger quantization and generally increases distortion.
After quantization, the latent representation consists of discrete symbols.
The project includes a simple Gaussian entropy model that estimates the probability mass of quantized latent values.
For a discrete symbol s with probability p(s), entropy is:
H(S) = −Σₛ p(s) log₂ p(s)
An estimated coding rate can be written as:
R = −Σᵢ log₂ p(ŷᵢ)
and normalized by image pixels:
BPP = R / (H × W)
The Gaussian entropy model is used as an experimental probability and rate-estimation component.
The actual binary demonstration uses Huffman coding over the quantized latent symbols.
The final prototype converts quantized latent symbols into an actual binary stream using Huffman coding.
The process is:
Quantized Latent
↓
Symbol Frequency Counting
↓
Huffman Tree
↓
Binary Codes
↓
Encoded Bitstream
↓
.bin File
The decoding process is lossless with respect to the quantized latent symbols.
Therefore:
Quantized Latent
↓
Huffman Encode
↓
Binary Bitstream
↓
Huffman Decode
↓
Same Quantized Latent
The image compression itself remains lossy because information is discarded during the learned transformation and quantization stages.
A central idea in image compression is the trade-off between bitrate and reconstruction quality.
The general rate-distortion objective is:
L = R + λD
where:
- R represents the coding rate
- D represents reconstruction distortion
- λ controls the rate-distortion trade-off
A system may accept greater distortion in exchange for a lower bitrate, or spend more bits to preserve more image information.
The project investigates this relationship using multiple quantization strengths.
The project uses the DIV2K high-resolution image dataset.
The dataset is used for:
- Model development
- Training
- Validation
- Compression experiments
The training pipeline extracts random:
256 × 256
image crops.
The dataset is not included in this repository because of its size.
The current workflow is:
DIV2K Image
↓
Random 256 × 256 Crop
↓
RGB Image
↓
Normalize to [0, 1]
↓
Mini-batch
↓
CNN Encoder
↓
Latent Representation
↓
CNN Decoder
↓
Reconstructed Image
↓
MSE Loss
↓
Backpropagation
After the reconstruction baseline is established, the latent representation is quantized and entropy-coded.
MSE = (1/N) Σᵢ (xᵢ − x̂ᵢ)²
Lower values indicate smaller pixel-wise reconstruction error.
The images are normalized to the range [0, 1].
Therefore:
PSNR = 10 log₁₀(1 / MSE)
Higher values indicate better reconstruction fidelity.
SSIM measures structural similarity between the original and reconstructed images.
Higher values indicate greater structural similarity.
The bitrate is normalized by image area:
BPP = coded bits / (H × W)
Lower BPP indicates a more compact representation.
For raw 8-bit RGB data:
CR = original raw size / compressed size
The compression ratios reported by this prototype are measured relative to raw RGB data, not against an already compressed JPEG or PNG file.
A full validation evaluation produced the following current baseline:
| Metric | Result |
|---|---|
| BPP | 1.2898 |
| PSNR | 23.10 dB |
| SSIM | 0.6250 |
| MSE | 0.007518 |
These values represent the current learned-compression baseline on the DIV2K validation set.
They should be interpreted as experimental baseline results rather than state-of-the-art results.
The current validation experiment produced the following operating points:
| Quantization Step | BPP | PSNR | SSIM |
|---|---|---|---|
| 0.05 | 2.2148 | 22.70 dB | 0.6053 |
| 0.10 | 1.2837 | 23.26 dB | 0.6258 |
| 0.20 | 1.0483 | 22.72 dB | 0.5980 |
| 0.50 | 1.0232 | 21.78 dB | 0.5852 |
| 1.00 | 1.0358 | 18.58 dB | 0.4919 |
As quantization becomes more aggressive, reconstruction quality decreases substantially.
The BPP values in this experiment come from the current experimental probability/rate-estimation setup and should not be interpreted as a production codec bitrate.
The project also generates an actual Huffman-coded .bin file.
For one unseen test image, one run produced:
| Metric | Result |
|---|---|
| Compressed size | 10,254 bytes |
| BPP | 1.2517 |
| Compression ratio | 19.17× |
| PSNR | 21.02 dB |
| SSIM | 0.6274 |
The compression ratio is measured against raw 8-bit RGB data for a 256 × 256 image.
The Huffman decoder was verified to recover the exact same quantized latent symbols.
JPEG was evaluated on the same DIV2K validation images.
| JPEG Quality | BPP | PSNR | SSIM |
|---|---|---|---|
| 10 | 0.3961 | 27.34 dB | 0.8091 |
| 25 | 0.6377 | 31.34 dB | 0.8946 |
| 50 | 0.9719 | 33.34 dB | 0.9238 |
| 75 | 1.4307 | 35.05 dB | 0.9475 |
| 90 | 2.3661 | 38.53 dB | 0.9707 |
The current learned baseline does not outperform JPEG in this experiment.
For example:
Learned baseline
≈ 1.29 BPP
≈ 23.10 dB PSNR
JPEG
≈ 0.97 BPP
≈ 33.34 dB PSNR
JPEG therefore provides substantially better rate-distortion performance for the current experimental setup.
This result is useful because it establishes a working learned-compression baseline and identifies the areas that require further improvement.
The current rate-distortion comparison is stored at:
experiments/plots/rate_distortion_comparison.png
The plot compares the learned compression baseline and JPEG using BPP and PSNR.
The prototype can process an unseen image through the complete pipeline:
Input Image
↓
CNN Encoder
↓
Latent Representation
↓
Quantization
↓
Huffman Coding
↓
Compressed .bin File
↓
Huffman Decoding
↓
Quantized Latent
↓
CNN Decoder
↓
Reconstructed Image
A demonstration image and associated metrics are included in the project outputs.
learned-image-compression/
│
├── README.md
├── LICENSE
├── requirements.txt
├── .gitignore
│
├── notebooks/
│ └── image_compression.ipynb
│
├── src/
│ ├── __init__.py
│ ├── model.py
│ ├── entropy.py
│ ├── huffman.py
│ ├── metrics.py
│ └── inference.py
│
├── experiments/
│ ├── results/
│ │ ├── final_results.json
│ │ ├── final_comparison.json
│ │ └── final_demo_metrics.json
│ │
│ └── plots/
│ └── rate_distortion_comparison.png
│
└── demo/
Contains:
- Encoder
- Decoder
- Autoencoder
Contains:
- Huffman tree construction
- Code generation
- Encoding
- Decoding
- Bit-to-byte conversion
Contains the experimental Gaussian entropy model used for probability and rate estimation.
Contains:
- MSE
- PSNR
- SSIM
Contains reusable utilities for:
- Image loading
- Latent encoding
- Quantization
- Huffman compression
- Latent decoding
- Reconstruction
Clone the repository:
git clone https://github.com/doitmuna/learned-image-compression.git
cd learned-image-compressionCreate a virtual environment:
python -m venv .venvActivate it on Windows:
.venv\Scripts\activateInstall the dependencies:
pip install -r requirements.txtThe complete experimental development notebook is available at:
notebooks/image_compression.ipynb
The notebook contains the development and evaluation workflow used to build the prototype.
The dataset itself is not included in the repository.
To reproduce the experiments:
- Install the dependencies from
requirements.txt. - Download the required DIV2K training and validation images.
- Place the images in the directories expected by the notebook.
- Open
notebooks/image_compression.ipynb. - Run the preprocessing, training, evaluation, and compression experiments.
Large datasets, trained checkpoints, and generated binary files are intentionally excluded from Git version control.
This project demonstrates the path from neural image representation to an actual compressed bitstream:
Image
↓
Learned Representation
↓
Quantization
↓
Entropy Modeling
↓
Entropy Coding
↓
Binary Stream
It also demonstrates an important practical lesson in compression research:
A smaller representation does not automatically imply better compression.
A useful compression system must jointly consider:
- Representation quality
- Quantization
- Probability modeling
- Entropy coding
- Reconstruction quality
- Actual coded size
The current implementation is intentionally a baseline research prototype.
Important limitations include:
- The encoder and decoder are relatively small compared with modern learned image-compression architectures.
- The final training schedule is limited compared with a full-scale research run.
- The entropy model is simplified.
- The Huffman codebook is generated from the symbols being encoded.
- The binary format is a prototype format rather than a standardized image codec.
- The current pipeline primarily operates on 256 × 256 images/crops.
- The probability model does not use a hyperprior or autoregressive context model.
- The current learned model does not outperform JPEG in the evaluated experiment.
- The current BPP experiments should not be interpreted as a production codec benchmark.
- The implementation is optimized for understanding the compression pipeline rather than maximum compression performance.
Several extensions can make the system substantially stronger.
Use deeper and more expressive encoder and decoder architectures.
Introduce a learned hyperprior to improve probability estimation of the latent representation.
A future architecture could follow:
Image
↓
Analysis Transform
↓
Latent y
↓
Hyperprior
↓
Probability Model
↓
Entropy Coding
↓
Bitstream
Use neighboring latent values to improve symbol probability prediction.
Train the model directly with:
L = R + λD
for multiple values of λ.
This can produce dedicated operating points across a broader bitrate range.
Replace the current prototype Huffman implementation with a more sophisticated learned entropy coder.
Investigate perceptual losses and additional quality metrics beyond MSE, PSNR, and SSIM.
Extend inference beyond the current 256 × 256 workflow.
Train the final architecture for substantially longer schedules using larger compute resources.
Compare against additional traditional and learned image codecs using standardized evaluation protocols.
The main outcome of this project is not that the current model beats JPEG.
It does not.
Instead, the project demonstrates the construction of a learned image-compression system from basic components:
Image
↓
Encoder
↓
Latent Representation
↓
Quantization
↓
Entropy Coding
↓
Compressed Bitstream
↓
Entropy Decoding
↓
Decoder
↓
Reconstructed Image
The experiments show how representation learning, quantization, entropy, bitrate, and reconstruction quality interact in a practical compression pipeline.
Munna Kumar Sah
GitHub: doitmuna
This project is licensed under the MIT License.
See the LICENSE file for details.
