A from-scratch image captioning system using a custom ResNet-18-style CNN encoder and an LSTM-based caption decoder trained on Flickr8k.
Project status: Baseline model completed, evaluated, and packaged for local inference and Streamlit deployment.
Image captioning is the task of generating a natural-language description of an image.
This project implements the complete image captioning pipeline without using pretrained model weights:
Image
↓
Custom ResNet Encoder
↓
512-dimensional Image Feature
↓
Feature Projection
↓
LSTM Decoder
↓
Vocabulary Probabilities
↓
Generated Caption
The project was built to understand the complete system from the model architecture to training, evaluation, inference, and deployment rather than relying on a pretrained image-captioning pipeline.
The main goals of the project were to:
- Implement the CNN image encoder from scratch
- Implement the LSTM caption decoder from scratch
- Understand residual learning and skip connections
- Build the vocabulary and caption-tokenization pipeline
- Train the system using teacher forcing
- Evaluate generated captions quantitatively
- Analyze qualitative failure cases
- Build a reusable inference pipeline
- Package the trained model for deployment
- Build a Streamlit application
IMAGE
│
▼
┌──────────────────────┐
│ Custom ResNet Encoder│
└──────────────────────┘
│
▼
512-D Image Feature
│
┌─────────┴─────────┐
▼ ▼
Feature → Hidden Feature → Cell
│ │
└─────────┬─────────┘
▼
Initial LSTM State
│
▼
Word Embedding
│
▼
LSTM
│
▼
Linear Projection
│
▼
Vocabulary Probabilities
│
▼
Generated Caption
The image encoder is a custom ResNet-18-style architecture implemented without pretrained weights.
Input
224 × 224 × 3
↓
Stem
↓
64 × 56 × 56
↓
Residual Layer 1
64 × 56 × 56
↓
Residual Layer 2
128 × 28 × 28
↓
Residual Layer 3
256 × 14 × 14
↓
Residual Layer 4
512 × 7 × 7
↓
Adaptive Average Pooling
↓
512-dimensional feature vector
The encoder contains approximately 11.18 million parameters.
The decoder uses:
| Component | Dimension |
|---|---|
| Image feature | 512 |
| Word embedding | 256 |
| LSTM hidden state | 512 |
| Vocabulary | 2,541 |
The image feature is projected into the initial hidden and cell states of the LSTM.
A standard convolutional network can learn increasingly abstract visual representations as depth increases.
ResNet introduces residual connections that allow a block to learn a residual transformation:
Input
│
├──────────────┐
│ │
▼ │
Convolution │
↓ │
BatchNorm │
↓ │
ReLU │
↓ │
Convolution │
↓ │
BatchNorm │
│ │
└────── + ─────┘
│
▼
ReLU
The skip connection allows the input to bypass the convolutional transformation.
In this project, the ResNet is used as a visual feature extractor rather than a classification network.
An image alone does not produce a sentence.
The decoder must generate a sequence:
<START>
↓
a
↓
boy
↓
is
↓
running
↓
through
↓
a
↓
field
↓
<END>
The LSTM maintains a hidden state and cell state that carry information through the sequence.
At each timestep:
Previous Token
+
Previous Hidden State
+
Previous Cell State
↓
LSTM
↓
Next Token Probability
The project uses the Flickr8k image-captioning dataset.
| Split | Images | Caption Samples |
|---|---|---|
| Training | 6,000 | 30,000 |
| Validation | 1,000 | 5,000 |
| Test | 1,000 | 5,000 |
Each image has five human-written reference captions.
The splits are performed at the image level so that images from the test set do not appear in the training set.
Captions are tokenized and surrounded by special tokens:
<START> caption words <END>
The vocabulary uses:
<PAD> = 0
<START> = 1
<END> = 2
<UNK> = 3
A minimum frequency threshold of 5 was used when constructing the vocabulary.
Final vocabulary size:
2,541 tokens
Words that do not occur in the final vocabulary are mapped to <UNK>.
Input images are resized to:
224 × 224
Training preprocessing includes random horizontal flipping.
Validation and test preprocessing is deterministic.
Images are normalized using:
mean = [0.5, 0.5, 0.5]
std = [0.5, 0.5, 0.5]
The baseline model was trained from scratch.
| Parameter | Value |
|---|---|
| Epochs | 15 |
| Batch Size | 32 |
| Optimizer | Adam |
| Learning Rate | 1e-4 |
| Weight Decay | 1e-5 |
| Loss | Cross-Entropy |
| Gradient Clipping | 5.0 |
Teacher forcing is used during training.
During training, the decoder receives the correct previous token and learns to predict the next token.
For example:
Input:
<START> a boy is
Target:
a boy is running
The objective is next-token prediction.
Captions have different lengths, so sequences are padded to a common length.
Padding tokens are ignored by the loss function:
CrossEntropyLoss(ignore_index=<PAD>)This prevents artificial learning from padded positions.
The baseline training loss decreased from:
Epoch 1:
4.4402
to:
Epoch 15:
2.4043
Validation loss decreased from:
Epoch 1:
3.8613
to:
Epoch 15:
2.9881
2.9881
The training and validation curves are available in:
results/training_validation_loss.png
The trained model was evaluated on:
1,000 unique test images
The generated captions were compared against the original human-written reference captions.
The project reports:
- BLEU-1
- BLEU-2
- BLEU-3
- BLEU-4
- METEOR
- ROUGE-L
| Metric | Score |
|---|---|
| BLEU-1 | 0.5136 |
| BLEU-2 | 0.3213 |
| BLEU-3 | 0.1955 |
| BLEU-4 | 0.1223 |
| METEOR | 0.3151 |
| ROUGE-L | 0.3833 |
These numbers represent the measured performance of this implementation on the specified Flickr8k test split.
They are not presented as state-of-the-art results.
Quantitative metrics alone do not fully describe captioning quality.
The project also includes qualitative inspection of generated captions.
The model generally learns broad scene-level information, but several recurring failure modes were observed.
- Wrong object
- Wrong action
- Wrong object count
- Wrong color or attribute
- Confusion between visually similar objects
- Difficulty representing interactions between multiple people
- Reduced accuracy on images outside the training distribution
Reference descriptions describe two boys riding a bicycle.
Generated caption:
a man in a blue shirt is standing on a <UNK> road
Reference descriptions describe a man driving an ATV.
Generated caption:
a man in a red shirt is riding a bike through the woods
Reference descriptions describe two football players interacting during a tackle.
Generated caption:
a football player is running on the field
These examples show that the model can capture the general scene while still failing to identify fine-grained objects, attributes, and interactions.
The model was also tested on images outside the Flickr8k test distribution.
For example, a modern photograph of a child running through a field produced:
a little girl in a pink shirt is running through a field
while the image actually showed a boy wearing a yellow shirt.
This demonstrates an important limitation of the current baseline:
The model can produce a plausible sentence that matches the general scene while getting specific visual attributes wrong.
This is one reason why qualitative analysis is included alongside automatic captioning metrics.
The inference pipeline uses greedy decoding.
Input Image
↓
Resize + Normalize
↓
ResNet Encoder
↓
Image Feature
↓
Initialize LSTM
↓
<START>
↓
Predict next token
↓
Feed token back into LSTM
↓
Repeat
↓
<END>
↓
Generated Caption
At each step, the token with the highest predicted probability is selected.
The inference pipeline can also expose the probability assigned to each generated token.
Example:
two 69.50%
black 75.98%
dogs 87.25%
are 25.77%
running 54.36%
This makes the inference pipeline useful not only as a demo but also as a simple way to inspect model decisions.
The project includes an interactive Streamlit interface.
The application provides:
- Image upload
- Image preview
- Caption generation
- Token-level probability display
- Model/device information
Supported image formats include:
JPG
JPEG
PNG
WEBP
BMP
TIF
TIFF
GIF
The deployment application loads the trained checkpoint and vocabulary directly from the project repository.
image-captioning-resnet/
│
├── app.py
├── README.md
├── requirements.txt
├── .gitignore
│
├── config/
│ ├── vocabulary.json
│ └── inference_config.json
│
├── models/
│ └── final_model.pth
│
├── results/
│ ├── evaluation_metrics.json
│ ├── failure_cases.json
│ ├── generated_captions.json
│ └── training_validation_loss.png
│
└── src/
├── decoder.py
├── inference.py
├── model.py
├── resnet.py
└── vocabulary.py
Clone the repository:
git clone https://github.com/doitmuna/image-captioning-resnet.git
cd image-captioning-resnetCreate a virtual environment:
python -m venv .venv.venv\Scripts\activatepip install -r requirements.txtstreamlit run app.pyThe trained model is stored in:
models/final_model.pth
Vocabulary:
config/vocabulary.json
Inference configuration:
config/inference_config.json
The Flickr8k dataset itself is intentionally not included in the repository.
Important artifacts were saved separately during development, including:
- trained model checkpoints
- vocabulary
- inference configuration
- training history
- evaluation metrics
- generated captions
- qualitative failure cases
The project can therefore be inspected and run without retraining the baseline model.
The current architecture compresses the entire image into a single 512-dimensional feature vector before passing the representation to the LSTM decoder.
This creates a bottleneck for fine-grained spatial understanding.
The baseline therefore has difficulty with:
- object identity
- object counting
- object attributes
- object interactions
- visually similar objects
- detailed actions
- images substantially different from the training distribution
These limitations are intentionally documented rather than hidden.
Potential improvements include:
Instead of collapsing the CNN representation immediately into a single vector:
Current:
ResNet
↓
Global Average Pooling
↓
512-D Vector
↓
LSTM
an attention-based extension could preserve spatial information:
ResNet Feature Map
↓
Spatial Attention
↓
Weighted Visual Features
↓
LSTM Decoder
This would allow the decoder to focus on different regions of the image while generating different words.
The current inference implementation uses greedy decoding.
A future version can compare greedy decoding with beam search.
The system can be trained on larger datasets such as Flickr30k to investigate whether additional visual-language diversity improves generalization.
Future experiments can include additional captioning metrics such as CIDEr and more systematic human evaluation.
The qualitative failure cases provide specific directions for further experimentation:
Wrong object
↓
Investigate visual representation
Wrong attribute
↓
Investigate spatial information
Wrong interaction
↓
Investigate multi-object relationships
The current implementation can serve as a reproducible baseline for future experiments.
A natural next experiment is:
Baseline
ResNet Encoder
↓
Global Feature Vector
↓
LSTM
↓
Caption
versus:
Attention Extension
ResNet Feature Map
↓
Spatial Attention
↓
LSTM
↓
Caption
The experimental question becomes:
Does preserving spatial visual information through attention improve image captioning compared with the global-feature baseline?
This provides a clear baseline → hypothesis → implementation → evaluation workflow.
This project demonstrates practical experience with:
- Convolutional neural networks
- Residual learning
- LSTM sequence modeling
- Word embeddings
- Teacher forcing
- Cross-entropy optimization
- Sequence padding
- Vocabulary construction
- Greedy decoding
- Model checkpointing
- Quantitative evaluation
- Qualitative error analysis
- PyTorch
- Streamlit
- Reproducible inference
The project reports the actual measured performance of the implemented baseline.
No claim of state-of-the-art performance is made.
The model's limitations and failure cases are explicitly documented as part of the project analysis.
Munna Kumar Shah
Computer Science and Engineering
GitHub:
This project is intended for educational and research purposes.