image-caption-transformer kuanmol · COMPLETED
This project implements an image captioning system that extracts features from images using a pre-trained Vision Transformer (ViT-B-16) and generates captions using a Transformer model with beam search. It is trained and evaluated on the Flickr30k dataset, achieving descriptive and contextually relevant captions.
github.com/kuanmol/image-caption-transformer · ★ 1 · Forks 0 · Size 360 KB
SUMMARY
Technologies 4
Scored 4
Observed 4
Practices 6
Evidence 4
Skips 0
COVERAGE
Analyzed 13 files · 10 commits · 0 API calls
TECHNOLOGIES & DEPTH
Markdown LANGUAGE Depth 70
1 files · PRODUCTION
Python LANGUAGE Depth 70
7 files · PRODUCTION
NumPy LIBRARY Depth 62
2 files · PRODUCTION
pandas LIBRARY Depth 62
2 files · PRODUCTION
PRACTICES
documentation · observedautomated_tests · observedcontinuous_integration · absentcontainerization · absentlinting · absentformatting · absent
ACTIVITY & OWNERSHIP
First commit 2025-08-13
Last commit 2025-08-13
Active months 1
Commits 10