🔊 Audio Description of Images using Deep Learning

Vision Transformer (ViT) encoder + GPT-2 decoder image captioning with text-to-speech narration — improving accessibility of visual content for visually impaired users. The model runs entirely in your browser.

ViT-base encoderGPT-2 decoderONNX / Transformers.jsWeb Speech API TTS
Click or drop an image here
JPEG / PNG — processed locally, never uploaded
preview
Loading model (~120 MB, first visit only)…

Architecture: the image is split into 16×16 patches encoded by a Vision Transformer whose self-attention captures global image context; a GPT-2 decoder attends to these features via cross-attention and generates the caption autoregressively; the Web Speech API narrates it as audio. Fine-tuned on Flickr8k-style captions (checkpoint: vit-gpt2-image-captioning, ONNX). Presented at ICAITPR-2024 · Full training & BLEU evaluation code on GitHub.