Vision Transformer (ViT) encoder + GPT-2 decoder image captioning with text-to-speech narration — improving accessibility of visual content for visually impaired users. The model runs entirely in your browser.
Architecture: the image is split into 16×16 patches encoded by a Vision Transformer whose self-attention captures global image context; a GPT-2 decoder attends to these features via cross-attention and generates the caption autoregressively; the Web Speech API narrates it as audio. Fine-tuned on Flickr8k-style captions (checkpoint: vit-gpt2-image-captioning, ONNX). Presented at ICAITPR-2024 · Full training & BLEU evaluation code on GitHub.