From Pixels to Sentences: A Deep Learning-Based Image Captioning Framework Using CNN and LSTM for Intelligent Visual Interpretation
Authors: Muhammad Inam ul Haq, Sayyed Talha Gohar Naqvi, Shehryar Qamar Paracha, Ubaidullah Akram, Abid Munir, Shahab Ahmad Niazi
Abstract
Quickly identifying the picture, its contents, and the actions of its objects is one of our strongest visual processing abilities. The development of AI has led us to the goal of having computers accomplish the same tasks automatically. Using Neural Networks and Natural Language Processing (NLP) techniques for image recognition is necessary for automatically generating captions for each photo. In recent years, natural language processing and computer vision researchers have focused heavily on the problem of creating textual descriptions for images. There have been other approaches proposed on this topic that rely on deep learning. These techniques use images annotated by humans to train and test the algorithms. For these models to function at their best, a large training dataset is necessary. Long Short-Term Memory (LSTM) and a Convolutional Neural Network (CNN) based deep learning model are the backbone of our technology that extracts image characteristics and deciphers visual descriptions.
