Developed a VQA model integrating BERT and VIT for processing textual and visual inputs, respectively