
1
1
India
3+ years of experience
I am a 4th-year Computer Engineering student passionate about technology and innovation. Currently, I am building a strong foundation in programming, problem-solving, and software development, while exploring fields such as web development, artificial intelligence, machine learning, and cybersecurity. I enjoy continuous learning, collaborating on projects, and contributing to impactful solutions. I am actively seeking internships, mentorship, and networking opportunities to grow and make meaningful strides in the tech industry.

This project is an AI-powered video captioning solution designed to automatically generate high-quality captions and summaries for short video clips ranging from 30 seconds to 2 minutes. Instead of producing only a single description, the system creates four unique caption styles for every video: formal, sarcastic, humorous-tech, and humorous-non-tech. This allows the same content to be presented for different audiences, platforms, and use cases. The solution combines multimodal video understanding with modern language models to analyze visual scenes, actions, objects, temporal events, and contextual information before generating style-specific captions. Each generated caption preserves the core meaning of the video while adapting vocabulary, tone, and writing style according to the requested format. Formal captions focus on clarity and accuracy, sarcastic captions add witty commentary, humorous-tech captions incorporate software and engineering references, and humorous-non-tech captions provide lighthearted jokes that are easy for a general audience to understand. The system is designed to process multiple videos efficiently through an automated pipeline that accepts video inputs, performs inference, and returns structured outputs suitable for evaluation. Prompt engineering, video preprocessing, and output validation help maintain consistency across different caption styles while reducing hallucinations and preserving factual correctness. The architecture can also be extended with fine-tuned models or custom-trained captioning components using open datasets.
13 Jul 2026