Browse applications built on modern technologies. Explore PoC and MVP applications created by our community and discover innovative use cases for modern technologies.
SovereignSwarm is an autonomous video captioning engine using Fireworks AI and OpenCV to extract chronological frames and generate structured, multi-tone captions (Formal, Sarcastic, Humorous Tech, Everyday) in an enterprise-ready JSON format.
TRANSDUCER watches any video clip and captions it in four voices - formal, sarcastic, humorous_tech, humorous_non_tech. One 1024px look, four persona briefs, grounded in what's on screen
This project is a simple video captioner. Uses a VLM to extract the text from a video and another model to transform the description in the required styles. As a fallback, it contains a fine-tunned local model based on Gemma 4.
A video captioning agent that watches a clip once, then rewrites what it saw as Formal, Sarcastic, or Humorous captions
Claption turns short videos into accurate captions in four tones: formal, sarcastic, humorous-tech, and humorous-non-tech, using Fireworks AI with a judge-and-repair loop.
Vision-based video captioning pipeline using Gemma 4 31B. Analyzes video frames,nbuilds scene understanding, generates 4 caption styles with automatic quality scoring.
An AI-powered video captioning platform that generates 4 distinct caption styles (Formal, Sarcastic, Tech-Humor, Casual-Humor) and burns them using FFmpeg, displayed in a synchronised 4-quadrant player with rich export options.
FourVoice generates four distinct caption styles — formal, sarcastic, tech-humor, and everyday-humor — for any short video clip, automatically grounding every caption in real audio and visual evidence rather than guessing.
An adaptive AI video agent that intelligently extracts frames and generates highly creative, persona-driven captions (formal, sarcastic, tech, and meme humor) using the Qwen Vision model.
Clip Insights is a Chrome extension, helps you create notes from YouTube videos, summarize, extracts key points, interactive video chat, and lets you capture screenshots and notes. Export it all into a PDF for fast, efficient video-based learning.
CaptionDB is an AI-powered video caption and scene analysis platform that automatically detects scenes, extracts keyframes, understands visual context, and generates high-quality captions using a scalable asynchronous AI pipeline.
Docker agent for Track 2: Kimi K2.6 builds JSON fact sheets from video frames; GLM 5.2 writes formal, sarcastic, humorous-tech & humorous-non-tech captions. Same pipeline runs Gemma-3-12B end-to-end on AMD MI300X via ROCm/vLLM (86s/8 clips).
EduWorld is an AI-powered 3D interactive learning platform that revolutionizes STEM education through immersive 3D models, AI tutoring, and AMD-accelerated compute for engaging, personalized learning.
A robust, Dockerized AI agent that automatically extracts visual frames from videos and leverages the Fireworks AI API to generate dynamic captions across multiple stylistic tones (formal, sarcastic, tech-humor, and everyday humor).
Reelbudget is an AI pre-production tool for indie filmmakers and enterprise marketing teams. Powered by the AMD MI300X, it instantly turns text briefs into true animated video storyboards and budget estimations, saving weeks of planning.
ReachIQ AI's Video RAG engine ingests competitor YouTube videos, builds a vector knowledge base, and generates strategic competitive intelligence reports — powered by Gemma 3 4B running natively on AMD Developer Cloud via ROCm + vLLM.
Every dense paper you've ever given up on, rewritten as the 3Blue1Brown video that would have made it click. It is generated automatically, covering every page, not just the flashy parts. Weeks worth of animations done in minutes.
An AI agent that watches a short video and captions it in four distinct voices: formal, sarcastic, tech-humor, and everyday-humor
An automated tool that watches video clips and writes creative captions for them in different writing styles.
An AI-powered video captioning system that extracts audio and key frames, uses Whisper for speech transcription and a multimodal model for visual analysis, and generates formal, sarcastic, and humorous captions.
An AI platform that creates video parodies by re-voicing existing footage or generating 9:16 vertical short-form videos from scratch using Moonshot Kimi K2.6, FLUX.1, and Google Gemma models with an asynchronous multi-threaded Streamlit interface.
A multi-model AI pipeline that analyzes video frames chronologically and synthesizes verified captions across four distinct tonal registers.
aionVIS turns one plain-English sentence into a trained, deployable object-detection model. An agent swarm generates its own training images, labels them, verifies its own labels, and trains the detector on a single AMD MI300X. Zero human annotation.
An AI launch assistant that transforms product images into story.
OmniCaption is a Dockerized, dual-model hybrid video captioning pipeline extracting visual and auditory evidence to synthesize grounded, stylized descriptions (formal, sarcastic, humorous) powered natively by AMD Instinct MI300X and local ROCm compute.
An AI-powered tool that samples frames from video files and generates captions in four distinct styles — formal, sarcastic, humorous-tech, and humorous-non-tech — using a vision-language model via the Fireworks AI API.
PersonaStudio AI maps a video's core meaning via text transcription or Gemma 4 vision into a reusable "Content DNA" payload. Users can instantly transform this single blueprint into multi-platform content without ever re-analyzing the video.
CaptionForge AI is a production-ready multimodal video captioning agent that analyzes video, speech, OCR text, scene transitions, and temporal context to generate accurate captions in formal, sarcastic, humorous-tech, and humorous-non-tech styles.
A containerised agent that samples keyframes from any video clip and generates four tonally distinct captions — formal, sarcastic, humorous-tech, and humorous-non-tech — grounded in what is actually on screen.
An AI-powered Video Captioning Agent that automatically generates accurate, context-aware captions, timestamps, and scene descriptions for videos, improving accessibility, searchability, and content understanding.