
Built for Track 2 of the AMD Developer Hackathon, Nami is an autonomous video captioning agent engineered for speed and precision. Instead of taking a shotgun approach to video analysis, Nami surgically extracts exactly 1 frame per second to build a dense visual context array. Powered by the Gemma 3 vision model via Fireworks AI, Nami processes this context to generate formal, sarcastic, and humorous captions simultaneously in a single, cost-effective inference pass. Featuring both a headless batch runner for strict evaluation constraints and a sleek minimalist Streamlit UI, Nami delivers unparalleled visual understanding with zero latency bloat.
13 Jul 2026