OmniCaption by VoxPersona: High-performance video intelligence built for style matching and contextual accuracy in the AMD ACT II challenge.
OmniCaption is a Dockerized, dual-model hybrid video captioning pipeline extracting visual and auditory evidence to synthesize grounded, stylized descriptions (formal, sarcastic, humorous) powered natively by AMD Instinct MI300X and local ROCm compute.