Browse applications built on OpenAI Whisper technology. Explore PoC and MVP applications created by our community and discover innovative use cases for OpenAI Whisper technology.
Vision-based video captioning pipeline using Gemma 4 31B. Analyzes video frames,nbuilds scene understanding, generates 4 caption styles with automatic quality scoring.
TruLens is a real-time AI overlay that listens to live media streams, instantly transcribes speech, and surfaces context-aware fact-checks directly to your screen to combat misinformation on the fly.
AI-powered emergency dispatch for Thailand's 191: live Thai transcription (Whisper), DeepSeek triage via Fireworks AI, location from speech alone (RAG landmarks, no GPS), one-click dispatch to the exact zone officer. Validated on AMD Instinct MI300X.
ReachIQ AI's Video RAG engine ingests competitor YouTube videos, builds a vector knowledge base, and generates strategic competitive intelligence reports — powered by Gemma 3 4B running natively on AMD Developer Cloud via ROCm + vLLM.
Upload a clip. Get the top 5 titles, descriptions, and hashtag sets — picked by an AI discriminator from 10 independently generated candidates per category, grounded in what your video actually says and shows and whats currently trending on youtube.
PersonaStudio AI maps a video's core meaning via text transcription or Gemma 4 vision into a reusable "Content DNA" payload. Users can instantly transform this single blueprint into multi-platform content without ever re-analyzing the video.
Two-stage video captioning agent: Chain-of-Verification grounding feeds four distinct-tone captions, with a self-judge regeneration pass and a fail-safe pipeline that guarantees valid, complete output on every run, even under model or infra failure.
Turn any travel idea into a complete trip. AITinerary uses AI to convert your preferences and Instagram/YouTube inspiration into personalized itineraries, hidden gems, budget planning, recommendations, and an intelligent travel companion.
Watches a clip and writes four captions in four different voices: formal, sarcastic, tech humor, casual humor. All four come from the same fact-checked scene description, using Gemma 4 through Hugging Face.
Containerized video captioning agent with AMD ROCm PyTorch support.
MonadLabs converts all data into data specialized for agents, integrates directly in company workflows, and reduces hallucination and token spending up to 85%. As API Bills surge into billions, it is the only way for enterprises to use AI efficiently.
Jardo is a voice-first assistant that supervises your coding agents (Claude Code, Gemini CLI): it reads your terminal, approves safe work, steers them back on track to user's goals, and talks in your language. Runs on Gemma on AMD Instinct GPUs via ROCm.
VoxFrame is a multimodal video captioning engine that turns clips into accurate, style-specific narratives. It uses FFmpeg, Groq Whisper, and AIMLAPI-powered Gemini models to verify visual context before generating polished captions.
Haven is an AI mental wellness companion: empathetic chat, journaling insights, voice therapy, breathing exercises, and clinical tools like PHQ-9 and GAD-7 — powered by open-source Llama & Qwen models with layered crisis detection.
Building Human-like AI companions with human-inspired cognition and self-learning capabilities that grow, adapt, and evolve through experience.
The only video captioning studio that QCs itself. Gemma locks the facts, then writes four calibrated voices — formal, sarcastic, tech & non-tech humor — each caption pre-judged for accuracy and tone before you see it. Gemma-first on Fireworks AI.
AI-conducted technical interviews that screen, proctor, and grade candidates live - on a fully self-hosted stack with no per-candidate vendor fees and no third-party audio/video pipeline.
AuraOS is an AI operating layer that seamlessly combines conversation, automation, and contextual intelligence to make interacting with your computer faster, smarter, and more natural.
CLARIS AI is an explainable multimodal video captioning system that combines vision, speech recognition, OCR, motion analysis, and AI reasoning to generate four natural, grounded caption styles from a single video with evidence-backed explanations.
Voice-control a swarm of headless AI coding agents from your phone. QuietVoice returns intent + emotion, not blind transcription — speech-to-intent and text-to-speech self-hosted on your own AMD GPU. Open source, runs on ROCm.
When towers fall, phones don't have to. OmniMesh tracks every victim, guides every responder, and gives command teams a live map of the whole disaster, all offline, with Gemma reasoning on AMD hardware two different ways so nothing goes unanswered.
A video captioning pipeline that extracts frames, transcribes audio, &generates captions in 4 distinct tones formal, sarcastic, humorous-tech, and humorous-non-tech using Minimax M3 via Fireworks AI, with a built-in LLM self-judge & Streamlit demo UI.
ASTRA is a hybrid Windows transcription app that filters audio locally, supports offline Whisper, and connects online to a license-protected server that batches jobs, routes them across speech providers, and automatically falls back when one fails.
Personalized hearing assistance tool that uses a patient's clinical audiogram to customize AI-powered voice enhancement — built for the SOCIELSA audiology ecosystem, running on AMD GPUs.
An on-device wellness co-pilot — task breakdown, energy prediction, hydration, and receipt scanning — with AI that runs entirely on your phone. No cloud, no accounts, nothing leaves your device. Works fully in airplane mode.
SellerKavach is an AI order intelligence layer for chat-based sellers. It intercepts messy WhatsApp/IG chats, extracts structured orders, predicts RTO risk before shipping, and builds a shared buyer trust network to stop fraud and losses.
An autonomous, multi-agent video captioning pipeline with edge-compute resilience. It combines Fireworks AI speed with a reliable local AMD CPU fallback to ensure zero-downtime, production-ready performance for automated content creation.
Real-time shield against phone scams and AI voice clones. Whisper + an open LLM score what's said; a wav2vec2 detector fine-tuned on AMD ROCm scores who's speaking. 98.8% detection on 3,393 real internet deepfakes, evaluated speaker and engine-disjoint.
A private, on-device memory assistant for cognitive decline. Using ExecuTorch on Snapdragon 8 Elite NPU, it identifies faces, transcribes conversations via Whisper, and generates warm recall cues via Llama 3.2 1B - fully offline, zero cloud.
Hermes is an Android translation assistant that scans text or speech and converts it into translated, accessible output on-device.