Top Builders

Explore the top contributors showcasing the highest number of app submissions within our community.

Stability AI's Stable Video

Stable Video is Stability AI’s pioneering venture into generative AI video models, designed for a broad spectrum of video applications in media, entertainment, education, and marketing. It enables users to convert text and image inputs into dynamic scenes, bringing concepts to life in a cinematic format.

General
Relese dateNovember 21, 2023
AuthorStability AI
RepositoryLink
TypeGenerative image-to-video model

Stable Video Specifications

Stable Video comes in two models, generating 14 and 25 frames respectively, at frame rates ranging from 3 to 30 FPS. These models have shown to surpass leading closed models in user preference studies.

  • Video duration: 2-5 seconds
  • Frame rate: Up to 30 FPS
  • Processing time: 2 minutes or less

Stable Video License

Stable Video Diffusion is available under a non-commercial community license. Review the full License and Stability’s Acceptable Use Policy here.

Model Sources

Evaluation

A user preference study comparing SVD Image-to-Video with GEN-2 and PikaLabs shows a higher preference for SVD in terms of video quality. Details are available in the research paper.

Uses

Intended for research, Stable Video can be utilized in generative model research, model safety, understanding model limitations, artistic creation, and educational tools. It is not designed for factual or true representations and should adhere to Stability AI's Acceptable Use Policy.

Limitations and Recommendations

The model produces short videos with potential limitations in motion, photorealism, and text rendering. It is primarily for research purposes.

Stability AI Stable Video AI technology Hackathon projects

Discover innovative solutions crafted with Stability AI Stable Video AI technology, developed by our community members during our engaging hackathons.

Gemma Captioner: Four Voices, One Truth

Gemma Captioner: Four Voices, One Truth

Most video captioning agents write all four styles in a single pass, so one visual mistake poisons every caption. And the funnier a caption tries to be, the more it tends to invent: a tech joke reaches for packets and uptime that were never there, a sarcastic line flips what actually happened. Track 2 scores both accuracy and tone, so a caption that reads well can still be wrong. We built it the other way around. Instead of adding style at the end, the agent uses Google Gemma-4-31B to look at the video once. It turns sampled frames into a grounded record of what is actually there: it reads clearly legible on-screen text like signs and logos, names concrete colors and counts, and flags whatever it cannot be sure of. Any detail a caption asserts that the frozen facts do not support is rejected and rewritten, so unsupported claims never survive. Only then do four parallel Gemma-4 text passes turn the same verified facts into formal, sarcastic, humorous-tech, and humorous-non-tech voices. A joke can change the tone but never the events, because the humor voice has to map its metaphor onto a real visible detail. The other hard part is speed. Long UHD clips are the trap, because a clip that finishes too late is graded as a placeholder. So the agent decodes only about ten frames per clip, aligned to the exact timestamps the judge samples for accuracy, rather than the whole video, and it captions multiple clips concurrently. A complete results.json sits on disk from the first moment and each clip upgrades its own entry as it finishes, so a slow or failed call degrades to a grounded fallback instead of a zero. The full set of clips finishes well inside the 10-minute budget. Gemma does the real work. Every scene fact and every caption comes from Gemma-4-31B: take it out and there is no product. It ships as a public linux/amd64 Docker image that runs headless, reading /input/tasks.json and writing /output/results.json.

ELIA Empowered Learning  Interaction Assistant

ELIA Empowered Learning Interaction Assistant

In early childhood (ages 3–10), teachers and parents play a vital role in shaping a child’s growth. But they often face daily challenges from creating engaging content and managing routines to handling emotional moments and simplifying complex topics for young minds. Many lack the time, tools, or creative energy to do it all. That’s where ELIA (Empowered Learning & Interaction Assistant) steps in — an AI-powered assistant designed to support both teachers and parents in nurturing children through smarter, more meaningful engagement. Using the Groq API, LLaMA models, and Blackbox.ai, ELIA provides real-time, personalized help for a wide range of teaching and parenting needs. Whether you're preparing for class or facing a parenting dilemma, ELIA is your creative, intelligent companion. 🔍 Key Features: 🧩 Weekly learning plans & fun activities based on child’s age/class 🍎 Nutrition, sleep, and emotional wellness support 📚 AI-generated stories, quizzes, and classroom content 👩‍🏫 Tips for simplifying complex ideas into child-friendly explanations 🏠 On-demand parenting advice to handle real-life situations with care 📝 Worksheets, rubrics, and teaching resources ready to use 🎯 What Makes ELIA Unique: Combines teaching and parenting support in one simple tool Built on Groq-powered LLaMA models for fast and smart content generation Developed with help from Blackbox.ai to enhance productivity and code quality Designed by students, inspired by real-life classroom and home challenges Warm, friendly, and safe — for both mentors and little learners ELIA isn’t just an assistant — it’s a bridge between care and curriculum, helping the people who shape tomorrow do it with ease, empathy, and excellence.