Top Builders

Explore the top contributors showcasing the highest number of app submissions within our community.

LLaVA: Large Language and Vision Assistant

LLaVA represents a novel end-to-end trained large multimodal model that combines a vision encoder and Vicuna for general-purpose visual and language understanding, achieving impressive chat capabilities mimicking spirits of the multimodal GPT-4 and setting a new state-of-the-art accuracy on Science QA.

General
Relese dateNovember 20, 2023
Repositoryhttps://github.com/haotian-liu/LLaVA
TypeMultimodal Language and Vision Model

What is LLaVA?

Visual Instruction Tuning: LLaVA, short for Large Language-and-Vision Assistant, represents a significant leap in multimodal AI models.

With a focus on visual instruction tuning, LLaVA has been engineered to rival the capabilities of GPT-4V, demonstrating its exceptional prowess in understanding both language and vision. This state-of-the-art model excels in tasks ranging from impressive chatbot interactions to setting a new standard in science question-answering accuracy, achieving a remarkable 92.53%. With LLaVA's innovative approach to instruction-following data and the effective combination of vision and language models, it promises a versatile solution for diverse applications, marking a significant milestone in the field of multimodal AI.

LLaVA Tutorials


LLaVA Libraries

A curated list of libraries and technologies to help you build great projects with 'technology'.


LLavA AI technology page Hackathon projects

Discover innovative solutions crafted with LLavA AI technology page, developed by our community members during our engaging hackathons.

LUMINA Vision Node β€” Alzheimer AI Monitor

LUMINA Vision Node β€” Alzheimer AI Monitor

LUMINA Vision Node is an AI-powered patient safety system for Alzheimer and dementia patients. 55 million people worldwide live with Alzheimer's β€” families ask daily: "Is my loved one safe?" LUMINA answers automatically, in real time. CORE TECHNOLOGY: LUMINA uses a hybrid computer vision pipeline. Primary detection: face recognition via dlib-based 128-dimensional encodings. Fallback: clothing color fingerprinting using HSV histogram back-projection β€” tracks the patient even when face is not visible. This dual-layer approach minimizes false negatives without requiring wearable devices. KEY FEATURES: - Live CCTV Monitoring β€” MJPEG stream at 25 FPS with real-time bounding box overlay - Face Recognition β€” Patient confirmed via face encoding; prevents false positives from guests - Color Fingerprint Fallback β€” HSV back-projection tracks patient when face is not visible - GPS Safe Zone β€” Haversine-formula geofencing alerts family when patient leaves home radius - Emotion Detection β€” Automatic inference: calm, anxious, or confused from activity patterns - SOS Emergency β€” Single-tap emergency with GPS coordinates sent to family dashboard - Memory Gallery β€” Family uploads photos to help patient recall loved ones - Stage Adaptation β€” AI reports adapt to Alzheimer stage (1=Mild, 2=Moderate, 3=Severe) AMD INTEGRATION: LUMINA uses a provider-based architecture for vision inference. The system routes snapshot analysis to AMD Developer Cloud running LLaVA vision model via vLLM and ROCm. Switching to AMD hardware requires only updating two environment variables (AMD_VISION_URL, AMD_MODEL_NAME) β€” zero code changes needed. Architecture is deployment-ready for AMD MI300X with 192GB HBM3. IMPACT: Built for 160 million family caregivers worldwide who cannot afford 24/7 professional care. LUMINA works with existing CCTV β€” no new hardware required. All patient data stored locally for maximum privacy. Because dignity should not depend on wealth.

Sentinel Generalist:

Sentinel Generalist:

Sentinel Generalist is a zero-shot agricultural intelligence system powered by a large Vision-Language Model (Qwen2.5-VL-7B-Instruct) running on an AMD MI300X GPU. Upload a single photograph of any plant β€” indoors, outdoors, healthy, or dying β€” and Sentinel performs 11 simultaneous agricultural analyses in real time, streaming its step-by-step reasoning trace like an expert agronomist thinking out loud. Reasoning trace β€” watch the AI examine shadows, soil texture, leaf color, and turgor in real time Species identification with confidence scores (50,000+ species, zero-shot) Geospatial inference β€” USDA hardiness zone, latitude, climate class deduced from shadows and soil Light assessment β€” estimated daily sun hours + adequacy check Watering analysis β€” visual turgor + soil moisture cues Nutrient deficiency β€” N-P-K + micronutrients with organic remedy doses Pest & disease β€” IPM-first diagnosis (organic β†’ chemical β†’ prevention) Companion planting β€” recommended neighbors + antagonists + placement instructions Harvest intelligence β€” days to harvest, visual cues, succession planting Seasonal planning β€” frost risk, crop rotation, next-season prep Beginner garden planner β€” "What should I plant right now?" with 7 curated plants for your zone and season Garden harvest preview β€” interactive top-down layout + AI-generated image of your garden at peak harvest Why AMD MI300X? The 192GB HBM3 VRAM loads the full 7B model with room to spare. ROCm 7.0 + gfx942 optimization delivers ~23% faster inference than comparable cloud APIs. Your plant photos never leave the Droplet β€” privacy-first by design. The demo differentiator: Tech stack: Qwen2.5-VL-7B-Instruct Β· AMD MI300X Β· ROCm 7.0 Β· HuggingFace Optimum-AMD Β· FastAPI Β· React 19 Β· GitHub Pages Β· NDJSON streaming Β· 100/100 Lighthouse Β· zero trackers Β· no CDNs