An AI agent development platform with all-in-one visual tools, simplifying agent creation, debugging, and deployment like never before. Coze your way to AI Agent creation.
-
Updated
Jul 29, 2026 - TypeScript
An AI agent development platform with all-in-one visual tools, simplifying agent creation, debugging, and deployment like never before. Coze your way to AI Agent creation.
🌌 Orion AI Workspace – A free intelligent workspace platform that combines advanced AI models with real-time collaboration tools, designed with privacy-first principles and user-controlled API keys.
Render agent context as images. Same text, ~60–75% fewer tokens.
AI StoryTeller is a multimodal AI application that converts images into creative short stories by combining computer vision and natural language generation. The system uses a pretrained image captioning model to understand visual content and Google Gemini to generate context-aware narratives grounded in the image.
A high-performance, multimodal AI desktop assistant for Windows powered by Gemini 2.5 Flash, featuring real-time audio processing, a reactive holographic UI, and system-level command execution.
An Agentic AI-powered visual storytelling system that generates rich, creative narratives from images using Multimodal LLMs + Retrieval-Augmented Generation (RAG).
Production-grade semantic video search engine - search across video content using natural language. Powered by Whisper, GPT-4o Vision, vector embeddings, and Pinecone.
AI-powered multimodal fashion analysis platform using CLIP embeddings, vector search, and Groq-powered RAG pipelines.
AI research initiative exploring computer vision, temporal reasoning, and multimodal AI for intelligent animal emergency detection.
Open-source dual-mode AI assistant for chat and voice, featuring conversational intelligence, multimodal interaction, and productivity-focused automation.
Agentic multimodal medical AI pipeline combining clinical speech, imaging, and text preprocessing with MedGemma-based structured radiology report generation.
GenAI turns waste (peels, grounds) into drugs <60s. Upload img/txt → fragments → structures → ADMET/EcoScore → RAG validate → PDF. Built: GPT-4o, Llama-3, LangChain, RDKit. Guided: Dr. Hammad Majeed (UMT Lahore). Hackathon 2025.
A Streamlit-based Multimodal AI Generator using Google's Gemini API for text and image generation.
MindTrack is an AI-powered multimodal emotion detection system using both text and images to monitor emotional well-being in real time.
An immersive, real-time voice interface for ship crews and naval command. Combines Gemini Live WebSocket streaming, function-calling RAG, and autonomous visual asset generation into a sleek React frontend.
RAG MCP Frontend — a lightweight React/TypeScript frontend for interacting with Retrieval-Augmented Generation (RAG) services and the MCP (Multi-Channel Processing) backend. This project offers a clean UI for document ingestion, query/response flows, conversation history.
Hệ thống Hỏi đáp trực quan (VQA). Mô hình AI đa phương thức kết hợp Thị giác máy tính (CNN) và Xử lý ngôn ngữ tự nhiên (LSTM) để trả lời câu hỏi dựa trên nội dung hình ảnh.
Add a description, image, and links to the multimodel-ai topic page so that developers can more easily learn about it.
To associate your repository with the multimodel-ai topic, visit your repo's landing page and select "manage topics."