Topic Hub
Best Any-to-Any Open Source Projects
This hub collects the best open source projects in Any-to-Any and ranks them by both momentum and authority.
Data window: Last 7 days (with 24h tie-breakers for Trending now)
Last updated: Sep 10, 2026
Projects
21
GitHub stars
5K
Hugging Face likes
8.9K
Replicate runs
0
Trending now
Fastest-growing projects in Any-to-Any over the last 7 days, with 24h change as a tie-breaker.
parlor
Enables natural voice and vision conversations with an AI that runs entirely on your machine.
Lance
Generate videos from images and edit videos with a unified multimodal model.
openbmb/MiniCPM-o-4_5
Multimodal language model for vision, speech, and live streaming applications.
Qwen/Qwen3.5-Omni-Offline-Demo
Interact with a multimodal AI using text, images, audio, or video to get responses and process files.
google/gemma-4-E2B-it
Multimodal AI model for text and image processing and generation.
inclusionAI/Ming-flash-omni-2.0
Advanced multimodal understanding and synthesis model for various applications.
google/gemma-4-12B-it
Multimodal AI model for text, audio, image, and video input and text output generation.
google/gemma-4-31B-it-assistant
Multimodal AI model for text and image input and text output generation.
google/gemma-4-E4B-it
Multimodal AI model for text and image input and text output generation.
sensenova/SenseNova-U1-8B-MoT
Generate coherent text and images with a unified multimodal model.
Hidden gems
Smaller projects with unusually strong momentum. We look for lower total metrics plus positive 7-day growth.
Qwen/Qwen3.5-Omni-Offline-Demo
Interact with a multimodal AI using text, images, audio, or video to get responses and process files.
inclusionAI/Ming-flash-omni-2.0
Advanced multimodal understanding and synthesis model for various applications.
Qwen/Qwen3.5-Omni-Online-Demo
Interact with a multimodal AI using text, image, audio, or video to get responses and generate content.
google/gemma-4-26B-A4B-it-assistant
Multimodal AI model for text and image input with text output, suitable for low-latency applications.
google/gemma-4-31B-it-assistant
Multimodal AI model for text and image input and text output generation.
sensenova/SenseNova-U1-8B-MoT
Generate coherent text and images with a unified multimodal model.
nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
Multimodal large language model for enterprise-grade content analysis and understanding.
unsloth/gemma-4-12B-it-qat-GGUF
Multimodal model for text and image input with text output generation.
meituan-longcat/LongCat-Next
Multimodal model for text, vision, and audio processing.
google/gemma-4-12B-it-qat-q4_0-gguf
Multimodal model for text and image input and text output generation.