Topic Hub

Best Any-to-Any Open Source Projects

This hub collects the best open source projects in Any-to-Any and ranks them by both momentum and authority.

Data window: Last 7 days (with 24h tie-breakers for Trending now)

Last updated: Sep 10, 2026

Projects

21

GitHub stars

5K

Hugging Face likes

8.9K

Replicate runs

0

Authority projects

Projects with the strongest long-term signal in Any-to-Any, ranked by total stars, likes, or runs on their primary platform.

minimind-o

Multimodal model for text, audio, and visual input and output.

GitHub
1.9Kstars

parlor

Enables natural voice and vision conversations with an AI that runs entirely on your machine.

GitHub
1.6Kstars

google/gemma-4-12B-it

Multimodal AI model for text, audio, image, and video input and text output generation.

Hugging Face Model
1.5Klikes

google/gemma-4-E4B-it

Multimodal AI model for text and image input and text output generation.

Hugging Face Model
1.3Klikes

Lance

Generate videos from images and edit videos with a unified multimodal model.

GitHub
1.2Kstars

bytedance-research/Lance

Unified multimodal model for image and video understanding, generation, and editing.

Hugging Face Model
1Klikes

openbmb/MiniCPM-o-4_5

Multimodal language model for vision, speech, and live streaming applications.

Hugging Face Model
912likes

ideogram-ai/ideogram-4-fp8

Hugging Face Model

Hugging Face Model
687likes

google/gemma-4-12B

Multimodal model for text, audio, image, and video inputs and text output.

Hugging Face Model
660likes

google/gemma-4-E2B-it

Multimodal AI model for text and image processing and generation.

Hugging Face Model
635likes

unsloth/gemma-4-12B-it-qat-GGUF

Multimodal model for text and image input with text output generation.

Hugging Face Model
348likes

kimi-k3-mlx

GitHub

GitHub
316stars

Hidden gems

Smaller projects with unusually strong momentum. We look for lower total metrics plus positive 7-day growth.

Qwen/Qwen3.5-Omni-Offline-Demo

Interact with a multimodal AI using text, images, audio, or video to get responses and process files.

Hugging Face Space
116likes
427d

inclusionAI/Ming-flash-omni-2.0

Advanced multimodal understanding and synthesis model for various applications.

Hugging Face Model
254likes
367d

Qwen/Qwen3.5-Omni-Online-Demo

Interact with a multimodal AI using text, image, audio, or video to get responses and generate content.

Hugging Face Space
55likes
147d

google/gemma-4-26B-A4B-it-assistant

Multimodal AI model for text and image input with text output, suitable for low-latency applications.

Hugging Face Model
130likes
197d

google/gemma-4-31B-it-assistant

Multimodal AI model for text and image input and text output generation.

Hugging Face Model
262likes
307d

sensenova/SenseNova-U1-8B-MoT

Generate coherent text and images with a unified multimodal model.

Hugging Face Model
274likes
277d

nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16

Multimodal large language model for enterprise-grade content analysis and understanding.

Hugging Face Model
296likes
247d

unsloth/gemma-4-12B-it-qat-GGUF

Multimodal model for text and image input with text output generation.

Hugging Face Model
348likes
277d

meituan-longcat/LongCat-Next

Multimodal model for text, vision, and audio processing.

Hugging Face Model
152likes
77d

google/gemma-4-12B-it-qat-q4_0-gguf

Multimodal model for text and image input and text output generation.

Hugging Face Model
273likes
77d