Topic Hub

Best Image-Text-to-Text Open Source Projects

This hub collects the best open source projects in Image-Text-to-Text and ranks them by both momentum and authority.

Data window: Last 7 days (with 24h tie-breakers for Trending now)

Last updated: Sep 18, 2026

Projects

95

GitHub stars

722

Hugging Face likes

91.3K

Replicate runs

77.8K

Authority projects

Projects with the strongest long-term signal in Image-Text-to-Text, ranked by total stars, likes, or runs on their primary platform.

kimi-k2.5

Multimodal language model for image and text understanding and generation.

Replicate
53.1Kruns

qwen3-7-plus

Replicate

Replicate
24.6Kruns

Qwen/Qwen3.8-27B

Hugging Face Model

Hugging Face Model
15.6Klikes

moonshotai/Kimi-K3

Hugging Face Model

Hugging Face Model
11.4Klikes

Qwen/Qwen3.8-Flash-Next

Hugging Face Model

Hugging Face Model
5.4Klikes

google/gemma-4-31B-it

Generate text output from image and text input using AI models.

Hugging Face Model
3.9Klikes

HauhauCS/Qwen3.6-35B-A3B-Uncensored-HauhauCS-Aggressive

Generates text based on image inputs with aggressive uncensoring.

Hugging Face Model
3.6Klikes

deepseek-ai/DeepSeek-V4.1-Flash

Hugging Face Model

Hugging Face Model
3.1Klikes

nvidia/LocateAnything-3B

Precise object localization and dense detection in images using vision-language models.

Hugging Face Model
3Klikes

moonshotai/Kimi-K2.5

Multimodal model integrating vision and language understanding for advanced agentic capabilities.

Hugging Face Model
2.4Klikes

zai-org/GLM-5.3-Flash

Hugging Face Model

Hugging Face Model
2.4Klikes

Qwen/Qwen3.6-27B

AI model for coding and natural language processing tasks with improved stability and utility.

Hugging Face Model
2.3Klikes

Hidden gems

Smaller projects with unusually strong momentum. We look for lower total metrics plus positive 7-day growth.

gemma-4-31b-it

Multimodal model for image and text understanding and generation.

Replicate
72runs
477d

ukisai/Swift-Qwen3.8-27b

Hugging Face Model

Hugging Face Model
412likes
4127d

LightVLM

Reconstruction of 3D scenes from images for simulation and visualization purposes.

GitHub
192stars
1927d

ukisai/Swift-Qwen3.8-27B-GGUF

Hugging Face Model

Hugging Face Model
248likes
2487d

microsoft/mage-vl-demo

Hugging Face Space

Hugging Face Space
98likes
477d

huggingface-projects/gemma-4-12b-it

Interact with a multimodal AI using text and multimedia inputs to generate text responses.

Hugging Face Space
76likes
367d

coreai-model-zoo

Community-driven repository for Apple Core AI models.

GitHub
143stars
937d

PRISM-VL

Improving vision-language models with measurement-domain observations.

GitHub
226stars
1077d

dealignai/DeepSeek-V4.1-Flash-UNCENSORED-FP8

Hugging Face Model

Hugging Face Model
269likes
1337d

webml-community/Qwen3.5-WebGPU

Chat with images using a local AI assistant

Hugging Face Space
51likes
217d

nvidia/Qwen3.8-Flash-Next-NVFP4

Hugging Face Model

Hugging Face Model
260likes
697d

Agnes-AI/Agnes-3.0-Flash

Hugging Face Model

Hugging Face Model
217likes
637d