Tencent Releases WeMM-Embedding for Multimodal Retrieval
Tencent’s WeChat Vision team has published WeMM-Embedding , a family comprising 2B, 4B and 9B embedding models. Each variant supports text, images, videos, visual documents and interleaved multimodal inputs; audio is not supported. For developers, the immediate consequence is a single repository containing model options, inference examples, serving instructions and evaluation code for retrieval…
We haven't written up this one. Dev.to has the full story — the link below goes straight to it.