Tag: multimodal
Alibaba's Wan3.0 Video Model Takes PowerPoint and PDF Files as Reference Material - Up to 30 Seconds at 30fps
Google DeepMind Brings Sign-Language-to-Text SL2T to Gboard and Live Transcribe on Pixel 11
Alibaba Qwen 3.8: 2.4T-Parameter Model Claims 'Second Only to Fable 5', Open Weights Promised
Google Gemma: The Evolution and Strategy of Lightweight, High-Performance Open Models