Tag
#multimodal
From News
News · 2026-07-31
Seedance 2.5 ships native 30-second clips and dense multimodal references
ByteDance is rolling Seedance 2.5 into Dreamina and partner tools as a video model that can hold a continuous clip around 30 seconds in one pass and take a dense stack of image, video and audio references. For creators, that shifts AI video away from stitched short segments toward one-shot continuity with local fixes.
Read →News · 2026-04-21
ChatGPT Images 2.0 finally handles text in graphics, but production needs independent testing
ChatGPT Images 2.0 brings improved image generation focused on text accuracy in graphics, multilingual support, and advanced visual reasoning for production workflows.
Read →News · 2025-11-18
Gemini 3 Pro in practice: decent transcription, wrong timestamps, and no model knows the pelican
Simon Willison tested Gemini 3 Pro on a three-hour city council recording and a revised pelican benchmark. Result: a structured transcript for $1.42, but timestamps are off by tens of minutes. And none of the models tested understood that a California brown pelican is not actually brown.
Read →From the Library