The model supports multimodal inputs including text, images and video, and generates high-fidelity 1080p videos of 5 to 20 seconds with natively synchronized audioBeijing, China--(Newsfile Corp. -) - ...
Microsoft has introduced a new AI model that, it says, can process speech, vision, and text locally on-device using less compute capacity than previous models. Innovation in generative artificial ...