i
DATAIST
Back to feed

Multimodality

Text, sound, image and video in a single model.

8 articles

Reka AI’s Rho-1 puts robot control inside a multimodal model

Reka AI has introduced Rho-1, a model that handles text, images and video while also controlling robots. Rather than routing each task to a separate system, it represents different data types as tokens in one context window. It can generate continuous video in real time and respond to new instructions without restarting—an approach that puts a shared model, rather than a collection of specialist tools, at the center of the system.

OpenAI prices GPT-6 Sol and Luna for an agentic split

OpenAI has released GPT-6 Sol and GPT-6 Luna with permanent API prices starting at $0.10 per 1 million input tokens. Sol costs $2 per 1 million input tokens and $10 per 1 million output tokens, half the price of GPT-5.6 Sol. Luna targets simpler, high-volume work at $0.10/$0.50, while Astra remains the company’s top-end model for difficult, multimodal and scientific tasks. The release is less about one model replacing another than about making model routing part of the product.

Tencent’s Gander targets the gap between talking and doing

Tencent’s Hunyuan Speech team and researchers from several universities have introduced Gander, a multimodal AI model designed to keep talking while another system works in the background. It processes speech, images and text continuously, including during its own replies, so a user can interrupt it at any moment. That targets a basic weakness in today’s voice assistants: they usually take turns, while real conversations involve interruptions, quick reactions and updates while a longer task is still running.

Meta opens Muse Spark 1.1 to developers through a new model API

Meta has released Muse Spark 1.1, a multimodal model built for agentic work, and opened it to outside developers for the first time through a new Meta Model API now in public preview. The model manages a one-million-token context window on its own, runs as either the lead agent or a subagent inside parallel multi-agent setups, and operates a computer by switching between writing scripts and…

DeepSeek's V4.1-Flash targets the memory bill, not the leaderboard

DeepSeek has released V4.1-Flash, a multimodal model whose pitch is a memory bill rather than a benchmark. The KV cache — the buffer that holds already-processed context so the model does not recompute it at every step — now occupies roughly a quarter of the fast GPU memory that DeepSeek-V4-Flash needed, and the portion permanently offloaded to SSD or host memory falls to about an eighth. The…

An agent recovers physics from video by writing simulator code

Large multimodal models can already describe what happens in a video: a ball rolls, a cup falls, a car turns. Physics is where the old problem persists. They see the phenomenon, not the mechanism. They can give a tidy account of a clip but don't always understand why the object moved the way it did, how fast it was going, what would change if you…

Predicting the answer's latent image beats text-only chain of thought

Multimodal LLMs have learned to recognize objects, but how do you give them visual imagination? A look at the Cognitive Supersensing idea. Over the past few years, multimodal LLMs (MLLMs) have learned to recognize objects, read captions, answer questions about an image and even give a decent account of what is happening in a frame. But they have…

Self-written explanations are what let models read dark humor in memes

Not all jokes work the same way. Clean humor runs on wordplay and harmless incongruity; dark humor runs on painful subjects, cultural references and fine contrasts between the image and the caption. In memes this is especially visible: the picture says one thing, the text says another, and the meaning appears where they meet. Until recently there was no good multimodal dataset for dark humor…