xAI
Grok Imagine 1.0 Video
From $0.009 per second
FLUX 3 is Black Forest Labs' multimodal foundation model, trained jointly on images, video, and audio in a single architecture. FLUX 3 writes video and sound together — up to 20 seconds in one generation — and follows keyframes, reference clips, and multilingual dialogue.
Real-world workflows and production use cases you can build and ship with this model.

Text is the starting point, not the whole instruction set. Hand FLUX 3 a first frame to animate, reference images that pin down a character, a source clip whose elements should carry into a new scene, or two keyframes that define where a transition begins and ends. FLUX 3 can also continue an existing video and its audio, so a shot can be extended instead of regenerated from scratch.
Read the FLUX 3 docs
FLUX 3 handles multilingual dialogue and renders readable text inside the frame, so one creative direction can travel between markets. Signage, titles and animated typography are generated as part of the scene rather than composited on afterwards, and the spoken track is produced with the picture — which is where FLUX 3 is already strongest, according to Black Forest Labs' own early evaluations of facial expression and multilingual work.

The same FLUX 3 backbone synthesizes and edits images across styles, aspect ratios and resolutions, then extends those stills into motion. Shape, colour, material and branding stay recognisable between the packshot and the clip, because both come out of one model instead of a still generator bolted to a separate video generator.
FLUX 3 learns from images, video and audio inside a unified architecture built on Self-Flow, Black Forest Labs' approach to aligning multimodal generation and understanding. The modalities constrain each other: the sound has to match the impact, the motion has to obey the mass.
Every FLUX 3 video output carries native audio from the same generation — dialogue, ambience and impacts land on the frame they belong to, instead of being timed back onto silent footage in post.
FLUX 3 chains individual generations into multi-shot sequences that run for minutes, with visual references holding characters and locations steady from scene to scene. Twenty seconds is the ceiling for one generation, not for the finished piece.
Black Forest Labs' own early evaluations put these two closest together of everything it compared — viewers preferred FLUX 3 in 52% of matchups, effectively a coin flip. The split is availability and scope: Seedance 2.0 is callable on reAPI right now, while FLUX 3 is a broader multimodal foundation still in phased rollout.
Comparison reflects publicly documented behavior at the time of writing. The 52% preference figure is from Black Forest Labs' own preliminary evaluation of 10-second 720p text-to-video clips with audio, on a model it describes as still in development; sample size and methodology were not published, and it is not an independent benchmark.
Sign up at reAPI and generate a key. One key, one balance, and the same request shape across every model on the platform.
OpenThe FLUX 3 page tracks what Black Forest Labs has actually published — capabilities, input types and rollout stage — and gets the endpoint, parameters and rates the moment they are confirmed.
OpenVideo models with native audio are already callable on reAPI. Build the integration against one of those now, and switching to FLUX 3 later is a model id and a parameter map.
OpenCommon questions about this model.
Explore more models in the same category.
xAI
From $0.009 per second
PixVerse
From $0.018 per second
Topaz Labs
From $0.044 per second
ByteDance
From $0.029 per second
Try it in the playground or grab an API key to integrate now.