reapi-video-1
reAPI Video 1 — async video generation built on MiniMax H3. 5–15 second clips with matching dialogue and sound from text, a first frame or your own references. Full parameter reference and billing rules.
reAPI's video generation model, built on MiniMax H3. One request makes the
whole scene, picture and sound together: a 5–15 second clip with its own
dialogue, ambience and effects. Unlike most video models on reAPI, the mode
is an explicit field — mode picks text to video, image to video or
reference to video. See pricing on the
model page.
Quick example
curl https://reapi.ai/api/v1/videos/generations \
-H "Authorization: Bearer YOUR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "reapi-video-1",
"mode": "text_to_video",
"prompt": "A lighthouse keeper in a wool coat stands on a wet stone pier at dawn and says, \"The fog lifts at seven.\" Locked-off shot, waves slapping the stones, no music.",
"duration": 5,
"resolution": "768p",
"aspect_ratio": "16:9"
}'import requests
resp = requests.post(
"https://reapi.ai/api/v1/videos/generations",
headers={
"Authorization": "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
json={
"model": "reapi-video-1",
"mode": "text_to_video",
"prompt": "A lighthouse keeper in a wool coat stands on a wet stone pier at dawn and says, \"The fog lifts at seven.\" Locked-off shot, waves slapping the stones, no music.",
"duration": 5,
"resolution": "768p",
"aspect_ratio": "16:9",
},
timeout=30,
)
print(resp.json())const r = await fetch("https://reapi.ai/api/v1/videos/generations", {
method: "POST",
headers: {
Authorization: "Bearer YOUR_API_KEY",
"Content-Type": "application/json",
},
body: JSON.stringify({
model: "reapi-video-1",
mode: "text_to_video",
prompt:
'A lighthouse keeper in a wool coat stands on a wet stone pier at dawn and says, "The fog lifts at seven." Locked-off shot, waves slapping the stones, no music.',
duration: 5,
resolution: "768p",
aspect_ratio: "16:9",
}),
});
console.log(await r.json());package main
import (
"bytes"
"encoding/json"
"fmt"
"io"
"net/http"
)
func main() {
body, _ := json.Marshal(map[string]any{
"model": "reapi-video-1",
"mode": "text_to_video",
"prompt": `A lighthouse keeper in a wool coat stands on a wet stone pier at dawn and says, "The fog lifts at seven." Locked-off shot, waves slapping the stones, no music.`,
"duration": 5,
"resolution": "768p",
"aspect_ratio": "16:9",
})
req, _ := http.NewRequest("POST",
"https://reapi.ai/api/v1/videos/generations", bytes.NewReader(body))
req.Header.Set("Authorization", "Bearer YOUR_API_KEY")
req.Header.Set("Content-Type", "application/json")
resp, _ := http.DefaultClient.Do(req)
defer resp.Body.Close()
out, _ := io.ReadAll(resp.Body)
fmt.Println(string(out))
}Authentication
Every call needs a Bearer token. Generate keys at reapi.ai/settings/apikeys.
Authorization: Bearer YOUR_API_KEYEndpoint
POST /api/v1/videos/generations
GET /api/v1/tasks/{id}Submission is async. The POST returns immediately with a task id; the task
endpoint returns the same envelope until completion. Polling does not consume
credits.
Request body
| Field | Type | Default | Notes |
|---|---|---|---|
model | string | — | reapi-video-1. Required. |
mode | enum | reference_to_video | text_to_video, image_to_video or reference_to_video. The short forms t2v, i2v and ref2va are accepted. See Modes. |
prompt | string | — | 1 to 32,000 characters. Describes the picture and the sound. Required. |
duration | integer | 5 | Any whole number of seconds from 5 to 15. |
resolution | enum | 768p | 480p or 768p. |
aspect_ratio | enum | by mode | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus adaptive in reference_to_video only. text_to_video defaults to 16:9, reference_to_video to adaptive; image_to_video always follows the first frame. |
prompt_enhancement | enum | turbo | turbo (fast expansion), quality (more thorough) or disabled (the prompt is used word for word). |
seed | integer | random | 0–4294967295. The same prompt and seed return the same clip; omitted → a new random seed per request. |
image | object | — | image_to_video only, required there. The literal first frame: { "type": "url", "url": "https://…" }. |
reference_images | object[] | [] | reference_to_video only. Up to 9 { "type": "url", "url" } objects. |
reference_videos | object[] | [] | reference_to_video only. Up to 3. Only the first 5 seconds of each clip are read. |
reference_audio | object[] | [] | reference_to_video only. Up to 3. |
The schema is strict: unknown fields are rejected rather than ignored.
callback_url / callback_id are not accepted — poll the task, or set a
webhook on your API key.
Checked before anything is submitted or charged:
promptis 1–32,000 characters;durationis a whole number 5–15imageis present inimage_to_videoand absent in the other modesreference_images,reference_videosandreference_audiocarry values only inreference_to_video(an empty list ornullpasses in any mode)reference_to_videohas at least one reference image or video, and at most 9 images, 3 videos, 3 audio clips and 12 references in totaladaptiveis used only inreference_to_video- every media object is exactly
{ "type": "url", "url": "<public http(s) URL>" } - each reference video's length can be measured (it is billed — see Pricing); an unreadable file is rejected with nothing charged
Modes
| Mode | Requires | Takes |
|---|---|---|
text_to_video | prompt | the shared fields |
image_to_video | prompt and image | image as the first frame |
reference_to_video (default) | prompt and at least one reference image or video | reference_images, reference_videos, reference_audio |
image_to_video treats the image as the literal first frame, so the clip
opens exactly as that still looks. To place a product or a person into a scene
of your own, use reference_to_video and describe the scene around it.
Referencing media in the prompt
In reference_to_video, list order becomes the label you use in the prompt.
The first entry of reference_images is <Picture 1>, the second
<Picture 2>; the first entry of reference_videos is <Video 1>; the first
entry of reference_audio is <Audio 1>. Images, videos and audio are
numbered independently. Prompt enhancement can reword the rest of the prompt
but not these labels; set prompt_enhancement to disabled to keep your text
exactly as written.
{
"model": "reapi-video-1",
"mode": "reference_to_video",
"prompt": "A supervisor wearing the harness in <Picture 1> stands still and speaks to camera.",
"reference_images": [{ "type": "url", "url": "https://example.com/harness.jpg" }],
"duration": 10
}Media inputs
Every reference, and the image_to_video first frame, is an object with a
public HTTP(S) URL:
{ "type": "url", "url": "https://example.com/photo.jpg" }| Input | Maximum size |
|---|---|
| Image | 16 MB |
| Video or audio | 32 MB |
The URL must be reachable without redirects.
No data: URIs or asset IDs. reAPI rejects base64 inputs platform-wide —
upload the file to your own object storage (S3, R2, OSS, …) and pass its URL.
Output
The video is MP4 (H.264) at 24 fps with 32 kHz stereo AAC audio. With an
explicit aspect_ratio, the size is fixed:
aspect_ratio | 768p | 480p |
|---|---|---|
21:9 | 1536 × 672 | 960 × 416 |
16:9 | 1344 × 768 | 832 × 480 |
4:3 | 1024 × 768 | 640 × 480 |
1:1 | 768 × 768 | 480 × 480 |
3:4 | 768 × 1024 | 480 × 640 |
9:16 | 768 × 1344 | 480 × 832 |
adaptive follows the first reference image, or the first reference video
when there are no images, scaled to the short edge of the resolution with each
side rounded to a multiple of 32. image_to_video follows the first frame,
including its EXIF orientation — crop the image to change the shape.
Response envelope
Submit and poll share the same shape — only status and output fill in over
time.
{
"id": "task_01a10ee0e3827079aaf4165bf54fd954",
"model": "reapi-video-1",
"status": "completed",
"created_at": 1791250981,
"output": {
"video_urls": ["https://cdn.reapi.ai/media/tasks/.../0.mp4"]
},
"error": null
}Poll GET /api/v1/tasks/{id} (see the Tasks reference)
every 10–20 seconds until status is completed or failed. A network
timeout while polling does not mean generation failed: keep the task id and
resume checking it.
Troubleshooting
| Failure | What it means | What to do |
|---|---|---|
| Synchronous 400 on submit | A field violates the schema (a field from another mode, adaptive outside reference mode, more than 12 references, …) | The error names the field — fix and resubmit; nothing was charged |
| Synchronous 400 about a reference video's length | The file could not be measured for billing | Use a public MP4 or MOV URL that is reachable without redirects |
| Task fails after starting | The generation was rejected or failed | Fully refunded — see the errors catalog for the code |
402 on submit | Not enough credits for the reserve | Top up, or lower duration / resolution |
Pricing
reAPI Video 1 bills per second of video, at a rate set by the mode (text / image to video vs. reference to video) and the resolution:
credits = ceil(per_second_rate × billable_seconds × 1000) 1 credit = $0.001Live per-second rates are on the model page — that table is generated from the current price, so it is always the authoritative number.
billable_seconds = duration (text_to_video, image_to_video)
billable_seconds = duration + ceil(Σ min(each reference video's seconds, 5)) (reference_to_video)Each reference video adds its length up to 5 seconds; a longer clip still adds
5. Reference images and audio are free. reAPI measures every reference video
server-side when the request is accepted, so the amount reserved is the amount
charged. An 8.6-second reference with duration: 5 bills 10 seconds.
Failed tasks are always refunded in full.