· reAPI Team
Hotel Lobby AI pets: upright vs four-paw cat test
We tested two cat poses in Hotel Lobby AI with the same Seedance 2.5 motion reference. See the inputs, silent drafts, visible defects and experiment limits.

In this Hotel Lobby AI pets test, the four-paw reference produced an upright upper body without an upright preparatory image. Both conditions also produced hand-like paw gestures; this does not establish a universal requirement or a guarantee of natural anatomy. We tested that preparation question with two orange-tabby inputs: one standing on four paws and one standing upright. Each went into a single Seedance 2.5 video request with the same woman on the right, the same motion reference and the same prompt. The experiment is dated October 10, 2026.
This is a small, inspectable comparison rather than a recipe for perfect animal choreography. The cats were generated independently, so their proportions and markings differ. We also retained the human-oriented template prompt. Those conditions matter when interpreting the outputs. If you are making your first clip, the complete Hotel Lobby AI tutorial covers the ordinary two-photo workflow; this article concentrates on animal pose preparation.
TL;DR
- We generated one four-paw cat and one upright cat with GPT Image 2, then submitted one video per input with Seedance 2.5. No repeated attempts were selected to create a winning example.[1]
- Both video requests used 15 seconds, 480p, 16:9,
draft: trueandgenerate_audio: false. The left image URL was the only field changed between the video requests.[1] - Both sampled outputs showed an upright cat on the left, including the four-paw input. Both also produced hand-like paw gestures; this is not evidence of anatomically exact motion transfer.[1]
- The two independently generated images are not a strict same-animal posture control. There was no fixed seed or repeated sample per condition.[1]
- The pet outputs below are silent raw drafts. The Hotel Lobby AI tool adds the original soundtrack after generation; the adult example demonstrates that separate step.[2]
The cover is concept artwork. It was created separately and is not a frame from either experimental video. The reference photos and videos below are the actual returned assets.
What changed in the Hotel Lobby AI pets experiment
The practical question is whether a naturally posed animal can be mapped onto a performance that was designed around two upright people. Rather than change the motion and animal at the same time, we held the video request constant and changed the left reference photo. This lets you see two concrete outcomes without treating them as a general model benchmark.
| Setting | Both requests |
|---|---|
| Video model | Seedance 2.5, doubao-seedance-2.5-face |
| Left performer | Orange tabby cat; posture image changes |
| Right performer | The same fictional adult woman reference |
| Motion reference | The same supplied line-art duet video |
| Requested length | 15 seconds |
| Resolution and frame shape | 480p, 16:9 |
| Mode | omni_reference_task_type: "reference" |
| Draft | draft: true |
| Generated audio | generate_audio: false |
| Content filter | content_filter: true |
| Attempts | One submission per posture |
These are submitted settings, recorded with task responses. The presence of a field in a successful request does not independently prove that every setting changed the model's behavior. Our claims about movement come from viewing the returned clips, not from the HTTP status. We did not submit a 1080p final, compare other models or generate a new reference performance.[1]
The prompt deliberately remained identical to the adult test. It assigns Image 1 to the left performer and Image 2 to the right; it asks Video 1 to guide choreography, gestures, mouth movements, framing and timing. It also asks the model to preserve “face, hair, clothing and identity” and produce “consistent hands.” Those words were written for human performers. They are an important limitation of this experiment, not a validated cat prompt.
The actual cat references
Both GPT Image 2 requests described an orange tabby with orange stripes, a white chin and chest, green eyes and a plain black sleeveless vest. They asked for a complete body in a portrait 3:4 image against a warm beige studio background, without additional animals, people, microphones or text. We did not generate a new woman, motion clip or soundtrack.
Four paws on the ground

The returned cat has a horizontal body, four visible paws on the ground and its face turned toward the camera. The black vest follows the feline torso. Its tail is visible. This is the naturally posed input used for the first video, not an illustration of a supposed video result.
An upright cat before video generation

The second image already has a vertical, elongated body, two hind paws on the ground and two forelimbs hanging at its sides. Those forelimbs look arm-like before animation begins. The photo therefore changes more than the location of the paws: it gives the video model a different body plan and vest shape to preserve.
Shared descriptive wording did not produce identical animals. Facial details, paw markings, scale and garment fit differ. We did not edit the first image into the second or photograph the same real pet in two poses. A claim that posture alone caused a difference would go beyond this setup.[1]
Watch the two silent drafts
Four-paw input: an upright upper body appeared
Open the original four-paw-input video.
The opening sample already shows an upright cat on the left, with the woman on the right. The original horizontal body was not retained. At the middle and late sampled times, the white front paws form raised gestures with elongated, digit-like shapes and visible pink pads. The orange fur and black vest remain recognizable. The crop excludes the hind legs, so we cannot assess whether the complete body has a correct number of limbs.
Upright input: pointing gestures with hand-like paws
Open the original upright-input video.
The upright-input samples also keep the orange cat on the left and the woman on the right. The cat raises its forelimbs into pointing and thumb-like shapes. Several early and middle sequence samples retain a pointing pose. These are recognizable performance gestures, but the paws are not consistently natural feline paws. We do not rank this output above the other one.
Both files were probed as H.264 video, 854×480, approximately 15.04 seconds, with no audio stream. The frame review covered 0, 5, 10 and 14.5 seconds plus a one-frame-per-second sequence. This was a frame-sequence review, not continuous playback or a quantitative timing assessment.[1]
We inspected opening, middle and late frames and a sequence across the returned clips. A still image is useful for checking a limb silhouette, but a full playback is needed to see whether a pose persists or changes during a gesture. We did not assign an accuracy percentage, a lip-sync score or a universal success rate. A visually pleasing excerpt would not establish that the entire performance copied the reference exactly.
When comparing the clips yourself, concentrate on the cat's forelimbs as they move toward the microphone, its relationship to the woman on the right, and the transition between gestures. Check the end as carefully as the opening. Identity preservation, body shape and reference timing are different questions; a clip can satisfy one and still disappoint on another.
What one attempt per pose can tell you
A single successful task establishes that a result was returned for that request. It does not establish that another cat photo will behave the same way. Two clips also cannot tell us the frequency of malformed paws, whether a particular prompt is best, or whether another model would improve the result.
For a creator choosing an input, our recommendation is to inspect the intended silhouette before paying for video generation. If you want a human-like upright duet, an upright concept image explicitly describes that intention. If you want your pet to remain recognizably quadrupedal, inspect whether the returned video honors that requirement rather than assuming the human motion template will preserve it. These are preparation choices, not guarantees of the animation result.
The input should also match your desired identity. A generated orange cat is suitable for testing the workflow, but it is not evidence that a real pet's exact face or markings will remain consistent. To test your own pet, start with an image you have permission to use and compare the output to that image. Keep the original request and task ID so that a failed result is distinguishable from a new attempt.
We intentionally did not rewrite the prompt between conditions. Changing pose, anatomy instructions, motion reference and model simultaneously would make a nicer example hard to interpret. A separate future test could use animal-specific language, but this article does not present that unperformed experiment as an improvement.
Prepare and review an animal duet
- Decide whether your intended character is a natural animal or an upright fictional performer. Record that choice before generation so the result is judged against your goal.
- Inspect the input at full size. Check the complete body, tail, paw separation, face visibility and any clothing. These are editorial review checks, not guaranteed model requirements.
- Assign the animal and partner to explicit sides. Our cat was Image 1 on the left; the woman stayed Image 2 on the right.
- Keep the motion reference and other settings stable for a first diagnostic attempt. Note any human-oriented anatomy words in the prompt.
- Review the complete output. Pause during a large gesture, a turn and the closing pose; watch for limbs that merge, change shape or appear to multiply.
- Check the finished soundtrack and visual result before sharing. A correct soundtrack cannot repair a malformed paw or a character that has moved to the wrong side.
The Seedance 2.5 model page is the place to check current options and prices. The tool's generation estimate should be checked before starting a new attempt.
For the soundtrack, distinguish the experiment from the website flow. We requested silent pet videos to isolate the visual inspection and did not merge audio into them. The recorded adult example includes a separate composition step that adds the supplied original track. The website performs that step automatically after generation; turning on model-generated audio is not the same operation and is not evidence that a song has been copied precisely.[2]
FAQ
Must I generate an upright animal first?
In this one four-paw-input result, the sampled upper body became upright without an upright preparatory image. That does not establish a universal rule. The test shows what two particular cat references produced with one shared request and one attempt each. Choose an upright input if that is your intended character design, then judge the full output.
Were these two photos of the same cat?
No. They are two independently generated fictional orange tabbies with shared descriptive wording. Posture, proportions, markings and vest fit differ, so this is not a strict same-identity control.
Did you test Seedance Mini?
No. Both experimental video requests used Seedance 2.5 with the exact API ID doubao-seedance-2.5-face. These results should not be attributed to Mini or any other model.
Was the prompt optimized for pets?
No. We retained the adult template wording, including “hair” and “hands,” in both requests. The experiment does not compare animal-specific prompting or establish the best wording.
Why do the pet examples have no soundtrack?
They are raw drafts requested with generate_audio: false, and we did not add audio afterward. The website's final soundtrack composition is a separate step demonstrated by the existing adult example.
Do these outputs prove perfect motion transfer?
No. We report visible behavior under the stated conditions, without a quantitative timing analysis or repeated samples. Completing a task is not proof that every gesture matches the reference.
Can I use my own pet photo?
The tool accepts reference images, but this experiment used fictional generated cats. It does not establish how your particular photo will perform. Inspect your input, check the estimate and evaluate the returned clip against the identity and pose you want.
Choose the silhouette you actually want
Use these two outputs as examples to inspect, rather than as a rule about all animals. Save a clear input, state which side it belongs on and review the complete movement before sharing or making another attempt. For Hotel Lobby AI pets, the useful decision is whether the reference character and the final body shape match your intended duet. Start with the Hotel Lobby AI workspace, or read the complete two-photo tutorial before generating.
For the next step, see the prompt and reference guide, cost and one-time-use guide, or template and generation troubleshooting.
References
- reAPI Team. Pet pose experiment, October 10, 2026. Actual input images and raw video outputs are linked in this article. Two GPT Image 2 tasks and two Seedance 2.5 requests, one video per posture. Public experiment record.
- reAPI. Hotel Lobby AI tool and recorded adult example. Checked October 10, 2026. reapi.ai/tools/hotel-lobby-ai.