What is the AI Hotel Lobby Video?
The Hotel Lobby trend is named after a rap song, not a place. The clip everyone copies is a live studio performance: two rappers in a plain, bright-orange booth with a single microphone hanging from the ceiling, trading lines of the verse back and forth, leaning toward each other, pointing, and stopping for a finger-to-lips "shh" before the next line. In 2026 people started recasting that clip with AI, putting themselves and a friend in the two performers' places while the original song plays.
This template does exactly that from two photos. The first photo becomes the performer on the left, the second the performer on the right. The model follows the original performance shot for shot: the same orange booth and hanging mic, the same framing and cut, the same gestures and timing, with both of you mouthing the verse. Each of you keeps your own face, hair and clothes.
The result is a 15-second vertical (9:16) video, ready for TikTok, Reels and Shorts: the performance runs across the middle of the frame over a blurred backdrop of itself, exactly like the viral posts. You also get the full landscape (16:9) copy. Both carry the original song, lined up with the performance, and no watermark. Generation takes about 15 minutes, because the model has to match a whole 15-second performance frame by frame.
How the recast works
Most photo-to-video templates start from one picture and invent the motion. This one works the other way round. The model is given a 15-second reference performance, the song, and your two photos, and is asked to redo that performance with you in it. That is why the timing, the camera cut and the gestures match the trend clip so closely, and why the faces are yours rather than the original performers'.
The model also generates its own vocals, which tend to drift a little off the beat. So once the video is ready, we replace its audio with the original 15-second song section. The picture is untouched; only the sound changes, which keeps the lip-sync and the music lined up.
Getting the best result
Everything depends on the two photos. Clear, front-facing pictures with even light, one person each, give the model the most to work with, and the faces stay steadier through the fast gestures. A photo where the face is small, turned away or half hidden is the most common reason a result looks less like you.
If a run comes out with a wrong hand or a face that drifts in the wide shot, generate again: each take is a little different. Keep in mind that a run takes about 15 minutes, so you can leave the page and find the video in Creations when it is done.






