Hotel Lobby AI Prompt Guide

The five-part formula behind the Hotel Lobby AI trend prompt — booth scene, subject anchoring, locked camera, alternating verses, room-tone audio — plus ready-to-use variants.
Oct 2, 2026

The Hotel Lobby AI trend is a scene, and the scene has a formula. Once you see the five parts — anchored subjects, the orange booth, the locked camera, alternating verses, room-tone audio — you can rewrite the prompt for any duo: friends, pets, characters, historical figures you have the rights to use.

The five-part formula

  1. Subjects, anchored to sides. "Subject A on the LEFT, Subject B on the RIGHT" — plus one sentence forcing each to keep the exact face, hairstyle, clothing, and skin tone from the input photos. This is the single strongest fix against side-swapping and face blending.
  2. The booth. A seamless warm orange studio backdrop and one black condenser microphone suspended at the exact center between the two subjects. That mic is the visual anchor of the whole trend.
  3. A locked camera. Tripod, eye-level, medium shot, no pan, no zoom, no cuts. Movement kills the live-session feel instantly.
  4. Alternating verses. Subject A performs a short verse with small hand gestures while Subject B nods and reacts; then they swap. Explicitly forbid both performing at the same time — models love to sync them.
  5. Room tone, not music. "Ambient studio room tone, mic proximity breath, no background music." The silence is what makes the mic feel real; the song gets added later in your editor.

A ready-to-run prompt

Create a vertical 9:16 image-to-video of a COLORS booth style performance. Subject A from the first photo stands on the left, Subject B from the second photo on the right; both keep the exact face, hairstyle, clothing, and skin tone from their input photos — no face blending, no identity swap. Setting: seamless warm burnt-orange studio backdrop, one black condenser microphone suspended at the exact center, soft even studio light. Camera: locked tripod, eye-level medium shot, no pan, no zoom, no cuts. Motion: Subject A performs a short verse with subtle hand gestures while Subject B nods and reacts; then Subject B performs and Subject A reacts — never both at once. Audio: ambient studio room tone, mic proximity breath, no background music. Avoid extra people, background movement, or camera shake.

Run it in this site's generator or in any tool that takes reference photos plus a prompt.

Duo variants that work

Swap the photos and adjust one line — everything else stays:

  • Two friends — the default; keep the verse split even.
  • A person and a pet — make the pet the reactor: head tilts and blinks while the human performs.
  • Two characters (anime, game) — keep the identity-preservation line; it matters even more for stylized faces.
  • Solo mode — one subject on the right, the empty side implied; ask for a full verse with a self-assured outro.

Fixes for the four common failures

  1. Subjects swap sides — strengthen the LEFT/RIGHT anchoring and re-render; do not re-upload.
  2. Faces blend — add "the two subjects never merge or morph into each other" and use sharper source photos.
  3. Both move at once — restate "only one subject performs at a time; the other only reacts."
  4. Camera drifts — repeat "locked tripod, no zoom, no pan" at the end of the prompt; models weight late instructions.

Two rules before you post

Don't publish with the original song's audio — add your own licensed track or keep the room-tone cut. And don't render real celebrities or artists without permission; the trend took heat for exactly that. Original duos, AI-generated label, your own audio — that's the clean version of the trend.