# Bar Couple Photo Gen

> Generate 9:16 photorealistic candid man-woman interaction images, selected Gary hotel lounge prompts, or approved single-female hotel CCTV prompts. Use when the user asks to create images with a fixed male character, replace only the female character, generate dating/couple/bar/social interaction photos, generate a saved single-female CCTV scene, or produce GPT Image 2 high-fidelity portrait consistency prompts.

- Skill: `kinkwanc22/bar-couple-photo-gen` (Agent Skill, multi-file: 10 files)
- Install (CLI): `npx skillmds@latest add kinkwanc22/bar-couple-photo-gen`
- Raw SKILL.md: https://api.skillmd.com/api/skills/kinkwanc22/bar-couple-photo-gen/raw
- Safety review: pending
- Works with: Claude Code, Claude.ai, OpenAI Codex
- Category: Coding & Dev Tools
- Author: kinkwanc22 (https://skillmd.com/u/kinkwanc22)
- Updated: 2026-09-21
- Page: https://skillmd.com/skills/kinkwanc22/bar-couple-photo-gen

---


# Bar Couple Photo Generator

Use this skill to generate vertical 9:16 realistic lifestyle images with a fixed male lead and a variable female lead. The default style is candid smartphone nightlife photography: direct phone flash, casual friend snapshot, realistic skin texture, non-commercial, non-cinematic, non-AI-polished.

For the current user's recurring Gary/female-lead series, preserve the tested prompt structure: change only the `{scene}` and `{ambiguous_interaction}` blocks when requested, keep the fixed texture/identity/negative wording stable, and avoid full-body compositions. The strongest tested look is close, imperfect, handheld indoor/nightlife framing where the woman is near the man and the man is not posing for the camera.

## Fixed Male Lead

Use `assets/fixed-male-lead.png` as the fixed male identity reference.

For this user's Gary/Lovart workflow, use these synced Windows reference paths by default:

- Fixed male lead: `D:\工作用（同步）\图\人物设定\2896d6707751ecfc30a82a7e9bc680b61f8ccd0efbf28a005f3674f5036a6c0c (1).png`
- Female reference library: `D:\工作用（同步）\图\人物设定\精品`

Preserve the male lead's real facial identity, facial proportions, short black hairstyle, mature understated temperament, age impression, body type, and natural skin texture. Do not beautify him into a model, celebrity, influencer, or fashion editorial face.

## Inputs

Require only one of these female-lead inputs:

- A female reference image.
- A text description of the female character's appearance, age impression, hair, body type, outfit, temperament, and desired vibe.

If the user also gives a specific interaction scene, use it. Otherwise choose one plausible casual social interaction from the interaction bank.

## Prompt Choice

Before generating Gary-series or related single-female media, ask the user which approved prompt to use unless they already specify it clearly:

1. `真实抓拍照片提示词` - the existing candid 9:16/16:9 phone-photo prompt used for image generation.
2. `高级酒店酒廊沙发抓拍照片提示词` - the fixed-Gary, random-female hotel lounge sofa candid photo prompt below.
3. `高级餐厅后方抓拍` - the fixed two-lead restaurant candid preset.
4. `情侣打闹视频首帧` - the fixed first-person boyfriend POV pillow-play opening frame.
5. `暧昧互动` - a stable-random first-person boyfriend POV preset. It randomly changes indoor, semi-outdoor, outdoor shaded scenes, time, light, position, and restrained intimate action while preserving the approved candid phone-shot feeling.
7. `单女主酒店走廊敲门CCTV` - a one-reference security-camera frame of only the female lead approaching the left-side room door. Gary must not be uploaded or appear.
8. `双人酒店走廊开门CCTV` - a two-reference security-camera frame with Gary waiting beside the female lead as she prepares to open the left-side room door.
10. `单女主西餐厅超低机位聊天抓拍` - uses only the current random or user-provided female as 图1. Gary is not uploaded or shown.

If the user chooses the hotel lounge sofa prompt, use the following prompt text exactly as the core prompt. Keep it in Chinese; do not translate it into English. Use the fixed Gary male reference as 图1 and the current female reference as 图2. The male lead remains fixed; the female lead may change by replacing 图2.

```text
高级酒店酒廊里的真实朋友圈抓拍照片，参考图1和图2的人物身份、面貌、服装、人物关系保持一致，不改变角色长相。男主和女生并肩坐在浅色沙发上，距离很近，像正在和朋友聊天时被随手拍下。男主坐在左侧，身体微微后靠，一只手自然搭在女生身后或沙发靠背上，另一只手朝画面外轻轻比划，像正在说话；男主看向画面左侧，不看镜头，表情自然沉稳。
负面提示词：
网红脸，模特感，明星脸，过度精修，磨皮，美颜滤镜，塑料皮肤，脸部过于完美，鼻子过高，五官过度立体，夸张妆容，时尚大片，杂志封面，棚拍，商业摄影，专业布光，电影感过强，光线过于干净，背景虚假，摆拍感，刻意看镜头，夸张姿势，不自然表情，人物重复，脸部结构错误，多余手指，手部畸形，肢体畸形，腿部畸形，比例错误，脚放在桌上，腿搭在桌上，脚踩茶几，脚接触酒桌，脚悬在桌面上，CG感，3D感，动漫感，AI感过强，过度锐化，低质量脸部

女生坐在男主右侧，身体贴近男主，姿态放松，双腿自然向镜头前方伸展或在脚踝处轻轻交叠，高跟鞋和腿部形成明显前景透视；双脚始终自然落在地面上，不得接触、踩踏或搭在酒桌、茶几及任何家具上。镜头正对两人但人物不刻意看镜头，像朋友聚会中临时拍下的一张原图。女生表情松弛自然，不是模特感，不要网红脸，鼻子不要太高，五官保持普通真实，真实皮肤纹理、轻微瑕疵和面部细节保留。
竖图 9:16，近距离低机位拍摄，镜头高度接近茶几，从人物正前方略偏右的位置拍摄，人物主要位于画面下三分之二。前景有黑色酒桌、威士忌酒杯、白色杯子、玻璃反光；背景有灰色墙面、深色柜门、玻璃门、落地灯。业余手机摄影风格，近距离广角，轻微歪斜构图，轻微运动模糊，局部失焦，夜拍噪点，机顶直闪，人物皮肤有闪光灯高光，背景偏暗，高光轻微溢出。非摆拍，非商业摄影，非精修，强烈生活流纪实感，像朋友聚会时无意拍到的一张真实照片。
```

## Default Output

Generate one vertical 9:16 photorealistic image unless the user asks for another ratio or multiple variants. For local Codex generation, use GPT Image 2 with high-fidelity portrait consistency. GPT Image 2.5 is disabled for all new generation and must never be selected as a fallback. Treat reference images as identity references, not edit targets. Do not generate full-body compositions unless the user explicitly asks. Prefer waist-up, half-body, close three-quarter, or tight social-photo framing with natural cropping. Slightly cut-off hands, shoulders, drinks, or table edges are acceptable when they make the photo feel like a real phone snapshot.

## GPT Image 2 Routing Lock

The user's standing rule is that all future Gary/couple image generation uses the legacy GPT Image 2 model in both the local Codex and Lovart lanes. Each lane keeps its own callable tool identifier, quality, size, thread, upload, and download behavior.

- Select or explicitly request GPT Image 2 / `gpt-image-2` in the built-in image-generation entry and in the generation prompt.
- Never select, request, or fall back to GPT Image 2.5 Sunburst or Flare for new output.
- If GPT Image 2 is unavailable or cannot be selected with confidence, stop and report the limitation instead of silently using or relabeling another model.
- Label provenance as `GPT Image 2`. If the tool does not return a backend model ID, disclose that separately as `工具未回传后端模型ID`; do not turn that absence into permission to use another route.
- Historical GPT Image 2.5 files remain valid read-only comparison/archive material; never reuse them as new-generation sources or relabel them as GPT Image 2.
- For Lovart, use the callable tool id `generate_image_gpt_image_2_medium` as the GPT Image 2 Medium route.

## Local GPT Image 2 Windows GPU 2K Finishing

This section applies only to locally generated GPT Image 2 / `image_gen` results. It does not apply to Lovart, and it must not change any Lovart quality, size, thread, upload, download, or saving behavior.

After a local GPT Image 2 result is generated, automatically finish it on the connected Windows host `win-codex` with the installed RTX GPU Real-ESRGAN pipeline:

```text
scripts/upscale_image2_windows_gpu.py <local-image-path> --batch-dir <dated-batch-folder>
```

Rules:

- Treat the GPT Image 2 file as the source image and preserve it under `<batch-folder>/.records/raw/`.
- Use Real-ESRGAN on the Windows GPU, then resize the enhanced result to the exact target dimensions.
- To avoid slow cross-device transfer of a large lossless image, return a temporary JPEG at quality 96 and convert it to the final PNG on the Mac. The temporary transfer file is not kept in the visible batch folder.
- Vertical `9:16` final: `W 2016 / H 3584`.
- Horizontal `16:9` final: `W 2048 / H 1152`.
- Save only the final requested images in the visible dated batch-folder root. Keep raw files and the processing manifest under `.records`.
- Record every result in `.records/image2_windows_gpu_manifest.jsonl`, including source size, target size, output path, Windows job, elapsed time, and failure reason.
- Label new provenance as `GPTImage2_WindowsGPU超分2K`. Do not call it native GPT Image 2 2K.
- Verify the returned file's actual pixel dimensions before reporting success.
- If Windows, SSH, the RTX GPU, or Real-ESRGAN is unavailable, report the local image as generated but the 2K finishing stage as failed. Do not substitute Mac CPU upscaling or ordinary resizing without the user's explicit instruction.
- Never send a Lovart output through this helper. Lovart continues to use the independent defaults and batch rules below, unchanged.

## Fixed Lovart Defaults

For Gary/couple image generation through Lovart, use these defaults unless the user explicitly overrides them:

- Model family: `GPT Image 2`; GPT Image 2.5 is disabled and is not an allowed fallback.
- Quality: `medium`, routed through Lovart's callable GPT Image 2 Medium tool id `generate_image_gpt_image_2_medium`.
- Number of images per generation: `1`.
- Vertical `9:16`: `W 1008 / H 1792`.
- Horizontal `16:9`: `W 1792 / H 1008`.
- Do not silently use `auto`, `high`, `2k`, `4k`, `2 img`, or `4 img`.

For this user's standing Gary batch preference, requests phrased as `一批`, `批量`, or an equivalent multi-image/multi-preset batch use the quality-stable batch profile unless the user explicitly overrides it:

- Keep quality at `medium` and generate exactly `1 img` per Lovart call.
- Use the Lovart UI `2K` sizes: vertical `2016x3584`, horizontal `2048x1152`.
- Run every requested image as its own `chat` call without a reused Lovart thread id, so each image starts from a clean conversation and does not inherit accumulated prompt context.
- This standing batch preference is the only automatic `2K` exception. Ordinary single-image tests keep the non-2K defaults above.

The Mac batch runner is `scripts/run_gary_batch_lovart_mac.py`. It owns these technical defaults and supports an offline `--dry-run` / `--print-prompt` mode that must not read credentials, upload references, or call Lovart.

On macOS, the runner resolves Lovart credentials in this order: current process environment, the current `launchctl` login environment, macOS Keychain services `codex-lovart-access-key` and `codex-lovart-secret-key`, then the user-only fallback file `~/.lovart/credentials.json`. Keep the fallback file at permission `0600`, never commit it, and keep credential values out of prompts, logs, manifests, skill files, and shell output.

For local saving, group every same-character, same-request batch into one dated folder under the user's long-video image directory. Name it with a readable date, female source stem, and batch purpose, for example `2026-07-19_IMG_9757_三套横版2K`. Pass that exact folder with `--batch-dir` to every preset invocation in the batch. The visible batch-folder root must contain only the requested final images. Put manifests, current-state files, run summaries, and Lovart historical downloads under the hidden `.records` subdirectory; do not create one visible output folder per preset and do not leave historical thread images beside the current batch outputs.

When the user explicitly asks for the Lovart UI `2K` size presets, or when the standing quality-stable batch profile applies, pass `--resolution-profile 2k`:

- `9:16（2K）`: `W 2016 / H 3584`.
- `16:9（2K）`: `W 2048 / H 1152`.
- Keep quality at `medium` and image count at `1` unless the user explicitly overrides them.
- This is an explicit override only; it does not change the ordinary default sizes above.

Use `--female-path` when a previously selected female reference must be excluded and a newly randomized top-level library file has already been chosen. The path must belong directly to `--female-dir`.

After the mandatory Lovart `config --json` and `threads --json` checks, pass related existing aspect threads with `--vertical-thread-id` and `--horizontal-thread-id` so ordinary follow-up tests reuse their prior conversations. For the standing quality-stable batch profile, deliberately omit both thread-id arguments; this user-approved batch exception makes every requested image a clean independent generation.
When a reused thread returns historical artifacts together with the new result, the runner keeps only the latest artifacts matching the requested dimensions in the manifest and success count.

## Coffee Candid Universal Preset

Preset id: `coffee_candid_universal`

Chinese name: `高级餐厅后方抓拍`

This preset works for both `9:16` and `16:9`. Keep aspect wording and dimensions in the generated parameter line, outside the fixed core prompt. The core prompt is fixed: do not randomize the scene, camera angle, position, interaction, background evidence, or lighting. For `16:9`, the first line must say that the frame keeps more of the high-end restaurant environment while the people remain a close candid crop and are not full-body.

Fixed core prompt:

```text
使用图1和图2作为人物参考，保持两位人物的真实面貌、五官比例、发型、年龄感、体型和气质一致，不要美化成模特或网红脸。高级餐厅的真实抓拍照片，男主和女主在沙发上聊天，从男主后方拍摄，能清楚看到女主的脸，距离很近，表情放松，不看镜头，真实到像朋友圈原图的手机抓拍，像正在聊天或临时合影时被朋友随手拍下，背景里有其他人，主角不看镜头，动作不统一，表情自然松弛，人物都是普通人，人物在图片下三分之二位置，女生鼻子不要太高，不是模特，不精修，不夸张打扮，保持与参考图角色一致，不改变面貌，真实皮肤纹理和轻微瑕疵保留。业余手机摄影风格，构图歪斜，轻微运动模糊，局部失焦，真实感灯光，高光，直闪 ，机顶闪光人像，非摆拍，非商业感，非宣传照，强烈生活流纪实感，像朋友聚会时无意拍到的一张照片，画面合规。 负面提示词： 网红脸，磨皮，美颜，时尚大片，棚拍，专业打光，AI感过强，塑料皮肤，脸部过于完美，多余手指，肢体畸形 负面提示词： 摆拍，棚拍，商业摄影，时尚大片，杂志封面，网红风，网红脸，模特感，过度美颜，磨皮，滤镜感，塑料皮肤，过度锐化，人物过于完美，夸张姿势，不自然表情，刻意看镜头，电影感过强，CG感，3D感，动漫感，AI感过强，光线过于干净，背景虚假，人物重复，多余手指，手部畸形，肢体畸形，脸部结构错误
```

## Couple Pillow Play First-Frame Preset

Preset id: `couple_pillow_play_first_frame`

Chinese name: `情侣打闹视频首帧`

Use this preset to generate the opening frame for a later image-to-video sequence. It supports both `9:16` and `16:9`. Treat 图1 as the boyfriend identity and 图2 as the adult female lead. All generated prompt text must remain Chinese.

For both the fixed and randomized couple-play first-frame workflows, the boyfriend entry wording is fixed and must be used verbatim:

```text
只允许画面边缘自然出现一只刚伸向抱枕的手或少量衣袖，不要出现男主的脸、头部、身体或第三人。
```

Random interaction must never be organized around a prop or object. Do not use interactions such as grabbing a phone, remote control, paper bag, sunglasses, coffee cup, room card, blanket, toy, clothing, or any other item. Randomize only human behavior and composition:

- Body distance: slightly closer, maintaining a natural close distance, or the woman leaning back slightly.
- Eye line: looking toward the boyfriend, briefly looking down and then back up, or turning her head while suppressing a smile.
- Expression: relaxed smile, restrained laughter, playful but ordinary expression.
- Small movement: slight sideways dodge, raising a hand to stop him, leaning closer or back, or turning away gently.
- Position relationship: woman centered or slightly off-center, seated at a slight angle, with usable continuation space preserved.

The pillow is the only fixed safe scene prop. It may remain in the woman's arms or against her torso, but the randomized interaction must not become a “grab an item” interaction. The boyfriend's hand is only just reaching toward the pillow and the action remains incomplete.

The fixed prompt is implemented in `scripts/run_gary_batch_lovart_mac.py` as `COUPLE_PILLOW_PLAY_FIRST_FRAME_CORE`. Keep these continuity anchors stable:

- The action is not yet complete: the boyfriend's hand has not touched the pillow.
- The woman is seated sideways in the sofa corner with the pillow as a fixed scene anchor; her randomized reaction uses only body distance, eye line, expression, slight dodging, leaning, turning, or a raised hand.
- The pillow may cover the torso but must not cover her face.
- Keep usable action space for a later turn, lean, or gentle dodge.
- Preserve a close half-body to three-quarter crop; do not turn it into a full-body wide shot.
- Keep the interaction playful, non-sexual, non-suggestive, and non-violent.
- For `16:9`, retain more hotel-suite environment and right-side action space without shrinking the female lead.

Call it with:

```text
--prompt-preset couple_pillow_play_first_frame
```

## Single-Female Hotel Door CCTV Preset

Preset id: `single_female_hotel_door_cctv`

Chinese name: `单女主酒店走廊敲门CCTV`

This preset uses only the current female reference as 图1. Do not upload, attach, mention, or depict Gary/the male lead. Do not reuse a two-lead Lovart thread for this preset. Preserve the following user-authored core prompt verbatim; only prepend the requested aspect ratio, dimensions, quality, and image-count parameter line.

```text
使用图1作为人物参考，保持人物的真实面貌、五官比例、发型、年龄感、体型和气质一致，不要美化成模特或网红脸。一张极其真实的酒店安防监控摄像头截图。固定在天花板墙角的高机位监控视角，略微向下俯拍，广角镜头。五星级酒店的走廊，暖灰色墙面，深色大理石门框，灰色地毯地面，暖黄色顶灯，走廊具有很强的纵深感，人物与场景比例正确。 女主正在走廊准备敲左侧的房间门，从前方拍摄，人物没有摆拍，没有看镜头，处于自然走路状态。在上三分之二向下走。 强烈的真实CCTV监控录像质感，普通安防摄像头成像，而不是电影摄影。轻微鱼眼广角畸变，轻微监控锐化，低码率视频压缩痕迹，细微噪点，轻微运动模糊，普通自动曝光，人物皮肤和衣服保留真实监控画面的细节损失，构图略显随意，像真实酒店监控系统随机截取的一帧。
```

Call it with:

```text
--prompt-preset single_female_hotel_door_cctv
```

## Couple Hotel Door CCTV Preset

Preset id: `couple_hotel_door_cctv`

Chinese name: `双人酒店走廊开门CCTV`

This preset uses the fixed Gary reference as 图1 and the current female reference as 图2. Both leads must appear. Preserve the following user-authored core prompt verbatim; only prepend the requested aspect ratio, dimensions, quality, and image-count parameter line.

```text
使用图1和图2作为人物参考，保持两位人物的真实面貌、五官比例、发型、年龄感、体型和气质一致，不要美化成模特或网红脸。一张极其真实的酒店安防监控摄像头截图。固定在天花板墙角的高机位监控视角，略微向下俯拍，广角镜头。现代高档酒店的狭长走廊，暖灰色墙面，深色大理石门框，浅灰色光滑反光的大理石地面，暖白色顶灯，走廊具有很强的纵深感。 人物与场景比例正确。 男女主并排，女主正在走廊准备开左侧的房间门，男主正在旁边等待，从前方拍摄，人物没有摆拍，没有看镜头，处于自然走路状态。在上三分之二向下走。 强烈的真实CCTV监控录像质感，普通安防摄像头成像，而不是电影摄影。轻微鱼眼广角畸变，轻微监控锐化，低码率视频压缩痕迹，细微噪点，轻微运动模糊，普通自动曝光，人物皮肤和衣服保留真实监控画面的细节损失，构图略显随意，像真实酒店监控系统随机截取的一帧。
```

Call it with:

```text
--prompt-preset couple_hotel_door_cctv
```

## Single-Female Restaurant Low-Angle Candid Preset

Preset id: `single_female_restaurant_low_angle`

Chinese name: `单女主西餐厅超低机位聊天抓拍`

- Upload only the randomly selected or user-provided female reference as 图1. Do not upload Gary or any other character reference.
- Read the full fixed Chinese prompt from [references/single-female-restaurant-low-angle-prompt.txt](references/single-female-restaurant-low-angle-prompt.txt) without rewriting or shortening it. Prepend only `生成竖屏9:16图片。` or `生成横屏16:9图片。` to lock the requested format.
- Only vertical `9:16` and horizontal `16:9` are allowed. Default to vertical `9:16` when the user does not specify one. Reject and do not report any other aspect ratio as a successful final output.

Call it with:

```text
--prompt-preset single_female_restaurant_low_angle
```

## Ambiguous Interaction Preset

Preset id: `ambiguous_interaction`

Chinese name: `暧昧互动`

Use this preset when the user says `暧昧互动` or asks to preserve the latest stable random feeling. This preset is not tied to one outdoor scene. It randomly varies the scene, time, lighting, background evidence, position relationship, camera phrasing, and a restrained intimate action while preserving the stable candid-phone look approved in the July 25 tests.

Stable core:

- Treat 图1 as the boyfriend identity reference and 图2 as the adult female lead identity reference.
- The generated image uses a first-person boyfriend POV. The boyfriend must not fully appear. Only one hand or a bit of sleeve may naturally appear at the edge of the frame, just reaching toward the female lead. No boyfriend face, head, body, third person, or extra character.
- The interaction must not be organized around props or objects. Do not use grabbing a phone, sunglasses, paper bag, remote control, coffee cup, room card, blanket, pillow, clothing, or any item as the random action.
- Randomize only human behavior and composition: distance, eye line, dodging, turning, restrained laughter, a hand raised to block, leaning away, leaning back, or briefly looking back at the boyfriend.
- Keep the action unfinished as a first frame for video: the hand has not reached the female lead yet, and her reaction has just started.
- Keep the tone intimate but restrained, playful, ordinary, non-sexual, non-suggestive, and non-violent.
- Keep the strongest visual feel: close half-body or three-quarter candid crop, slightly tilted phone composition, subject in the lower two-thirds, uneven light, local phone flash/fill, realistic skin highlights, darker or uneven background exposure, slight highlight spill, phone compression, mild motion blur, local defocus, real skin texture, and ordinary-person realism.
- The random scene may be indoor, semi-outdoor, or outdoor, but avoid clean travel-photo scenery, big blue sky, empty scenic background, soft portrait lighting, commercial outdoor portrait lighting, and overly polished composition.

Call it with:

```text
--prompt-preset ambiguous_interaction
```

The Mac runner implements this preset as `AMBIGUOUS_INTERACTION_CORE` with the `AMBIGUOUS_STABLE_*` random variable pools.

## Prompt Assembly

When generating, assemble the prompt with:

1. **Identity roles:** Fixed male lead from this skill; female lead from the user's current input.
2. **Interaction:** One natural couple/social interaction.
3. **Setting:** Real-life nightlife or social venue unless the user specifies another setting.
4. **Camera style:** Amateur smartphone photo, direct phone flash, slight handheld tilt, central composition, natural handheld crop, mild background defocus, slight motion blur.
5. **Realism constraints:** ordinary people, real skin, no beauty filter, no commercial polish.
6. **Negative prompt:** Use the fixed negative list in this skill and append any user-specific avoid items.

For Gary series variations:

- Keep the first line: `生成9:16竖版手机照片。`
- Replace only the scene phrase (for example: `酒吧`, `高级酒廊`, `高级餐厅`, `半私密卡座`, `包厢`, `KTV包间`, `朋友家客厅`, `酒店大堂酒廊`, `夜宵店/大排档室内`, `私房菜餐厅`) and the ambiguous interaction phrase.
- Keep the fixed identity, texture, and negative prompt wording unchanged unless the user explicitly asks to change them.
- Favor indoor or semi-indoor night scenes with mixed ambient light plus phone flash.
- Do not ask for full-body framing. Use close waist-up or half-body candid framing by default.
- If saving locally in this user's workflow, save outputs to `D:\工作用（同步）\图\长视频用图`.

## Interaction Bank

Use these only when the user does not specify an interaction:

- Woman hugs the man from behind while laughing toward the camera; man talks to someone outside frame and does not look at camera.
- Woman leans on the man's shoulder in a crowded bar booth; man turns slightly sideways mid-conversation.
- Woman pulls the man's sleeve while laughing; man holds a drink and looks off-frame naturally.
- Woman stands close beside the man with one arm loosely around his waist; both look caught mid-moment, not posed.
- Woman leans close beside the man while laughing toward the camera; man smiles lightly while looking away.
- Woman rests her chin near the man's shoulder from behind; man is relaxed and not directly posing.
- Woman raises a glass near the man; they are close together but with imperfect, spontaneous body language.

## Ambiguous Indoor Interaction Bank

Use these when the user asks for a more ambiguous/flirtatious but non-explicit indoor mood. Keep it natural, not erotic, not staged, and not full-body.

- In a private restaurant booth, the woman leans from behind near the man's shoulder with her face visible; one hand rests naturally on his shoulder or forearm; the man turns slightly and smiles low, not looking at the camera.
- In an upscale lounge booth, the woman loosely hugs the man from behind and laughs near his cheek; the man is mid-conversation with someone off-frame.
- At a dim restaurant table, the woman leans over the man's shoulder to look at something on his phone; their faces are close, and the man smiles while looking down.
- In a KTV/private room sofa corner, the woman rests her chin near the man's shoulder with a teasing smile; the man holds a glass and glances away.
- In a hotel lobby lounge, the woman lightly pulls the man's sleeve while leaning close; he turns sideways mid-sentence, relaxed and not posing.
- In a late-night indoor food stall or private dining room, the woman sits close beside him with one arm behind his back; both look caught mid-moment with imperfect body angles.

## Fixed Style Prompt

Use this style block in every image prompt:

```text
Photorealistic candid vertical 9:16 smartphone snapshot, real nightlife/social venue, casual friend gathering, not a staged shoot. Slightly tilted handheld composition, subjects near the center, direct phone flash / on-camera flash portrait look, mild local highlight clipping, realistic flash falloff, warm mixed ambient background light, mild background defocus, slight edge and hair motion blur. Natural skin texture, visible pores, minor blemishes, natural skin tone, ordinary-person realism. Looks like an original phone gallery social photo, strong life documentary feeling. No beauty filter, no retouching, no skin smoothing, no over-sharpening, no cinematic color grade, no commercial-photo polish.
```

Use natural social-photo framing. The image should usually be waist-up, half-body, close three-quarter, or a tight seated/booth crop. Avoid full-body shots by default. Do not force both characters' entire bodies into frame; tighter crops usually look more realistic for this series.

## Fixed Negative Prompt

```text
influencer face, model face, celebrity face, excessive beauty retouching, skin smoothing, plastic skin, waxy skin, overly perfect face, nose too high, nose too sharp, exaggerated makeup, fashion editorial, studio shoot, professional lighting, commercial photography, advertising photo, promotional photo, magazine cover, strong cinematic look, CG look, 3D look, anime look, obvious AI look, filter look, over-sharpened, overly clean lighting, fake background, staged pose, exaggerated pose, unnatural expression, everyone deliberately staring into camera, duplicated people, extra people, extra faces, extra fingers, malformed hands, deformed limbs, wrong facial structure, distorted facial features, inconsistent identity, changed face, low resolution, severe noise, dirty artifacts, watermark, text
```

## Generation Workflow

1. Before Gary-series generation, ask the user whether to use `真实抓拍照片提示词` or `高级酒店酒廊沙发抓拍照片提示词`, unless the choice is already explicit.
2. If the user provides a female reference image, use it as the female identity reference.
3. If the user provides only text, describe the female lead explicitly in the prompt and keep her ordinary, natural, and consistent with the user's description.
4. Use `assets/fixed-male-lead.png` as the male identity reference.
5. Use the built-in `image_gen` tool by default. If the user explicitly asks for CLI/API/model control, follow the system imagegen skill fallback rules.
6. For local generation, select and prompt for GPT Image 2 high-fidelity portrait consistency. GPT Image 2.5 Sunburst and Flare are not allowed fallbacks. Ask for vertical 9:16 output in the prompt.
7. Save project-bound final outputs under the current thread's `outputs` directory when possible; for the current user's Gary series, prefer `D:\工作用（同步）\图\长视频用图` when available. Otherwise show the generated image inline and report where it was saved.

## TeamoRouter Single-Female Route

Use `scripts/run_gary_single_teamorouter.py` when the user asks to run the existing heroine-reference workflow through TeamoRouter GPT Image 2.5 Sunburst.

- This route preserves the current Mac heroine library, random selection, fixed single-female prompt presets, dated output folder, manifest, and actual-dimension verification.
- The documented TeamoRouter edit endpoint accepts one `image` field. Therefore this adapter supports only single-female presets and uploads only the selected heroine reference. Do not claim that it preserves two independent identities.
- `--prompt-preset random` randomly chooses between `single_female_hotel_door_cctv` and `single_female_restaurant_low_angle`.
- Run `--dry-run` first. A live request requires the user's authorization and `--confirm-spend`; make one submission and never retry automatically.
- The TeamoRouter key must already be configured through the `teamorouter-image` skill. Never read or print it.

Example dry run:

```bash
python3 scripts/run_gary_single_teamorouter.py --female-count 1 --prompt-preset random --dry-run
```

## Prompt Template

```text
Use GPT Image 2 high-fidelity portrait identity consistency through the Codex built-in image-generation entry. Do not use GPT Image 2.5 Sunburst or Flare. Use the fixed male lead reference from this skill as the man. Use the user-provided female character as the woman. Generate a new photorealistic candid vertical 9:16 image, not an edit of the source references.

For the Gary series, use this Chinese structure when the user wants to vary only scene and flirtatious interaction:

生成9:16竖版手机照片。
{scene}的真实抓拍照片，{ambiguous_interaction}。使用图1和图2作为人物参考，保持两位人物的真实面貌、五官比例、发型、年龄感、体型和气质一致，不要美化成模特或网红脸。两人的动作不完全统一，像朋友聚会或临时合影时被旁边朋友随手拍下的一瞬间。场景是真实生活化背景。整体氛围像朋友圈原图、生活流纪实照片，不是商业摄影，不是宣传照。业余手机摄影风格，轻微歪斜构图，人物在图片中间位置，手机直闪效果，机顶闪光人像质感，局部高光轻微溢出，背景有轻微失焦，人物边缘有一点点运动模糊，保留真实皮肤纹理、毛孔、轻微瑕疵和自然肤色。画面不要过度锐化，不要精修，不要磨皮，不要滤镜感。人物都是普通人，不夸张打扮，女生鼻子不要过高，不要变成精致网红脸。真实手机夜拍质感，轻微颗粒感可以保留，但不要严重脏噪点。整体像朋友随手拍到的一张真实照片，生活感强，非摆拍，非电影感，非AI精修感。
负面提示词：网红脸，模特脸，过度美颜，磨皮，塑料皮肤，脸部过于完美，鼻子过高，夸张妆容，时尚大片，棚拍，专业打光，商业摄影，宣传照，杂志封面，电影感过强，CG感，3D感，动漫感，AI感过强，滤镜感，过度锐化，光线过于干净，背景虚假，刻意摆拍，夸张姿势，不自然表情，刻意直视镜头，人物重复，多余人物，多余手指，手部畸形，肢体畸形，脸部结构错误，五官变形，身份不一致，改变面貌，低清，严重噪点，画面脏污，水印，文字

Do not add full-body wording to this template. If a scene tends to produce standing full-body images, add only a minimal crop constraint such as `近距离半身抓拍，不要全身照，人物可以被自然裁切。`

Female lead:
{female_lead_description_or_reference_role}

Scene and interaction:
{interaction_scene}

Identity preservation:
Preserve the fixed male lead's real facial identity, facial proportions, short black hairstyle, mature understated age impression, body type, and temperament. Preserve the female lead's described/reference facial identity, hairstyle, age impression, body type, outfit direction, and temperament. Keep both people ordinary and natural. Do not beautify either person into a model, influencer, celebrity, or AI-polished face. Keep the woman's nose natural, not too high or sharp, unless the user specifically describes otherwise.

Style:
{fixed_style_prompt}

Negative prompt:
{fixed_negative_prompt}
```

