ai-photos
ai-photos turns one or more local photo sources into a searchable AI photo album for OpenClaw.
Supported formats:
- macOS:
jpg, jpeg, png, webp, heic
- Linux:
jpg, jpeg, png, webp
- Linux
heic: best-effort only; do not promise captioning or preview support
When talking to users:
- try to match the user's language
- explain the outcome simply: choose local folders now, then use OpenClaw to search and organize them
- stay focused on the current ai-photos request
- keep user-facing replies short and product-level: progress, readiness, and what the user can do next
- keep implementation details internal unless the user asks or troubleshooting requires them
- once indexing is complete and the backend is confirmed ready, say the album is ready and invite the user to try a search
- when the user asks what ai-photos can do, or when handing off a ready album, briefly describe the product in user terms:
- natural-language search across captions, scene labels, and tags
- date-based browsing and filtering
- a local web gallery for thumbnail browsing and large-photo viewing
- photo detail view with caption, scene, tags, capture time, device, location, orientation, and file info when available
- opening the original local file from the web UI
- manual sync now, optional automatic indexing later
- when introducing the web UI, describe it as a local searchable gallery rather than an API or server unless implementation details are needed
- keep these capability descriptions short, concrete, and user-facing; do not drift into backend details
Suggested user-facing capability summary:
- "You can search your photos in plain language, filter by date, and browse everything in a local gallery."
- "The web UI shows thumbnails, opens large previews, and lets you inspect captions, tags, time, device, location, and other file details when available."
- "You can also open the original local file directly, and later either sync changes manually or turn on automatic indexing."
Required outcome
This task is not complete until all of the following are true:
- at least one photo source is chosen and readable for a new album
- image analysis is verified to work in the current OpenClaw runtime
- the album backend is created or reconnected and writable
- the first import succeeds, or an existing album is verified reachable
- the user explicitly approved automatic indexing or explicitly declined it
- if automatic indexing was approved, OpenClaw heartbeat is configured without breaking existing heartbeat tasks, the ai-photos block is present in
HEARTBEAT.md, and one verification heartbeat has run
- the user has been told the album is ready and has been invited to try a search
- the user has been sent the final handoff
Internal terms
Use these terms for agent reasoning, troubleshooting, or recovery only.
Do not introduce them to the user unless needed.
photo sources: one or more local paths scanned into the same album
album backend: where the searchable photo index is stored
album profile: saved reconnect information, stored automatically under ~/.openclaw/ai-photos/albums/default.json
caption input JSONL: the manifest file that still needs vision captions and import
If the user asks what to save for later, explain that OpenClaw saves the reconnect information automatically at ~/.openclaw/ai-photos/albums/default.json, and that they only need to keep that file if they want a manual backup.
Caption schema
Each captioned JSONL line should contain the original manifest fields plus vision-model output.
Required base fields:
file_path
filename
sha256
mime_type
size_bytes
width
height
taken_at
exif
Vision fields:
caption: one short factual sentence
tags: array of 5-12 short tags
scene: short scene label
objects: array of the main visible objects
text_in_image: visible text or null
Optional fields:
metadata: free-form JSON object
search_text: concatenated retrieval text; if omitted, the importer builds it
Example:
{
"file_path": "/photos/2026/03/cat.jpg",
"filename": "cat.jpg",
"sha256": "abc123",
"mime_type": "image/jpeg",
"size_bytes": 231231,
"width": 3024,
"height": 4032,
"taken_at": "2026-03-12T09:12:00+00:00",
"exif": {"Make": "Apple", "Model": "iPhone 15 Pro"},
"caption": "A white cat resting on a gray sofa near a sunlit window.",
"tags": ["cat", "sofa", "indoor", "sunlight", "pet"],
"scene": "living room",
"objects": ["cat", "sofa", "window"],
"text_in_image": null,
"metadata": {"source": "demo"}
}
CLI runtime
This skill does not depend on a local Python environment or a checked-out Go source tree.
It uses the latest published ai-photos CLI release from:
- repository:
https://github.com/zoubingwu/openclaw-ai-photos
- install dir:
~/.openclaw/ai-photos/bin
- binary path:
~/.openclaw/ai-photos/bin/ai-photos
At the start of every ai-photos task, run the bootstrap flow exactly once and reuse the resulting binary path for the rest of the task.
Bootstrap flow
Run this shell block and capture its stdout as AI_PHOTOS_BIN:
ensure_ai_photos() {
AI_PHOTOS_REPO="zoubingwu/openclaw-ai-photos"
AI_PHOTOS_BIN_DIR="$HOME/.openclaw/ai-photos/bin"
AI_PHOTOS_BIN="$AI_PHOTOS_BIN_DIR/ai-photos"
mkdir -p "$AI_PHOTOS_BIN_DIR"
os="$(uname -s | tr '[:upper:]' '[:lower:]')"
case "$os" in
darwin) goos="darwin" ;;
linux) goos="linux" ;;
*)
echo "unsupported platform: $os" >&2
return 1
;;
esac
arch="$(uname -m)"
case "$arch" in
x86_64|amd64) goarch="amd64" ;;
arm64|aarch64) goarch="arm64" ;;
*)
echo "unsupported architecture: $arch" >&2
return 1
;;
esac
archive_name="ai-photos_${goos}_${goarch}.tar.gz"
archive_url="https://github.com/${AI_PHOTOS_REPO}/releases/latest/download/${archive_name}"
tmp_dir="$(mktemp -d)"
had_existing_binary=0
if [ -x "$AI_PHOTOS_BIN" ]; then
had_existing_binary=1
fi
if curl -fL "${archive_url}" -o "$tmp_dir/${archive_name}" \
&& tar -xzf "$tmp_dir/${archive_name}" -C "$tmp_dir" \
&& install -m 0755 "$tmp_dir/ai-photos" "$AI_PHOTOS_BIN"; then
rm -rf "$tmp_dir"
printf '%s\n' "$AI_PHOTOS_BIN"
return 0
fi
rm -rf "$tmp_dir"
if [ "$had_existing_binary" -eq 1 ]; then
printf '%s\n' "$AI_PHOTOS_BIN"
return 0
fi
echo "could not download ai-photos release archive" >&2
return 1
}
AI_PHOTOS_BIN="$(ensure_ai_photos)"
Rules:
- always run the bootstrap flow before using the CLI
- the bootstrap flow downloads the latest stable release asset from
releases/latest/download/... and does not call api.github.com
- if the latest asset download or unpack step fails, continue with the cached binary when one already exists
- if the latest asset download fails and no cached binary exists, setup is blocked
- do not tell the user to clone the repository or build the CLI locally unless troubleshooting requires it
- if you need command details, use
"$AI_PHOTOS_BIN" help or "$AI_PHOTOS_BIN" help <subcommand>
Onboarding
Step 0 - Choose mode
User-facing:
- Ask whether the user wants to create a new photo album, reconnect an existing one, or search an already configured album.
- If they want to reconnect, explain that you will try the saved connection first and only ask for more details if needed.
[AGENT] Branching:
1: continue to Step 1
2: continue to Step 3 and Step 4
3: go directly to Search flow
- if the user wants search but no configured album exists, tell them setup is required first
Step 1 - Ask for photo folders
User-facing:
- Ask for one or more local folder paths that contain photos.
[AGENT]
Do not continue until the user has provided at least one photo source.
Step 2 - Run preflight
User-facing:
- Tell the user you will quickly verify that the folders are readable and that image analysis works before importing anything.
[AGENT]
Before indexing anything, verify:
- each photo source exists and is readable
- the selected sources contain supported image files
agents.defaults.imageModel is vision-capable
- image analysis actually works on a real image in the current OpenClaw runtime
- the installed CLI runs successfully
- local image preparation works on a real sample image through
"$AI_PHOTOS_BIN" prepare-image
Suggested preflight sequence:
- choose one real sample image from the provided sources
- run
"$AI_PHOTOS_BIN" prepare-image --mode caption <sample-file>
- on macOS, also run
"$AI_PHOTOS_BIN" prepare-image --mode preview <sample-file>
- inspect the JSON result
If the image backend check fails:
- on macOS, treat this as blocking because
heic and local preview preparation depend on sips
- on Linux, do not block setup for
jpg, jpeg, png, or webp; OpenClaw can still caption those files directly from the original path
- on Linux, explain that preview preparation and large-image downscaling are reduced without a local backend
- only suggest installing ImageMagick when the user wants better local image preparation or troubleshooting requires it
If preflight fails:
- tell the user setup is blocked in plain language
- explain exactly what must be fixed without exposing unnecessary implementation details
- stop
Step 3 - Choose the backend
[AGENT]
- if reconnecting, keep the existing backend
- otherwise use
db9 if it is installed and usable
- if
db9 is not available, use TiDB Cloud Zero
- if using
TiDB Cloud Zero, tell the user to claim it if they want to keep it, but do not lead with backend details unless they matter
Step 4 - Create or reconnect the album
User-facing for a new album:
- Tell the user setup is in progress and that the selected folders will be searchable through OpenClaw when it finishes.
- If useful, add one short product sentence such as: "You'll be able to search in plain language or browse everything in the local gallery once import finishes."
[AGENT]
For a new album, run exactly one setup command:
# db9
"$AI_PHOTOS_BIN" setup --source <photo-source-a> --source <photo-source-b> --backend db9 --target <db>
# TiDB
"$AI_PHOTOS_BIN" setup --source <photo-source-a> --source <photo-source-b> --backend tidb --target /path/to/tidb-target.json
Read the JSON output:
profile_path tells you where the default album profile was saved
caption_input_jsonl is the input for the first record ingestion pass
sync.to_caption tells you how many records still need captions and import
[AGENT] For reconnect:
- try the saved default album profile first
- verify the backend is reachable
- verify the album can be searched or written
- ask only for missing backend details
Suggested reconnect check:
"$AI_PHOTOS_BIN" search --recent --limit 1
Do not continue until the backend is confirmed reachable.
Step 5 - Run the shared record ingestion flow
Use this same flow for:
- the first album import
- later incremental updates
User-facing:
- Tell the user photos are being imported and that large libraries may take some time.
[AGENT]
Input:
- first import:
caption_input_jsonl from ai-photos setup
- later updates:
incremental_manifest_jsonl from ai-photos sync
Before generating records, read the Caption schema section in this file.
[AGENT] For each record in the input manifest:
- run
"$AI_PHOTOS_BIN" prepare-image --mode caption <file_path>
- send the returned
output_path to the vision-capable model
- preserve the original manifest fields from the source image
- add
caption, tags, scene, objects, and text_in_image
- write one JSON object per line into a captioned JSONL file
- import it with:
"$AI_PHOTOS_BIN" import /tmp/photos.captioned.jsonl
Rules:
- keep captions short, factual, retrieval-oriented, and visually grounded
prepare-image prefers macOS sips when available and also supports ImageMagick for Linux-friendly setups
- if
prepare-image returns the original file path in caption mode, continue with that file instead of blocking the batch
- on Linux, allow direct caption fallback for
jpg, jpeg, png, and webp when no local image backend is available
- do not promise Linux
heic captioning or preview support
- do not invent names, sensitive traits, or stories
- do not replace the original
file_path with the temporary derived image path
- if one file still cannot be captioned, skip only that file and continue the rest of the batch
- if there is nothing to caption, skip this step
Step 6 - Enable automatic indexing
User-facing:
- Offer automatic indexing in plain language.
- Explain that OpenClaw can periodically check the selected folders for new or changed photos and update the album index.
- Ask whether the user wants to enable that now.
[AGENT]
If the user declines:
- skip this step
- do not change heartbeat config
- do not change
HEARTBEAT.md
If the user says yes:
- inspect the existing heartbeat config before changing anything
- do not overwrite or replace existing heartbeat tasks
- do not tell the user to manually restart Gateway for heartbeat-only changes
- let OpenClaw handle heartbeat configuration using its normal mechanisms unless debugging requires lower-level manual steps
- reuse the existing heartbeat scope and workspace whenever possible
- if there is more than one reasonable heartbeat-enabled scope, do not guess; ask the user which one should own ai-photos automatic indexing
- do not convert an existing per-agent heartbeat setup back into a defaults-based setup
- preserve existing heartbeat behavior unless a missing setting must be filled with a reasonable default
- do not spell out or rely on a fixed command recipe unless the current environment requires debugging
Then update <workspace>/HEARTBEAT.md without removing unrelated content:
- if the file does not exist, create it
- if the file exists, preserve all existing user content
- manage only one ai-photos block delimited by stable markers
- if the ai-photos block already exists, replace only that block
- if the ai-photos block does not exist, append it to the end of the file
<!-- ai-photos:auto-indexing:start -->
## ai-photos automatic indexing
- Read and learn how to use `ai-photos` skill
- Use `~/.openclaw/ai-photos/bin/ai-photos sync` to scan the configured photo folders for changes.
- Check the configured photo folders for changes and keep the album index up to date.
- If `to_caption` is `0`, it means nothing needs attention, reply `HEARTBEAT_OK`.
- If `to_caption` is greater than `0`, run the shared record ingestion flow using `incremental_manifest_jsonl`.
- Stay quiet unless indexing failed or user action is needed.
<!-- ai-photos:auto-indexing:end -->
Do not rewrite the whole file just to add this block.
Then verify once:
- trigger one heartbeat run if it is safe and practical in the current environment, otherwise wait for the next scheduled run
- check the heartbeat result and make sure the ai-photos task completed as intended
- do not claim success until the verification result is clear
Then tell the user the result:
- success: explain that automatic indexing is active and the verification run succeeded
- declined: explain that the album is ready, but future changes require a manual re-index
- failed: explain that the album is usable, but automatic indexing is not active yet
Step 7 - Final handoff
User-facing handoff should include:
- that the album is ready to use
- how the user can use it now: search in plain language or ask OpenClaw to help organize photos
- whether automatic indexing is on or off, in one short sentence only when it matters
- when useful, mention the local gallery capabilities in one short sentence: browse thumbnails, open large previews, inspect metadata, and open the original file
Keep the handoff short and user-facing.
Default to readiness, status, and next actions.
Only include implementation details when the user asks or recovery requires them.
[AGENT]
Immediately after setup:
- hand off directly once setup is ready
- tell the user the album is ready to search
- invite the user to search in plain language or ask OpenClaw to help organize photos
- if the user declined automatic indexing, say clearly that the album is in manual-only indexing mode
Search flow
When the user asks to find photos, run:
"$AI_PHOTOS_BIN" search --text "cat on sofa"
"$AI_PHOTOS_BIN" search --tag cat
"$AI_PHOTOS_BIN" search --date 2026-03
"$AI_PHOTOS_BIN" search --recent
When answering:
- summarize the best matches clearly and in plain language
- mention filenames, dates, or captions when useful
- answer at the product level unless the user asks for implementation details
- before sending an image file, run
"$AI_PHOTOS_BIN" prepare-image --mode preview <matched-file>
- send the returned
output_path when possible
- if preview preparation fails on Linux without a local image backend, say so briefly and fall back to the original file only when it is safe to send as-is
- if results are weak, say so and suggest a better query
Local web search
When the user asks to open a browser view for the album:
- start the local web service
- prefer the saved album profile; use environment variables only to fill missing backend fields
- wait for the JSON startup line and return the local URL to the user
- keep the process running while the user is browsing
If the user wants to open the gallery from another device:
- recommend Tailscale as the default remote access path
- run
"$AI_PHOTOS_BIN" serve --host 0.0.0.0 only when they explicitly want remote access
- explain that the startup JSON still prints a browser URL for the machine running
ai-photos; for remote access, share the machine's Tailscale IP or MagicDNS name instead
- do not recommend exposing the port directly to the public internet unless the user explicitly asks for that tradeoff
- clarify that "open original" opens the file on the machine running
ai-photos, not on the remote client
Run:
"$AI_PHOTOS_BIN" serve
If the user wants a specific album profile:
"$AI_PHOTOS_BIN" serve --profile default
The web service provides:
- a local search page
- search/filter/detail APIs for the page
- thumbnail and preview endpoints
- an action to open the original local file on the machine running
ai-photos
When handing the web UI to the user:
- describe the page in product terms, for example:
- "The page lets you search in plain language, filter by date, scroll the gallery, open a large preview, and inspect metadata on the right."
- "When a photo has metadata, the detail panel can show caption, scene, tags, capture time, device, location, orientation, and file info."
- prefer this product summary over technical endpoint descriptions unless the user is debugging
Heartbeat run behavior
When a heartbeat arrives for a configured album:
- run:
"$AI_PHOTOS_BIN" sync
- read the JSON output
- if
to_caption is 0, return HEARTBEAT_OK
- if
to_caption is greater than 0, run the shared record ingestion flow using incremental_manifest_jsonl
- stay quiet unless indexing failed or user attention is needed
1---2name: ai-photos3description: Personal AI photo album for OpenClaw. Use when users say: - "index my photos" - "set up an AI photo album" - "search my photo library" - "reconnect my photo album" - "find photos of ..."4---56# ai-photos78ai-photos turns one or more local photo sources into a searchable AI photo album for OpenClaw.910Supported formats:11- macOS: `jpg`, `jpeg`, `png`, `webp`, `heic`12- Linux: `jpg`, `jpeg`, `png`, `webp`13- Linux `heic`: best-effort only; do not promise captioning or preview support1415When talking to users:16- try to match the user's language17- explain the outcome simply: choose local folders now, then use OpenClaw to search and organize them18- stay focused on the current ai-photos request19- keep user-facing replies short and product-level: progress, readiness, and what the user can do next20- keep implementation details internal unless the user asks or troubleshooting requires them21- once indexing is complete and the backend is confirmed ready, say the album is ready and invite the user to try a search22- when the user asks what ai-photos can do, or when handing off a ready album, briefly describe the product in user terms:23 - natural-language search across captions, scene labels, and tags24 - date-based browsing and filtering25 - a local web gallery for thumbnail browsing and large-photo viewing26 - photo detail view with caption, scene, tags, capture time, device, location, orientation, and file info when available27 - opening the original local file from the web UI28 - manual sync now, optional automatic indexing later29- when introducing the web UI, describe it as a local searchable gallery rather than an API or server unless implementation details are needed30- keep these capability descriptions short, concrete, and user-facing; do not drift into backend details3132Suggested user-facing capability summary:3334- "You can search your photos in plain language, filter by date, and browse everything in a local gallery."35- "The web UI shows thumbnails, opens large previews, and lets you inspect captions, tags, time, device, location, and other file details when available."36- "You can also open the original local file directly, and later either sync changes manually or turn on automatic indexing."3738## Required outcome3940This task is not complete until all of the following are true:41421. at least one photo source is chosen and readable for a new album432. image analysis is verified to work in the current OpenClaw runtime443. the album backend is created or reconnected and writable454. the first import succeeds, or an existing album is verified reachable465. the user explicitly approved automatic indexing or explicitly declined it476. if automatic indexing was approved, OpenClaw heartbeat is configured without breaking existing heartbeat tasks, the ai-photos block is present in `HEARTBEAT.md`, and one verification heartbeat has run487. the user has been told the album is ready and has been invited to try a search498. the user has been sent the final handoff5051## Internal terms5253Use these terms for agent reasoning, troubleshooting, or recovery only.54Do not introduce them to the user unless needed.5556- `photo sources`: one or more local paths scanned into the same album57- `album backend`: where the searchable photo index is stored58- `album profile`: saved reconnect information, stored automatically under `~/.openclaw/ai-photos/albums/default.json`59- `caption input JSONL`: the manifest file that still needs vision captions and import6061If the user asks what to save for later, explain that OpenClaw saves the reconnect information automatically at `~/.openclaw/ai-photos/albums/default.json`, and that they only need to keep that file if they want a manual backup.6263## Caption schema6465Each captioned JSONL line should contain the original manifest fields plus vision-model output.6667Required base fields:68- `file_path`69- `filename`70- `sha256`71- `mime_type`72- `size_bytes`73- `width`74- `height`75- `taken_at`76- `exif`7778Vision fields:79- `caption`: one short factual sentence80- `tags`: array of 5-12 short tags81- `scene`: short scene label82- `objects`: array of the main visible objects83- `text_in_image`: visible text or `null`8485Optional fields:86- `metadata`: free-form JSON object87- `search_text`: concatenated retrieval text; if omitted, the importer builds it8889Example:9091```json92{93 "file_path": "/photos/2026/03/cat.jpg",94 "filename": "cat.jpg",95 "sha256": "abc123",96 "mime_type": "image/jpeg",97 "size_bytes": 231231,98 "width": 3024,99 "height": 4032,100 "taken_at": "2026-03-12T09:12:00+00:00",101 "exif": {"Make": "Apple", "Model": "iPhone 15 Pro"},102 "caption": "A white cat resting on a gray sofa near a sunlit window.",103 "tags": ["cat", "sofa", "indoor", "sunlight", "pet"],104 "scene": "living room",105 "objects": ["cat", "sofa", "window"],106 "text_in_image": null,107 "metadata": {"source": "demo"}108}109```110111## CLI runtime112113This skill does not depend on a local Python environment or a checked-out Go source tree.114It uses the latest published `ai-photos` CLI release from:115116- repository: `https://github.com/zoubingwu/openclaw-ai-photos`117- install dir: `~/.openclaw/ai-photos/bin`118- binary path: `~/.openclaw/ai-photos/bin/ai-photos`119120At the start of every ai-photos task, run the bootstrap flow exactly once and reuse the resulting binary path for the rest of the task.121122### Bootstrap flow123124Run this shell block and capture its stdout as `AI_PHOTOS_BIN`:125126```bash127ensure_ai_photos() {128 AI_PHOTOS_REPO="zoubingwu/openclaw-ai-photos"129 AI_PHOTOS_BIN_DIR="$HOME/.openclaw/ai-photos/bin"130 AI_PHOTOS_BIN="$AI_PHOTOS_BIN_DIR/ai-photos"131132 mkdir -p "$AI_PHOTOS_BIN_DIR"133134 os="$(uname -s | tr '[:upper:]' '[:lower:]')"135 case "$os" in136 darwin) goos="darwin" ;;137 linux) goos="linux" ;;138 *)139 echo "unsupported platform: $os" >&2140 return 1141 ;;142 esac143144 arch="$(uname -m)"145 case "$arch" in146 x86_64|amd64) goarch="amd64" ;;147 arm64|aarch64) goarch="arm64" ;;148 *)149 echo "unsupported architecture: $arch" >&2150 return 1151 ;;152 esac153154 archive_name="ai-photos_${goos}_${goarch}.tar.gz"155 archive_url="https://github.com/${AI_PHOTOS_REPO}/releases/latest/download/${archive_name}"156 tmp_dir="$(mktemp -d)"157 had_existing_binary=0158 if [ -x "$AI_PHOTOS_BIN" ]; then159 had_existing_binary=1160 fi161162 if curl -fL "${archive_url}" -o "$tmp_dir/${archive_name}" \163 && tar -xzf "$tmp_dir/${archive_name}" -C "$tmp_dir" \164 && install -m 0755 "$tmp_dir/ai-photos" "$AI_PHOTOS_BIN"; then165 rm -rf "$tmp_dir"166 printf '%s\n' "$AI_PHOTOS_BIN"167 return 0168 fi169170 rm -rf "$tmp_dir"171 if [ "$had_existing_binary" -eq 1 ]; then172 printf '%s\n' "$AI_PHOTOS_BIN"173 return 0174 fi175176 echo "could not download ai-photos release archive" >&2177 return 1178}179180AI_PHOTOS_BIN="$(ensure_ai_photos)"181```182183Rules:184- always run the bootstrap flow before using the CLI185- the bootstrap flow downloads the latest stable release asset from `releases/latest/download/...` and does not call `api.github.com`186- if the latest asset download or unpack step fails, continue with the cached binary when one already exists187- if the latest asset download fails and no cached binary exists, setup is blocked188- do not tell the user to clone the repository or build the CLI locally unless troubleshooting requires it189- if you need command details, use `"$AI_PHOTOS_BIN" help` or `"$AI_PHOTOS_BIN" help <subcommand>`190191## Onboarding192193### Step 0 - Choose mode194195User-facing:196197- Ask whether the user wants to create a new photo album, reconnect an existing one, or search an already configured album.198- If they want to reconnect, explain that you will try the saved connection first and only ask for more details if needed.199200`[AGENT]` Branching:201202- `1`: continue to Step 1203- `2`: continue to Step 3 and Step 4204- `3`: go directly to Search flow205- if the user wants search but no configured album exists, tell them setup is required first206207### Step 1 - Ask for photo folders208209User-facing:210211- Ask for one or more local folder paths that contain photos.212213`[AGENT]`214215Do not continue until the user has provided at least one photo source.216217### Step 2 - Run preflight218219User-facing:220221- Tell the user you will quickly verify that the folders are readable and that image analysis works before importing anything.222223`[AGENT]`224225Before indexing anything, verify:226- each photo source exists and is readable227- the selected sources contain supported image files228- `agents.defaults.imageModel` is vision-capable229- image analysis actually works on a real image in the current OpenClaw runtime230- the installed CLI runs successfully231- local image preparation works on a real sample image through `"$AI_PHOTOS_BIN" prepare-image`232233Suggested preflight sequence:2341. choose one real sample image from the provided sources2352. run `"$AI_PHOTOS_BIN" prepare-image --mode caption <sample-file>`2363. on macOS, also run `"$AI_PHOTOS_BIN" prepare-image --mode preview <sample-file>`2374. inspect the JSON result238239If the image backend check fails:240- on macOS, treat this as blocking because `heic` and local preview preparation depend on `sips`241- on Linux, do not block setup for `jpg`, `jpeg`, `png`, or `webp`; OpenClaw can still caption those files directly from the original path242- on Linux, explain that preview preparation and large-image downscaling are reduced without a local backend243- only suggest installing ImageMagick when the user wants better local image preparation or troubleshooting requires it244245If preflight fails:246- tell the user setup is blocked in plain language247- explain exactly what must be fixed without exposing unnecessary implementation details248- stop249250### Step 3 - Choose the backend251252`[AGENT]`253254- if reconnecting, keep the existing backend255- otherwise use `db9` if it is installed and usable256- if `db9` is not available, use `TiDB Cloud Zero`257- if using `TiDB Cloud Zero`, tell the user to claim it if they want to keep it, but do not lead with backend details unless they matter258259### Step 4 - Create or reconnect the album260261User-facing for a new album:262263- Tell the user setup is in progress and that the selected folders will be searchable through OpenClaw when it finishes.264- If useful, add one short product sentence such as: "You'll be able to search in plain language or browse everything in the local gallery once import finishes."265266`[AGENT]`267268For a new album, run exactly one setup command:269270```bash271# db9272"$AI_PHOTOS_BIN" setup --source <photo-source-a> --source <photo-source-b> --backend db9 --target <db>273274# TiDB275"$AI_PHOTOS_BIN" setup --source <photo-source-a> --source <photo-source-b> --backend tidb --target /path/to/tidb-target.json276```277278Read the JSON output:279- `profile_path` tells you where the default album profile was saved280- `caption_input_jsonl` is the input for the first record ingestion pass281- `sync.to_caption` tells you how many records still need captions and import282283`[AGENT]` For reconnect:284- try the saved default album profile first285- verify the backend is reachable286- verify the album can be searched or written287- ask only for missing backend details288289Suggested reconnect check:290291```bash292"$AI_PHOTOS_BIN" search --recent --limit 1293```294295Do not continue until the backend is confirmed reachable.296297### Step 5 - Run the shared record ingestion flow298299Use this same flow for:300- the first album import301- later incremental updates302303User-facing:304305- Tell the user photos are being imported and that large libraries may take some time.306307`[AGENT]`308309Input:310- first import: `caption_input_jsonl` from `ai-photos setup`311- later updates: `incremental_manifest_jsonl` from `ai-photos sync`312313Before generating records, read the `Caption schema` section in this file.314315`[AGENT]` For each record in the input manifest:3161. run `"$AI_PHOTOS_BIN" prepare-image --mode caption <file_path>`3172. send the returned `output_path` to the vision-capable model3183. preserve the original manifest fields from the source image3194. add `caption`, `tags`, `scene`, `objects`, and `text_in_image`3205. write one JSON object per line into a captioned JSONL file3216. import it with:322323```bash324"$AI_PHOTOS_BIN" import /tmp/photos.captioned.jsonl325```326327Rules:328- keep captions short, factual, retrieval-oriented, and visually grounded329- `prepare-image` prefers macOS `sips` when available and also supports ImageMagick for Linux-friendly setups330- if `prepare-image` returns the original file path in caption mode, continue with that file instead of blocking the batch331- on Linux, allow direct caption fallback for `jpg`, `jpeg`, `png`, and `webp` when no local image backend is available332- do not promise Linux `heic` captioning or preview support333- do not invent names, sensitive traits, or stories334- do not replace the original `file_path` with the temporary derived image path335- if one file still cannot be captioned, skip only that file and continue the rest of the batch336- if there is nothing to caption, skip this step337338### Step 6 - Enable automatic indexing339340User-facing:341342- Offer automatic indexing in plain language.343- Explain that OpenClaw can periodically check the selected folders for new or changed photos and update the album index.344- Ask whether the user wants to enable that now.345346`[AGENT]`347348If the user declines:349- skip this step350- do not change heartbeat config351- do not change `HEARTBEAT.md`352353If the user says yes:354- inspect the existing heartbeat config before changing anything355- do not overwrite or replace existing heartbeat tasks356- do not tell the user to manually restart Gateway for heartbeat-only changes357- let OpenClaw handle heartbeat configuration using its normal mechanisms unless debugging requires lower-level manual steps358- reuse the existing heartbeat scope and workspace whenever possible359- if there is more than one reasonable heartbeat-enabled scope, do not guess; ask the user which one should own ai-photos automatic indexing360- do not convert an existing per-agent heartbeat setup back into a defaults-based setup361- preserve existing heartbeat behavior unless a missing setting must be filled with a reasonable default362- do not spell out or rely on a fixed command recipe unless the current environment requires debugging363364Then update `<workspace>/HEARTBEAT.md` without removing unrelated content:365- if the file does not exist, create it366- if the file exists, preserve all existing user content367- manage only one ai-photos block delimited by stable markers368- if the ai-photos block already exists, replace only that block369- if the ai-photos block does not exist, append it to the end of the file370371```md372<!-- ai-photos:auto-indexing:start -->373## ai-photos automatic indexing374375- Read and learn how to use `ai-photos` skill376- Use `~/.openclaw/ai-photos/bin/ai-photos sync` to scan the configured photo folders for changes.377- Check the configured photo folders for changes and keep the album index up to date.378- If `to_caption` is `0`, it means nothing needs attention, reply `HEARTBEAT_OK`.379- If `to_caption` is greater than `0`, run the shared record ingestion flow using `incremental_manifest_jsonl`.380- Stay quiet unless indexing failed or user action is needed.381<!-- ai-photos:auto-indexing:end -->382```383384Do not rewrite the whole file just to add this block.385386Then verify once:387- trigger one heartbeat run if it is safe and practical in the current environment, otherwise wait for the next scheduled run388- check the heartbeat result and make sure the ai-photos task completed as intended389- do not claim success until the verification result is clear390391Then tell the user the result:392- success: explain that automatic indexing is active and the verification run succeeded393- declined: explain that the album is ready, but future changes require a manual re-index394- failed: explain that the album is usable, but automatic indexing is not active yet395396### Step 7 - Final handoff397398User-facing handoff should include:399- that the album is ready to use400- how the user can use it now: search in plain language or ask OpenClaw to help organize photos401- whether automatic indexing is on or off, in one short sentence only when it matters402- when useful, mention the local gallery capabilities in one short sentence: browse thumbnails, open large previews, inspect metadata, and open the original file403404Keep the handoff short and user-facing.405Default to readiness, status, and next actions.406Only include implementation details when the user asks or recovery requires them.407408`[AGENT]`409410Immediately after setup:411- hand off directly once setup is ready412- tell the user the album is ready to search413- invite the user to search in plain language or ask OpenClaw to help organize photos414- if the user declined automatic indexing, say clearly that the album is in manual-only indexing mode415416## Search flow417418When the user asks to find photos, run:419420```bash421"$AI_PHOTOS_BIN" search --text "cat on sofa"422"$AI_PHOTOS_BIN" search --tag cat423"$AI_PHOTOS_BIN" search --date 2026-03424"$AI_PHOTOS_BIN" search --recent425```426427When answering:428- summarize the best matches clearly and in plain language429- mention filenames, dates, or captions when useful430- answer at the product level unless the user asks for implementation details431- before sending an image file, run `"$AI_PHOTOS_BIN" prepare-image --mode preview <matched-file>`432- send the returned `output_path` when possible433- if preview preparation fails on Linux without a local image backend, say so briefly and fall back to the original file only when it is safe to send as-is434- if results are weak, say so and suggest a better query435436## Local web search437438When the user asks to open a browser view for the album:4394401. start the local web service4412. prefer the saved album profile; use environment variables only to fill missing backend fields4423. wait for the JSON startup line and return the local URL to the user4434. keep the process running while the user is browsing444445If the user wants to open the gallery from another device:446- recommend Tailscale as the default remote access path447- run `"$AI_PHOTOS_BIN" serve --host 0.0.0.0` only when they explicitly want remote access448- explain that the startup JSON still prints a browser URL for the machine running `ai-photos`; for remote access, share the machine's Tailscale IP or MagicDNS name instead449- do not recommend exposing the port directly to the public internet unless the user explicitly asks for that tradeoff450- clarify that "open original" opens the file on the machine running `ai-photos`, not on the remote client451452Run:453454```bash455"$AI_PHOTOS_BIN" serve456```457458If the user wants a specific album profile:459460```bash461"$AI_PHOTOS_BIN" serve --profile default462```463464The web service provides:465- a local search page466- search/filter/detail APIs for the page467- thumbnail and preview endpoints468- an action to open the original local file on the machine running `ai-photos`469470When handing the web UI to the user:471- describe the page in product terms, for example:472 - "The page lets you search in plain language, filter by date, scroll the gallery, open a large preview, and inspect metadata on the right."473 - "When a photo has metadata, the detail panel can show caption, scene, tags, capture time, device, location, orientation, and file info."474- prefer this product summary over technical endpoint descriptions unless the user is debugging475476## Heartbeat run behavior477478When a heartbeat arrives for a configured album:4794801. run:481482```bash483"$AI_PHOTOS_BIN" sync484```4854862. read the JSON output4873. if `to_caption` is `0`, return `HEARTBEAT_OK`4884. if `to_caption` is greater than `0`, run the shared record ingestion flow using `incremental_manifest_jsonl`4895. stay quiet unless indexing failed or user attention is needed