When to Use
Use when a task needs a REAL or ARCHIVAL photo / clip (hero, texture, reference, historical footage) or a reaction / animated GIF, rather than a generated one — fan out across free image/video/GIF sources in one query and download license-tagged results.
Source: connerkward/web-media-getter-skill (MIT).
web-media
Query many free image/video sources in one fan-out, get a normalized result list,
optionally download top-K with an attribution sidecar. Zero-dep stdlib script.
Script: webmedia.py (in this dir). Keys: PEXELS_API_KEY, PIXABAY_API_KEY
in central/.env (optional — the 5 no-key sources work without them).
Sources
| Source |
Key? |
Best for |
Media |
| openverse |
none |
CC web images (Flickr, museums) |
image |
| wikimedia |
none |
factual / historical / landmark photos |
image |
| internetarchive |
none |
historical/archival images + films |
image, video |
| loc |
none |
historical US prints/photos |
image |
| nasa |
none |
space imagery + video |
image, video |
| pexels |
free key |
modern stock photos + short video clips |
image, video |
| pixabay |
free key |
modern photos/illustrations + short clips |
image, video |
| klipy |
free key |
GIFs — recommended (free, unlimited, Tenor drop-in) |
gif |
| giphy |
free key |
GIFs — biggest library (prod key needs approval) |
gif |
GIF sources fire only with --type gif. Keys: KLIPY_API_KEY, GIPHY_API_KEY
in central/.env. (tenor adapter removed — Google EOL'd the API 2026-06-30.)
klipy is the one to get (free + unlimited);
its adapter is unverified — assumes Tenor-compatible request/response;
verify against docs.klipy.com when you key it. webmedia.py "shrug" --type gif --count 6 --json
Usage
webmedia.py "1950s street scene" --type image --count 8 --json
webmedia.py "rocket launch" --type video --source nasa,internetarchive
webmedia.py "car factory 1930s" --source all --download --out /tmp/cars
--source all (default) | nokey (no-key only) | comma list (wikimedia,pexels)
--type image|video · --count N · --json · --download --out DIR
--download fetches each result's direct media URL and writes attribution.json
(source, author, license, url, page_url) alongside the files.
Record schema
{source, title, url, thumb, dl, page_url, author, license, w, h, type} —
dl is the directly-downloadable media URL (None when only a page exists).
The video caveat (important)
Archival sources (Internet Archive, Europeana, LoC) host whole films/documentaries,
not single shots. So:
- Modern single clip →
pexels / pixabay (born as short clips, direct MP4). Done.
- Historical single shot → retrieve the IA film here, then extract the shot:
- Twelve Labs Marengo search (free 600 min) — pass the IA public MP4 URL, get a
timestamped moment for "car on assembly line", clip with ffmpeg. Semantic, cheap.
- or PySceneDetect (free, local) to cut the film into shots, then rank keyframes
with CLIP via the
muser skill. Fully offline.
Audio: freesound + audio QA
webmedia.py is image/video. For sound effects (real, CC-licensed) and for
judging audio (since Claude can't hear), two sibling scripts live in
central/scripts/:
freesound-fetch.py "<query>" [count] [max_sec] [out_dir] — searches freesound.org
and downloads short hq-mp3 previews. Prints one JSON line per file with
license/user for attribution. Key: FREESOUND_API_KEY in central/.env
(token-based read; full originals would need OAuth — previews suffice for SFX).
audio-judge.py <file> "<target>" — sends the clip to OpenAI gpt-audio
(audio-native) and returns JSON {heard, score, matches, suggestion}, enabling a
generate/fetch → judge → iterate loop. Auto-sources a real sk- OPENAI_API_KEY
from .env (ignores a local lm-studio stub env var). Pads sub-2s clips so the
speech-tuned model doesn't refuse. Caveat: it reliably describes audio and
filters obvious mismatches, but it is NOT a trustworthy judge of subjective qualities
like "grating" — it labels nearly any beep "sharp/high-pitched". Use it to cull, not
to make the final aesthetic call; confirm by ear.
Where this fits
This is the internet-retrieval capability — peer to muser (local semantic search)
and fal (generate). A future media router would fan out across all three and rank
candidates by relevance (CLIP), handing aesthetic spreads to lookdev. Don't build that
router until the model demonstrably mis-routes without it.
Limitations
- Results depend on third-party API availability, quotas, credentials, and license metadata quality.
- License tags and attribution fields must still be reviewed before commercial or public use.
- Relevance ranking can find plausible assets, but final aesthetic fit, brand safety, and audio suitability require human inspection.
1---2name: web-media-getter3description: One query across free image / video / GIF APIs (stock + historical/archival + GIF engines), returning normalized, license-tagged results with optional top-K download + attribution sidecar. The retrieval peer to local semantic search and generative...4license: MIT5---67## When to Use89Use when a task needs a REAL or ARCHIVAL photo / clip (hero, texture, reference, historical footage) or a reaction / animated GIF, rather than a generated one — fan out across free image/video/GIF sources in one query and download license-tagged results.1011_Source: [connerkward/web-media-getter-skill](https://github.com/connerkward/web-media-getter-skill) (MIT)._1213# web-media1415Query many free image/video sources in one fan-out, get a normalized result list,16optionally download top-K with an attribution sidecar. Zero-dep stdlib script.1718**Script:** `webmedia.py` (in this dir). **Keys:** `PEXELS_API_KEY`, `PIXABAY_API_KEY`19in `central/.env` (optional — the 5 no-key sources work without them).2021## Sources2223| Source | Key? | Best for | Media |24|--------|------|----------|-------|25| openverse | none | CC web images (Flickr, museums) | image |26| wikimedia | none | factual / historical / landmark photos | image |27| internetarchive | none | **historical/archival** images + films | image, video |28| loc | none | historical US prints/photos | image |29| nasa | none | space imagery + video | image, video |30| pexels | free key | modern stock photos + **short video clips** | image, video |31| pixabay | free key | modern photos/illustrations + **short clips** | image, video |32| klipy | free key | **GIFs** — recommended (free, unlimited, Tenor drop-in) | gif |33| giphy | free key | **GIFs** — biggest library (prod key needs approval) | gif |3435GIF sources fire only with `--type gif`. Keys: `KLIPY_API_KEY`, `GIPHY_API_KEY`36in `central/.env`. (tenor adapter removed — Google EOL'd the API 2026-06-30.)37**klipy** is the one to get (free + unlimited);38its adapter is **unverified — assumes Tenor-compatible** request/response;39verify against docs.klipy.com when you key it. `webmedia.py "shrug" --type gif --count 6 --json`4041## Usage4243```bash44webmedia.py "1950s street scene" --type image --count 8 --json45webmedia.py "rocket launch" --type video --source nasa,internetarchive46webmedia.py "car factory 1930s" --source all --download --out /tmp/cars47```4849- `--source all` (default) | `nokey` (no-key only) | comma list (`wikimedia,pexels`)50- `--type image|video` · `--count N` · `--json` · `--download --out DIR`51- `--download` fetches each result's direct media URL and writes `attribution.json`52 (source, author, license, url, page_url) alongside the files.5354## Record schema5556`{source, title, url, thumb, dl, page_url, author, license, w, h, type}` —57`dl` is the directly-downloadable media URL (None when only a page exists).5859## The video caveat (important)6061Archival sources (Internet Archive, Europeana, LoC) host **whole films/documentaries**,62not single shots. So:63- **Modern single clip** → `pexels` / `pixabay` (born as short clips, direct MP4). Done.64- **Historical single shot** → retrieve the IA film here, then extract the shot:65 - **Twelve Labs** Marengo search (free 600 min) — pass the IA public MP4 URL, get a66 timestamped moment for "car on assembly line", clip with ffmpeg. Semantic, cheap.67 - or **PySceneDetect** (free, local) to cut the film into shots, then rank keyframes68 with CLIP via the `muser` skill. Fully offline.6970## Audio: freesound + audio QA7172`webmedia.py` is image/video. For **sound effects** (real, CC-licensed) and for73**judging audio** (since Claude can't hear), two sibling scripts live in74`central/scripts/`:7576- **`freesound-fetch.py "<query>" [count] [max_sec] [out_dir]`** — searches freesound.org77 and downloads short hq-mp3 previews. Prints one JSON line per file with78 `license`/`user` for attribution. Key: `FREESOUND_API_KEY` in `central/.env`79 (token-based read; full originals would need OAuth — previews suffice for SFX).80- **`audio-judge.py <file> "<target>"`** — sends the clip to OpenAI `gpt-audio`81 (audio-native) and returns JSON `{heard, score, matches, suggestion}`, enabling a82 generate/fetch → judge → iterate loop. Auto-sources a real `sk-` `OPENAI_API_KEY`83 from `.env` (ignores a local `lm-studio` stub env var). Pads sub-2s clips so the84 speech-tuned model doesn't refuse. **Caveat:** it reliably *describes* audio and85 filters obvious mismatches, but it is NOT a trustworthy judge of subjective qualities86 like "grating" — it labels nearly any beep "sharp/high-pitched". Use it to cull, not87 to make the final aesthetic call; confirm by ear.8889## Where this fits9091This is the **internet-retrieval** capability — peer to `muser` (local semantic search)92and `fal` (generate). A future `media` router would fan out across all three and rank93candidates by relevance (CLIP), handing aesthetic spreads to `lookdev`. Don't build that94router until the model demonstrably mis-routes without it.9596## Limitations9798- Results depend on third-party API availability, quotas, credentials, and license metadata quality.99- License tags and attribution fields must still be reviewed before commercial or public use.100- Relevance ranking can find plausible assets, but final aesthetic fit, brand safety, and audio suitability require human inspection.