AI Tech RSS Fetch
Core Goal
- Subscribe to RSS/Atom sources.
- Persist feed and entry metadata to SQLite.
- Deduplicate entries with layered identity keys plus content fingerprints.
- Keep only metadata; do not fetch full article bodies and do not summarize.
Triggering Conditions
- Receive a request to subscribe RSS feeds from URLs or OPML.
- Receive a request to run incremental RSS sync reliably.
- Need stable metadata persistence for downstream processing.
- Need dedupe-safe storage of feed items over repeated runs.
Workflow
- Prepare runtime and database.
- Ensure dependency is installed:
python3 -m pip install feedparser.
- In multi-agent runtimes, pin DB to an absolute path before any command:
export AI_RSS_DB_PATH="/absolute/path/to/workspace-rss-bot/ai_rss.db"
- Initialize SQLite schema once:
python3 scripts/rss_subscribe.py init-db --db "$AI_RSS_DB_PATH"
- Add feed subscriptions.
python3 scripts/rss_subscribe.py add-feed --db "$AI_RSS_DB_PATH" --url "https://example.com/feed.xml"
python3 scripts/rss_subscribe.py import-opml --db "$AI_RSS_DB_PATH" --opml assets/hn-popular-blogs-2025.opml
- Run incremental sync.
- Fetch active feeds and store metadata:
python3 scripts/rss_subscribe.py sync --db "$AI_RSS_DB_PATH" --max-feeds 20 --max-items-per-feed 100
python3 scripts/rss_subscribe.py sync --db "$AI_RSS_DB_PATH" --feed-url "https://example.com/feed.xml"
- Query persisted metadata.
python3 scripts/rss_subscribe.py list-feeds --db "$AI_RSS_DB_PATH" --limit 50
python3 scripts/rss_subscribe.py list-entries --db "$AI_RSS_DB_PATH" --limit 100
Input Requirements
- Supported inputs:
- RSS XML feed URLs.
- OPML feed list files.
Output Contract (Metadata Only)
- Persist
feeds metadata to SQLite:
feed_url, feed_title, site_url, etag, last_modified, status fields.
- Persist
entries metadata to SQLite:
id, dedupe_key (compat primary identity snapshot), guid, url,
canonical_url, title, author, published_at, updated_at, summary,
categories, content_hash, match_confidence, timestamps.
- Persist
entry_identities mapping table to SQLite:
entry_id, key_type, key_value, created_at.
- Supported key types:
guid, canonical_url, legacy_guid, fallback_hash.
- Do not store generated summaries and do not create archive markdown files.
Configurable Parameters
db_path
AI_RSS_DB_PATH (recommended absolute path in multi-agent runtime)
opml_path
feed_urls
max_feeds_per_run
max_items_per_feed
user_agent
seen_ttl_days
enable_conditional_get
- Example config:
assets/config.example.json
Error and Boundary Handling
- Feed HTTP/network failure: keep syncing other feeds and record
last_error.
- Feed
304 Not Modified: skip entry parsing and keep state.
- Missing
guid and link: use hashed fallback identity and set match_confidence=low.
- Dependency missing (
feedparser): return install guidance.
Final Output Checklist (Required)
- core goal
- trigger conditions
- input requirements
- metadata schema
- dedupe and sync rules
- command workflow
- configurable parameters
- error handling
Use the following simplified checklist verbatim when the user requests it:
核心目标
输入需求
触发条件
元数据模型
去重与同步规则
命令流程
可配置参数
错误处理
References
references/input-model.md
references/output-rules.md
references/time-range-rules.md
Assets
assets/hn-popular-blogs-2025.opml (candidate feed pool)
assets/config.example.json
Scripts
1---2name: ai-tech-rss-fetch3description: Subscribe to AI and tech RSS feeds and persist normalized metadata into SQLite using mature Python tooling (feedparser + sqlite3). Use when adding feed URLs/OPML sources, running incremental sync with deduplication, and storing entry metadata without full-text extraction or summarization.4---56# AI Tech RSS Fetch78## Core Goal9- Subscribe to RSS/Atom sources.10- Persist feed and entry metadata to SQLite.11- Deduplicate entries with layered identity keys plus content fingerprints.12- Keep only metadata; do not fetch full article bodies and do not summarize.1314## Triggering Conditions15- Receive a request to subscribe RSS feeds from URLs or OPML.16- Receive a request to run incremental RSS sync reliably.17- Need stable metadata persistence for downstream processing.18- Need dedupe-safe storage of feed items over repeated runs.1920## Workflow211. Prepare runtime and database.22- Ensure dependency is installed: `python3 -m pip install feedparser`.23- In multi-agent runtimes, pin DB to an absolute path before any command:2425```bash26export AI_RSS_DB_PATH="/absolute/path/to/workspace-rss-bot/ai_rss.db"27```2829- Initialize SQLite schema once:3031```bash32python3 scripts/rss_subscribe.py init-db --db "$AI_RSS_DB_PATH"33```34352. Add feed subscriptions.36- Add one feed URL:3738```bash39python3 scripts/rss_subscribe.py add-feed --db "$AI_RSS_DB_PATH" --url "https://example.com/feed.xml"40```4142- Import feeds from OPML:4344```bash45python3 scripts/rss_subscribe.py import-opml --db "$AI_RSS_DB_PATH" --opml assets/hn-popular-blogs-2025.opml46```47483. Run incremental sync.49- Fetch active feeds and store metadata:5051```bash52python3 scripts/rss_subscribe.py sync --db "$AI_RSS_DB_PATH" --max-feeds 20 --max-items-per-feed 10053```5455- Optional one-feed sync:5657```bash58python3 scripts/rss_subscribe.py sync --db "$AI_RSS_DB_PATH" --feed-url "https://example.com/feed.xml"59```60614. Query persisted metadata.62- List feeds:6364```bash65python3 scripts/rss_subscribe.py list-feeds --db "$AI_RSS_DB_PATH" --limit 5066```6768- List recent entries:6970```bash71python3 scripts/rss_subscribe.py list-entries --db "$AI_RSS_DB_PATH" --limit 10072```7374## Input Requirements75- Supported inputs:76 - RSS XML feed URLs.77 - OPML feed list files.7879## Output Contract (Metadata Only)80- Persist `feeds` metadata to SQLite:81 - `feed_url`, `feed_title`, `site_url`, `etag`, `last_modified`, status fields.82- Persist `entries` metadata to SQLite:83 - `id`, `dedupe_key` (compat primary identity snapshot), `guid`, `url`,84 `canonical_url`, `title`, `author`, `published_at`, `updated_at`, `summary`,85 `categories`, `content_hash`, `match_confidence`, timestamps.86- Persist `entry_identities` mapping table to SQLite:87 - `entry_id`, `key_type`, `key_value`, `created_at`.88 - Supported key types: `guid`, `canonical_url`, `legacy_guid`, `fallback_hash`.89- Do not store generated summaries and do not create archive markdown files.9091## Configurable Parameters92- `db_path`93- `AI_RSS_DB_PATH` (recommended absolute path in multi-agent runtime)94- `opml_path`95- `feed_urls`96- `max_feeds_per_run`97- `max_items_per_feed`98- `user_agent`99- `seen_ttl_days`100- `enable_conditional_get`101- Example config: `assets/config.example.json`102103## Error and Boundary Handling104- Feed HTTP/network failure: keep syncing other feeds and record `last_error`.105- Feed `304 Not Modified`: skip entry parsing and keep state.106- Missing `guid` and `link`: use hashed fallback identity and set `match_confidence=low`.107- Dependency missing (`feedparser`): return install guidance.108109## Final Output Checklist (Required)110- core goal111- trigger conditions112- input requirements113- metadata schema114- dedupe and sync rules115- command workflow116- configurable parameters117- error handling118119Use the following simplified checklist verbatim when the user requests it:120121```text122核心目标123输入需求124触发条件125元数据模型126去重与同步规则127命令流程128可配置参数129错误处理130```131132## References133- `references/input-model.md`134- `references/output-rules.md`135- `references/time-range-rules.md`136137## Assets138- `assets/hn-popular-blogs-2025.opml` (candidate feed pool)139- `assets/config.example.json`140141## Scripts142- `scripts/rss_subscribe.py`