Scraper Skill
Work with PriceTrack's price scrapers — add new sites, debug existing ones, or modify scraping logic.
Architecture
Scrapers are split into two parts:
Config (
functions/config/<site>.js) — Site-specific selectors and URL patterns. Exports an object with:urlPattern— regex to match product URLsselectors— CSS selectors for price, title, image, etc.transform— optional data transformation functions
Module (
functions/modules/pullData.js) — Generic pull logic that reads configs and fetches/scrapes product pages using:- Puppeteer-core + @sparticuz/chromium for JS-rendered pages (serverless Chromium)
- jsdom for static HTML parsing
- axios for simple HTTP fetches
Adding a New Scraper
- Create
functions/config/<site>.jswith the site's URL pattern and selectors - Register it in the config index (check how existing sites are imported)
- Test with:
firebase functions:shell → pullData({url: '<product-url>'}) - Verify the scraped data structure matches what
modules/pullData.jsexpects
Debugging
- Check Firebase Functions logs:
firebase functions:log - Puppeteer scrapers run headless on serverless Chromium — no GPU, limited memory
- Some sites require specific headers or cookie consent handling
@sparticuz/chromiumhas size limits; check if the binary is up to date
Key Files
functions/config/— per-site scraper configurationsfunctions/modules/pullData.js— main scraping orchestrationfunctions/modules/updateInfo.js— product info updatesfunctions/utils/fetch.js— HTTP fetch utilities
Source: duyet/pricetrack — distributed by TomeVault.