Search Engine Setup
Overview
This skill helps AI agents implement production-quality search in applications. It covers index design with custom analyzers, database-to-index sync pipelines, search APIs with faceting and highlights, autocomplete, and relevance tuning based on real query data.
Instructions
Index Design (Elasticsearch)
Map source database columns to Elasticsearch field types:
- Text columns users search →
text with custom analyzer
- Enum/category columns for filtering →
keyword
- Numeric columns for range filters →
integer, float
- Boolean flags →
boolean
- Dates →
date
- Fields for autocomplete →
completion
Custom analyzer template for product/content search:
{
"analyzer": {
"content_analyzer": {
"tokenizer": "standard",
"filter": ["lowercase", "synonym_filter", "edge_ngram_filter"]
}
},
"filter": {
"synonym_filter": { "type": "synonym", "synonyms_path": "synonyms.txt" },
"edge_ngram_filter": { "type": "edge_ngram", "min_gram": 3, "max_gram": 15 }
}
}
Boost fields by search importance: title/name (3-5x), tags (2x), description (1x).
Always add a suggest field of type completion for typeahead.
Index Design (Algolia)
- Set
searchableAttributes in priority order: ["name", "category", "description"].
- Set
attributesForFaceting: prefix filterable attributes with filterOnly() for non-displayed facets.
- Configure
customRanking: ["desc(popularity)", "desc(rating)"].
- Enable typo tolerance (on by default) and set
minWordSizefor1Typo: 3.
Sync Pipeline
- Full re-index: On first run or manual trigger, paginate through all source records (1000 per batch), transform to index documents, bulk insert.
- Incremental sync: Poll
updated_at > last_sync_time every 10 seconds, or use database triggers/CDC.
- Deletions: Track soft-deleted records. Remove from index when detected.
- Idempotency: Use source record ID as document ID. Upsert, never blind insert.
- Error handling: Log failed documents, continue batch. Retry failures in next cycle.
Search API
Build an endpoint that accepts:
q — full-text query string
- Filter params —
category, brand, min_price, max_price, rating, in_stock
sort — relevance (default), price_asc, price_desc, newest, rating
page / per_page or cursor-based pagination
Query construction (Elasticsearch):
{
"query": {
"bool": {
"must": [{ "multi_match": { "query": "q", "fields": ["name^5", "description"], "fuzziness": "AUTO" }}],
"filter": [
{ "term": { "category": "electronics" }},
{ "range": { "price_cents": { "gte": 2000, "lte": 10000 }}},
{ "term": { "in_stock": true }}
],
"should": [{ "term": { "in_stock": { "value": true, "boost": 2 }}}]
}
},
"highlight": { "fields": { "name": {}, "description": {} }},
"aggs": {
"categories": { "terms": { "field": "category", "size": 20 }},
"brands": { "terms": { "field": "brand", "size": 20 }},
"price_ranges": { "range": { "field": "price_cents", "ranges": [
{ "to": 2500 }, { "from": 2500, "to": 10000 }, { "from": 10000 }
]}}
}
}
Autocomplete
- Use completion suggester for prefix-based typeahead (fastest).
- Return top 5 suggestions with category context.
- Add "did you mean" using phrase suggester for low-result queries.
Relevance Tuning
Analyze search logs to improve quality:
- Zero-result queries: Check for misspellings → add synonyms. Check for missing data → flag content gaps.
- Low CTR queries: Top results don't match intent → adjust boost weights or add synonyms.
- Position bias: If users consistently click result #3+, the ranking formula needs tuning.
- Apply changes iteratively: synonyms first, then boost adjustments, then custom scoring.
Examples
Example 1 — Blog search index
Input: "Set up search for a blog with 10K articles."
Output:
{
"mappings": {
"properties": {
"title": { "type": "text", "analyzer": "content_analyzer", "boost": 5.0 },
"body": { "type": "text", "analyzer": "content_analyzer" },
"author": { "type": "keyword" },
"tags": { "type": "keyword" },
"published_at": { "type": "date" },
"suggest": { "type": "completion", "contexts": [{ "name": "tag", "type": "category" }] }
}
}
}
Example 2 — Algolia configuration for an e-commerce store
Input: "Configure Algolia for a store with products."
Output:
index.setSettings({
searchableAttributes: ['name', 'brand', 'category', 'description'],
attributesForFaceting: ['category', 'brand', 'filterOnly(price_cents)', 'rating'],
customRanking: ['desc(sales_count)', 'desc(rating)'],
typoTolerance: true,
minWordSizefor1Typo: 3,
minWordSizefor2Typos: 6,
hitsPerPage: 20,
snippetEllipsisText: '…',
attributesToSnippet: ['description:30'],
});
Guidelines
- Start with Elasticsearch for control, Algolia for speed-to-market. Elasticsearch gives full tuning power; Algolia is faster to set up but costs more at scale.
- Never search the primary database. Always sync to a dedicated search index. SQL
LIKE does not scale.
- Fuzziness AUTO is almost always correct. It allows 1 typo for 3-5 char words and 2 typos for 6+ chars.
- Synonyms are the highest-ROI tuning. Most zero-result queries are fixed by adding 10-20 synonym pairs.
- Monitor query performance. Set an alert if p95 search latency exceeds 200ms.
1---2name: search-engine-setup3description: Search Engine Setup4---5# Search Engine Setup67## Overview89This skill helps AI agents implement production-quality search in applications. It covers index design with custom analyzers, database-to-index sync pipelines, search APIs with faceting and highlights, autocomplete, and relevance tuning based on real query data.1011## Instructions1213### Index Design (Elasticsearch)14151. Map source database columns to Elasticsearch field types:16 - Text columns users search → `text` with custom analyzer17 - Enum/category columns for filtering → `keyword`18 - Numeric columns for range filters → `integer`, `float`19 - Boolean flags → `boolean`20 - Dates → `date`21 - Fields for autocomplete → `completion`22232. Custom analyzer template for product/content search:24 ```json25 {26 "analyzer": {27 "content_analyzer": {28 "tokenizer": "standard",29 "filter": ["lowercase", "synonym_filter", "edge_ngram_filter"]30 }31 },32 "filter": {33 "synonym_filter": { "type": "synonym", "synonyms_path": "synonyms.txt" },34 "edge_ngram_filter": { "type": "edge_ngram", "min_gram": 3, "max_gram": 15 }35 }36 }37 ```38393. Boost fields by search importance: title/name (3-5x), tags (2x), description (1x).40414. Always add a `suggest` field of type `completion` for typeahead.4243### Index Design (Algolia)44451. Set `searchableAttributes` in priority order: `["name", "category", "description"]`.462. Set `attributesForFaceting`: prefix filterable attributes with `filterOnly()` for non-displayed facets.473. Configure `customRanking`: `["desc(popularity)", "desc(rating)"]`.484. Enable typo tolerance (on by default) and set `minWordSizefor1Typo: 3`.4950### Sync Pipeline51521. **Full re-index**: On first run or manual trigger, paginate through all source records (1000 per batch), transform to index documents, bulk insert.532. **Incremental sync**: Poll `updated_at > last_sync_time` every 10 seconds, or use database triggers/CDC.543. **Deletions**: Track soft-deleted records. Remove from index when detected.554. **Idempotency**: Use source record ID as document ID. Upsert, never blind insert.565. **Error handling**: Log failed documents, continue batch. Retry failures in next cycle.5758### Search API5960Build an endpoint that accepts:61- `q` — full-text query string62- Filter params — `category`, `brand`, `min_price`, `max_price`, `rating`, `in_stock`63- `sort` — `relevance` (default), `price_asc`, `price_desc`, `newest`, `rating`64- `page` / `per_page` or cursor-based pagination6566Query construction (Elasticsearch):67```json68{69 "query": {70 "bool": {71 "must": [{ "multi_match": { "query": "q", "fields": ["name^5", "description"], "fuzziness": "AUTO" }}],72 "filter": [73 { "term": { "category": "electronics" }},74 { "range": { "price_cents": { "gte": 2000, "lte": 10000 }}},75 { "term": { "in_stock": true }}76 ],77 "should": [{ "term": { "in_stock": { "value": true, "boost": 2 }}}]78 }79 },80 "highlight": { "fields": { "name": {}, "description": {} }},81 "aggs": {82 "categories": { "terms": { "field": "category", "size": 20 }},83 "brands": { "terms": { "field": "brand", "size": 20 }},84 "price_ranges": { "range": { "field": "price_cents", "ranges": [85 { "to": 2500 }, { "from": 2500, "to": 10000 }, { "from": 10000 }86 ]}}87 }88}89```9091### Autocomplete92931. Use completion suggester for prefix-based typeahead (fastest).942. Return top 5 suggestions with category context.953. Add "did you mean" using phrase suggester for low-result queries.9697### Relevance Tuning9899Analyze search logs to improve quality:1001. **Zero-result queries**: Check for misspellings → add synonyms. Check for missing data → flag content gaps.1012. **Low CTR queries**: Top results don't match intent → adjust boost weights or add synonyms.1023. **Position bias**: If users consistently click result #3+, the ranking formula needs tuning.1034. Apply changes iteratively: synonyms first, then boost adjustments, then custom scoring.104105## Examples106107### Example 1 — Blog search index108109**Input:** "Set up search for a blog with 10K articles."110111**Output:**112```json113{114 "mappings": {115 "properties": {116 "title": { "type": "text", "analyzer": "content_analyzer", "boost": 5.0 },117 "body": { "type": "text", "analyzer": "content_analyzer" },118 "author": { "type": "keyword" },119 "tags": { "type": "keyword" },120 "published_at": { "type": "date" },121 "suggest": { "type": "completion", "contexts": [{ "name": "tag", "type": "category" }] }122 }123 }124}125```126127### Example 2 — Algolia configuration for an e-commerce store128129**Input:** "Configure Algolia for a store with products."130131**Output:**132```js133index.setSettings({134 searchableAttributes: ['name', 'brand', 'category', 'description'],135 attributesForFaceting: ['category', 'brand', 'filterOnly(price_cents)', 'rating'],136 customRanking: ['desc(sales_count)', 'desc(rating)'],137 typoTolerance: true,138 minWordSizefor1Typo: 3,139 minWordSizefor2Typos: 6,140 hitsPerPage: 20,141 snippetEllipsisText: '…',142 attributesToSnippet: ['description:30'],143});144```145146## Guidelines147148- **Start with Elasticsearch for control, Algolia for speed-to-market.** Elasticsearch gives full tuning power; Algolia is faster to set up but costs more at scale.149- **Never search the primary database.** Always sync to a dedicated search index. SQL `LIKE` does not scale.150- **Fuzziness AUTO is almost always correct.** It allows 1 typo for 3-5 char words and 2 typos for 6+ chars.151- **Synonyms are the highest-ROI tuning.** Most zero-result queries are fixed by adding 10-20 synonym pairs.152- **Monitor query performance.** Set an alert if p95 search latency exceeds 200ms.