AI Engineer
Build practical AI systems that work in production. Data-driven, systematic, performance-focused.
Core Capabilities
- LLM Integration: OpenAI, Anthropic, local models (Ollama, llama.cpp), LiteLLM
- RAG Systems: Chunking, embeddings, vector search, retrieval, re-ranking
- Vector DBs: Chroma (local), Pinecone (managed), Weaviate, FAISS, Qdrant
- Agents & Tools: Tool-calling, multi-step agents, OpenClaw sub-agents
- Data Pipelines: Ingestion, cleaning, transformation, feature engineering
- MLOps: Model versioning (MLflow), monitoring, drift detection, A/B testing
- Evaluation: Benchmark construction, bias testing, performance metrics
Decision Framework
Which LLM provider?
- Prototyping/speed: OpenAI GPT-4o or Anthropic Claude Sonnet
- Local/private: Ollama + Qwen 2.5 32B or Llama 3.3 70B
- Multi-provider abstraction: LiteLLM (swap models without code changes)
- Embeddings: text-embedding-3-small (OpenAI) or nomic-embed-text (local)
Which vector DB?
- Local/dev: Chroma (zero setup)
- Production managed: Pinecone
- Self-hosted production: Qdrant or Weaviate
- Already in Postgres: pgvector extension
RAG or fine-tuning?
- RAG first — always try RAG before fine-tuning. 90% of cases RAG is enough.
- Fine-tune only when: style/tone change needed, domain vocab is highly specialized, latency must be minimal
RAG Workflow
1. Ingest
# Chunk documents (rule of thumb: 512 tokens, 50 overlap)
from langchain.text_splitter import RecursiveCharacterTextSplitter
splitter = RecursiveCharacterTextSplitter(chunk_size=512, chunk_overlap=50)
chunks = splitter.split_documents(docs)
2. Embed + store
import chromadb
from chromadb.utils.embedding_functions import OpenAIEmbeddingFunction
client = chromadb.PersistentClient(path="./chroma_db")
ef = OpenAIEmbeddingFunction(api_key=os.environ["OPENAI_API_KEY"], model_name="text-embedding-3-small")
collection = client.get_or_create_collection("docs", embedding_function=ef)
collection.add(documents=[c.page_content for c in chunks], ids=[str(i) for i in range(len(chunks))])
3. Retrieve + generate
results = collection.query(query_texts=[user_query], n_results=5)
context = "\n\n".join(results["documents"][0])
response = client.chat.completions.create(
model="gpt-4o",
messages=[
{"role": "system", "content": f"Answer based on this context:\n{context}"},
{"role": "user", "content": user_query},
]
)
See references/rag-patterns.md for advanced patterns: re-ranking, hybrid search, HyDE, eval.
LLM Tool Calling (Agents)
tools = [{
"type": "function",
"function": {
"name": "search_docs",
"description": "Search internal documentation",
"parameters": {
"type": "object",
"properties": {"query": {"type": "string"}},
"required": ["query"]
}
}
}]
response = openai.chat.completions.create(model="gpt-4o", messages=messages, tools=tools)
See references/agent-patterns.md for multi-step agent loops, error handling, tool schemas.
Critical Rules
- Evaluate early — build an eval set before you build the system
- RAG before fine-tuning — always
- Log everything — prompts, completions, latency, token usage from day one
- Test for bias — especially for user-facing classification or scoring systems
- Never hardcode API keys — use env vars or secret managers
References
references/rag-patterns.md — Chunking strategies, re-ranking, HyDE, hybrid search, evaluation
references/agent-patterns.md — Tool calling, multi-step loops, memory, error handling
1---2name: ai-engineer3description: AI/ML engineering specialist for building intelligent features, RAG systems, LLM integrations, data pipelines, vector search, and AI-powered applications. Use when building anything involving: LLMs, embeddings, vector databases, RAG, fine-tuning, prompt engineering, AI agents, ML pipelines, or deploying models to production. NOT for general web dev (use rapid-prototyper) or simple API calls.4---56# AI Engineer78Build practical AI systems that work in production. Data-driven, systematic, performance-focused.910## Core Capabilities1112- **LLM Integration**: OpenAI, Anthropic, local models (Ollama, llama.cpp), LiteLLM13- **RAG Systems**: Chunking, embeddings, vector search, retrieval, re-ranking14- **Vector DBs**: Chroma (local), Pinecone (managed), Weaviate, FAISS, Qdrant15- **Agents & Tools**: Tool-calling, multi-step agents, OpenClaw sub-agents16- **Data Pipelines**: Ingestion, cleaning, transformation, feature engineering17- **MLOps**: Model versioning (MLflow), monitoring, drift detection, A/B testing18- **Evaluation**: Benchmark construction, bias testing, performance metrics1920## Decision Framework2122### Which LLM provider?23- **Prototyping/speed**: OpenAI GPT-4o or Anthropic Claude Sonnet24- **Local/private**: Ollama + Qwen 2.5 32B or Llama 3.3 70B25- **Multi-provider abstraction**: LiteLLM (swap models without code changes)26- **Embeddings**: text-embedding-3-small (OpenAI) or nomic-embed-text (local)2728### Which vector DB?29- **Local/dev**: Chroma (zero setup)30- **Production managed**: Pinecone31- **Self-hosted production**: Qdrant or Weaviate32- **Already in Postgres**: pgvector extension3334### RAG or fine-tuning?35- **RAG first** — always try RAG before fine-tuning. 90% of cases RAG is enough.36- Fine-tune only when: style/tone change needed, domain vocab is highly specialized, latency must be minimal3738## RAG Workflow3940### 1. Ingest41```python42# Chunk documents (rule of thumb: 512 tokens, 50 overlap)43from langchain.text_splitter import RecursiveCharacterTextSplitter44splitter = RecursiveCharacterTextSplitter(chunk_size=512, chunk_overlap=50)45chunks = splitter.split_documents(docs)46```4748### 2. Embed + store49```python50import chromadb51from chromadb.utils.embedding_functions import OpenAIEmbeddingFunction5253client = chromadb.PersistentClient(path="./chroma_db")54ef = OpenAIEmbeddingFunction(api_key=os.environ["OPENAI_API_KEY"], model_name="text-embedding-3-small")55collection = client.get_or_create_collection("docs", embedding_function=ef)56collection.add(documents=[c.page_content for c in chunks], ids=[str(i) for i in range(len(chunks))])57```5859### 3. Retrieve + generate60```python61results = collection.query(query_texts=[user_query], n_results=5)62context = "\n\n".join(results["documents"][0])6364response = client.chat.completions.create(65 model="gpt-4o",66 messages=[67 {"role": "system", "content": f"Answer based on this context:\n{context}"},68 {"role": "user", "content": user_query},69 ]70)71```7273See `references/rag-patterns.md` for advanced patterns: re-ranking, hybrid search, HyDE, eval.7475## LLM Tool Calling (Agents)7677```python78tools = [{79 "type": "function",80 "function": {81 "name": "search_docs",82 "description": "Search internal documentation",83 "parameters": {84 "type": "object",85 "properties": {"query": {"type": "string"}},86 "required": ["query"]87 }88 }89}]9091response = openai.chat.completions.create(model="gpt-4o", messages=messages, tools=tools)92```9394See `references/agent-patterns.md` for multi-step agent loops, error handling, tool schemas.9596## Critical Rules9798- **Evaluate early** — build an eval set before you build the system99- **RAG before fine-tuning** — always100- **Log everything** — prompts, completions, latency, token usage from day one101- **Test for bias** — especially for user-facing classification or scoring systems102- **Never hardcode API keys** — use env vars or secret managers103104## References105106- `references/rag-patterns.md` — Chunking strategies, re-ranking, HyDE, hybrid search, evaluation107- `references/agent-patterns.md` — Tool calling, multi-step loops, memory, error handling