ChatQnA Helm Deploy
Deploy the Chat Question and Answer Core sample application Helm chart at sample-applications/chat-question-and-answer-core/chart/ to Kubernetes using
Helm. The chart's dependencies are chatqna-core and chatqna-ui, which are built from the same source code as the Docker Compose deployment. Also it includes nginx as a reverse proxy for the backend and UI.
Environment setup (run first)
This skill operates on real ChatQnA source files, so the ChatQnA application must be present and commands must run from the app root. Do this before any Helm workflow, whether or not source is already in your workspace.
Run the bundled bootstrap. It searches for an existing ChatQnA checkout by
walking up from the current directory and checking the enclosing git repo, then
reuses it without re-cloning. Only when no checkout is found does it do a
shallow, single-branch, sparse checkout of just
sample-applications/chat-question-and-answer-core from main.
It prints the resolved app root on stdout:
# SKILL_DIR is this skill directory. In-repo it is:
# .github/skills/chatqna-helm-deploy
SKILL_DIR=".github/skills/chatqna-helm-deploy"
APP_ROOT="$(bash "$SKILL_DIR/scripts/chatqna-bootstrap.sh")"
cd "$APP_ROOT"
Every command below assumes the working directory is this APP_ROOT.
To use a fork/branch or a specific clone path, override these before running the bootstrap script:
CHATQNA_REPO_URLCHATQNA_REPO_BRANCHCHATQNA_CLONE_DIRCHATQNA_FORCE_CLONE(set to1to force clone)
Codebase root: sample-applications/chat-question-and-answer-core/
Prerequisites
Confirm a reachable Kubernetes cluster is available and
kubectlis configured to access it.kubectl get nodesFor GPU, discover resource keys before writing values:
kubectl get nodes -o json | jq -r '.items[] | "\(.metadata.name):\n" + (.status.allocatable | to_entries | map(select(.key | test("gpu|npu|vpu|accel";"i"))) | map(" \(.key): \(.value)") | join("\n"))'Common Intel keys are
gpu.intel.com/i915,gpu.intel.com/xe
What This Skill Produces
- A running ChatQnA Core Helm release in a target namespace for one runtime:
- OpenVINO CPU
- OpenVINO GPU
- Ollama
- A generated override file (
values-override.yaml) that translates Docker Compose andsetup_env.shstyle inputs to Helm values keys. - A verified deployment state using pods, services, and health endpoint checks.
- A concise deployment report containing:
- runtime selected and GPU mode
- chart source (local path or OCI chart)
- values files used and major override keys
- access URL and API docs URL
- warnings (missing token, GPU key, model constraints)
When to Use
- "Deploy chatqna core to Kubernetes"
- "Helm install chatqna-core"
- "Configure values.yaml for chatqna core"
- "Translate docker compose setup_env.sh into helm values"
- "Deploy OpenVINO GPU profile with Helm"
- "Deploy Ollama with chart values"
Inputs To Confirm
Before running commands, confirm or infer these values:
- Runtime:
openvinoorollama - Device target:
cpuorgpu(GPU valid only for OpenVINO) - Namespace and release name (default release:
chatqna-core) - Chart source:
- local chart path (
./chart), or - OCI chart (
oci://registry-1.docker.io/intel/chat-question-and-answer-core)
- local chart path (
- Image source and tags:
- prebuilt registry tags, or
- custom/private registry and tags
- Model settings (
EMBEDDING_MODEL,LLM_MODEL, optionalRERANKER_MODEL) - Optional Hugging Face token (
HUGGINGFACEHUB_API_TOKEN) for OpenVINO - Optional proxy values (
http_proxy,https_proxy,no_proxy)
If runtime/device values are missing, default to openvino + cpu.
If prebuilt images are used and tags are not specified by the user, default to
the tags in chart/values.yaml.
Use Helm and kubectl commands for deployment actions in this skill.
Decision Logic
- If runtime is
ollama:- select
-f values.yaml -f values-ollama.yaml - force CPU-only devices
- select
- If runtime is
openvinoand device isgpu:- select
-f values.yaml -f values-openvino.yaml - set
gpu.enabled=true - require
gpu.keyfrom cluster labels
- select
- If runtime is
openvinoand device iscpu:- select
-f values.yaml -f values-openvino.yaml - set
gpu.enabled=false
- select
- If requested values conflict with chart validation (for example GPU model
device with
gpu.enabled=false), correct values before install.
Compose and setup_env.sh Translation
Use the reference mapping in
./references/compose-setupenv-to-helm-mapping.md if the user asks to map Compose or setup_env.sh inputs to Helm values.
Deployment Workflow
Run from sample-applications/chat-question-and-answer-core.
1. Preflight
kubectl version --client
helm version
kubectl config current-context
If using local source chart:
cd chart
helm dependency build
If using OCI chart:
helm pull oci://registry-1.docker.io/intel/chat-question-and-answer-core --version <version>
tar -xvf chat-question-and-answer-core-<version>.tgz
cd chat-question-and-answer-core
helm dependency build
Ensure namespace exists:
kubectl create namespace <namespace> --dry-run=client -o yaml | kubectl apply -f -
2. Build values override file
Create or update values-override.yaml by translating user intent or
Compose/setup_env style inputs using the reference mapping. Do not commit filled secrets or tokens.
If running behind a proxy, include these keys in values-override.yaml using
the values from your current system environment:
global:
http_proxy: "${http_proxy}"
https_proxy: "${https_proxy}"
no_proxy: "${no_proxy}"
Select base files by runtime:
- OpenVINO:
values.yaml+values-openvino.yaml+values-override.yaml - Ollama:
values.yaml+values-ollama.yaml+values-override.yaml
3. Validate rendered manifests
helm template chatqna-core \
-f values.yaml \
-f values-<runtime>.yaml \
-f values-override.yaml \
.
4. Install or upgrade release
helm upgrade --install chatqna-core \
-f values.yaml \
-f values-<runtime>.yaml \
-f values-override.yaml \
. \
--namespace <namespace>
5. Verify deployment
kubectl get pods -n <namespace>
kubectl get services -n <namespace>
kubectl get events -n <namespace> --sort-by=.lastTimestamp | tail -n 30
kubectl rollout status deploy/chatqna-core -n <namespace>
kubectl rollout status deploy/chatqna-core-nginx -n <namespace>
Health endpoint evidence:
chatqna_hostip=$(kubectl get pods -l app=chatqna-core-nginx -n <namespace> -o jsonpath='{.items[0].status.hostIP}')
chatqna_port=$(kubectl get service chatqna-core-nginx -n <namespace> -o jsonpath='{.spec.ports[0].nodePort}')
curl -sS -w "\nHTTP_STATUS:%{http_code}\n" "http://${chatqna_hostip}:${chatqna_port}/v1/chatqna/health"
6. Access and teardown
# UI
echo "http://${chatqna_hostip}:${chatqna_port}"
# API docs
echo "http://${chatqna_hostip}:${chatqna_port}/v1/chatqna/docs"
# Uninstall
helm uninstall chatqna-core -n <namespace>
Failure Handling
- Helm template validation fails:
- report exact key causing failure and propose corrected key/value.
- GPU requested but
gpu.keymissing:- instruct user to run
kubectl describe nodeand provide device plugin key, then re-run withgpu.enabled=true.
- instruct user to run
- Pods not ready:
- collect
kubectl describe podandkubectl logsfor failing pods.
- collect
- Health check non-200:
- inspect
chatqna-corelogs for model download/config issues. - note first startup can take longer due to model pull/conversion.
- inspect
- PVC stuck:
- list and optionally delete stuck PVC only when explicitly requested.
- Need larger storage:
- increase PVC size in
values-override.yamland re-runhelm upgrade.
- increase PVC size in
Completion Criteria
- Runtime-specific install command is executed with correct values files.
- Compose/setup_env inputs (if provided) are translated into a concrete
values-override.yaml. - Pods/services are healthy in the target namespace.
- Health endpoint returns
HTTP_STATUS:200. - User receives UI URL, API docs URL, release/namespace, and uninstall command.
- Response includes raw verification evidence (
kubectl get, rollout status, health check output).