Level AI Platform Help
Step 1 — Gather context
If references/learnings.md exists, read it first for accumulated platform knowledge.
What do you need help with?
- A) Setting up InstaScore QA scorecards and evaluation criteria
- B) Configuring Real-Time Agent Assist (AgentGPT) for live guidance
- C) Voice of Customer analytics — sentiment, iCSAT, trends
- D) AI Virtual Agent setup (voice/chat automation)
- E) CCaaS integration (Five9, Talkdesk, Zendesk, Genesys, etc.)
- F) API integration — pushing/pulling interaction data
- G) Comparing Level AI to another tool (Observe.AI, Cresta, CallMiner, Balto)
- H) Compliance monitoring (PCI, HIPAA, script adherence)
- I) Screen recording and agent coaching workflows
- J) Other
What's your current setup?
- A) Evaluating whether to buy
- B) New — haven't started implementation
- C) In implementation
- D) Running but having issues
- E) Expanding to new modules
What's your CCaaS/telephony?
- A) Five9
- B) Talkdesk
- C) Zendesk
- D) Genesys
- E) Amazon Connect
- F) Twilio
- G) Freshworks
- H) Kustomer
- I) Other
Contact center size?
- A) Mid-size (50-200 agents)
- B) Large (200-1,000 agents)
- C) Enterprise (1,000+ agents)
Skip-ahead rule: if the user's prompt already contains enough context, skip to Step 2.
Step 2 — Route or answer directly
| Problem domain |
Route to |
| Building a coaching program or training cadence |
/sales-coaching {user's question} |
| Reviewing a specific call transcript for coaching |
/sales-call-review {user's question} |
| Choosing between note-taker/conversation intelligence platforms |
/sales-note-taker {user's question} |
| CCaaS platform comparison/selection |
/sales-ccaas-selection {user's question} |
| General CRM/tool integration patterns (Zapier, webhooks) |
/sales-integration {user's question} |
Otherwise, answer directly using the platform reference below.
Step 3 — Level AI platform reference
Read references/platform-guide.md for the full platform reference — modules, pricing, integrations, data model, workflows.
Answer the user's question using only the relevant section. Don't dump the full reference.
Step 4 — Actionable guidance
You no longer need the platform guide — focus on the user's specific situation.
Implementation priority order:
- Connect your CCaaS first — call data must flow before anything else works
- Configure InstaScore with a starter scorecard (5-8 criteria) — validate transcription accuracy on 50+ calls before trusting scores
- Calibrate QA — have human reviewers score the same 20-30 calls and compare to InstaScore. Adjust criteria until alignment is acceptable
- Roll out coaching workflows — use screen recordings and QA findings to identify coaching priorities
- Enable AgentGPT for real-time guidance once post-call QA is stable
- AI Virtual Agent last — requires the most tuning and governance
When comparing to competitors:
- vs Observe.AI: Similar scope. Level AI differentiates on semantic intelligence (intent-based, not keyword). Observe.AI has more mature VoiceAI/ChatAI agents and 250+ integrations. Both ~$150-200/agent/mo.
- vs Cresta: Cresta targets larger enterprises with Knowledge Agent and deeper virtual agent capabilities. Level AI is more accessible for mid-market.
- vs CallMiner: CallMiner has deeper post-call analytics and compliance for regulated industries. Level AI has stronger real-time agent assist. CallMiner averages ~$102K/yr.
- vs Balto: Balto deploys in weeks and leads in real-time guidance. Level AI combines real-time with post-call QA in one platform. Balto is ~$100-150/agent/mo.
If you discover a gotcha, workaround, or tip not covered in references/learnings.md, append it there.
Gotchas
Best-effort from research — review these, especially items about plan-gated features and integration gotchas that may be outdated.
- Call ingestion can be delayed 24+ hours. G2 reviewers report calls not appearing until the next day, making same-day QA monitoring unreliable. Ask about ingestion SLAs during evaluation.
- InstaScore accuracy requires calibration. AI QA scores won't match human QA out of the box — scoring issues tied to language barriers and subjective criteria. Start with binary/objective criteria and run a calibration period against manual scores.
- Sentiment analysis misclassifies neutral conversations. Users report neutral conversations flagged as negative. Don't rely on sentiment alone for coaching decisions — cross-reference with QA scores.
- No public pricing. Estimated ~$185/agent/month. Custom quotes required — pricing varies by agent count, modules, and contract length.
- API exists but is not publicly documented. GraphQL-based per their engineering blog. Enterprise customers get access — ask your account team for the API spec during evaluation.
- Mid-market focus may limit scalability for very large operations. The platform excels at 100-1,000 agent contact centers. For 5,000+ agents, also evaluate Observe.AI, Cresta, or NICE CXone.
- Self-improving: If you discover something not covered here, append it to
references/learnings.md with today's date.
Related skills
/sales-call-review — Review specific sales calls and extract coaching insights
/sales-coaching — Build coaching programs, onboarding, role-plays, certifications
/sales-note-taker — Compare AI note-takers and conversation intelligence tools or wire APIs into CRM
/sales-observe-ai — Observe.AI platform help (enterprise contact center QA, deepest AI agent capabilities)
/sales-cresta — Cresta platform help (enterprise contact center AI, Fortune 500 focus)
/sales-callminer — CallMiner platform help (enterprise post-call analytics, strongest compliance)
/sales-balto — Balto platform help (real-time agent guidance, fastest deployment)
/sales-enthu — Enthu.AI platform help (affordable contact center QA for smaller teams)
/sales-ccaas-selection — Compare CCaaS platforms (Genesys, NICE, Talkdesk, Five9, etc.)
/sales-do — Not sure which skill to use? The router matches any sales objective to the right skill. Install: npx skills add sales-skills/sales --skill sales-do -a claude-code -y
Examples
Example 1: Setting up automated QA
User says: "We manually review 3% of calls. How do I set up InstaScore in Level AI to auto-score everything?"
Skill does:
- Reads platform guide for InstaScore module
- Explains 100% auto-scoring with custom scorecards — define criteria, assign weights, auto-evaluate all interactions
- Walks through scorecard creation with binary criteria (compliance, verification) and scaled criteria (empathy, resolution)
- Recommends a 2-week calibration period comparing InstaScore to manual reviews
Result: User has a plan to move from 3% manual sampling to 100% automated QA
Example 2: Comparing Level AI vs Observe.AI
User says: "We're evaluating Level AI and Observe.AI for our 300-agent insurance contact center"
Skill does:
- Compares both platforms on QA, real-time assist, compliance, integrations, and pricing
- Notes Level AI's semantic intelligence vs Observe.AI's post-call analytics depth
- Highlights Observe.AI's more mature VoiceAI agents vs Level AI's stronger mid-market positioning
- Recommends piloting both with 20-30 agents over 30 days
Result: Side-by-side comparison tailored to insurance contact center requirements
Troubleshooting
InstaScore QA scores don't match manual evaluations
Symptom: Auto-scores diverge significantly from human QA reviewer scores
Cause: Scorecard criteria may be too subjective, transcription errors affecting scoring, or language barriers causing misclassification
Solution: Make criteria binary where possible ("Did agent verify identity?" not "Was the opening professional?"). Run calibration on 30 calls with human reviewers. Check transcription quality — poor transcription = unreliable scoring regardless of criteria.
Calls not appearing or delayed ingestion
Symptom: Calls from today aren't visible in the platform, 24+ hour delay
Cause: CCaaS integration pipeline processing lag — known issue in G2 reviews
Solution: Check CCaaS integration status in admin panel. Verify call recordings are flowing from your telephony system. Contact Level AI support if delays exceed 24 hours consistently. For same-day monitoring needs, evaluate Balto for real-time visibility.
AgentGPT not surfacing relevant guidance
Symptom: Real-time assist shows generic or irrelevant knowledge articles during calls
Cause: Knowledge base not properly configured or not mapped to the right call contexts
Solution: Review your knowledge base content in Level AI — articles need clear topic tagging. Test with specific call scenarios. AgentGPT uses semantic matching, so ensure knowledge articles use language similar to how agents and customers discuss topics.
1---2name: sales-level-ai3description: Level AI platform help — contact center intelligence with Naviant semantic AI, InstaScore 100% automated QA, Real-Time Agent Assist (AgentGPT), Voice of Customer analytics (iCSAT), AI Virtual Agent, screen recording, omnichannel analysis. Use when setting up Level AI InstaScore QA scorecards for contact center agents, AgentGPT real-time assist not surfacing correct knowledge articles during calls, QA scores seem inaccurate or don't match manual evaluations, Level AI call ingestion delayed more than 24 hours, comparing Level AI vs Observe.AI or Cresta or CallMiner for contact center QA, integrating Level AI with Five9 or Talkdesk or Amazon Connect or Zendesk, configuring compliance monitoring for PCI or HIPAA, or evaluating Level AI for a mid-market contact center. Do NOT use for building a general coaching program (use /sales-coaching), reviewing a specific call transcript (use /sales-call-review), or CCaaS platform selection (use /sales-ccaas-selection).4license: MIT5---6
7# Level AI Platform Help
8
9## Step 1 — Gather context
10
11If `references/learnings.md` exists, read it first for accumulated platform knowledge.
12
131. **What do you need help with?**
14 - A) Setting up InstaScore QA scorecards and evaluation criteria
15 - B) Configuring Real-Time Agent Assist (AgentGPT) for live guidance
16 - C) Voice of Customer analytics — sentiment, iCSAT, trends
17 - D) AI Virtual Agent setup (voice/chat automation)
18 - E) CCaaS integration (Five9, Talkdesk, Zendesk, Genesys, etc.)
19 - F) API integration — pushing/pulling interaction data
20 - G) Comparing Level AI to another tool (Observe.AI, Cresta, CallMiner, Balto)
21 - H) Compliance monitoring (PCI, HIPAA, script adherence)
22 - I) Screen recording and agent coaching workflows
23 - J) Other
24
252. **What's your current setup?**
26 - A) Evaluating whether to buy
27 - B) New — haven't started implementation
28 - C) In implementation
29 - D) Running but having issues
30 - E) Expanding to new modules
31
323. **What's your CCaaS/telephony?**
33 - A) Five9
34 - B) Talkdesk
35 - C) Zendesk
36 - D) Genesys
37 - E) Amazon Connect
38 - F) Twilio
39 - G) Freshworks
40 - H) Kustomer
41 - I) Other
42
434. **Contact center size?**
44 - A) Mid-size (50-200 agents)
45 - B) Large (200-1,000 agents)
46 - C) Enterprise (1,000+ agents)
47
48Skip-ahead rule: if the user's prompt already contains enough context, skip to Step 2.
49
50## Step 2 — Route or answer directly
51
52| Problem domain | Route to |
53|---|---|
54| Building a coaching program or training cadence | `/sales-coaching {user's question}` |
55| Reviewing a specific call transcript for coaching | `/sales-call-review {user's question}` |
56| Choosing between note-taker/conversation intelligence platforms | `/sales-note-taker {user's question}` |
57| CCaaS platform comparison/selection | `/sales-ccaas-selection {user's question}` |
58| General CRM/tool integration patterns (Zapier, webhooks) | `/sales-integration {user's question}` |
59
60Otherwise, answer directly using the platform reference below.
61
62## Step 3 — Level AI platform reference
63
64**Read `references/platform-guide.md`** for the full platform reference — modules, pricing, integrations, data model, workflows.
65
66Answer the user's question using only the relevant section. Don't dump the full reference.
67
68## Step 4 — Actionable guidance
69
70You no longer need the platform guide — focus on the user's specific situation.
71
72**Implementation priority order:**
731. Connect your CCaaS first — call data must flow before anything else works
742. Configure InstaScore with a starter scorecard (5-8 criteria) — validate transcription accuracy on 50+ calls before trusting scores
753. Calibrate QA — have human reviewers score the same 20-30 calls and compare to InstaScore. Adjust criteria until alignment is acceptable
764. Roll out coaching workflows — use screen recordings and QA findings to identify coaching priorities
775. Enable AgentGPT for real-time guidance once post-call QA is stable
786. AI Virtual Agent last — requires the most tuning and governance
79
80**When comparing to competitors:**
81- vs **Observe.AI**: Similar scope. Level AI differentiates on semantic intelligence (intent-based, not keyword). Observe.AI has more mature VoiceAI/ChatAI agents and 250+ integrations. Both ~$150-200/agent/mo.
82- vs **Cresta**: Cresta targets larger enterprises with Knowledge Agent and deeper virtual agent capabilities. Level AI is more accessible for mid-market.
83- vs **CallMiner**: CallMiner has deeper post-call analytics and compliance for regulated industries. Level AI has stronger real-time agent assist. CallMiner averages ~$102K/yr.
84- vs **Balto**: Balto deploys in weeks and leads in real-time guidance. Level AI combines real-time with post-call QA in one platform. Balto is ~$100-150/agent/mo.
85
86If you discover a gotcha, workaround, or tip not covered in `references/learnings.md`, append it there.
87
88## Gotchas
89
90> *Best-effort from research — review these, especially items about plan-gated features and integration gotchas that may be outdated.*
91
92- **Call ingestion can be delayed 24+ hours.** G2 reviewers report calls not appearing until the next day, making same-day QA monitoring unreliable. Ask about ingestion SLAs during evaluation.
93- **InstaScore accuracy requires calibration.** AI QA scores won't match human QA out of the box — scoring issues tied to language barriers and subjective criteria. Start with binary/objective criteria and run a calibration period against manual scores.
94- **Sentiment analysis misclassifies neutral conversations.** Users report neutral conversations flagged as negative. Don't rely on sentiment alone for coaching decisions — cross-reference with QA scores.
95- **No public pricing.** Estimated ~$185/agent/month. Custom quotes required — pricing varies by agent count, modules, and contract length.
96- **API exists but is not publicly documented.** GraphQL-based per their engineering blog. Enterprise customers get access — ask your account team for the API spec during evaluation.
97- **Mid-market focus may limit scalability for very large operations.** The platform excels at 100-1,000 agent contact centers. For 5,000+ agents, also evaluate Observe.AI, Cresta, or NICE CXone.
98- **Self-improving**: If you discover something not covered here, append it to `references/learnings.md` with today's date.
99
100## Related skills
101
102- `/sales-call-review` — Review specific sales calls and extract coaching insights
103- `/sales-coaching` — Build coaching programs, onboarding, role-plays, certifications
104- `/sales-note-taker` — Compare AI note-takers and conversation intelligence tools or wire APIs into CRM
105- `/sales-observe-ai` — Observe.AI platform help (enterprise contact center QA, deepest AI agent capabilities)
106- `/sales-cresta` — Cresta platform help (enterprise contact center AI, Fortune 500 focus)
107- `/sales-callminer` — CallMiner platform help (enterprise post-call analytics, strongest compliance)
108- `/sales-balto` — Balto platform help (real-time agent guidance, fastest deployment)
109- `/sales-enthu` — Enthu.AI platform help (affordable contact center QA for smaller teams)
110- `/sales-ccaas-selection` — Compare CCaaS platforms (Genesys, NICE, Talkdesk, Five9, etc.)
111- `/sales-do` — Not sure which skill to use? The router matches any sales objective to the right skill. Install: `npx skills add sales-skills/sales --skill sales-do -a claude-code -y`
112
113## Examples
114
115### Example 1: Setting up automated QA
116**User says**: "We manually review 3% of calls. How do I set up InstaScore in Level AI to auto-score everything?"
117**Skill does**:
1181. Reads platform guide for InstaScore module
1192. Explains 100% auto-scoring with custom scorecards — define criteria, assign weights, auto-evaluate all interactions
1203. Walks through scorecard creation with binary criteria (compliance, verification) and scaled criteria (empathy, resolution)
1214. Recommends a 2-week calibration period comparing InstaScore to manual reviews
122**Result**: User has a plan to move from 3% manual sampling to 100% automated QA
123
124### Example 2: Comparing Level AI vs Observe.AI
125**User says**: "We're evaluating Level AI and Observe.AI for our 300-agent insurance contact center"
126**Skill does**:
1271. Compares both platforms on QA, real-time assist, compliance, integrations, and pricing
1282. Notes Level AI's semantic intelligence vs Observe.AI's post-call analytics depth
1293. Highlights Observe.AI's more mature VoiceAI agents vs Level AI's stronger mid-market positioning
1304. Recommends piloting both with 20-30 agents over 30 days
131**Result**: Side-by-side comparison tailored to insurance contact center requirements
132
133## Troubleshooting
134
135### InstaScore QA scores don't match manual evaluations
136**Symptom**: Auto-scores diverge significantly from human QA reviewer scores
137**Cause**: Scorecard criteria may be too subjective, transcription errors affecting scoring, or language barriers causing misclassification
138**Solution**: Make criteria binary where possible ("Did agent verify identity?" not "Was the opening professional?"). Run calibration on 30 calls with human reviewers. Check transcription quality — poor transcription = unreliable scoring regardless of criteria.
139
140### Calls not appearing or delayed ingestion
141**Symptom**: Calls from today aren't visible in the platform, 24+ hour delay
142**Cause**: CCaaS integration pipeline processing lag — known issue in G2 reviews
143**Solution**: Check CCaaS integration status in admin panel. Verify call recordings are flowing from your telephony system. Contact Level AI support if delays exceed 24 hours consistently. For same-day monitoring needs, evaluate Balto for real-time visibility.
144
145### AgentGPT not surfacing relevant guidance
146**Symptom**: Real-time assist shows generic or irrelevant knowledge articles during calls
147**Cause**: Knowledge base not properly configured or not mapped to the right call contexts
148**Solution**: Review your knowledge base content in Level AI — articles need clear topic tagging. Test with specific call scenarios. AgentGPT uses semantic matching, so ensure knowledge articles use language similar to how agents and customers discuss topics.