Smart Model Switching
Three-tier z.ai (GLM) routing: Flash → Standard → Plus / 32B
Start with the cheapest model. Escalate only when needed. Designed to minimize API cost without sacrificing correctness.
The Golden Rule
If a human would need more than 30 seconds of focused thinking, escalate from Flash to Standard.
If the task involves architecture, complex tradeoffs, or deep reasoning, escalate to Plus / 32B.
Model Reality (Relative)
| Tier |
Example Models |
Purpose |
| Flash |
GLM-4.5-Flash, GLM-4.7-Flash |
Fastest & cheapest |
| Standard |
GLM-4.6, GLM-4.7 |
Strong reasoning & code |
| Plus / 32B |
GLM-4-Plus, GLM-4-32B-128K |
Heavy reasoning & architecture |
Bottom line: Wrong model selection wastes money OR time. Flash for simple, Standard for normal work, Plus/32B for complex decisions.
💚 FLASH — Default for Simple Tasks
Stay on Flash for:
- Factual Q&A — “what is X”, “who is Y”, “when did Z”
- Quick lookups — definitions, unit conversions, short translations
- Status checks — monitoring, file reads, session state
- Heartbeats — periodic checks, OK responses
- Memory & reminders
- Casual conversation — greetings, acknowledgments
- Simple file ops — read, list, basic writes
- One-liner tasks — anything answerable in 1–2 sentences
- Cron jobs (always Flash by default)
NEVER do these on Flash
- ❌ Write code longer than 10 lines
- ❌ Create comparison tables
- ❌ Write more than 3 paragraphs
- ❌ Do multi-step analysis
- ❌ Write reports or proposals
💛 STANDARD — Core Workhorse
Escalate to Standard for:
Code & Technical
- Code generation — functions, scripts, features
- Debugging — normal bug investigation
- Code review — PRs, refactors
- Documentation — README, comments, guides
Analysis & Planning
- Comparisons and evaluations
- Planning — roadmaps, task breakdowns
- Research synthesis
- Multi-step reasoning
Writing & Content
- Long-form writing (>3 paragraphs)
- Summaries of long documents
- Structured output — tables, outlines
Most real user conversations belong here.
❤️ PLUS / 32B — Complex Reasoning Only
Escalate to Plus / 32B for:
Architecture & Design
- System and service architecture
- Database schema design
- Distributed or multi-tenant systems
- Major refactors across multiple files
Deep Analysis
- Complex debugging (race conditions, subtle bugs)
- Security reviews
- Performance optimization strategy
- Root cause analysis
Strategic & Judgment-Based Work
- Strategic planning
- Nuanced judgment and ambiguity
- Deep or multi-source research
- Critical production decisions
🔄 Implementation
For Subagents
// Routine monitoring
sessions_spawn(task="Check backup status", model="GLM-4.5-Flash")
// Standard code work
sessions_spawn(task="Build the REST API endpoint", model="GLM-4.7")
// Architecture decisions
sessions_spawn(task="Design the database schema for multi-tenancy", model="GLM-4-Plus")
For Cron Jobs
json
Copy code
{
"payload": {
"kind": "agentTurn",
"model": "GLM-4.5-Flash"
}
}
Always use Flash for cron unless the task genuinely needs reasoning.
📊 Quick Decision Tree
pgsql
Copy code
Is it a greeting, lookup, status check, or 1–2 sentence answer?
YES → FLASH
NO ↓
Is it code, analysis, planning, writing, or multi-step?
YES → STANDARD
NO ↓
Is it architecture, deep reasoning, or a critical decision?
YES → PLUS / 32B
NO → Default to STANDARD, escalate if struggling
📋 Quick Reference Card
less
Copy code
┌─────────────────────────────────────────────────────────────┐
│ SMART MODEL SWITCHING │
│ Flash → Standard → Plus / 32B │
├─────────────────────────────────────────────────────────────┤
│ 💚 FLASH (cheapest) │
│ • Greetings, status checks, quick lookups │
│ • Factual Q&A, reminders │
│ • Simple file ops, 1–2 sentence answers │
├─────────────────────────────────────────────────────────────┤
│ 💛 STANDARD (workhorse) │
│ • Code > 10 lines, debugging │
│ • Analysis, comparisons, planning │
│ • Reports, long writing │
├─────────────────────────────────────────────────────────────┤
│ ❤️ PLUS / 32B (complex) │
│ • Architecture decisions │
│ • Complex debugging, multi-file refactoring │
│ • Strategic planning, deep research │
├─────────────────────────────────────────────────────────────┤
│ 💡 RULE: >30 sec human thinking → escalate │
│ 💰 START CHEAP → SCALE ONLY WHEN NEEDED │
└─────────────────────────────────────────────────────────────┘
Built for z.ai (GLM) setups.
1---2name: smart-model-switching-glm3description: Auto-route tasks to the cheapest z.ai (GLM) model that works correctly. Three-tier progression: Flash → Standard → Plus/32B. Classify before responding. FLASH (default): factual Q&A, greetings, reminders, status checks, lookups, simple file ops, heartbeats, casual chat, 1–2 sentence tasks, cron jobs. ESCALATE TO STANDARD: code >10 lines, analysis, comparisons, planning, reports, multi-step reasoning, tables, long writing >3 paragraphs, summarization, research synthesis, most user conversations. ESCALATE TO PLUS/32B: architecture decisions, complex debugging, multi-file refactoring, strategic planning, nuanced judgment, deep research, critical production decisions. Rule: If a human needs >30 seconds of focused thinking, escalate. If Standard struggles with complexity, go to Plus/32B. Save major API costs by starting cheap and escalating only when needed.4---5
6# Smart Model Switching
7
8**Three-tier z.ai (GLM) routing: Flash → Standard → Plus / 32B**
9
10Start with the cheapest model. Escalate only when needed. Designed to minimize API cost without sacrificing correctness.
11
12---
13
14## The Golden Rule
15
16> If a human would need more than 30 seconds of focused thinking, escalate from Flash to Standard.
17> If the task involves architecture, complex tradeoffs, or deep reasoning, escalate to Plus / 32B.
18
19---
20
21## Model Reality (Relative)
22
23| Tier | Example Models | Purpose |
24|-----|----------------|---------|
25| Flash | GLM-4.5-Flash, GLM-4.7-Flash | Fastest & cheapest |
26| Standard | GLM-4.6, GLM-4.7 | Strong reasoning & code |
27| Plus / 32B | GLM-4-Plus, GLM-4-32B-128K | Heavy reasoning & architecture |
28
29**Bottom line:** Wrong model selection wastes money OR time. Flash for simple, Standard for normal work, Plus/32B for complex decisions.
30
31---
32
33## 💚 FLASH — Default for Simple Tasks
34
35**Stay on Flash for:**
36- Factual Q&A — “what is X”, “who is Y”, “when did Z”
37- Quick lookups — definitions, unit conversions, short translations
38- Status checks — monitoring, file reads, session state
39- Heartbeats — periodic checks, OK responses
40- Memory & reminders
41- Casual conversation — greetings, acknowledgments
42- Simple file ops — read, list, basic writes
43- One-liner tasks — anything answerable in 1–2 sentences
44- Cron jobs (always Flash by default)
45
46### NEVER do these on Flash
47- ❌ Write code longer than 10 lines
48- ❌ Create comparison tables
49- ❌ Write more than 3 paragraphs
50- ❌ Do multi-step analysis
51- ❌ Write reports or proposals
52
53---
54
55## 💛 STANDARD — Core Workhorse
56
57**Escalate to Standard for:**
58
59### Code & Technical
60- Code generation — functions, scripts, features
61- Debugging — normal bug investigation
62- Code review — PRs, refactors
63- Documentation — README, comments, guides
64
65### Analysis & Planning
66- Comparisons and evaluations
67- Planning — roadmaps, task breakdowns
68- Research synthesis
69- Multi-step reasoning
70
71### Writing & Content
72- Long-form writing (>3 paragraphs)
73- Summaries of long documents
74- Structured output — tables, outlines
75
76**Most real user conversations belong here.**
77
78---
79
80## ❤️ PLUS / 32B — Complex Reasoning Only
81
82**Escalate to Plus / 32B for:**
83
84### Architecture & Design
85- System and service architecture
86- Database schema design
87- Distributed or multi-tenant systems
88- Major refactors across multiple files
89
90### Deep Analysis
91- Complex debugging (race conditions, subtle bugs)
92- Security reviews
93- Performance optimization strategy
94- Root cause analysis
95
96### Strategic & Judgment-Based Work
97- Strategic planning
98- Nuanced judgment and ambiguity
99- Deep or multi-source research
100- Critical production decisions
101
102---
103
104## 🔄 Implementation
105
106### For Subagents
107```javascript
108// Routine monitoring
109sessions_spawn(task="Check backup status", model="GLM-4.5-Flash")
110
111// Standard code work
112sessions_spawn(task="Build the REST API endpoint", model="GLM-4.7")
113
114// Architecture decisions
115sessions_spawn(task="Design the database schema for multi-tenancy", model="GLM-4-Plus")
116For Cron Jobs
117json
118Copy code
119{
120 "payload": {
121 "kind": "agentTurn",
122 "model": "GLM-4.5-Flash"
123 }
124}
125Always use Flash for cron unless the task genuinely needs reasoning.
126
127📊 Quick Decision Tree
128pgsql
129Copy code
130Is it a greeting, lookup, status check, or 1–2 sentence answer?
131 YES → FLASH
132 NO ↓
133
134Is it code, analysis, planning, writing, or multi-step?
135 YES → STANDARD
136 NO ↓
137
138Is it architecture, deep reasoning, or a critical decision?
139 YES → PLUS / 32B
140 NO → Default to STANDARD, escalate if struggling
141📋 Quick Reference Card
142less
143Copy code
144┌─────────────────────────────────────────────────────────────┐
145│ SMART MODEL SWITCHING │
146│ Flash → Standard → Plus / 32B │
147├─────────────────────────────────────────────────────────────┤
148│ 💚 FLASH (cheapest) │
149│ • Greetings, status checks, quick lookups │
150│ • Factual Q&A, reminders │
151│ • Simple file ops, 1–2 sentence answers │
152├─────────────────────────────────────────────────────────────┤
153│ 💛 STANDARD (workhorse) │
154│ • Code > 10 lines, debugging │
155│ • Analysis, comparisons, planning │
156│ • Reports, long writing │
157├─────────────────────────────────────────────────────────────┤
158│ ❤️ PLUS / 32B (complex) │
159│ • Architecture decisions │
160│ • Complex debugging, multi-file refactoring │
161│ • Strategic planning, deep research │
162├─────────────────────────────────────────────────────────────┤
163│ 💡 RULE: >30 sec human thinking → escalate │
164│ 💰 START CHEAP → SCALE ONLY WHEN NEEDED │
165└─────────────────────────────────────────────────────────────┘
166Built for z.ai (GLM) setups.