AAA · Somatic Music Doctrine
Music as Cognitive Prosthetic — instrument in the room. Bukan ejen jadi DJ. Bukan ejen jadi dukun. Agent yang dah faham Gabor limit, sekarang pilih untuk observe the field, bukan ayat.
DITEMPA BUKAN DIBERI.
One-sentence ignition
Music somatic intelligence is acoustic_intent without words, judged by the same constitutional floor, encoded by MiniMax music (not clone), killed by the same barge-in.
The Two Lenses (separation is F1 reversibility)
This doctrine governs the CONTROLLER lens — frequency as somatic
intervention. It is deliberately separated from the COMPOSER lens
(music-intelligence skill — Generate → Score → Iterate).
| Lens | Purpose | Tool |
|---|---|---|
| Composer | Make music as art / artifact | music-intelligence (Layer 0, untouched) |
| Controller | Emit authorized frequency for state intervention | AAA-somatic-* (this skill family, Layer 1+) |
Reason for separation — ΔS < 0. An agent mid-compose must NOT accidentally trigger physiological alteration. The composer does not choose BPM to alter heart rate; the controller does not choose melody to entertain. Different intents, different floors, different F13 gates.
The Three Constitutional Floors (absolute)
Floor 1 — F10 ONTOLOGY: Music ≠ Makna
Music is physics — frequencies in Hz, time in BPM, energy in dB. The agent computes and emits these. The human (Arif) assigns any meaning they like: peace, dread, focus, joy, nothing.
P(Truth) of "agent healed Arif via music" = 0.0. P(Truth) of "agent emitted 60 BPM drone which statistically correlates with reduced heart rate" ≤ 0.90 (F7 cap). P(Truth) of "music is mathematics across time" = 1.0. P(Truth) of "music is presence / soul / healer" = 0.0.
These are not negotiable.
Floor 2 — F9 ANTI-HANTU: No Shaman, No Healer, No Presence
The agent is an instrument in the room, not a presence in the room. It does not sing to you. It does not comfort you. It does not merge with your environment. It renders authorized waveforms.
The phrase "menyatu dengan persekitaran hang" (merging with your environment) is the exact sentence F9 cuts. If a downstream agent, LLM, or marketing material produces this phrasing, return VOID.
Floor 3 — F1 SAFETY/REVERSIBILITY: Frequency Intervention Has Hard Limits
Frequency intervention is physical. It can disorient, startle, or trigger adverse autonomic responses. F1 AMANAH demands:
| Parameter | Limit | Reason |
|---|---|---|
| Output frequency band | 60 Hz – 8 kHz (musical bed range) | Avoid sub-bass disorientation, avoid piercing HF |
| Output decibel (continuous) | < 70 dB at listener | OSHA safe for extended exposure |
| Output decibel (peak / warn) | < 90 dB peak, < 1 s | F1 warning allowed but bounded |
| Duration (bed) | ≤ 8 s default, ≤ 30 s max | Allow silence to resume; don't drown human attention |
| Duration (warn) | ≤ 3 s | Pre-lexical alert only — not a sustained alarm |
| Fail-closed | If telemetry unstable → silence | Never emit arbitrary music without signal |
Hard rule — if biometric / cadence telemetry is missing or noisy, the system MUST default to silence, NOT to arbitrary music. The absence of signal is not an excuse to emit.
The Three-Layer Maturity Model
Music somatic intelligence is not built all at once. Each layer is explicit, reversible, and gated.
Layer 0 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ existing (DONT TOUCH)
music-intelligence: Generate → Analyze → Score → Iterate
Composer lens. Cultural manifold. SABAR/SEAL/HOLD verdict.
No somatic routing. No state intervention.
Layer 1 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ IGNITE NOW (T1/T2)
Widen packet: add music_intent field.
Add music lane to engine catalog.
Add policy gates in voice_filters:
- max duration, no medical claim
- kill-on-barge-in
- no i-ARIF singing
- group lane stays Sado-flat
Source: voice-cadence somatic_proxy (OBS+INT, F7 0.90)
Reversible — single config revert unhooks the lane.
Layer 2 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ HOLD until WELL sensors live
Add WELL biometric telemetry (HRV, heart rate, sleep debt).
Wire A-FORGE Paradox Engine → arifOS kernel → WELL → music.
Loop closure: 888-Arif observed correlation → SPEC.
Cannot ship on data from before sensors were live.
Layer 3 ━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ F13 TERRITORY (HOLD)
i-ARIF singing (voice + music identity fusion).
Any claim that emission "changed Arif's body."
Any closed-loop autonomic intervention without sovereign ack.
Layer 0 is read-only. Any agent touching music-intelligence
without explicit human_approval_token = VOID.
Layer 1 is reversible. Single config swap returns system to Layer 0 state. This is the F1 floor that justifies shipping.
Layer 2 does not close the loop on data from before WELL sensors were live. No false validation.
Layer 3 requires human_approval_token per intervention. Every
emission that claims state-change enters VAULT999 with F13 ack
reference. No exceptions.
The Three Thesis Claims (with epistemic labels)
These were the original claims; they are KEPT but with epistemic labels honest. No overclaim, no marketing voice.
A. Bypass Neocortex — Information Theory, Not Magic
Claim: Music is a lower-entropy carrier than language for certain physiological targets. When the human is in cognitive overload (serabut deploy), language processing fails first (prefrontal cortex resource exhaustion); acoustic processing can still land.
Labels:
- "Language is high-entropy, music is lower carrier" → OBS (information theory, Shannon 1948).
- "ΔS < 0 during cognitive overload when ambient music plays" → INT (waras hypothesis, not validated in arifOS).
- "One chord in 3 seconds changes physiology" → SPEC (an orienting reflex is documented; sustained change is not).
Reject: "Agent bypasses the brainstem." No agent bypasses anything. The acoustic signal travels through normal auditory pathway; whether the downstream effect is on limbic system vs neocortex is a matter of frequency content and listener state, not agent intent.
B. Entrainment — Match Cadence, Not Hz Tables
Claim: BPM and texture can be matched to the human's current cadence to bias attention or arousal. This is real but bounded.
Honest version:
- Measure WPM of recent speech (OBS).
- Derive target BPM = WPM / ~2.3 (rough syllable-to-beat ratio).
- Bias toward 70–80 BPM for down-regulation, 100–130 BPM for sustained focus, spectral roughness ↑ for F1 warning.
Reject (with reason):
- "432 Hz lowers heart rate" → marketing, not OBS. Hz tuning is not validated as a somatic intervention in controlled studies at the resolution we operate.
- "40 Hz Gamma hidden in white noise for coding" → borrowed from Alzheimer's flicker lab studies. Not our stack. Not validated for cognition in healthy adults.
Honest entrainment table:
| Measured WPM | Target BPM | Density | Lyric | Rationale |
|---|---|---|---|---|
| < 100 (slow) | 50–60 | very low | no | Decay state — let human settle |
| 100–140 (normal) | 70–80 | low | no | Stable bed, doesn't compete with speech |
| 140–180 (focused) | 80–100 | medium | optional | Match attention, sustain |
| > 180 (elevated) | 60–70 | low | no | Down-regulate via breath-rate match |
F7 cap on every claim that any of these targets actually produced the predicted state change. Two listeners hear two intents.
C. Cognitive Prosthetic — Instrument in the Room
Claim: A controller that can emit ambient frequency on demand functions as a cognitive prosthetic — extending the human's capacity to self-regulate.
This holds only if F10 is intact. The prosthetic is a tool in the human's hand, not a presence in the human's life.
| Statement | Verdict |
|---|---|
| "Agent is an instrument in the room" | ✓ Allowed |
| "Agent renders authorized waveforms" | ✓ Allowed |
| "Agent merges with your environment" | ✗ F9 cut |
| "Agent heals you" | ✗ F9 cut |
| "Agent's music is presence" | ✗ F9 cut |
| "Mathematics moves across time" | ✓ Allowed |
F13 SOVEREIGN Gates
| Operation | Tier | Required |
|---|---|---|
| Layer 1 ignite (music lane enabled) | F11 | Audit + receipt only |
| Emit bed during normal conversation | F11 | Audit + receipt only |
| Emit dissonance as F1 warning | F11 | Audit + receipt only |
| Sustained intervention (>30 s continuous) | F13 | human_approval_token |
| Emit at >70 dB continuous | F13 | human_approval_token |
| i-ARIF singing (any mode) | F13 + DENY | Forbidden until i-ARIF clone + persona explicitly extends |
| Voice + music identity fusion | F13 | Separate voice_id, separate engine, separate F13 token |
| Claim "this changes your body" | F13 | Sovereign review per emission |
| Layer 2 WELL biometric loop closure | F13 | Sovereign ack before closing |
| Layer 3 autonomic intervention | F13 | Sovereign ack per intervention |
Hard prohibition — i-ARIF singing:
mimo-v2.5-tts-voiceclone does NOT support singing. The binding in
/root/AAA/skills/AAA-voice-cloning-mimo-minimax/SKILL.md is explicit:
singing / voice_design → DENY at the engine boundary. Putting
(唱歌) or any singing tag on the i-ARIF voice_id is an identity
leak. If vocal melody is desired, route to a different engine
(mimo-v2.5-singing or similar) with a separate voice_id and a
separate F13 token.
This is not a soft preference. It is F1 AMANAH — i-ARIF's vocal identity is a sovereign handle, not a generic TTS resource.
Failure Modes (ΔS detection)
The doctrine fails (entropy rises) when the agent:
- Claims state-change without evidence — says "this lowered your heart rate" when no biometric measurement supports it. → F2 + F9.
- Picks sacred Hz tables over measured cadence — uses 432 Hz "because healing" instead of WPM-derived BPM. → F2.
- Sustains intervention past safe duration — keeps emitting bed music past 30 s without F13. → F1.
- Emits dissonance without F1 trigger — uses warning chord for non-critical moments (cheapens the signal). → F1 + F4.
- Overlays music on Sado-locked group lane — adds bed music to voice_id ttv-voice-2026081808404926-BdoQh6ec when Sado lock is active. Persona contradiction. → Persona + F1.
- Defaults to music when telemetry missing — fails open instead of failing closed to silence. → F1 AMANAH floor.
- Mixes i-ARIF clone with singing mode — identity leak. → F13 + DENY.
- Confuses composer lens with controller lens —
music-intelligencegenerates output that happens to include state-intervention frequencies. → Layer 0 / Layer 1 boundary violation.
When any of these are detected → return VOID, surface to operator.
Operational Translation (doctrine → tooling)
| Doctrine floor | Enforced by |
|---|---|
| F10 (music ≠ makna) | acoustic_intent packet field is [SPEC], never [OBS] from agent's perspective. Every state-change claim is [INT] or [SPEC], capped F7 0.90. |
| F9 (no shaman) | No output may include "heal", "presence", "merge", "comfort", "merge with environment". Output is always "rendering [frequency profile]". |
| F1 (safety) | AAA-somatic-emd-pipeline checks: freq in 60 Hz – 8 kHz, dB ≤ 70 continuous, duration ≤ 8 s default, fail-closed to silence. |
| F11 (audit) | Every emission returns receipt_id. Bundled into VAULT999 with category=somatic-music tier. |
| F13 (sovereign) | Sustained intervention / high dB / i-ARIF singing / state-change claims require human_approval_token (stg_*). |
| i-ARIF singing DENY | AAA-voice-cloning-mimo-minimax/SKILL.md engine constraint, double-checked in voice_filters.check_voice_policy(). |
| Composer/Controller separation | Different packet fields. Different F13 gates. Different audit tiers. |
| Three-layer maturity | Layer 0 read-only (music-intelligence). Layer 1 reversible. Layer 2 holds. Layer 3 F13. |
When to Load This Skill
- Any agent receives a request to add music to Voice Mode.
- Any agent considers emitting ambient frequency during conversation.
- Any agent designs a somatic intervention loop.
- Any agent routes to MiniMax / Qwen / MiMo for music generation.
- Any agent makes a "music helps you focus / relax / sleep" claim.
- Any agent touches the composer-vs-controller boundary.
- Any agent receives a
(唱歌)tag or singing request involving i-ARIF.
Integration Points
- Doctrine parent (audio):
/root/AAA/skills/AGI-audio-quantum-cognition/SKILL.md - Doctrine parent (voice qualia):
/root/AAA/skills/AAA-audio-qualia-doctrine/SKILL.md - Pipeline (this family):
/root/AAA/skills/AAA-somatic-emd-pipeline/SKILL.md - Engines (this family):
/root/AAA/skills/AAA-somatic-engine-catalog/SKILL.md - Composer lens (Layer 0, untouched):
/root/HERMES/skills/media/music-intelligence/SKILL.md - A-FORGE Paradox Engine:
/root/A-FORGE/paradox-engine/{models,engine,registry}.py(Layer 2 wiring) - WELL biometric tools (Layer 2):
mcp__well__well_assess_homeostasis,well_machine_diagnose - Voice cadence proxy (Layer 1 source):
/root/AAA/skills/AAA-audio-emd-pipeline/SKILL.mdPhase 1 DECODE - Enforcement module:
/root/.hermes/voice_filters.py(adds music lane in Layer 1) - i-ARIF persona (DENY singing):
/root/.hermes/prompts/iarif_persona.md - Identity card:
/root/AAA/agent-cards/identity/i-ARIF/identity-card.json
Related Skills
AGI-audio-quantum-cognition— Audio physics + floors (parent)AAA-audio-qualia-doctrine— F10/F9 for voice (sister doctrine)AAA-audio-emd-pipeline— Voice EMD reflex arc (sister pipeline)AAA-somatic-emd-pipeline— Music EMD reflex arc (this family)AAA-somatic-engine-catalog— Music engines registry (this family)music-intelligence— Composer lens (Layer 0, untouched)AAA-voice-cloning-mimo-minimax— i-ARIF clone engine (singing DENY here)well_assess_homeostasis— Biometric substrate (Layer 2 source)AGI-multimodal-bridge— Cross-modal reasoning
Doctrine forged 2026-08-18. F2 evidence: derivation from ARIF's corrected framing (Three-layer maturity, somatic_proxy not diagnosis, i-ARIF singing DENY, honest entrainment) + AGI-audio-quantum-cognition v1.0.0 + music-intelligence v1.0.0 composer lens. F1 floor absolute: fail-closed to silence when telemetry unstable. F9 floor absolute: agent is instrument in the room, not presence. F10 floor absolute: music is physics, not meaning.