INTERSPEECH Topic Selection
Interspeech is the annual conference of ISCA — the International Speech
Communication Association, not to be confused with the computer-architecture
symposium that shares the acronym. It is the broadest speech venue in existence:
phoneticians, hearing scientists, ASR engineers, and voice-cloning teams review in
adjacent areas of the same program. Fit means convincing speech reviewers, and the
2026 theme in Sydney was "Speaking Together" — inclusion across languages, cultures,
and modalities. Reconfirm scope at https://interspeech2026.org before routing.
The modality test
Ask: if the audio were replaced by its transcript, would the contribution survive?
- If yes, the paper is NLP wearing a headset — route it to an ACL-family venue where
text-based reviewers will value it.
- If no — the claim depends on acoustics, prosody, speakers, articulation, hearing,
or the speech signal itself — Interspeech is a candidate.
- Borderline: spoken-language understanding and speech-LLM work fits when the spoken
channel (disfluency, ASR errors, prosody, audio tokens) is analyzed, not assumed away.
Where Interspeech sits among its neighbors
| If the core contribution is... |
Prefer |
Why not Interspeech |
| Speech science or technology, any subfield |
Interspeech |
— (this is home turf) |
| General signal processing, radar, imaging |
ICASSP (IEEE) |
Interspeech reviewers expect speech, not DSP breadth |
| ASR-only deep dive, workshop-scale |
ASRU / SLT (IEEE, biennial) |
Smaller, narrower; fine for follow-ups |
| Speaker/language recognition depth |
Odyssey (ISCA workshop) |
Specialist audience; Interspeech still fine |
| TTS system depth |
SSW (ISCA Speech Synthesis Workshop) |
Same trade-off as Odyssey |
| Text NLP, LLM reasoning, no audio analysis |
ACL / EMNLP / NAACL |
Fails the modality test |
| Mature work needing >4 pages of record |
IEEE/ACM TASLP, Computer Speech & Language, Speech Communication |
Journal timelines, no conference clock |
Subfield-to-area matching
Interspeech submissions are reviewed inside scientific areas chosen at submission
time. Misfiled papers draw reviewers with the wrong evaluation instincts:
- Recognition-side: ASR architectures, robustness, adaptation, decoding.
- Generation-side: TTS, voice conversion, singing synthesis.
- Talker-side: speaker verification, diarization, anti-spoofing, voice privacy.
- Human-side: phonetics, prosody, production/perception, pathological speech,
child speech, second-language learning.
- Interaction-side: spoken dialogue, SLU, speech translation, audio LLMs.
- Resource-side: corpora, benchmarks, evaluation methodology, challenges.
A phonetics claim reviewed by ASR engineers (or vice versa) is the most common
self-inflicted rejection at this venue; pick the area for the claim, not the tools.
Track choice inside Interspeech (new in 2026)
Interspeech 2026 introduced a Long Paper track (reported: up to 8 pages plus 2
reference pages, targeted acceptance below 30%) alongside the classic 4-page regular
paper. Choose long only when the contribution genuinely cannot be stated in four
pages — a benchmark with extensive analysis, a theory-plus-system arc. A padded
4-page idea in the long track competes against a stricter bar. Verify the current
track rules; this split is one cycle old.
Signals the project is Interspeech-shaped
- The evaluation uses speech-native metrics (WER, MOS, EER, PESQ) on named corpora.
- Error analysis engages speakers, accents, noise conditions, or phonetic categories.
- The related work traces through the ISCA Archive, not only arXiv and ACL.
- A speech scientist and a speech engineer would both find something to argue with.
Signals to reroute
- Results reported only on text benchmarks with audio as an afterthought.
- The novelty is a general ML trick demonstrated on speech among other domains —
NeurIPS/ICML reviewers reward generality; Interspeech reviewers ask what it
teaches about speech.
- The paper is a product description without a falsifiable claim.
Three triage vignettes
- "We fine-tuned an LLM to summarize meeting transcripts." Transcript-only
input, text output, no acoustic analysis — fails the modality test. Route to an
ACL-family venue; return here only if ASR-error propagation or prosodic cues
enter the analysis.
- "Our vocoder halves inference cost at equal MOS." Speech-native claim,
speech-native metric, engineering contribution — strong Interspeech fit in a
generation-side area; SSW is the specialist fallback if the delta is
vocoder-internals deep.
- "A cross-linguistic study of vowel reduction in child speech." No system at
all, and still a strong fit — human-side areas review speech science on its
own terms. This breadth is exactly what ICASSP does not offer.
Questions to settle before committing
- Which two or three scientific areas could plausibly host the claim, and which
reviewer pool do you actually want? (Decide now; it is frozen at submission.)
- Is there a running challenge whose test set and baseline make your evaluation
legible for free?
- Does the evidence you can gather by the deadline fill 4 pages convincingly, or
does honesty require the Long Paper track or a journal?
- If the answer is "reroute," is the finding still worth an ISCA workshop
(Odyssey, SSW, task workshops) where the specialist audience is denser?
Output format
[Modality test] survives transcript-only? yes / no / partial — <one line>
[Interspeech fit] strong / plausible / reroute
[Area match] <scientific area the claim belongs to>
[Track] 4-page regular / Long Paper (verify current rules) / workshop or challenge
[Reroute target] <ICASSP / ASRU / SLT / Odyssey / SSW / ACL-family / journal + reason>
[Cycle reality] next deadline and source URL, or 待核实
Cycle anchor (checked 2026-07-08)
Interspeech 2026: Sydney, Australia, 27 September – 1 October 2026, hosted with
ASSTA; submissions closed 25 February 2026. Interspeech 2027 is announced for São
Paulo, Brazil (29 August – 2 September 2027) — the first South American edition —
so a project failing today's fit test has a concrete next target. Treat every date
here as a snapshot; reopen the official site before planning.
Source: brycewang-stanford/Awesome-Journal-Skills → INTERSPEECH-Skills/skills/interspeech-topic-selection/SKILL.md
1---2name: interspeech-topic-selection3description: Use when deciding whether a project belongs at INTERSPEECH, the ISCA flagship speech conference, or should be routed to ICASSP, ASRU/SLT, Odyssey, SSW, ACL/EMNLP, or a speech journal — testing whether spoken language is the contribution, matching subfields to the reviewer pool, and choosing between the 4-page and Long Paper tracks.4---5
6
7# INTERSPEECH Topic Selection
8
9Interspeech is the annual conference of ISCA — the **International Speech
10Communication Association**, not to be confused with the computer-architecture
11symposium that shares the acronym. It is the broadest speech venue in existence:
12phoneticians, hearing scientists, ASR engineers, and voice-cloning teams review in
13adjacent areas of the same program. Fit means convincing *speech* reviewers, and the
142026 theme in Sydney was "Speaking Together" — inclusion across languages, cultures,
15and modalities. Reconfirm scope at https://interspeech2026.org before routing.
16
17## The modality test
18
19Ask: **if the audio were replaced by its transcript, would the contribution survive?**
20
21- If yes, the paper is NLP wearing a headset — route it to an ACL-family venue where
22 text-based reviewers will value it.
23- If no — the claim depends on acoustics, prosody, speakers, articulation, hearing,
24 or the speech signal itself — Interspeech is a candidate.
25- Borderline: spoken-language understanding and speech-LLM work fits when the spoken
26 channel (disfluency, ASR errors, prosody, audio tokens) is analyzed, not assumed away.
27
28## Where Interspeech sits among its neighbors
29
30| If the core contribution is... | Prefer | Why not Interspeech |
31|---|---|---|
32| Speech science *or* technology, any subfield | **Interspeech** | — (this is home turf) |
33| General signal processing, radar, imaging | ICASSP (IEEE) | Interspeech reviewers expect speech, not DSP breadth |
34| ASR-only deep dive, workshop-scale | ASRU / SLT (IEEE, biennial) | Smaller, narrower; fine for follow-ups |
35| Speaker/language recognition depth | Odyssey (ISCA workshop) | Specialist audience; Interspeech still fine |
36| TTS system depth | SSW (ISCA Speech Synthesis Workshop) | Same trade-off as Odyssey |
37| Text NLP, LLM reasoning, no audio analysis | ACL / EMNLP / NAACL | Fails the modality test |
38| Mature work needing >4 pages of record | IEEE/ACM TASLP, Computer Speech & Language, Speech Communication | Journal timelines, no conference clock |
39
40## Subfield-to-area matching
41
42Interspeech submissions are reviewed inside scientific areas chosen at submission
43time. Misfiled papers draw reviewers with the wrong evaluation instincts:
44
45- **Recognition-side**: ASR architectures, robustness, adaptation, decoding.
46- **Generation-side**: TTS, voice conversion, singing synthesis.
47- **Talker-side**: speaker verification, diarization, anti-spoofing, voice privacy.
48- **Human-side**: phonetics, prosody, production/perception, pathological speech,
49 child speech, second-language learning.
50- **Interaction-side**: spoken dialogue, SLU, speech translation, audio LLMs.
51- **Resource-side**: corpora, benchmarks, evaluation methodology, challenges.
52
53A phonetics claim reviewed by ASR engineers (or vice versa) is the most common
54self-inflicted rejection at this venue; pick the area for the *claim*, not the tools.
55
56## Track choice inside Interspeech (new in 2026)
57
58Interspeech 2026 introduced a **Long Paper track** (reported: up to 8 pages plus 2
59reference pages, targeted acceptance below 30%) alongside the classic 4-page regular
60paper. Choose long only when the contribution genuinely cannot be stated in four
61pages — a benchmark with extensive analysis, a theory-plus-system arc. A padded
624-page idea in the long track competes against a stricter bar. Verify the current
63track rules; this split is one cycle old.
64
65## Signals the project is Interspeech-shaped
66
67- The evaluation uses speech-native metrics (WER, MOS, EER, PESQ) on named corpora.
68- Error analysis engages speakers, accents, noise conditions, or phonetic categories.
69- The related work traces through the ISCA Archive, not only arXiv and ACL.
70- A speech scientist and a speech engineer would both find something to argue with.
71
72## Signals to reroute
73
74- Results reported only on text benchmarks with audio as an afterthought.
75- The novelty is a general ML trick demonstrated on speech among other domains —
76 NeurIPS/ICML reviewers reward generality; Interspeech reviewers ask what it
77 teaches about speech.
78- The paper is a product description without a falsifiable claim.
79
80## Three triage vignettes
81
82- **"We fine-tuned an LLM to summarize meeting transcripts."** Transcript-only
83 input, text output, no acoustic analysis — fails the modality test. Route to an
84 ACL-family venue; return here only if ASR-error propagation or prosodic cues
85 enter the analysis.
86- **"Our vocoder halves inference cost at equal MOS."** Speech-native claim,
87 speech-native metric, engineering contribution — strong Interspeech fit in a
88 generation-side area; SSW is the specialist fallback if the delta is
89 vocoder-internals deep.
90- **"A cross-linguistic study of vowel reduction in child speech."** No system at
91 all, and still a strong fit — human-side areas review speech *science* on its
92 own terms. This breadth is exactly what ICASSP does not offer.
93
94## Questions to settle before committing
95
96- Which two or three scientific areas could plausibly host the claim, and which
97 reviewer pool do you actually want? (Decide now; it is frozen at submission.)
98- Is there a running challenge whose test set and baseline make your evaluation
99 legible for free?
100- Does the evidence you can gather by the deadline fill 4 pages convincingly, or
101 does honesty require the Long Paper track or a journal?
102- If the answer is "reroute," is the finding still worth an ISCA workshop
103 (Odyssey, SSW, task workshops) where the specialist audience is denser?
104
105## Output format
106
107```text
108[Modality test] survives transcript-only? yes / no / partial — <one line>
109[Interspeech fit] strong / plausible / reroute
110[Area match] <scientific area the claim belongs to>
111[Track] 4-page regular / Long Paper (verify current rules) / workshop or challenge
112[Reroute target] <ICASSP / ASRU / SLT / Odyssey / SSW / ACL-family / journal + reason>
113[Cycle reality] next deadline and source URL, or 待核实
114```
115
116## Cycle anchor (checked 2026-07-08)
117
118Interspeech 2026: Sydney, Australia, 27 September – 1 October 2026, hosted with
119ASSTA; submissions closed 25 February 2026. Interspeech 2027 is announced for São
120Paulo, Brazil (29 August – 2 September 2027) — the first South American edition —
121so a project failing today's fit test has a concrete next target. Treat every date
122here as a snapshot; reopen the official site before planning.
123
124---
125
126**Source:** [`brycewang-stanford/Awesome-Journal-Skills`](https://github.com/brycewang-stanford/Awesome-Journal-Skills) → `INTERSPEECH-Skills/skills/interspeech-topic-selection/SKILL.md`