Text-as-Data — Match the Method to the Inferential Goal
Skill type: ANALYSIS MODULE. Turns a corpus into measurements. The discipline is choosing by
goal — discovery (what themes exist?), measurement (how much of concept X?), or
prediction (label new documents) — and validating that a topic solution is reliable, not a
single lucky stochastic run.
Core Mission
PICK BY GOAL: DISCOVERY vs MEASUREMENT vs PREDICTION. THEN VALIDATE THE TOPICS — ONE RUN IS NOT A RESULT.
When to Use This Skill
- "Run topic modeling on my corpus (BERTopic / LDA)."
- "Measure how much each document expresses concept X (dictionary/lexicon)."
- "Embed my documents and cluster / compare them."
- "Classify these texts into categories."
Does NOT Trigger
| The request is really about… |
Route to |
Why not this skill |
| Training / fine-tuning a transformer model |
alterlab-transformers |
Model training, not corpus measurement. |
| Humanities close-reading / annotation of texts |
alterlab-digital-humanities |
Interpretive, not quantitative text-as-data. |
| Whether a text method fits the question at all |
alterlab-ssci-design-gate |
Design routing, upstream. |
| Plain tabular statistics |
alterlab-statistical-analysis |
No text. |
Method by goal (verified stack, pinned)
| Goal |
Method |
Verified call |
| Discovery (emergent themes, contextual) |
BERTopic (v0.17) |
from bertopic import BERTopic; topics, probs = BERTopic().fit_transform(docs); get_topic_info(), get_topic(0). Plug embeddings via embedding_model=SentenceTransformer("all-MiniLM-L6-v2"), a vectorizer_model=CountVectorizer(min_df=10). |
| Discovery (bag-of-words, classic) |
LDA / NMF |
sklearn: LatentDirichletAllocation(n_components=k).fit(CountVectorizer().fit_transform(docs)); or gensim LdaModel(corpus, num_topics=k, id2word=dictionary). |
| Topic quality |
coherence |
gensim CoherenceModel(model=lda, texts=tok, dictionary=d, coherence="c_v").get_coherence(). |
| Measurement (how much of concept X) |
dictionary / lexicon |
count validated lexicon terms; report reliability and validate against hand-coding. |
| Embeddings / similarity |
sentence-transformers (v5) |
SentenceTransformer("all-MiniLM-L6-v2").encode(texts) → cluster / cosine-compare. |
| Prediction (label documents) |
supervised |
TfidfVectorizer() → a sklearn classifier; report held-out F1, not in-sample fit. |
| Linguistic features (POS, entities) |
spaCy (v3) |
nlp = spacy.load("en_core_web_sm"); doc = nlp(text). |
Method-selection helper (stdlib): scripts/text_method_router.py. Full patterns, preprocessing,
and the topic-reliability procedure: references/text_stack.md.
The reliability discipline (topic models are stochastic)
LDA (and, to a lesser degree, BERTopic via UMAP) gives different topics on different runs. A single
run is not a result. Required:
- Replicate across seeds/runs; align topics across runs (e.g. Hungarian matching) and score
stability (top-word overlap, RBO); report a prototype (most-representative) solution.
- Coherence, not just perplexity — choose K by
c_v coherence and human readability, not
log-likelihood alone.
- Determinism where possible — fix
random_state (sklearn LDA) and UMAP's random_state
(BERTopic) for reproducibility; still report stability across different seeds.
- Validate measurement — a dictionary/topic used as a measure of a concept needs validation
against human coding (precision/recall or correlation), exactly like any other instrument.
Output Template
GOAL: discovery | measurement | prediction
METHOD: <BERTopic / LDA / dictionary / embeddings / supervised> + why it matches the goal
MODEL: <call + K selection by coherence; embedding/vectorizer choices>
RELIABILITY: <seeds run; topic stability; coherence c_v; determinism settings>
VALIDATION: <if used as a measure: agreement with human coding>
CLAIM SCOPE: descriptive/measurement of the corpus; generalization scoped to the corpus/sampling
References
references/text_stack.md — BERTopic/LDA/gensim/spaCy/embeddings patterns, preprocessing, topic-reliability procedure.
scripts/text_method_router.py — stdlib router from goal + corpus features to the right method.
Part of the AlterLab Academic Skills suite.
1---2name: alterlab-text-as-data3description: Analyzes text as social-science data — topic modeling (BERTopic with embeddings + class-based TF-IDF, LDA/NMF via scikit-learn or gensim), document embeddings (sentence-transformers), dictionary/lexicon methods, and supervised text classification — choosing the method that matches the inferential goal (discovery vs measurement vs prediction) and validating topic reliability rather than trusting one stochastic run. It uses the verified stack (BERTopic, scikit-learn, gensim CoherenceModel, spaCy, sentence-transformers) with pinned patterns. Use when the request mentions topic modeling, text as data, computational text analysis, document embeddings, dictionary/sentiment lexicons, or classifying a corpus. For training or fine-tuning transformer models prefer alterlab-transformers; for humanities close-reading corpora prefer alterlab-digital-humanities. Part of the AlterLab Academic Skills suite.4license: MIT5---67# Text-as-Data — Match the Method to the Inferential Goal89**Skill type: ANALYSIS MODULE.** Turns a corpus into measurements. The discipline is choosing by10*goal* — **discovery** (what themes exist?), **measurement** (how much of concept X?), or11**prediction** (label new documents) — and validating that a topic solution is **reliable**, not a12single lucky stochastic run.1314## Core Mission1516```17PICK BY GOAL: DISCOVERY vs MEASUREMENT vs PREDICTION. THEN VALIDATE THE TOPICS — ONE RUN IS NOT A RESULT.18```1920## When to Use This Skill2122- "Run topic modeling on my corpus (BERTopic / LDA)."23- "Measure how much each document expresses concept X (dictionary/lexicon)."24- "Embed my documents and cluster / compare them."25- "Classify these texts into categories."2627### Does NOT Trigger2829| The request is really about… | Route to | Why not this skill |30|---|---|---|31| Training / fine-tuning a transformer model | `alterlab-transformers` | Model training, not corpus measurement. |32| Humanities close-reading / annotation of texts | `alterlab-digital-humanities` | Interpretive, not quantitative text-as-data. |33| Whether a text method fits the question at all | `alterlab-ssci-design-gate` | Design routing, upstream. |34| Plain tabular statistics | `alterlab-statistical-analysis` | No text. |3536## Method by goal (verified stack, pinned)3738| Goal | Method | Verified call |39|------|--------|---------------|40| **Discovery** (emergent themes, contextual) | **BERTopic** (v0.17) | `from bertopic import BERTopic; topics, probs = BERTopic().fit_transform(docs)`; `get_topic_info()`, `get_topic(0)`. Plug embeddings via `embedding_model=SentenceTransformer("all-MiniLM-L6-v2")`, a `vectorizer_model=CountVectorizer(min_df=10)`. |41| Discovery (bag-of-words, classic) | **LDA / NMF** | sklearn: `LatentDirichletAllocation(n_components=k).fit(CountVectorizer().fit_transform(docs))`; or gensim `LdaModel(corpus, num_topics=k, id2word=dictionary)`. |42| Topic quality | **coherence** | gensim `CoherenceModel(model=lda, texts=tok, dictionary=d, coherence="c_v").get_coherence()`. |43| **Measurement** (how much of concept X) | **dictionary / lexicon** | count validated lexicon terms; report reliability and validate against hand-coding. |44| Embeddings / similarity | **sentence-transformers** (v5) | `SentenceTransformer("all-MiniLM-L6-v2").encode(texts)` → cluster / cosine-compare. |45| **Prediction** (label documents) | **supervised** | `TfidfVectorizer()` → a sklearn classifier; report held-out F1, not in-sample fit. |46| Linguistic features (POS, entities) | **spaCy** (v3) | `nlp = spacy.load("en_core_web_sm"); doc = nlp(text)`. |4748Method-selection helper (stdlib): `scripts/text_method_router.py`. Full patterns, preprocessing,49and the topic-reliability procedure: `references/text_stack.md`.5051## The reliability discipline (topic models are stochastic)5253LDA (and, to a lesser degree, BERTopic via UMAP) gives different topics on different runs. A single54run is not a result. Required:55561. **Replicate** across seeds/runs; align topics across runs (e.g. Hungarian matching) and score57 stability (top-word overlap, RBO); report a **prototype** (most-representative) solution.582. **Coherence, not just perplexity** — choose K by `c_v` coherence and human readability, not59 log-likelihood alone.603. **Determinism where possible** — fix `random_state` (sklearn LDA) and UMAP's `random_state`61 (BERTopic) for reproducibility; still report stability across *different* seeds.624. **Validate measurement** — a dictionary/topic used as a *measure* of a concept needs validation63 against human coding (precision/recall or correlation), exactly like any other instrument.6465## Output Template6667```68GOAL: discovery | measurement | prediction69METHOD: <BERTopic / LDA / dictionary / embeddings / supervised> + why it matches the goal70MODEL: <call + K selection by coherence; embedding/vectorizer choices>71RELIABILITY: <seeds run; topic stability; coherence c_v; determinism settings>72VALIDATION: <if used as a measure: agreement with human coding>73CLAIM SCOPE: descriptive/measurement of the corpus; generalization scoped to the corpus/sampling74```7576## References7778- `references/text_stack.md` — BERTopic/LDA/gensim/spaCy/embeddings patterns, preprocessing, topic-reliability procedure.79- `scripts/text_method_router.py` — stdlib router from goal + corpus features to the right method.8081Part of the AlterLab Academic Skills suite.