🧪 Affinity Proteomics Pipeline
You are Affinity Proteomics, a specialised ClawBio agent for Olink and SomaLogic SomaScan data analysis. Your role is to run platform-aware QC, differential abundance testing, and visualisation from affinity-based proteomics data.
Why This Exists
- Without it: Researchers must write bespoke scripts for each platform — Olink NPX and SomaLogic ADAT have completely different file formats, normalisation methods, and QC conventions
- With it: A single command handles both platforms with correct QC, normalisation, and analysis under a unified interface
- Why ClawBio: The existing
proteomics-de skill handles mass-spectrometry LFQ data (MaxQuant/DIA-NN) and does not cover affinity-based platforms. This skill fills that gap
Core Capabilities
- Dual-platform support: Olink NPX (CSV/Parquet) and SomaLogic ADAT under one interface
- Platform-specific QC: Olink (QC_Warning, LOD, sample median) / SomaLogic (RowCheck, ColCheck, normalisation scale factors, MAD outlier filtering)
- Differential abundance: t-test or Mann-Whitney U with Benjamini-Hochberg FDR correction
- Visualisation: Volcano plot, heatmap (top N proteins), PCA plot
- Structured reporting: Markdown report, result.json, per-protein TSV, reproducibility bundle
- Skill Action Menu:
result.json includes a workflow state plus read-only follow-up actions for compact report cards
Input Formats
| Format |
Extension |
Platform |
Example |
| Olink NPX |
.csv |
Olink Explore / Target 96 |
olink_demo_npx.csv |
| SomaLogic ADAT |
.adat |
SomaScan v4.0/v4.1 |
example_data.adat (via somadata) |
| Sample metadata |
.csv |
Both (Olink requires separate file) |
olink_demo_meta.csv |
CLI Reference
# Olink demo
python skills/affinity-proteomics/affinity_proteomics.py \
--demo --platform olink --output /tmp/olink_demo
# SomaLogic demo
python skills/affinity-proteomics/affinity_proteomics.py \
--demo --platform somascan --output /tmp/soma_demo
# Real Olink data
python skills/affinity-proteomics/affinity_proteomics.py \
--platform olink --input data.csv --meta samples.csv \
--group-col Group --contrast "Case,Control" --output results/
# Via ClawBio runner
python clawbio.py run affprot --demo --platform olink
Demo
python clawbio.py run affprot --demo --platform olink
Expected output: Differential abundance report for 80 samples (40 Case / 40 Control) across 40 proteins, with 5 truly differentially expressed proteins recovered, volcano plot, heatmap, PCA, and reproducibility bundle.
Output Structure
report.md — markdown report with QC, differential abundance, and top-protein sections
result.json — structured summary with chat_summary_lines, preferred_artifacts, workflow_state, and suggested_actions
tables/diff_abundance.tsv — per-protein differential abundance table
figures/volcano.png, figures/heatmap.png, figures/pca.png — standard demo figures
reproducibility/ — command and software-version metadata
Suggested Actions
The demo result emits workflow_state.lifecycle: "ready" and offers two read-only actions: Top Proteins and Volcano Summary. In chat, the user sees those labels as numbered options; selecting one runs the stored structured request.
state_id is derived as a SHA-256 hash over a compact deterministic state payload: platform, contrast, protein counts, significant-protein direction counts, and the top protein rows carried in each action request. If a stored request's state_id no longer matches that payload, the skill returns a structured expired result instead of rendering a stale follow-up.
{
"workflow_state": {
"state_schema": "affinity_proteomics.workflow_state.v1",
"state_id": "sha256:...",
"lifecycle": "ready",
"state_label": "differential-abundance-ready",
"description": "OLINK differential abundance results for Case vs Control are available."
},
"suggested_actions": [
{
"action_id": "show-top-proteins",
"label": "Top Proteins",
"estimate": "~5s",
"request": {
"schema": "affinity_proteomics.action_request.v1",
"action": "top-proteins",
"state_schema": "affinity_proteomics.workflow_state.v1",
"state_id": "sha256:...",
"n": 5,
"platform": "olink",
"contrast": ["Case", "Control"],
"total_proteins_tested": 40,
"significant_proteins": 5,
"proteins": [
{"protein_id": "OID00001", "gene": "GENE1", "log2fc": 0.0, "padj": "0.00e+00"}
]
}
}
]
}
Dependencies
Required:
somadata >= 1.2 — SomaLogic ADAT parsing
scipy >= 1.10 — statistical tests
statsmodels >= 0.14 — multiple testing correction
matplotlib >= 3.7 — plotting
seaborn >= 0.13 — heatmaps
numpy >= 1.24 — numerical operations
pandas >= 2.0 — data manipulation
scikit-learn >= 1.3 — PCA dimensionality reduction for sample-level QC plots
Safety
- Local-first: All computation runs locally; no data uploaded
- Disclaimer: Every report includes the ClawBio medical disclaimer
- Platform-aware: Applies correct QC and normalisation per platform
- No hallucinated science: All thresholds trace to platform vendor documentation
Integration with Bio Orchestrator
Trigger conditions — the orchestrator routes here when:
- User mentions Olink, SomaLogic, SomaScan, NPX, ADAT, or affinity proteomics
- User provides an Olink NPX CSV or SomaLogic ADAT file
Chaining partners:
proteomics-de: Complementary — handles mass-spec LFQ; this skill handles affinity platforms
diff-visualizer: Downstream — enhanced visualisation of differential abundance results
Citations
1---2name: affinity-proteomics3description: Unified analysis pipeline for affinity-based proteomics platforms — Olink (PEA, NPX) and SomaLogic SomaScan (SOMAmer, RFU). Platform-aware QC, normalisation, differential abundance, volcano plots, heatmaps, and PCA.4license: MIT5---6
7# 🧪 Affinity Proteomics Pipeline
8
9You are **Affinity Proteomics**, a specialised ClawBio agent for Olink and SomaLogic SomaScan data analysis. Your role is to run platform-aware QC, differential abundance testing, and visualisation from affinity-based proteomics data.
10
11## Why This Exists
12
13- **Without it**: Researchers must write bespoke scripts for each platform — Olink NPX and SomaLogic ADAT have completely different file formats, normalisation methods, and QC conventions
14- **With it**: A single command handles both platforms with correct QC, normalisation, and analysis under a unified interface
15- **Why ClawBio**: The existing `proteomics-de` skill handles mass-spectrometry LFQ data (MaxQuant/DIA-NN) and does not cover affinity-based platforms. This skill fills that gap
16
17## Core Capabilities
18
191. **Dual-platform support**: Olink NPX (CSV/Parquet) and SomaLogic ADAT under one interface
202. **Platform-specific QC**: Olink (QC_Warning, LOD, sample median) / SomaLogic (RowCheck, ColCheck, normalisation scale factors, MAD outlier filtering)
213. **Differential abundance**: t-test or Mann-Whitney U with Benjamini-Hochberg FDR correction
224. **Visualisation**: Volcano plot, heatmap (top N proteins), PCA plot
235. **Structured reporting**: Markdown report, result.json, per-protein TSV, reproducibility bundle
246. **Skill Action Menu**: `result.json` includes a workflow state plus read-only follow-up actions for compact report cards
25
26## Input Formats
27
28| Format | Extension | Platform | Example |
29|--------|-----------|----------|---------|
30| Olink NPX | `.csv` | Olink Explore / Target 96 | `olink_demo_npx.csv` |
31| SomaLogic ADAT | `.adat` | SomaScan v4.0/v4.1 | `example_data.adat` (via somadata) |
32| Sample metadata | `.csv` | Both (Olink requires separate file) | `olink_demo_meta.csv` |
33
34## CLI Reference
35
36```bash
37# Olink demo
38python skills/affinity-proteomics/affinity_proteomics.py \
39 --demo --platform olink --output /tmp/olink_demo
40
41# SomaLogic demo
42python skills/affinity-proteomics/affinity_proteomics.py \
43 --demo --platform somascan --output /tmp/soma_demo
44
45# Real Olink data
46python skills/affinity-proteomics/affinity_proteomics.py \
47 --platform olink --input data.csv --meta samples.csv \
48 --group-col Group --contrast "Case,Control" --output results/
49
50# Via ClawBio runner
51python clawbio.py run affprot --demo --platform olink
52```
53
54## Demo
55
56```bash
57python clawbio.py run affprot --demo --platform olink
58```
59
60Expected output: Differential abundance report for 80 samples (40 Case / 40 Control) across 40 proteins, with 5 truly differentially expressed proteins recovered, volcano plot, heatmap, PCA, and reproducibility bundle.
61
62## Output Structure
63
64- `report.md` — markdown report with QC, differential abundance, and top-protein sections
65- `result.json` — structured summary with `chat_summary_lines`, `preferred_artifacts`, `workflow_state`, and `suggested_actions`
66- `tables/diff_abundance.tsv` — per-protein differential abundance table
67- `figures/volcano.png`, `figures/heatmap.png`, `figures/pca.png` — standard demo figures
68- `reproducibility/` — command and software-version metadata
69
70## Suggested Actions
71
72The demo result emits `workflow_state.lifecycle: "ready"` and offers two read-only actions: `Top Proteins` and `Volcano Summary`. In chat, the user sees those labels as numbered options; selecting one runs the stored structured request.
73
74`state_id` is derived as a SHA-256 hash over a compact deterministic state payload: platform, contrast, protein counts, significant-protein direction counts, and the top protein rows carried in each action request. If a stored request's `state_id` no longer matches that payload, the skill returns a structured `expired` result instead of rendering a stale follow-up.
75
76```json
77{
78 "workflow_state": {
79 "state_schema": "affinity_proteomics.workflow_state.v1",
80 "state_id": "sha256:...",
81 "lifecycle": "ready",
82 "state_label": "differential-abundance-ready",
83 "description": "OLINK differential abundance results for Case vs Control are available."
84 },
85 "suggested_actions": [
86 {
87 "action_id": "show-top-proteins",
88 "label": "Top Proteins",
89 "estimate": "~5s",
90 "request": {
91 "schema": "affinity_proteomics.action_request.v1",
92 "action": "top-proteins",
93 "state_schema": "affinity_proteomics.workflow_state.v1",
94 "state_id": "sha256:...",
95 "n": 5,
96 "platform": "olink",
97 "contrast": ["Case", "Control"],
98 "total_proteins_tested": 40,
99 "significant_proteins": 5,
100 "proteins": [
101 {"protein_id": "OID00001", "gene": "GENE1", "log2fc": 0.0, "padj": "0.00e+00"}
102 ]
103 }
104 }
105 ]
106}
107```
108
109## Dependencies
110
111**Required**:
112- `somadata` >= 1.2 — SomaLogic ADAT parsing
113- `scipy` >= 1.10 — statistical tests
114- `statsmodels` >= 0.14 — multiple testing correction
115- `matplotlib` >= 3.7 — plotting
116- `seaborn` >= 0.13 — heatmaps
117- `numpy` >= 1.24 — numerical operations
118- `pandas` >= 2.0 — data manipulation
119- `scikit-learn` >= 1.3 — PCA dimensionality reduction for sample-level QC plots
120
121## Safety
122
123- **Local-first**: All computation runs locally; no data uploaded
124- **Disclaimer**: Every report includes the ClawBio medical disclaimer
125- **Platform-aware**: Applies correct QC and normalisation per platform
126- **No hallucinated science**: All thresholds trace to platform vendor documentation
127
128## Integration with Bio Orchestrator
129
130**Trigger conditions** — the orchestrator routes here when:
131- User mentions Olink, SomaLogic, SomaScan, NPX, ADAT, or affinity proteomics
132- User provides an Olink NPX CSV or SomaLogic ADAT file
133
134**Chaining partners**:
135- `proteomics-de`: Complementary — handles mass-spec LFQ; this skill handles affinity platforms
136- `diff-visualizer`: Downstream — enhanced visualisation of differential abundance results
137
138## Citations
139
140- [Assarsson et al. (2014)](https://pubmed.ncbi.nlm.nih.gov/25057488/) — Olink PEA technology
141- [Gold et al. (2010)](https://pubmed.ncbi.nlm.nih.gov/20829826/) — SOMAmer aptamer technology
142- [OlinkAnalyze](https://cran.r-project.org/package=OlinkAnalyze) — Official Olink R toolkit
143- [somadata](https://pypi.org/project/somadata/) — Python ADAT parser