Cheminformatics
Cheminformatics skill. RDKit molecular property calculation, SMILES/InChI handling, molecular fingerprints, substructure search, chemical similarity, and ADMET prediction.
Use This Skill When
- RDKit molecular property calculation.
- SMILES/InChI handling.
- Molecular fingerprints.
- Substructure search.
- Chemical similarity.
Required Inputs
- Research objective, decision target, or hypothesis.
- Available data, source constraints, and domain assumptions.
- Required outputs, success metrics, and deadline or reproducibility constraints.
Workflow
- Confirm scope, assumptions, and the exact artifact set to save.
- Apply the narrowest domain method that answers the request with defensible evidence.
- Save code, tables, figures, and intermediate outputs to files instead of chat-only output.
- State limitations, uncertainty, and any validation or sensitivity checks performed.
- Append skill selection, handoff I/O, and file writes to
logs/process-log.jsonl.
Deliverables
report.md: concise method, results, interpretation, and file inventory in the user's language.
results/: structured outputs, metrics, model artifacts, or extracted findings.
figures/: English-only charts, diagrams, or panels when visual output is needed.
data/: processed or derived datasets when transformation occurs.
Available Tools (MCP)
External tools available via ToolUniverse MCP server.
Falls back to Python requests + public REST APIs when MCP is unavailable.
| Source |
Tool |
Description |
| PubChem |
PubChem_search_compound |
PubChem API |
| PubChem |
PubChem_get_compound |
PubChem API |
| ChEMBL |
ChEMBL_search_compound |
ChEMBL API |
| ChEMBL |
ChEMBL_get_activity |
ChEMBL API |
Quality Gates
If any gate fails: identify the specific failing check, fix the issue, and re-validate before proceeding.
Gotchas
- SMILES strings may represent different stereoisomers. Canonicalize SMILES before database lookups
- Assay results from different sources use different activity units (IC50, Ki, EC50). Standardize before comparison
- Chemical similarity metrics (Tanimoto, Dice) give different rankings. Report fingerprint and metric used
Validation Loop
- Execute analysis and generate outputs
- Check:
- Method selection matches the research question and stated assumptions
- All outputs are saved to files (no chat-only results)
- Limitations and uncertainty are explicitly stated
logs/process-log.jsonl is updated with execution trace
- If any check fails:
- Identify the failing gate
- Fix the specific issue
- Re-run validation
- Proceed only after all gates pass
1---2name: co-scientist-cheminformatics3description: Cheminformatics skill. RDKit molecular property calculation, SMILES/InChI handling, molecular fingerprints, substructure search, chemical similarity, and ADMET prediction. Use when working with rdkit molecular property calculation, smiles/inchi handling, molecular fingerprints.4---56# Cheminformatics78Cheminformatics skill. RDKit molecular property calculation, SMILES/InChI handling, molecular fingerprints, substructure search, chemical similarity, and ADMET prediction.910## Use This Skill When1112- RDKit molecular property calculation.13- SMILES/InChI handling.14- Molecular fingerprints.15- Substructure search.16- Chemical similarity.1718## Required Inputs1920- Research objective, decision target, or hypothesis.21- Available data, source constraints, and domain assumptions.22- Required outputs, success metrics, and deadline or reproducibility constraints.2324## Workflow25261. Confirm scope, assumptions, and the exact artifact set to save.272. Apply the narrowest domain method that answers the request with defensible evidence.283. Save code, tables, figures, and intermediate outputs to files instead of chat-only output.294. State limitations, uncertainty, and any validation or sensitivity checks performed.305. Append skill selection, handoff I/O, and file writes to `logs/process-log.jsonl`.3132## Deliverables3334- `report.md`: concise method, results, interpretation, and file inventory in the user's language.35- `results/`: structured outputs, metrics, model artifacts, or extracted findings.36- `figures/`: English-only charts, diagrams, or panels when visual output is needed.37- `data/`: processed or derived datasets when transformation occurs.3839## Available Tools (MCP)4041> External tools available via [ToolUniverse](https://github.com/mims-harvard/ToolUniverse) MCP server.42> Falls back to Python `requests` + public REST APIs when MCP is unavailable.4344| Source | Tool | Description |45|--------|------|-------------|46| PubChem | `PubChem_search_compound` | PubChem API |47| PubChem | `PubChem_get_compound` | PubChem API |48| ChEMBL | `ChEMBL_search_compound` | ChEMBL API |49| ChEMBL | `ChEMBL_get_activity` | ChEMBL API |5051## Quality Gates5253- [ ] The selected method matches the scientific question and stated assumptions.54- [ ] Outputs are reproducible, saved to files, and traceable from inputs to conclusions.55- [ ] Missing data, uncertainty, bias, and hard limits are made explicit.56- [ ] `report.md` and `logs/process-log.jsonl` reference the generated artifacts.57- [ ] No essential result remains chat-only.5859If any gate fails: identify the specific failing check, fix the issue, and re-validate before proceeding.6061## Gotchas6263- SMILES strings may represent different stereoisomers. Canonicalize SMILES before database lookups64- Assay results from different sources use different activity units (IC50, Ki, EC50). Standardize before comparison65- Chemical similarity metrics (Tanimoto, Dice) give different rankings. Report fingerprint and metric used6667## Validation Loop68691. Execute analysis and generate outputs702. Check:71 - Method selection matches the research question and stated assumptions72 - All outputs are saved to files (no chat-only results)73 - Limitations and uncertainty are explicitly stated74 - `logs/process-log.jsonl` is updated with execution trace753. If any check fails:76 - Identify the failing gate77 - Fix the specific issue78 - Re-run validation794. Proceed only after all gates pass