ai4protein
- 36 skills
- 0 followers
- 9 hours ago last updated
- ▌ Fda · ai4protein bundleQuery openFDA via VenusFactory for drugs, devices, adverse events, recalls, and regulatory submissions (510k, PMA). Use when the user needs FDA pharmacovigilance, labeling, NDC/UNII, or openFDA analytics. Do NOT use for ChEMBL bioactivity (chembl_database) or general biomedical literature (pubmed).
- ▌ Arxiv · ai4proteinarXiv preprint server — keyword search the official API and download papers as PDF / HTML / source tarball. Use whenever the user mentions an arXiv ID (e.g. 2106.04559) or wants preprints on a topic in CS / physics / math / quantitative biology. For biomedical preprints specifically, prefer biorxiv skill.
- ▌ Pymol · ai4protein bundleHeadless PyMOL rendering of protein structures (PNG + PSE session) and structural superposition with RMSD. Use to produce static publication-style images, color a structure by pLDDT/B-factor/chain/secondary-structure, or compare two structures by cealign. Do NOT use for interactive 3D exploration (use MolstarViewer in the frontend), docking, MD simulation, or sequence-only analysis.
- ▌ Rdkit · ai4protein bundleRDKit cheminformatics via VenusFactory bioinfo scripts and agent_generated_code. Use for SMILES/SDF, descriptors, fingerprints, substructure filters, similarity. Do NOT use for ChEMBL bioactivity download (chembl_database) or openFDA (fda). No dedicated rdkit_* LangChain tool — run scripts or import modules.
- ▌ Pubmed · ai4proteinPubMed — NCBI's biomedical literature database (>35M citations). Keyword search inline, or batch-fetch full title + structured abstract + authors + DOI for a known list of PMIDs. Use for medical / biological literature lookup, citation resolution, or building an abstract corpus for downstream NLP. Honors NCBI_API_KEY env var.
- ▌ Biorxiv · ai4proteinbioRxiv & medRxiv — biology / medicine preprint servers. Search by keyword + date window, fetch a specific preprint's full metadata (with all version history + abstract + JATS XML link) by DOI. Use for cutting-edge biology/medicine work that hasn't gone through peer review yet, or to find the original preprint version of a paper you only have the DOI for.
- ▌ Seaborn · ai4protein bundleSeaborn statistical plots for exploratory analysis via agent_generated_code. Use for quick relational/distribution/categorical charts. Do NOT use for submission-grade Nature figures (nature_figure) or low-level artists control (matplotlib).
- ▌ Openalex · ai4proteinOpenAlex — free, comprehensive scholarly graph (works, authors, sources/journals, institutions, topics, concepts, funders). Search papers by keyword/filter/sort, fetch a single entity by ID, look up author profiles, institutions, citation networks. Use whenever the user asks about papers, citation counts, author publications, institution outputs, or research topics. Lower-friction alternative to Semantic Scholar / Web of Science / Scopus.
- ▌ Biopython · ai4protein bundleBiopython guidance for sequence I/O, alignments, Bio.PDB, and Entrez parsing. Use for custom bioinformatics code via agent_generated_code. Prefer VenusFactory download tools for NCBI/UniProt/AlphaFold bulk fetches so large payloads stay on disk. Do NOT use as a substitute for protein_sequence_similarity_search or alphafold_database hub tools.
- ▌ Ncbi Gene · ai4protein bundleQuery NCBI Gene via E-utilities/Datasets API. Search by symbol/ID, retrieve gene info (RefSeqs, GO, locations, phenotypes), batch lookups, for gene annotation and functional analysis.
- ▌ Matplotlib · ai4protein bundleMatplotlib OO/pyplot guidance for custom plots via agent_generated_code. Use for fine-grained control. Prefer nature_figure for manuscript figures and seaborn for quick statistical EDA.
- ▌
- ▌ Ncbi Clinvar · ai4protein bundleQuery NCBI ClinVar for variant clinical significance. Search by gene/condition/CLNSIG, interpret pathogenicity, use E-utilities or FTP; annotate VCFs. Use project tools in src.tools.database.ncbi.
- ▌ Kegg Database · ai4protein bundleKEGG REST access via VenusFactory download tools (academic use). Use for pathway/gene/compound lookups, ID conversion, and DDI. Do NOT use for PPI networks (string_database) or enzyme kinetics (brenda_database). Non-academic use of KEGG requires a commercial license.
- ▌ Nature Figure · ai4protein bundleSubmission-grade Nature/high-impact journal figure workflow for Python or R. Use whenever the user asks to create, revise, audit, or polish manuscript figures, multi-panel scientific plots, figures4papers-style matplotlib plots, or journal-ready SVG/PDF/TIFF outputs, especially for Nature-family or other high-impact journals. Before plotting, define the figure's conclusion, evidence logic, export needs, and review risks. If the user has not chosen Python or R, ask "Python or R?" and stop. Use only the selected backend for figure generation, previewing, exporting, and QA. Supports matplotlib/seaborn and ggplot2/patchwork/ComplexHeatmap. Not for dashboards or Illustrator/Figma-first infographics. Also trigger on general academic-writing figure needs even without the word "Nature", such as making figures/plots for a paper, scientific/academic plotting, data visualization for a manuscript, and Chinese phrasings like 论文配图、学术写作配图、科研绘图、科研作图、画图、作图、出图、论文图表、可视化.
- ▌ Ncbi Sequence · ai4proteinNCBI E-utilities for biological sequences — fetch protein/nucleotide FASTA by accession, run BLAST, translate CDS to protein, search NCBI Protein by gene+organism. Use when the user provides an NCBI accession (NP_, XP_, NM_, NR_, etc.), asks for a sequence by gene name + species, or needs to translate a coding sequence. Don't use for ClinVar variants (use ncbi_clinvar) or gene metadata lookup (use ncbi_gene).
- ▌ Rcsb Database · ai4proteinRCSB Protein Data Bank (PDB) — experimentally determined 3D biomolecular structures. Search by full-text/sequence/structure/attribute, fetch entry metadata, download coordinate files (PDB/mmCIF). Use when the user provides a PDB ID, asks for structures of a protein, wants to find similar structures by sequence, or needs experimental (not predicted) coordinates. Don't use for AlphaFold predictions (use alphafold_database).
- ▌ Nature Writing · ai4protein bundleDraft, restructure, or plan Nature-style manuscript sections from author-provided claims, results, figures, notes, or Chinese drafts. Use when the user wants to write or rebuild an abstract, introduction, related-work, method, experiments, discussion, conclusion, title, or full manuscript argument rather than only polish finished prose. Also trigger on general academic-writing requests even without the word "Nature", such as writing a paper from scratch, drafting a manuscript/section, structuring a paper, and Chinese phrasings like 学术写作、科研写作、论文写作、写论文、写paper、SCI写作、帮我写论文、搭论文框架、起草论文、写引言/摘要/讨论.
- ▌ Brenda Database · ai4protein bundleBRENDA enzyme kinetics via VenusFactory download tools (SOAP). Use for Km/kcat, reactions, organism comparison, environmental optima by EC number. Do NOT use for pathway maps alone (kegg_database) or protein sequence fetch (uniprot_database). Requires BRENDA_EMAIL and BRENDA_PASSWORD.
- ▌ Chembl Database · ai4protein bundleChEMBL bioactive molecules and drugs via VenusFactory download tools. Use for molecule/drug by ID, similarity/substructure by SMILES, SAR starting points. Do NOT use for openFDA regulatory data (fda) or RDKit-only local chemistry (rdkit).
- ▌ String Database · ai4protein bundleSTRING PPI networks and enrichment via VenusFactory download tools. Use for interaction networks, partners, GO/KEGG enrichment, homology across 5000+ species. Do NOT use for sequence homology search (protein_sequence_similarity_search) or KEGG pathway entries alone (kegg_database).
- ▌ Nature Polishing · ai4protein bundlePolish, restructure, or translate academic prose into Nature-leaning English using writing-strategy principles, curated Nature/Nature Communications article patterns, and phrase-level support from Academic Phrasebank. Use whenever the user asks to polish a manuscript paragraph, abstract, introduction, results, discussion, conclusion, title, methods section, or Chinese academic draft for publication-quality English. Also covers LaTeX layout/typesetting (排版) fixes — loose or sparse pages, stranded section headings, figures that don't fill the page or split across pages, "Float too large", multi-panel figure arrangement, and Supplementary Information that looks empty — via references/latex-layout.md. Also trigger on general academic/scientific writing requests even without the word "Nature", including academic writing, scientific writing, SCI/paper writing, English manuscript polishing, language editing, proofreading, and Chinese phrasings such as 学术写作、科研写作、论文润色、写paper、SCI写作、英文论文润色、语言润色、润色、改写、学术英语、英文写作.
- ▌ Uniprot Database · ai4proteinUniProt — protein sequence, function, taxonomy, cross-references. Search proteins by query, retrieve a UniProt entry, map IDs between databases (PDB↔UniProt etc.), pull FASTA sequence, fetch metadata, run SPARQL against sparql.uniprot.org. Use whenever the user mentions a UniProt accession (e.g. P04637), asks for protein function/sequence/family info, or needs cross-DB ID mapping. Don't use for AlphaFold structures (use alphafold_database) or PDB structures (use rcsb_database).
- ▌ Alphafold Database · ai4protein bundleAlphaFold DB structures and confidence analytics via VenusFactory tools. Use when the user needs predicted structures by UniProt ID, pLDDT/PAE analysis, or PDB/mmCIF download. Do NOT use for experimental PDB (rcsb_database), local ESMFold without UniProt (predict_structure_esmfold / protein_structure_pipeline), or sequence annotation (uniprot_database).
- ▌ Structure File Prep · ai4proteinStructure/sequence file preparation with VenusFactory file tools. Use for FASTA parsing, PDB chain extraction, PDB↔mmCIF conversion (MAXIT), apo checks, batch PDB→FASTA, and UniProt ID from RCSB metadata. Do NOT use for structure prediction (protein_structure_pipeline) or homology search.
- ▌ Hpa Expression Context · ai4proteinHuman Protein Atlas expression and localization via VenusFactory download tools. Use when the user needs tissue expression, subcellular location, single-cell type, blood expression, or protein summary by gene symbol for therapeutic/target context. Do NOT use for mouse/non-human expression atlases or PPI networks (string_database).
- ▌ Workflow Skill Creator · ai4protein bundleDistills a completed user workflow or interaction into a reusable VenusFactory agent skill. Use when the user says "make this a skill", "create a skill from what we just did", "package this workflow" or similar. Adapts the workflow into the VenusFactory tools wiring + SKILL.md pattern. Do not use for creating skills from scratch without an existing workflow.
- ▌ Venus Finetune Workflow · ai4proteinFine-tune and run custom protein models on VenusFactory (CSV/HF → config → train → predict). Use when the user brings labeled sequences, wants adapter training (ProtT5/ESM2/Ankh/QLoRA notes), or batch inference with a trained config. Do NOT use for zero-shot mutation without labels (zero_shot_mutation_workflow) or built-in function heads already covered by predict_protein_function.
- ▌ Interpro Domain Annotation · ai4proteinInterPro domain/family annotation via VenusFactory download tools. Use when the user needs domain boundaries, family membership, or UniProt→InterPro annotations for engineering target selection. Do NOT use for pathway enrichment (string_database / kegg_database) or kinetic parameters (brenda_database).
- ▌ Protein Structure Pipeline · ai4proteinProtein structure obtain → confidence → visualize pipeline. Use when the user needs a 3D structure from sequence or UniProt ID, AlphaFold/ESMFold retrieval, pLDDT/PAE analysis, or structure rendering. Do NOT use for mutation ranking (zero_shot_mutation_workflow), FoldSeek search (foldseek_structural_similarity), or experimental PDB-only metadata without structure needs (rcsb_database alone may suffice).
- ▌ Protein Property Prediction · ai4proteinPhysicochemical properties, surface/SS features, and finetuned protein/residue function prediction. Use when the user asks for solubility, optimal temperature, activity/binding/conserved sites, RSA/SASA/secondary structure, or property tables from FASTA/PDB. Do NOT use for zero-shot mutation ranking (zero_shot_mutation_workflow) or custom model training (train tools / future finetune workflow).
- ▌ Proteinmpnn Design Workflow · ai4proteinProteinMPNN inverse folding: design or score sequences on a fixed backbone. Use when the user wants sequence design from PDB, interface/binder design, homomer symmetry, or fixed catalytic residues. Do NOT use for zero-shot mutation ranking on a wild-type sequence (zero_shot_mutation_workflow) or de novo fold hallucination without a backbone.
- ▌ Zero Shot Mutation Workflow · ai4proteinZero-shot mutation engineering with VenusFactory PLMs. Use when the user wants beneficial mutations, directed evolution candidates, or stability/fitness ranking from a FASTA sequence or PDB structure. Do NOT use for ProteinMPNN inverse folding (proteinmpnn_design_workflow), sequence homology search (protein_sequence_similarity_search), or experimental wet-lab protocols alone.
- ▌ Foldseek Structural Similarity · ai4proteinFoldSeek structural similarity search against PDB with optional protected-region masking. Use when the user has a PDB and wants fold-level homologs, structural neighbors, or to protect an active site while searching. Do NOT use for sequence BLAST/MMseqs2 (protein_sequence_similarity_search) or MSA (clustalo_msa).
- ▌ Protein Engineering Hypothesis · ai4proteinEvidence-bounded hypothesis and experiment planning for protein engineering. Use when the user asks what to mutate next, how to prioritize variants, how to falsify a mechanism, or how to design a directed-evolution round. Do NOT invent wet-lab results; chain VenusFactory tools for computational evidence and label model scores as hypotheses.
- ▌ Protein Sequence Similarity Search · ai4proteinFind homologous protein sequences from a query sequence using MMseqs2 (fast, ColabFold web API) or BLAST (comprehensive, EBI). Use when the user provides a protein sequence or FASTA file and wants homologs, function inference by sequence similarity, or input for an MSA. Do NOT use for structural similarity (use foldseek) or DNA/RNA queries.