← all publishers

ai4protein

@ai4protein source repo

36 published skills

  1. Fda · ai4protein bundle
    Query openFDA via VenusFactory for drugs, devices, adverse events, recalls, and regulatory submissions (510k, PMA). Use when the user needs FDA pharmacovigilance, labeling, NDC/UNII, or openFDA analytics. Do NOT use for ChEMBL bioactivity (chembl_database) or general biomedical literature (pubmed).
    0
    installs
  2. Arxiv · ai4protein
    arXiv preprint server — keyword search the official API and download papers as PDF / HTML / source tarball. Use whenever the user mentions an arXiv ID (e.g. 2106.04559) or wants preprints on a topic in CS / physics / math / quantitative biology. For biomedical preprints specifically, prefer biorxiv skill.
    0
    installs
  3. Pymol · ai4protein bundle
    Headless PyMOL rendering of protein structures (PNG + PSE session) and structural superposition with RMSD. Use to produce static publication-style images, color a structure by pLDDT/B-factor/chain/secondary-structure, or compare two structures by cealign. Do NOT use for interactive 3D exploration (use MolstarViewer in the frontend), docking, MD simulation, or sequence-only analysis.
    0
    installs
  4. Rdkit · ai4protein bundle
    RDKit cheminformatics via VenusFactory bioinfo scripts and agent_generated_code. Use for SMILES/SDF, descriptors, fingerprints, substructure filters, similarity. Do NOT use for ChEMBL bioactivity download (chembl_database) or openFDA (fda). No dedicated rdkit_* LangChain tool — run scripts or import modules.
    0
    installs
  5. Pubmed · ai4protein
    PubMed — NCBI's biomedical literature database (>35M citations). Keyword search inline, or batch-fetch full title + structured abstract + authors + DOI for a known list of PMIDs. Use for medical / biological literature lookup, citation resolution, or building an abstract corpus for downstream NLP. Honors NCBI_API_KEY env var.
    0
    installs
  6. Biorxiv · ai4protein
    bioRxiv & medRxiv — biology / medicine preprint servers. Search by keyword + date window, fetch a specific preprint's full metadata (with all version history + abstract + JATS XML link) by DOI. Use for cutting-edge biology/medicine work that hasn't gone through peer review yet, or to find the original preprint version of a paper you only have the DOI for.
    0
    installs
  7. Seaborn · ai4protein bundle
    Seaborn statistical plots for exploratory analysis via agent_generated_code. Use for quick relational/distribution/categorical charts. Do NOT use for submission-grade Nature figures (nature_figure) or low-level artists control (matplotlib).
    0
    installs
  8. Openalex · ai4protein
    OpenAlex — free, comprehensive scholarly graph (works, authors, sources/journals, institutions, topics, concepts, funders). Search papers by keyword/filter/sort, fetch a single entity by ID, look up author profiles, institutions, citation networks. Use whenever the user asks about papers, citation counts, author publications, institution outputs, or research topics. Lower-friction alternative to Semantic Scholar / Web of Science / Scopus.
    0
    installs
  9. Biopython · ai4protein bundle
    Biopython guidance for sequence I/O, alignments, Bio.PDB, and Entrez parsing. Use for custom bioinformatics code via agent_generated_code. Prefer VenusFactory download tools for NCBI/UniProt/AlphaFold bulk fetches so large payloads stay on disk. Do NOT use as a substitute for protein_sequence_similarity_search or alphafold_database hub tools.
    0
    installs
  10. Ncbi Gene · ai4protein bundle
    Query NCBI Gene via E-utilities/Datasets API. Search by symbol/ID, retrieve gene info (RefSeqs, GO, locations, phenotypes), batch lookups, for gene annotation and functional analysis.
    0
    installs
  11. Matplotlib · ai4protein bundle
    Matplotlib OO/pyplot guidance for custom plots via agent_generated_code. Use for fine-grained control. Prefer nature_figure for manuscript figures and seaborn for quick statistical EDA.
    0
    installs
  12. Clustalo Msa · ai4protein
    Clustal Omega MSA (EBI)
    0
    installs
  13. Ncbi Clinvar · ai4protein bundle
    Query NCBI ClinVar for variant clinical significance. Search by gene/condition/CLNSIG, interpret pathogenicity, use E-utilities or FTP; annotate VCFs. Use project tools in src.tools.database.ncbi.
    0
    installs
  14. Kegg Database · ai4protein bundle
    KEGG REST access via VenusFactory download tools (academic use). Use for pathway/gene/compound lookups, ID conversion, and DDI. Do NOT use for PPI networks (string_database) or enzyme kinetics (brenda_database). Non-academic use of KEGG requires a commercial license.
    0
    installs
  15. Nature Figure · ai4protein bundle
    Submission-grade Nature/high-impact journal figure workflow for Python or R. Use whenever the user asks to create, revise, audit, or polish manuscript figures, multi-panel scientific plots, figures4papers-style matplotlib plots, or journal-ready SVG/PDF/TIFF outputs, especially for Nature-family or other high-impact journals. Before plotting, define the figure's conclusion, evidence logic, export needs, and review risks. If the user has not chosen Python or R, ask "Python or R?" and stop. Use only the selected backend for figure generation, previewing, exporting, and QA. Supports matplotlib/seaborn and ggplot2/patchwork/ComplexHeatmap. Not for dashboards or Illustrator/Figma-first infographics. Also trigger on general academic-writing figure needs even without the word "Nature", such as making figures/plots for a paper, scientific/academic plotting, data visualization for a manuscript, and Chinese phrasings like 论文配图、学术写作配图、科研绘图、科研作图、画图、作图、出图、论文图表、可视化.
    0
    installs
  16. Ncbi Sequence · ai4protein
    NCBI E-utilities for biological sequences — fetch protein/nucleotide FASTA by accession, run BLAST, translate CDS to protein, search NCBI Protein by gene+organism. Use when the user provides an NCBI accession (NP_, XP_, NM_, NR_, etc.), asks for a sequence by gene name + species, or needs to translate a coding sequence. Don't use for ClinVar variants (use ncbi_clinvar) or gene metadata lookup (use ncbi_gene).
    0
    installs
  17. Rcsb Database · ai4protein
    RCSB Protein Data Bank (PDB) — experimentally determined 3D biomolecular structures. Search by full-text/sequence/structure/attribute, fetch entry metadata, download coordinate files (PDB/mmCIF). Use when the user provides a PDB ID, asks for structures of a protein, wants to find similar structures by sequence, or needs experimental (not predicted) coordinates. Don't use for AlphaFold predictions (use alphafold_database).
    0
    installs
  18. Nature Writing · ai4protein bundle
    Draft, restructure, or plan Nature-style manuscript sections from author-provided claims, results, figures, notes, or Chinese drafts. Use when the user wants to write or rebuild an abstract, introduction, related-work, method, experiments, discussion, conclusion, title, or full manuscript argument rather than only polish finished prose. Also trigger on general academic-writing requests even without the word "Nature", such as writing a paper from scratch, drafting a manuscript/section, structuring a paper, and Chinese phrasings like 学术写作、科研写作、论文写作、写论文、写paper、SCI写作、帮我写论文、搭论文框架、起草论文、写引言/摘要/讨论.
    0
    installs
  19. Brenda Database · ai4protein bundle
    BRENDA enzyme kinetics via VenusFactory download tools (SOAP). Use for Km/kcat, reactions, organism comparison, environmental optima by EC number. Do NOT use for pathway maps alone (kegg_database) or protein sequence fetch (uniprot_database). Requires BRENDA_EMAIL and BRENDA_PASSWORD.
    0
    installs
  20. Chembl Database · ai4protein bundle
    ChEMBL bioactive molecules and drugs via VenusFactory download tools. Use for molecule/drug by ID, similarity/substructure by SMILES, SAR starting points. Do NOT use for openFDA regulatory data (fda) or RDKit-only local chemistry (rdkit).
    0
    installs
  21. String Database · ai4protein bundle
    STRING PPI networks and enrichment via VenusFactory download tools. Use for interaction networks, partners, GO/KEGG enrichment, homology across 5000+ species. Do NOT use for sequence homology search (protein_sequence_similarity_search) or KEGG pathway entries alone (kegg_database).
    0
    installs
  22. Nature Polishing · ai4protein bundle
    Polish, restructure, or translate academic prose into Nature-leaning English using writing-strategy principles, curated Nature/Nature Communications article patterns, and phrase-level support from Academic Phrasebank. Use whenever the user asks to polish a manuscript paragraph, abstract, introduction, results, discussion, conclusion, title, methods section, or Chinese academic draft for publication-quality English. Also covers LaTeX layout/typesetting (排版) fixes — loose or sparse pages, stranded section headings, figures that don't fill the page or split across pages, "Float too large", multi-panel figure arrangement, and Supplementary Information that looks empty — via references/latex-layout.md. Also trigger on general academic/scientific writing requests even without the word "Nature", including academic writing, scientific writing, SCI/paper writing, English manuscript polishing, language editing, proofreading, and Chinese phrasings such as 学术写作、科研写作、论文润色、写paper、SCI写作、英文论文润色、语言润色、润色、改写、学术英语、英文写作.
    0
    installs
  23. Uniprot Database · ai4protein
    UniProt — protein sequence, function, taxonomy, cross-references. Search proteins by query, retrieve a UniProt entry, map IDs between databases (PDB↔UniProt etc.), pull FASTA sequence, fetch metadata, run SPARQL against sparql.uniprot.org. Use whenever the user mentions a UniProt accession (e.g. P04637), asks for protein function/sequence/family info, or needs cross-DB ID mapping. Don't use for AlphaFold structures (use alphafold_database) or PDB structures (use rcsb_database).
    0
    installs
  24. Alphafold Database · ai4protein bundle
    AlphaFold DB structures and confidence analytics via VenusFactory tools. Use when the user needs predicted structures by UniProt ID, pLDDT/PAE analysis, or PDB/mmCIF download. Do NOT use for experimental PDB (rcsb_database), local ESMFold without UniProt (predict_structure_esmfold / protein_structure_pipeline), or sequence annotation (uniprot_database).
    0
    installs
  25. Structure File Prep · ai4protein
    Structure/sequence file preparation with VenusFactory file tools. Use for FASTA parsing, PDB chain extraction, PDB↔mmCIF conversion (MAXIT), apo checks, batch PDB→FASTA, and UniProt ID from RCSB metadata. Do NOT use for structure prediction (protein_structure_pipeline) or homology search.
    0
    installs
  26. Hpa Expression Context · ai4protein
    Human Protein Atlas expression and localization via VenusFactory download tools. Use when the user needs tissue expression, subcellular location, single-cell type, blood expression, or protein summary by gene symbol for therapeutic/target context. Do NOT use for mouse/non-human expression atlases or PPI networks (string_database).
    0
    installs
  27. Workflow Skill Creator · ai4protein bundle
    Distills a completed user workflow or interaction into a reusable VenusFactory agent skill. Use when the user says "make this a skill", "create a skill from what we just did", "package this workflow" or similar. Adapts the workflow into the VenusFactory tools wiring + SKILL.md pattern. Do not use for creating skills from scratch without an existing workflow.
    0
    installs
  28. Venus Finetune Workflow · ai4protein
    Fine-tune and run custom protein models on VenusFactory (CSV/HF → config → train → predict). Use when the user brings labeled sequences, wants adapter training (ProtT5/ESM2/Ankh/QLoRA notes), or batch inference with a trained config. Do NOT use for zero-shot mutation without labels (zero_shot_mutation_workflow) or built-in function heads already covered by predict_protein_function.
    0
    installs
  29. Interpro Domain Annotation · ai4protein
    InterPro domain/family annotation via VenusFactory download tools. Use when the user needs domain boundaries, family membership, or UniProt→InterPro annotations for engineering target selection. Do NOT use for pathway enrichment (string_database / kegg_database) or kinetic parameters (brenda_database).
    0
    installs
  30. Protein Structure Pipeline · ai4protein
    Protein structure obtain → confidence → visualize pipeline. Use when the user needs a 3D structure from sequence or UniProt ID, AlphaFold/ESMFold retrieval, pLDDT/PAE analysis, or structure rendering. Do NOT use for mutation ranking (zero_shot_mutation_workflow), FoldSeek search (foldseek_structural_similarity), or experimental PDB-only metadata without structure needs (rcsb_database alone may suffice).
    0
    installs
  31. Protein Property Prediction · ai4protein
    Physicochemical properties, surface/SS features, and finetuned protein/residue function prediction. Use when the user asks for solubility, optimal temperature, activity/binding/conserved sites, RSA/SASA/secondary structure, or property tables from FASTA/PDB. Do NOT use for zero-shot mutation ranking (zero_shot_mutation_workflow) or custom model training (train tools / future finetune workflow).
    0
    installs
  32. Proteinmpnn Design Workflow · ai4protein
    ProteinMPNN inverse folding: design or score sequences on a fixed backbone. Use when the user wants sequence design from PDB, interface/binder design, homomer symmetry, or fixed catalytic residues. Do NOT use for zero-shot mutation ranking on a wild-type sequence (zero_shot_mutation_workflow) or de novo fold hallucination without a backbone.
    0
    installs
  33. Zero Shot Mutation Workflow · ai4protein
    Zero-shot mutation engineering with VenusFactory PLMs. Use when the user wants beneficial mutations, directed evolution candidates, or stability/fitness ranking from a FASTA sequence or PDB structure. Do NOT use for ProteinMPNN inverse folding (proteinmpnn_design_workflow), sequence homology search (protein_sequence_similarity_search), or experimental wet-lab protocols alone.
    0
    installs
  34. Foldseek Structural Similarity · ai4protein
    FoldSeek structural similarity search against PDB with optional protected-region masking. Use when the user has a PDB and wants fold-level homologs, structural neighbors, or to protect an active site while searching. Do NOT use for sequence BLAST/MMseqs2 (protein_sequence_similarity_search) or MSA (clustalo_msa).
    0
    installs
  35. Protein Engineering Hypothesis · ai4protein
    Evidence-bounded hypothesis and experiment planning for protein engineering. Use when the user asks what to mutate next, how to prioritize variants, how to falsify a mechanism, or how to design a directed-evolution round. Do NOT invent wet-lab results; chain VenusFactory tools for computational evidence and label model scores as hypotheses.
    0
    installs
  36. Protein Sequence Similarity Search · ai4protein
    Find homologous protein sequences from a query sequence using MMseqs2 (fast, ColabFold web API) or BLAST (comprehensive, EBI). Use when the user provides a protein sequence or FASTA file and wants homologs, function inference by sequence similarity, or input for an MSA. Do NOT use for structural similarity (use foldseek) or DNA/RNA queries.
    0
    installs