Research & Search
Research agent skills teach AI agents to gather and synthesize information properly: literature reviews, competitive analysis, web research with citations, and structured summaries. Each SKILL.md encodes a method, not just a prompt, so results stay consistent across runs.
-
phoroth Skill Scientific WritingThis is the core skill for the deep research and writing tool—combining AI-driven deep research with well-formatted written outputs. Every document produced is backed by comprehensive literature search and verified citations through the research-lookup skill.
3 -
phoroth Skill Citation ManagementManage citations systematically throughout the research and writing process.
3 -
phoroth Bundle Xvary Stock ResearchThesis-driven equity analysis from public SEC EDGAR and market data; /analyze, /score, /compare workflows with bundled Python tools (Claude Code, Cursor, Codex).
3 -
phoroth Bundle Apify Market ResearchAnalyze market conditions, geographic opportunities, pricing, consumer behavior, and product validation across Google Maps, Facebook, Instagram, Booking.com, and TripAdvisor.
3 -
phoroth Skill Efficient Web ResearchProtocol for token-efficient web research. Use when accessing URLs, GitHub repos, or running search queries. Prevents full-page fetching waste.
3 -
phoroth Bundle Pubmed DatabaseSearch PubMed for scientific literature, including published clinical trials. Fetch abstracts and full text. Link published research to biological databases (gene, protein, nucleotide, PubChem) to discover associations between papers and specific compounds or genes. Verify medical spelling, match raw citations, and cache result sets for bulk processing. Interfaces NCBI E-utilities and PMC BioC APIs.
3 -
phoroth Bundle Comprehensive Review Pr EnhanceGenerate structured PR descriptions from diffs, add review checklists, risk assessments, and test coverage summaries. Use when the user says "write a PR description", "improve this PR", "summarize my changes", "PR review", "pull request", or asks to document a diff for reviewers.
3 -
phoroth Skill Uspto DatabaseUSPTO patent and trademark data workflow for official record lookup, PatentSearch queries, TSDR checks, assignment data, and reproducible IP research logs.
3 -
phoroth Bundle Literature Search BiorxivBrowse, filter, and download life sciences, biology, and medical preprints from bioRxiv and medRxiv. Supports fetching paper metadata by DOI, and browsing by date range with category and keyword filters. Keyword filtering is local, so date ranges MUST be narrow (1-4 weeks) with a category to prevent timeouts.
3 -
phoroth Bundle Literature Search OpenalexQuery the OpenAlex scholarly database for research papers, authors, institutions, topics, sources, publishers, funders, geo-locations, and keywords. Use when searching academic papers, resolving DOIs, downloading open-access PDFs, finding an author's publications, aggregating bibliometric data (citation counts, h-index, impact factor), exploring the research taxonomies, or performing DOI lookups.
3 -
phoroth Bundle Literature Search EuropepmcSearch Europe PMC for scientific literature and download open-access full texts and PDFs. Retrieve full-text XML/plain text by PMCID, get citation lists and bibliography.
3 -
qhjqhj00 Skill Bwor EvalEvaluates LLMs' ability to automate operations research problem solving through mathematical modeling, code generation, and solver-based optimization. It probes whether reasoning agents can correctly translate natural language OR problems into executable models and compute optimal solutions. Use when the user wants to benchmark on BWOR, or asks about evaluating this task. Reports accuracy.
3 -
qhjqhj00 Skill Care EvalEvaluates information extraction systems on clinical literature for fine-grained extraction of experimental findings, including entities, attributes, and complex n-ary relations with discontinuous spans and variable arity. Use when the user wants to benchmark on CARE, or asks about evaluating this task. Reports relaxed overlap F1.
3 -
qhjqhj00 Skill Iclr PointsQuantifies the average research effort required to produce one publication at top-tier conferences across 27 computer science subfields. It enables cross-area comparisons of faculty productivity and publication effort by normalizing faculty headcounts against publication counts. Use when the user has predictions and gold and needs to compute ICLR points.
3 -
qhjqhj00 Skill Litqa2 EvalEvaluates an agentic system's ability to retrieve relevant scientific literature, re-rank passages, and answer multiple-choice questions based on the retrieved text. It probes retrieval coverage, passage localization, and question-answering accuracy under information loss constraints. Use when the user wants to benchmark on LitQA2, or asks about evaluating this task. Reports Accuracy.
3 -
qhjqhj00 Skill Kbp Auc EvalEvaluates the ability of distantly supervised relation extraction and knowledge base validation systems to correctly predict and rank triples in web-scale knowledge graphs. It probes how well global graph structure and confidence scoring can refine noisy extractions and reduce logical inconsistencies. Use when the user wants to benchmark on NYT-FB, CC-DBP, NELL-165, or asks about evaluating this task. Reports AUC.
3 -
qhjqhj00 Skill Medhelm EvalEvaluates large language models on a comprehensive taxonomy of real-world clinical workflows, covering tasks like clinical note generation, patient communication, medical research assistance, clinical decision support, and administration/workflow. It probes both closed-ended factual/reasoning tasks and open-ended free-text generation capabilities in medical domains. Use when the user wants to benchmark on MedHELM, or asks about evaluating this task. Reports Macro-average performance.
3 -
qhjqhj00 Skill Openxai EvalEvaluates the faithfulness, stability, and fairness of post-hoc feature attribution explanation methods (e.g., LIME, SHAP, gradient-based) on tabular datasets to enable reproducible and transparent comparisons. Use when the user wants to benchmark on Popular tabular datasets for XAI and fairness research, or asks about evaluating this task. Reports faithfulness.
3 -
qhjqhj00 Skill Scicode EvalProbes large language models' ability to perform scientific reasoning, domain-specific knowledge recall, and code synthesis on real-world research problems. It evaluates performance on both decomposed subproblems and full main problems under varying conditions of background knowledge and context carry-over. Use when the user wants to benchmark on SciCode, or asks about evaluating this task. Reports pass@1.
3 -
qhjqhj00 Skill Wildsci EvalEvaluates language models' scientific reasoning capabilities by testing their ability to answer domain-specific multiple-choice questions derived from peer-reviewed literature and established scientific benchmarks. Use when the user wants to benchmark on WildSci-Val, GPQA-Aug, SuperGPQA, MMLU-Pro, or asks about evaluating this task. Reports accuracy.
3 -
qhjqhj00 Skill Captrack EvalThis evaluation framework probes systematic capability drift and forgetting in large language models after post-training. It measures degradation across latent competence (knowledge, reasoning), default behavioral preferences (refusal, verbosity, formatting), and protocol compliance (instruction following, tool use, citation) in legal and medical domains. Use when the user wants to benchmark on CapTrack Evaluation Suite, or asks about evaluating this task. Reports average forgetting.
3 -
qhjqhj00 Skill Forc2025 EvalEvaluates hierarchical multi-label classification of academic papers into a taxonomy of 170 research fields. It tests zero-shot/few-shot prompting and weakly-labeled data integration for field-of-research prediction. Use when the user wants to benchmark on FoRC4CL 2025, or asks about evaluating this task. Reports Micro-F1.
3 -
qhjqhj00 Skill Infoseek EvalEvaluates LLMs on complex, multi-step reasoning and agentic search tasks, including single-hop and multi-hop question answering as well as deep research benchmarks requiring web search and synthesis. Use when the user wants to benchmark on NQ, TQA, PopQA, HQA, 2Wiki, MSQ, Bamb, BrowseComp-Plus, or asks about evaluating this task. Reports Accuracy.
3 -
qhjqhj00 Skill Innoeval EvalEvaluates an AI system's ability to assess scientific research ideas across classification, selection, ranking, and comparison tasks. It probes knowledge-grounded reasoning and multi-perspective evaluation aligned with human expert judgments and conference acceptance standards. Use when the user wants to benchmark on D_point, D_group, D_pair, or asks about evaluating this task. Reports Accuracy.
3 -
qhjqhj00 Skill Citebench EvalEvaluates the capability of citation recommendation models to identify relevant academic references given local citation contexts. It probes robustness across varying contextual features, including context length, reference position, academic field, publication year, citation count, and part-of-speech tags. Use when the user wants to benchmark on S2ORC/S2AG Diagnostic Datasets, or asks about evaluating this task. Reports MRR.
3 -
qhjqhj00 Skill Exp Bench EvalEvaluates AI agents' end-to-end capability to conduct real AI research experiments, including designing methodologies, implementing code, executing experiments, and drawing conclusions. Use when the user wants to benchmark on EXP-Bench, or asks about evaluating this task. Reports All·E✓.
3 -
qhjqhj00 Skill Lab Bench EvalThis benchmark evaluates large language models' capabilities in biology research, including literature retrieval, figure/table interpretation, database querying, protocol troubleshooting, and DNA/protein sequence manipulation. It probes whether models can perform multi-step, tool-dependent scientific reasoning or rely on memorization and heuristic guesswork. Use when the user wants to benchmark on LAB-Bench, or asks about evaluating this task. Reports accuracy.
3 -
qhjqhj00 Skill Oag Bench EvalThis benchmark evaluates models on academic graph mining tasks, including author disambiguation, scholar profiling, entity tagging, academic recommendation, question answering, paper source tracing, and influence prediction. It probes the ability of graph neural networks, retrieval systems, and LLMs to handle structured academic data, extract attributes from long texts, and perform ranking or prediction tasks on citation networks. Use when the user wants to benchmark on OAG-Bench, or asks about evaluating this task. Reports MAP.
3 -
qhjqhj00 Skill Papermind EvalEvaluates multimodal LLMs' ability to perform integrated agentic reasoning and critical assessment over scientific papers. It probes capabilities in multimodal grounding, experimental interpretation, cross-source evidence synthesis via tool use, and critical evaluation of research claims. Use when the user wants to benchmark on PaperMind, or asks about evaluating this task. Reports F1 score.
3 -
qhjqhj00 Skill Airs Bench EvalEvaluates AI research agents across the full scientific lifecycle, including idea generation, experiment design, and iterative refinement. Agents must generate and execute code to train models on specified datasets without baseline code, testing reasoning, generalization, and solution exploration capabilities. Use when the user wants to benchmark on AIRS-Bench, or asks about evaluating this task. Reports average normalized score.
3 -
qhjqhj00 Skill Benchie Fl EvalEvaluates Open Information Extraction (OIE) systems on their ability to extract fact-based triples from text. It uses a conservative exact-matching function with synset-based clustering to penalize non-informative copies and reward precise fact extraction, while also measuring correlation with downstream QA and knowledge base tasks. Use when the user wants to benchmark on BenchIE^FL, or asks about evaluating this task. Reports exact-match.
3 -
qhjqhj00 Skill Dbench Bio EvalEvaluates whether large language models can discover genuinely new biological knowledge by generating correct scientific hypotheses or mechanisms from post-release literature, enforcing strict temporal separation to prevent data leakage. Use when the user wants to benchmark on DBench-Bio, or asks about evaluating this task. Reports Score.
3 -
qhjqhj00 Skill Fire Bench EvalEvaluates autonomous coding agents on their ability to rediscover established scientific findings by autonomously planning, implementing, and executing experiments from scratch based only on high-level research questions. It probes end-to-end research workflow capabilities, including experimental design, code generation, and evidence-based conclusion formation. Use when the user wants to benchmark on FIRE-Bench, or asks about evaluating this task. Reports F1.
3 -
qhjqhj00 Skill Mattermech EvalEvaluates large language models' ability to reason about physicochemical principles in nanomaterial synthesis. It probes whether models can generate scientifically valid hypotheses and understand conceptual mechanisms from literature abstracts, rather than relying on abstract logic alone. Use when the user wants to benchmark on MatterMech, or asks about evaluating this task. Reports accuracy.
3 -
qhjqhj00 Skill Trec Covid EvalEvaluates information retrieval systems on pandemic-related queries using a dynamically evolving corpus. It probes a system's ability to retrieve relevant medical literature under real-world conditions where terminology and document availability change rapidly. Use when the user wants to benchmark on TREC-COVID Round 1, or asks about evaluating this task. Reports NDCG@10.
3 -
qhjqhj00 Skill Repro Bench EvalEvaluates whether agentic AI systems can accurately assess the computational reproducibility of social science research by comparing original paper findings against results reproduced from provided raw data and code. It probes end-to-end agentic reasoning, including command execution, debugging, and result interpretation in a simulated research environment. Use when the user wants to benchmark on REPRO-Bench, or asks about evaluating this task. Reports accuracy.
3
Frequently asked questions
What are Research & Search agent skills?
Research agent skills teach AI agents to gather and synthesize information properly: literature reviews, competitive analysis, web research with citations, and structured summaries. Each SKILL.md encodes a method, not just a prompt, so results stay consistent across runs.
Which Research & Search skills are most installed?
Popular Research & Search skills on SkillMD right now include scientific-writing, citation-management, apify-market-research. Rankings shift as installs change; sort this page by "Most installs" for the live list.
Do Research & Search skills work with Claude Code and Cursor?
Yes. Every skill here ships as a SKILL.md file, an open format that works in Claude Code, Claude.ai, Cursor, Codex, Windsurf, and 60+ other agents. Install one with npx skillmds@latest add <owner>/<name>, or copy the file into your agent's skills directory.