sentencepiece

orchestra-research/sentencepiece · Agent Skill (multi-file)

by Orchestra Research · bundle

Published · Last updated


Train and use SentencePiece tokenizers for multilingual NLP, supporting BPE and Unigram algorithms with raw Unicode text.

SKILL.md

Files

This skill is a package of 3 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁references
  • 📄algorithms.md 4.1 KB
  • 📄training.md 6.1 KB

Related

  1. huggingface-tokenizers · orchestra-research bundle
    Fast tokenization for NLP using Rust-based tokenizers supporting BPE, WordPiece, and Unigram algorithms, with training, alignment tracking, and padding/truncation.
    10.4k
    repo stars
  2. gtars · lingxling bundle
    High-performance toolkit for genomic interval analysis in Rust with Python bindings. Use when working with genomic regions, BED files, coverage tracks, overlap detection, tokenization for ML models, or fragment analysis in computational genomics and machine learning applications.
    253
    repo stars
  3. aeon · k-dense-ai bundle
    Perform time series machine learning tasks including classification, regression, clustering, forecasting, anomaly detection, segmentation, and similarity search using a scikit-learn compatible Python toolkit.
    30.2k
    repo stars
  4. holoscan-install-conda · nvidia bundle
    Install Holoscan SDK v4.3+ via Conda in a CUDA 13 environment, including Python bindings and C++ development headers.
    2.2k
    repo stars
  5. cuopt-developer · nvidia bundle
    Modify, build, test, debug, and contribute to the NVIDIA cuOpt codebase (C++/CUDA, Python, server, CI). Includes guidance for solver internals, pull requests, DCO signoff, and code conventions.
    2.2k
    repo stars
  6. gtars · k-dense-ai bundle
    High-performance toolkit for genomic interval analysis in Rust with Python bindings. Use when working with genomic regions, BED files, coverage tracks, overlap detection, tokenization for ML models, or fragment analysis in computational genomics and machine learning applications.
    30.2k
    repo stars

Frequently asked questions

How do I install the sentencepiece skill?

Run npx skillmds add orchestra-research/sentencepiece in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the sentencepiece skill do?

Train and use SentencePiece tokenizers for multilingual NLP, supporting BPE and Unigram algorithms with raw Unicode text. It is listed under AI & ML, Coding & Dev Tools, Data & Analytics, Model Training & Fine-tuning on SkillMD.

Is sentencepiece safe to use?

SkillMD's automated safety review verdict for this skill is PASS. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: makes network calls, reads secrets. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with sentencepiece?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is sentencepiece free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under MIT.

Who published sentencepiece?

Orchestra Research (@orchestra-research) published this skill. Their other Agent Skills are listed on their SkillMD profile.