huggingface-tokenizers

orchestra-research/huggingface-tokenizers · Agent Skill (multi-file)

by Orchestra Research · bundle

Published · Last updated


Fast tokenization for NLP using Rust-based tokenizers supporting BPE, WordPiece, and Unigram algorithms, with training, alignment tracking, and padding/truncation.

SKILL.md

Files

This skill is a package of 5 files. Install with the command above, or download the folder.

  • 📄SKILL.md entry
  • 📁references
  • 📄algorithms.md 14.8 KB
  • 📄integration.md 15.0 KB
  • 📄pipeline.md 16.4 KB
  • 📄training.md 14.2 KB

Related

  1. sentencepiece · orchestra-research bundle
    Train and use SentencePiece tokenizers for multilingual NLP, supporting BPE and Unigram algorithms with raw Unicode text.
    10.4k
    repo stars
  2. gtars · k-dense-ai bundle
    High-performance toolkit for genomic interval analysis in Rust with Python bindings. Use when working with genomic regions, BED files, coverage tracks, overlap detection, tokenization for ML models, or fragment analysis in computational genomics and machine learning applications.
    30.2k
    repo stars
  3. gtars · lingxling bundle
    High-performance toolkit for genomic interval analysis in Rust with Python bindings. Use when working with genomic regions, BED files, coverage tracks, overlap detection, tokenization for ML models, or fragment analysis in computational genomics and machine learning applications.
    253
    repo stars
  4. transformers-js · huggingface bundle
    Run state-of-the-art machine learning models directly in JavaScript/TypeScript across browsers and server-side runtimes using Transformers.js.
    10.8k
    repo stars
  5. huggingface-tool-builder · huggingface bundle
    Creates reusable command-line scripts and utilities for the Hugging Face API, enabling chaining, piping, and intermediate data processing.
    10.8k
    repo stars
  6. earth2studio-create-datasource · nvidia bundle
    Create and validate Earth2Studio data source wrappers (DataSource, ForecastSource, DataFrameSource, ForecastFrameSource) from remote stores like S3, GCS, Azure, HTTP, or HuggingFace.
    2.2k
    repo stars

Frequently asked questions

How do I install the huggingface-tokenizers skill?

Run npx skillmds add orchestra-research/huggingface-tokenizers in your terminal (requires Node.js), paste this page's agent-chat prompt into Claude, Cursor, or any MCP-connected agent, or download the SKILL.md file and copy it into your agent's skills directory.

What does the huggingface-tokenizers skill do?

Fast tokenization for NLP using Rust-based tokenizers supporting BPE, WordPiece, and Unigram algorithms, with training, alignment tracking, and padding/truncation. It is listed under AI & ML, Prompt Engineering on SkillMD.

Is huggingface-tokenizers safe to use?

SkillMD's automated safety review verdict for this skill is PASS. Independent scanners report: SkillSpector: PASS, Skill Scanner: PASS. Capability flags: makes network calls, reads secrets. SkillMD never runs a skill's scripts for you; review the SKILL.md before installing.

Which AI agents work with huggingface-tokenizers?

This skill is tagged as working with Claude Code, Claude.ai, OpenAI Codex. SKILL.md is an open format, so most agents that read a skills directory can load it too.

Is huggingface-tokenizers free to use?

Yes. Installing skills from SkillMD is free. This skill is licensed under MIT.

Who published huggingface-tokenizers?

Orchestra Research (@orchestra-research) published this skill. Their other Agent Skills are listed on their SkillMD profile.