本文へ移動
cccskills

「bpe」の検索結果

18 件 ・ 関連度順

概要と使いどころ

huggingface-tokenizers

無料日本語概要

文章をAIが扱う小さな単位に分け、大量のテキスト処理や独自の語彙学習を支援するスキル。分割結果と元の文章の位置対応、入力の長さ調整も扱います。

  • 大量の文章をモデル入力用に処理したいとき
  • 文章ファイルから独自の語彙を学習
  • 予測結果を元の文章の位置に対応づける
NousResearch/hermes-agent25.3万2026年10月11日 更新

Fast tokenizers optimized for research and production. Rust-based implementation tokenizes 1GB in <20 seconds. Supports BPE, WordPiece, and Unigram algorithms. Train custom vocabularies, track alignments, handle padding/truncation. Integrates seamlessly with transformers. Use when you need high-performance tokenization or custom tokenizer training.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Language-independent tokenizer treating text as raw Unicode. Supports BPE and Unigram algorithms. Fast (50k sentences/sec), lightweight (6MB memory), deterministic vocabulary. Used by T5, ALBERT, XLNet, mBART. Train on raw text without pre-tokenization. Use when you need multilingual support, CJK languages, or reproducible tokenization.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Language-independent tokenizer treating text as raw Unicode. Supports BPE and Unigram algorithms. Fast (50k sentences/sec), lightweight (6MB memory), deterministic vocabulary. Used by T5, ALBERT, XLNet, mBART. Train on raw text without pre-tokenization. Use when you need multilingual support, CJK languages, or reproducible tokenization.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年10月11日 更新

Fast tokenizers optimized for research and production. Rust-based implementation tokenizes 1GB in <20 seconds. Supports BPE, WordPiece, and Unigram algorithms. Train custom vocabularies, track alignments, handle padding/truncation. Integrates seamlessly with transformers. Use when you need high-performance tokenization or custom tokenizer training.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年10月11日 更新

Build direct Bluetooth Low Energy workflows with Core Bluetooth. Use when implementing BLE central or peripheral GATT communication, scanning or connecting with CBCentralManager, discovering services and characteristics, reading/writing/subscribing with CBPeripheral, publishing local services with CBPeripheralManager, handling Bluetooth authorization, background BLE modes, state restoration, write flow control, or CBUUID-based workflows. For privacy-preserving accessory setup/picker flows, use accessorysetupkit first and return here for post-setup GATT communication.

日本語の概要は準備中です。原文の説明を表示しています。

dpearson2699/swift-ios-skills1,1882026年8月1日 更新

Tokenizador agnóstico de linguagem que trata texto como Unicode bruto. Suporta algoritmos BPE e Unigram. Rápido (50k sentenças/seg), leve (6MB de memória), vocabulário determinístico. Usado por T5, ALBERT, XLNet, mBART. Treina em texto bruto sem pré-tokenização. Use quando você precisar de suporte multilíngue, linguagens CJK ou tokenização reproduzível.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

Tokenizadores rápidos otimizados para pesquisa e produção. Implementação em Rust que tokeniza 1GB em menos de 20 segundos. Suporta algoritmos BPE, WordPiece e Unigram. Treine vocabulários customizados, rastreie alinhamentos, gerencie padding/truncagem. Integração perfeita com transformers. Use quando precisar de tokenização de alto desempenho ou treinamento de tokenizadores customizados.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

Scan, connect, and communicate with Bluetooth Low Energy peripherals and publish local peripheral services using Core Bluetooth. Use when implementing BLE central or peripheral roles, discovering services and characteristics, reading and writing characteristic values, subscribing to notifications, configuring background BLE modes, restoring state after app relaunch, or working with CBCentralManager, CBPeripheral, CBPeripheralManager, CBService, CBCharacteristic, CBUUID, or Bluetooth Low Energy workflows.

日本語の概要は準備中です。原文の説明を表示しています。

JordanCoin/ios-skills-collection62026年9月10日 更新

Fast tokenizers optimized for research and production. Rust-based implementation tokenizes 1GB in <20 seconds. Supports BPE, WordPiece, and Unigram algorithms. Train custom vocabularies, track alignments, handle padding/truncation. Integrates seamlessly with transformers. Use when you need high-performance tokenization or custom tokenizer training.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

Language-independent tokenizer treating text as raw Unicode. Supports BPE and Unigram algorithms. Fast (50k sentences/sec), lightweight (6MB memory), deterministic vocabulary. Used by T5, ALBERT, XLNet, mBART. Train on raw text without pre-tokenization. Use when you need multilingual support, CJK languages, or reproducible tokenization.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

Fast tokenizers optimized for research and production. Rust-based implementation tokenizes 1GB in <20 seconds. Supports BPE, WordPiece, and Unigram algorithms. Train custom vocabularies, track alignments, handle padding/truncation. Integrates seamlessly with transformers. Use when you need high-performance tokenization or custom tokenizer training.

日本語の概要は準備中です。原文の説明を表示しています。

Lord1Egypt/awesome-skill-forge22026年6月10日 更新

Audits any AI instruction set for over-prompting using the core test — would a smarter model make this rule unnecessary? Applies Five Questions to every rule (Claude already does this? Contradiction? Redundant? One-off fix? Vague?) then classifies as CUT/RESOLVE/MERGE/EVALUATE/SHARPEN/MOVE/KEEP. Workflows: Audit (full system, token savings), QuickCheck (single file). Principle: less scaffolding = better output. USE WHEN BPE, bitter pill, audit setup, over-prompting, trim instructions, dead weight, simplify setup, clean up CLAUDE.md. NOT FOR attacking logical flaws in ideas (use RedTeam).

日本語の概要は準備中です。原文の説明を表示しています。

danielmiessler/LifeOS1.9万2026年9月4日 更新

Screens biomedical / life-science papers for signs of data fabrication, image manipulation, and statistical anomalies, using the detection techniques distilled from the field's canonical exposure platforms (PubPeer, Data Colada, Science Integrity Digest, For Better Science) and tools (ImageTwin/Proofig, statcheck, GRIM/GRIMMER, Problematic Paper Screener, Seek & Blastn). Use when asked to check a paper/figure for image duplication, blot splicing, impossible statistics, paper-mill or tortured-phrase signals, research integrity, or "is this data faked"; or when a user shares a figure, Western blot, supplementary dataset, or DOI and asks whether it looks manipulated. Reports observable anomalies as questions for clarification — it never accuses anyone of fraud.

日本語の概要は準備中です。原文の説明を表示しています。

agentsope/SkillAlchemy4412026年10月9日 更新

Builds a GPT and BPE tokenizer from scratch. Use when implementing autograd, attention, a nanoGPT loop, muP hyperparameter transfer, WSD annealing, loss spikes, or mid-training.

日本語の概要は準備中です。原文の説明を表示しています。

vasilyu1983/AI-Agents-public912026年10月5日 更新

Use this skill to configure a Lightning Service Console app in Salesforce — Console Navigation (split view), workspace tabs and subtabs, the utility bar with Omni-Channel, Macros and History, Quick Text, keyboard shortcuts, and navigation rules per object. Trigger keywords: Service Console, console app, workspace tabs, subtabs, utility bar macros, Omni-Channel utility, split view, Quick Text, console navigation rules, keyboard shortcuts service console, navType Console, workspaceConfig mappings, WorkspaceMapping fieldName, UtilityBar FlexiPage, tabLimitConfig, listPlacement, MacroInstruction, QuickText Channel multipicklist, isNavTabPersistenceDisabled. NOT for opening or refreshing console tabs from code — use lwc/lwc-console-workspace-api. NOT for Omni-Channel routing setup — use admin/omni-channel-routing-setup.

日本語の概要は準備中です。原文の説明を表示しています。

PranavNagrecha/AwesomeSalesforceSkills192026年10月4日 更新

Fast BPE/WordPiece tokenization and custom vocab training.

日本語の概要は準備中です。原文の説明を表示しています。

openamer/openamer52026年10月11日 更新

Fast tokenizers optimized for research and production. Rust-based implementation tokenizes 1GB in <20 seconds. Supports BPE, WordPiece, and Unigram algorithms. Train custom vocabularies, track alignments, handle padding/truncation. Integrates seamlessly with transformers. Use when you need high-performance tokenization or custom tokenizer training.

日本語の概要は準備中です。原文の説明を表示しています。

ibragimov-oasis/vibe-coder22026年6月24日 更新