文章をAIが扱う小さな単位に分け、大量のテキスト処理や独自の語彙学習を支援するスキル。分割結果と元の文章の位置対応、入力の長さ調整も扱います。
- 大量の文章をモデル入力用に処理したいとき
- 文章ファイルから独自の語彙を学習
- 予測結果を元の文章の位置に対応づける
18 件 ・ 関連度順
概要と使いどころ
文章をAIが扱う小さな単位に分け、大量のテキスト処理や独自の語彙学習を支援するスキル。分割結果と元の文章の位置対応、入力の長さ調整も扱います。
Fast tokenizers optimized for research and production. Rust-based implementation tokenizes 1GB in <20 seconds. Supports BPE, WordPiece, and Unigram algorithms. Train custom vocabularies, track alignments, handle padding/truncation. Integrates seamlessly with transformers. Use when you need high-performance tokenization or custom tokenizer training.
日本語の概要は準備中です。原文の説明を表示しています。
Language-independent tokenizer treating text as raw Unicode. Supports BPE and Unigram algorithms. Fast (50k sentences/sec), lightweight (6MB memory), deterministic vocabulary. Used by T5, ALBERT, XLNet, mBART. Train on raw text without pre-tokenization. Use when you need multilingual support, CJK languages, or reproducible tokenization.
日本語の概要は準備中です。原文の説明を表示しています。
Language-independent tokenizer treating text as raw Unicode. Supports BPE and Unigram algorithms. Fast (50k sentences/sec), lightweight (6MB memory), deterministic vocabulary. Used by T5, ALBERT, XLNet, mBART. Train on raw text without pre-tokenization. Use when you need multilingual support, CJK languages, or reproducible tokenization.
日本語の概要は準備中です。原文の説明を表示しています。
Fast tokenizers optimized for research and production. Rust-based implementation tokenizes 1GB in <20 seconds. Supports BPE, WordPiece, and Unigram algorithms. Train custom vocabularies, track alignments, handle padding/truncation. Integrates seamlessly with transformers. Use when you need high-performance tokenization or custom tokenizer training.
日本語の概要は準備中です。原文の説明を表示しています。
Build direct Bluetooth Low Energy workflows with Core Bluetooth. Use when implementing BLE central or peripheral GATT communication, scanning or connecting with CBCentralManager, discovering services and characteristics, reading/writing/subscribing with CBPeripheral, publishing local services with CBPeripheralManager, handling Bluetooth authorization, background BLE modes, state restoration, write flow control, or CBUUID-based workflows. For privacy-preserving accessory setup/picker flows, use accessorysetupkit first and return here for post-setup GATT communication.
日本語の概要は準備中です。原文の説明を表示しています。
Tokenizador agnóstico de linguagem que trata texto como Unicode bruto. Suporta algoritmos BPE e Unigram. Rápido (50k sentenças/seg), leve (6MB de memória), vocabulário determinístico. Usado por T5, ALBERT, XLNet, mBART. Treina em texto bruto sem pré-tokenização. Use quando você precisar de suporte multilíngue, linguagens CJK ou tokenização reproduzível.
日本語の概要は準備中です。原文の説明を表示しています。
Tokenizadores rápidos otimizados para pesquisa e produção. Implementação em Rust que tokeniza 1GB em menos de 20 segundos. Suporta algoritmos BPE, WordPiece e Unigram. Treine vocabulários customizados, rastreie alinhamentos, gerencie padding/truncagem. Integração perfeita com transformers. Use quando precisar de tokenização de alto desempenho ou treinamento de tokenizadores customizados.
日本語の概要は準備中です。原文の説明を表示しています。
Scan, connect, and communicate with Bluetooth Low Energy peripherals and publish local peripheral services using Core Bluetooth. Use when implementing BLE central or peripheral roles, discovering services and characteristics, reading and writing characteristic values, subscribing to notifications, configuring background BLE modes, restoring state after app relaunch, or working with CBCentralManager, CBPeripheral, CBPeripheralManager, CBService, CBCharacteristic, CBUUID, or Bluetooth Low Energy workflows.
日本語の概要は準備中です。原文の説明を表示しています。
Fast tokenizers optimized for research and production. Rust-based implementation tokenizes 1GB in <20 seconds. Supports BPE, WordPiece, and Unigram algorithms. Train custom vocabularies, track alignments, handle padding/truncation. Integrates seamlessly with transformers. Use when you need high-performance tokenization or custom tokenizer training.
日本語の概要は準備中です。原文の説明を表示しています。
Language-independent tokenizer treating text as raw Unicode. Supports BPE and Unigram algorithms. Fast (50k sentences/sec), lightweight (6MB memory), deterministic vocabulary. Used by T5, ALBERT, XLNet, mBART. Train on raw text without pre-tokenization. Use when you need multilingual support, CJK languages, or reproducible tokenization.
日本語の概要は準備中です。原文の説明を表示しています。
Fast tokenizers optimized for research and production. Rust-based implementation tokenizes 1GB in <20 seconds. Supports BPE, WordPiece, and Unigram algorithms. Train custom vocabularies, track alignments, handle padding/truncation. Integrates seamlessly with transformers. Use when you need high-performance tokenization or custom tokenizer training.
日本語の概要は準備中です。原文の説明を表示しています。
Audits any AI instruction set for over-prompting using the core test — would a smarter model make this rule unnecessary? Applies Five Questions to every rule (Claude already does this? Contradiction? Redundant? One-off fix? Vague?) then classifies as CUT/RESOLVE/MERGE/EVALUATE/SHARPEN/MOVE/KEEP. Workflows: Audit (full system, token savings), QuickCheck (single file). Principle: less scaffolding = better output. USE WHEN BPE, bitter pill, audit setup, over-prompting, trim instructions, dead weight, simplify setup, clean up CLAUDE.md. NOT FOR attacking logical flaws in ideas (use RedTeam).
日本語の概要は準備中です。原文の説明を表示しています。
Screens biomedical / life-science papers for signs of data fabrication, image manipulation, and statistical anomalies, using the detection techniques distilled from the field's canonical exposure platforms (PubPeer, Data Colada, Science Integrity Digest, For Better Science) and tools (ImageTwin/Proofig, statcheck, GRIM/GRIMMER, Problematic Paper Screener, Seek & Blastn). Use when asked to check a paper/figure for image duplication, blot splicing, impossible statistics, paper-mill or tortured-phrase signals, research integrity, or "is this data faked"; or when a user shares a figure, Western blot, supplementary dataset, or DOI and asks whether it looks manipulated. Reports observable anomalies as questions for clarification — it never accuses anyone of fraud.
日本語の概要は準備中です。原文の説明を表示しています。
Builds a GPT and BPE tokenizer from scratch. Use when implementing autograd, attention, a nanoGPT loop, muP hyperparameter transfer, WSD annealing, loss spikes, or mid-training.
日本語の概要は準備中です。原文の説明を表示しています。
Use this skill to configure a Lightning Service Console app in Salesforce — Console Navigation (split view), workspace tabs and subtabs, the utility bar with Omni-Channel, Macros and History, Quick Text, keyboard shortcuts, and navigation rules per object. Trigger keywords: Service Console, console app, workspace tabs, subtabs, utility bar macros, Omni-Channel utility, split view, Quick Text, console navigation rules, keyboard shortcuts service console, navType Console, workspaceConfig mappings, WorkspaceMapping fieldName, UtilityBar FlexiPage, tabLimitConfig, listPlacement, MacroInstruction, QuickText Channel multipicklist, isNavTabPersistenceDisabled. NOT for opening or refreshing console tabs from code — use lwc/lwc-console-workspace-api. NOT for Omni-Channel routing setup — use admin/omni-channel-routing-setup.
日本語の概要は準備中です。原文の説明を表示しています。
Fast BPE/WordPiece tokenization and custom vocab training.
日本語の概要は準備中です。原文の説明を表示しています。
Fast tokenizers optimized for research and production. Rust-based implementation tokenizes 1GB in <20 seconds. Supports BPE, WordPiece, and Unigram algorithms. Train custom vocabularies, track alignments, handle padding/truncation. Integrates seamlessly with transformers. Use when you need high-performance tokenization or custom tokenizer training.
日本語の概要は準備中です。原文の説明を表示しています。