本文へ移動
cccskills
無料GitHub で公開

gemma-trainer

Trigger this skill when the user wants to train, fine-tune, or adapt Gemma models (e.g. SFT, DPO, RLHF, Reward Modeling) on local hardware. Covers TRL, Unsloth, dataset preparation, validation, and GGUF/LiteRT conversion.

インストール方法を見る

含まれるファイル(6)

  • SKILL.md7.5 KB
  • assets/dataset_prep.py11.2 KB
  • assets/distill_dataset.py17.2 KB
  • assets/dpo_train.py8.0 KB
  • assets/reward_train.py5.4 KB
  • assets/sft_train.py13.6 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Gemma Training and Fine-Tuning Skill

1. Core Principles: Local Fine-Tuning Setup

When training locally, memory efficiency and execution speed are huge. Always guide the user to follow these best practices:

  • Prioritize Unsloth: For local single-GPU training, always recommend Unsloth. It supports Gemma 4 natively, uses up to 70% less memory, and is up to 2x faster than standard Hugging Face PEFT training.
  • Fall Back to TRL: For multi-GPU environments (using DDP/FSDP) or when Unsloth is unavailable, use Hugging Face TRL (SFTTrainer, DPOTrainer) coupled with PEFT and bitsandbytes (for QLoRA).
  • Always use QLoRA (4-bit Quantization): Crucial for fitting Gemma models (like Gemma 4 12B/31B) into consumer VRAM.
  • Manage Context Window & Max Length: Although Gemma 4 supports up to a 256K context window, recommend training with a context window of 2048 to 8192 tokens locally to prevent Out-Of-Memory (OOM) errors.

2. Choosing the Right Training Method

Help the user choose the correct workflow based on their goal:

  • Supervised Fine-Tuning (SFT): Teaching new domains, specialized task instructions, or custom output structures.
    • Prerequisites: Raw text, instruction pairs, or chat logs.
    • Output: Adaptor trained on prompt/completion pairs.
  • Direct Preference Optimization (DPO): Aligning model style, behavior, tone, or safety with human preferences.
    • Prerequisites: A previously SFT-trained Gemma model and preferred pairwise datasets.
    • Output: Aligning model weights directly without a separate reward head.
  • Reward Modeling (RM): Training a scoring system to evaluate response quality.
    • Prerequisites: Binary preference pairwise datasets.
    • Output: A classification-style reward head on top of Gemma.

3. Dataset Preparation & Validation

Formatting issues are the #1 cause of poor training runs. Ensure you validate files using the utility script [assets/dataset_prep.py].

Gemma Chat Prompt Format

Ensure the dataset matches Gemma's official chat template:

<|turn>system
Your instruction here<turn|>
<|turn>user
Your query here<turn|>
<|turn>model
Your response here<turn|>

To avoid formatting drift, use the tokenizer's apply_chat_template during dataset tokenization.

Format Specifications

  • SFT (Supervised Fine-Tuning): Format as a list of conversation turns.
    {
      "messages": [
        {"role": "user", "content": "Tell me a joke."},
        {"role": "model", "content": "Why did the computer go to the doctor? It had a virus!"}
      ]
    }
    
  • DPO (Direct Preference Optimization): Requires pairwise samples containing a prompt, a chosen (better) response, and a rejected (worse) response.
    {
      "prompt": "Write a python function to compute factorial.",
      "chosen": "def factorial(n):\n    return 1 if n <= 1 else n * factorial(n - 1)",
      "rejected": "factorial is computed using recursion or loops. Just import math."
    }
    
  • Reward Modeling: Format identically to DPO datasets. The RewardTrainer evaluates the pair and learns to output a higher logit score for chosen than for rejected.

Dataset Distillation & Synthesis (Teacher-Student)

Local knowledge distillation allows you to train small, lightweight student models (such as Gemma 4 E2B) using high-quality dataset outputs generated by larger, highly capable teacher models (such as Gemma 4 31B or Gemma 4 26B A4B).

Use the [assets/distill_dataset.py] utility script to generate fine-tuning datasets on your local machine.

4. Fine-Tuning Workflows

Supervised Fine-Tuning (SFT)

Use the [assets/sft_train.py] asset to launch a local QLoRA fine-tuning session.

  • LoRA Hyperparameters:
    • Rank (r): 16 or 32 (Higher rank captures complex behaviors but consumes more memory).
    • Alpha (lora_alpha): 32 or 64 (Rule of thumb: lora_alpha = 2 * r).
    • Dropout (lora_dropout): 0.05 or 0.1 (Forcing the model to learn more robust features rather than relying on specific paths).
    • Target Modules: Use PEFT's Gemma 4 defaults scope to the LM layers.
    • Learning Rate: 2e-4 for QLoRA; 2e-5 for full fine-tuning.

Direct Preference Optimization (DPO)

Use the [assets/dpo_train.py] template to execute alignment.

  • Rules for DPO:
    • Always perform SFT on the base model using your instruction format before running DPO. Running DPO directly on an out-of-domain base model usually fails or degrades output formatting.
    • Set beta (DPO temperature parameter) to 0.1. Values between 0.1 and 0.5 control how strictly the model adheres to the reference policy.

Reward Modeling (RM)

Use the [assets/reward_train.py] template to train an evaluation model.

  • Initializes the model with a sequence classification head (AutoModelForSequenceClassification with num_labels=1).
  • Trains the single scalar reward value to distinguish preferred responses.

5. Multimodal Fine-Tuning (Vision & Audio)

Gemma 4 models are natively multimodal. To fine-tune them on images or audio:

Vision SFT

  • Use Gemma 4 E2B/E4B/12B/26B/31B models.
  • Use standard Hugging Face SFTTrainer with a custom visual data collator.
  • Prepare your dataset containing local image paths or PIL images, alongside corresponding conversation instructions:
    {
      "messages": [
        {"role": "user", "content": [
            {"type": "image", "url": "path/to/image.png"},
            {"type": "text", "text": "Describe this image."}
        ]},
        {"role": "assistant", "content": [
            {"type": "text", "text": "An abstract oil painting with vibrant warm gradients."}
        ]}
      ]
    }
    
    

Audio SFT

  • Use Gemma 4 E2B/E4B/12B models.
  • Feed raw audio arrays (sampled at 16kHz) through the model processor to produce input_features.
  • Maintain conversational formatting, replacing image type with audio type in the message format.
    {
      "messages": [
        {"role": "user", "content": [
            {"type": "text", "text": "Describe this audio."},
            {"type": "audio", "url": "path/to/audio.wav"}
        ]},
        {"role": "assistant", "content": [
            {"type": "text", "text": "This is an audio file of a bird chirping."}
        ]}
      ]
    }
    

6. Post-Training & Deployment Utilities

GGUF Conversion

Once your LoRA training is finished, you can convert your model to GGUF format.

Option 1: Native Export via Unsloth (Recommended)

If you trained your model using Unsloth, you can export directly to GGUF natively. This automatically handles merging and quantization.

Fetch Saving to GGUF for the best practice.

Option 2: Manual Conversion with llama.cpp

If you did not use Unsloth, you can convert your merged Hugging Face model directory manually using llama.cpp.

On-Device Deployment with LiteRT-LM (.litertlm)

LiteRT-LM is optimized for running models like Gemma 4 E2B and Gemma 4 E4B on mobile, web, and IoT hardware with hardware acceleration (CPU, GPU, NPU). Fetch LiteRT-LM guide for the best practice.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Use this skill whenever the user wants to create or improve a presentation for an academic context — conference papers, seminar talks, thesis defenses, grant briefings, lab meetings, invited lectures, or any presentation where the audience will evaluate reasoning and evidence. Triggers include: 'conference talk', 'seminar slides', 'thesis defense', 'research presentation', 'academic deck', 'academic presentation'. Also triggers when the user asks to 'make slides' in combination with academic content (e.g., 'make slides for my paper on X', 'create a presentation for my dissertation defense', 'build a deck for my grant proposal'). This skill governs CONTENT and STRUCTURE decisions. For the technical work of creating or editing the .pptx file itself, also read the pptx SKILL.md.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd712026年10月9日 更新

adp

無料

Redpanda's Agentic Data Plane: governance infrastructure for building, running, and governing AI agents and MCP servers, plus a proxying AI Gateway for LLM providers, operated via `rpk ai` and the ADP API. Use when creating or managing AI agents (managed or self-managed) via `rpk ai agent` or `AgentRegistryService`; configuring MCP servers (remote or managed catalog, code mode, auth); setting up LLM providers or querying models via `rpk ai llm`/`rpk ai model` or the AI Gateway proxy; or configuring budgets, guardrails, or Cedar access-control policies through the governance APIs. Also covers reading agent transcripts and spending insights, and wiring OAuth clients or providers to the aigw Authorization Server. For the separate rpk cloud mcp control-plane MCP server, see `/redpanda:rpk-cloud`.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd712026年10月9日 更新

ads

無料

Operate professional paid advertising across Google, Meta, YouTube, LinkedIn, TikTok, Microsoft, Apple, Amazon, Reddit, Pinterest, Snapchat, and X. Use for account intake, source-grounded audits, strategy, budget and measurement planning, creative production, experiments, reporting, monitoring, and explicitly approved campaign changes. Also trigger on PPC, paid social, retail media, attribution, tracking, landing pages, cross-platform conversion totals, negative keywords or search terms, beta-feature scoring, stale platform claims, API-token or credential setup, campaign deletion, and safe Claude Ads installation or uninstall.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd712026年10月9日 更新

ads-apple

無料

Audit Apple Ads measurement, AdServices and AdAttributionKit, campaign and keyword structure, Search Match, App Store placements, custom product pages, bidding, budgets, MMP reconciliation, and policy. Use for Apple Ads, Apple Search Ads, App Store ads, Search Match, custom product pages, AdServices, or Apple app-install campaigns.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd712026年10月9日 更新

Research competitor paid-ad presence, messaging, creative, formats, landing pages, keyword and auction signals, transparent ad libraries, and strategic gaps across supported platforms. Use for competitor ads, ad libraries, ad spy, competitive PPC analysis, competitor creative, Google Ads Transparency, Meta Ad Library, or paid-media competitor research.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd712026年10月9日 更新

ads-dna

無料

Extract a public-safe brand and offer profile for paid advertising from an authorized website and operator input. Triggers on: brand DNA, brand profile, brand identity, brand style, brand colors, brand voice, visual identity, style guide, website brand analysis.

日本語の概要は準備中です。原文の説明を表示しています。

skillmds/skillmd712026年10月9日 更新

skillmds のスキルをすべて見る

このスキルの問題を報告する