Create aesthetically beautiful interfaces following proven design principles. Use when building UI/UX, analyzing designs from inspiration sites, generating design images with ai-multimodal, implementing visual hierarchy and color theory, adding micro-interactions, or creating design documentation. Includes workflows for capturing and analyzing inspiration screenshots with chrome-devtools and ai-multimodal, iterative design image generation until aesthetic standards are met, and comprehensive design system guidance covering BEAUTIFUL (aesthetic principles), RIGHT (functionality/accessibility), SATISFYING (micro-interactions), and PEAK (storytelling) stages. Integrates with chrome-devtools, ai-multimodal, media-processing, ui-styling, and web-frameworks skills.
日本語の概要は準備中です。原文の説明を表示しています。
mrgoonie/claudekit-skills☆ 2,2272026年4月3日 更新
Patterns for building multimodal AI applications that combine text, images, audio, and video. Covers vision APIs, audio transcription, and unified pipelines. Use when "multimodal AI, vision API, image understanding, GPT-4V, Claude vision, audio transcription, Whisper, document extraction, image to text, " mentioned.
日本語の概要は準備中です。原文の説明を表示しています。
omer-metin/skills-for-antigravity☆ 1642026年1月22日 更新
对 DiT(扩散/多模态生成)模型执行 W8A8 动态(data-free)或 W8A8 MXFP8(data-free)量化。覆盖 HunyuanVideo、Wan2.2-T2V/I2V/TI2V、FLUX.1-dev(已迁移)、SD3、Sana、HunyuanDiT、CogViewX 等。YAML apiversion 固定 multimodal_sd_modelslim_v1,统一 MultimodalPipelineInterface 重构路径(不再保留 LegacyMultimodalPipelineInterface)。模板与字段以 msmodelslim/model/hunyuan_video/model_adapter.py(唯一已迁移到重构路径且 lab_practice YAML 已 validated 的 DiT)为参考;其它 DiT 需按各自推理仓的 `parse_args` 调整字段集合。data-free 指不引入外部激活校准数据,spec.dataset(校准集短名,短名如 wan2_2_t2v;enable_dump: false 时不参与 dump)仍必填;精度调优通过下游经验库 L2 §7 + 直接调 `quantization-expert-experience-tuning-rules/scripts/apply_rollback.py`(整层回退)扩展 exclude 列表,闭环由 `quantization-accuracy-tuning-orchestrator` 调度。
日本語の概要は準備中です。原文の説明を表示しています。
kali20gakki/msAgent☆ 312026年10月10日 更新
Build multimodal AI applications with vision and text. TRIGGERS - Use when user needs help with ai-multimodal-app related tasks.
日本語の概要は準備中です。原文の説明を表示しています。
lionelsimai/claude-skills-collection☆ 292026年2月8日 更新
Create aesthetically beautiful interfaces following proven design principles. Use when building UI/UX, analyzing designs from inspiration sites, generating design images with ai-multimodal, implementing visual hierarchy and color theory, adding micro-interactions, or creating design documentation. Includes workflows for capturing and analyzing inspiration screenshots with chrome-devtools and ai-multimodal, iterative design image generation until aesthetic standards are met, and comprehensive design system guidance covering BEAUTIFUL (aesthetic principles), RIGHT (functionality/accessibility), SATISFYING (micro-interactions), and PEAK (storytelling) stages. Integrates with chrome-devtools, ai-multimodal, media-processing, ui-styling, and web-frameworks skills.
日本語の概要は準備中です。原文の説明を表示しています。
VoDaiLocz/kilo-kit-mcp☆ 272026年9月13日 更新
Guides agents to interactively discover customer requirements for live, bidirectional multi-agent AI systems that process continuous streams of multimodal data for real-time technical guidance and safety monitoring. Generates a custom Google Cloud solution that uses opinionated best practices and architecture guidance. Use when users need agentic assistance to design and create a multi-product solution in the cloud for live bidirectional multimodal streaming workloads. Don't use for simple text-based chat applications or workloads without real-time streaming requirements.
日本語の概要は準備中です。原文の説明を表示しています。
vaila-multimodaltoolbox/vaila☆ 192026年10月8日 更新
Create aesthetically beautiful interfaces following proven design principles. Use when building UI/UX, analyzing designs from inspiration sites, generating design images with ai-multimodal, implementing visual hierarchy and color theory, adding micro-interactions, or creating design documentation. Includes workflows for capturing and analyzing inspiration screenshots with chrome-devtools and ai-multimodal, iterative design image generation until aesthetic standards are met, and comprehensive design system guidance covering BEAUTIFUL (aesthetic principles), RIGHT (functionality/accessibility), SATISFYING (micro-interactions), and PEAK (storytelling) stages. Integrates with chrome-devtools, ai-multimodal, media-processing, ui-styling, and web-frameworks skills.
日本語の概要は準備中です。原文の説明を表示しています。
kettleofketchup/DraftForge☆ 152026年9月14日 更新
Create aesthetically beautiful interfaces following proven design principles. Use when building UI/UX, analyzing designs from inspiration sites, generating design images with ai-multimodal, implementing visual hierarchy and color theory, adding micro-interactions, or creating design documentation. Includes workflows for capturing and analyzing inspiration screenshots with chrome-devtools and ai-multimodal, iterative design image generation until aesthetic standards are met, and comprehensive design system guidance covering BEAUTIFUL (aesthetic principles), RIGHT (functionality/accessibility), SATISFYING (micro-interactions), and PEAK (storytelling) stages. Integrates with chrome-devtools, ai-multimodal, media-processing, ui-styling, and web-frameworks skills.
日本語の概要は準備中です。原文の説明を表示しています。
EthanYoQ/Skill-hub☆ 112026年10月5日 更新
Build multimodal AI applications with vision and text. TRIGGERS - Use when user needs help with ai-multimodal-app related tasks.
日本語の概要は準備中です。原文の説明を表示しています。
Winbda/claude-skills-collection☆ 42026年4月6日 更新
Framework for state-of-the-art sentence, text, and image embeddings. Provides 5000+ pre-trained models for semantic similarity, clustering, and retrieval. Supports multilingual, domain-specific, and multimodal models. Use for generating embeddings for RAG, semantic search, or similarity tasks. Best for production embedding generation.
日本語の概要は準備中です。原文の説明を表示しています。
davila7/claude-code-templates☆ 3.3万2026年10月11日 更新
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
日本語の概要は準備中です。原文の説明を表示しています。
davila7/claude-code-templates☆ 3.3万2026年10月11日 更新
Expert guidance for fine-tuning LLMs with LLaMA-Factory - WebUI no-code, 100+ models, 2/3/4/5/6/8-bit QLoRA, multimodal support
日本語の概要は準備中です。原文の説明を表示しています。
davila7/claude-code-templates☆ 3.3万2026年10月11日 更新
Guides agents to interactively discover customer requirements for live, bidirectional multi-agent AI systems that process continuous streams of multimodal data for real-time technical guidance and safety monitoring. Generates a custom Google Cloud solution that uses opinionated best practices and architecture guidance. Use when users need agentic assistance to design and create a multi-product solution in the cloud for live bidirectional multimodal streaming workloads. Don't use for simple text-based chat applications or workloads without real-time streaming requirements.
日本語の概要は準備中です。原文の説明を表示しています。
google/skills☆ 2.1万2026年10月10日 更新
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
日本語の概要は準備中です。原文の説明を表示しています。
Orchestra-Research/AI-Research-SKILLs☆ 1.3万2026年6月16日 更新
Framework for state-of-the-art sentence, text, and image embeddings. Provides 5000+ pre-trained models for semantic similarity, clustering, and retrieval. Supports multilingual, domain-specific, and multimodal models. Use for generating embeddings for RAG, semantic search, or similarity tasks. Best for production embedding generation.
日本語の概要は準備中です。原文の説明を表示しています。
Orchestra-Research/AI-Research-SKILLs☆ 1.3万2026年6月16日 更新
Expert guidance for fine-tuning LLMs with LLaMA-Factory - WebUI no-code, 100+ models, 2/3/4/5/6/8-bit QLoRA, multimodal support
日本語の概要は準備中です。原文の説明を表示しています。
Orchestra-Research/AI-Research-SKILLs☆ 1.3万2026年6月16日 更新
Generates BYO custom safety policies for NVIDIA Nemotron content-safety guardrails — Nemotron-Content-Safety-Reasoning-4B (text) and multimodal Nemotron-3-Content-Safety. Produces a Markdown policy, JSON taxonomy, and drop-in inference prompts. Maps rough words or an existing policy to V2 categories, adding custom categories or topic-following rules.
日本語の概要は準備中です。原文の説明を表示しています。
NVIDIA/skills☆ 3,5592026年10月10日 更新
Azure AI Content Understanding SDK for Python. Use for multimodal content extraction from documents, images, audio, and video. Triggers: "azure-ai-contentunderstanding", "ContentUnderstandingClient", "multimodal analysis", "document extraction", "video analysis", "audio transcription".
日本語の概要は準備中です。原文の説明を表示しています。
microsoft/skills☆ 3,1002026年10月10日 更新
Process and generate multimedia content using Google Gemini API. Capabilities include analyze audio files (transcription with timestamps, summarization, speech understanding, music/sound analysis up to 9.5 hours), understand images (captioning, object detection, OCR, visual Q&A, segmentation), process videos (scene detection, Q&A, temporal analysis, YouTube URLs, up to 6 hours), extract from documents (PDF tables, forms, charts, diagrams, multi-page), generate images (text-to-image, editing, composition, refinement). Use when working with audio/video files, analyzing images or screenshots, processing PDF documents, extracting structured data from media, creating images from text prompts, or implementing multimodal AI features. Supports multiple models (Gemini 2.5/2.0) with context windows up to 2M tokens.
日本語の概要は準備中です。原文の説明を表示しています。
mrgoonie/claudekit-skills☆ 2,2272026年4月3日 更新
SkillsBench task PR review — classifies the task track (standard / research / multimodal), runs static policy checks against the track-specific rubric, benchmarks the task across oracle plus Claude and Codex (with and without skills), audits trajectories for cheating and skill invocation, and produces a `pr-N-task-timestamp-run.txt` review report alongside a `prN.zip` bundle of trajectories. Use when reviewing a SkillsBench task PR (by number, branch, or local task path), when the user asks to review a task, run benchmarks on a PR, audit a submission, classify a task as research or multimodal track, or prepare a comment to post on a SkillsBench PR.
日本語の概要は準備中です。原文の説明を表示しています。
benchflow-ai/skillsbench☆ 1,8372026年7月24日 更新
Process and generate multimedia content using Google Gemini API. Capabilities include analyze audio files (transcription with timestamps, summarization, speech understanding, music/sound analysis up to 9.5 hours), understand images (captioning, object detection, OCR, visual Q&A, segmentation), process videos (scene detection, Q&A, temporal analysis, YouTube URLs, up to 6 hours), extract from documents (PDF tables, forms, charts, diagrams, multi-page), generate images (text-to-image, editing, composition, refinement). Use when working with audio/video files, analyzing images or screenshots, processing PDF documents, extracting structured data from media, creating images from text prompts, or implementing multimodal AI features. Supports multiple models (Gemini 2.5/2.0) with context windows up to 2M tokens.
日本語の概要は準備中です。原文の説明を表示しています。
Microck/ordinary-claude-skills☆ 4042026年9月7日 更新
Analyze and extract information from images using multimodal AI recognition. Triggered when users want to analyze, describe, or extract information from an image URL — image recognition, image analysis, image description, visual content understanding, OCR text recognition, or visual Q&A. When a user provides an image URL and asks questions about its visual content, this skill should be triggered even if they do not explicitly say "image recognition."
日本語の概要は準備中です。原文の説明を表示しています。
nexscope-ai/nexscope-ecommerce-skills☆ 852026年10月9日 更新
Master optimization protocol for Gemini Agent (Antigravity) to unlock native 2M+ to 10M+ long-context reasoning, Gemini 4 Pro & Gemini 4 Flash dynamic thinking budget control, native context caching, Multimodal Live API, and high-speed problem solving / Protokol optimasi utama untuk Gemini Agent (Antigravity) untuk mengaktifkan pemikiran long-context 2M+ hingga 10M+, kontrol thinking budget dinamis Gemini 4 Pro & Gemini 4 Flash, context caching native, Multimodal Live API, dan pemecahan masalah kecepatan tinggi.
日本語の概要は準備中です。原文の説明を表示しています。
roedyrustam/vibes-plug☆ 752026年10月9日 更新
Expert guide for Affective Computing, emotional AI, and real-time sentiment analysis through native multimodal tokens (voice intonation and facial micro-expressions) / Panduan ahli komputasi afektif, AI emosional, dan analisis sentimen real-time melalui token multimodal native.
日本語の概要は準備中です。原文の説明を表示しています。
roedyrustam/vibes-plug☆ 752026年10月9日 更新