本文へ移動
cccskills

「multimodal」の検索結果

394 件 ・ 関連度順

概要と使いどころ

aesthetic

無料

Create aesthetically beautiful interfaces following proven design principles. Use when building UI/UX, analyzing designs from inspiration sites, generating design images with ai-multimodal, implementing visual hierarchy and color theory, adding micro-interactions, or creating design documentation. Includes workflows for capturing and analyzing inspiration screenshots with chrome-devtools and ai-multimodal, iterative design image generation until aesthetic standards are met, and comprehensive design system guidance covering BEAUTIFUL (aesthetic principles), RIGHT (functionality/accessibility), SATISFYING (micro-interactions), and PEAK (storytelling) stages. Integrates with chrome-devtools, ai-multimodal, media-processing, ui-styling, and web-frameworks skills.

日本語の概要は準備中です。原文の説明を表示しています。

mrgoonie/claudekit-skills2,2272026年4月3日 更新

Patterns for building multimodal AI applications that combine text, images, audio, and video. Covers vision APIs, audio transcription, and unified pipelines. Use when "multimodal AI, vision API, image understanding, GPT-4V, Claude vision, audio transcription, Whisper, document extraction, image to text, " mentioned.

日本語の概要は準備中です。原文の説明を表示しています。

omer-metin/skills-for-antigravity1642026年1月22日 更新

对 DiT(扩散/多模态生成)模型执行 W8A8 动态(data-free)或 W8A8 MXFP8(data-free)量化。覆盖 HunyuanVideo、Wan2.2-T2V/I2V/TI2V、FLUX.1-dev(已迁移)、SD3、Sana、HunyuanDiT、CogViewX 等。YAML apiversion 固定 multimodal_sd_modelslim_v1,统一 MultimodalPipelineInterface 重构路径(不再保留 LegacyMultimodalPipelineInterface)。模板与字段以 msmodelslim/model/hunyuan_video/model_adapter.py(唯一已迁移到重构路径且 lab_practice YAML 已 validated 的 DiT)为参考;其它 DiT 需按各自推理仓的 `parse_args` 调整字段集合。data-free 指不引入外部激活校准数据,spec.dataset(校准集短名,短名如 wan2_2_t2v;enable_dump: false 时不参与 dump)仍必填;精度调优通过下游经验库 L2 §7 + 直接调 `quantization-expert-experience-tuning-rules/scripts/apply_rollback.py`(整层回退)扩展 exclude 列表,闭环由 `quantization-accuracy-tuning-orchestrator` 调度。

日本語の概要は準備中です。原文の説明を表示しています。

kali20gakki/msAgent312026年10月10日 更新

Build multimodal AI applications with vision and text. TRIGGERS - Use when user needs help with ai-multimodal-app related tasks.

日本語の概要は準備中です。原文の説明を表示しています。

lionelsimai/claude-skills-collection292026年2月8日 更新

aesthetic

無料

Create aesthetically beautiful interfaces following proven design principles. Use when building UI/UX, analyzing designs from inspiration sites, generating design images with ai-multimodal, implementing visual hierarchy and color theory, adding micro-interactions, or creating design documentation. Includes workflows for capturing and analyzing inspiration screenshots with chrome-devtools and ai-multimodal, iterative design image generation until aesthetic standards are met, and comprehensive design system guidance covering BEAUTIFUL (aesthetic principles), RIGHT (functionality/accessibility), SATISFYING (micro-interactions), and PEAK (storytelling) stages. Integrates with chrome-devtools, ai-multimodal, media-processing, ui-styling, and web-frameworks skills.

日本語の概要は準備中です。原文の説明を表示しています。

VoDaiLocz/kilo-kit-mcp272026年9月13日 更新

Guides agents to interactively discover customer requirements for live, bidirectional multi-agent AI systems that process continuous streams of multimodal data for real-time technical guidance and safety monitoring. Generates a custom Google Cloud solution that uses opinionated best practices and architecture guidance. Use when users need agentic assistance to design and create a multi-product solution in the cloud for live bidirectional multimodal streaming workloads. Don't use for simple text-based chat applications or workloads without real-time streaming requirements.

日本語の概要は準備中です。原文の説明を表示しています。

vaila-multimodaltoolbox/vaila192026年10月8日 更新

aesthetic

無料

Create aesthetically beautiful interfaces following proven design principles. Use when building UI/UX, analyzing designs from inspiration sites, generating design images with ai-multimodal, implementing visual hierarchy and color theory, adding micro-interactions, or creating design documentation. Includes workflows for capturing and analyzing inspiration screenshots with chrome-devtools and ai-multimodal, iterative design image generation until aesthetic standards are met, and comprehensive design system guidance covering BEAUTIFUL (aesthetic principles), RIGHT (functionality/accessibility), SATISFYING (micro-interactions), and PEAK (storytelling) stages. Integrates with chrome-devtools, ai-multimodal, media-processing, ui-styling, and web-frameworks skills.

日本語の概要は準備中です。原文の説明を表示しています。

kettleofketchup/DraftForge152026年9月14日 更新

aesthetic

無料

Create aesthetically beautiful interfaces following proven design principles. Use when building UI/UX, analyzing designs from inspiration sites, generating design images with ai-multimodal, implementing visual hierarchy and color theory, adding micro-interactions, or creating design documentation. Includes workflows for capturing and analyzing inspiration screenshots with chrome-devtools and ai-multimodal, iterative design image generation until aesthetic standards are met, and comprehensive design system guidance covering BEAUTIFUL (aesthetic principles), RIGHT (functionality/accessibility), SATISFYING (micro-interactions), and PEAK (storytelling) stages. Integrates with chrome-devtools, ai-multimodal, media-processing, ui-styling, and web-frameworks skills.

日本語の概要は準備中です。原文の説明を表示しています。

EthanYoQ/Skill-hub112026年10月5日 更新

Build multimodal AI applications with vision and text. TRIGGERS - Use when user needs help with ai-multimodal-app related tasks.

日本語の概要は準備中です。原文の説明を表示しています。

Winbda/claude-skills-collection42026年4月6日 更新

Framework for state-of-the-art sentence, text, and image embeddings. Provides 5000+ pre-trained models for semantic similarity, clustering, and retrieval. Supports multilingual, domain-specific, and multimodal models. Use for generating embeddings for RAG, semantic search, or similarity tasks. Best for production embedding generation.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Expert guidance for fine-tuning LLMs with LLaMA-Factory - WebUI no-code, 100+ models, 2/3/4/5/6/8-bit QLoRA, multimodal support

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Guides agents to interactively discover customer requirements for live, bidirectional multi-agent AI systems that process continuous streams of multimodal data for real-time technical guidance and safety monitoring. Generates a custom Google Cloud solution that uses opinionated best practices and architecture guidance. Use when users need agentic assistance to design and create a multi-product solution in the cloud for live bidirectional multimodal streaming workloads. Don't use for simple text-based chat applications or workloads without real-time streaming requirements.

日本語の概要は準備中です。原文の説明を表示しています。

google/skills2.1万2026年10月10日 更新

Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Framework for state-of-the-art sentence, text, and image embeddings. Provides 5000+ pre-trained models for semantic similarity, clustering, and retrieval. Supports multilingual, domain-specific, and multimodal models. Use for generating embeddings for RAG, semantic search, or similarity tasks. Best for production embedding generation.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Expert guidance for fine-tuning LLMs with LLaMA-Factory - WebUI no-code, 100+ models, 2/3/4/5/6/8-bit QLoRA, multimodal support

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年6月16日 更新

Generates BYO custom safety policies for NVIDIA Nemotron content-safety guardrails — Nemotron-Content-Safety-Reasoning-4B (text) and multimodal Nemotron-3-Content-Safety. Produces a Markdown policy, JSON taxonomy, and drop-in inference prompts. Maps rough words or an existing policy to V2 categories, adding custom categories or topic-following rules.

日本語の概要は準備中です。原文の説明を表示しています。

NVIDIA/skills3,5592026年10月10日 更新

Azure AI Content Understanding SDK for Python. Use for multimodal content extraction from documents, images, audio, and video. Triggers: "azure-ai-contentunderstanding", "ContentUnderstandingClient", "multimodal analysis", "document extraction", "video analysis", "audio transcription".

日本語の概要は準備中です。原文の説明を表示しています。

microsoft/skills3,1002026年10月10日 更新

Process and generate multimedia content using Google Gemini API. Capabilities include analyze audio files (transcription with timestamps, summarization, speech understanding, music/sound analysis up to 9.5 hours), understand images (captioning, object detection, OCR, visual Q&A, segmentation), process videos (scene detection, Q&A, temporal analysis, YouTube URLs, up to 6 hours), extract from documents (PDF tables, forms, charts, diagrams, multi-page), generate images (text-to-image, editing, composition, refinement). Use when working with audio/video files, analyzing images or screenshots, processing PDF documents, extracting structured data from media, creating images from text prompts, or implementing multimodal AI features. Supports multiple models (Gemini 2.5/2.0) with context windows up to 2M tokens.

日本語の概要は準備中です。原文の説明を表示しています。

mrgoonie/claudekit-skills2,2272026年4月3日 更新

SkillsBench task PR review — classifies the task track (standard / research / multimodal), runs static policy checks against the track-specific rubric, benchmarks the task across oracle plus Claude and Codex (with and without skills), audits trajectories for cheating and skill invocation, and produces a `pr-N-task-timestamp-run.txt` review report alongside a `prN.zip` bundle of trajectories. Use when reviewing a SkillsBench task PR (by number, branch, or local task path), when the user asks to review a task, run benchmarks on a PR, audit a submission, classify a task as research or multimodal track, or prepare a comment to post on a SkillsBench PR.

日本語の概要は準備中です。原文の説明を表示しています。

benchflow-ai/skillsbench1,8372026年7月24日 更新

Process and generate multimedia content using Google Gemini API. Capabilities include analyze audio files (transcription with timestamps, summarization, speech understanding, music/sound analysis up to 9.5 hours), understand images (captioning, object detection, OCR, visual Q&A, segmentation), process videos (scene detection, Q&A, temporal analysis, YouTube URLs, up to 6 hours), extract from documents (PDF tables, forms, charts, diagrams, multi-page), generate images (text-to-image, editing, composition, refinement). Use when working with audio/video files, analyzing images or screenshots, processing PDF documents, extracting structured data from media, creating images from text prompts, or implementing multimodal AI features. Supports multiple models (Gemini 2.5/2.0) with context windows up to 2M tokens.

日本語の概要は準備中です。原文の説明を表示しています。

Microck/ordinary-claude-skills4042026年9月7日 更新

Analyze and extract information from images using multimodal AI recognition. Triggered when users want to analyze, describe, or extract information from an image URL — image recognition, image analysis, image description, visual content understanding, OCR text recognition, or visual Q&A. When a user provides an image URL and asks questions about its visual content, this skill should be triggered even if they do not explicitly say "image recognition."

日本語の概要は準備中です。原文の説明を表示しています。

nexscope-ai/nexscope-ecommerce-skills852026年10月9日 更新

Master optimization protocol for Gemini Agent (Antigravity) to unlock native 2M+ to 10M+ long-context reasoning, Gemini 4 Pro & Gemini 4 Flash dynamic thinking budget control, native context caching, Multimodal Live API, and high-speed problem solving / Protokol optimasi utama untuk Gemini Agent (Antigravity) untuk mengaktifkan pemikiran long-context 2M+ hingga 10M+, kontrol thinking budget dinamis Gemini 4 Pro & Gemini 4 Flash, context caching native, Multimodal Live API, dan pemecahan masalah kecepatan tinggi.

日本語の概要は準備中です。原文の説明を表示しています。

roedyrustam/vibes-plug752026年10月9日 更新

Expert guide for Affective Computing, emotional AI, and real-time sentiment analysis through native multimodal tokens (voice intonation and facial micro-expressions) / Panduan ahli komputasi afektif, AI emosional, dan analisis sentimen real-time melalui token multimodal native.

日本語の概要は準備中です。原文の説明を表示しています。

roedyrustam/vibes-plug752026年10月9日 更新