本文へ移動
cccskills

「vision-language」の検索結果

23 件 ・ 関連度順

概要と使いどころ

Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年10月11日 更新

clip

無料

OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

llava

無料

Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.

日本語の概要は準備中です。原文の説明を表示しています。

davila7/claude-code-templates3.3万2026年10月11日 更新

llava

無料

Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年10月11日 更新

clip

無料

OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.

日本語の概要は準備中です。原文の説明を表示しています。

Orchestra-Research/AI-Research-SKILLs1.3万2026年10月11日 更新

Framework de pré-treinamento visão-linguagem que conecta codificadores de imagem congelados e LLMs. Use quando você precisar de legendagem de imagens, resposta a perguntas visuais, recuperação imagem-texto ou chat multimodal com desempenho zero-shot de última geração.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

clip

無料

OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.

日本語の概要は準備中です。原文の説明を表示しています。

huang-sh/DeepScience42026年7月15日 更新

llava

無料

Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.

日本語の概要は準備中です。原文の説明を表示しています。

Lord1Egypt/awesome-skill-forge22026年6月10日 更新

clip

無料

OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.

日本語の概要は準備中です。原文の説明を表示しています。

Lord1Egypt/awesome-skill-forge22026年6月10日 更新

llava

無料日本語概要

画像の内容を質問したり、説明文を作ったりするために、LLaVAの導入と利用を支援します。画像を見ながら続ける対話や、画像付き文書の理解にも使えるスキルです。

  • 写真の説明文を作りたいとき
  • 画像内の物や場面について質問したいとき
  • 画像について話せるチャットの試作
NousResearch/hermes-agent25.3万2026年10月11日 更新

Fine-tune vision-language models (VLMs) with supervised learning on image+text data. Use when adapting a VLM to a visual domain or task, configuring frozen-vision-tower LoRA, or debugging a VLM fine-tune that trains without learning.

日本語の概要は準備中です。原文の説明を表示しています。

wshobson/agents4万2026年10月5日 更新

dexbotic

無料

Use Dexbotic to prepare DexData, train and serve vision-language-action policies, evaluate checkpoints, and integrate explicitly external RL or robot backends.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3322026年9月3日 更新

Builds ML pipelines: EDA, features, leakage checks, evaluation. Use when doing data science or explaining fairness, privacy, speech, vision-language, or diffusion mechanics.

日本語の概要は準備中です。原文の説明を表示しています。

vasilyu1983/AI-Agents-public912026年10月5日 更新

vlm-ocr

無料

OCRs scanned or image-only corpora with vision-language models in three phases. Evaluate selects a system by measured CER or WER against stratified human ground truth. Run builds the production pipeline. Clean corrects raw OCR with LLM and rule-based passes, quality diagnostics, and span-level provenance. Use when choosing an OCR model for a corpus, measuring OCR accuracy, transcribing scans at scale, or correcting raw OCR. Born-digital documents with text layers go to doc-to-markdown.

日本語の概要は準備中です。原文の説明を表示しています。

scdenney/open-science-skills642026年10月8日 更新

vlm-ocr

無料

OCR scanned or image-only corpora with vision-language models in three phases: `evaluate` picks a system by measured CER/WER against stratified human ground truth, `run` builds the production pipeline, and `clean` corrects raw OCR with LLM and rule-based passes, diagnostics, and span-level provenance. Born-digital documents with a text layer go to $doc-to-markdown.

日本語の概要は準備中です。原文の説明を表示しています。

scdenney/open-science-skills642026年10月8日 更新

smolvlm

無料

Local vision-language model for image analysis using SmolVLM-2B

日本語の概要は準備中です。原文の説明を表示しています。

tdimino/claude-code-minoan412026年9月28日 更新

Semantic image-text matching with CLIP and alternatives. Use for image search, zero-shot classification, similarity matching. NOT for counting objects, fine-grained classification (celebrities, car models), spatial reasoning, or compositional queries. Activate on "CLIP", "embeddings", "image similarity", "semantic search", "zero-shot classification", "image-text matching".

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/windags-skills132026年10月1日 更新

llava

無料

Assistente de Linguagem Grande e Visão. Habilita ajuste de instrução visual e conversas baseadas em imagem. Combina encoder de visão CLIP com modelos de linguagem Vicuna/LLaMA. Suporta chat multi-turno com imagem, resposta a perguntas visuais e seguimento de instruções. Use para chatbots visão-linguagem ou tarefas de compreensão de imagem. Melhor para análise conversacional de imagem.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

clip

無料

Modelo da OpenAI que conecta visão e linguagem. Permite classificação de imagens com zero-shot, correspondência imagem-texto e recuperação cross-modal. Treinado em 400M pares imagem-texto. Use para busca de imagens, moderação de conteúdo ou tarefas visão-linguagem sem fine-tuning. Melhor para compreensão geral de imagens.

日本語の概要は準備中です。原文の説明を表示しています。

artubss/SKILLS-CLAUDE-CODE112026年5月17日 更新

clip

無料

OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.

日本語の概要は準備中です。原文の説明を表示しています。

ibragimov-oasis/vibe-coder22026年6月24日 更新

llava

無料

Large Language and Vision Assistant. Enables visual instruction tuning and image-based conversations. Combines CLIP vision encoder with Vicuna/LLaMA language models. Supports multi-turn image chat, visual question answering, and instruction following. Use for vision-language chatbots or image understanding tasks. Best for conversational image analysis.

日本語の概要は準備中です。原文の説明を表示しています。

ibragimov-oasis/vibe-coder22026年6月24日 更新

Semantic image-text matching with CLIP and alternatives. Use for image search, zero-shot classification, similarity matching. NOT for counting objects, fine-grained classification (celebrities, car models), spatial reasoning, or compositional queries. Activate on "CLIP", "embeddings", "image similarity", "semantic search", "zero-shot classification", "image-text matching".

日本語の概要は準備中です。原文の説明を表示しています。

curiositech/port-daddy22026年10月8日 更新