追加の学習データを用意せず、文章で指定した候補に画像を分類したり、内容に合う画像を検索したりします。画像と文章の類似度比較や、複数画像の一括処理も扱うスキルです。
- 文章のラベルで写真を分類したいとき
- 画像の内容を言葉で指定して探したいとき
- 画像と説明文の対応を比較したいとき
36 件 ・ 関連度順
概要と使いどころ
追加の学習データを用意せず、文章で指定した候補に画像を分類したり、内容に合う画像を検索したりします。画像と文章の類似度比較や、複数画像の一括処理も扱うスキルです。
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
日本語の概要は準備中です。原文の説明を表示しています。
Foundation model for image segmentation with zero-shot transfer. Use when you need to segment any object in images using points, boxes, or masks as prompts, or automatically generate all object masks in an image.
日本語の概要は準備中です。原文の説明を表示しています。
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
日本語の概要は準備中です。原文の説明を表示しています。
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
日本語の概要は準備中です。原文の説明を表示しています。
Vision-language pre-training framework bridging frozen image encoders and LLMs. Use when you need image captioning, visual question answering, image-text retrieval, or multimodal chat with state-of-the-art zero-shot performance.
日本語の概要は準備中です。原文の説明を表示しています。
Foundation model for image segmentation with zero-shot transfer. Use when you need to segment any object in images using points, boxes, or masks as prompts, or automatically generate all object masks in an image.
日本語の概要は準備中です。原文の説明を表示しています。
Semantic image-text matching with CLIP and alternatives. Use for image search, zero-shot classification, similarity matching. NOT for counting objects, fine-grained classification (celebrities, car models), spatial reasoning, or compositional queries. Activate on "CLIP", "embeddings", "image similarity", "semantic search", "zero-shot classification", "image-text matching".
日本語の概要は準備中です。原文の説明を表示しています。
Modelo de fundação para segmentação de imagens com transferência zero-shot. Use quando precisar segmentar qualquer objeto em imagens usando pontos, caixas ou máscaras como prompts, ou gerar automaticamente todas as máscaras de objetos em uma imagem.
日本語の概要は準備中です。原文の説明を表示しています。
Framework de pré-treinamento visão-linguagem que conecta codificadores de imagem congelados e LLMs. Use quando você precisar de legendagem de imagens, resposta a perguntas visuais, recuperação imagem-texto ou chat multimodal com desempenho zero-shot de última geração.
日本語の概要は準備中です。原文の説明を表示しています。
Modelo da OpenAI que conecta visão e linguagem. Permite classificação de imagens com zero-shot, correspondência imagem-texto e recuperação cross-modal. Treinado em 400M pares imagem-texto. Use para busca de imagens, moderação de conteúdo ou tarefas visão-linguagem sem fine-tuning. Melhor para compreensão geral de imagens.
日本語の概要は準備中です。原文の説明を表示しています。
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
日本語の概要は準備中です。原文の説明を表示しています。
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
日本語の概要は準備中です。原文の説明を表示しています。
Semantic image-text matching with CLIP and alternatives. Use for image search, zero-shot classification, similarity matching. NOT for counting objects, fine-grained classification (celebrities, car models), spatial reasoning, or compositional queries. Activate on "CLIP", "embeddings", "image similarity", "semantic search", "zero-shot classification", "image-text matching".
日本語の概要は準備中です。原文の説明を表示しています。
SAM: zero-shot image segmentation via points, boxes, masks.
日本語の概要は準備中です。原文の説明を表示しています。
Performs zero-shot time-series forecasting with Google's TimesFM, including regular-grid CSV preparation, quantile forecasts, XReg covariates, and held-out evaluation. Uses the Apache-licensed TimesFM 2.5 checkpoint by default and documents the distinct TimesFM 3.0 multivariate API and weight-license requirements.
日本語の概要は準備中です。原文の説明を表示しています。
Zero-shot time series forecasting with Google's TimesFM foundation model. Use this skill when forecasting ANY univariate time series — sales, sensor readings, stock prices, energy demand, patient vitals, weather, or scientific measurements — without training a custom model. Supports both basic forecasting and advanced covariate forecasting (XReg) with dynamic and static exogenous variables. Automatically checks system RAM/GPU before loading the model, validates dataset fit before processing, supports CSV/DataFrame/array inputs, and returns point forecasts with calibrated prediction intervals. Includes a preflight system checker script that MUST be run before first use to verify the machine can load the model and handle your specific dataset.
日本語の概要は準備中です。原文の説明を表示しています。
Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology. Use this skill when: (1) Producing cell embeddings from an AnnData for clustering/integration, (2) Zero-shot or fine-tuned cell-type annotation, (3) Gene-level representation for perturbation/GRN tasks. For probabilistic single-cell models (scVI etc.), use the scvi-tools library.
日本語の概要は準備中です。原文の説明を表示しています。
Zero-shot time series forecasting with Google's TimesFM foundation model. Use for any univariate time series (sales, sensors, energy, vitals, weather) without training a custom model. Supports CSV/DataFrame/array inputs with point forecasts and prediction intervals. Includes a preflight system checker script to verify RAM/GPU before first use.
日本語の概要は準備中です。原文の説明を表示しています。
Zero-shot time series forecasting with Google's TimesFM foundation model. Use this skill when forecasting ANY univariate time series — sales, sensor readings, stock prices, energy demand, patient vitals, weather, or scientific measurements — without training a custom model. Automatically checks system RAM/GPU before loading the model, supports CSV/DataFrame/array inputs, and returns point forecasts with calibrated prediction intervals. Includes a preflight system checker script that MUST be run before first use to verify the machine can load the model. For classical statistical time series models (ARIMA, SARIMAX, VAR) use statsmodels; for time series classification/clustering use aeon.
日本語の概要は準備中です。原文の説明を表示しています。
LLM-based zero-shot and few-shot classification for flexible intent detection
日本語の概要は準備中です。原文の説明を表示しています。
Use when tackling complex reasoning tasks requiring step-by-step logic, multi-step arithmetic, commonsense reasoning, symbolic manipulation, or problems where simple prompting fails - provides comprehensive guide to Chain-of-Thought and related prompting techniques (Zero-shot CoT, Self-Consistency, Tree of Thoughts, Least-to-Most, ReAct, PAL, Reflexion) with templates, decision matrices, and research-backed patterns
日本語の概要は準備中です。原文の説明を表示しています。
Launch an intelligent sub-agent with automatic model selection based on task complexity, specialized agent matching, Zero-shot CoT reasoning, and mandatory self-critique verification
日本語の概要は準備中です。原文の説明を表示しています。
Engineer effective LLM prompts using zero-shot, few-shot, chain-of-thought, and structured output techniques. Use when building LLM applications requiring reliable outputs, implementing RAG systems, creating AI agents, or optimizing prompt quality and cost. Covers OpenAI, Anthropic, and open-source models with multi-language examples (Python/TypeScript).
日本語の概要は準備中です。原文の説明を表示しています。