追加の学習データを用意せず、文章で指定した候補に画像を分類したり、内容に合う画像を検索したりします。画像と文章の類似度比較や、複数画像の一括処理も扱うスキルです。
- 文章のラベルで写真を分類したいとき
- 画像の内容を言葉で指定して探したいとき
- 画像と説明文の対応を比較したいとき
43 件 ・ 関連度順
概要と使いどころ
追加の学習データを用意せず、文章で指定した候補に画像を分類したり、内容に合う画像を検索したりします。画像と文章の類似度比較や、複数画像の一括処理も扱うスキルです。
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
日本語の概要は準備中です。原文の説明を表示しています。
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
日本語の概要は準備中です。原文の説明を表示しています。
Perform visual similarity search for design patents using an image URL, with filtering by country, legal status, date ranges, Locarno classification, and assignee. Supports design patent types only (type D). Triggered when users mention patent image search, design patent search, search patent by image, visual patent lookup, patent similarity detection, patent image matching, or design patent infringement check. Even if the user does not explicitly mention "patent image," this skill should be triggered whenever the need involves searching for similar design patents through an image. For utility model patents, use ecommerce.Nexscope-utility-patent-image-search instead.
日本語の概要は準備中です。原文の説明を表示しています。
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
日本語の概要は準備中です。原文の説明を表示しています。
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
日本語の概要は準備中です。原文の説明を表示しています。
Modelo da OpenAI que conecta visão e linguagem. Permite classificação de imagens com zero-shot, correspondência imagem-texto e recuperação cross-modal. Treinado em 400M pares imagem-texto. Use para busca de imagens, moderação de conteúdo ou tarefas visão-linguagem sem fine-tuning. Melhor para compreensão geral de imagens.
日本語の概要は準備中です。原文の説明を表示しています。
Semantic image-text matching with CLIP and alternatives. Use for image search, zero-shot classification, similarity matching. NOT for counting objects, fine-grained classification (celebrities, car models), spatial reasoning, or compositional queries. Activate on "CLIP", "embeddings", "image similarity", "semantic search", "zero-shot classification", "image-text matching".
日本語の概要は準備中です。原文の説明を表示しています。
Semantic image-text matching with CLIP and alternatives. Use for image search, zero-shot classification, similarity matching. NOT for counting objects, fine-grained classification (celebrities, car models), spatial reasoning, or compositional queries. Activate on "CLIP", "embeddings", "image similarity", "semantic search", "zero-shot classification", "image-text matching".
日本語の概要は準備中です。原文の説明を表示しています。
Build image generation pipelines with Stable Diffusion, FLUX, ControlNet, LoRA, and ComfyUI workflows. Activate on: image generation pipeline, ComfyUI workflow, ControlNet, LoRA training, diffusion model. NOT for: video generation (ai-video-production-master), image classification (computer-vision-pipeline).
日本語の概要は準備中です。原文の説明を表示しています。
OpenAI's model connecting vision and language. Enables zero-shot image classification, image-text matching, and cross-modal retrieval. Trained on 400M image-text pairs. Use for image search, content moderation, or vision-language tasks without fine-tuning. Best for general-purpose image understanding.
日本語の概要は準備中です。原文の説明を表示しています。
Trains and fine-tunes vision models for object detection (D-FINE, RT-DETR v2, DETR, YOLOS), image classification (timm models — MobileNetV3, MobileViT, ResNet, ViT/DINOv3 — plus any Transformers classifier), and SAM/SAM2 segmentation using Hugging Face Transformers on Hugging Face Jobs cloud GPUs. Covers COCO-format dataset preparation, Albumentations augmentation, mAP/mAR evaluation, accuracy metrics, SAM segmentation with bbox/point prompts, DiceCE loss, hardware selection, cost estimation, Trackio monitoring, and Hub persistence. Use when users mention training object detection, image classification, SAM, SAM2, segmentation, image matting, DETR, D-FINE, RT-DETR, ViT, timm, MobileNet, ResNet, bounding box models, or fine-tuning vision models on Hugging Face Jobs.
日本語の概要は準備中です。原文の説明を表示しています。
Work with hyperspectral and multispectral images in MATLAB. Covers reading/writing (ENVI, NITF, TIFF, Sentinel-2, Landsat, ASTER), ECOSTRESS spectral libraries, processing (calibration, atmospheric correction, denoising, sharpening, dimensionality reduction, endmember extraction, unmixing, target/anomaly detection, spectral indices, segmentation), labeling (Spectral Image Labeler app, ground truth objects, automation algorithms), and deep learning (pixel classification CNNs, unmixing autoencoders, transfer learning). Use when reading, writing, processing, analyzing, classifying, labeling, or applying deep learning to hyperspectral or multispectral images.
日本語の概要は準備中です。原文の説明を表示しています。
Trains and fine-tunes vision models for object detection (D-FINE, RT-DETR v2, DETR, YOLOS), image classification (timm models — MobileNetV3, MobileViT, ResNet, ViT/DINOv3 — plus any Transformers classifier), and SAM/SAM2 segmentation using Hugging Face Transformers on Hugging Face Jobs cloud GPUs. Covers COCO-format dataset preparation, Albumentations augmentation, mAP/mAR evaluation, accuracy metrics, SAM segmentation with bbox/point prompts, DiceCE loss, hardware selection, cost estimation, Trackio monitoring, and Hub persistence. Use when users mention training object detection, image classification, SAM, SAM2, segmentation, image matting, DETR, D-FINE, RT-DETR, ViT, timm, MobileNet, ResNet, bounding box models, or fine-tuning vision models on Hugging Face Jobs.
日本語の概要は準備中です。原文の説明を表示しています。
Use Transformers.js to run state-of-the-art machine learning models directly in JavaScript/TypeScript. Supports NLP (text classification, translation, summarization), computer vision (image classification, object detection), audio (speech recognition, audio classification), and multimodal tasks. Works in browsers and server-side runtimes (Node.js, Bun, Deno) with WebGPU/WASM using pre-trained models from Hugging Face Hub.
日本語の概要は準備中です。原文の説明を表示しています。
Use when creating high-quality scientific or medical infographic images or prompts, including review figures, disease maps, teaching diagrams, mechanism summaries, classification overviews, timelines, matrices, and academic visual explainers. If the user asks to generate the image, use ChatGPT image generation.
日本語の概要は準備中です。原文の説明を表示しています。
Use Transformers.js to run state-of-the-art machine learning models directly in JavaScript/TypeScript. Supports NLP (text classification, translation, summarization), computer vision (image classification, object detection), audio (speech recognition, audio classification), and multimodal tasks. Works in browsers and server-side runtimes (Node.js, Bun, Deno) with WebGPU/WASM using pre-trained models from Hugging Face Hub.
日本語の概要は準備中です。原文の説明を表示しています。
This skill should be used when working with pre-trained transformer models for natural language processing, computer vision, audio, or multimodal tasks. Use for text generation, classification, question answering, translation, summarization, image classification, object detection, speech recognition, and fine-tuning models on custom datasets.
日本語の概要は準備中です。原文の説明を表示しています。
Azure AI Content Safety SDK for Python. Use for detecting harmful content in text and images with multi-severity classification. Triggers: "azure-ai-contentsafety", "ContentSafetyClient", "content moderation", "harmful content", "text analysis", "image analysis".
日本語の概要は準備中です。原文の説明を表示しています。
Core ML, Create ML, Vision framework, Natural Language framework, on-device ML integration. Use when user wants image classification, text analysis, object detection, sound classification, model optimization, or custom model integration. Covers Core ML vs Foundation Models decision.
日本語の概要は準備中です。原文の説明を表示しています。
Azure AI Content Safety SDK for Python. Use for detecting harmful content in text and images with multi-severity classification. Triggers: "azure-ai-contentsafety", "ContentSafetyClient", "content moderation", "harmful content", "text analysis", "image analysis".
日本語の概要は準備中です。原文の説明を表示しています。
Zero-shot image classification and image-text search.
日本語の概要は準備中です。原文の説明を表示しています。
Azure AI Content Safety SDK for Python. Use for detecting harmful content in text and images with multi-severity classification.
日本語の概要は準備中です。原文の説明を表示しています。
Performs deep Root Cause Analysis (RCA) on NVIDIA TAO Visual ChangeNet classification experiments with image-evidence-driven investigation. Use when analyzing ChangeNet model failures, investigating poor recall / FAR / PASS-NO_PASS metrics, auditing visual inspection pipeline quality, or running an RCA report for an AOI defect-detection model. Trigger phrases include "RCA on my ChangeNet model", "why is my AOI model failing", "audit ChangeNet predictions", "investigate FAR regressions", "root cause analysis on visual-changenet".
日本語の概要は準備中です。原文の説明を表示しています。