本文へ移動
cccskills
無料GitHub で公開

vision

Query images with a local Ollama vision model without loading the image into the main agent context. Use when you need to describe a screenshot, check whether rendered content is present, detect overlapping elements, or ask any visual question about a PNG/JPEG/WebP file. Requires Ollama running locally with the Gemma 4 multimodal model (`gemma4` on Ollama). Script: .agents/skills/vision/scripts/ask.py. Trigger phrases: "describe image", "what does this screenshot show", "does the canvas contain content", "check screenshot visually", "look at this image", "any overlapping elements", "vision query".

インストール方法を見る

含まれるファイル(2)

  • SKILL.md5.8 KB
  • scripts/ask.py9.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Vision — Local Image Querying via Ollama

Ask natural-language questions about images without passing them to the main agent as visual input. Useful for verifying screenshots, annotating assets, or building automated checks around visual output.

When to Use This Skill

  • Describing a screenshot for a PR description or user-facing document
  • Checking whether an automated browser run produced visible canvas content
  • Asking "do any elements overlap?" on a rendered output
  • Any question where the answer is in the pixels but you don't want to use vision tokens in the main context

Quick Reference

All commands use uv run — dependencies are installed automatically.

SCRIPT=.agents/skills/vision/scripts/ask.py

# health check (fast, no image, confirms Ollama + model respond)
uv run $SCRIPT --ping

# system info — memory, storage, installed models
uv run $SCRIPT --info
uv run $SCRIPT --memory
uv run $SCRIPT --storage

# describe an image (default prompt)
uv run $SCRIPT path/to/image.png

# explicit shortcut
uv run $SCRIPT path/to/image.png describe

# custom question
uv run $SCRIPT path/to/image.png \
  --prompt "Do you see any overlapping UI elements?"

uv run $SCRIPT canvas.png \
  --prompt "Does this canvas contain any designed content, or is it empty?"

# optional: pin a specific Gemma 4 tag (default is any installed gemma4)
uv run $SCRIPT image.png --model gemma4:e4b

# list installed Gemma 4 vision models
uv run $SCRIPT --list-models

Prerequisites

Ollama must be running locally. The script connects to http://localhost:11434 and fails immediately if it cannot reach it.

# start Ollama (if not already running)
ollama serve

# install Gemma 4 (multimodal — required for this skill)
ollama pull gemma4

The script does not install models. If Gemma 4 is not installed it prints the list of installed models and a pull suggestion, then exits.

uv is required to run the script (handles dependency installation automatically). No requirements.txt or manual pip install needed.


Model Selection

This skill uses only Gemma 4 on Ollama (gemma4 and tags such as gemma4:latest, gemma4:e4b). Other multimodal models are ignored so agents do not silently fall back to a different family.

When --model is omitted, the script picks any installed gemma4 tag (for example gemma4:latest). Use --model gemma4:e4b (or another tag) to pin a specific variant.


System Info

Before running a heavy query, check whether the machine has enough resources. This is optional — the script does not enforce limits — but useful context for deciding whether to proceed or skip.

uv run $SCRIPT --info      # memory + storage + model list
uv run $SCRIPT --memory    # just memory
uv run $SCRIPT --storage   # just storage

Tip: on machines with ≤8 GB RAM, large vision models may cause swapping or OOM. Consider a smaller Gemma 4 variant (for example gemma4:e2b) or skip the query.


Behavior

  • Fails fast if Ollama is unreachable or Gemma 4 is not installed. Exit code is non-zero; the error message includes a hint or pull command.
  • Sequential only — Ollama is a single-worker process. Never call ask.py in parallel (e.g. two concurrent tool calls). Queue calls one at a time.
  • No side effects beyond the local Ollama process.
  • Auto-installs deps via uv inline script metadata (PEP 723). Only dependency is the ollama Python package.
  • Supported formats: .png, .jpg, .jpeg, .webp, .gif, .bmp.

Typical Agent Workflow

  1. A tool (browser automation, screenshot capture, golden renderer) writes an image to disk.
  2. Call ask.py with a targeted prompt suited to the task.
  3. Parse the text response to decide the next action.
# Quick sanity check first
uv run $SCRIPT --ping

# Verify a browser screenshot has content before including it in a doc
uv run $SCRIPT /tmp/preview.png \
  --prompt "Answer with YES or NO: does this screenshot show any visible UI content, shapes, or text?"

# Describe a captured screenshot for a PR description
uv run $SCRIPT /tmp/canvas-screenshot.png \
  --prompt "Describe what visual effect is shown. Be specific about blur, colors, and shapes."

Troubleshooting

SymptomCauseFix
cannot reach OllamaOllama not runningollama serve
no Gemma 4 vision model foundGemma 4 not installedollama pull gemma4
model 'X' is not availableModel name typo or not installed--list-models to see what's installed
Slow responseLarge model on CPUTry a smaller tag (e.g. gemma4:e2b)
Vague or wrong answerGeneric promptWrite a more specific --prompt
'ollama' package not foundNot using uv runRun with uv run ask.py instead

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Grida AI agent system work: `@grida/daemon` (DaemonServer, loopback HTTP perimeter, files/workspaces, secrets store, daemon discovery) and `@grida/agent` (the agent tenant: sessions, providers/BYOK, runtime/tool execution, skills discovery, prompts, tiers, sandbox hosts). Use for `packages/grida-daemon/**`, `packages/grida-ai-agent/**`, desktop sidecar protocol changes, agent chat transport, and bugs in agent state or streams. For pure Electron window, preload, menu, deep-link, or CDP work, use `desktop`.

日本語の概要は準備中です。原文の説明を表示しています。

gridaco/grida2,6652026年10月9日 更新

ai-models

無料

Research, compare, and update shared AI model JSON for TypeScript, web, and Rust consumers. Covers text model tiers, image and video generation models, image tool models, release provenance, pricing data sourcing, and provider-cost metering against prepaid org credit. Use when bumping model versions, adding new models, updating pricing, or auditing model specs against provider documentation.

日本語の概要は準備中です。原文の説明を表示しています。

gridaco/grida2,6652026年10月9日 更新

React-specific code shape in the Grida editor. Hooks cannot be tested or benchmarked and silently break tuned UX under layered composition, so they are barred from the engine and main system — load-bearing logic lives in classes and namespaces, hooks only as thin edge wires. `data-testid` follows component-root discipline: one per significant component, not scattered. Use when authoring React in `editor/grida-canvas-react/`, `editor/components/`, `editor/scaffolds/`, or `editor/app/*`.

日本語の概要は準備中です。原文の説明を表示しています。

gridaco/grida2,6652026年10月9日 更新

code-ts

無料

TypeScript code shape inside a well-named module — taste, not lint. Prefer one class or namespace per file (the unit a test targets) over scattered free exports; consolidate related code, don't fragment. The unit of code should be the unit of spec. Use when authoring TS in `editor/grida-canvas*`, `editor/lib/`, or `packages/*`. Sibling to the `naming` skill; React-specific shape lives in `code-react`.

日本語の概要は準備中です。原文の説明を表示しています。

gridaco/grida2,6652026年10月9日 更新

database

無料

Use BEFORE editing any file in `supabase/migrations/` or `supabase/schemas/`, OR when the user runs a `/database` subcommand (`compact local migration`, `rls scenarios`, `align`). Encodes the three contracts that protect the Grida database layer: applied migrations are immutable, RLS implementation mirrors tests (never the reverse), `schemas/*.sql` is the human-readable end-state. Companion to `supabase/AGENTS.md` (RLS, grants, security boundaries).

日本語の概要は準備中です。原文の説明を表示しています。

gridaco/grida2,6652026年10月9日 更新

desktop

無料

Grida Desktop Electron shell and release-impact work: BrowserWindow, preload, `window.grida`, menus, protocol/deep links, file associations, Forge, path-scoped bridge security, Electron-only UI bugs, and CDP / Playwright verification. Use for `desktop/`, `editor/app/desktop/**`, `editor/scaffolds/desktop/**`, `editor/lib/desktop/**`, `/desktop/*` CSP, GRIDA-SEC-004, and deciding whether linked-package or hosted-renderer changes require a native Desktop version bump or coordinated release. For implementing daemon/agent-tenant core behavior, use `agent-system` as well.

日本語の概要は準備中です。原文の説明を表示しています。

gridaco/grida2,6652026年10月9日 更新

gridaco のスキルをすべて見る

このスキルの問題を報告する