Build integrated IS/BS/CF financial workbooks in Excel.
日本語の概要は準備中です。原文の説明を表示しています。
vision-audio: screenshot, describe, transcribe, diagram.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
scripts/vision-audio.py ist ein CLI-Tool für:
--screenshot)--describe)--diagram)--transcribe)--record)Alle Ausgaben erfolgen als strukturiertes JSON.
pip install mss PyAudio openai-whisper Pillow
usage: vision-audio.py [-h] [--screenshot] [--describe BILD]
[--diagram DIAGRAMM] [--transcribe AUDIO] [--record]
[--seconds SECONDS] [--json]
Vision + Audio Pipeline — Multimodal Understanding
options:
-h, --help show this help message and exit
--screenshot Screenshot aufnehmen + beschreiben
--describe BILD Bilddatei beschreiben (Pixel-Analyse)
--diagram DIAGRAMM Code-Diagramm-Screenshot analysieren
--transcribe AUDIO Audiodatei transkribieren
--record Mikrofon aufnehmen + transkribieren
--seconds SECONDS Aufnahmedauer in Sekunden (Default: 10)
--json Nur JSON ausgeben (keine Stderr-Info)
# Screenshot + beschreiben
python scripts/vision-audio.py --screenshot
# Bild beschreiben
python scripts/vision-audio.py --describe screenshot.png
# Code-Diagramm analysieren
python scripts/vision-audio.py --diagram code-flow.png
# Audio transkribieren
python scripts/vision-audio.py --transcribe meeting.wav
# Mikrofon aufnehmen + transkribieren
python scripts/vision-audio.py --record
python scripts/vision-audio.py --record --seconds 30
vision-audio.py
├── Vision
│ ├── vision_screenshot() → MSS / PowerShell-Fallback
│ ├── vision_describe() → Pillow Pixelanalyse + Farbextraktion
│ └── vision_diagram() → Kontrast-basierte Text-Bereichs-Detektion
├── Audio
│ ├── audio_transcribe() → openai-whisper (Model: base)
│ └── audio_record() → PyAudio → WAV → whisper
└── CLI (argparse)
{
"action": "describe|screenshot|transcribe|diagram|record",
"status": "ok|error",
"timestamp": "2026-08-22T00:40:21.551751",
"text": "Bild-Deskription / Transkription / Diagramm …",
"file_path": "path/to/file.png",
"metadata": {
"width": 1920,
"height": 1200,
"format": "PNG",
"mode": "RGB",
"avg_brightness": 0.635,
"contrast": 0.883,
"total_pixels": 26666,
"dominant_colors": [["rgb(240,240,245)", 820], ...],
"thumbnail_base64": "iVBOR..."
}
}
--transcribe lädt Whisper das base-Modell herunter (~150 MB).~/.openamer/vision-audio/.まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Build integrated IS/BS/CF financial workbooks in Excel.
日本語の概要は準備中です。原文の説明を表示しています。
Run and grow the OpenAmer Agent-to-Agent swarm: identity, trust, node-to-node ask, signed skill/insight sharing, and the autonomous self-learning loop.
日本語の概要は準備中です。原文の説明を表示しています。
Use for A/B experiments on OpenAmer configs and skills.
日本語の概要は準備中です。原文の説明を表示しています。
Roleplay a hostile user to find and triage UX pain points.
日本語の概要は準備中です。原文の説明を表示しています。
Agent mesh: master/worker nodes, HTTP delegation, heartbeat.
日本語の概要は準備中です。原文の説明を表示しています。
Use for agentic commerce payments tasks. Candidate synthesized from trend radar; awaiting trial promotion.
日本語の概要は準備中です。原文の説明を表示しています。