本文へ移動
cccskills
無料GitHub で公開

docarray

Use DocArray for multimodal Pydantic-style documents, typed DocList and DocVec batches, serialization and local storage, FastAPI payloads, and vector retrieval indexes.

インストール方法を見る

含まれるファイル(21)

  • SKILL.md5.0 KB
  • references/repo-provenance.md3.0 KB
  • references/repo-routing-metadata.json398 B
  • references/troubleshooting.md2.8 KB
  • scripts/check_env.py2.2 KB
  • sub-skills/document-modeling/references/api-reference.md11.2 KB
  • sub-skills/document-modeling/references/troubleshooting.md7.8 KB
  • sub-skills/document-modeling/references/workflows.md8.3 KB
  • sub-skills/document-modeling/scripts/schema_smoke.py6.7 KB
  • sub-skills/document-modeling/SKILL.md2.4 KB
  • sub-skills/serialization-storage/references/serialization-reference.md11.0 KB
  • sub-skills/serialization-storage/references/storage-reference.md8.2 KB
  • sub-skills/serialization-storage/references/troubleshooting.md8.8 KB
  • sub-skills/serialization-storage/scripts/roundtrip_smoke.py11.4 KB
  • sub-skills/serialization-storage/SKILL.md2.9 KB
  • sub-skills/vector-indexing/references/api-reference.md4.0 KB
  • sub-skills/vector-indexing/references/index-workflows.md4.8 KB
  • sub-skills/vector-indexing/references/optional-backends.md3.8 KB
  • sub-skills/vector-indexing/references/troubleshooting.md4.5 KB
  • sub-skills/vector-indexing/scripts/inmemory_index_smoke.py4.3 KB
  • sub-skills/vector-indexing/SKILL.md2.6 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

DocArray

Use this repo skill when a task names DocArray or asks for a Python data model for multimodal records, typed document batches, document serialization, local document storage, FastAPI document payloads, or vector-index retrieval.

Verified baseline

The verified default is a CPU workflow using DocArray base dependencies plus the proto, pandas, and web extras. Install the public package for normal use:

pip install -U docarray
# add only the surfaces you need:
pip install -U "docarray[proto,pandas,web]"
python -c "import docarray; print(docarray.__version__)"

For the current verified DocVec path, prefer numpy<2 and run the bundled smoke helpers before trusting a new environment. Optional tensor frameworks, media loaders, cloud stores, and external vector databases are not implied by the base install.

Route by task

  • Model a record or multimodal schema: read document-modeling for BaseDoc, predefined modality docs, dynamic schemas, DocList, DocVec, nested fields, and typed tensor shapes.
  • Serialize, store, stream, or serve documents: read serialization-storage for JSON, protobuf, bytes, base64, binary files, CSV/DataFrame exchange, file://, S3 boundaries, and DocArrayResponse.
  • Index, search, filter, or persist vectors: read vector-indexing for InMemoryExactNNIndex, query builders, subindexes, persistence, and optional backend selection.

Tasks that span multiple surfaces should start here, choose the schema in document-modeling, then hand the resulting typed documents to serialization-storage or vector-indexing.

Fast decisions

NeedFirst choiceWatch for
One validated data pointBaseDoc subclassRequired fields, nested docs, and per-document tensor shapes.
Mutable/reorderable/streaming collectionDocList[MyDoc]Typed lists are homogeneous; bare DocList may be heterogeneous.
Contiguous ML batchDocVec[MyDoc]Homogeneous fields; optional doc/tensor columns must be all present or all missing.
Human-readable transportJSONRich tensor/list unions may need explicit schema handling.
Compact trusted transportprotobuf or protobuf-arrayInstall docarray[proto]; document unions are not protobuf-safe.
Local retrieval prototypeInMemoryExactNNIndex[MyDoc]Dimensioned vector field and backend-specific filter/query behavior.
Production/vector serviceOptional backendInstall and verify the chosen client, service, credentials, schema, and metric separately.

Common failure boundaries

  • Missing google.protobuf, pandas, fastapi, or backend client: install the matching narrow extra; do not install full by default.
  • DocVec fails with a NumPy device error: check NumPy compatibility and try numpy<2 for the current verified CPU path.
  • File-store push raises FileNotFoundError: create the explicit parent namespace directory first.
  • CSV cannot rebuild tensor fields: use JSON, protobuf-array, binary, or DataFrame instead; CSV is safest for scalar rows.
  • In-memory query builder does not support text search, and equal-score result ties can expose a comparison edge case; see the vector troubleshooting reference.

Read references/troubleshooting.md for cross-cutting recovery guidance and references/repo-provenance.md before deciding whether this skill matches a changed checkout.

Routing metadata path convention

In references/repo-routing-metadata.json, every useful_entry_points value is a path relative to this DocArray skill root (the directory containing this SKILL.md), not a repository- or bundle-prefixed path.

Safe bundled checks

These helpers are self-contained and do not require the original repository checkout:

python sub-skills/document-modeling/scripts/schema_smoke.py --help
python sub-skills/serialization-storage/scripts/roundtrip_smoke.py --help
python sub-skills/vector-indexing/scripts/inmemory_index_smoke.py --help
python scripts/check_env.py --help

Run the relevant helper after installing the package and selected extras. Helpers use tiny in-memory or temporary-file fixtures; they do not start databases, use credentials, or download models.

Public capability boundaries

DocArray also exposes Torch, TensorFlow, JAX, image/audio/video/mesh loaders, S3, HNSWLib, Qdrant, Weaviate, Elasticsearch, Redis, Milvus, MongoDB Atlas, Epsilla, and Jina/FastAPI integrations. Those are routed by the sub-skills but remain optional until their exact dependency variant, service/network/credential plan, and native smoke have been verified.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Routes 3D ResNets PyTorch video action-recognition workflows across training, inference, and data preparation.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

3ddfa

無料

Guide 3DDFA Python inference, geometry rendering, training/evaluation, and optional C++ ONNX workflows for 3D dense face alignment.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

3ddfa-v2

無料

Routes 3DDFA_V2 face-alignment setup, still-image demos, video tracking, and ONNX benchmarking workflows.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

ab3dmot

無料

Operate AB3DMOT 3D multi-object tracking workflows for KITTI and nuScenes data, tracking, evaluation, and visualization.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

Use Hugging Face Accelerate for PyTorch training-loop migration, distributed launch/configuration, DeepSpeed/FSDP/TPU backend setup, big-model inference/offload, checkpointing, tracking, and troubleshooting.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

acme

無料

Route Acme reinforcement-learning framework tasks across core loops, replay/data, JAX agents, and TensorFlow/Sonnet agents.

日本語の概要は準備中です。原文の説明を表示しています。

VectorSpaceLab/AREX-Skill3312026年9月3日 更新

VectorSpaceLab のスキルをすべて見る

このスキルの問題を報告する