本文へ移動
cccskills
無料GitHub で公開

scgpt

Embed and annotate single-cell expression data with scGPT, a foundation model for single-cell biology. Use this skill when: (1) Producing cell embeddings from an AnnData for clustering/integration, (2) Zero-shot or fine-tuned cell-type annotation, (3) Gene-level representation for perturbation/GRN tasks. For probabilistic single-cell models (scVI etc.), use the scvi-tools library.

インストール方法を見る

含まれるファイル(2)

  • SKILL.md5.8 KB
  • .catalog_stamp13 B

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

scGPT — Single-Cell Foundation Model

Prerequisites

RequirementMinimumRecommended
Python3.10+3.11
CUDA12.1+12.4+
GPU VRAM16 GB24 GB+

How to run

Loading the vocabulary and checkpoint

scGPT checkpoints are raw directories (args.json, best_model.pt, vocab.json) — not Hugging Face hub repos. Point at the directory, not an HF repo id.

from scgpt.tokenizer.gene_tokenizer import GeneVocab
gv = GeneVocab.from_file("/path/to/scgpt-human/vocab.json")
print(len(gv))   # 60697 for the released human checkpoint

Embedding an AnnData

import anndata as ad
from scgpt.tasks import embed_data

adata = ad.read_h5ad("dataset.h5ad")        # var must contain a gene-name column
emb = embed_data(
    adata,
    model_dir="/path/to/scgpt-human",
    gene_col="feature_name",
    use_fast_transformer=False,             # see Gotchas
)
# emb is an AnnData with .obsm["X_scGPT"]

Output format

embed_data returns an AnnData whose .obsm["X_scGPT"] is the per-cell embedding (n_cells × emb_dim, 512 by default). Downstream: feed to scanpy.pp.neighbors / scanpy.tl.umap.

Remote compute

Needs ≥24 GB VRAM and the released human checkpoint (~200 MB: args.json, best_model.pt, vocab.json). Read compute_details({provider, mode:'read'}) for an environment with scgpt and a pre-cached checkpoint directory, then:

c = host.compute.create(provider)
job = c.submitJob(
    intent="scGPT embed 50k cells — 1×GPU, ~5 min",
    inputs=[
        {"src": "dataset.h5ad", "dstFilename": "dataset.h5ad"},
        {"src": "embed.py", "dstFilename": "embed.py"},
    ],
    command="python3 embed.py",
    environment=...,   # env name from compute_details
    outputs=["embedded.h5ad"],
    timeoutSeconds=1800,
)
print(job.job_id)   # cell ends here — kernel never blocks on compute

Retain the exact returned job_id. Query that saved ID with the non-blocking host.compute.create(provider).attachJob(job_id).status() or .result() when its state or result is relevant; do not scan Job history. A final .result() read reports whether its follow-up was suppressed or had already been committed; otherwise the app starts the later analysis turn for an unread final result. See the remote-compute-ssh skill for the orchestration details.

See the remote-compute-ssh / remote-compute-modal skill for the orchestration details.

In embed.py, pass model_dir= the checkpoint path from compute_details. If flash-attn is unavailable in that environment, set use_fast_transformer=False.

Gotchas

  • use_fast_transformer default is True but resolves to a FlashAttention path that may not import in every env. Pass use_fast_transformer=False unless you've confirmed flash_attn loads cleanly.
  • The package historically depended on torchtext.vocab.Vocab; in environments without torchtext a pure-Python shim provides Vocab — functionally identical for GeneVocab, but if you hit AttributeError: 'Vocab' object has no attribute …, you're on a stale shim.
  • Gene names must match the vocab; unmatched genes are dropped. Set gene_col to the column in adata.var that holds symbols.

Troubleshooting

SymptomFix
flash_attn is not installed warning at importHarmless; pass use_fast_transformer=False
'Vocab' object has no attribute 'vocab'Env has an old torchtext shim — update the env
Nearly all genes droppedWrong gene_col; check adata.var.columns
"scgpt not in manifest" / env-detection misses scGPTThe baked env manifest lists the distribution as scGPT (and flash_attn), pip's canonical casing — normalize manifest keys before lookup: name.lower().replace('-', '_')

Next: cluster/annotate the embedding with the scanpy library (sc.pp.neighbors → sc.tl.leiden / sc.tl.umap), or compare to an scvi-tools latent space on the same data.

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Predict protein structure for monomers and multimers with AlphaFold2 via the ColabFold runner (Mirdita et al. 2022, github.com/sokrypton/ColabFold; AlphaFold2 Jumper et al. 2021). Reach for this skill to fold a sequence or complex with the AF2/AF2-Multimer evoformer, to validate designed sequences by self-consistency pLDDT, ipTM, and RMSD, or to run a quick MSA-backed prediction using the public MMseqs2 server.

日本語の概要は準備中です。原文の説明を表示しています。

aipoch/open-science5,5132026年10月11日 更新

boltz

無料

Structure prediction for protein, nucleic-acid, and small-molecule complexes with Boltz-2 (Passaro & Wohlwend et al. 2025, github.com/jwohlwend/boltz). Reach for this skill to validate designed binders against a target, to co-fold a protein with a SMILES or CCD ligand, or to get an open-source AlphaFold3 alternative with optional binding-affinity prediction.

日本語の概要は準備中です。原文の説明を表示しています。

aipoch/open-science5,5132026年10月11日 更新

borzoi

無料

Predict genome-wide functional tracks (RNA-seq, CAGE, DNase, ChIP) from DNA sequence with Borzoi. Use this skill when: (1) Scoring the regulatory effect of a variant on expression/accessibility, (2) Generating predicted coverage tracks for a locus, (3) Prioritising non-coding variants by predicted track delta.

日本語の概要は準備中です。原文の説明を表示しています。

aipoch/open-science5,5132026年10月11日 更新

chai1

無料

Structure prediction for protein, nucleic-acid, and small-molecule complexes with the Chai-1 foundation model (Chai Discovery 2024, github.com/chaidiscovery/chai-lab). Reach for this skill to predict an antibody-antigen or protein-ligand complex from a single FASTA, to re-fold designed binders as an AlphaFold-multimer alternative, or to drive co-folding from Python for batched campaigns on a GPU.

日本語の概要は準備中です。原文の説明を表示しています。

aipoch/open-science5,5132026年10月11日 更新

Prepare reproducible setup instructions and validate a user-managed named software environment on an Open-Science SSH Compute Host, including direct SSH and Slurm hosts. Use when a remote job needs packages, modules, cache variables, or a repeatable activation that the host does not already provide.

日本語の概要は準備中です。原文の説明を表示しています。

aipoch/open-science5,5132026年10月11日 更新

customize

無料

Use when the user wants to create or manage a Specialist agent or create, revise, publish, or delete a Skill through the conversational `/Customize` entry. Routes Skill work to the internal skill-creator and handles Specialist work through the JavaScript host.agents SDK.

日本語の概要は準備中です。原文の説明を表示しています。

aipoch/open-science5,5132026年10月11日 更新

aipoch のスキルをすべて見る

このスキルの問題を報告する