本文へ移動
cccskills
無料GitHub で公開

executing-spark

Execute arbitrary Python or PySpark code on Fabric Spark compute without creating a notebook artifact; ephemeral Livy sessions with full Delta table access. Automatically invoke when the user asks to "run PySpark in Fabric", "create a Livy session", "execute Python on Fabric compute", "run Spark without a notebook", "submit code to Fabric", "ephemeral Spark execution", "run ETL in Fabric".

インストール方法を見る

含まれるファイル(3)

  • SKILL.md6.5 KB
  • references/example-script.md3.6 KB
  • references/livy-api.md4.9 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Executing Spark Code in Fabric (No Notebook)

Run arbitrary PySpark or Python code on Fabric Spark compute via the Livy API. No notebook artifact is created or persisted; sessions are ephemeral. Full read/write access to lakehouse Delta tables via Spark SQL.

Prerequisites

  • Azure CLI authenticated (az login)
  • A lakehouse in the target workspace (the Livy session runs against it)
  • Fabric capacity (F or trial)

Critical: Authentication

The Livy API requires a token from az account get-access-token --resource https://api.fabric.microsoft.com. Tokens from fab auth do not work for OneLake storage access inside the Spark session.

import subprocess, json

result = subprocess.run(
    ["az", "account", "get-access-token", "--resource", "https://api.fabric.microsoft.com"],
    capture_output=True, text=True
)
token = json.loads(result.stdout)["accessToken"]

Do not output or log the token. Pass it directly to the API call.

Lifecycle

1. Create session   POST .../sessions              {"kind": "pyspark"}
2. Wait for idle    GET  .../sessions/{id}          poll until state: "idle" (~30-90s)
3. Submit code      POST .../sessions/{id}/statements   {"code": "...", "kind": "pyspark"}
4. Get result       GET  .../sessions/{id}/statements/{n}   poll until state: "available"
5. Delete session   DELETE .../sessions/{id}        ALWAYS do this

Base URL: https://api.fabric.microsoft.com/v1/workspaces/{wsId}/lakehouses/{lhId}/livyapi/versions/2023-12-01

CRITICAL: Always delete sessions when done. Idle sessions consume Fabric capacity units (CUs). A forgotten session burns compute until it times out (default: 20 minutes). In automation, wrap cleanup in a finally block.

Getting IDs

WS_ID=$(fab get "Workspace.Workspace" -q "id" | tr -d '"')
LH_ID=$(fab get "Workspace.Workspace/Lakehouse.Lakehouse" -q "id" | tr -d '"')

Submitting Code

Submit PySpark or pure Python as statements. The spark object is available automatically.

# Statement payload
{"code": "df = spark.sql('SELECT * FROM products LIMIT 10')\ndf.show()", "kind": "pyspark"}

Results are in output.data["text/plain"] when state: "available" and output.status: "ok".

What Works

  • spark.sql("SELECT ...") ; full Spark SQL against lakehouse tables
  • spark.sql("SHOW TABLES") ; metastore access
  • df.write.mode("overwrite").saveAsTable(...) ; write Delta tables
  • Pure Python (pandas, numpy, pyarrow); runs on Spark container
  • In-memory Spark DataFrames and transformations
  • Multiple sequential statements in one session

What Does Not Work

  • deltalake (delta-rs) is not pre-installed; use Spark SQL instead
  • notebookutils has limited functionality (no FUSE mount at /lakehouse/default/)
  • Tokens from fab auth ; must use az CLI token
  • Tokens expire after ~60 minutes; long sessions need token refresh

When to Use This vs Alternatives

ScenarioApproach
Quick read-only explorationDuckDB locally (fastest; see using-duckdb skill)
Write data back to lakehouseLivy session or notebook
Ephemeral transform; no artifactLivy session (this skill)
Complex multi-cell workflowNotebook (nb exec or portal)
Scheduled ETLNotebook via fab job run
Agent-driven compute (Dagster, orchestrators)Livy session

Persisting code as a notebook: poll the definition LRO tightly

This skill is for ephemeral execution with no artifact. When you instead want to persist or change a notebook (deploy new code, iterate on an existing one), that is an item-definition change, and the poll interval is the single biggest performance lever. fab import, nb create, and nb cell edit take 25-60s because they poll the create/update long-running operation at the server's advertised Retry-After: 20; the work itself finishes in ~1s, and neither CLI lets you change that interval. Poll the LRO at ~0.3s and the same deploy takes ~1-2s. The fabric-cli skill ships scripts/deploy_notebook.py which does this (auto-detects create vs update, --poll-interval default 0.3s); strongly prefer it over fab import / nb for any notebook definition change.

Sessions vs Batch Jobs

A Livy session (this skill) is interactive: create it, submit statements, read output as it runs, delete it. It stays alive and you pay for idle time until you delete it or it times out (~20 min).

A Livy batch is one-shot: submit a single job (a file or inline job spec), poll it to a terminal state, done. No idle-CU footgun, nothing to remember to delete. For scheduled or fire-and-forget agent ETL, prefer a batch over a session; keep sessions for interactive, multi-statement work. Same base URL, /batches instead of /sessions -- see references/livy-api.md.

Livy vs Notebook Jobs: reading the outcome

A Livy statement returns its result directly in the response (output.status = ok/error), so you always know whether it worked. A notebook run via fab job run does not -- its job status reports Completed even when the notebook caught an exception and exited a failure payload. If you run notebooks as batch jobs instead of Livy, you must read the notebook's exit value to get its real verdict. The fabric-cli skill (in the fabric-cli plugin) documents that endpoint and ships scripts/run_notebook_checked.py for it.

References

  • references/livy-api.md -- Full API reference with endpoints (sessions + batches), request/response formats, and error handling
  • references/example-script.md -- Complete working script that creates a session, queries data, writes results, and cleans up

Related

  • using-duckdb skill (same etl plugin) -- read-only Delta querying, local or in-notebook, when you don't need Spark compute
  • fabric-cli skill (fabric-cli plugin) -- nb exec / fab job run for notebooks, reading a notebook's exit value, the SQL-endpoint metadata sync after a Spark write, and scripts/deploy_notebook.py for fast notebook definition changes (tight LRO polling)

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Automatically invoke this skill whenever the user asks about Fabric tenant settings or Power BI tenant settings or auditing tenant settings. You can use this skill if the user mentions "Fabric administration".

日本語の概要は準備中です。原文の説明を表示しています。

data-goblin/power-bi-agentic-development1,0352026年10月11日 更新

bpa-rules

無料

Interactive BPA rule generation for Power BI semantic models; guided discovery, model investigation, and expert rule authoring. Automatically invoke when the user mentions "BPA rule", "Best Practice Analyzer", or asks to "create a BPA rule", "audit BPA rules", "recommend BPA rules", "set up BPA for my team", "check model for best practices", "validate BPA rules", "improve a BPA expression".

日本語の概要は準備中です。原文の説明を表示しています。

data-goblin/power-bi-agentic-development1,0352026年10月11日 更新

Writing and executing C# scripts and macros against Power BI semantic models using Tabular Editor 2/3. Automatically invoke when the user mentions "C# script", "Tabular Editor script", "TOM scripting", "MacroActions.json", "XMLA", or asks to "automate model changes", "bulk update measures", "create calculation groups", "write a macro", "format DAX expressions", "manage model metadata".

日本語の概要は準備中です。原文の説明を表示しています。

data-goblin/power-bi-agentic-development1,0352026年10月11日 更新

TOM and ADOMD.NET guidance via PowerShell for connecting to Power BI Desktop's local Analysis Services instance. Covers model enumeration, DAX queries, metadata modification, annotations, calendar definitions, field parameters, query tracing, DAX library package management (daxlib.org), and the Desktop Bridge for reloading and screenshotting the report canvas. Automatically invoke when the user mentions "Power BI Desktop", "Analysis Services port", "TOM", "ADOMD", "daxlib", "DAX library", "DAX UDF package", or asks to "connect to PBI Desktop", "query PBI Desktop with DAX", "modify PBI Desktop model", "add a measure to PBI", "capture visual queries", "create a field parameter", "validate DAX", "intercept DAX queries", "install daxlib", "add DAX SVG", "add IBCS", "reload the report canvas", "screenshot a report page", "Desktop Bridge", or to work with the model and report in Power BI Desktop together.

日本語の概要は準備中です。原文の説明を表示しています。

data-goblin/power-bi-agentic-development1,0352026年10月11日 更新

Step-by-step workflow for creating complete Power BI reports from scratch using pbir CLI. Covers model discovery, report creation, page layout, theme setup, visual placement, field binding, filtering, formatting, validation, and publishing. Automatically invoke when the user asks to "create a new report", "build a report from scratch", "make a dashboard", "set up a report with KPIs", "create an executive dashboard", "add pages and visuals to a new report".

日本語の概要は準備中です。原文の説明を表示しています。

data-goblin/power-bi-agentic-development1,0352026年10月11日 更新

dax

無料

DAX performance optimization for semantic models. Automatically invoke when the user asks to "optimize DAX", "fix slow DAX", "DAX performance", "tune a measure", "debug a measure", "DAX anti-patterns", or mentions slow queries, server timings, or DAX authoring.

日本語の概要は準備中です。原文の説明を表示しています。

data-goblin/power-bi-agentic-development1,0352026年10月11日 更新

data-goblin のスキルをすべて見る

このスキルの問題を報告する