本文へ移動
cccskills
無料GitHub で公開

databricks-serverless

Serverless compute for Databricks jobs, Lakeflow pipelines and Declarative Automation Bundles (DABs), with STANDARD or PERFORMANCE_OPTIMIZED and classic only where serverless cannot run the workload. Use when creating, deploying, scheduling or editing a job or pipeline, or deciding its compute. Invoke BEFORE writing a job spec. For migrating existing classic workloads, use databricks-serverless-migration.

インストール方法を見る

含まれるファイル(4)

  • SKILL.md10.6 KB
  • agents/openai.yaml335 B
  • assets/databricks.png15.0 KB
  • assets/databricks.svg582 B

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Deploying Databricks jobs and pipelines on serverless

FIRST: Use the parent databricks-core skill for CLI basics, authentication and profile selection.

This skill decides compute. Where another skill or example (for example databricks-jobs) shows cluster configuration, follow this skill instead.

Serverless is the default compute for every job, pipeline and bundle you create. If the user explicitly asks for a cluster, do what they ask and mention the serverless option in one line. When editing an existing job, keep its compute unless the user asks to change it; tasks you add follow this skill.

1. Serverless or classic

Use classic compute only for:

  • R code, or Scala in notebook cells (%scala): serverless notebooks run Python and SQL only;
  • custom Docker images;
  • OS packages (apt-get) or native libraries with no pip equivalent;
  • custom Spark data source JARs;
  • a stated hard requirement for an instance type or Databricks Runtime version;
  • an explicit user request;
  • existing code over the budget in section 5;
  • a workspace without serverless: if creating or running the job fails because serverless is not available there, put the whole job on classic and say so.

Then give only the blocked task a cluster (one small job_clusters entry with the latest LTS runtime and autoscale: {min_workers: 1, max_workers: 4}), keep every other task serverless, and state the blocker in one line. For a blocked pipeline, give the pipeline a clusters block instead of serverless: true.

Not reasons for classic: Scala or Java in a JAR task (JAR tasks run on serverless), GPUs (serverless GPU), large data, long runtimes, Kafka, cost (serverless STANDARD is on par with or cheaper than on-demand classic in most cases), or habit.

2. Put every task on serverless

  • New jobs: leave out new_cluster, job_clusters, job_cluster_key, existing_cluster_id, instance_pool_id, node_type_id, num_workers, autoscale, spark_version. A task without a cluster runs on serverless. In an existing job, keep the cluster settings it has.
  • Notebook tasks need nothing else. Python script, wheel and JAR tasks take their libraries from a job-level environment. Use the latest environment version the workspace offers, as a quoted string:
    environments:
      - environment_key: default
        spec:
          environment_version: "4"                # example; use the latest
          dependencies: ["requests==2.32.3"]   # pin versions; JARs go in java_dependencies
    tasks:
      - task_key: main
        spark_python_task: {python_file: /Workspace/.../main.py}
        environment_key: default
    
  • SQL tasks: a serverless SQL warehouse. Pipelines: serverless: true, no clusters block.
  • Streaming: .trigger(availableNow=True), the trigger serverless supports (no processingTime or continuous triggers). For always-on processing, run that stream in a continuous job (continuous: {pause_status: UNPAUSED}), which starts the next run as soon as one finishes, or use a continuous pipeline.
  • Tag every job you create for resource tracking, as other Databricks skills do: job-level tags: {"aidevkit_project": "ai-dev-kit"}. Keep tags the user already has.

3. Choose the performance mode

Set the job-level performance_target:

  • STANDARD when nothing is time-critical and cost matters more than speed: scheduled batch and ETL, nightly or weekly runs, or any job where the user states no latency need. Runs start within minutes instead of seconds and cost less. This is the default choice.
  • PERFORMANCE_OPTIMIZED only when the user names an SLA or a deadline that is tight relative to the runtime, a person waits on the result, or the job runs every 30 minutes or more often.

performance_target applies to the whole job, notebook tasks included (only interactive notebooks always run performance-optimized). Say which mode you chose and why in one line, e.g. "STANDARD: nightly batch, no deadline." Serverless GPU tasks ignore performance_target.

Pipelines have no performance_target of their own. To schedule a pipeline, trigger it from a job with a pipeline_task and set the mode on that job; the job's mode applies to the pipeline update. A continuous pipeline runs in STANDARD only when a continuous job runs it.

4. Code you write runs on serverless from the start

Serverless reads Unity Catalog tables, /Volumes/... paths, cloud paths behind a Unity Catalog external location, and the sample data under dbfs:/databricks-datasets/. It cannot read DBFS mounts (dbfs:/mnt/...) or other DBFS root paths; only those count as a DBFS blocker.

DataFrame and SQL APIs only (no RDDs, sc.*, spark.sparkContext); Unity Catalog three-part table names; /Volumes/... paths instead of DBFS mounts (dbfs:/mnt/...); availableNow streaming triggers; no .cache() or .persist(); no spark.conf.set for executor, driver or memory settings; pinned dependencies in the environment, never init scripts.

5. Existing code: a quick check, not a migration

The user is waiting for the deployment. Spend one short pass, then deploy.

  1. Scan only the files the job runs (a text search such as grep or rg; do not read the whole repo) for: sc., sparkContext, .rdd, parallelize, mapPartitions, reduceByKey, groupByKey, dbfs:/, /dbfs/, dbutils.fs.mount, .cache(, .persist(, spark.conf.set, processingTime, continuous=, writeStream, hive_metastore., GLOBAL TEMP, hivevar, init_scripts, apt-get, docker, and other languages: %scala or %r as the whole magic (# MAGIC %r, not %run), .r files and R notebooks, SparkR and sparklyr. R and Scala cells go to classic (section 1).

  2. Minor blockers: fix them and deploy on serverless if every fix is mechanical and local, keeps the logic, needs no new infrastructure, and there are about 5 small edits or fewer in total:

    • delete .cache(), .persist(), .unpersist(), setLogLevel, executor/driver/AQE spark.conf.set lines
    • move pip installs or requirements.txt into the environment, pinned
    • add .trigger(availableNow=True) to a stream that sets no trigger, in a job that runs on a schedule anyway (a stream with a processingTime or continuous trigger is a step 4 blocker)
    • replace a dbfs:/mnt/... path with a /Volumes/... path only when that volume's storage location is the mount's source, so the job reads the same data; a volume with a similar name is not enough. Otherwise the mount is a step 4 blocker.
    • CREATE GLOBAL TEMP VIEW becomes CREATE TEMP VIEW, and its global_temp.<view> reads become <view>, when the same file creates and reads the view (serverless has no global temp views); a global temp view that another task or notebook reads is a step 4 blocker
    • add the catalog to a two-part table name only when you know which Unity Catalog catalog holds the table (for example the job's default catalog); a table in hive_metastore is a step 4 blocker

    List every edit you made in your reply.

  3. RDD or SparkContext code: translate it to DataFrames and deploy on serverless only if all of these hold:

    • it is one file of about 200 lines or fewer;
    • the change is a pure API translation: the same logic, expressed with DataFrame or SQL;
    • every input path, table, output table and parameter stays exactly as it was. Never swap in another dataset, even a sample one;
    • every input is already readable on serverless (section 4: a Unity Catalog table, an existing /Volumes/... path, a cloud path behind an external location, or dbfs:/databricks-datasets/), so nothing new has to be set up.

    Save the original next to it as <name>.classic.py and point the job at the translated file. Tell the user that the translation has not been run yet and suggest one test run that compares the output with the classic version.

  4. Anything bigger: do not rewrite. Put that task on classic now (section 1), keep the other tasks serverless, and end with: "Kept <task> on classic because <blocker>. To move it to serverless, use the databricks-serverless-migration skill." Bigger means: RDD code that fails any condition in step 3 (for example it reads a DBFS mount), mounts with no volume on the same storage, Hive Metastore to Unity Catalog, recompiling a JAR, networking changes, changing what a stream does (an always-on processingTime or continuous stream), or more than about 5 edits besides a step 3 translation.

Do not stop to ask whether to migrate. Deploy, then report what you did.

Common Issues

IssueSolution
Libraries field is not supported for serverless taskMove libraries into environments[].spec.dependencies (JARs: java_dependencies) and set environment_key on the task
Serverless job starts too slowlyperformance_target: PERFORMANCE_OPTIMIZED, only if the user needs fast starts (section 3)
Only remote Spark sessions ... are supported or another RDD errorRDD/SparkContext code: section 5, step 3 or 4
[PATH_NOT_FOUND] or no access on dbfs:/mnt/...DBFS mounts do not work on serverless: use a volume on the same storage (section 5, step 2), otherwise step 4
Serverless is not enabled or not available in the workspacePut the job on classic and tell the user (section 1)
A sibling skill's example uses job_clusters for a new jobLeave the cluster block out; this skill decides compute. Existing jobs keep their compute

Related Skills

  • databricks-core - CLI basics, authentication, profiles
  • databricks-jobs - job structure, task types, triggers, notifications
  • databricks-pipelines - Lakeflow Spark Declarative Pipelines
  • databricks-serverless-migration - migrating existing classic workloads to serverless

Documentation

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Create Agent Bricks: Knowledge Assistants (KA) for document Q&A and Supervisor Agents for multi-agent orchestration (MAS).

日本語の概要は準備中です。原文の説明を表示しています。

databricks/databricks-agent-skills3452026年10月10日 更新

Use Databricks built-in AI Functions (ai_classify, ai_extract, ai_summarize, ai_mask, ai_translate, ai_fix_grammar, ai_gen, ai_analyze_sentiment, ai_similarity, ai_parse_document, ai_prep_search, ai_query, ai_forecast) to add AI capabilities directly to SQL and PySpark pipelines without managing model endpoints. Also covers document parsing and building custom RAG pipelines (parse → prep_search → index → query).

日本語の概要は準備中です。原文の説明を表示しています。

databricks/databricks-agent-skills3452026年10月10日 更新

Databricks AI Runtime, the `databricks air` CLI commands for submitting and managing GPU training workloads on Databricks serverless compute. Use for: writing and submitting `databricks air` workload YAML, passing hyperparameters and secrets, checking run status, listing/cancelling runs, streaming a run's logs and watching its progress, custom Docker image setup, and environment configuration.

日本語の概要は準備中です。原文の説明を表示しています。

databricks/databricks-agent-skills3452026年10月10日 更新

Create Databricks AI/BI dashboards. Must use when creating, updating, or deploying Lakeview dashboards as Databricks Dashboard have a unique json structure. CRITICAL: You MUST test ALL SQL queries via CLI BEFORE deploying. Follow guidelines strictly.

日本語の概要は準備中です。原文の説明を表示しています。

databricks/databricks-agent-skills3452026年10月10日 更新

Design the UX of custom-code Databricks Apps (AppKit/React) data screens — KPI/overview pages, reports, charts, tables, and Genie/chat data assistants — mapped to concrete AppKit components. Use when BUILDING or reviewing the UI of an AppKit/React app that displays data or answers data questions: choosing genre, layout, charts, KPIs, semantic color, required states (loading/empty/error), IBCS notation, and AI-result trust (showing generated SQL/sources for Genie/chat). A plain "create a dashboard" request means a managed AI/BI (Lakeview) dashboard → use databricks-aibi-dashboards, NOT this skill. Also NOT for non-data frontend (forms, settings, auth, marketing) or scaffolding/build/deploy (→ databricks-apps). Complements databricks-apps; use it alongside whenever a custom app has a chart, table, KPI, report, or Genie/chat/AI surface.

日本語の概要は準備中です。原文の説明を表示しています。

databricks/databricks-agent-skills3452026年10月10日 更新

Build apps on Databricks Apps platform. Use when asked to create data apps, analytics tools, or custom interactive visualizations. A plain "create a dashboard" request means a managed AI/BI (Lakeview) dashboard → use databricks-aibi-dashboards, not this skill. Evaluates data access patterns (analytics vs Lakebase synced tables) before scaffolding. Invoke BEFORE starting implementation.

日本語の概要は準備中です。原文の説明を表示しています。

databricks/databricks-agent-skills3452026年10月10日 更新

databricks のスキルをすべて見る

このスキルの問題を報告する