本文へ移動
cccskills
無料GitHub で公開

matlab-connect-databricks

Connect MATLAB to Databricks via Spark (Databricks Connect) or JDBC (Database Toolbox). Use when setting up the MATLAB Interface for Databricks, configuring authentication (OauthU2M, OauthM2M, PAT), creating Spark sessions with getDatabricksSession(), reading Unity Catalog tables with server-side filtering, creating JDBC connections with databricks.JDBCConnection or StandaloneJDBCConnection, connecting to SQL Warehouses, selecting JDBC drivers (Simba/OSS), or writing data back to Databricks. Triggers on: Databricks Connect, Spark from MATLAB, getDatabricksSession, .databrickscfg, databricks.JDBCConnection, StandaloneJDBCConnection, SQLWarehouse, Databricks JDBC, Databricks cluster, large table server-side filtering.

インストール方法を見る

含まれるファイル(5)

  • SKILL.md13.6 KB
  • manifest.yaml450 B
  • references/authentication.md4.6 KB
  • references/driver-selection.md3.1 KB
  • references/standalone-jdbc.md3.4 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Connect MATLAB to Databricks

Connect MATLAB to Databricks via Spark (Databricks Connect) or JDBC (Database Toolbox). This skill covers first-time setup, authentication, path selection, session/connection creation, and data operations through both paths.

When to Use

  • First-time setup of the MATLAB Interface for Databricks package
  • Configuring .databrickscfg and authentication (OauthU2M, OauthM2M, PAT)
  • Choosing between Spark and JDBC for a Databricks workflow
  • Reading large tables via Spark with server-side filtering (getDatabricksSession)
  • Creating JDBC connections to clusters or SQL Warehouses (databricks.JDBCConnection, SQLWarehouse.connect())
  • Standalone JDBC connectivity without the full package (StandaloneJDBCConnection)
  • Writing data back to Databricks (via Spark DataFrames or JDBC sqlwrite)
  • Selecting and configuring JDBC drivers (Simba vs. OSS)

When NOT to Use

  • Generic Database Toolbox operations after connection is established — use matlab-use-database
  • ODBC connections (databricks.ODBCConnection)
  • Databricks REST APIs (Clusters, Jobs, DBFS, Unity Catalog admin)
  • MLflow from MATLAB
  • Statement Execution REST API
  • Deploying compiled MATLAB code to Databricks clusters (Job workflow)
  • DuckDB — use matlab-use-duckdb
  • Databricks notebooks, Databricks CLI, or databricks-sdk Python workflows
  • PySpark without MATLAB context (pure Python Spark usage)

Decision Framework: Spark vs. JDBC

ScenarioPathWhy
Large table (millions of rows), need server-side filtering before pulling locallySparkDataFrame operations run on cluster; only filtered results transfer
SQL queries on small-to-medium datasetsJDBCDirect SQL via Database Toolbox; simpler setup
Need DataFrame transformations (withColumn, select, filter chains)SparkNative DataFrame API; operations stay on cluster
Writing large data with performance optimizationJDBCSimba driver's UseNativeQuery optimizes sqlwrite
No MATLAB Interface for Databricks package installedJDBCStandaloneJDBCConnection works with just Database Toolbox + driver jar
Need to use Database Explorer appJDBCsaveSource() + copyToken() integration
Reading files from Unity Catalog Volumes (CSV, Parquet, JSON)Sparkspark.read.format().load() with Volumes paths
Interactive exploration with sqlread/fetchJDBCStandard Database Toolbox workflow on j.Connection

Default: Use Spark when the user mentions large data, server-side filtering, or DataFrames. Use JDBC when the user mentions SQL queries, Database Toolbox, or small/medium datasets. If unclear, ask the user about their data size and preferred workflow.

Workflow

  1. Obtain the package — Ask if the user has the MATLAB Interface for Databricks. If not, direct them to https://www.mathworks.com/solutions/partners/databricks.html (or use StandaloneJDBCConnection for JDBC-only without the package)
  2. Setup — Run setup from the package's Software/MATLAB directory. It configures settings, .databrickscfg, and installs the Databricks Connect library via pip into a venv
  3. Startup — Run startup to add package paths (required once per MATLAB session)
  4. Choose path — Use the Decision Framework above to select Spark or JDBC
  5. Connect — Create a session (getDatabricksSession) or connection (databricks.JDBCConnection)
  6. Verify — Spark: table(spark.range(1)). JDBC: check j.Connection.Message is empty
  7. Operate — Read, filter, write data using the appropriate path's API
  8. Close — JDBC: close(j). Spark: sessions are managed automatically

Key Functions

Spark Path

FunctionPurpose
getDatabricksSession()Creates a Spark session (classic compute)
getDatabricksSession(serverless=true)Creates a serverless session (no cluster, 10-min timeout)
spark.read().table("catalog.schema.table")Reads a Unity Catalog table as a DataFrame
spark.read.format(fmt).load(path)Reads files from Volumes (csv, parquet, json)
DF.filter(expr)Server-side row filtering
DF.select(cols)Server-side column selection
DF.limit(n)Server-side row limiting
DF.withColumn(name, col)Adds/transforms a column (requires Column objects)
table(DF)Converts DataFrame to MATLAB table (pulls data locally)
matlab.sparkutils.table2dataset(T, spark)Converts MATLAB table back to Spark DataFrame
DF.write.mode(m).format(f).saveAsTable(name)Writes DataFrame to Unity Catalog table
matlab.pyspark.sql.functions.col(name)Creates a Column reference
matlab.pyspark.sql.functions.lit(value)Creates a literal Column constant

JDBC Path

FunctionPurpose
databricks.JDBCConnection()Creates a JDBC connection (full package)
StandaloneJDBCConnection()Creates a JDBC connection (no package dependencies)
databricks.SQLWarehouse.connect()Connects to a SQL Warehouse by ID
j.ConnectionThe database.jdbc.connection object for Database Toolbox functions
j.testConnection()Verifies connection is working
j.saveSource()Saves connection as a Database Toolbox data source
close(j)Closes connection and releases resources

Patterns

First-Time Setup

Run setup from the package's Software/MATLAB directory. It is interactive — follow prompts for host URL, auth method, cluster ID, and Databricks Connect library installation.

The Databricks Connect library is downloaded via pip into Software/MATLAB/Connect/<version>/venv/. Requires Python 3.10-3.12. If the download fails repeatedly, ask the user to check with their IT team — do not retry with modified arguments.

cd('/path/to/databricks-package/Software/MATLAB')
setup
startup

After setup, MATLAB's pyenv must point to the venv Python. If getDatabricksSession fails with "databricks.connect package is not installed":

terminate(pyenv);
pyenv(Version="/path/to/databricks-package/Software/MATLAB/Connect/17.3/venv/bin/python");

Warning: If Python is already loaded InProcess, terminate(pyenv) fails. A full MATLAB restart is required. Do not attempt to switch pyenv mid-session after Python has been used.

For authentication configuration details (.databrickscfg format, profiles, environment variables, token caching), see references/authentication.md — consult when configuring auth methods or troubleshooting credential issues.

Spark: Read and Filter a Table

Data stays on the cluster until explicitly collected. Filter server-side first, then collect.

spark = getDatabricksSession(authMethod="PAT");
DF = spark.read().table("catalog.schema.sensor_readings");
filtered = DF.filter("temperature > 100 AND event_date > '2024-01-01'");
T = table(filtered);

Spark: Serverless Session

No cluster needed. Starts instantly with 10-minute inactivity timeout. Requires Python 3.11-3.12 and Databricks Connect >= 15.4.

spark = getDatabricksSession(serverless=true);
DF = spark.read().table("catalog.schema.events");
T = table(DF.limit(50));

Spark: Add Computed Columns

withColumn requires Column objects — not raw scalars. Use col() for references and lit() for constants.

import matlab.pyspark.sql.functions.col
import matlab.pyspark.sql.functions.lit

DF = spark.read().table("catalog.schema.measurements");
DF2 = DF.withColumn("temp_fahrenheit", col("temp_celsius") * lit(9/5) + lit(32));

Spark: Write Back to Databricks

Always specify .mode() — without it, writes fail if the target exists.

DF_new = matlab.sparkutils.table2dataset(T, spark);
DF_new.write.mode("overwrite").format("delta").saveAsTable("catalog.schema.output_table");

Spark: Read Files from Volumes

DF = spark.read.format("csv").option("header", "true").load("/Volumes/catalog/schema/volume/data.csv");
T = table(DF.limit(1000));

JDBC: Cluster Connection

j = databricks.JDBCConnection(catalog="mycatalog", schema="myschema");
data = sqlread(j.Connection, "mytable");
close(j);

JDBC: SQL Warehouse Connection

warehouse.connect() returns a database.jdbc.connection directly — pass it to sqlread/fetch without .Connection.

warehouse = databricks.SQLWarehouse;
warehouse.id = "abc123def456";
conn = warehouse.connect();
data = fetch(conn, "SELECT * FROM mycatalog.myschema.mytable LIMIT 10");
close(conn);

JDBC: Standalone (No Package)

When the user does NOT have the MATLAB Interface for Databricks, use StandaloneJDBCConnection. Requires Database Toolbox and the Simba driver jar only. See references/standalone-jdbc.md for setup and JSON template.

j = StandaloneJDBCConnection(schema="myschema", catalog="mycatalog");
data = fetch(j.Connection, "SELECT * FROM mytable LIMIT 10");
close(j);

JDBC: On-Databricks (Browser MATLAB)

The JDBC driver's OAuth flow cannot open a browser when MATLAB runs on a Databricks cluster. Use package-managed auth instead.

if databricks.internal.isOnDatabricks()
    j = databricks.JDBCConnection(authMethod="OauthU2M", useDriverAuth=false);
else
    j = databricks.JDBCConnection();
end
data = fetch(j.Connection, "SELECT * FROM mycatalog.myschema.mytable LIMIT 10");
close(j);

JDBC: Write-Optimized Connection

Simba driver write performance improves with native query mode (enabled by default).

j = databricks.JDBCConnection(catalog="main", schema="telemetry");
sqlwrite(j.Connection, "measurements", data);
close(j);

Connection Cleanup

Use onCleanup to guarantee closure even when operations fail.

j = databricks.JDBCConnection(catalog="main", schema="analytics");
cleanup = onCleanup(@() close(j));
data = fetch(j.Connection, "SELECT * FROM large_table WHERE id > 1000");

For Spark stale sessions:

clear spark
spark = getDatabricksSession(forceNewSession=true);

Conventions

  • Always filter DataFrames server-side before calling table(DF) — pulling millions of unfiltered rows wastes bandwidth and memory
  • Use three-level names for Unity Catalog tables: "catalog.schema.table"
  • Use getDatabricksSession() for Spark — never construct sessions manually or mimic PySpark builder patterns
  • Use databricks.JDBCConnection or StandaloneJDBCConnection for JDBC — never manually construct JDBC URLs with database()
  • Always pass authMethod explicitly — the default chain may trigger unexpected browser prompts
  • Never hardcode tokens or secrets in MATLAB code — use .databrickscfg or environment variables
  • Always call close(j) when done with JDBC connections
  • Use forceNewSession=true for stale Spark sessions, not clear all
  • For JDBC driver selection details, see references/driver-selection.md — consult when choosing between Simba and OSS drivers or configuring Java

Common Mistakes

MistakeWhy It's WrongCorrect Approach
Manually building JDBC URLs with database()Fragile, error-prone, exposes secrets, misses automatic URL construction, driver classpath management, unified auth chain, and write optimizationUse databricks.JDBCConnection() or StandaloneJDBCConnection()
Inventing PySpark builder: databricks.spark.Session.builder().remote()Does not exist in MATLABUse getDatabricksSession()
Calling table(DF) on unfiltered large tablesTransfers entire dataset locallyApply .filter() / .select() / .limit() first
Passing scalars to withColumn: DF.withColumn("x", 2)Second argument must be a Column objectUse DF.withColumn("x", lit(2))
Writing without .mode(): DF.write.save(path)Fails if target existsUse .mode("overwrite") or .mode("append")
Using driver auth on-clusterBrowser OAuth fails in browser-based MATLABUse useDriverAuth=false
Hardcoding tokens in source codeSecurity risk; tokens expireStore in .databrickscfg or environment variables
Calling getDatabricksSession() without authMethodDefault chain may trigger unwanted browser promptPass authMethod="PAT" or other method explicitly
Mismatched Python version for SparkSession creation failsMatch local Python to runtime (e.g., 3.12 for runtime 17.3)
Using OSS driver with Java 8OSS requires Java 11+Set Java 11+ via jenv first (requires MATLAB restart)
Passing DataFrame args to chained methods: DF.unionAll(DF2)MATLAB Spark wrapper doesn't support DataFrame method argumentsUse spark.sql() with SQL UNION ALL instead

Copyright 2026 The MathWorks, Inc.


レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Guide for accessing financial and economic data in MATLAB using the Datafeed Toolbox. Covers Bloomberg (market data via bloomberg/blp/bloombergHypermedia), FRED (Federal Reserve economic data via fredrs), Haver Analytics (economic data via haver/haverdirect/haverview), and LSEG Datastream (historical data via datastreamws). Use when connecting to any of these data providers from MATLAB.

日本語の概要は準備中です。原文の説明を表示しています。

matlab/matlab-agentic-toolkit1,1492026年10月9日 更新

Read BEFORE writing any code that adds Additive White Gaussian Noise (AWGN) to signals and converts between SNR, Eb/No, Es/No, and per-subcarrier SNR for communications simulations, using awgn(), convertSNR(), berawgn(). The default MATLAB patterns for AWGN (e.g., 'measured' option, manual SNR formulas) produce subtly incorrect results. This skill specifies the correct calling conventions, required function usage, and critical anti-patterns that must be avoided.

日本語の概要は準備中です。原文の説明を表示しています。

matlab/matlab-agentic-toolkit1,1492026年10月9日 更新

Analyze AMS waveform data using Mixed-Signal Blockset utilities: phase noise measurement, clock jitter, anti-aliased resampling, timing measurements, lock time, INL/DNL, ADC/DAC calibration, HSpice import. Use when analyzing time-domain voltage from PLL/VCO/clock simulations, measuring phase noise from variable-step solver output, computing jitter, or resampling non-uniform data.

日本語の概要は準備中です。原文の説明を表示しています。

matlab/matlab-agentic-toolkit1,1492026年10月9日 更新

Design and analyze electrically large antenna structures using MATLAB Antenna Toolbox. Covers reflector antennas (parabolic, Cassegrain, Gregorian, offset, corner, cylindrical, spherical, custom STL), reflectarrays and reconfigurable intelligent surfaces (RIS), antennas installed on platforms (vehicles, aircraft, ships, satellites), and radar cross section (RCS) analysis. Includes solver selection (MoM-PO, PO, MoM, FMM), mesh control, and GPU acceleration. Use when the user wants to design a dish/reflector antenna, reflectarray, analyze an antenna on a platform, or compute RCS.

日本語の概要は準備中です。原文の説明を表示しています。

matlab/matlab-agentic-toolkit1,1492026年10月9日 更新

Analyze data using MATLAB. Use when the task involves tables, timetables, time-series data, numeric arrays, sensor matrices, or gridded data — including but not limited to exploring, row filtering, sorting, cleaning, transforming, aggregating, smoothing, padding, trimming, and answering questions about data. MATLAB provides extensive, easy-to-use built-in functions for these workflows with no additional products required.

日本語の概要は準備中です。原文の説明を表示しています。

matlab/matlab-agentic-toolkit1,1492026年10月9日 更新

S-parameters, insertion loss, fields, currents, mesh control, and solver selection for RF PCB performance validation. TRIGGER: user asks to compute S-parameters, analyze insertion/return loss, extract fields or currents, compare MoM vs FEM, or control mesh for any RF PCB component. Invoke BEFORE writing sparameters() or solver code — API is non-obvious. SKIP: designing or creating components (use the specific matlab-design-pcb-* skill), material/stackup setup only (use matlab-manage-pcb-material), optimization sweeps (use matlab-optimize-pcb-design), PDN/IR-drop analysis (use matlab-analyze-pcb-pdn).

日本語の概要は準備中です。原文の説明を表示しています。

matlab/matlab-agentic-toolkit1,1492026年10月9日 更新

matlab のスキルをすべて見る

このスキルの問題を報告する