本文へ移動
cccskills
無料GitHub で公開

spark-best-practices

Apache Spark 4.0.2 best practices for PySpark and Scala distributed data processing

インストール方法を見る

含まれるファイル(1)

  • SKILL.md1.8 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Apache Spark Best Practices

Version: Spark 4.0.2. Key changes from Spark 3.x: ANSI mode is now default (stricter SQL type coercion and overflow checks), and Spark Connect provides a decoupled client-server protocol for remote Spark access.

Performance Optimization

Broadcast Joins (CRITICAL)

  • Use broadcast(small_df) for small-large table joins
  • Default broadcast threshold: 10MB (spark.sql.autoBroadcastJoinThreshold)
  • Avoid broadcast for tables > 100MB

Shuffles (CRITICAL)

  • Minimize shuffles: expensive operations
  • Use coalesce() to reduce partitions without shuffle
  • Use repartition() only when necessary (causes shuffle)
  • Predicate pushdown: filter before joins

Caching

  • Cache DataFrames used multiple times: df.cache() or df.persist()
  • Choose storage level: MEMORY_ONLY, MEMORY_AND_DISK, DISK_ONLY
  • Unpersist when done: df.unpersist()

Resource Management

Executor Configuration

  • Executor memory: 80% of available memory per executor
  • Executor cores: 4-5 cores per executor (optimal)
  • Dynamic allocation: enable for varying workloads

Partitioning

  • Optimal partition size: 100-200MB
  • Too few partitions: underutilized cluster
  • Too many partitions: task overhead

Data Processing

UDFs

  • Prefer built-in functions over UDFs
  • Use Pandas UDF for vectorized operations
  • Avoid Python UDFs (serialization overhead)

Storage Formats

  • Parquet: default for analytics (columnar, compression)
  • ORC: alternative to Parquet
  • Delta/Iceberg: ACID transactions, time travel

References

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Pre-action boundary checking — validates agent tool calls against declared capabilities and task contracts

日本語の概要は準備中です。原文の説明を表示しています。

baekenough/second-brain152026年10月8日 更新

Auto-detect project context and optimize harness — deactivate unused agents/skills, suggest missing experts, generate project profile

日本語の概要は準備中です。原文の説明を表示しています。

baekenough/second-brain152026年10月8日 更新

Adversarial code review using attacker mindset — trust boundary, attack surface, business logic, and defense evaluation

日本語の概要は準備中です。原文の説明を表示しています。

baekenough/second-brain152026年10月8日 更新

Apache Airflow best practices for DAG authoring, testing, and production deployment

日本語の概要は準備中です。原文の説明を表示しています。

baekenough/second-brain152026年10月8日 更新

Alembic migration patterns for naming conventions, safety checks, expand-contract, env.py configuration, and CI integration

日本語の概要は準備中です。原文の説明を表示しています。

baekenough/second-brain152026年10月8日 更新

Pre-routing ambiguity analysis — scores request clarity and asks clarifying questions when needed (inspired by ouroboros)

日本語の概要は準備中です。原文の説明を表示しています。

baekenough/second-brain152026年10月8日 更新

baekenough のスキルをすべて見る

このスキルの問題を報告する