本文へ移動
cccskills
無料GitHub で公開

multi-cluster-api-data-mismatch

Debug "API returns data that doesn't exist in database" when multiple Kubernetes clusters exist (production, staging, POC). Use when: (1) REST API returns records that direct database queries can't find, (2) Data counts match approximately but specific records don't overlap, (3) kubectl context points to a different cluster than the one serving the public domain, (4) ClickHouse/Postgres queries return stale or different data than the API, (5) "completely different datasets" despite same table names and schemas. Root cause: querying the wrong cluster's database.

インストール方法を見る

含まれるファイル(1)

  • SKILL.md5.0 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

Multi-Cluster API Data Mismatch

Problem

When investigating why an API returns data that doesn't match direct database queries, the root cause may be that your kubectl context points to a different cluster (e.g., POC) than the one actually serving the public API (e.g., production). Both clusters have the same schemas and similar data volumes, making the mismatch non-obvious.

Context / Trigger Conditions

  • API returns records that direct database queries (kubectl exec ... clickhouse-client) can't find
  • Data counts are similar (e.g., both have ~20k records) but specific records differ
  • You're debugging a data pipeline that "should work" but labels/counts never update
  • Multiple kubectl contexts exist (production, staging, POC, test)
  • The API response header shows server: nginx but no ingress exists in the current cluster's namespace

Solution

Step 1: Verify which cluster serves the domain

# Check DNS for the public API domain
dig relay.divine.video +short
# → 34.58.27.79

# Check your current cluster's gateway/ingress IP
kubectl get gateway -A  # or kubectl get ingress -A
# → 35.184.10.63  (different IP = different cluster!)

Step 2: Check all available contexts

kubectl config get-contexts
# Look for production vs staging vs POC contexts

Step 3: Switch to the correct cluster

kubectl config use-context <production-context-name>

Step 4: Verify the database connection

# Check what database the API actually connects to
kubectl get secret -n <namespace> <db-credentials> -o jsonpath='{.data.DATABASE_URL}' | base64 -d
# Production might use managed services (ClickHouse Cloud, Cloud SQL)
# POC might use in-cluster pods

Step 5: Also check for ExternalSecrets when updating secrets

If secrets are managed by ExternalSecrets operator:

# Check if ExternalSecrets manages the secret
kubectl get externalsecret -n <namespace>

# If yes, update the SOURCE (e.g., GCP Secret Manager), not the k8s secret
gcloud secrets versions add <secret-name> --project=<project> --data-file=-

# Force sync after updating
kubectl annotate externalsecret <name> -n <namespace> force-sync=$(date +%s) --overwrite

# Verify the k8s secret updated
kubectl get secret <name> -n <namespace> -o jsonpath='{.data.<key>}' | base64 -d

Verification

After switching to the correct cluster:

  1. Re-run the same database query — the "missing" records should now appear
  2. The latest timestamp in the database should match recent API activity
  3. kubectl get pods -n <namespace> should show the same pods the API logs reference

Key Indicators You're On the Wrong Cluster

SymptomExplanation
API video count ~21k, DB count ~20kSimilar but not identical = different datasets
Latest DB record is days oldAPI is receiving writes on a different cluster
No ingress/gateway routes match the public domainTraffic routes elsewhere
Managed DB URL (e.g., ClickHouse Cloud) vs in-cluster podProduction vs POC architecture
API pod logs show only health checks, no real trafficReal traffic goes to another cluster

Example

# You query POC ClickHouse and find 20,371 videos, latest from Feb 26
# But the API at relay.divine.video returns videos from today
# The API returns a video with d_tag that doesn't exist in your ClickHouse

# Fix: switch to production
kubectl config use-context connectgateway_dv-platform-prod_...

# Now query production ClickHouse Cloud
CH_URL=$(kubectl get secret -n funnelcake funnelcake-clickhouse-credentials \
  -o jsonpath='{.data.CLICKHOUSE_URL}' | base64 -d)
# → https://z1hzyismdt.us-central1.gcp.clickhouse.cloud:8443
# (managed service, not in-cluster pod)

# Query returns 21,221 videos with latest from minutes ago — matches the API

Notes

  • Production often uses managed database services (ClickHouse Cloud, Cloud SQL) while POC/staging use in-cluster pods — same schemas, different data
  • The similar-but-not-identical record counts are the most misleading symptom
  • Always check max(created_at) or max(indexed_at) — if it's days old on a live service, you're querying the wrong instance
  • When updating k8s secrets in production, check for ExternalSecrets first — manual kubectl create secret will be immediately overwritten by the operator syncing from GCP Secret Manager / AWS Secrets Manager / Vault

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Fix ArgoCD ExternalSecret deployment failing with "namespace X is not permitted in project Y". Use when: (1) ExternalSecret shows OutOfSync in ArgoCD but won't sync, (2) ArgoCD application status shows "namespace X is not permitted in project 'infrastructure'", (3) ExternalSecret targets a namespace managed by a different ArgoCD project, (4) Using apps-of-apps pattern with separate infrastructure and application projects.

日本語の概要は準備中です。原文の説明を表示しています。

divinevideo/divine-mobile2662026年10月10日 更新

Art direction for any content — reads text, PDF, Word, HTML, PPT, then proposes 2-3 creative directions with photography style, mood, and visual language. After selection, generates AI image prompts and visual briefs section-by-section. Use when the user shares content and needs visual direction, image sourcing, or creative direction for any material.

日本語の概要は準備中です。原文の説明を表示しています。

divinevideo/divine-mobile2662026年10月10日 更新

Fix "Null check operator used on a null value" errors when an object is set to null during an async await. Use when: (1) Object reference is nullified while awaiting, (2) Code accesses object with ! after await returns, (3) Cancel/dispose operations run concurrently with async operations on same object. Solution: capture local reference before await.

日本語の概要は準備中です。原文の説明を表示しています。

divinevideo/divine-mobile2662026年10月10日 更新

Add custom metadata headers (x-amz-meta-*) to AWS v4 signed requests for GCS S3-compatible API. Use when: (1) Adding custom metadata to GCS uploads via S3 API, (2) Getting signature mismatch errors after adding new headers, (3) x-amz-meta-* headers being ignored or causing 403 errors. Custom headers MUST be included in canonical headers and signed headers list.

日本語の概要は準備中です。原文の説明を表示しています。

divinevideo/divine-mobile2662026年10月10日 更新

Fix password/secret authentication failures caused by trailing newlines when creating Google Cloud secrets (or similar) with bash here-strings. Use when: (1) Password authentication fails with correct password, (2) Secret created with `<<< "value"` syntax, (3) Error like "password authentication failed" or "invalid token" despite correct value. Bash here-strings (`<<<`) add a trailing newline that corrupts secrets.

日本語の概要は準備中です。原文の説明を表示しています。

divinevideo/divine-mobile2662026年10月10日 更新

Fix silent video/media processing failures caused by URL extraction code that filters on file extensions (.mp4, .webm, .webp). Use when: (1) Media moderation, transcoding, or analysis silently skips files from Blossom or content-addressed storage servers, (2) URL extraction from Nostr event tags (imeta, r tags) drops URLs without recognized extensions, (3) CDN fallback URLs append .mp4 but the actual server uses extensionless content-addressed paths like /{sha256}. Common in Nostr video events (kind 34236) where different clients use different URL formats.

日本語の概要は準備中です。原文の説明を表示しています。

divinevideo/divine-mobile2662026年10月10日 更新

divinevideo のスキルをすべて見る

このスキルの問題を報告する