本文へ移動
cccskills
無料GitHub で公開

mhc-algorithm

Implement mHC (Manifold-Constrained Hyper-Connections) for stabilizing deep network training. Use when implementing residual connection improvements with doubly stochastic matrices via Sinkhorn-Knopp algorithm. Based on DeepSeek's 2025 paper (arXiv:2512.24880).

インストール方法を見る

含まれるファイル(6)

  • SKILL.md4.0 KB
  • references/core-concepts.md1.7 KB
  • references/gpt-integration.md2.9 KB
  • references/module-implementation.md3.4 KB
  • references/pitfalls.md3.4 KB
  • references/sinkhorn-knopp.md2.3 KB

SKILL.md(原文)

インストールする前に、エージェントに与えられる指示の中身を確認できます。

mHC: Manifold-Constrained Hyper-Connections

Overview

mHC (Manifold-Constrained Hyper-Connections) stabilizes deep network training by constraining residual mixing matrices to be doubly stochastic. It provides:

  • Stable Training: Lower gradient norm variance via doubly stochastic constraints
  • Multiple Streams: Hyper-Connections with learnable mixing across residual streams
  • Sinkhorn Projection: Log-space Sinkhorn-Knopp algorithm for doubly stochastic projection
  • GPT Integration: Pattern for wrapping attention and MLP layers

Two components:

  • HyperConnections Module: Core PyTorch module with H_res, H_pre, H_post matrices
  • Sinkhorn-Knopp: Log-space projection to doubly stochastic manifold

Quick Reference

TopicReference
Core Concepts & MathCore Concepts
Sinkhorn AlgorithmSinkhorn-Knopp
HyperConnections ModuleModule Implementation
GPT IntegrationGPT Integration
Common PitfallsPitfalls

Installation

# Required packages
pip install torch einops numpy

Minimal Example

import torch
import torch.nn as nn
from einops import rearrange, einsum

def sinkhorn_knopp(logits, num_iters=20, tau=0.05):
    log_alpha = logits / tau
    for _ in range(num_iters):
        log_alpha = log_alpha - torch.logsumexp(log_alpha, dim=-1, keepdim=True)
        log_alpha = log_alpha - torch.logsumexp(log_alpha, dim=-2, keepdim=True)
    return torch.exp(log_alpha)

class HyperConnections(nn.Module):
    def __init__(self, num_streams, dim, branch=None, layer_idx=0):
        super().__init__()
        self.num_streams = num_streams
        self.branch = branch

        # Initialize H_res near identity (use small negative for gradient flow)
        init_h_res = torch.full((num_streams, num_streams), -0.1)
        init_h_res.fill_diagonal_(0.0)
        self.H_res_logits = nn.Parameter(init_h_res)

        # H_pre/H_post for depth connections
        init_h_pre = torch.full((1, num_streams), -0.1)
        init_h_pre[0, layer_idx % num_streams] = 0.0
        self.H_pre_logits = nn.Parameter(init_h_pre)
        self.H_post_logits = nn.Parameter(torch.zeros(1, num_streams))

    def forward(self, x):
        s = self.num_streams
        x = rearrange(x, "(b s) t d -> b t s d", s=s)

        h_res = sinkhorn_knopp(self.H_res_logits)
        x_mixed = einsum(h_res, x, "s t, b n s d -> b n t d")

        h_pre = self.H_pre_logits.softmax(dim=-1)
        branch_in = einsum(h_pre, x, "v s, b n s d -> b n v d").squeeze(-2)

        branch_out = self.branch(branch_in) if self.branch else branch_in

        h_post = self.H_post_logits.softmax(dim=-1)
        depth_out = einsum(branch_out, h_post, "b t d, v s -> b t s d")

        output = x_mixed + depth_out
        return rearrange(output, "b t s d -> (b s) t d")

Common Imports

import torch
import torch.nn as nn
import torch.nn.functional as F
from einops import rearrange, einsum, repeat, reduce

When to Use What

ScenarioApproach
Standard residual connectionNo mHC needed
Deep networks (>12 layers) with stability issuesUse mHC with num_streams=4
GPT/Transformer trainingWrap both attention and MLP with HyperConnections
Custom Sinkhorn iterationsAdjust num_iters (20 default) and tau (0.05 default)
Memory-constrained trainingReduce num_streams or batch size

External Resources

レビュー

まだレビューはありません。使ってみた感想をお寄せください。

同じリポジトリのスキル

概要と使いどころ

Perform various data analysis on SEC 13-F and obtain some insights of fund activities such as number of holdings, AUM, and change of holdings between two quarters.

日本語の概要は準備中です。原文の説明を表示しています。

benchflow-ai/skillsbench1,8372026年7月24日 更新

AC branch pi-model power flow equations (P/Q and |S|) with transformer tap ratio and phase shift, matching `acopf-math-model.md` and MATPOWER branch fields. Use when computing branch flows in either direction, aggregating bus injections for nodal balance, checking MVA (rateA) limits, computing branch loading %, or debugging sign/units issues in AC power flow.

日本語の概要は準備中です。原文の説明を表示しています。

benchflow-ai/skillsbench1,8372026年7月24日 更新

Redact text from PDF documents for blind review anonymization

日本語の概要は準備中です。原文の説明を表示しています。

benchflow-ai/skillsbench1,8372026年7月24日 更新

Use when checking simplified ADA-derived plan-view bathroom accessibility constraints such as turning space, door clear width, toilet centerline, grab bars, and lavatory knee/toe clearance.

日本語の概要は準備中です。原文の説明を表示しています。

benchflow-ai/skillsbench1,8372026年7月24日 更新

Analyze failed GitHub Action jobs for a pull request.

日本語の概要は準備中です。原文の説明を表示しています。

benchflow-ai/skillsbench1,8372026年7月24日 更新

Use when extracting plan-view architectural geometry from DXF files with semantic CAD layers, especially when outputs must normalize rooms, doors, fixtures, clearances, and grab bars into machine-checkable JSON.

日本語の概要は準備中です。原文の説明を表示しています。

benchflow-ai/skillsbench1,8372026年7月24日 更新

benchflow-ai のスキルをすべて見る

このスキルの問題を報告する