Route gh-aw workflow design/create/debug/upgrade requests to the right prompts.
日本語の概要は準備中です。原文の説明を表示しています。
Design, implement, optimize, and review SIMD code in .NET. USE FOR: vectorizing scalar loops with TensorPrimitives, Vector64/128/256/512, or platform hardware intrinsics; reviewing existing SIMD code, including the generic Vector type, for contract equivalence, tail handling, memory safety, portability, fallbacks, and measured performance. DO NOT USE FOR: performance work unrelated to SIMD or vectorization.
インストール方法を見るインストールする前に、エージェントに与えられる指示の中身を確認できます。
Produce a portable optimization that preserves the scalar contract, remains memory-safe at every length, and earns its complexity with measured results. Read the official SIMD and hardware-intrinsics guidance first and follow its comprehensive implementation templates. In particular, use its self-contained per-width dispatch, dedicated small-input handling, loop, and remainder shapes rather than reducing them to a chain of width checks. This skill supplies the decision rules and validation checks to apply while changing real code.
Discover these from the repository before asking the user:
| Input | Required | What to establish |
|---|---|---|
| Scalar implementation and tests | Yes | Existing contract, representative call sites, and supported overlap |
| Target frameworks and platforms | Yes | Available SIMD APIs and architectures that must behave consistently |
| Build and test workflow | Yes | The repository's normal commands and how to launch separate test processes |
| Representative workload or benchmark | For optimization | Typical input sizes and the baseline to beat |
Do not add a package merely because an API exists there. First check the target framework and the project's existing dependency/versioning policy.
Span<T> and string
operations, TensorPrimitives, and tensor types already accelerate many operations. LINQ
reductions such as Sum, Min, Max, and Average can also accelerate when the source exposes
its underlying span. Verify empty-input and floating-point behavior rather than assuming similarly
named operations are interchangeable. Once an existing API preserves the contract, use it instead
of continuing into handwritten SIMD. Before writing an explicit loop, name the framework APIs
considered and why none applies. Fixed-shape System.Numerics types remain appropriate for
graphics and similar domains.Vector128<T>. It is accelerated across the broadest
hardware set. Add wider fixed-width paths only when measurements justify them.(vector & mask) == Vector128<byte>.Zero becomes ptest on x86/x64. Use
architecture-specific intrinsics only for a measured gap, guard them with IsSupported, and
retain equivalent portable or scalar behavior.IsHardwareAccelerated, IsSupported, and Count directly. The JIT treats them as
constants, so caching them adds no value and obscures which branches disappear.If the task is review-only, do not rewrite the code. Report correctness and memory-safety defects before performance opportunities.
Vector128<T> and scalar first. Only after
measurements justify wider paths, check Vector512<T>, then Vector256<T>, optional Vector<T>,
Vector128<T>, and finally scalar. Omit paths the implementation does not need. Each outer
fixed-width guard checks only its IsHardwareAccelerated property and, for generic element types,
IsSupported. Inside that block, run the width-specific helper when the input has at least
Count elements; otherwise run a dedicated small-input helper, then return. Do not put the length
check in the outer guard and fall through to repeat dispatch at narrower widths. Keeping each
supported-width block self-contained lets the JIT remove unsupported blocks and avoids redundant
work on common small inputs.Vector128.Create(span) and CopyTo; the JIT keeps them
efficient and they require no pinning or reference arithmetic. Unsafe loads and stores are largely
unnecessary. When a path genuinely must walk a buffer by managed reference, use the element-offset
LoadUnsafe(ref T, nuint) and StoreUnsafe overloads rather than pointers or manually advanced
references.MemoryMarshal.GetReference(span) or MemoryMarshal.GetArrayDataReference(array), not by indexing
element 0.char or bool. Reinterpret with MemoryMarshal.Cast or As<TFrom, TTo>; reinterpretation
changes only the type, not the bits. Keep Boolean data as 0 or 1 and characters as valid
UTF-16, normalizing results before storing when necessary.Count or converting an
index to nuint; otherwise a negative value becomes a huge unsigned offset.0, Count - 1, Count, Count + 1, and
nonmultiples of each width. Once the input contains a full vector, keep the tail vectorized by
reprocessing the last full vector. An idempotent operation can fold that overlap in directly. A
non-idempotent operation must use ConditionalSelect to replace repeated lanes with the
operation's identity before folding them in. This is the JIT-recognized general pattern; it can
reduce a zero-identity selection to a bitwise mask while retaining broader optimization
opportunities. For in-place transforms, preserve the original tail values before overlapping
stores and write only valid results.Native and Estimate operations can intentionally relax precision or IEEE edge-case behavior;
use them only when the contract permits it and measurements justify them.The official guidance contains the complete dispatch, small-input, unrolling, and remainder
templates; use those for the full implementation. The following excerpt illustrates only the inner
safe Vector128<T> loop for an in-place elementwise transform, after its self-contained dispatch
block has established at least one full vector. Transform represents the operation being
implemented:
Span<int> tail = data.Slice(data.Length - Vector128<int>.Count);
Vector128<int> end = Vector128.Create<int>(tail);
Span<int> remaining = data;
while (remaining.Length >= Vector128<int>.Count)
{
Vector128<int> values = Vector128.Create<int>(remaining);
Transform(values).CopyTo(remaining);
remaining = remaining.Slice(Vector128<int>.Count);
}
if (!remaining.IsEmpty)
{
Transform(end).CopyTo(tail);
}
The early end load preserves original values before overlapping stores. For a read-only reduction,
load the same final span after the main loop and use ConditionalSelect to replace already-processed
lanes with the operation's identity. Do not substitute LoadUnsafe/StoreUnsafe or a scalar
epilogue merely to avoid span bounds checks.
DOTNET_EnableAVX2=0 disables AVX2 and DOTNET_EnableHWIntrinsic=0 disables hardware
intrinsics. Use the repository's normal test command and do not change these process-wide
settings inside a unit test. These settings do not change code already compiled as ReadyToRun or
ahead of time, so confirm the target code is JIT-compiled when using them to force a path.Use BenchmarkDotNet to measure representative small and large inputs before keeping the added
complexity. Compare scalar, Vector128<T>, and each wider implemented path in the same run. Small
inputs can be slower because setup dominates, and speedups are rarely the theoretical vector-width
multiple because memory throughput, alignment, and latency still apply. Report throughput or time
with noise context and, when relevant, generated code size or instruction counts. Control allocation
alignment for stable measurements or randomize it to observe the distribution. A wider vector is
not automatically faster.
If the project cannot target the required framework, run the relevant architecture, or execute the fallback configuration, state exactly which path remains unverified. Do not claim success from a default-hardware test alone.
Review in this order:
まだレビューはありません。使ってみた感想をお寄せください。
概要と使いどころ
Route gh-aw workflow design/create/debug/upgrade requests to the right prompts.
日本語の概要は準備中です。原文の説明を表示しています。
Use a repo-root `.editorconfig` to configure free .NET analyzer and style rules. Use when a .NET repo needs rule severity, code-style options, section layout, or analyzer ownership made explicit. USE FOR: the repo needs a root .editorconfig; analyzer severity and style ownership are unclear; the team wants one source of truth for rule configuration. DO NOT USE FOR: choosing analyzers with no config change; formatting-only execution with no config ownership question. INVOKES: inspect the repository context, edit targeted files, and run relevant build, test, lint, or validation commands when changes are made.
日本語の概要は準備中です。原文の説明を表示しています。
Scans .NET code for ~50 performance anti-patterns across async, memory, strings, collections, LINQ, regex, serialization, and I/O with tiered severity classification. Use when analyzing .NET code for optimization opportunities, reviewing hot paths, or auditing allocation-heavy patterns.
日本語の概要は準備中です。原文の説明を表示しています。
Symbolicate the .NET runtime frames in an Android tombstone file. Extracts BuildIds and PC offsets from the native backtrace, downloads debug symbols from the Microsoft symbol server, and runs llvm-symbolizer to produce function names with source file and line numbers. USE FOR triaging a .NET MAUI or Mono Android app crash from a tombstone, resolving native backtrace frames in libmonosgen-2.0.so or libcoreclr.so to .NET runtime source code, or investigating SIGABRT, SIGSEGV, or other native signals originating from the .NET runtime on Android. DO NOT USE FOR pure Java/Kotlin crashes, managed .NET exceptions that are already captured in logcat, or iOS crash logs. INVOKES Symbolicate-Tombstone.ps1 script, llvm-symbolizer, Microsoft symbol server.
日本語の概要は準備中です。原文の説明を表示しています。
Symbolicate .NET runtime frames in Apple platform .ips crash logs (iOS, tvOS, Mac Catalyst, macOS). Extracts UUIDs and addresses from the native backtrace, locates dSYM debug symbols, and runs atos to produce function names with source file and line numbers. Automatically downloads .dwarf symbols from the Microsoft symbol server using Mach-O UUIDs. USE FOR triaging a .NET MAUI or Mono app crash from an .ips file on any Apple platform, resolving native backtrace frames in libcoreclr or libmonosgen-2.0 to .NET runtime source code, retrieving .ips crash logs from a connected iOS device or iPhone, or investigating EXC_CRASH, EXC_BAD_ACCESS, SIGABRT, or SIGSEGV originating from the .NET runtime. DO NOT USE FOR pure Swift/Objective-C crashes with no .NET components, or Android tombstone files. INVOKES Symbolicate-Crash.ps1 script, atos, dwarfdump, idevicecrashreport.
日本語の概要は準備中です。原文の説明を表示しています。
Design or review .NET solution architecture across modular monoliths, clean architecture, vertical slices, microservices, DDD, CQRS, and cloud-native boundaries without over-engineering. USE FOR: .NET architecture choices; layer and domain boundary review; service decomposition; clean architecture, vertical slice, DDD, CQRS, and modular monolith decisions. DO NOT USE FOR: unrelated stacks; generic tasks that do not need this specific guidance. INVOKES: inspect the repository context, edit targeted files, and run relevant build, test, lint, or validation commands when changes are made.
日本語の概要は準備中です。原文の説明を表示しています。