Table of Contents

Hardware acceleration and the SIMD opt-out

Several primitives in Bodu.Security.Cryptography ship a vectorised kernel alongside their scalar reference implementation, and dispatch to it automatically when the host CPU supports the required instructions. This page documents exactly which algorithms are accelerated, when the fast path engages, and how to force the scalar path process-wide.

Which algorithms are accelerated

Every accelerated primitive has a scalar reference implementation and one or more vector kernels in sibling files - *.Avx512.cs for the AVX-512 kernels, Blake2bCore.Vector256.cs with its Avx512 and Avx2 shims and Blake2bCore.Vector128.cs over Argon2's 128-bit shims for BLAKE2b, Blake2sCore.Vector128.cs with its Avx512, Ssse3 and AdvSimd shims for BLAKE2s, Blake3Core.Vector512.cs, Blake3Core.Vector256.cs with its Avx512 and Avx2 shims and Blake3Core.Vector128.cs over the BLAKE2s shims for BLAKE3, Argon2Core.Avx2.cs, Argon2Core.Vector128.cs with its Ssse3 and AdvSimd shims for Argon2, ScryptCore.Vector128.cs with its Sse2 and AdvSimd shims for scrypt, ChaCha20Core.Vector512.cs, ChaCha20Core.Vector256.cs with its Avx512 and Avx2 shims and ChaCha20Core.Vector128.cs with its Avx512, Ssse3 and AdvSimd shims for ChaCha20, the matching Salsa20Core.Vector*.cs kernels over the same shims for Salsa20, SerpentCore.Vector256.cs and SerpentCore.Vector128.cs over ChaCha20's shims for Serpent-128, CubeHashCore.Vector512.cs, CubeHashCore.Vector256.cs and CubeHashCore.Vector128.cs over ChaCha20's shims for CubeHash, Poly1305Core.Vector512.cs, Poly1305Core.Vector256.cs and Poly1305Core.Vector128.cs for Poly1305, and Ghash.Clmul.cs with its PclmulqdqIsa and PmullIsa shims for GHASH and POLYVAL. The dispatch is per-operation and transparent - the public API and output are identical either way.

Primitive Accelerated operation Instruction set gate
Blake2b Block compression AVX-512F + VL, else AVX2, else SSSE3 (x64); none on ARM64, where the scalar kernel ran faster
Blake2s Block compression AVX-512F + VL, else SSSE3 (x64); none on ARM64, where the scalar kernel ran faster
Blake3 Chunks and parents, many at once; single blocks AVX-512F where the runtime prefers 512-bit vectors (16 chunks at once), else AVX-512F + VL or AVX2 (8), else SSSE3 (x64); AdvSimd (ARM64) (4), with single blocks on the scalar kernel
Threefish256 Encrypt / decrypt block AVX-512F + VL
Threefish512 Encrypt / decrypt block AVX-512F
Threefish1024 Encrypt / decrypt block AVX-512F
CubeHash Round permutation AVX-512F, else AVX2, else SSSE3 (x64); AdvSimd (ARM64)
Argon2d / Argon2i / Argon2id Compression function G AVX2, else SSSE3 (x64); AdvSimd with the general registers (ARM64), two rows or columns at once, one in each
Scrypt scryptBlockMix (Salsa20/8) SSE2 (x64); none on ARM64, where the scalar kernel ran faster
ChaCha20 / XChaCha20, XChaCha20Poly1305 Keystream, many blocks at once AVX-512F where the runtime prefers 512-bit vectors (16 blocks at once), else AVX-512F + VL or AVX2 (8), else SSSE3 (x64); AdvSimd (ARM64) (4)
Salsa20 / XSalsa20, XSalsa20Poly1305 / XSalsa20Poly1305Aead Keystream, many blocks at once As ChaCha20
Serpent128Cipher (and Serpent128) EncryptBlocks / DecryptBlocks: many blocks at once; CTR, EAX and SIV: the counter blocks formed and their keystream applied in the kernel AVX-512F + VL or AVX2 (8 blocks at once), else SSSE3 (x64); AdvSimd (ARM64) (4)
Poly1305, and the Poly1305 in XChaCha20Poly1305 / XSalsa20Poly1305 / XSalsa20Poly1305Aead Runs of whole blocks, many at once, from 512 bytes on x64, 256 bytes on ARM64, and 128 bytes on Apple silicon AVX-512F where the runtime prefers 512-bit vectors (8 blocks at once, from 4 KiB), else AVX2 (4; two groups of four per step from 1 KiB where AVX-512VL is available) (x64); AdvSimd (ARM64) (2; two groups of two per step)
GcmModeTransform / GcmSivModeTransform GHASH / POLYVAL PCLMULQDQ + SSSE3 (x64); PMULL (ARM64)

There are two gate forms. The 128- and 256-bit-lane AVX-512 kernels (BLAKE2b, BLAKE2s, BLAKE3, ChaCha20, Salsa20, Serpent-128, Threefish-256) require the AVX-512 Vector Length extension (Avx512F.VL.IsSupported); the 512-bit-lane kernels (Threefish-512/1024, CubeHash) require only AVX-512 Foundation (Avx512F.IsSupported). The 512-bit kernels of BLAKE3, ChaCha20 and Salsa20 also wait for Vector512.IsHardwareAccelerated, which the runtime clears on processors whose clock drops under sustained 512-bit work; there their 256-bit AVX-512VL kernels run instead, and DOTNET_PreferredVectorBitWidth=512 opts such a processor in. GHASH, which GcmModeTransform authenticates with, and POLYVAL, which GcmSivModeTransform does, share one carry-less kernel that folds four blocks into each reduction. It runs on PCLMULQDQ with SSSE3 for the byte shuffles on x64 (Pclmulqdq.IsSupported && Ssse3.IsSupported) and on the cryptography extension's PMULL on ARM64 (Aes.IsSupported, which .NET uses to expose it); like Argon2's kernel it is written once and differs between the two architectures only in a small shim. Elsewhere, including ARM64 processors without the extension, a constant-time scalar multiply built from integer multiplications runs instead. Argon2 takes the wider of two kernels an x64 processor offers: a 256-bit AVX2 kernel, else a 128-bit kernel over SSSE3. The 128-bit kernel is written once over portable Vector128 arithmetic, and a shim of five operations over AdvSimd runs it on ARM64 too, but alone it ran slower than the scalar kernel on a Neoverse N2, whose two vector pipes it keeps busy while four integer pipes wait. A block's eight rows are independent of each other, as are its eight columns, so on ARM64 a hybrid kernel takes them in pairs: one of each pair in eight vector registers through the 128-bit kernel's half-steps, the other in sixteen general registers, the two interleaved so that both kinds of pipe work at once. It ran 1.24 to 1.28 times as fast as the scalar kernel on the N2 and 1.3 to 1.6 times as fast as the AdvSimd kernel on an Apple M1, under .NET 8 and .NET 10 alike, and took 37 to 42 percent of 1.0.0's CPU on the two. ARM64's thirty-one general registers hold the sixteen words that x64's sixteen registers cannot, and it rotates a 64-bit word in one instruction. Argon2 uses no AVX-512 kernel: its cost is dominated by memory, which wider registers do not help. BLAKE2b and BLAKE2s each write their kernel once over portable vectors and specialize it with a shim of rotations: BLAKE2b's 256-bit kernel runs with AVX-512VL's single-instruction rotations or, on AVX2, with byte shuffles and shifts, and a 128-bit kernel over Argon2's SSSE3 shim covers the rest of x64; BLAKE2s's whole state fits one 128-bit kernel, over AVX-512VL or SSSE3. Both 128-bit kernels have AdvSimd shims too, but ARM64 runs BLAKE2's scalar kernels, which keep the working vector in registers with the message schedule resolved at compile time; they ran BLAKE2b 2.4 to 3.3 times, and BLAKE2s 1.6 to 2.0 times, as fast as the AdvSimd kernels on the ARM64 processors measured. Processors with none of these instruction sets run them as well. BLAKE3's tree makes every chunk independent until its chaining value joins the tree, so its kernels compress many chunks at once, one to each lane of a vector: sixteen over 512 bits, eight over 256 bits with AVX-512VL's rotations or AVX2's byte shuffles, and four over 128 bits with BLAKE2s's SSSE3 and AdvSimd shims - BLAKE3's G is BLAKE2s's, rotations included - with transposes carrying message words in and chaining values out. Parent nodes are compressed many at once the same way, and single blocks, such as a message's last chunk, take a BLAKE2s-style 128-bit kernel on x64 and the scalar kernel elsewhere, ARM64 included, where it compressed a block about 1.5 times as fast as the 128-bit AdvSimd kernel. The threads Blake3.MaxDegreeOfParallelism allows are separate from this dispatch: each thread runs the same kernels on its own part of the input. scrypt's scryptBlockMix runs the Salsa20/8 core on four 128-bit vectors, one per diagonal of its state, over SSE2 on x64 - part of the x64 baseline, so every x64 processor takes it. An AdvSimd shim, which differs only in the three lane rotations it supplies, runs the same kernel on ARM64, but dispatch selects the scalar kernel there, which ran 1.2 to 2.1 times as fast on a Neoverse N2, and on an Apple M1 about 1.6 times as fast under .NET 10 and about as fast under .NET 8. A BlockMix chain is sequential, so the kernel stays at 128 bits: one Salsa20/8 state fills a vector exactly, and nothing wider would have independent work to fill it. The ChaCha20 and Salsa20 keystreams, by contrast, are counter mode - every block independent - so their kernels produce many blocks at once, one to each lane: each vector holds one state word of sixteen, eight or four consecutive blocks, with the block counters in successive lanes, and a transpose carries the words back into block order as the keystream meets the data. Rotations are single instructions under AVX-512, byte shuffles for 16 and 8 bits elsewhere, and shift pairs otherwise; ChaCha20 and Salsa20 share the shims and the transposes. StreamCipherTransform hands every whole block to these kernels in one call; a partial last block, and runs of fewer than four blocks, take the scalar block function. The Poly1305 AEADs draw a message's keystream the same way unless a kernel is estimated to draw it faster: a short message can then take all of its keystream, the block that keys Poly1305 included, in one run of up to sixteen blocks through a buffer, rounded up to a whole group where that costs no more, and a longer message its last few blocks in one step of four. The estimates were measured per kernel: with AVX-512VL, a step of four or eight blocks takes about as long as the scalar function over one block, and without it about twice as long, because the kernels then spill from 16 registers. Serpent-128's EncryptBlocks and DecryptBlocks - which ECB, XTS, OCB and the other batched modes call - work the same way: each vector holds one word of eight or four blocks, the S-box circuits and the linear transform run on all of them at once, and transposes move the blocks in and out of word order. In CTR, and in EAX's and SIV's counter mode, the same kernels form the counter blocks themselves: each lane's big-endian counter is held as four words, one vector per word, advanced by masks rather than branches and byte-reversed into the words Serpent reads, and the keystream is XORed into the data as it is stored, so no run of counter blocks is laid out in memory and no second pass applies the keystream. The rotations are AVX-512's single instruction or shift pairs, and the four-block kernels are the fallback for runs of four to seven blocks. A single Encrypt, as CBC encryption makes, and the wide-block Serpent variants run the scalar circuits. Poly1305's scalar loop runs on three 64-bit limbs, taking the high half of each 64-by-64-bit product from mulx on x64 with BMI2 or umulh on ARM64, and from three ordinary 64-bit multiplies on processors with neither, which the limbs' bounds make exact. That loop takes one block at a time, because each block's multiplication by r waits on the one before. Its vector kernels break that chain the way OpenSSL's and BoringSSL's do: each 64-bit lane takes every fourth (or eighth) block and multiplies by r⁴ (or r⁸), and after the last group each lane multiplies by the power of r its last block needs, so the lanes sum to exactly the scalar loop's result. In the lanes the numbers are five 26-bit limbs, since the only vector multiply is 32 by 32 bits (vpmuludq); .NET exposes no AVX-512 IFMA. Each run computes its powers of r afresh, and its 256-bit multiplies can lower the processor's clock for the rest of an AEAD message, so the kernels start at 512 bytes: below that, the scalar loop finishes an AEAD message first. Where AVX-512VL gives the AVX2 kernel 32 registers, it takes two groups per step from 1 KiB, (h + m)·r⁸ + m′·r⁴, so the second group's products fill the multiplier while the first group's carries run; on AVX2 alone that loop spills from 16 registers, so it is not used there. On ARM64 a 128-bit kernel over AdvSimd gives each 64-bit lane every second block, from 256 bytes, and from 128 bytes on Apple silicon. UMULL and UMLAL multiply 32 by 32 bits, taking one factor from a lane of another register, so a power of r and five times its upper limbs fill three registers. The kernel takes two groups per step, (h + m)·r⁴ + m′·r²: it splits a step's four blocks into limbs together, one to each 32-bit lane, multiplies the second group in the registers' upper halves with UMULL2 and UMLAL2, and carries with USRA, which shifts a limb's carry and adds it to the next in one instruction. On a Neoverse N2 it ran Poly1305 over 1 MiB 1.6 to 1.8 times as fast as the scalar loop, which made XChaCha20-Poly1305 15 to 19 percent faster there, and on an Apple M1 about 3.8 to 4.0 times as fast. Below 256 bytes the N2's scalar loop, whose umulh multiplies 64 by 64 bits, finished first, but the M1's kernel caught it at 128 bytes, so on Apple silicon - macOS, iOS, tvOS and Mac Catalyst on ARM64 - the kernel starts there. The Curve25519 field arithmetic that X25519 and Ed25519 share works the same way on five 51-bit limbs, and forms the split from four ordinary 64-bit multiplies on processors with neither instruction. CubeHash's round is one long dependency chain over its 1024-bit state, so its kernels hold the whole state in registers - two 512-bit, four 256-bit or eight 128-bit - for a run of rounds rather than working on several inputs at once. The state's exchanges of words 8 or 4 apart move whole registers or register halves, which the narrower kernels write into which register each result lands in; its exchanges of words 2 or 1 apart are one in-lane shuffle each. The AVX2 kernel comes close to the AVX-512 one, since the chain, not the register width, bounds the round.

The gates are evaluated where the code is compiled. Under the JIT that is the running processor; a NativeAOT application is compiled ahead of time for a baseline instruction set, which on x64 excludes AVX2 and AVX-512, so it takes the narrower kernels unless the project raises the target with <IlcInstructionSet> (for example x86-64-v3 for AVX2) - only where every machine it will run on supports it.

When the fast path engages

By default the vectorised path is taken whenever the CPU reports the required instruction set. The check is a JIT intrinsic: on a host without AVX-512 it folds to a compile-time constant and the vectorised branch is eliminated entirely, so there is no runtime probing cost.

Forcing the scalar path

For scenarios that need a single, deterministic code path - reproducing a result bit-for-bit across heterogeneous hardware, differential testing against the scalar reference, or an audit that wants one implementation to reason about - the vectorised paths can be disabled process-wide with the feature switch:

Bodu.Security.Cryptography.DisableSimd = true

When set, every accelerated primitive runs its scalar reference implementation regardless of the host CPU. The switch is read once, the first time any accelerated primitive is used, so it must be set before that point. Any of the standard mechanisms works:

  • runtimeconfig.json / project file (recommended - applied before any managed code runs):

    <ItemGroup>
      <RuntimeHostConfigurationOption Include="Bodu.Security.Cryptography.DisableSimd" Value="true" Trim="false" />
    </ItemGroup>
    
  • In code, at startup, before touching any hashing or cipher type:

AppContext.SetSwitch("Bodu.Security.Cryptography.DisableSimd", true);

As a coarser alternative, the .NET runtime's own knobs make the intrinsics report false, which also forces the scalar paths - DOTNET_EnableAVX512F=0 on .NET 8 or DOTNET_EnableAVX512=0 on .NET 10 for the AVX-512 kernels (each runtime ignores the other's name, so set both where either may run), DOTNET_EnableAVX2=0 to move Argon2, BLAKE2b, BLAKE3, ChaCha20, Salsa20 and Serpent-128 from AVX2 to SSSE3, or DOTNET_EnableHWIntrinsic=0 for everything - but they affect the whole process including the BCL, not just this library. Prefer the library switch when you only want to pin these primitives.

Note

The opt-out exists for determinism, reproducibility, and audit, not because the vectorised paths are unsafe. BLAKE2, BLAKE3, ChaCha20, Salsa20, and Threefish are ARX constructions - addition, rotation, and XOR only, with no secret-dependent branches or table lookups - so the scalar and vector paths are both inherently constant-time and produce bit-identical output; the byte shuffles that stand in for rotations index by constants, not data. Serpent's S-boxes are Boolean circuits and its linear transform rotations, shifts and XOR, so the same holds for it. Argon2's kernels add a 32-bit multiply, which has a fixed latency on every targeted processor, and byte shuffles whose indices are constants, not data; which blocks a derivation reads is decided outside the kernels, from public inputs in Argon2i and the first half of Argon2id's first pass. scrypt's kernels are Salsa20/8, again addition, rotation, and XOR, with lane rotations by constants; scrypt's own data-dependent read of V[j] is part of its design, made outside the kernels, and identical on both paths. The GHASH and POLYVAL kernels are carry-less or integer multiplications, shifts, and XORs, which likewise run in fixed time, with no branch on or index by the key or the data. The switch only selects which of two equivalent implementations runs; it makes no additional constant-time guarantee beyond what the algorithms already provide, and this library is not independently audited.

Where to go next