Abstract
Your CPU can add eight numbers in a single instruction, but a normal Go loop adds them one at a time. For years, the only way to use that hardware from Go was hand-written assembly. Go 1.26 added the experimental simd/archsimd package, and Go 1.27 added a portable simd package on top of it. Now you can write SIMD code in plain Go.
But to understand these packages, with functions like LoadFloat32s, or BroadcastInt8s, you need to know what's actually happening in the CPU. So we'll start there. We'll look at the vector registers and how the same bits can be read as different types and sizes. Then we'll follow the basic cycle: load data, work on all the lanes at once, and get the results back out.
With that in place, the Go packages read almost like an API to magage that machinery of the hardware. We'll see a couple of small examples and jump into a real scenario, a hot path in VictoriaLogs that was already heavily optimized, and how rewriting it with the simd package still made it noticeably faster.
You'll leave knowing what SIMD is and how to tell when your code is a good fit for it.