# Low-Level Performance: Layout, SIMD and Buffered I/O — Zig

Source: https://www.geekswithgeeks.com/en/zig/p-performance

> Control memory layout, use vectors and write efficient I/O.

## Fast by being explicit

Zig gives direct control over **data layout**. A normal `struct` lets the compiler reorder fields for efficiency; **`extern struct`** uses the C ABI layout; **`packed struct`** packs fields bit by bit (great for flags, hardware registers and binary formats) and can be given a backing integer such as `packed struct(u8)`. **`align(n)`** sets alignment. **`std.MultiArrayList`** stores a list of structs as separate arrays per field (struct of arrays), improving cache use when loops touch only a few fields. **SIMD** is available portably through **`@Vector(N, T)`**: arithmetic operates on all lanes at once, `@splat` broadcasts a scalar, and `@reduce` combines lanes. **I/O** should be **buffered**: in Zig 0.15 the standard library introduced non-generic `std.Io.Writer` and `std.Io.Reader` interfaces where **you provide the buffer**, and you must **`flush`** before exiting. Measure with benchmarks and profilers (perf, Tracy, Valgrind's Callgrind) in ReleaseFast or ReleaseSafe, and remember that algorithm choice and memory access patterns matter more than micro-optimisations.

## Packed flags, SIMD and buffered stdout

Explicit layout, vector arithmetic and the Zig 0.15 writer interface.

```zig
const std = @import("std");

const OrderFlags = packed struct(u8) {
    paid: bool = false,
    shipped: bool = false,
    gift_wrap: bool = false,
    express: bool = false,
    _reserved: u4 = 0,
};

fn sumOfSquares(values: []const f32) f32 {
    const lanes = 8;
    const V = @Vector(lanes, f32);
    var acc: V = @splat(0);
    var i: usize = 0;
    while (i + lanes <= values.len) : (i += lanes) {
        const chunk: V = values[i..][0..lanes].*;   // load 8 floats as a vector
        acc += chunk * chunk;                        // 8 multiplications at once
    }
    var total = @reduce(.Add, acc);
    while (i < values.len) : (i += 1) total += values[i] * values[i];   // remainder
    return total;
}

pub fn main() !void {
    const flags = OrderFlags{ .paid = true, .express = true };
    const byte: u8 = @bitCast(flags);                // fits exactly in one byte

    const data = [_]f32{ 1, 2, 3, 4, 5, 6, 7, 8, 9, 10 };

    // Buffered stdout in Zig 0.15: you own the buffer and must flush
    var stdout_buffer: [4096]u8 = undefined;
    var stdout_writer = std.fs.File.stdout().writer(&stdout_buffer);
    const stdout = &stdout_writer.interface;

    try stdout.print("flags byte = 0b{b:0>8}\n", .{byte});          // 0b00001001
    try stdout.print("sum of squares = {d}\n", .{sumOfSquares(&data)}); // 385
    try stdout.flush();                              // without this, output may be lost
}
```

## Packing a suitcase

A packed struct is a carefully packed suitcase where every item fits snugly; SIMD is carrying eight bags in one trip instead of eight trips; buffering is collecting all your mail and posting it together instead of one letter at a time.

**Quiz:** Why must you call flush on a buffered writer in Zig 0.15?

- [ ] To free memory
- [x] Because buffered data is only written when the buffer fills or flush is called
- [ ] To close the file
- [ ] It is optional and does nothing

*Answer:* Because buffered data is only written when the buffer fills or flush is called. Unflushed data stays in the buffer and can be lost when the program exits.
