Latency x-ray for undocumented hardware
  • C 82.9%
  • Python 16.9%
  • Makefile 0.2%
Find a file
xoreaxeaxeax 009eb6f772 fix typo
History 2026-08-11 09:49:03 -04:00
examples add examples 2026-08-11 09:20:23 -04:00
.gitignore initial commit 2026-08-06 13:42:25 -04:00
arch_timing.h initial commit 2026-08-06 13:42:25 -04:00
cpu.c initial commit 2026-08-06 13:42:25 -04:00
cpu.h initial commit 2026-08-06 13:42:25 -04:00
cycles2time.c initial commit 2026-08-06 13:42:25 -04:00
cycles2time.h initial commit 2026-08-06 13:42:25 -04:00
LICENSE initial commit 2026-08-06 13:42:25 -04:00
Makefile initial commit 2026-08-06 13:42:25 -04:00
mmio.c initial commit 2026-08-06 13:42:25 -04:00
mmio.h initial commit 2026-08-06 13:42:25 -04:00
mmiotic.c initial commit 2026-08-06 13:42:25 -04:00
pci.c initial commit 2026-08-06 13:42:25 -04:00
pci.h initial commit 2026-08-06 13:42:25 -04:00
plot_mmio.py add examples 2026-08-11 09:20:23 -04:00
progress.c initial commit 2026-08-06 13:42:25 -04:00
progress.h initial commit 2026-08-06 13:42:25 -04:00
README.md fix typo 2026-08-11 09:49:03 -04:00
scan.c initial commit 2026-08-06 13:42:25 -04:00
scan.h initial commit 2026-08-06 13:42:25 -04:00

mmiotic

An exploratory tool for MMIO timing — time any physical address, and dissect the hardware off the latency.

MMIO latency heatmap

Unexpectedly useful for hardware reverse-engineering, hypervisor fingerprinting, side-channelling device activity, register characterization, calibrating malicious triggers, and crafting obscenely long machine instructions.

Examples

Reverse engineering memory ranges

The first megabyte of physical memory is an easy example to start with, even if its structure is already well-known.

first megabyte heatmap

The latency boundaries let us carve up the address space, without relying on reported layouts.

  1. Start with a simple latency scan: sudo ./mmiotic --start-address 0x0 --end-address 0x100000 --stride 4 --count 20 --quiet-mmap
  2. Latency lands on the classic boundaries — DRAM to 0x9FFFF at ~380 cy, then the 128 KB video hole, which splits in two at 0xB0000. The change in variance at 0x20000 may indicate a different page walk for the null page.
  3. A0000–AFFFF: slow and jittery (996 cy, stdev 327), routed to the powered-but-idle Radeon. B0000–BFFFF: fast and rigid (780 cy, stdev 3.2, reads ffffffff), claimed by nobody. Similar-looking data, but a 100x variance gap.
  4. C0000–FFFFF: /proc/iomem says "System ROM", but scan shows DRAM latency, not flash speed — the BIOS is shadowed into DRAM.

Undocumented mailbox registers

Scan a root-complex config space (00:00.0, 4 KB) for latency outliers against a flat ~675-cycle floor:

mailbox register heatmap

The mailbox latency is significantly higher than the surrounding data, and the duplicated spikes and equivalent variances let us match functionally similar registers.

  1. Scan the root-complex config space: sudo ./mmiotic --start-address 0xf8000000 --end-address 0xf8001000 --stride 4 --count 200
  2. Two dwords are the outliers: offset 0xe4 (1218 cy, value 80e3110b) and 0xa4 (1199 cy, value deadbeef). e4 is a documented doorbell; a4 is undocumented.
  3. Near-identical latency, 0x40 apart, both far above the floor — one register family reaching the same off-die target, of which only e4 is public.
  4. The values differ: deadbeef reads as an uninitialized/error sentinel — same mailbox, a4 unimplemented or fed invalid input.

Interrupt calibration

Carefully selected latencies can sometimes be used in unusual exploitation scenarios (MCHAMMER, smiiiiiiiiiiiiiiii):

interrupt calibration heatmap

Here an idle, driverless Radeon turns one aligned 4-byte read into a ~100,000-cycle stall.

  1. sudo ./mmiotic --start-address 0x90e00000 --end-address 0x90e00100 --stride 4 --count 50 — sweep the head of the GPU BAR (0x90e00000, 256 KB). 0x90e00008 costs 110,152 cy; 0x90e00018, 16 bytes on, is fast again at 1,523 — per-register, not per-page.
  2. sudo ./mmiotic --start-address 0x90e00000 --end-address 0x90e40000 --stride 4 --count 1 --binary — --binary probes in bit-reversal order, so the full 256 KB aperture is uniformly covered after 1,012 probes, mapping a contiguous 48 KB block where every register is slow.
  3. Likely, the latency is due to an SMU round-trip to ungate a clock/power-gated domain — so the latency map extracts the GPU's power layout, no driver or datasheet needed.

Killer peek

Ideally, reading a physical address would be a relatively safe operation, but mmiotic is especially good at accidentally finding the exceptions:

killer peek heatmap

For example, on the above platform, a 1-byte read of 0xdc5003b0 — offset 0x3b0 of an undriven Zen 4 iGPU register BAR (Region 5, 0xdc500000, 512 KB) — reliably resets the system.

  1. sudo ./mmiotic --address 0xdc5003b0 --size 1 --count 1. No output; the box reboots.
  2. Exactly one dword: 0x3ac and 0x3b4 on either side read 00000000 at ~3,200 cy and are harmless.
  3. A per-dword sweep of 0x3a0–0x3fc finds 7 poison dwords interleaved with 17 safe ones.
  4. sudo busybox devmem 0xdc5003b0 32 and sudo dd if=/dev/mem bs=1 count=1 skip=$((0xdc5003b0)) take the box down if you'd rather skip the middleman.

Usage

Warning

Some platforms will reset when specific address regions are probed by mmiotic. Be careful where you use this.

Targeting

Option Names
-b/-d/-f/-r PCI --bus/--device/--function/--register — decoded against the ECAM base
-o, --offset <n> offset from the MMIO base (skips B/D/F decode)
-A, --address <n> one physical address; sets the region base implicitly
-a/-z --start-address/--end-address — a physical range; implies --scan

The ECAM base and size are auto-detected from ACPI MCFG (/sys/firmware/acpi/tables/MCFG). Override with -M, --mmio-base / -Z, --mmio-size.

Scanning

Option Effect
-S, --scan walk all addresses in the target region
-t, --stride <n> step between addresses (default: 4)
-n, --limit <n> stop after this many addresses
--binary probe in bit-reversal order — after k probes the range has uniform coverage at ~range/k granularity (see the warning above)
-x, --limited limited register range (0x00–0xff vs 0x000–0xfff)
-e, --skip skip functions that read 0xffffffff at offset 0 (absent devices)
-I, --iomem scan every non–System-RAM top-level region in /proc/iomem
-R, --ioregion <name> scan /proc/iomem regions whose label matches <name>, any depth

--find-target/--find-longest turn scanning into search:

Option Effect
-F, --find-target <s> find an address whose access time reaches s seconds, then escalate width (unaligned 4b → 8/16/32/64/512b) to push it higher
-G, --find-longest same search, but exhaust every candidate instead of bailing on the first hit
-B, --fallbacks <n> candidates carried into the escalation phase (default: 10)

Access method

Option Effect
-s, --size <n> bytes per access: 1 / 2 / 4 (default) / 8 / 16 (XMM) / 32 (YMM) / 64 (ZMM) / 512 (fxrstor). Sizes > 4 are outside spec but work on tested hardware; AVX widths need AVX / AVX-512F
-c, --count <n> samples per address (default: 1); the minimum is reported
-L, --lock time a locked RMW: lock xadd (sizes 1/2/4/8) or lock cmpxchg16b (size 16). Writes the value back
-g, --gather time a vectored gather; --size 16/32/64 picks XMM/YMM/ZMM (AVX-512)
-w, --write-read capture write latency by timing a read that must observe a prior posted write — see below
-E, --enter (--enter-2, --enter-3) drive a burst of up to 30 accesses from one enter $0, $31 — see below
-k, --continue on an fxrstor fault, drop that sample and keep scanning

Output / execution

Option Effect
-C, --min-cycles <n> only print results with min_cycles >= n
-T, --time show nanoseconds alongside cycles (detects TSC frequency)
-N, --nanosecond with --time, print raw nanoseconds, no unit scaling
-P, --progress live scan progress on stderr
-p, --processor <n> pin to logical CPU n (default: 0) — reduces timer noise
-u, --unbound don't pin to a core
--lazy-mmap map in 2 GB windows on demand; required for regions > 2 GB
--quiet-mmap suppress "mmap failed" noise from inaccessible regions

Result lines

b/d/f.o.l: 00:02:00.00.4  address: 00000000f0100000  min:      312  max:      480  stdev:      2.1  dynamic: 0  value: 8086abcd
Field Meaning
b/d/f.o.l bus/device/function, register offset, access size
offset offset from the MMIO base — replaces b/d/f.o.l when the target came from -o/-A/-a+-z
address full 64-bit physical address
min/max fastest and slowest access seen, in cycles
stdev standard deviation of the sampled cycle counts
dynamic 1 if the value changed at any point across samples
value value read — for sizes over 8 bytes, the first 32 bytes as space-separated qwords
faults appended only if an access faulted

-T suffixes the cycle counts with cy and inserts a scaled min/max time pair (-N prints those as raw nanoseconds):

b/d/f.o.l: 00:02:00.00.4  address: 00000000f0100000  min:      312 cy  max:      480 cy  (   104 ns/   160 ns)  stdev:      2.1  dynamic: 0  value: 8086abcd

Setup

Build:

make

Runs as root to get access to /dev/mem.

If you see: mmap ... Operation not permitted, try booting with iomem=relaxed to lift CONFIG_STRICT_DEVMEM:

# /etc/default/grub
GRUB_CMDLINE_LINUX="iomem=relaxed nopat"

Then sudo update-grub and reboot.

In-the-wild

mmiotic has been used to advance a variety of internal and external research, sometimes in unexpected ways. Some examples of mmiotic applications:

  • MCHAMMER: mmiotic calibrates a delayed machine check exception for precision delivery of MC# signals into protected environments.

  • smiiiiiiiiiiiiiiii: mmiotic resolves the high-latency instruction for breaking the SMM rendezvous.

  • The Assembly Hall of Shame: mmiotic is used extensively for the important problem of performance deoptimization.

Author

mmiotic is a research effort from Christopher Domas (@xoreaxeaxeax).