← All writing
Strix Halo

I Let an AI Run ML Experiments on My Laptop While I Slept

Originally published on Medium. This archived article reflects the projects, opinions, and versions at the time of publication. View the Medium original ↗

In this article
The PortWhat It Actually DoesWhy This MattersThe Fork
AMD 395+ is not a toy
AMD 395+ is not a toy

Last week, Andrej Karpathy open-sourced autoresearch — a framework that lets an AI agent autonomously run machine learning experiments. Modify the model. Train for five minutes. Check if it got smarter. Repeat. No human needed. You go to sleep, you wake up to results.

It requires “a single NVIDIA GPU.” Tested on an H100.

I don’t have an H100. I have an AMD Strix Halo laptop. A consumer APU with integrated graphics sharing 64 GB of memory with the CPU. The GPU that AMD’s own ROCm team barely acknowledges exists.

I ported it anyway. It works. And it works better than the datacenter AMD version.

The Port

The original autoresearch uses Flash Attention 3, which is NVIDIA-only. It uses torch.compile, which has dtype bugs on older ROCm PyTorch. It assumes you have 80 GB of dedicated HBM sitting on a PCI-e bus. None of that applies to a Strix Halo.

Here’s what I changed:

Flash Attention 3 got replaced with PyTorch’s built-in scaled dot product attention. On ROCm, this dispatches to AOTriton — AMD’s answer to NVIDIA’s cuDNN kernels. It works. It’s fast. AMD doesn’t mention it in any consumer documentation. I found out it existed by accident three months ago while running my own ML benchmark suite.

The dtype bugs in torch.compile got fixed by explicit casting on the lerp_() calls in the optimizer. The existing AMD fork — which targets MI308X datacenter GPUs costing tens of thousands of dollars — gave up on torch.compile entirely and runs in eager mode. My laptop doesn’t give up that easily.

The result: 24.4% MFU on a consumer APU. The datacenter fork gets 3.18%.

Read that again. A laptop GPU running eight times the computational efficiency of a server GPU’s port. The difference is one flag: torch.compile. The datacenter fork turned it off. I kept it on.

What It Actually Does

The number people fixate on is val_bpb — validation bits per byte. Lower means the model is better at predicting text. Karpathy’s H100 hits about 0.998. My Halo hits 1.602. The datacenter AMD fork hits 1.521.

That gap isn’t about intelligence. It’s about throughput. The H100 gets 953 training steps in five minutes. I get 127. The model itself is identical — same architecture, same optimizer, same data. I’m just doing fewer reps in the same time.

So I doubled the time budget to ten minutes. About 250 steps. Enough for the AI agent to tell whether an architecture change actually helped or was just noise.

And here’s the thing nobody talks about: I’m using 6.2 GB out of 64 GB available. The H100 baseline uses 44 GB out of 80 GB. I have ten times my current footprint in headroom. The first thing the agent is going to try tonight is scaling up — bigger model, bigger batches, fill the pool.

Why This Matters

There are three Strix Halo users on record filing bugs with AMD’s ROCm team. Three. Between us, we’ve run over 250 ML experiments characterizing how this hardware actually behaves. We’ve found critical bf16 precision bugs, undocumented 19x performance flags, and compiler issues that AMD’s own test suite doesn’t catch.

We’re not doing this because AMD asked. We’re doing it because this hardware is absurdly capable and nobody’s using it for what it can do.

The Strix Halo has 64 GB of unified memory. Not 8 GB of VRAM like a gaming card. Not 24 GB like an RTX 4090. Sixty-four gigabytes shared between CPU and GPU in a single address space. That’s enough to run models that don’t fit on any consumer discrete GPU. And unlike an H100, it costs under two thousand dollars and fits in a backpack.

Karpathy’s autoresearch was designed for H100s. The AMD community ported it to MI308X datacenter GPUs. I ported it to a laptop. Each step made it more accessible. Each step proved the same point: you don’t need a datacenter to do ML research. You need working software, a capable GPU, and enough stubbornness to debug the parts nobody documented.

The Fork

It’s open source. MIT license, same as Karpathy’s original.

https://github.com/bkpaine1/autoresearch-halo

If you have a Strix Halo and you’ve been told it can’t do ML work — it can. If you’ve been told ROCm doesn’t work on consumer hardware — it does, if you use TheROCk nightly builds instead of the official releases that pretend your GPU doesn’t exist. If you’ve been told torch.compile doesn’t work on AMD — it does, on PyTorch 2.11 nightlies with one dtype cast.

The AI agent is running experiments on my laptop right now. I’m going to go look at the house I’m buying on Wednesday. When I get back, the results will be waiting.

That’s the point. The research doesn’t stop because you have a life.

Thank you to Andrej Karpathy for building autoresearch and making it open. Thank you to andyluo7 for the first AMD port. Thank you to the TheROCk team for keeping ROCm alive on hardware that the official release notes forgot.

The future of AI research isn’t locked in a datacenter. It’s in your backpack. It just needs someone stubborn enough to make it work.

More from the studio

Explore all 13 articles →See the current work →