Writing

Neanderthal: Coaxing LLM Inference from a 2004 Pentium 4

Original writing, AI-assisted editing.

While scrounging through cabinets and basement corners for viable PC hardware that could help populate my homelab, I found a machine that didn’t make the cut – a 2004 HP Pavilion a450n. Once upon a time this was my daily driver, a Windows XP tower with a 32-bit Intel Pentium 4 3.00 Ghz (Prescott) processor, originally equipped with 512 MB PC3200 DDR SDRAM (Note: Today we’d call it DDR1, but back then we didn’t need to distinguish between DDR1-DDR4.), which at some point I maxed it out at 2.0 GB (2x1GB). It has Integrated Intel Extreme Graphics 2 for video, enough for 1024x768 or 1280x1024 resolutions and DVD playback.

Nothing about this PC is suggestive of AI or LLMs. The CPU offers maybe 6 FP32 GFLOPS, and it pulls a lot of juice to acheive even that much compute with a massive amount of waste heat thrown off. I’ll need a 32-bit operating system that can run on not just 2GB RAM but have as small of a footprint as possible to leave room for even a tiny LLM. Plus the integrated video has fixed-function 2D/3D acceleration – there is no modern compute drivers or tensor math processing. In fact it will steal 64 MB of system RAM because it doesn’t have any of its own!

The full toolchain build log, every antiX gotcha, and the raw benchmark numbers behind this post are public: legacy-hardware-llm-benchmark on GitHub.

The Machine

Host Model Era CPU (Architecture) TDP Cores/Threads RAM Peak FP32 GFLOPS
pavilion HP Pavilion a450n Q1 2004 Intel Pentium 4 (Prescott) 84W 1/2 2 GB ~6

To make this work at all, Windows XP will consume too many resources, and options are limited for a supported 32-bit OS. Most Linux distros have ceased 32-bit support. Fortunately Linux comes in numerous varieties, and a few are still geared for this type archaic hardware. One that I’ve explored previously is antiX Linux, which is based on Debian. It still supports 32-bit systems and even has minimal installations that skip a desktop environment and other conveniences that users typically want. Another OS option that I considered was Alpine Linux since it also supports 32-bit installations and has a slightly smaller footprint by using “musl libc” rather than “glibc”. However, llama.ccp and other ML toolchains require “glibc” so antiX-26 Core Linux aligned better with the needs of the project.

The Constraints

First Things First

Though this machine has both CD-ROM and DVD Writer drives, plus a 3.5-inch floppy drive, I don’t have that kind of media lying around anymore. I was going to be installing a new OS via a USB flach drive. Of course, I needed to enter the BIOS to change the boot order, but it kept freezing on the splash screen no matter how fast I tried to hit F2. After some research, I learned that hardware from this era shared hardware interupts across legacy IDE, SATA, and PCI/USB 2 controllers. The kicker was that I also didn’t have any wired keyboards lying around either, so I had a USB receiver (a.k.a., dongle) plugged in to connect my wireless keyboard. This machine had both an IDE HDD and a SATA HHD, and the combination of the two was clearly not working with the USB keyboard. The IDE drive alone didn’t work. The SATA drive alone didn’t work. Fortunately, antiX Live can run off a USB flash drive.

With both HDDs disconnected, there was nothing to interupt the USB controller so I could boot to antiX Live on the USB flash drive. However, a minimal live Linux OS alone was not sufficient. I needed to add things like llama.cpp and specific LLMs, and I didn’t want to start from scratch each time I booted the machine. I needed a peristent live USB. I provisioned /dev/sda2 with persistent loop mounts, /live/persist-root at 30 GB and /live/persist-home at 60 GB, and a dedicated USB swap partition of 8GB on /dev/sda3 so that any I/O was limited to the USB alone. I’ll concede that I don’t have a lot of Linux administration experience (Software engineering is my forte.), so getting this configuration took a fair bit of trial and error.

Building for Archaic Silicon

Since this is a 32-bit machine, I can’t just run a standard version of llama.cpp. I had to build it on the host using GNU Compiler Collections (GCC) 14.2.0 and CMake 3.31.6 targeting “i686-linux-gnu”. I explicitly enabled SSE3 vector instructions, the most recent supported by the Prescott core, and ran the compilation overnight with the 2 threads I had available. When I checked on it the following day, it had completed successly. Using “vmstat -s” and “dmesg”, I could see that there was no thermal throttling with only 11.6 MB of peak swap utilized and a total disk footprint of 665 MB

To validate the compiled binaries against the Prescott instruction set, I ran “time ./llama-cli –version”. This cold start initialized and exited cleanly in 0.074 seconds thus verifiying that the compiler flags correctly restricted vector instructions to SSE3, and avoiding any unsupported AVX/AVX2 routines. Additionally it verified that there were no loader anomalies associated with the dynamic runtime linkage against Debian’s 32-bit glibc.

Choosing the Models

TinyLlama:1.1B was the obvious first choice for a model to test. It is a well know model that has been trained on over 3 trillion tokens and been shown to be computationally efficient, outperforms other comparably sized models. The 4-bit quantized version should take up only around 0.6 GB. I also wanted a smaller model and a larger model for comparison. I selected Qwen2.5:0.5b and SmolLM2:1.7B at approximately 0.4 GB and 1.0 GB respectively.

I downloaded the GGUF binaries directly from https://huggingface.co/.

curl -L -o qwen2.5-0.5b-instruct-q4_k_m.gguf
https://huggingface.co/Qwen/Qwen2.5-0.5B-Instruct-GGUF/resolve/main/qwen2.5-0.5b-instruct-q4_k_m.gguf

curl -L -o tinyllama-1.1b-chat-v1.0.Q4_K_M.gguf
https://huggingface.co/TheBloke/TinyLlama-1.1B-Chat-v1.0-GGUF/resolve/main/tinyllama-1.1b-chat-v1.0.Q4_K_M.gguf

curl -L -o smollm2-1.7b-instruct-q4_k_m.gguf
https://huggingface.co/bartowski/SmolLM2-1.7B-Instruct-GGUF/resolve/main/SmolLM2-1.7B-Instruct-Q4_K_M.gguf

graph LR
model["Quantized model<br/>15M – 1.1B params"]
build["llama2.c<br/>-m32 -march=prescott"]
cpu["Pentium 4 530<br/>1C/2T · SSE3 · 2 GB DDR"]
out["Tokens<br/>TBD tok/s"]

model --> build --> cpu --> out

classDef box fill:#eef2fb,stroke:#1f4b8f,color:#1a1d21
classDef old fill:#fdf0e1,stroke:#c2410c,color:#1a1d21

class model,build,out box
class cpu old

Results

I ran the same prompt against all three models: “Explain the concept of a mathematical limit in three simple sentences.” Single-user CLI, both logical threads enabled, no other load on the box.

Model Params Quant Prompt Processing Token Generation Wall Clock Status
qwen2.5:0.5b 494M Q4_K_M 1.1 t/s 0.9 t/s 2m 49.8s Success
tinyllama:1.1b 1.1B Q4_K_M 0.5 t/s 0.5 t/s 5m 56.1s Success
smollm2:1.7b 1.71B Q4_K_M 0.3 t/s 0.0 t/s 46m 13.0s Failed — swap bound

Qwen2.5-0.5B was the first successful inference on the Prescott core, and the only one of the three fast enough to feel usable at all. CPU user time (5m 24.5s) ran to nearly twice the wall-clock time, confirming Hyper-Threading was keeping both logical threads saturated, with no thermal throttling.

TinyLlama-1.1B showed the scaling tax you’d expect — roughly half the throughput of the 0.5B model — but it stayed inside the ~1.78 GB RAM budget and finished without touching swap. Still functional for one isolated query, not for anything resembling a conversation.

SmolLM2-1.7B is where the machine broke. The weights technically fit in RAM plus swap, but that swap partition lives on the USB 2.0 boot drive, and once the kernel started paging, I/O wait dominated the run: wall-clock time (46m 13s) came out to more than 3x the CPU time, meaning the CPU spent most of the run idle, waiting on the bus rather than computing. The output was accurate. The speed was not usable for anything.

Was It Worth It?

Yes, with a caveat: this was never about making the a450n useful, it was about finding exactly where a 2004 machine’s limits are and what actually causes them. Raw clock speed was never the bottleneck. Instruction-set support — SSE3 versus the AVX2 every modern inference kernel assumes — decided whether the toolchain would even compile. Memory bandwidth, and specifically a swap partition stuck behind USB 2.0, decided where inference collapsed. The 0.5B model landing just under 1 token/sec is a genuinely usable floor for local, air-gapped, non-real-time inference. Past roughly 1.5B parameters, the bottleneck stops being the CPU entirely and becomes I/O.

If nothing else, this project corrected an assumption I’d been carrying around unexamined: “old hardware can’t run local AI” is really shorthand for “old hardware’s instruction set and memory subsystem can’t.” The CPU was willing the whole time. Everything around it gave out first.