- Published on
The Origin of Compusential: From a Bricked RIVA 128 to CUDA
- Topics & Classification
Format
Technical Journal Entry
The name Compusential came to me during a morning walk to work, when I was a junior web developer at NVIDIA in Theale, Berkshire.
But the story behind it had begun almost a decade earlier—with a graphics card, a failed firmware update and a lesson in German customer service.
1997: The MediaMarkt Incident
In 1997, I was a teenager at an international school in Beijing. During a trip to Germany, I found myself in the graphics-card aisle of a MediaMarkt, ready to spend my carefully saved money on an upgrade for my Compaq PC.
I had been following the hardware section of PC Games—the precursor to what later became PC Games Hardware—and thought I knew what I wanted.
The obvious darling of 3D gaming was the 3dfx Voodoo Graphics. I chose the dark horse instead: NVIDIA’s RIVA 128, or NV3, an ambitious all-in-one 2D and 3D accelerator equipped with 4MB of SGRAM.
Back in Beijing, I decided that the card needed a firmware upgrade.
One botched ROM flash later, the screen went black.
On our next trip to Germany, I carried the dead graphics card roughly 7,000 kilometres in my luggage, hoping MediaMarkt would replace it. At the customer-service counter, I made the mistake of honestly explaining that I had flashed the BIOS myself.
The clerk responded with classic German directness:
“Leider ist das Ihre Schuld, und deshalb können wir Ihnen nicht weiterhelfen.”
“Unfortunately, that is your own fault, and therefore we cannot help you.”
I left without a replacement, but the experience did not discourage me from taking computers apart. If anything, it had the opposite effect.
A Decade of Tinkering
Through my teenage years in Beijing, I worked my way through the RIVA TNT2, GeForce 256 and GeForce4 Ti generations. I also destroyed a few Socket A AMD Athlons along the way.
Those processors had exposed silicon dies and very little tolerance for cooling mistakes. A badly seated heatsink could turn an expensive processor into a small piece of electronic archaeology remarkably quickly.
The experimentation was sometimes costly, but it taught me how computers behaved beyond the specifications printed on the box: how cooling, power delivery, memory and physical construction could matter as much as the processor itself.
I completed my Computer Science degree at the University of Reading in 2006, including a C++ project on face recognition. My familiarity with graphics hardware—and my ability to talk enthusiastically about GPU specifications—helped me secure a role at NVIDIA.
The company whose graphics card I had destroyed as a teenager had become my first employer after university.
2006: The G80 and CUDA Moment
In late 2006, we were preparing to update the website with the technical details of the new GeForce 8800 GTX, based on NVIDIA’s G80 architecture.
G80 was NVIDIA’s first unified DirectX 10 GPU. Earlier programmable graphics processors used separate vertex and pixel pipelines. General-purpose GPU computing was possible, but developers often had to disguise data as textures and express calculations through graphics-oriented APIs and shader programs.
G80 replaced those separate pipelines with a unified, massively parallel shader architecture. At the same time, CUDA was beginning to take shape as a more approachable way to use that hardware for general-purpose computing. Its initial public beta followed in February 2007, with CUDA 1.0 arriving that June.
Instead of forcing every problem through the language of computer graphics, CUDA allowed developers to write C-like kernels and distribute work across thousands of lightweight GPU threads.
The early software was not especially intuitive—as is often the case with the first generation of a new platform—but the underlying idea was compelling.
I bought an 8800 GTX directly from add-in-board partner PNY. Whatever discount I received was quickly absorbed by the cost of a larger power supply and the dual six-pin PCIe power connections required to run it.
Internal engineering demonstrations made the potential of GPU computing tangible. Workloads that had traditionally been expressed as long sequences of CPU instructions could be reorganised into thousands of calculations running in parallel.
It felt like more than another generation of faster gaming hardware. The graphics processor was becoming a programmable parallel computer.
Finding the Name
It was against that background that Compusential came to me on my walk to the NVIDIA office.
I wanted a place to write about the computational essence of technology: not only what a new piece of hardware promised, but what was happening underneath it. Memory hierarchies, power limits, parallelism, silicon constraints and the mathematics that made the system work.
The name also captured the kind of practical curiosity that had followed me since the failed RIVA 128 flash. I wanted to understand technology by taking it apart, experimenting with it and occasionally breaking it.
That same thread would later run through my early GPU mining experiments, local compute systems with 128GB of unified memory, and the production software built on top of them.
The hardware has changed dramatically since 1997. The instinct behind it has not.
Compusential is about looking beyond the product announcement or the prevailing hype, getting close to the machinery and working out what it can actually do.