
CUDA and AI
The strangest winning bet in chip history: a platform for a market that didn't exist — until it did.
The insight
By the mid-2000s, NVIDIA’s engineers noticed something: the graphics chips they’d built for games were astonishingly good at a kind of math — massive parallel computation — that had nothing to do with graphics. Scientists were already hacking game GPUs to run simulations, a scrappy movement called GPGPU (general-purpose computing on graphics processing units).
Huang’s decision: instead of letting researchers hack the hardware, NVIDIA would build them a proper platform — a programming model, a compiler, libraries — so anyone could write ordinary code that ran on the GPU’s hundreds of cores. The project was called CUDA (Compute Unified Device Architecture).
The launch (2007)
CUDA launched in 2007 alongside the Tesla line of compute GPUs. The pitch was radical: your graphics card is a supercomputer.
The market’s response was a shrug. CUDA lost money for years. Wall Street analysts questioned why a gaming company was pouring R&D into a platform with no customers. Inside NVIDIA, the project survived because Huang protected it — the same founder authority that had survived the NV1 disaster now shielded a bet with no revenue.
Key figure: Ian Buck, the Stanford researcher whose GPGPU work helped inspire the effort, joined NVIDIA and became the driving force behind CUDA’s software stack.
The bet’s logic
Huang’s reasoning, as he’s explained it since:
- Moore’s Law was slowing. General-purpose CPUs were hitting physical limits; the industry needed a new engine of performance.
- Parallelism was the answer. Many of the hardest problems — simulation, and eventually machine learning — are “embarrassingly parallel”: the same operation on vast data, exactly what GPUs do.
- Software moats beat hardware moats. Chips get copied; a decade of developer tools, libraries, and trained engineers does not. CUDA wasn’t a chip — it was an ecosystem.
2012: the world catches up
In 2012, researchers at the University of Toronto trained AlexNet, a neural network that crushed the ImageNet competition — on two NVIDIA GTX 580 gaming GPUs. The result electrified AI research: deep learning worked, and it needed GPUs.
Suddenly CUDA had a market. Every AI lab in the world standardized on NVIDIA hardware and the CUDA software stack. The money-losing platform became the toll road of the AI era: to train a serious AI model, you bought NVIDIA.
The data-center pivot
NVIDIA reorganized around the opportunity. The data center — selling GPUs by the rack and the supercomputer to cloud providers — went from a side business to the company’s core. Each generation (Kepler, Pascal, Volta, Ampere, Hopper, Blackwell) widened the lead, and each CUDA release deepened the software moat: libraries for every AI framework, every scientific domain.
By the 2020s, the pattern was set: AI researchers learned CUDA in school, wrote CUDA in the lab, and deployed on NVIDIA in production. Competitors could match the silicon; matching the decade of software was the hard part.
The AI factory
Huang’s current framing, repeated at GTC keynotes: data centers are becoming “AI factories” — facilities that take in electricity and produce intelligence, the way 19th-century factories took in coal and produced goods. Dynamo, NVIDIA’s software for orchestrating inference at scale, is pitched as “the operating system of the AI factory.”
Whether the metaphor outlives the hype cycle is an open question. The installed base is not: as of 2026, the overwhelming majority of AI training compute runs on NVIDIA hardware and CUDA.
What CUDA proved
- Build for the market that doesn’t exist yet. CUDA’s wilderness years (2007–2012) are Huang’s favorite case study in conviction.
- The platform is the product. NVIDIA’s moat is as much software as silicon.
- Timing beats brilliance. GPGPU researchers had the idea; NVIDIA had the patience to productize it for five revenue-free years.
Source notes
CUDA history draws on NVIDIA’s corporate records, Ian Buck’s published accounts, contemporaneous coverage of the 2007 launch and the 2012 AlexNet result, and Huang’s GTC keynotes (including the 2025 “AI factory”/Dynamo segment covered by Engadget and others). Characterizations of CUDA’s early unprofitability and the competitive dynamics are reported in business press profiles.