September 29, 2026

Every gigawatt of new AI infrastructure now costs $50 to $60 billion to build, roughly $45 billion of which goes straight into GPUs. With more than 400 gigawatts of new compute needed by 2030, the industry faces a roughly $24 trillion capital bill. But according to this week's guest, the real threat to Western AI competitiveness isn't Nvidia, Google, or Broadcom out-maneuvering one another. It's China Inc., which is targeting a cost of under $10 billion per gigawatt, a five- to six-times cost advantage that could reshape who wins the AI infrastructure race.
In this episode of TechSurge, host David Goldman speaks with Raja Koduri, one of the most influential architects in the history of graphics computing. Raja twice led graphics at AMD, directed graphics architecture at Apple, and served as chief architect of Intel's Core and Visual Computing Group. His team helped bring high bandwidth memory to market for AMD's GPUs back in 2015, a technology that now underpins nearly every AI accelerator on the market.
The conversation centers on Raja's new startup, Oxmiq, which he describes as converting "electrons to tokens super efficiently." Raja and David dig into why the bottleneck in AI infrastructure has quietly shifted away from raw compute and toward memory hierarchy, how data moves between SRAM, HBM, and NAND, and between chips, now that models are too large to fit on a single die. Oxmiq's approach uses 3D-stacked, hybrid-bonded memory to unlock 10x the bandwidth of today's HBM and a 10x increase in token generation rate, even on older process nodes.
From there, the discussion turns to how AI coding agents are changing the economics of chip design. Raja argues that spec-writing, once the "boring" 90 percent of the job, is now the hardest and most valuable skill, while writing the code itself is increasingly something agents can handle. He points to OpenAI and Broadcom's Jalapeño chip as evidence that AI-assisted design can compress timelines that used to take multiple generations of custom silicon. Raja also reflects on lessons from building products alongside Steve Jobs at Apple and Lisa Su at AMD, and explains why he believes Intel's decision to kill 3D XPoint memory came at exactly the wrong moment.
The episode closes with Raja's outlook on where AI infrastructure goes next: a future of both massive "token factories" and smaller, personal "token banks," a shift he compares to the transition offices went through with server rooms, becoming invisible infrastructure that simply works. As he puts it, "the more boring you make it, the more it becomes fabulous."
Sign up for our newsletter at techsurgepodcast.com for updates on upcoming TechSurge Live Summits and future episodes.
Speaker Profiles and Links
David Goldman: Partner, Celesta Capital
Raja Koduri: Founder, Oxmiq; twice Corporate VP/Chief Architect of Graphics at AMD; former VP and Head of Graphics Architecture at Apple; former Chief Architect, Intel Core and Visual Computing Group
LinkedIn: https://www.linkedin.com/in/raja-koduri-3a51611
X: https://x.com/RajaXg
Oxmiq: https://oxmiq.ai
Further Reading and Resources
OpenAI and Broadcom — "OpenAI and Broadcom Unveil LLM-Optimized Inference Chip": https://openai.com/index/openai-broadcom-jalapeno-inference-chip/
High Bandwidth Memory (HBM) — the memory technology Raja's AMD team helped bring to market with HBM1 and HBM2, and the benchmark Oxmiq's 3D-stacked approach aims to beat by 10x: https://en.wikipedia.org/wiki/High_Bandwidth_Memory
Intel 3D XPoint (Optane) — the discontinued memory technology Raja says could have made Intel a major player in the inference era: https://en.wikipedia.org/wiki/3D_XPoint
Timestamps
00:00 - A Gigawatt of AI Now Costs $60 Billion
01:58 - Introducing Raja Koduri
02:02 - What Oxmiq Builds: Electrons In, Tokens Out
05:18 - The $24 Trillion AI Infrastructure Bill
06:05 - China vs. the Rest of the World
09:10 - Memory Is the New Bottleneck
22:03 - How AI Agents Are Changing Chip Design
36:30 - Lessons From Steve Jobs and Lisa Su
44:54 - Advanced Packaging, Memory Prices, and Intel's Mistake
55:38 - Boom or Bust: The Future of Token Factories
.png)
Almost 2% of U.S. GDP will be spent on AI infrastructure this year, nearly double 2025's figure. But beneath those headline numbers, the composition of that spending has quietly flipped: for the first time, dollars spent on running models in production now outweigh dollars spent training them.
In this episode of TechSurge, host David Goldman speaks with Austin Lyons, a semiconductor analyst at Creative Strategies, co-host of the Semi Doped podcast, and author of the Chipstrat newsletter. Lyons previously worked as a hardware engineer at Intel and as a product manager on John Deere's autonomous tractor and Blue River Technology teams before turning to full-time chip industry analysis.
The conversation opens with why AI buyers have moved from assembling commoditized parts to buying entire pre-integrated systems, tracing how Nvidia's rack-scale approach, exemplified by its 72-GPU Grace Blackwell racks, made turnkey deployment the default, and why that raises the bar for any chip startup trying to compete. Lyons and Goldman then unpack how inference workloads have split into two distinct problems, prefill and decode, and how that split created an opening for SRAM-based challengers to outperform general-purpose GPUs on decode speed.
From there, the discussion turns to the rise of neoclouds, the GPU-rental companies that grew into public businesses worth well over $100 billion combined, and why so many traditional investors missed them. Lyons and Goldman work through the circular financing debate head-on: the mechanics of Nvidia's equity stakes, GPU-backed debt, and hyperscaler off-take agreements that critics compare to dot-com-era vendor financing, and the counterargument that demand is simply outrunning fixed supply.
The episode closes on Lyons's own framework for identifying the next trillion-dollar chip company, built on four conditions including the ability to run trillion-parameter models at rack scale, beat an incumbent on a key performance metric, and land a frontier anchor customer, along with a look at how AI-assisted chip design is lowering the barrier for more companies, from OpenAI to electric vehicle makers, to design their own custom silicon.
Sign up for our newsletter at techsurgepodcast.com for updates on upcoming TechSurge Live Summits and future episodes.
.png)
In this episode, Nobel Prize-winning physicist Dr. John Martinis reveals how his breakthrough in superconducting qubits made quantum physics real at macroscopic scale and what it means for the future of technology. The former lead of Google's Quantum AI lab explains why quantum computing is so fragile, why a lot of hype has a low chance to work, and why his fabless company Qolab could be the Nvidia of quantum computing.
In this conversation, Dr. Martinis joins Tech Surge to explain the science behind macroscopic quantum coherence, the engineering challenges of scaling quantum computers, and how hybrid quantum-classical computing will shape the future of technology.
The conversation covers:
✅ How the superconducting qubit breakthrough won the Nobel Prize in Physics
✅ Why Nature wants to destroy quantum coherence and why quantum is fragile
✅ From academic physics to building Google's quantum computer
✅ The engineering challenge of scaling quantum computing beyond the lab
✅ Why a lot of quantum computing hype has a low chance to work
✅ How Qolab's fabless model could scale quantum hardware
Sign up for our newsletter at techsurgepodcast.com for updates on upcoming TechSurge Live Summits and future episodes.

Silicon Valley was built on semiconductors, but for nearly two decades, venture capital shifted its attention towards software. Today, AI is changing that as the demand for compute, memory and networking explodes, hardware is once again at the centre of the industry's biggest bets.
In this episode of TechSurge, host Michael Marks speaks with Lip-Bu Tan, CEO of Intel and one of the semiconductor industry's most influential investors and executives. The conversation traces Tan's journey from studying nuclear engineering at MIT to leading Cadence's turnaround, investing in more than 500 technology companies, and now steering Intel through one of the most significant transformations in its history.
Tan shares his VC conviction on backing semiconductor startups when most venture investors favored software, and why he believes AI's next breakthroughs will come from advances in memory, packaging, photonics, cooling and high-speed connectivity. He also opens up on the leadership philosophy that defined his time at Cadence, where listening to customers and building a culture of responsiveness became the foundation of the company's revival.
Wearing his CEO hat, Tan explains Intel's long-term strategy, why vertical integration still matters, how the company plans to reconnect with the startup ecosystem, and why missing another technology wave is not an option.
Sign up for our newsletter at techsurgepodcast.com for updates on upcoming TechSurge Live Summits and future episodes.