At our latest YC Paper Club, researchers and builders presented on multi-GPU kernel optimization, intelligence per watt for local inference, AI-generated GPU kernels and benchmarking, heterogeneous inference infrastructure design, and GPU-accelerated game engines for reinforcement learning. Thanks to the following presenters:
Stuart Sol (Stanford / Cursor), John (Stanford), Mark (PyTorch / GPU Mode / CoreAuto), Misha (Marlo), and Brennan (Stanford)
Chapters:
0:00 – Francois Chaubard: The case for chip and kernel specialization
7:16 – Stuart Sul: Parallel Kittens – Systematic and Practical Simplification of Multi-GPU Al Kernels (https://arxiv.org/abs/2511.13940)
21:29 – Jon Saad-Falcon: Intelligence per Watt – Measuring the Intelligence Efficiency of Local and Cloud AI (https://arxiv.org/abs/2511.07885)
31:05 – Mark Saroufim: When Al Starts Writing Systems Code
47:04 – Misha Smelyanskiy: Why AI Inference Needs Heterogeneous Hardware
1:04:33 – Brennan Shacklett: Building a High-Throughput Game Engine that Runs ENTIRELY on the GPU (https://madrona-engine.github.io/shacklett_siggraph23.pdf)
1:15:35 – Wrap-up & what’s next
Apply to Y Combinator: https://www.ycombinator.com/apply
Work at a startup: https://www.ycombinator.com/jobs


