Skip to main content
Ryan Orban

Ryan Orban

Subject
5 entries

Compute

Bookmarks

  1. CubeCL: Multi-Platform GPU Compute for Rust

    CubeCL is a multi-platform GPU compute language extension for Rust — write one kernel that targets CUDA, ROCm, Vulkan/WebGPU, and Metal. The foundation for the Burn deep learning framework's backend abstraction.

  2. Compute Watch: LLM Compute Costs and GPU Availability Tracker

    Compute Watch is a tracker for LLM compute costs and GPU availability — benchmarking inference costs across providers and tracking H100/A100 spot availability. Useful for anyone making infrastructure decisions around model serving costs.

  3. Nvidia H100 GPUs: Supply and Demand

    A detailed analysis of the Nvidia H100 GPU supply and demand situation in mid-2023 — how constrained supply was, where the demand was coming from, and what the bottlenecks were. Essential context for understanding the AI infrastructure market that year.

  4. Compute-Optimal LLMs: Chinchilla Scaling Calculator

    howmanyparams.com is a calculator for compute-optimal LLM training based on Chinchilla scaling laws — given a compute budget, it tells you the optimal model size and token count. A practical tool for applying the Hoffmann et al. scaling law findings.

  5. Dheeraj Pandey on Making Computing and Storage Converge

    TechCrunch 'In the Studio' interview with Nutanix CEO Dheeraj Pandey explaining the hyper-converged infrastructure vision — collapsing the traditional three-tier data center into a single software-defined appliance. Ryan noted it was a fantastic video at the time.

All bookmarks