[{"data":1,"prerenderedAt":28},["ShallowReactive",2],{"nr-en-nvidia-cuda-rust-gpu-kernels":3},{"slug":4,"title":5,"dek":6,"date":7,"time":8,"publishedAt":9,"updated":10,"updatedAt":10,"dateFmt":11,"updatedFmt":10,"kind":12,"tier":13,"author":14,"authorName":15,"topics":16,"tracker":10,"trackerLabel":10,"headlineStat":22,"image":23,"ogImage":24,"imageAlt":5,"csv":10,"minutes":25,"words":26,"html":27},"nvidia-cuda-rust-gpu-kernels","Nvidia Introduces CUDA Rust – GPU Kernels Now Native in Rust","Nvidia opens GPU kernel development to Rust programmers. Two new frameworks – cuda-oxide and cutile-rs – let developers write and compile GPU kernels directly in Rust to PTX, eliminating the need for C++ or Python wrappers.","2026-09-09","07:36","2026-09-09T07:36:00+02:00","","September 9, 2026","news","standard","ideal-syka","Ideal Syka",[17,18,19,20,21],"GPU programming","Rust","CUDA","AI infrastructure","Nvidia","Two new Rust frameworks for GPU kernels","\u002Fnewsroom\u002Fimg\u002Fnvidia-cuda-rust-gpu-kernels.webp","\u002Fog-nr\u002Fnvidia-cuda-rust-gpu-kernels.en.png",3,512,"\u003Cp>In September 2026, Nvidia unveiled two new Rust frameworks for GPU programming: \u003Cstrong>cuda-oxide\u003C\u002Fstrong> and \u003Cstrong>cutile-rs\u003C\u002Fstrong>. This closes a long-standing gap – while more of AI infrastructure is written in Rust (drivers, inference engines, agent runtimes), GPU kernels have had to be developed in other languages. Now developers can write kernels directly in Rust and compile them natively to PTX (Parallel Thread Execution).\u003C\u002Fp>\n\u003Ch2>The essentials\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Cstrong>Two programming models\u003C\u002Fstrong>: cuda-oxide follows the \u003Cstrong>SIMT model\u003C\u002Fstrong> (Single Instruction, Multiple Threads); cutile-rs uses the newer \u003Cstrong>Tile model\u003C\u002Fstrong>, which delegates architecture decisions to the compiler\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Different maturity levels\u003C\u002Fstrong>: cutile-rs runs on \u003Cstrong>stable Rust 1.89+\u003C\u002Fstrong> and is already published on crates.io; cuda-oxide requires a \u003Cstrong>pinned Nightly toolchain\u003C\u002Fstrong> and remains in early alpha\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Early adopters\u003C\u002Fstrong>: cutile-rs is already used in \u003Cstrong>HuggingFace&#39;s Grout inference engine\u003C\u002Fstrong> and \u003Cstrong>mistral.rs\u003C\u002Fstrong>\u003C\u002Fli>\n\u003Cli>\u003Cstrong>Memory safety\u003C\u002Fstrong>: Both frameworks enforce memory safety at compile time – cuda-oxide via DisjointSlice and launch contracts, cutile-rs via tensor partitioning and ownership\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Ch2>Two paths, one language\u003C\u002Fh2>\n\u003Cp>The two frameworks serve different developer needs. \u003Cstrong>cuda-oxide\u003C\u002Fstrong> targets those who need full control over thread mapping and memory layout – the classic SIMT model familiar from CUDA C++ or Numba. You specify what a single thread does, then launch thousands of them.\u003C\u002Fp>\n\u003Cp>\u003Cstrong>cutile-rs\u003C\u002Fstrong> takes a more abstract approach: you define what a data tile should do, and the compiler decides how to map those tiles onto the target GPU architecture. This makes code more portable and less prone to architecture-specific optimization mistakes. Nvidia recommends starting with Tile first, dropping to SIMT only when you need that control.\u003C\u002Fp>\n\u003Cdiv class=\"tbl-scroll\">\u003Ctable>\n\u003Cthead>\n\u003Ctr>\n\u003Cth>Criterion\u003C\u002Fth>\n\u003Cth>cuda-oxide\u003C\u002Fth>\n\u003Cth>cutile-rs\u003C\u002Fth>\n\u003C\u002Ftr>\n\u003C\u002Fthead>\n\u003Ctbody>\u003Ctr>\n\u003Ctd>Model\u003C\u002Ftd>\n\u003Ctd>SIMT (classic)\u003C\u002Ftd>\n\u003Ctd>Tile (modern)\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Rust version\u003C\u002Ftd>\n\u003Ctd>Nightly (pinned)\u003C\u002Ftd>\n\u003Ctd>Stable 1.89+\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Dependencies\u003C\u002Ftd>\n\u003Ctd>LLVM required\u003C\u002Ftd>\n\u003Ctd>No custom LLVM\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Status\u003C\u002Ftd>\n\u003Ctd>Early alpha\u003C\u002Ftd>\n\u003Ctd>Production (crates.io)\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003Ctr>\n\u003Ctd>Known users\u003C\u002Ftd>\n\u003Ctd>–\u003C\u002Ftd>\n\u003Ctd>HuggingFace, mistral.rs\u003C\u002Ftd>\n\u003C\u002Ftr>\n\u003C\u002Ftbody>\u003C\u002Ftable>\u003C\u002Fdiv>\n\u003Ch2>Memory safety without performance trade-offs\u003C\u002Fh2>\n\u003Cp>Both frameworks enforce memory safety at compile time – a major Rust advantage. \u003Cstrong>cuda-oxide\u003C\u002Fstrong> uses DisjointSlice and launch contracts to prevent aliasing issues. \u003Cstrong>cutile-rs\u003C\u002Fstrong> relies on tensor partitioning and ownership rules to guarantee exclusive access. This means entire classes of bugs (race conditions, buffer overflows) never occur in the first place, rather than being debugged at runtime.\u003C\u002Fp>\n\u003Ch2>Interoperability planned\u003C\u002Fh2>\n\u003Cp>Nvidia plans for CUDA Rust, CUDA C++, and CUDA Python to work seamlessly together in the future. This means: choosing your frontend language won&#39;t lock you out of other ecosystems. You can use Rust kernels alongside C++ kernels in the same project.\u003C\u002Fp>\n\u003Ch2>What this means for enterprises\u003C\u002Fh2>\n\u003Cp>For teams building AI infrastructure, inference engines, or specialized GPU software, the barrier to entry is lowering. Instead of onboarding Rust developers to CUDA C++, they can stay in their language. This could be particularly valuable for startups and mid-market companies building Rust-first stacks. However: cuda-oxide is still alpha, cutile-rs is production-ready – if you&#39;re starting today, begin with Tile and cutile-rs, not SIMT.\u003C\u002Fp>\n\u003Ch2>Sources\u003C\u002Fh2>\n\u003Cul>\n\u003Cli>\u003Ca href=\"https:\u002F\u002Fdeveloper.nvidia.com\u002Fblog\u002Fintroducing-cuda-rust-two-tracks-for-writing-gpu-kernels\u002F\">developer.nvidia.com\u003C\u002Fa>\u003C\u002Fli>\n\u003C\u002Ful>\n\u003Cp>\u003Cem>Editorially owned by \u003Ca href=\"\u002Fen\u002Fautor\u002Fideal-syka\">Ideal Syka\u003C\u002Fa>. Sources and method: \u003Ca href=\"\u002Fen\u002Fredaktion\">Newsroom &amp; method\u003C\u002Fa>. Tips and corrections: \u003Ca href=\"mailto:ai@i6eal.de\">ai@i6eal.de\u003C\u002Fa>.\u003C\u002Fem>\u003C\u002Fp>\n",1788941028154]