platoseed
Making AI run fast on any hardware.
Luminal builds an ML framework and compiler that generates GPU code. Our stack 10x's model speed while simplifying deployment and cutting idle GPU costs Github: https://github.com/luminal-ai/luminal Discord: https://discord.gg/APjuwHAbGy
Luminal is an AI inference compiler that pre-compiles models into optimized native code for GPUs and ASICs to deliver high-throughput, low-latency inference. It emphasizes avoiding runtime overhead and enables dynamic, scalable deployment across heterogeneous hardware.
Luminal compiles models into a minimal dataflow graph, applies hardware-aware optimizations (fusion, tiling, memory planning, scheduling), and emits native code directly to GPU kernels or ASIC instructions with zero runtime overhead. It supports PyTorch and Hugging Face models, offers a compiler-first inference path, and provides an Inference OS that dynamically schedules workloads across CPUs/GPUs/ASICs with load balancing and scalable deployment options (Cloud serverless or On-Prem).
Who itβs for: AI teams and organizations needing fast, scalable AI inference across heterogeneous hardware (GPUs, ASICs, CPUs), including deployment in cloud or on-prem environments.
Hiring/traction mentioned via product growth features and deployment options; active roadmap and enterprise/SLA mentions; official collaboration language suggests commercialization and enterprise focus.
Generating GPU kernels automatically to speed up ML models. Ex-Intel, worked on CPU microcode and ML accelerators.
Co-founder at Luminal AI. Ex-Amazon engineer, with globally deployed projects automatically finding issues in the Amazon fulfillment network and cost effectively fixing them
Cofounder at Luminal: generating GPU kernels automatically to speed up ML models. Ex-Apple. Talk to me about donuts or compilers or both :)
Luminal automatically speeds up and deploys your PyTorch models
Luminal is an open-source ML compiler that automatically generates optimized CUDA kernels from PyTorch models to speed up production inference. It enables drop-in, serverless deployments with no idle costs, aiming to replace costly GPU-engineering work and accelerate model serving across research and production environments.

We automatically monetize idle GPUs

The Fastest Multimodal Inference OS