Overview
ML Drift is the on-device GPU compute engine the Google AI Edge team released under Apache 2.0 on October 8, 2026, built specifically for AI inference on the device itself. Its formal place is as the GPU acceleration layer inside LiteRT, taking over from the earlier TFLite GPU delegate, but it also stands alone as a library that custom graphics or inference runtimes can build on.
To see why it exists, look at what makes on-device GPU inference hard. Datacenter inference runs models on homogeneous, predictable accelerator clusters, while the edge is defined by hardware diversity: GPU architectures, driver versions and low-level APIs all vary, and developers cannot know in advance which specific chip an application will land on. The TFLite GPU delegate laid the foundation for GPU acceleration, but it hardcodes 4D tensors and maintains a separate shader codebase per backend, which leaves compute and memory as bottlenecks once workloads stretch from real-time vision, audio and depth processing to high-parameter generative models. ML Drift is meant to be the universal engineering foundation that gives classical models portability and next-generation models peak performance in one framework.
Key Features
- Tensor virtualization and unified shaders: Tensor virtualization decouples a tensor logical representation from its physical allocation on the GPU, with dynamic shader templates resolving and translating coordinates during compiler initialization. The need to maintain separate backend-specific shader codebases disappears, runtime overhead stays minimal, and cross-platform model portability is preserved
- One abstraction over four low-level APIs: OpenGL ES, OpenCL, Metal and WebGPU sit behind a single abstraction, so developers do not adapt separately to each GPU architecture and driver version, and do not need to know ahead of time which hardware the application will run on
- 5D tensors out of the box: The legacy delegate was structurally hardcoded to 4D, forcing layout hacks for models needing 5D. ML Drift supports 5D natively in the new LiteRT GPU accelerator, so volumetric and spatial AI such as 3D convolutional networks, along with spatiotemporal models like YOLO 11n, MobileViT v2 and Swin Transformer v2, run directly on edge GPUs
- Stage-aware optimization for edge LLMs: Autoregressive LLM inference splits into a compute-bound KV cache prefill stage and a memory-bandwidth-bound decode stage, and ML Drift switches kernels and layout configurations based on the active phase. During decode it uses a convolution-aligned KV cache layout with aggressive in-kernel activation quantization to bypass redundant memory roundtrips
- Extensible custom op framework: Direct registration APIs give low-level shading language access, and an agentic SKILL.md guide lets coding agents author, register and verify performant custom shaders in minutes, sharply lowering the barrier to integrating proprietary model blocks into the execution graph
- Desktop previews and lower memory overhead: The WebGPU backend was originally built for in-browser acceleration, and Chromium Dawn lets that exact codebase compile natively outside the browser, bypassing DirectX and Vulkan fragmentation on Windows and Linux while macOS keeps a native Metal backend. Gemma benchmarks show up to 12 percent lower memory overhead than other frameworks
Use Cases
- [object Object]
- [object Object]
- [object Object]
- [object Object]
- [object Object]
Pros
- Apache 2.0 open source, so commercial use, modification and self-hosting are all permitted with no quota or vendor lock-in
- One codebase covers OpenGL ES, OpenCL, Metal and WebGPU, cutting cross-backend maintenance sharply
- Native 5D tensor support frees volumetric and spatiotemporal models from layout hacks
- Prefill and decode are optimized separately for on-device LLMs, with a decode-stage KV cache layout and in-kernel quantization aimed straight at the memory bandwidth bottleneck
- Already battle-tested in production across Chrome, YouTube Shorts, Google Photos and Google Meet on millions of devices daily rather than only on paper
- The bundled agentic SKILL.md guide lets custom shader authoring and verification be handed to coding agents
- Up to 12 percent lower memory overhead than other frameworks, freeing memory for other local tasks running concurrently
Pricing
Released under the Apache 2.0 license, free for commercial use, modification and self-hosting. Documentation is maintained alongside LiteRT, and when shipped as the LiteRT GPU acceleration layer it comes under the same open source license at no separate charge.
Summary
The value of ML Drift is not a new algorithm but a rebuilt engineering foundation for on-device GPU inference. Tensor virtualization attacks the cost of maintaining cross-backend shaders, 5D tensor support removes a structural ceiling on model shapes, and stage-aware optimization targets the memory bandwidth bottleneck that actually limits on-device LLMs. It already runs in a set of Google products with daily reach in the millions, and numbers like the 40 percent frame latency reduction on YouTube Shorts and the 2 second speedup in Google Photos were measured in production rather than on a benchmark table.
For developers the practical meaning is this: if you build on-device vision, on-device generative models or local coding agents, there is now a unified, open and modifiable GPU foundation available, with no need to tune a separate shader set per backend. Used as a standalone library it can also be embedded into custom graphics or inference runtimes. One caveat is that desktop support is still in preview, and the team primary focus remains mobile and edge devices where resource constraints are tightest.
Version History
- ML Drift open-sourced under Apache 2.0 (2026-10-08): The Google AI Edge team released this cross-platform on-device GPU compute engine as the core GPU acceleration layer inside LiteRT, succeeding the TFLite GPU delegate while remaining available as a standalone library. It unifies OpenGL ES, OpenCL, Metal and WebGPU behind one API, introduces tensor virtualization for a unified shader model, adds 5D tensor support and an extensible custom op framework, and optimizes the prefill and decode stages of on-device LLMs separately. It already runs in production across Chrome, YouTube Shorts, Google Photos, Google Meet and AI Edge Gallery, with Adobe Lightroom and Photoshop also adopting it, and desktop support arrives as a preview by compiling the same codebase through Dawn