ChengRang

ML Drift

AI Platforms Open Source

The Google AI Edge team open-sourced this cross-platform on-device GPU compute engine under Apache 2.0 on October 8, 2026. Built for on-device AI inference, it serves as the core GPU acceleration layer inside LiteRT while also standing alone as a library. It abstracts OpenGL ES, OpenCL, Metal and WebGPU behind one API, uses tensor virtualization to decouple logical tensors from physical GPU allocation so a single shader codebase covers every backend, adds 5D tensor support and an extensible custom op framework, and switches kernels and layouts between the prefill and decode phases of on-device LLMs. It already runs daily across millions of devices in Chrome, YouTube Shorts, Google Photos and Meet

On-device InferenceGPU AccelerationApache 2.0LiteRTCross-platform
Visit ML Drift

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

ML Drift is the on-device GPU compute engine the Google AI Edge team released under Apache 2.0 on October 8, 2026, built specifically for AI inference on the device itself. Its formal place is as the GPU acceleration layer inside LiteRT, taking over from the earlier TFLite GPU delegate, but it also stands alone as a library that custom graphics or inference runtimes can build on.

To see why it exists, look at what makes on-device GPU inference hard. Datacenter inference runs models on homogeneous, predictable accelerator clusters, while the edge is defined by hardware diversity: GPU architectures, driver versions and low-level APIs all vary, and developers cannot know in advance which specific chip an application will land on. The TFLite GPU delegate laid the foundation for GPU acceleration, but it hardcodes 4D tensors and maintains a separate shader codebase per backend, which leaves compute and memory as bottlenecks once workloads stretch from real-time vision, audio and depth processing to high-parameter generative models. ML Drift is meant to be the universal engineering foundation that gives classical models portability and next-generation models peak performance in one framework.

Key Features

Use Cases

Pros

Pricing

Released under the Apache 2.0 license, free for commercial use, modification and self-hosting. Documentation is maintained alongside LiteRT, and when shipped as the LiteRT GPU acceleration layer it comes under the same open source license at no separate charge.

Summary

The value of ML Drift is not a new algorithm but a rebuilt engineering foundation for on-device GPU inference. Tensor virtualization attacks the cost of maintaining cross-backend shaders, 5D tensor support removes a structural ceiling on model shapes, and stage-aware optimization targets the memory bandwidth bottleneck that actually limits on-device LLMs. It already runs in a set of Google products with daily reach in the millions, and numbers like the 40 percent frame latency reduction on YouTube Shorts and the 2 second speedup in Google Photos were measured in production rather than on a benchmark table.

For developers the practical meaning is this: if you build on-device vision, on-device generative models or local coding agents, there is now a unified, open and modifiable GPU foundation available, with no need to tune a separate shader set per backend. Used as a standalone library it can also be embedded into custom graphics or inference runtimes. One caveat is that desktop support is still in preview, and the team primary focus remains mobile and edge devices where resource constraints are tightest.

Version History

Category
AI Platforms
Pricing
Open Source
Tags
On-device Inference · GPU Acceleration · Apache 2.0

Related Tools