How Arm and Google are improving ML Drift performance on Arm GPUs
Learn how Arm and Google optimize ML Drift for Arm GPUs, improving on-device AI inference performance through kernel, resource, and workgroup tuning
By Bala Gattu

Google AI Edge is open-sourcing ML Drift under the Apache 2.0 license.
ML Drift is Google's high-performance, cross-platform GPU compute framework for on-device AI/ML inference, serving as the core GPU acceleration engine behind the LiteRT ML Drift accelerator and as a standalone C++ library for custom runtimes.
Through architectural advances like GPU-optimized tensor layouts, device-aware workgroup selection, efficient data movement, and runtime-generated kernels, ML Drift accelerates both classical machine learning and generative AI workloads across mobile, web, and desktop platforms.
On-device AI depends on heterogeneous compute, with CPUs, GPUs, and other accelerators each playing an important role. For ML Drift, that means making the GPU execution path efficient across the hardware developers deploy on.
Arm collaborated closely with Google to enhance ML Drift capabilities across Arm GPUs, leveraging deep hardware architecture insight to unlock targeted optimizations beneath the framework layer.

Improving ML Drift on Arm GPUs
Getting good AI performance from a GPU is not just about the hardware. How kernels are implemented, how GPU resources are configured, and how workloads are tuned all make a difference.
Arm worked closely with Google to co-design optimizations for Mali GPUs in ML Drift, including:
- Co-optimized compute kernels engineered for high-throughput, power-efficient execution,
- Texture cache locality and GPU resource optimization to maximize memory bandwidth efficiency, and
- Device-aware workgroup tuning across Arm GPU architecture generations.
Together, these contributions are designed to improve how ML Drift executes inference workloads on Arm GPUs.
“Delivering great on-device AI performance requires close collaboration across the hardware and software stack. Through optimized kernels, device-aware GPU resource handling, and execution tuning, Arm's contributions to ML Drift have helped improve GPU inference performance on Arm-based mobile platforms, providing developers a strong foundation to build on for high-performance AI.”
- Matthias Grundmann, Google AI Edge Lead
This work matters because most developers should not spend their time tuning individual GPU kernels or working around hardware-specific behavior.
Arm can do that work further down the stack, working with partners like Google to optimize the frameworks and runtimes developers are already using.
What does this mean for developers?
ML Drift provides GPU-accelerated inference within Google’s on-device AI software stack.
For developers, the important question is how improvements at this layer translate into the software they build.
If an application uses LiteRT with GPU acceleration enabled (or migrates from the legacy TFLite GPU delegate to the LiteRT ML Drift accelerator), ML Drift provides the underlying GPU execution path. Improvements in ML Drift can therefore benefit applications using that path without requiring developers to reproduce the same hardware-specific tuning.
That is the direction we want to keep pushing: make the software developers already use work better on Arm.
As more AI moves onto devices, developers need practical ways to use the CPU, GPU, and other compute available to them without taking on unnecessary platform-specific work.
Arm’s collaboration with Google on ML Drift is one example of how hardware and software teams can move that complexity further down the stack.
Open source makes the optimization layer more accessible
Open-sourcing ML Drift also means developers can look below the higher-level APIs.
You can inspect how GPU inference is implemented, see how workloads are tuned, experiment with the code, and contribute back to the project.
For developers working on AI performance, GPU compute, or the lower layers of the ML stack, that makes ML Drift useful beyond the applications that depend on it. It is also a chance to see some of the optimization work happening underneath the frameworks developers use every day.
Try it yourself
With ML Drift becoming open source, developers can start exploring the project directly.
Explore ML Drift
Google ML Drift RepositoryGet started with ML Drift quickstarts
Want to go deeper on profiling and optimizing workloads for Arm GPUs?
Arm will continue working with Google and the broader software ecosystem to bring Arm optimization into the frameworks, runtimes, and tools developers already use, so they can spend more time building AI experiences and less time managing platform-specific complexity.
The goal is simple: make it easier for developers to build, optimize, and run AI workloads on Arm.
Arm team acknowledgements: Varun Chari, Albin Bernhardsson, Dan Wilson, Gene Chorba
Google team acknowledgements: Raman Sarokin, Juhyun Lee, Chintan Parikh, Meghna Johar
By Bala Gattu
Re-use is only permitted for informational and non-commercial or personal use only.
